All models
Alibaba · 8.2B · dense

Qwen3 8B

The 8B Qwen3 with thinking mode. Fits an 8 GB laptop GPU at Q4 and shares the 4096-wide RMSNorm with Llama-class models.

Verified
Verifiers re-ran it
Pending
No verdict yet
Rejected
Tested, but not proven faster

Kernel Code Efficiency Ranking

Kernels for Qwen3 8B, ranked by verified speedup. Buy a license for 0.1 SUI: 70% goes to the tuner, 20% to the kernel it improved on, 10% to Opti-om. Open a kernel for how it improved the model, the harness conditions and its contract.

Nothing verified for Qwen3 8B on LLM Inference yet. The first kernel verifiers confirm takes #1.
Nothing verified for Qwen3 8B on Coding yet. The first kernel verifiers confirm takes #1.
#KernelSpeedupVerified byTuner
1Winograd F(4,3) · tensor-core transformImage Gen · 7cc0a2fa65… · sample2.44×3/3 passedincl. platform harness0x7ea4…bcde
2GroupNorm · fused SiLUImage Gen · 09e258dae5… · sample1.62×4/4 passedincl. platform harness0x2b1d…79a6
3Conv2d · implicit GEMMImage Gen · e978805c08… · sample1.18×4/4 passedverified0x0340…3f1a
—Attention · flash v2 tilesImage Gen · c8da2acdb2… · sample1.30×1/3 passedrejected0xd069…ff66
Nothing verified for Qwen3 8B on Video yet. The first kernel verifiers confirm takes #1.
Nothing verified for Qwen3 8B on 3D Rendering yet. The first kernel verifiers confirm takes #1.
Nothing verified for Qwen3 8B on Scientific / HPC yet. The first kernel verifiers confirm takes #1.

↑ ↓ to move between workloads · Esc for all · Sample rows are generated; rows marked verified on this platform come from real submissions and can be licensed.