All models
Alibaba · 8.2B · dense
Qwen3 8B
The 8B Qwen3 with thinking mode. Fits an 8 GB laptop GPU at Q4 and shares the 4096-wide RMSNorm with Llama-class models.
- Weights
- Q4_K_M · 5 GB
- Context
- 128K tokens
- Your laptop
- Pair your agent to check your GPU
- Est. decode
- —
- Kernels
- 2 of 6 workloads
- Verified
- Verifiers re-ran it
- Pending
- No verdict yet
- Rejected
- Tested, but not proven faster
Kernel Code Efficiency Ranking
Kernels for Qwen3 8B, ranked by verified speedup. Buy a license for 0.1 SUI: 70% goes to the tuner, 20% to the kernel it improved on, 10% to Opti-om. Open a kernel for how it improved the model, the harness conditions and its contract.
Nothing verified for Qwen3 8B on LLM Inference yet. The first kernel verifiers confirm takes #1.
Nothing verified for Qwen3 8B on Coding yet. The first kernel verifiers confirm takes #1.
| # | Kernel | Speedup | Verified by | Tuner | |
|---|---|---|---|---|---|
| 1 | Winograd F(4,3) · tensor-core transformImage Gen · 7cc0a2fa65… · sample | 2.44× | 3/3 passedincl. platform harness | 0x7ea4…bcde | |
| 2 | GroupNorm · fused SiLUImage Gen · 09e258dae5… · sample | 1.62× | 4/4 passedincl. platform harness | 0x2b1d…79a6 | |
| 3 | Conv2d · implicit GEMMImage Gen · e978805c08… · sample | 1.18× | 4/4 passedverified | 0x0340…3f1a | |
| — | Attention · flash v2 tilesImage Gen · c8da2acdb2… · sample | 1.30× | 1/3 passedrejected | 0xd069…ff66 |
Nothing verified for Qwen3 8B on Video yet. The first kernel verifiers confirm takes #1.
Nothing verified for Qwen3 8B on 3D Rendering yet. The first kernel verifiers confirm takes #1.
Nothing verified for Qwen3 8B on Scientific / HPC yet. The first kernel verifiers confirm takes #1.
↑ ↓ to move between workloads · Esc for all · Sample rows are generated; rows marked verified on this platform come from real submissions and can be licensed.