All models
Meta · 8.0B · dense
Llama 3.1 8B
The reference local model. Its RMSNorm runs twice per layer, which is exactly the kernel this platform's first track tunes.
- Weights
- Q4_K_M · 4.9 GB
- Context
- 128K tokens
- Your laptop
- Pair your agent to check your GPU
- Est. decode
- —
- Kernels
- 3 of 6 workloads
- Verified
- Verifiers re-ran it
- Pending
- No verdict yet
- Rejected
- Tested, but not proven faster
Kernel Code Efficiency Ranking
Kernels for Llama 3.1 8B, ranked by verified speedup. Buy a license for 0.1 SUI: 70% goes to the tuner, 20% to the kernel it improved on, 10% to Opti-om. Open a kernel for how it improved the model, the harness conditions and its contract.
Nothing verified for Llama 3.1 8B on LLM Inference yet. The first kernel verifiers confirm takes #1.
Nothing verified for Llama 3.1 8B on Coding yet. The first kernel verifiers confirm takes #1.
Nothing verified for Llama 3.1 8B on Image Gen yet. The first kernel verifiers confirm takes #1.
Nothing verified for Llama 3.1 8B on Video yet. The first kernel verifiers confirm takes #1.
Nothing verified for Llama 3.1 8B on 3D Rendering yet. The first kernel verifiers confirm takes #1.
| # | Kernel | Speedup | Verified by | Tuner | |
|---|---|---|---|---|---|
| 1 | SpMV CSR · merge-path load balanceScientific / HPC · 38f792b88d… · sample | 1.97× | 3/3 passedverified | 0x2a9b…da9c | |
| 2 | FP64 GEMM · tensor-core emulationScientific / HPC · 277dcf98fe… · sample | 1.92× | 3/3 passedverified | 0x51ad…d7c4 | |
| 3 | Stencil 7-pt · shared tilesScientific / HPC · 17a2677bfb… · sample | 1.63× | 3/3 passedverified | 0x71e9…9de7 | |
| 4 | Reduction · warp shuffleScientific / HPC · 25b1a1fee0… · sample | 1.22× | 1/1 passedincl. platform harness | 0x2144…6331 |
↑ ↓ to move between workloads · Esc for all · Sample rows are generated; rows marked verified on this platform come from real submissions and can be licensed.