All models
Meta · 8.0B · dense

Llama 3.1 8B

The reference local model. Its RMSNorm runs twice per layer, which is exactly the kernel this platform's first track tunes.

Verified
Verifiers re-ran it
Pending
No verdict yet
Rejected
Tested, but not proven faster

Kernel Code Efficiency Ranking

Kernels for Llama 3.1 8B, ranked by verified speedup. Buy a license for 0.1 SUI: 70% goes to the tuner, 20% to the kernel it improved on, 10% to Opti-om. Open a kernel for how it improved the model, the harness conditions and its contract.

Nothing verified for Llama 3.1 8B on LLM Inference yet. The first kernel verifiers confirm takes #1.
Nothing verified for Llama 3.1 8B on Coding yet. The first kernel verifiers confirm takes #1.
Nothing verified for Llama 3.1 8B on Image Gen yet. The first kernel verifiers confirm takes #1.
Nothing verified for Llama 3.1 8B on Video yet. The first kernel verifiers confirm takes #1.
Nothing verified for Llama 3.1 8B on 3D Rendering yet. The first kernel verifiers confirm takes #1.
#KernelSpeedupVerified byTuner
1SpMV CSR · merge-path load balanceScientific / HPC · 38f792b88d… · sample1.97×3/3 passedverified0x2a9b…da9c
2FP64 GEMM · tensor-core emulationScientific / HPC · 277dcf98fe… · sample1.92×3/3 passedverified0x51ad…d7c4
3Stencil 7-pt · shared tilesScientific / HPC · 17a2677bfb… · sample1.63×3/3 passedverified0x71e9…9de7
4Reduction · warp shuffleScientific / HPC · 25b1a1fee0… · sample1.22×1/1 passedincl. platform harness0x2144…6331

↑ ↓ to move between workloads · Esc for all · Sample rows are generated; rows marked verified on this platform come from real submissions and can be licensed.