Local AI models
Pick the model you run. See whether it fits your GPU, how fast it should go, which workloads have verified kernels, and license the kernel on Sui.
Kimi K3
3 verifiedMoonshot's newest mixture-of-experts. Agentic coding and long-context reasoning; runs locally on a 512 GB unified-memory machine.
- Weights
- Q3_K_M · 452 GB · 256K context
- Est. decode
- —
Kimi K2
2 verifiedThe open-weights K2 mixture-of-experts. Strong at tool use and code; the model most people run on a Mac Studio.
- Weights
- Q3_K_M · 431 GB · 128K context
- Est. decode
- —
Qwen3 8B
2 verifiedThe 8B Qwen3 with thinking mode. Fits an 8 GB laptop GPU at Q4 and shares the 4096-wide RMSNorm with Llama-class models.
- Weights
- Q4_K_M · 5 GB · 128K context
- Est. decode
- —
Llama 3.1 8B
3 verifiedThe reference local model. Its RMSNorm runs twice per layer, which is exactly the kernel this platform's first track tunes.
- Weights
- Q4_K_M · 4.9 GB · 128K context
- Est. decode
- —
DeepSeek R1 14B
2 verifiedThe Qwen-distilled R1 reasoner. Long chains of thought, so decode throughput and p99 latency matter more than prefill.
- Weights
- Q5_K_M · 10.5 GB · 128K context
- Est. decode
- —
Gemma 3 27B
3 verifiedMultimodal Gemma at 27B. Image and video understanding on a single 24 GB card, with a vision tower the diffusion kernels also serve.
- Weights
- Q4_K_M · 16.5 GB · 128K context
- Est. decode
- —
Decode speeds are estimates from your GPU's memory bandwidth; workload numbers on the model pages are sample data until a track measures them. Cover orbs: Orbkit by zzzzshawn (MIT).