BLK · CODING BENCHMARKSfirst-party · per-quant · reproducible
Coding benchmark leaderboard
Reviewed HumanEval+ scores for open-weight local models at real-world quantizations. First-party rows were run by RunLocalAI; community rows render only after review. Every public row links to its raw test-run Gist.
Test runs
30
Unique models
25
Quant variants
1
Benchmark suites
3
METHODOLOGY
How these numbers are produced
- Each model loads into the listed runtime (Ollama / vLLM / llama.cpp) at the listed quantization on the listed hardware.
- EvalPlus is used as the official scorer. For OpenAI-compatible local runtimes, we first generate deterministic JSONL samples, then score them with
python -m evalplus.evaluate. - Greedy sampling (deterministic; one sample per problem). Score = pass@1 on HumanEval+ (164 problems with augmented tests) × 100.
- Raw stdout + stderr captured and posted to a public GitHub Gist before the row enters this table. Click any score to verify.
- Runner source:
scripts/run-humaneval-plus.ts(public repo) — reproducible end-to-end.
| # | Model | Benchmark | Quant | Rig | Score | Trust | Log |
|---|---|---|---|---|---|---|---|
| 1 | Muse Glimmer 30B tested 2026-08-10 | HumanEval | Q4_K_M | llama.cpp-4dee52f-vast5090 rtx-5090 | 97.6 | First-party command published | Gist → |
| 2 | Gemma 4 26B-A4B tested 2026-07-20 | HumanEval+ | Q4_K_M | ollama-0.32.1-vast5090 rtx-5090 | 93.3 | First-party command published | Gist → |
| 3 | Gemma 4 31B Dense tested 2026-07-20 | HumanEval+ | Q4_K_M | ollama-0.32.1-vast5090 rtx-5090 | 93.3 | First-party command published | Gist → |
| 4 | Muse Glimmer 30B tested 2026-08-10 | HumanEval+ | Q4_K_M | llama.cpp-4dee52f-vast5090 rtx-5090 | 91.5 | First-party command published | Gist → |
| 5 | North Mini Code 1.0 tested 2026-07-21 | HumanEval+ | Q4_K_M | ollama-0.32.1-vast5090 rtx-5090 | 89.6 | First-party command published | Gist → |
| 6 | Ornith 1.0 35B tested 2026-07-21 | HumanEval+ | Q4_K_M | ollama-0.32.1-vast5090 rtx-5090 | 88.4 | First-party command published | Gist → |
| 7 | Qwen3.5 35B-A3B tested 2026-07-21 | HumanEval+ | Q4_K_M | ollama-0.32.1-vast5090 rtx-5090 | 87.8 | First-party command published | Gist → |
| 8 | Qwen3 Coder 30B-A3B tested 2026-07-20 | HumanEval+ | Q4_K_M | ollama-0.32.1-vast5090 rtx-5090 | 87.8 | First-party command published | Gist → |
| 9 | Laguna XS 2.1 tested 2026-07-21 | HumanEval+ | Q4_K_M | ollama-0.32.1-vast5090 rtx-5090 | 87.2 | First-party command published | Gist → |
| 10 | Qwen3.5 27B tested 2026-07-20 | HumanEval+ | Q4_K_M | ollama-0.32.1-vast5090 rtx-5090 | 87.2 | First-party command published | Gist → |
| 11 | Qwen3.6 27B tested 2026-07-20 | HumanEval+ | Q4_K_M | ollama-0.32.1-vast5090 rtx-5090 | 86.6 | First-party command published | Gist → |
| 12 | Nemotron 3 Nano Omni 33B tested 2026-07-21 | HumanEval+ | Q4_K_M | ollama-0.32.1-vast5090 rtx-5090 | 85.4 | First-party command published | Gist → |
| 13 | Granite 4.1 30B tested 2026-07-20 | HumanEval+ | Q4_K_M | ollama-0.32.1-vast5090 rtx-5090 | 85.4 | First-party command published | Gist → |
| 14 | GLM-4.7-Flash tested 2026-07-21 | HumanEval+ | Q4_K_M | ollama-0.32.1-vast5090 rtx-5090 | 85.4 | First-party command published | Gist → |
| 15 | Qwen 2.5 Coder 7B Instruct tested 2026-05-28 | HumanEval+ | Q4_K_M | ollama-0.24 rtx-3080-16gb-mobile | 81.1 | First-party command published | Gist → |
| 16 | Qwen3.6 35B-A3B tested 2026-07-21 | HumanEval+ | Q4_K_M | ollama-0.32.1-vast5090 rtx-5090 | 81.1 | First-party command published | Gist → |
| 17 | Phi-4 14B tested 2026-05-28 | HumanEval+ | Q4_K_M | ollama-0.24 rtx-3080-16gb-mobile | 78.7 | First-party command published | Gist → |
| 18 | Mellum2 12B-A2.5B tested 2026-07-20 | HumanEval+ | Q4_K_M | ollama-0.32.1-vast5090 rtx-5090 | 76.8 | First-party command published | Gist → |
| 19 | Ministral 3 14B tested 2026-07-20 | HumanEval+ | Q4_K_M | ollama-0.32 rtx-3080-16gb-mobile | 76.8 | First-party command published | Gist → |
| 20 | Ornith 1.0 9B tested 2026-07-18 | HumanEval+ | Q4_K_M | ollama-0.32 rtx-3080-16gb-mobile | 73.2 | First-party command published | Gist → |
| 21 | Trendyol LLM Asure 12B tested 2026-05-27 | MBPP+ | Q4_K_M | ollama-0.24.0 rtx-5080 | 71.7 | First-party command published | Gist → |
| 22 | Trendyol LLM Asure 12B tested 2026-05-27 | HumanEval+ | Q4_K_M | ollama-0.24.0 rtx-5080 | 69.5 | First-party command published | Gist → |
| 23 | Qwen 2.5 Coder 7B Instruct tested 2026-05-29 | MBPP+ | Q4_K_M | ollama-0.24 rtx-3080-16gb-mobile | 66.9 | First-party command published | Gist → |
| 24 | Phi-4 14B tested 2026-05-29 | MBPP+ | Q4_K_M | ollama-0.24 rtx-3080-16gb-mobile | 60.3 | First-party command published | Gist → |
| 25 | Dolphin 3.0 8B tested 2026-07-20 | HumanEval+ | Q4_K_M | ollama-0.32.1-vast5090 rtx-5090 | 56.7 | First-party command published | Gist → |
| 26 | Llama 3.1 8B Instruct tested 2026-05-28 | HumanEval+ | Q4_K_M | ollama-0.24 rtx-3080-16gb-mobile | 56.1 | First-party command published | Gist → |
| 27 | Qwen3.5 9B tested 2026-07-19 | HumanEval+ | Q4_K_M | ollama-0.32 rtx-3080-16gb-mobile | 42.7 | First-party command published | Gist → |
| 28 | Hermes 3 Llama 3.1 8B tested 2026-07-20 | HumanEval+ | Q4_K_M | ollama-0.32.1-vast5090 rtx-5090 | 41.5 | First-party command published | Gist → |
| 29 | Llama 3.1 8B Instruct tested 2026-05-29 | MBPP+ | Q4_K_M | ollama-0.24 rtx-3080-16gb-mobile | 39.2 | First-party command published | Gist → |
| 30 | Qwen 3 8B tested 2026-05-29 | HumanEval+ | Q4_K_M | ollama-0.24 rtx-3080-16gb-mobile | 2.4 | First-party command published | Gist → |
A note on quant comparisons. The same model at different quantization levels can score materially differently — especially on code, where precision loss hurts. When you see Q4_K_M next to Q6_K for the same model, the gap is the “quant tax” on coding specifically. Take that into account when picking what to run on /will-it-run.