BLK · CODING BENCHMARKSfirst-party · per-quant · reproducible

Coding benchmark leaderboard

Reviewed HumanEval+ scores for open-weight local models at real-world quantizations. First-party rows were run by RunLocalAI; community rows render only after review. Every public row links to its raw test-run Gist.

Test runs
30
Unique models
25
Quant variants
1
Benchmark suites
3
METHODOLOGY

How these numbers are produced

  1. Each model loads into the listed runtime (Ollama / vLLM / llama.cpp) at the listed quantization on the listed hardware.
  2. EvalPlus is used as the official scorer. For OpenAI-compatible local runtimes, we first generate deterministic JSONL samples, then score them with python -m evalplus.evaluate.
  3. Greedy sampling (deterministic; one sample per problem). Score = pass@1 on HumanEval+ (164 problems with augmented tests) × 100.
  4. Raw stdout + stderr captured and posted to a public GitHub Gist before the row enters this table. Click any score to verify.
  5. Runner source: scripts/run-humaneval-plus.ts (public repo) — reproducible end-to-end.
#ModelBenchmarkQuantRigScoreTrustLog
1Muse Glimmer 30B
tested 2026-08-10
HumanEval
Q4_K_M
llama.cpp-4dee52f-vast5090
rtx-5090
97.6
First-party
command published
Gist →
2Gemma 4 26B-A4B
tested 2026-07-20
HumanEval+
Q4_K_M
ollama-0.32.1-vast5090
rtx-5090
93.3
First-party
command published
Gist →
3Gemma 4 31B Dense
tested 2026-07-20
HumanEval+
Q4_K_M
ollama-0.32.1-vast5090
rtx-5090
93.3
First-party
command published
Gist →
4Muse Glimmer 30B
tested 2026-08-10
HumanEval+
Q4_K_M
llama.cpp-4dee52f-vast5090
rtx-5090
91.5
First-party
command published
Gist →
5North Mini Code 1.0
tested 2026-07-21
HumanEval+
Q4_K_M
ollama-0.32.1-vast5090
rtx-5090
89.6
First-party
command published
Gist →
6Ornith 1.0 35B
tested 2026-07-21
HumanEval+
Q4_K_M
ollama-0.32.1-vast5090
rtx-5090
88.4
First-party
command published
Gist →
7Qwen3.5 35B-A3B
tested 2026-07-21
HumanEval+
Q4_K_M
ollama-0.32.1-vast5090
rtx-5090
87.8
First-party
command published
Gist →
8Qwen3 Coder 30B-A3B
tested 2026-07-20
HumanEval+
Q4_K_M
ollama-0.32.1-vast5090
rtx-5090
87.8
First-party
command published
Gist →
9Laguna XS 2.1
tested 2026-07-21
HumanEval+
Q4_K_M
ollama-0.32.1-vast5090
rtx-5090
87.2
First-party
command published
Gist →
10Qwen3.5 27B
tested 2026-07-20
HumanEval+
Q4_K_M
ollama-0.32.1-vast5090
rtx-5090
87.2
First-party
command published
Gist →
11Qwen3.6 27B
tested 2026-07-20
HumanEval+
Q4_K_M
ollama-0.32.1-vast5090
rtx-5090
86.6
First-party
command published
Gist →
12Nemotron 3 Nano Omni 33B
tested 2026-07-21
HumanEval+
Q4_K_M
ollama-0.32.1-vast5090
rtx-5090
85.4
First-party
command published
Gist →
13Granite 4.1 30B
tested 2026-07-20
HumanEval+
Q4_K_M
ollama-0.32.1-vast5090
rtx-5090
85.4
First-party
command published
Gist →
14GLM-4.7-Flash
tested 2026-07-21
HumanEval+
Q4_K_M
ollama-0.32.1-vast5090
rtx-5090
85.4
First-party
command published
Gist →
15Qwen 2.5 Coder 7B Instruct
tested 2026-05-28
HumanEval+
Q4_K_M
ollama-0.24
rtx-3080-16gb-mobile
81.1
First-party
command published
Gist →
16Qwen3.6 35B-A3B
tested 2026-07-21
HumanEval+
Q4_K_M
ollama-0.32.1-vast5090
rtx-5090
81.1
First-party
command published
Gist →
17Phi-4 14B
tested 2026-05-28
HumanEval+
Q4_K_M
ollama-0.24
rtx-3080-16gb-mobile
78.7
First-party
command published
Gist →
18Mellum2 12B-A2.5B
tested 2026-07-20
HumanEval+
Q4_K_M
ollama-0.32.1-vast5090
rtx-5090
76.8
First-party
command published
Gist →
19Ministral 3 14B
tested 2026-07-20
HumanEval+
Q4_K_M
ollama-0.32
rtx-3080-16gb-mobile
76.8
First-party
command published
Gist →
20Ornith 1.0 9B
tested 2026-07-18
HumanEval+
Q4_K_M
ollama-0.32
rtx-3080-16gb-mobile
73.2
First-party
command published
Gist →
21Trendyol LLM Asure 12B
tested 2026-05-27
MBPP+
Q4_K_M
ollama-0.24.0
rtx-5080
71.7
First-party
command published
Gist →
22Trendyol LLM Asure 12B
tested 2026-05-27
HumanEval+
Q4_K_M
ollama-0.24.0
rtx-5080
69.5
First-party
command published
Gist →
23Qwen 2.5 Coder 7B Instruct
tested 2026-05-29
MBPP+
Q4_K_M
ollama-0.24
rtx-3080-16gb-mobile
66.9
First-party
command published
Gist →
24Phi-4 14B
tested 2026-05-29
MBPP+
Q4_K_M
ollama-0.24
rtx-3080-16gb-mobile
60.3
First-party
command published
Gist →
25Dolphin 3.0 8B
tested 2026-07-20
HumanEval+
Q4_K_M
ollama-0.32.1-vast5090
rtx-5090
56.7
First-party
command published
Gist →
26Llama 3.1 8B Instruct
tested 2026-05-28
HumanEval+
Q4_K_M
ollama-0.24
rtx-3080-16gb-mobile
56.1
First-party
command published
Gist →
27Qwen3.5 9B
tested 2026-07-19
HumanEval+
Q4_K_M
ollama-0.32
rtx-3080-16gb-mobile
42.7
First-party
command published
Gist →
28Hermes 3 Llama 3.1 8B
tested 2026-07-20
HumanEval+
Q4_K_M
ollama-0.32.1-vast5090
rtx-5090
41.5
First-party
command published
Gist →
29Llama 3.1 8B Instruct
tested 2026-05-29
MBPP+
Q4_K_M
ollama-0.24
rtx-3080-16gb-mobile
39.2
First-party
command published
Gist →
30Qwen 3 8B
tested 2026-05-29
HumanEval+
Q4_K_M
ollama-0.24
rtx-3080-16gb-mobile
2.4
First-party
command published
Gist →

A note on quant comparisons. The same model at different quantization levels can score materially differently — especially on code, where precision loss hurts. When you see Q4_K_M next to Q6_K for the same model, the gap is the “quant tax” on coding specifically. Take that into account when picking what to run on /will-it-run.