BLK · CODING MODELScoder-specialised · HumanEval+ · MBPP+

Local coding models

Coder-specialised open-weight models that run on your own hardware — the Qwen2.5-Coder family, DeepSeek-Coder, StarCoder2, Yi-Coder, CodeLlama. Filtered by HumanEval+ and MBPP+ scores where we've benchmarked.

Models curated
52
Vendors
25
Commercial OK
52/52
Benchmarked
0/52

Coding-tuned LLMs aren't just smaller versions of general-purpose chat models — they're trained on much more code, often with fill-in-the-middle objectives, and they post-train with execution feedback. The result is dramatic per-parameter strength on coding benchmarks compared to general models at the same size.

Headline benchmark on this laptop (RTX 3080, 16GB): qwen-2.5-coder-7b-instruct scored 81.1 HumanEval+ pass@1 + 66.9 MBPP+ pass@1 — comparable to commercial models 5-10x its size. The pattern repeats across the coder family: smaller models hit much higher coding scores than their general-chat siblings.

Each row links to the model's full operator notes including the actual prompting kit, recommended quantization, and benchmark scores (HumanEval+, MBPP+) we've run. Filter by 'commercial OK' if the license matters.

FAM · QWEN

Qwen-based

15 models
Qwen3.6 27B
27B params · Alibaba
▸ best all-round pick on 24GB+ GPUs

Qwen3.6 27B, released April 22, 2026, is widely cited as the best all-round open model for consumer hardware in its size class, leading its size class on coding benchmarks per third-party trackers. Dense 27B under Apache

License
Apache-2.0 · OK
Context
256K
Qwen3 Coder 30B-A3B
30B params · Alibaba
▸ dedicated coding assistant on 24GB+ GPUs

Qwen3 Coder 30B-A3B is Alibaba's coding-specialized MoE (30B total, 3B active) under Apache-2.0, sitting at 7.4M Ollama pulls. It is the entry point into Alibaba's dedicated Qwen3-Coder line, whose smallest official vari

License
Apache-2.0 · OK
Context
128K
Qwen3.5 9B
9B params · Alibaba
▸ multilingual chat and coding on 8-12GB GPUs

Qwen3.5 9B is Alibaba's dense 9B release from March 2026 under Apache-2.0. It takes text and image input, runs a 256K context, and covers 201 languages. The qwen3.5 family is at 15.7M Ollama pulls (#14 in the library); t

License
Apache-2.0 · OK
Context
256K
Qwen3.6 35B-A3B
35B params · Alibaba
▸ fast MoE decode on 32GB GPUs

Qwen3.6 35B-A3B is the MoE sibling of the Qwen3.6 27B dense release (April 2026), 35B total with ~3B active, under Apache-2.0 with a 256K context.

License
Apache-2.0 · OK
Context
256K
Qwen3.5 27B
27B params · Alibaba
▸ general reasoning and coding on 24GB+ GPUs

Qwen3.5 27B is Alibaba's dense flagship-tier release in the Qwen3.5 family, under Apache-2.0 with a 256K context. Vendor-reported SWE-bench Verified sits at 72.4, positioning it above the smaller Qwen3.5 9B on coding tas

License
Apache-2.0 · OK
Context
256K
Qwen3.5 35B-A3B
35B params · Alibaba
▸ fast MoE decode on 32GB GPUs

Qwen3.5 35B-A3B is the MoE variant of the Qwen3.5 family (35B total, ~3B active), under Apache-2.0 with a 256K context. The low active-parameter count gives it decode speed closer to a much smaller dense model.

License
Apache-2.0 · OK
Context
256K
Qwen 3.6 35B-A3B (MTP)
35B params · Alibaba / Qwen team
judged 8.0/10
▸ high-throughput MoE inference at workstation tier

Qwen 3.6 35B-A3B with Multi-Token Prediction (MTP). The "A3B" suffix means ~3B activated parameters per token via Mixture-of-Experts — inference cost stays mid-tier while total parameter count climbs to 35B. MTP enables

License
Apache-2.0 · OK
Context
256K
Qwen 3.6 27B (MTP)
27B params · Alibaba / Qwen team
judged 8.0/10
▸ dense workstation model with throughput-acceleration

Qwen 3.6 27B dense (not MoE) with Multi-Token Prediction. Sits between the 14B and 35B-A3B as a "single dense model with MTP throughput acceleration." Targets workloads where the MoE activated-param dance isn't ideal but

License
Apache-2.0 · OK
Context
128K
VibeThinker-3B
3B params · WeiboAI (Sina Weibo)
▸ Compact MIT reasoning model that runs on a single consumer GPU (~6.7GB)

VibeThinker-3B is a compact open-weight reasoning model from WeiboAI (Sina Weibo), fine-tuned from Qwen2.5-Coder-3B (Hugging Face `WeiboAI/VibeThinker-3B`, 2026-06). It is a dense ~3B-parameter model (runs in roughly 6.7

License
MIT · OK
Context
128K
CodeQwen 1.5 7B
7B params · Alibaba
▸ historical reference — Qwen 2.5 Coder 7B is the modern pick

CodeQwen 1.5 — Qwen Coder predecessor. Superseded by Qwen 2.5 Coder for new deployments.

License
Tongyi Qianwen L · OK
Context
64K
Qwen 2.5 Coder 14B Instruct
14B params · Alibaba
▸ 16GB-VRAM coding

Coding-specialized Qwen 2.5 at 14B. The 16GB-VRAM tier coding model — fits comfortably with 8K context.

License
Apache 2.0 · OK
Context
128K
Qwen 2.5 Coder 3B
3B params · Alibaba
▸ Apple Silicon laptop coding autocomplete

Compact Qwen 2.5 Coder. Sweet spot for laptop autocomplete and small refactor agents.

License
Apache 2.0 · OK
Context
32K
Qwen 2.5 Coder 1.5B
1.5B params · Alibaba
▸ IDE autocomplete on integrated GPUs

Smallest Qwen 2.5 Coder. Targets edge / autocomplete on integrated GPUs and Apple Silicon laptops.

License
Apache 2.0 · OK
Context
32K
Qwen 2.5 Coder 7B Instruct
7B params · Alibaba
▸ consumer-tier coding at 8GB VRAM

Coding-specialized Qwen 2.5 at 7B. The 8-12GB-VRAM coding model — entry-tier autocomplete + IDE assistant. Smaller sibling of the 14B / 32B Coder line.

License
Apache 2.0 · OK
Context
128K
Qwen 3 Coder 32B
32B params · Alibaba
▸ coding-specialized agent workloads

Coding-specialized fine-tune of Qwen 3 32B. Curated coding corpus; outperforms Qwen 2.5 Coder 32B on SWE-Bench by ~6 points. Apache 2.0.

License
Apache 2.0 · OK
Context
128K
FAM · OTHER

Other / from-scratch

12 models
GPT-OSS 20B
20.9B params · OpenAI
▸ strongest general pick for 16GB cards

GPT-OSS 20B is OpenAI's open-weight 20.9B MoE, released August 2025 under Apache-2.0 and shipped natively in MXFP4 at 14GB. It is widely cited as the strongest general pick that fits a 16GB card, and sits at 11M Ollama p

License
Apache-2.0 · OK
Context
128K
MiniMax-M3
428B params · MiniMax
▸ Open-weight 1M-context multimodal MoE for agentic coding + video understanding

MiniMax-M3 is a native multimodal Mixture-of-Experts model from MiniMax (Hugging Face `MiniMaxAI/MiniMax-M3`, 2026-06), with ~428B total / ~23B active parameters per token. It accepts text, image, and video input and out

License
MiniMax Communit · OK
Context
1024K
Ornith 1.0 9B
9B params · DeepReinforce
▸ agentic coding on 8-12GB GPUs

Ornith 1.0 9B is DeepReinforce's agentic-coding model, RL post-trained on Qwen3.5 and Gemma 4 bases and released June 25, 2026 under MIT. It runs a 256K context, and the Q4_K_M build is 5.6GB. The Ollama tag drew 262.9K

License
MIT · OK
Context
256K
Laguna XS 2.1
33B params · Poolside
▸ agentic coding on 24GB+ GPUs

Laguna XS 2.1, released July 2, 2026, is Poolside's updated agentic-coding model (33B total, ~3B active MoE, 40 layers, 256 experts). It switched licensing from Apache-2.0 (the original Laguna XS.2) to OpenMDW-1.1, a per

License
OpenMDW-1.1 · OK
Context
256K
North Mini Code 1.0
30B params · Cohere
▸ agentic coding on 24GB+ GPUs

North Mini Code, announced June 9, 2026, is Cohere's first open-weight developer model: a 30B-A3B MoE for agentic coding, small enough in active parameters to run locally. Cohere's official release names Apache-2.0, thou

License
Apache-2.0 · OK
Context
256K
Ornith 1.0 35B
35B params · DeepReinforce
▸ agentic coding on 24GB+ GPUs

Ornith 1.0 35B is the larger MoE sibling of Ornith 9B (this catalog), DeepReinforce's RL-post-trained agentic-coding line, released June 25, 2026 under MIT with a 256K context.

License
MIT · OK
Context
256K
Mellum2 12B-A2.5B
12.15B params · JetBrains
▸ code completion and agentic SWE tasks on 12GB+ GPUs

Mellum2 12B-A2.5B is JetBrains' from-scratch SWE-focused MoE (12.15B total, 2.5B active), released June 2026 under Apache-2.0. Purpose-built for code completion and agentic SWE tasks rather than general chat; the Q4_K_M

License
Apache-2.0 · OK
Context
128K
Ring-2.6-1T
1000B params · InclusionAI / Ant Group
judged 8.0/10
▸ frontier reasoning at MoE serving cost

InclusionAI's Ring-2.6-1T is a 1 trillion parameter Mixture-of-Experts model with ~32B activated parameters per token, released by Ant Group's AI research arm. Targets frontier reasoning and code at MoE serving cost. Apa

License
Apache-2.0 · OK
Context
125K
Sarvam 105B
105B params · sarvamai
judged 9.3/10
▸ Hindi and Indian-language reasoning or agentic workflows

Sarvam-105B is a Mixture-of-Experts model with 105B total parameters but only 10.3B active at inference time. It targets reasoning, coding, and agentic tasks with coverage across 22 Indian languages. Apache 2.0 licensed

License
apache-2.0 · OK
Context
125K
StarCoder 2 7B
7B params · BigCode
▸ consumer-tier code completion at 8GB

Mid-size StarCoder 2. The 8GB-VRAM autocomplete pick.

License
BigCode OpenRAIL · OK
Context
16K
StarCoder 2 15B
15B params · BigCode
▸ permissively-licensed coding at 16GB-VRAM

StarCoder 2 flagship. The largest BigCode coder; 16k context with strong fill-in-middle.

License
BigCode OpenRAIL · OK
Context
16K
StarCoder 2 3B
3B params · BigCode
▸ edge-tier code completion

BigCode's StarCoder 2 at 3B. Trained on The Stack v2 with 600+ programming languages.

License
BigCode OpenRAIL · OK
Context
16K
FAM · DEEPSEEK

DeepSeek-based

9 models
DeepSeek V4 Pro (1.6T MoE)
1600B params · DeepSeek
▸ frontier-tier coding + reasoning serving — currently the open-weight ceiling

DeepSeek's April 2026 frontier flagship. 1.6T total / 49B active MoE with hybrid Compressed Sparse Attention + Heavily Compressed Attention. 1M context window. Closes most of the gap with Claude Opus 4.6 on coding while

License
MIT · OK
Context
1024K
DeepSeek V4 Flash (284B MoE)
284B params · DeepSeek
▸ datacenter MoE — V4 efficiency variant

The cost-efficient sibling of V4-Pro. 284B total / 13B active MoE, same hybrid CSA+HCA attention, same 1M context. The MoE active-param ratio (4.5%) makes it surprisingly fast for its nameplate size — practical on dual A

License
MIT · OK
Context
1024K
DeepSeek R1 Distill Qwen 3 32B
32B params · DeepSeek AI
▸ workstation reasoning with Qwen 3 base improvements

Newer R1 distill on a Qwen 3 base. Combines R1 reasoning with Qwen 3's reasoning-toggle architecture. Apache 2.0.

License
Apache 2.0 · OK
Context
128K
DeepSeek V2.5 236B
236B params · DeepSeek
▸ DeepSeek lineage reference — pre-V3

DeepSeek V2.5 — merged V2 chat + Coder. Pre-V3 baseline; 21B active MoE.

License
DeepSeek License · OK
Context
128K
DeepSeek Coder V3
33B params · DeepSeek AI
▸ workstation coding alternative to Qwen 2.5 Coder

DeepSeek's coder line successor. Dense 33B; competitive with Qwen 2.5 Coder 32B on SWE-Bench.

License
DeepSeek License · OK
Context
128K
DeepSeek R1 Distill Llama 8B
8B params · DeepSeek AI
▸ consumer-tier reasoning on 8GB+ GPUs

R1 reasoning distilled into a Llama 3 8B base. Smaller R1 distill; useful when 32B is too heavy. Reasoning quality is meaningfully below the 32B distill but still beats non-reasoning Llama 8B on math/code.

License
Apache 2.0 · OK
Context
128K
DeepSeek V3 Lite (16B MoE)
16B params · DeepSeek AI
▸ consumer-tier MoE inference

Distillation of DeepSeek V3 to a smaller MoE. 16B total / 2.4B active. Captures most of V3's reasoning at consumer-card-friendly memory.

License
DeepSeek License · OK
Context
128K
DeepSeek V4
745B params · DeepSeek AI
▸ frontier-tier reasoning on multi-machine clusters

DeepSeek's spring 2026 frontier MoE. 745B total / 38B active. The current open-weight benchmark leader on coding + math; closes the gap with closed-source flagships on reasoning.

License
DeepSeek License · OK
Context
128K
DeepSeek Coder V2 236B
236B params · DeepSeek
▸ datacenter-tier MoE coding

Full DeepSeek Coder V2. 236B total / 21B active MoE coder.

License
DeepSeek License · OK
Context
128K
FAM · LLAMA

Llama-based

4 models
FAM · MISTRAL

Mistral-based

3 models
FAM · GEMMA

Gemma-based

2 models
FAM · GLM

GLM-based

2 models
FAM · DOLPHIN

dolphin

1 model
FAM · MOONSHOT

moonshot

1 model
FAM · GRANITE

Granite-based

1 model
FAM · YI

Yi-based

1 model
FAM · OPENCODER

opencoder

1 model
COVERAGE

Coding agent setup?

Cross-reference the model rows here with the runtime guidance at /apps for Cline, Aider, Continue, OpenInterpreter, and Claude Code adapters. The 'best coding agent for local models' Q&A digs into the actual workflows. Read the coding-agent comparison.