Local coding models
Coder-specialised open-weight models that run on your own hardware — the Qwen2.5-Coder family, DeepSeek-Coder, StarCoder2, Yi-Coder, CodeLlama. Filtered by HumanEval+ and MBPP+ scores where we've benchmarked.
Coding-tuned LLMs aren't just smaller versions of general-purpose chat models — they're trained on much more code, often with fill-in-the-middle objectives, and they post-train with execution feedback. The result is dramatic per-parameter strength on coding benchmarks compared to general models at the same size.
Headline benchmark on this laptop (RTX 3080, 16GB): qwen-2.5-coder-7b-instruct scored 81.1 HumanEval+ pass@1 + 66.9 MBPP+ pass@1 — comparable to commercial models 5-10x its size. The pattern repeats across the coder family: smaller models hit much higher coding scores than their general-chat siblings.
Each row links to the model's full operator notes including the actual prompting kit, recommended quantization, and benchmark scores (HumanEval+, MBPP+) we've run. Filter by 'commercial OK' if the license matters.
Qwen-based
Qwen3.6 27B, released April 22, 2026, is widely cited as the best all-round open model for consumer hardware in its size class, leading its size class on coding benchmarks per third-party trackers. Dense 27B under Apache
- License
- Apache-2.0 · OK
- Context
- 256K
Qwen3 Coder 30B-A3B is Alibaba's coding-specialized MoE (30B total, 3B active) under Apache-2.0, sitting at 7.4M Ollama pulls. It is the entry point into Alibaba's dedicated Qwen3-Coder line, whose smallest official vari
- License
- Apache-2.0 · OK
- Context
- 128K
Qwen3.5 9B is Alibaba's dense 9B release from March 2026 under Apache-2.0. It takes text and image input, runs a 256K context, and covers 201 languages. The qwen3.5 family is at 15.7M Ollama pulls (#14 in the library); t
- License
- Apache-2.0 · OK
- Context
- 256K
Qwen3.6 35B-A3B is the MoE sibling of the Qwen3.6 27B dense release (April 2026), 35B total with ~3B active, under Apache-2.0 with a 256K context.
- License
- Apache-2.0 · OK
- Context
- 256K
Qwen3.5 27B is Alibaba's dense flagship-tier release in the Qwen3.5 family, under Apache-2.0 with a 256K context. Vendor-reported SWE-bench Verified sits at 72.4, positioning it above the smaller Qwen3.5 9B on coding tas
- License
- Apache-2.0 · OK
- Context
- 256K
Qwen3.5 35B-A3B is the MoE variant of the Qwen3.5 family (35B total, ~3B active), under Apache-2.0 with a 256K context. The low active-parameter count gives it decode speed closer to a much smaller dense model.
- License
- Apache-2.0 · OK
- Context
- 256K
Qwen 3.6 35B-A3B with Multi-Token Prediction (MTP). The "A3B" suffix means ~3B activated parameters per token via Mixture-of-Experts — inference cost stays mid-tier while total parameter count climbs to 35B. MTP enables
- License
- Apache-2.0 · OK
- Context
- 256K
Qwen 3.6 27B dense (not MoE) with Multi-Token Prediction. Sits between the 14B and 35B-A3B as a "single dense model with MTP throughput acceleration." Targets workloads where the MoE activated-param dance isn't ideal but
- License
- Apache-2.0 · OK
- Context
- 128K
VibeThinker-3B is a compact open-weight reasoning model from WeiboAI (Sina Weibo), fine-tuned from Qwen2.5-Coder-3B (Hugging Face `WeiboAI/VibeThinker-3B`, 2026-06). It is a dense ~3B-parameter model (runs in roughly 6.7
- License
- MIT · OK
- Context
- 128K
CodeQwen 1.5 — Qwen Coder predecessor. Superseded by Qwen 2.5 Coder for new deployments.
- License
- Tongyi Qianwen L · OK
- Context
- 64K
Coding-specialized Qwen 2.5 at 14B. The 16GB-VRAM tier coding model — fits comfortably with 8K context.
- License
- Apache 2.0 · OK
- Context
- 128K
Compact Qwen 2.5 Coder. Sweet spot for laptop autocomplete and small refactor agents.
- License
- Apache 2.0 · OK
- Context
- 32K
Smallest Qwen 2.5 Coder. Targets edge / autocomplete on integrated GPUs and Apple Silicon laptops.
- License
- Apache 2.0 · OK
- Context
- 32K
Coding-specialized Qwen 2.5 at 7B. The 8-12GB-VRAM coding model — entry-tier autocomplete + IDE assistant. Smaller sibling of the 14B / 32B Coder line.
- License
- Apache 2.0 · OK
- Context
- 128K
Coding-specialized fine-tune of Qwen 3 32B. Curated coding corpus; outperforms Qwen 2.5 Coder 32B on SWE-Bench by ~6 points. Apache 2.0.
- License
- Apache 2.0 · OK
- Context
- 128K
Other / from-scratch
GPT-OSS 20B is OpenAI's open-weight 20.9B MoE, released August 2025 under Apache-2.0 and shipped natively in MXFP4 at 14GB. It is widely cited as the strongest general pick that fits a 16GB card, and sits at 11M Ollama p
- License
- Apache-2.0 · OK
- Context
- 128K
MiniMax-M3 is a native multimodal Mixture-of-Experts model from MiniMax (Hugging Face `MiniMaxAI/MiniMax-M3`, 2026-06), with ~428B total / ~23B active parameters per token. It accepts text, image, and video input and out
- License
- MiniMax Communit · OK
- Context
- 1024K
Ornith 1.0 9B is DeepReinforce's agentic-coding model, RL post-trained on Qwen3.5 and Gemma 4 bases and released June 25, 2026 under MIT. It runs a 256K context, and the Q4_K_M build is 5.6GB. The Ollama tag drew 262.9K
- License
- MIT · OK
- Context
- 256K
Laguna XS 2.1, released July 2, 2026, is Poolside's updated agentic-coding model (33B total, ~3B active MoE, 40 layers, 256 experts). It switched licensing from Apache-2.0 (the original Laguna XS.2) to OpenMDW-1.1, a per
- License
- OpenMDW-1.1 · OK
- Context
- 256K
North Mini Code, announced June 9, 2026, is Cohere's first open-weight developer model: a 30B-A3B MoE for agentic coding, small enough in active parameters to run locally. Cohere's official release names Apache-2.0, thou
- License
- Apache-2.0 · OK
- Context
- 256K
Ornith 1.0 35B is the larger MoE sibling of Ornith 9B (this catalog), DeepReinforce's RL-post-trained agentic-coding line, released June 25, 2026 under MIT with a 256K context.
- License
- MIT · OK
- Context
- 256K
Mellum2 12B-A2.5B is JetBrains' from-scratch SWE-focused MoE (12.15B total, 2.5B active), released June 2026 under Apache-2.0. Purpose-built for code completion and agentic SWE tasks rather than general chat; the Q4_K_M
- License
- Apache-2.0 · OK
- Context
- 128K
InclusionAI's Ring-2.6-1T is a 1 trillion parameter Mixture-of-Experts model with ~32B activated parameters per token, released by Ant Group's AI research arm. Targets frontier reasoning and code at MoE serving cost. Apa
- License
- Apache-2.0 · OK
- Context
- 125K
Sarvam-105B is a Mixture-of-Experts model with 105B total parameters but only 10.3B active at inference time. It targets reasoning, coding, and agentic tasks with coverage across 22 Indian languages. Apache 2.0 licensed
- License
- apache-2.0 · OK
- Context
- 125K
Mid-size StarCoder 2. The 8GB-VRAM autocomplete pick.
- License
- BigCode OpenRAIL · OK
- Context
- 16K
StarCoder 2 flagship. The largest BigCode coder; 16k context with strong fill-in-middle.
- License
- BigCode OpenRAIL · OK
- Context
- 16K
BigCode's StarCoder 2 at 3B. Trained on The Stack v2 with 600+ programming languages.
- License
- BigCode OpenRAIL · OK
- Context
- 16K
DeepSeek-based
DeepSeek's April 2026 frontier flagship. 1.6T total / 49B active MoE with hybrid Compressed Sparse Attention + Heavily Compressed Attention. 1M context window. Closes most of the gap with Claude Opus 4.6 on coding while
- License
- MIT · OK
- Context
- 1024K
The cost-efficient sibling of V4-Pro. 284B total / 13B active MoE, same hybrid CSA+HCA attention, same 1M context. The MoE active-param ratio (4.5%) makes it surprisingly fast for its nameplate size — practical on dual A
- License
- MIT · OK
- Context
- 1024K
Newer R1 distill on a Qwen 3 base. Combines R1 reasoning with Qwen 3's reasoning-toggle architecture. Apache 2.0.
- License
- Apache 2.0 · OK
- Context
- 128K
DeepSeek V2.5 — merged V2 chat + Coder. Pre-V3 baseline; 21B active MoE.
- License
- DeepSeek License · OK
- Context
- 128K
DeepSeek's coder line successor. Dense 33B; competitive with Qwen 2.5 Coder 32B on SWE-Bench.
- License
- DeepSeek License · OK
- Context
- 128K
R1 reasoning distilled into a Llama 3 8B base. Smaller R1 distill; useful when 32B is too heavy. Reasoning quality is meaningfully below the 32B distill but still beats non-reasoning Llama 8B on math/code.
- License
- Apache 2.0 · OK
- Context
- 128K
Distillation of DeepSeek V3 to a smaller MoE. 16B total / 2.4B active. Captures most of V3's reasoning at consumer-card-friendly memory.
- License
- DeepSeek License · OK
- Context
- 128K
DeepSeek's spring 2026 frontier MoE. 745B total / 38B active. The current open-weight benchmark leader on coding + math; closes the gap with closed-source flagships on reasoning.
- License
- DeepSeek License · OK
- Context
- 128K
Full DeepSeek Coder V2. 236B total / 21B active MoE coder.
- License
- DeepSeek License · OK
- Context
- 128K
Llama-based
Salamandra 2B is a base-only transformer trained from scratch by Barcelona Supercomputing Center on 12.875 trillion tokens across 35 European languages and code. At 2.25B parameters and an 8192-token context window, it i
- License
- apache-2.0 · OK
- Context
- 8K
Salamandra 7B is a base language model from Barcelona Supercomputing Center, pretrained on 12.875 trillion tokens across 35 European languages and code. It is not instruction-tuned — this is a raw foundation model. Apach
- License
- apache-2.0 · OK
- Context
- 8K
Phind's CodeLlama-derived coder at 34B. Older release; retained for historical / continuity value. Newer Qwen Coder lineage has surpassed it.
- License
- Llama 2 Communit · OK
- Context
- 16K
Llama 4 dense at 70B. Drop-in successor to Llama 3.3 70B; same hardware envelope, better on reasoning benchmarks.
- License
- Llama 4 Communit · OK
- Context
- 128K
Mistral-based
Ministral 3 14B is Mistral AI's dense 14B from December 2025, released under Apache-2.0 with vision input and a 256K context. The Q4_K_M build is 9.1GB, which targets 16GB cards once KV cache is counted.
- License
- Apache-2.0 · OK
- Context
- 256K
Mistral's Mamba (state-space) architecture coding model. Linear inference cost — the architectural alternative to attention-based coding models. Apache 2.0.
- License
- Apache 2.0 · OK
- Context
- 250K
Mistral's coding-specialized Mistral Small 2 successor. Apache 2.0 — the rare commercial-OK Mistral coder.
- License
- Apache 2.0 · OK
- Context
- 128K
Gemma-based
Gemma 4 12B is Google's dense 12B release from April 2026, with the unified 12B variant landing in June 2026. It runs a 256K context, takes vision and audio input through an encoder-free multimodal design, and supports t
- License
- Apache-2.0 · OK
- Context
- 256K
Gemma 4 26B-A4B is Google's MoE variant of the Gemma 4 family (26B total, 4B active), released alongside the dense Gemma 4 lineup in 2026. The 4B active-parameter footprint gives faster decode than the dense 12B/31B sibl
- License
- Gemma Terms of U · OK
- Context
- 128K
GLM-based
GLM-5.2 is Zhipu AI's (Z.ai) flagship open-weight LLM, released on Hugging Face as `zai-org/GLM-5.2` (2026-06). It is a Mixture-of-Experts decoder-only transformer with a new "IndexShare" sparse-attention design (one ind
- License
- MIT · OK
- Context
- 1024K
GLM-4.7-Flash, released January 19, 2026, is Zhipu's lightweight sibling of the 355B-class GLM-4.7 flagship: a 31B-A3B MoE under MIT, built for local coding and agent workloads. Ollama describes it as the strongest model
- License
- MIT · OK
- Context
- 198K
dolphin
moonshot
Granite-based
Yi-based
opencoder
Coding agent setup?
Cross-reference the model rows here with the runtime guidance at /apps for Cline, Aider, Continue, OpenInterpreter, and Claude Code adapters. The 'best coding agent for local models' Q&A digs into the actual workflows. Read the coding-agent comparison.