The frontier of open-weight model releases
Open-weight model releases tracked by RunLocalAI — recent additions, rising families, distill chains, multimodal and reasoning waves. Each card links into the catalog with authority badges (L1.25 enriched · benchmark-backed · verdict) so you can scan editorial coverage at a glance.
Recent releases (12 newest)
Catalog entries with the most recent release dates. Use the authority badges to spot which have full editorial coverage (L1.25 enriched + benchmark) and which are catalog-only.
Muse Glimmer 30B
Hunyuan 3.0 (Hy3)
Apache-licensed agent/tool-calling backbone for teams with 8x datacenter GPUs
LongCat-2.0
MIT-licensed near-frontier agentic coding for multi-node datacenter deployments
Laguna XS 2.1
agentic coding on 24GB+ GPUs
Ornith 1.0 35B
agentic coding on 24GB+ GPUs
Ornith 1.0 9B
agentic coding on 8-12GB GPUs
LFM2.5-230M
On-device tool calling and data extraction on phones and Raspberry Pi-class hardware
GLM-5.2
Open-weight 1M-context MoE for long-horizon agentic coding
Kimi K2.7-Code
Open-weight 1T MoE for long-horizon software engineering
VibeThinker-3B
Compact MIT reasoning model that runs on a single consumer GPU (~6.7GB)
North Mini Code 1.0
agentic coding on 24GB+ GPUs
Nemotron 3 Ultra (550B-A55B)
Open-weight frontier-scale long-context reasoning for complex agentic workflows
New reasoning models
Models with explicit thinking-block emission — DeepSeek R1 family, QwQ, Kimi, Magistral, Qwen 3 reasoning-mode. /stacks/local-reasoning-model for the canonical deployment recipe.
Kimi K2.6
Moonshot frontier MoE — long-context specialist
Magistral 32B
research / non-commercial reasoning at 32B scale
Kimi K1.5
deep math + reasoning research
Qwen 3 Coder 32B
coding-specialized agent workloads
DeepSeek R1 Distill Qwen 3 32B
workstation reasoning with Qwen 3 base improvements
Qwen 3 235B-A22B
Qwen 3 MoE flagship — pre-3.5 baseline
New coding models
Coding-specialized fine-tunes. The Qwen Coder lineage is the current open-weight benchmark leader; DeepSeek Coder V3, Codestral, Devstral, OpenCoder are the credible alternatives. /stacks/local-coding-agent for the canonical deployment recipe.
DeepSeek Coder V3
workstation coding alternative to Qwen 2.5 Coder
Devstral Small 2 24B
Apache 2.0 coding alternative to Qwen 2.5 Coder
Yi Coder 9B
8GB-VRAM coding
Qwen 2.5 Coder 32B Instruct
single-user autonomous coding agents on RTX 4090 / 5090 / dual-A100 hardware
Qwen 2.5 Coder 14B Instruct
16GB-VRAM coding
Qwen 2.5 Coder 7B Instruct
consumer-tier coding at 8GB VRAM
New multimodal models
Vision-language models. The 2025-2026 wave: Llama 4 Scout / Maverick, Qwen 2.5-VL, Pixtral, Janus-Pro, Phi-4 Multimodal. /stacks/local-vision-model for the canonical deployment recipe.
Muse Glimmer 30B
MiniMax-M3
Open-weight 1M-context multimodal MoE for agentic coding + video understanding
Llama 4 Maverick
frontier-tier multimodal serving on multi-machine clusters
Gemma 4 31B Dense
workstation-tier multilingual chat with permissive license
Gemma 4 26B MoE
Gemma 4 MoE — workstation efficiency variant
Gemma 4 12B
multimodal general assistant on 12-16GB GPUs
New MoE models
Mixture-of-Experts releases. Active-parameter efficiency shapes the deployment economics. See /systems/distributed-inference for the architectural depth.
Hunyuan 3.0 (Hy3)
Apache-licensed agent/tool-calling backbone for teams with 8x datacenter GPUs
LongCat-2.0
MIT-licensed near-frontier agentic coding for multi-node datacenter deployments
LFM2.5-230M
On-device tool calling and data extraction on phones and Raspberry Pi-class hardware
Ring-2.6-1T
frontier reasoning at MoE serving cost
Qwen 3.6 35B-A3B (MTP)
high-throughput MoE inference at workstation tier
Qwen 3.5 235B-A17B (MoE)
frontier-tier reasoning + multilingual serving on multi-machine clusters
New edge / phone-tier models
Sub-4B models for phone / Pi / embedded deployment. Phi-4 Mini, Gemma 3 1B, MiniCPM 3 4B, SmolLM 3, Hermes 3 3B, Dolphin 3 3B, RWKV 7 Goose 1.5B.
LFM2.5-230M
On-device tool calling and data extraction on phones and Raspberry Pi-class hardware
Granite 4.1 3B Instruct
Edge-deployed function calling and RAG extraction where every GB of memory counts
Phi-4 Reasoning Mini 4B
edge-tier reasoning
Gemma 4 E4B (Effective 4B)
edge-tier Gemma 4 — laptop friendly
Gemma 4 E2B (Effective 2B)
phone-tier Gemma 4
Phi-4 Mini 4B
edge / embedded reasoning
Enrichment gaps — OPERATOR queue
High-relevance catalog entries (7B-100B) that lack L1.25 enrichment, verdict, AND benchmark. These render noindex today — the next sprint's editorial queue. Surfacing them here keeps the gap visible.
Gemma 4 12B
multimodal general assistant on 12-16GB GPUs
Qwen3.5 9B
multilingual chat and coding on 8-12GB GPUs
GPT-OSS 20B
strongest general pick for 16GB cards
Ministral 3 14B
general chat with vision on 16GB GPUs
Ornith 1.0 9B
agentic coding on 8-12GB GPUs
Turkish Gemma 9B T1
Trendyol LLM 7B Chat v0.1
Turkish Llama 8B Instruct v0.1
Going deeper
- Ecosystem maps — structured-landscape views (memory frameworks, inference runtimes, MCP, coding agents).
- Execution stacks — recipes that combine models with runtimes + hardware.
- Frontier index — broader ecosystem-momentum view across coding agents, inference runtimes, memory systems, MCP.
- Benchmarks — measured tokens-per-second + topology fields across hardware/model/runtime triples.