The frontier of open-weight model releases
Open-weight model releases tracked by RunLocalAI — recent additions, rising families, distill chains, multimodal and reasoning waves. Each card links into the catalog with authority badges (L1.25 enriched · benchmark-backed · verdict) so you can scan editorial coverage at a glance.
Filtered results (48)
Models matching your filters. Clear filters by clicking “Any” on each row above, or remove individual filters via the URL.
Hunyuan 3.0 (Hy3)
Apache-licensed agent/tool-calling backbone for teams with 8x datacenter GPUs
LongCat-2.0
MIT-licensed near-frontier agentic coding for multi-node datacenter deployments
Laguna XS 2.1
agentic coding on 24GB+ GPUs
Ornith 1.0 35B
agentic coding on 24GB+ GPUs
Ornith 1.0 9B
agentic coding on 8-12GB GPUs
LFM2.5-230M
On-device tool calling and data extraction on phones and Raspberry Pi-class hardware
GLM-5.2
Open-weight 1M-context MoE for long-horizon agentic coding
Kimi K2.7-Code
Open-weight 1T MoE for long-horizon software engineering
VibeThinker-3B
Compact MIT reasoning model that runs on a single consumer GPU (~6.7GB)
North Mini Code 1.0
agentic coding on 24GB+ GPUs
Nemotron 3 Ultra (550B-A55B)
Open-weight frontier-scale long-context reasoning for complex agentic workflows
Ring-2.6-1T
frontier reasoning at MoE serving cost
Qwen 3.6 35B-A3B (MTP)
high-throughput MoE inference at workstation tier
Qwen 3.6 27B (MTP)
dense workstation model with throughput-acceleration
Qwen 3.5 235B-A17B (MoE)
frontier-tier reasoning + multilingual serving on multi-machine clusters
Mistral Medium 3.5 (675B MoE)
frontier MoE — Mistral's response to the open MoE wave
Granite 4.1 30B Instruct
Workstation-class agent and RAG backends where tool-call accuracy matters more than throughput
Mistral Medium 3 24B (dense)
research / non-commercial workstation deployments
Granite 4.1 8B Instruct
Local agents and RAG on a single consumer GPU, with best-in-class tool calling per GB under Apache 2.0
Granite 4.1 3B Instruct
Edge-deployed function calling and RAG extraction where every GB of memory counts
DeepSeek V4 Pro (1.6T MoE)
frontier-tier coding + reasoning serving — currently the open-weight ceiling
DeepSeek V4 Flash (284B MoE)
datacenter MoE — V4 efficiency variant
Qwen3.6 27B
best all-round pick on 24GB+ GPUs
OLMo 2 32B
fully-open AI2 OLMo 2 — research provenance flagship
Phi-4 Reasoning Mini 4B
edge-tier reasoning
Llama 4 Scout
production multimodal serving — image + text at workstation-cluster scale
DeepSeek V4
frontier-tier reasoning on multi-machine clusters
Granite 3.3 8B
enterprise tool-calling on IBM stacks
Kimi K2.6
Moonshot frontier MoE — long-context specialist
Mistral Small 3.2 24B
consumer-tier multilingual instruction-following
Phi-4 Mini 4B
edge / embedded reasoning
GLM-5 Pro
Chinese-language enterprise serving
Nemotron 3 Super (120B-A12B)
NVIDIA-tuned datacenter-tier reasoning
Llama 4 405B
frontier-tier serving on cluster hardware
Llama 4 70B
production self-hosted serving on 2x A100 / H100
DeepSeek Coder V3
workstation coding alternative to Qwen 2.5 Coder
GLM-5
Zhipu GLM-5 frontier MoE
Nemotron 3 Super 49B
32GB-VRAM enterprise deployments
Nemotron 3 Nano 9B
NVIDIA-stack tool-calling agents
GLM-4.7-Flash
fast local coding and agents on 24GB+ GPUs
Nemotron 3 Nano (30B-A3B)
NVIDIA-tuned consumer-tier general
DeepSeek V3 Lite (16B MoE)
consumer-tier MoE inference
Hermes 4 Llama 3.3 70B
datacenter-tier instruction-tuned alternative to base Llama 3.3
Magistral 32B
research / non-commercial reasoning at 32B scale
Kimi K1.5
deep math + reasoning research
Qwen 3 Coder 32B
coding-specialized agent workloads
DeepSeek R1 Distill Qwen 3 32B
workstation reasoning with Qwen 3 base improvements
EXAONE 3.5 32B
Korean / Japanese / CJK workloads
Going deeper
- Ecosystem maps — structured-landscape views (memory frameworks, inference runtimes, MCP, coding agents).
- Execution stacks — recipes that combine models with runtimes + hardware.
- Frontier index — broader ecosystem-momentum view across coding agents, inference runtimes, memory systems, MCP.
- Benchmarks — measured tokens-per-second + topology fields across hardware/model/runtime triples.