The frontier of open-weight model releases
Open-weight model releases tracked by RunLocalAI — recent additions, rising families, distill chains, multimodal and reasoning waves. Each card links into the catalog with authority badges (L1.25 enriched · benchmark-backed · verdict) so you can scan editorial coverage at a glance.
Filtered results (48)
Models matching your filters. Clear filters by clicking “Any” on each row above, or remove individual filters via the URL.
Muse Glimmer 30B
Laguna XS 2.1
agentic coding on 24GB+ GPUs
Ornith 1.0 35B
agentic coding on 24GB+ GPUs
Ornith 1.0 9B
agentic coding on 8-12GB GPUs
North Mini Code 1.0
agentic coding on 24GB+ GPUs
Qwen 3.6 35B-A3B (MTP)
high-throughput MoE inference at workstation tier
Qwen 3.6 27B (MTP)
dense workstation model with throughput-acceleration
Granite 4.1 30B Instruct
Workstation-class agent and RAG backends where tool-call accuracy matters more than throughput
Mistral Medium 3 24B (dense)
research / non-commercial workstation deployments
Granite 4.1 8B Instruct
Local agents and RAG on a single consumer GPU, with best-in-class tool calling per GB under Apache 2.0
Qwen3.6 27B
best all-round pick on 24GB+ GPUs
OLMo 2 32B
fully-open AI2 OLMo 2 — research provenance flagship
Gemma 4 26B MoE
Gemma 4 MoE — workstation efficiency variant
Gemma 4 12B
multimodal general assistant on 12-16GB GPUs
Granite 3.3 8B
enterprise tool-calling on IBM stacks
Mistral Small 3.2 24B
consumer-tier multilingual instruction-following
Qwen3.5 9B
multilingual chat and coding on 8-12GB GPUs
Phi-4 Multimodal
16GB-consumer multimodal Q&A
Llama 4 70B
production self-hosted serving on 2x A100 / H100
DeepSeek Coder V3
workstation coding alternative to Qwen 2.5 Coder
Nemotron 3 Super 49B
32GB-VRAM enterprise deployments
Nemotron 3 Nano 9B
NVIDIA-stack tool-calling agents
GLM-4.7-Flash
fast local coding and agents on 24GB+ GPUs
Nemotron 3 Nano (30B-A3B)
NVIDIA-tuned consumer-tier general
DeepSeek V3 Lite (16B MoE)
consumer-tier MoE inference
Hermes 4 Llama 3.3 70B
datacenter-tier instruction-tuned alternative to base Llama 3.3
Magistral 32B
research / non-commercial reasoning at 32B scale
Qwen 3 Coder 32B
coding-specialized agent workloads
DeepSeek R1 Distill Qwen 3 32B
workstation reasoning with Qwen 3 base improvements
EXAONE 3.5 32B
Korean / Japanese / CJK workloads
EXAONE 3.5 8B
consumer-tier Korean workloads
InternLM 3 8B
Chinese-language consumer workloads
Dolphin 3 Llama 3.3 70B
datacenter creative / less-restricted generation
Devstral Small 2 24B
Apache 2.0 coding alternative to Qwen 2.5 Coder
Yi Coder 9B
8GB-VRAM coding
Qwen 3 7B
consumer-tier reasoning on 8GB+ GPUs
EVA Llama 3.3 70B
datacenter-tier creative / narrative generation
MiniCPM-V 3 8B
consumer multimodal document Q&A
Qwen 3 Embedding 8B
permissively-licensed embeddings at 8B
MedGemma 27B
medical-domain fine-tune of Gemma 3 27B
Phi-4 Reasoning 14B
consumer-tier reasoning via Phi-4 lineage
Qwen 3 30B-A3B
workstation MoE — 3B active, 30B total
Qwen 3 8B
consumer-tier reasoning toggle
Granite 3 MoE (3B active)
consumer-tier enterprise MoE
Llama 3.3 8B Instruct
consumer-tier chat — drop-in 3.1 8B replacement
Llama 3.1 Nemotron Nano 8B
consumer-tier Nemotron-Llama
DeepSeek R1 Distill Mistral 24B
consumer-tier reasoning with Mistral instruction lineage
Gemma 3 12B
consumer-tier multilingual chat with vision support in 'it' variant
Going deeper
- Ecosystem maps — structured-landscape views (memory frameworks, inference runtimes, MCP, coding agents).
- Execution stacks — recipes that combine models with runtimes + hardware.
- Frontier index — broader ecosystem-momentum view across coding agents, inference runtimes, memory systems, MCP.
- Benchmarks — measured tokens-per-second + topology fields across hardware/model/runtime triples.