The frontier of open-weight model releases
Open-weight model releases tracked by RunLocalAI — recent additions, rising families, distill chains, multimodal and reasoning waves. Each card links into the catalog with authority badges (L1.25 enriched · benchmark-backed · verdict) so you can scan editorial coverage at a glance.
Filtered results (41)
Models matching your filters. Clear filters by clicking “Any” on each row above, or remove individual filters via the URL.
Muse Glimmer 30B
Laguna XS 2.1
agentic coding on 24GB+ GPUs
Ornith 1.0 35B
agentic coding on 24GB+ GPUs
North Mini Code 1.0
agentic coding on 24GB+ GPUs
Qwen 3.6 35B-A3B (MTP)
high-throughput MoE inference at workstation tier
Qwen 3.6 27B (MTP)
dense workstation model with throughput-acceleration
Granite 4.1 30B Instruct
Workstation-class agent and RAG backends where tool-call accuracy matters more than throughput
Qwen3.6 27B
best all-round pick on 24GB+ GPUs
OLMo 2 32B
fully-open AI2 OLMo 2 — research provenance flagship
Gemma 4 31B Dense
workstation-tier multilingual chat with permissive license
Gemma 4 26B MoE
Gemma 4 MoE — workstation efficiency variant
DeepSeek Coder V3
workstation coding alternative to Qwen 2.5 Coder
Nemotron 3 Super 49B
32GB-VRAM enterprise deployments
GLM-4.7-Flash
fast local coding and agents on 24GB+ GPUs
Magistral 32B
research / non-commercial reasoning at 32B scale
Qwen 3 Coder 32B
coding-specialized agent workloads
DeepSeek R1 Distill Qwen 3 32B
workstation reasoning with Qwen 3 base improvements
EXAONE 3.5 32B
Korean / Japanese / CJK workloads
MedGemma 27B
medical-domain fine-tune of Gemma 3 27B
Qwen 3 32B
general-purpose reasoning + chat with toggle-style reasoning emission
Qwen 3 30B-A3B
workstation MoE — 3B active, 30B total
Gemma 3 27B
Google's open-weight workstation-tier multilingual flagship — pre-Gemma-4 baseline
DeepSeek R1 Distill Qwen 32B
single-machine reasoning — the canonical local R1 deployment
QwQ 32B Preview
workstation-tier reasoning — Qwen team alternative to R1
Qwen 2.5 Coder 32B Instruct
single-user autonomous coding agents on RTX 4090 / 5090 / dual-A100 hardware
Aya Expanse 32B
research / non-commercial multilingual workflows
Qwen 2.5 32B Instruct
workstation-tier multilingual general chat
Jamba 1.5 Mini
workstation long-context with hybrid SSM throughput
Codestral 22B
workstation coding at 22B class
Aya 23 35B
multilingual research at workstation tier
Yi 1.5 34B
workstation-tier multilingual
Command R 35B
workstation-tier RAG-tuned
Mixtral 8x7B Instruct
workstation MoE — 13B active, 47B total
Phind CodeLlama 34B v2
historical reference for Llama 2 coder lineage
Qwen3.6 35B-A3B
fast MoE decode on 32GB GPUs
Qwen3.5 35B-A3B
fast MoE decode on 32GB GPUs
Nemotron 3 Nano Omni 33B
multimodal video/audio/image/text tasks on 32GB+ GPUs
Qwen3 Coder 30B-A3B
dedicated coding assistant on 24GB+ GPUs
Granite 4.1 30B
enterprise-provenance general model on 24GB+ GPUs
Qwen3.5 27B
general reasoning and coding on 24GB+ GPUs
Gemma 4 26B-A4B
fast MoE decode on 24GB+ GPUs
Going deeper
- Ecosystem maps — structured-landscape views (memory frameworks, inference runtimes, MCP, coding agents).
- Execution stacks — recipes that combine models with runtimes + hardware.
- Frontier index — broader ecosystem-momentum view across coding agents, inference runtimes, memory systems, MCP.
- Benchmarks — measured tokens-per-second + topology fields across hardware/model/runtime triples.