RUNLOCALAIv38
->Will it run?Best GPUCompareTroubleshootStartLearnPulseModelsHardwareToolsBench
Run check
RUNLOCALAI

Independently operated catalog for local-AI hardware and software. Hand-written verdicts. Source-cited claims. Reproducible commands when we have them.

OP·Eruo Fredoline
DIR
  • Models
  • Hardware
  • Tools
  • Benchmarks
TOOLS
  • Will it run?
  • Compare hardware
  • Cost vs cloud
  • Choose my GPU
  • Prompting kits
  • Quick answers
REF
  • All buyer guides
  • Learn local AI
  • Methodology
  • Glossary
  • Errors KB
  • Trust
  • Suggest a feature
EDITOR
  • About
  • Author
  • How we make money
  • Editorial policy
  • Contact
LEGAL
  • Privacy
  • Terms
  • Sitemap
MAIL · MONTHLY DIGEST
Get monthly local AI changes
Monthly recap. No spam.
DISCLOSURE

Some links on this site are affiliate links (Amazon Associates and other first-class retailers). When you buy through them, we earn a small commission at no extra cost to you. Affiliate links do not influence our verdicts — there are cards we rate highly that we don't have affiliate relationships with, and cards that sell well that we refuse to recommend. Read more →

© 2026 runlocalai.coIndependently operated
RUNLOCALAI · v38
  1. >
  2. Home
  3. /Frontier
  4. /Models
Frontier zone · Model releases

The frontier of open-weight model releases

Open-weight model releases tracked by RunLocalAI — recent additions, rising families, distill chains, multimodal and reasoning waves. Each card links into the catalog with authority badges (L1.25 enriched · benchmark-backed · verdict) so you can scan editorial coverage at a glance.

By Eruo Fredoline · Refreshed continuously from catalog seed
Filter
Family
AnyQwenLlamaDeepSeekMistralGemmaPhiGLMOLMo
Deployment
AnyEdgeConsumerWorkstationDatacenterFrontier
Modality
AnyMultimodalText-only
Coverage
AnyL1.25 enrichedNeeds L1.25Needs benchmark

Filtered results (48)

Models matching your filters. Clear filters by clicking “Any” on each row above, or remove individual filters via the URL.

Hunyuan 3.0 (Hy3)

Tencent · 2026-07-06
295B/21B-Afrontier

Apache-licensed agent/tool-calling backbone for teams with 8x datacenter GPUs

Verdict

LongCat-2.0

Meituan · 2026-07-05
1600B/48B-Afrontier

MIT-licensed near-frontier agentic coding for multi-node datacenter deployments

Verdict

Laguna XS 2.1

Poolside · 2026-07-02
33Bworkstation

agentic coding on 24GB+ GPUs

Benchmark

Ornith 1.0 35B

DeepReinforce · 2026-06-25
35Bworkstation

agentic coding on 24GB+ GPUs

Benchmark

Ornith 1.0 9B

DeepReinforce · 2026-06-25
9Bconsumer

agentic coding on 8-12GB GPUs

LFM2.5-230M

Liquid AI · 2026-06-25
0.23B/0.23B-Aedge

On-device tool calling and data extraction on phones and Raspberry Pi-class hardware

Verdict

GLM-5.2

Zhipu AI (Z.ai) · 2026-06-16
753Bfrontier

Open-weight 1M-context MoE for long-horizon agentic coding

Verdict

Kimi K2.7-Code

Moonshot AI · 2026-06-12
1000Bfrontier

Open-weight 1T MoE for long-horizon software engineering

Verdict

VibeThinker-3B

WeiboAI (Sina Weibo) · 2026-06-12
3Bconsumer

Compact MIT reasoning model that runs on a single consumer GPU (~6.7GB)

Verdict

North Mini Code 1.0

Cohere · 2026-06-09
30Bworkstation

agentic coding on 24GB+ GPUs

Benchmark

Nemotron 3 Ultra (550B-A55B)

NVIDIA · 2026-06-04
550Bfrontier

Open-weight frontier-scale long-context reasoning for complex agentic workflows

Verdict

Ring-2.6-1T

InclusionAI / Ant Group · 2026-05-14
1000B/32B-Afrontier

frontier reasoning at MoE serving cost

Verdict

Qwen 3.6 35B-A3B (MTP)

Alibaba / Qwen team · 2026-05-11
35B/3B-Aworkstation

high-throughput MoE inference at workstation tier

Verdict

Qwen 3.6 27B (MTP)

Alibaba / Qwen team · 2026-05-11
27Bworkstation

dense workstation model with throughput-acceleration

Verdict

Qwen 3.5 235B-A17B (MoE)

Alibaba · 2026-05-01
397B/17B-Afrontier

frontier-tier reasoning + multilingual serving on multi-machine clusters

L1.25 enrichedVerdict

Mistral Medium 3.5 (675B MoE)

Mistral AI · 2026-04-29
675B/41B-Afrontier

frontier MoE — Mistral's response to the open MoE wave

Verdict

Granite 4.1 30B Instruct

IBM · 2026-04-29
28.87Bworkstation

Workstation-class agent and RAG backends where tool-call accuracy matters more than throughput

Verdict

Mistral Medium 3 24B (dense)

Mistral AI · 2026-04-29
24Bconsumer

research / non-commercial workstation deployments

Verdict

Granite 4.1 8B Instruct

IBM · 2026-04-29
8.79Bconsumer

Local agents and RAG on a single consumer GPU, with best-in-class tool calling per GB under Apache 2.0

Verdict

Granite 4.1 3B Instruct

IBM · 2026-04-29
3.4Bedge

Edge-deployed function calling and RAG extraction where every GB of memory counts

Verdict

DeepSeek V4 Pro (1.6T MoE)

DeepSeek · 2026-04-24
1600B/49B-Afrontier

frontier-tier coding + reasoning serving — currently the open-weight ceiling

L1.25 enrichedVerdict

DeepSeek V4 Flash (284B MoE)

DeepSeek · 2026-04-24
284B/13B-Adatacenter

datacenter MoE — V4 efficiency variant

Verdict

Qwen3.6 27B

Alibaba · 2026-04-22
27Bworkstation

best all-round pick on 24GB+ GPUs

Benchmark

OLMo 2 32B

AI2 (Allen AI) · 2026-04-12
32Bworkstation

fully-open AI2 OLMo 2 — research provenance flagship

Verdict

Phi-4 Reasoning Mini 4B

Microsoft · 2026-04-08
3.8Bedge

edge-tier reasoning

Verdict

Llama 4 Scout

Meta · 2026-04-05
109Bdatacenter

production multimodal serving — image + text at workstation-cluster scale

L1.25 enrichedVerdict

DeepSeek V4

DeepSeek AI · 2026-03-15
745B/38B-Afrontier

frontier-tier reasoning on multi-machine clusters

Verdict

Granite 3.3 8B

IBM · 2026-03-12
8Bconsumer

enterprise tool-calling on IBM stacks

Verdict

Kimi K2.6

Moonshot AI · 2026-03-10
1000Bfrontier

Moonshot frontier MoE — long-context specialist

Verdict

Mistral Small 3.2 24B

Mistral AI · 2026-03-08
24Bconsumer

consumer-tier multilingual instruction-following

Verdict

Phi-4 Mini 4B

Microsoft · 2026-02-25
3.8Bedge

edge / embedded reasoning

Verdict

GLM-5 Pro

Zhipu AI · 2026-02-18
144B/16B-Adatacenter

Chinese-language enterprise serving

Verdict

Nemotron 3 Super (120B-A12B)

NVIDIA · 2026-02-15
120Bdatacenter

NVIDIA-tuned datacenter-tier reasoning

Verdict

Llama 4 405B

Meta · 2026-02-10
405Bfrontier

frontier-tier serving on cluster hardware

Verdict

Llama 4 70B

Meta · 2026-02-10
70Bdatacenter

production self-hosted serving on 2x A100 / H100

Verdict

DeepSeek Coder V3

DeepSeek AI · 2026-02-08
33Bworkstation

workstation coding alternative to Qwen 2.5 Coder

Verdict

GLM-5

Zhipu AI (Z.AI) · 2026-02-05
200Bfrontier

Zhipu GLM-5 frontier MoE

Verdict

Nemotron 3 Super 49B

NVIDIA · 2026-01-22
49Bworkstation

32GB-VRAM enterprise deployments

Verdict

Nemotron 3 Nano 9B

NVIDIA · 2026-01-22
9Bconsumer

NVIDIA-stack tool-calling agents

Verdict

GLM-4.7-Flash

Zhipu AI (Z.ai) · 2026-01-19
31Bworkstation

fast local coding and agents on 24GB+ GPUs

Benchmark

Nemotron 3 Nano (30B-A3B)

NVIDIA · 2026-01-15
30Bconsumer

NVIDIA-tuned consumer-tier general

Verdict

DeepSeek V3 Lite (16B MoE)

DeepSeek AI · 2026-01-10
16B/2.4B-Aconsumer

consumer-tier MoE inference

Verdict

Hermes 4 Llama 3.3 70B

Nous Research · 2025-12-22
70Bdatacenter

datacenter-tier instruction-tuned alternative to base Llama 3.3

Verdict

Magistral 32B

Mistral AI · 2025-12-15
32Bworkstation

research / non-commercial reasoning at 32B scale

Verdict

Kimi K1.5

Moonshot AI · 2025-12-01
200Bdatacenter

deep math + reasoning research

Verdict

Qwen 3 Coder 32B

Alibaba · 2025-11-20
32Bworkstation

coding-specialized agent workloads

Verdict

DeepSeek R1 Distill Qwen 3 32B

DeepSeek AI · 2025-11-15
32Bworkstation

workstation reasoning with Qwen 3 base improvements

Verdict

EXAONE 3.5 32B

LG AI Research · 2025-11-10
32Bworkstation

Korean / Japanese / CJK workloads

Verdict

Going deeper

  • Ecosystem maps — structured-landscape views (memory frameworks, inference runtimes, MCP, coding agents).
  • Execution stacks — recipes that combine models with runtimes + hardware.
  • Frontier index — broader ecosystem-momentum view across coding agents, inference runtimes, memory systems, MCP.
  • Benchmarks — measured tokens-per-second + topology fields across hardware/model/runtime triples.