RUNLOCALAIv38
->Will it run?Best GPUCompareTroubleshootStartLearnPulseModelsHardwareToolsBench
Run check
RUNLOCALAI

Independently operated catalog for local-AI hardware and software. Hand-written verdicts. Source-cited claims. Reproducible commands when we have them.

OP·Eruo Fredoline
DIR
  • Models
  • Hardware
  • Tools
  • Benchmarks
TOOLS
  • Will it run?
  • Compare hardware
  • Cost vs cloud
  • Choose my GPU
  • Prompting kits
  • Quick answers
REF
  • All buyer guides
  • Learn local AI
  • Methodology
  • Glossary
  • Errors KB
  • Trust
  • Suggest a feature
EDITOR
  • About
  • Author
  • How we make money
  • Editorial policy
  • Contact
LEGAL
  • Privacy
  • Terms
  • Sitemap
MAIL · MONTHLY DIGEST
Get monthly local AI changes
Monthly recap. No spam.
DISCLOSURE

Some links on this site are affiliate links (Amazon Associates and other first-class retailers). When you buy through them, we earn a small commission at no extra cost to you. Affiliate links do not influence our verdicts — there are cards we rate highly that we don't have affiliate relationships with, and cards that sell well that we refuse to recommend. Read more →

© 2026 runlocalai.coIndependently operated
RUNLOCALAI · v38
  1. >
  2. Home
  3. /Frontier
  4. /Models
Frontier zone · Model releases

The frontier of open-weight model releases

Open-weight model releases tracked by RunLocalAI — recent additions, rising families, distill chains, multimodal and reasoning waves. Each card links into the catalog with authority badges (L1.25 enriched · benchmark-backed · verdict) so you can scan editorial coverage at a glance.

By Eruo Fredoline · Refreshed continuously from catalog seed
Filter
Family
AnyQwenLlamaDeepSeekMistralGemmaPhiGLMOLMo
Deployment
AnyEdgeConsumerWorkstationDatacenterFrontier
Modality
AnyMultimodalText-only
Coverage
AnyL1.25 enrichedNeeds L1.25Needs benchmark

Filtered results (46)

Models matching your filters. Clear filters by clicking “Any” on each row above, or remove individual filters via the URL.

VibeThinker-3B

WeiboAI (Sina Weibo) · 2026-06-12
3Bconsumer

Compact MIT reasoning model that runs on a single consumer GPU (~6.7GB)

Verdict

Qwen 3.6 35B-A3B (MTP)

Alibaba / Qwen team · 2026-05-11
35B/3B-Aworkstation

high-throughput MoE inference at workstation tier

Verdict

Qwen 3.6 27B (MTP)

Alibaba / Qwen team · 2026-05-11
27Bworkstation

dense workstation model with throughput-acceleration

Verdict

Qwen 3.5 235B-A17B (MoE)

Alibaba · 2026-05-01
397B/17B-Afrontier

frontier-tier reasoning + multilingual serving on multi-machine clusters

L1.25 enrichedVerdict

Qwen3.6 27B

Alibaba · 2026-04-22
27Bworkstation

best all-round pick on 24GB+ GPUs

Benchmark

Qwen3.5 9B

Alibaba · 2026-03-02
9Bconsumer

multilingual chat and coding on 8-12GB GPUs

Multimodal

Qwen 3 Coder 32B

Alibaba · 2025-11-20
32Bworkstation

coding-specialized agent workloads

Verdict

Qwen 3 7B

Alibaba · 2025-09-15
7Bconsumer

consumer-tier reasoning on 8GB+ GPUs

Verdict

Qwen 3 Embedding 8B

Alibaba · 2025-06-05
8Bconsumer

permissively-licensed embeddings at 8B

Verdict

Qwen 3 235B-A22B

Alibaba · 2025-04-29
235Bfrontier

Qwen 3 MoE flagship — pre-3.5 baseline

Verdict

Qwen 3 32B

Alibaba · 2025-04-29
32Bworkstation

general-purpose reasoning + chat with toggle-style reasoning emission

L1.25 enrichedVerdict

Qwen 3 30B-A3B

Alibaba · 2025-04-29
30Bworkstation

workstation MoE — 3B active, 30B total

Verdict

Qwen 3 14B

Alibaba · 2025-04-29
14Bconsumer

16GB-VRAM reasoning workloads with thinking-mode toggle

L1.25 enrichedBenchmarkVerdict

Qwen 3 8B

Alibaba · 2025-04-29
8Bconsumer

consumer-tier reasoning toggle

Verdict

Qwen 3 4B

Alibaba · 2025-04-29
4Bedge

edge-tier Qwen 3 — Apple Silicon laptop friendly

BenchmarkVerdict

Qwen 2.5-VL 72B

Alibaba · 2025-03-10
72Bdatacenter

frontier-tier multimodal serving

VerdictMultimodal

Qwen 2.5-VL 7B

Alibaba · 2025-03-10
7Bconsumer

consumer-tier OCR + image Q&A

VerdictMultimodal

Qwen 2.5-VL 3B

Alibaba · 2025-01-26
3Bedge

edge-tier multimodal

VerdictMultimodal

QwQ 32B Preview

Alibaba · 2024-11-27
32Bworkstation

workstation-tier reasoning — Qwen team alternative to R1

Verdict

Qwen 2.5 Coder 32B Instruct

Alibaba · 2024-11-12
32Bworkstation

single-user autonomous coding agents on RTX 4090 / 5090 / dual-A100 hardware

L1.25 enrichedVerdict

Qwen 2.5 Coder 14B Instruct

Alibaba · 2024-11-12
14Bconsumer

16GB-VRAM coding

BenchmarkVerdict

Qwen 2.5 Coder 7B Instruct

Alibaba · 2024-11-12
7Bconsumer

consumer-tier coding at 8GB VRAM

Verdict

Qwen 2.5 Coder 3B

Alibaba · 2024-11-12
3Bedge

Apple Silicon laptop coding autocomplete

Verdict

Qwen 2.5 Coder 1.5B

Alibaba · 2024-11-12
1.5Bedge

IDE autocomplete on integrated GPUs

Verdict

Qwen 2.5 Math 72B

Alibaba · 2024-09-19
72Bdatacenter

datacenter-tier math specialist

Verdict

Qwen 2.5 72B Instruct

Alibaba · 2024-09-19
72Bdatacenter

production multilingual at 70B-class

Verdict

Qwen 2.5 32B Instruct

Alibaba · 2024-09-19
32Bworkstation

workstation-tier multilingual general chat

Verdict

Qwen 2.5 14B Instruct

Alibaba · 2024-09-19
14Bconsumer

16GB-VRAM general chat with multilingual depth

Verdict

Qwen 2.5 7B Instruct

Alibaba · 2024-09-19
7Bconsumer

consumer-tier multilingual chat

BenchmarkVerdict

Qwen 2.5 Math 7B

Alibaba · 2024-09-19
7Bconsumer

consumer-tier math problem solving

Verdict

Qwen 2.5 3B Instruct

Alibaba · 2024-09-19
3Bedge

edge-tier Qwen 2.5 chat

Verdict

Qwen 2.5 1.5B Instruct

Alibaba · 2024-09-19
1.5Bedge

edge-tier Apache 2.0 chat

Verdict

Qwen 2.5 0.5B Instruct

Alibaba · 2024-09-19
0.5Bedge

phone-tier Qwen baseline

Verdict

Qwen 2-VL 7B

Alibaba · 2024-08-29
7Bconsumer

consumer-tier multimodal — pre-2.5-VL baseline

VerdictMultimodal

CodeQwen 1.5 7B

Alibaba · 2024-04-16
7Bconsumer

historical reference — Qwen 2.5 Coder 7B is the modern pick

Verdict

Qwen3.5 35B-A3B

Alibaba
35Bworkstation

fast MoE decode on 32GB GPUs

Benchmark

Qwen3.6 35B-A3B

Alibaba
35Bworkstation

fast MoE decode on 32GB GPUs

Benchmark

Qwen3 Swallow 32B RL v0.2

tokyotech-llm
32B

Japanese-English reasoning tasks: math, coding, structured analysis

Verdict

Qwen3 Coder 30B-A3B

Alibaba
30Bworkstation

dedicated coding assistant on 24GB+ GPUs

Benchmark

Qwen3.5 27B

Alibaba
27Bworkstation

general reasoning and coding on 24GB+ GPUs

Benchmark

Qwen3.5 9B Thai Law Base

Phonsiri
8.95B

Foundation for fine-tuning Thai legal NLP tools

Verdict

Qwen2-VL 2B Instruct

Alibaba
2B

Lightweight document and chart understanding on a consumer GPU

L1.25 enrichedVerdict

Qwen 3.5 2B Turkish SFT

Tuguberk
2B
L1.25 enriched

Qwen 3 1.7B

Alibaba
1.7B

Edge laptop assistant with reasoning that fits in 2GB VRAM

L1.25 enrichedVerdict

Qwen3 0.6B Hindi Instruct v1 GGUF

pankajpandey-dev
0.6B

Simple Hindi instruction following on CPU-only devices

L1.25 enrichedVerdict

Qwen 3 0.6B

Alibaba
0.6B

Sub-1B on-device chat and tool-calling agent on phones

L1.25 enrichedVerdict

Going deeper

  • Ecosystem maps — structured-landscape views (memory frameworks, inference runtimes, MCP, coding agents).
  • Execution stacks — recipes that combine models with runtimes + hardware.
  • Frontier index — broader ecosystem-momentum view across coding agents, inference runtimes, memory systems, MCP.
  • Benchmarks — measured tokens-per-second + topology fields across hardware/model/runtime triples.