# RunLocalAI — full content index > Complete machine-readable index of RunLocalAI: every indexable model page, > hardware verdict, tool review, guide, course, glossary term, and error fix. > The curated summary lives at https://www.runlocalai.co/llms.txt. Editorial standards, benchmark sourcing, and methodology: - [Editorial policy](https://www.runlocalai.co/editorial-policy) - [Methodology](https://www.runlocalai.co/methodology) - [Will-It-Run framework](https://www.runlocalai.co/resources/will-it-run-framework) Machine-readable endpoints: - [Sitemap](https://www.runlocalai.co/sitemap.xml) - [RSS feed — recently updated pages](https://www.runlocalai.co/feed.xml) - [Datasets](https://www.runlocalai.co/datasets) ## Models (195) - [Aya 23 35B](https://www.runlocalai.co/models/aya-23-35b): Aya 23 at 35B. Built on Cohere's Command-R lineage. Non-commercial. - [Aya 23 8B](https://www.runlocalai.co/models/aya-23-8b): Cohere's multilingual research model covering 23 languages. CC-BY-NC — research only. - [Aya Expanse 32B](https://www.runlocalai.co/models/aya-expanse-32b): Cohere's multilingual Aya at 32B. Covers 23 languages; strongest open-weight multilingual model in late 2024 — Apache-2.0 alternative is Qwen 2.5 32B but Aya has deeper coverage o… - [Baichuan 4 13B](https://www.runlocalai.co/models/baichuan-4-13b): Baichuan AI's 13B. Chinese-language ecosystem alternative to Qwen / GLM. Restricted commercial license. - [BGE M3](https://www.runlocalai.co/models/bge-m3): BAAI's multilingual embedding flagship. Dense + sparse + ColBERT-style multi-vector. The de-facto open multilingual embedding pick. - [CodeGemma 7B](https://www.runlocalai.co/models/codegemma-7b): Coding-specialist Gemma. Decent FIM completion. Now mostly historical with Qwen 2.5 Coder dominating. - [CodeQwen 1.5 7B](https://www.runlocalai.co/models/codeqwen-1.5-7b): CodeQwen 1.5 — Qwen Coder predecessor. Superseded by Qwen 2.5 Coder for new deployments. - [Codestral 22B](https://www.runlocalai.co/models/codestral-22b): Mistral's coding-specialist. Strong fill-in-the-middle for IDE autocompletion. Personal/research use only. - [Codestral Mamba 7B](https://www.runlocalai.co/models/codestral-mamba-7b): Mistral's Mamba (state-space) architecture coding model. Linear inference cost — the architectural alternative to attention-based coding models. Apache 2.0. - [Command R 35B](https://www.runlocalai.co/models/command-r-35b): Cohere's mid-tier — RAG and tool use. Non-commercial license. - [Command R+ (Aug 2024)](https://www.runlocalai.co/models/command-r-plus-08-2024): Cohere's August 2024 Command R+ refresh. RAG-optimized; non-commercial license. Strong tool-calling and citation discipline. - [Command R+ 104B](https://www.runlocalai.co/models/command-r-plus-104b): Cohere's flagship — RAG-tuned, multilingual. Open weights but non-commercial. - [DBRX Base](https://www.runlocalai.co/models/dbrx-base): DBRX base (non-instruct). 132B total / 36B active fine-grained MoE. - [DBRX Instruct](https://www.runlocalai.co/models/dbrx-instruct): Databricks' MoE. 132B total / 36B active. Designed for Mosaic ML pipelines; strong tool-calling discipline. Multi-GPU only. - [DeepSeek Coder V2 236B](https://www.runlocalai.co/models/deepseek-coder-v2-236b): Full DeepSeek Coder V2. 236B total / 21B active MoE coder. - [DeepSeek Coder V2 Lite (16B)](https://www.runlocalai.co/models/deepseek-coder-v2-lite): MoE coding specialist — 16B total / 2.4B active. Fast on 12GB cards. - [DeepSeek Coder V3](https://www.runlocalai.co/models/deepseek-coder-v3): DeepSeek's coder line successor. Dense 33B; competitive with Qwen 2.5 Coder 32B on SWE-Bench. - [DeepSeek MoE 16B Base](https://www.runlocalai.co/models/deepseek-moe-16b-base): DeepSeek's first MoE — 16B / 2.4B active. Older model retained for ecosystem-context value as the base of the V2/V3 lineage. - [DeepSeek R1 (671B reasoning)](https://www.runlocalai.co/models/deepseek-r1): Open reasoning model that closed the gap with frontier proprietary reasoners. Visible chain-of-thought, MIT license, and a family of distilled smaller variants. - [DeepSeek R1 Distill Llama 70B](https://www.runlocalai.co/models/deepseek-r1-distill-llama-70b): Reasoning distillation onto Llama 3.3 70B. Best-in-class open-weight reasoner you can actually fit on a workstation. - [DeepSeek R1 Distill Llama 8B](https://www.runlocalai.co/models/deepseek-r1-distill-llama-8b): R1 reasoning distilled into a Llama 3 8B base. Smaller R1 distill; useful when 32B is too heavy. Reasoning quality is meaningfully below the 32B distill but still beats non-reason… - [DeepSeek R1 Distill Mistral 24B](https://www.runlocalai.co/models/deepseek-r1-distill-mistral-24b): Community R1 distill onto a Mistral Small 3 base. Apache 2.0; combines R1 reasoning with Mistral instruction polish. - [DeepSeek R1 Distill Qwen 1.5B](https://www.runlocalai.co/models/deepseek-r1-distill-qwen-1.5b): Smallest R1 distill. Surprisingly capable reasoning at 1.5B for its size class; right pick when you need reasoning AND edge deployment. - [DeepSeek R1 Distill Qwen 14B](https://www.runlocalai.co/models/deepseek-r1-distill-qwen-14b): 14B reasoning distill. Fits on 12GB cards. - [DeepSeek R1 Distill Qwen 3 32B](https://www.runlocalai.co/models/deepseek-r1-distill-qwen-3-32b): Newer R1 distill on a Qwen 3 base. Combines R1 reasoning with Qwen 3's reasoning-toggle architecture. Apache 2.0. - [DeepSeek R1 Distill Qwen 32B](https://www.runlocalai.co/models/deepseek-r1-distill-qwen-32b): 32B distill — fits on a single 24GB card with reasoning capability. Best price-per-thinking-token combo for prosumers. - [DeepSeek R1 Distill Qwen 7B](https://www.runlocalai.co/models/deepseek-r1-distill-qwen-7b): Smallest practical R1 distill. Reasoning on a 6GB GPU. - [DeepSeek V2.5 236B](https://www.runlocalai.co/models/deepseek-v2.5-236b): DeepSeek V2.5 — merged V2 chat + Coder. Pre-V3 baseline; 21B active MoE. - [DeepSeek V3 (671B MoE)](https://www.runlocalai.co/models/deepseek-v3): DeepSeek's flagship MoE — 671B total / 37B active. Server-tier, but the smaller R1 distills make this lineage approachable. - [DeepSeek V3 Lite (16B MoE)](https://www.runlocalai.co/models/deepseek-v3-lite): Distillation of DeepSeek V3 to a smaller MoE. 16B total / 2.4B active. Captures most of V3's reasoning at consumer-card-friendly memory. - [DeepSeek V4](https://www.runlocalai.co/models/deepseek-v4): DeepSeek's spring 2026 frontier MoE. 745B total / 38B active. The current open-weight benchmark leader on coding + math; closes the gap with closed-source flagships on reasoning. - [DeepSeek V4 Flash (284B MoE)](https://www.runlocalai.co/models/deepseek-v4-flash): The cost-efficient sibling of V4-Pro. 284B total / 13B active MoE, same hybrid CSA+HCA attention, same 1M context. The MoE active-param ratio (4.5%) makes it surprisingly fast for… - [DeepSeek V4 Pro (1.6T MoE)](https://www.runlocalai.co/models/deepseek-v4-pro): DeepSeek's April 2026 frontier flagship. 1.6T total / 49B active MoE with hybrid Compressed Sparse Attention + Heavily Compressed Attention. 1M context window. Closes most of the… - [Devstral Small 2 24B](https://www.runlocalai.co/models/devstral-small-2-24b): Mistral's coding-specialized Mistral Small 2 successor. Apache 2.0 — the rare commercial-OK Mistral coder. - [Dolphin 3.0 Llama 3.2 3B](https://www.runlocalai.co/models/dolphin-3-llama-3.2-3b): Eric Hartford's Dolphin fine-tune at 3B. Less-censored than the base Llama; popular for unconstrained-generation use cases. - [Dolphin 3 Llama 3.3 70B](https://www.runlocalai.co/models/dolphin-3-llama-3.3-70b): Eric Hartford's Dolphin 3 at 70B Llama 3.3 base. Less-restricted alternative for creative / unconstrained workflows. - [Dolphin 3.0 Mistral 24B](https://www.runlocalai.co/models/dolphin-3.0-mistral-24b): Eric Hartford's Dolphin fine-tune of Mistral Small 3 — uncensored, function-calling, agent-friendly. - [EVA Llama 3.3 70B](https://www.runlocalai.co/models/eva-llama-3.3-70b): EVA community's storytelling-focused fine-tune of Llama 3.3 70B. Popular in the creative-writing / roleplay community. - [EXAONE 3.5 2.4B](https://www.runlocalai.co/models/exaone-3.5-2.4b): LG AI's edge-tier EXAONE. Strong Korean / English. Research-only license. - [EXAONE 3.5 32B](https://www.runlocalai.co/models/exaone-3.5-32b): LG AI Research's flagship Korean-ecosystem model. Strong on Korean/Japanese language tasks; competitive on English. License blocks commercial use without LG agreement. - [EXAONE 3.5 8B](https://www.runlocalai.co/models/exaone-3.5-8b): Smaller EXAONE for consumer-tier Korean / CJK workloads. - [Falcon 3 10B](https://www.runlocalai.co/models/falcon-3-10b): TII's Falcon 3 at the 10B tier. Strong on Arabic-language tasks; competitive on English. - [Falcon 3 7B Instruct](https://www.runlocalai.co/models/falcon-3-7b): Falcon 3 mid-size from TII. Permissive Falcon license; multilingual focus. - [Falcon Mamba 7B](https://www.runlocalai.co/models/falcon-mamba-7b): TII's Mamba (state-space) architecture model. Linear inference cost; the architectural alternative to attention-based models. - [Gemma 2 9B Instruct](https://www.runlocalai.co/models/gemma-2-9b-it): Mid-size Gemma 2. Strong chat quality with a different training mix from Llama family. - [Gemma 3 12B](https://www.runlocalai.co/models/gemma-3-12b): 12B Gemma 3. Fits on 12GB consumer cards. Multimodal. - [Gemma 3 1B](https://www.runlocalai.co/models/gemma-3-1b): Smallest text-only Gemma 3 for phones and IoT. - [Gemma 3 27B](https://www.runlocalai.co/models/gemma-3-27b): Pre-Gemma-4 flagship. Multimodal (4B+ variants), 128K context, 140 languages. Strong daily driver on 24GB cards. - [Gemma 3 4B](https://www.runlocalai.co/models/gemma-3-4b): 4B Gemma 3 for edge. Multimodal. - [Gemma 4 26B MoE](https://www.runlocalai.co/models/gemma-4-26b-moe): MoE variant of Gemma 4. Faster per-token than the 31B dense at similar quality on most tasks. - [Gemma 4 31B Dense](https://www.runlocalai.co/models/gemma-4-31b): Google's flagship dense Gemma 4. Beats some 400B-class proprietary models on benchmarks. Targets the 24GB single-GPU sweet spot. - [Gemma 4 E2B (Effective 2B)](https://www.runlocalai.co/models/gemma-4-e2b): Smallest Gemma 4. Designed for phones and Raspberry-Pi-class hardware. - [Gemma 4 E4B (Effective 4B)](https://www.runlocalai.co/models/gemma-4-e4b): Edge-class Gemma 4. The 'Effective 4B' branding signals it punches above its parameter count via training-data quality. - [GLM-4 9B](https://www.runlocalai.co/models/glm-4-9b): Zhipu's GLM-4 at 9B. Strong on Chinese-language tasks; tool-calling format slightly different from OpenAI convention. - [GLM-4V 9B](https://www.runlocalai.co/models/glm-4v-9b): GLM-4 with vision encoder. Strong on Chinese document Q&A; restricted commercial license. - [GLM-5](https://www.runlocalai.co/models/glm-5): Zhipu's GLM-5 currently leads the Open LLM Leaderboard 2026. Strong reasoning and bilingual EN/ZH capability. - [GLM-5 Pro](https://www.runlocalai.co/models/glm-5-pro): Zhipu's GLM-5 flagship. 144B total / 16B active MoE. Strong on Chinese-language tasks; competitive on English at the workstation-cluster tier. - [GLM-5.2](https://www.runlocalai.co/models/glm-5.2): GLM-5.2 is Zhipu AI's (Z.ai) flagship open-weight LLM, released on Hugging Face as zai-org/GLM-5.2 (2026-06). It is a Mixture-of-Experts decoder-only transformer with a new "Index… - [Granite 3 MoE (3B active)](https://www.runlocalai.co/models/granite-3-moe-3b-active): Granite MoE shape. 16B total / 3B active. Workstation-deployable; the IBM enterprise alternative to Qwen / DeepSeek small MoEs. - [Granite 3.0 2B Instruct](https://www.runlocalai.co/models/granite-3.0-2b): IBM Granite at 2B. Apache 2.0 enterprise-friendly small model with safety tuning. - [Granite 3.0 8B Instruct](https://www.runlocalai.co/models/granite-3.0-8b): Granite 3.0 8B — IBM's enterprise-tier baseline. Apache 2.0. - [Granite 3.2 8B](https://www.runlocalai.co/models/granite-3.2-8b): IBM's enterprise-tuned 8B. Apache 2.0. Strong on enterprise-shaped tool-calling and structured output. Watson + RHEL ecosystem alignment. - [Granite 3.3 8B](https://www.runlocalai.co/models/granite-3.3-8b): IBM Granite 3.3. Iterative refresh of 3.2 — same architecture; improved instruction following and tool-call reliability. Apache 2.0. - [Granite 4.1 30B Instruct](https://www.runlocalai.co/models/granite-4.1-30b-instruct): Granite 4.1 30B Instruct is the largest Granite 4.1 language model: a dense, decoder-only transformer with 28.9B parameters across 64 layers (grouped-query attention, RoPE, SwiGLU… - [Granite 4.1 3B Instruct](https://www.runlocalai.co/models/granite-4.1-3b-instruct): Granite 4.1 3B Instruct is the smallest of IBM's Granite 4.1 language models, released April 29, 2026 under Apache 2.0. It is a dense, decoder-only transformer — 3.4B parameters,… - [Granite 4.1 8B Instruct](https://www.runlocalai.co/models/granite-4.1-8b-instruct): Granite 4.1 8B Instruct is the mid-size model in IBM's Granite 4.1 family: a dense, decoder-only transformer with 8.79B parameters (40 layers, grouped-query attention, RoPE, SwiGL… - [Hermes 3 Llama 3.1 70B](https://www.runlocalai.co/models/hermes-3-llama-3.1-70b): Hermes 3 at 70B. Workstation-tier agent-tuned model. - [Hermes 3 Llama 3.1 8B](https://www.runlocalai.co/models/hermes-3-llama-3.1-8b): NousResearch's Hermes fine-tune of Llama 3.1 8B. Stronger system-prompt adherence, JSON output, role-play, and agent steering than the base Llama. - [Hermes 3 Llama 3.2 3B](https://www.runlocalai.co/models/hermes-3-llama-3.2-3b): Nous Research's Hermes 3 fine-tune of Llama 3.2 3B. Strong general-instruction following at the 3B tier. - [Hermes 4 Llama 3.3 70B](https://www.runlocalai.co/models/hermes-4-llama-3.3-70b): Nous Research's Hermes 4 fine-tune of Llama 3.3 70B. Strong on instruction following and creative tasks; community-favored alternative to base Llama. - [Hunyuan 3.0 (Hy3)](https://www.runlocalai.co/models/hunyuan-3.0): Hy3 is Tencent's flagship open-weights release: a 295B-parameter Mixture-of-Experts text model activating 21B parameters per token, plus a 3.8B-parameter multi-token-prediction (M… - [Hunyuan Large 389B MoE](https://www.runlocalai.co/models/hunyuan-large-moe): Tencent's frontier MoE. 389B total / 52B active. License permits commercial use with restrictions on companies above MAU thresholds. - [InternLM 2.5 7B Chat](https://www.runlocalai.co/models/internlm-2.5-7b-chat): InternLM 2.5 mid-size chat. Apache 2.0; strong on math and Chinese. - [InternLM 3 8B](https://www.runlocalai.co/models/internlm-3-8b): Shanghai AI Lab's open-research line. InternLM 3 at 8B; strong on Chinese-language tasks. - [InternVL 2.5 26B](https://www.runlocalai.co/models/internvl-2.5-26b): InternVL 2.5 mid-tier — Shanghai AI Lab vision-language model with strong document and chart understanding. - [InternVL 2.5 78B](https://www.runlocalai.co/models/internvl-2.5-78b): InternVL 2.5 flagship. Approaches frontier proprietary VLMs on document and OCR tasks. - [Jamba 1.5 Large](https://www.runlocalai.co/models/jamba-1.5-large): Jamba flagship at 398B total / 94B active. Frontier hybrid-architecture model with 256k context. - [Jamba 1.5 Mini](https://www.runlocalai.co/models/jamba-1.5-mini): AI21's hybrid Mamba-Transformer MoE. 256k context with the SSM throughput advantage. - [Janus-Pro 7B](https://www.runlocalai.co/models/janus-pro-7b): DeepSeek's multimodal 7B. Decoupled visual encoding for understanding vs generation — different from typical VLM design. - [Kimi K1.5](https://www.runlocalai.co/models/kimi-k1.5): Moonshot's reasoning model. Reasoning-token emission with very long thinking-block depth — sometimes 5000+ tokens per query. Strong on math; restricted commercial license. - [Kimi K2.6](https://www.runlocalai.co/models/kimi-k2.6): Moonshot's long-context, agent-oriented MoE. Optimized for stability under tool use and multi-step coding/planning workflows. - [Kimi K2.7-Code](https://www.runlocalai.co/models/kimi-k2.7-code): Kimi K2.7-Code is a coding-specialized open-weight Mixture-of-Experts model from Moonshot AI (Hugging Face moonshotai/Kimi-K2.7-Code, 2026-06), built on K2.6 with ~1T total / ~32B… - [LFM2.5-230M](https://www.runlocalai.co/models/lfm2.5-230m): LFM2.5-230M is Liquid AI's smallest LFM2.5 model: a 230M-parameter, text-only hybrid (14 layers — 8 double-gated LIV convolution blocks plus 6 GQA attention blocks) trained on a 1… - [Llama 3.1 70B Instruct](https://www.runlocalai.co/models/llama-3.1-70b-instruct): The 70B sibling of Llama 3.1 8B. Strong generalist reasoning with 128K context, popular base for agentic fine-tunes (Hermes 3, Nemotron). Mostly superseded by Llama 3.3 70B for ne… - [Llama 3.1 8B Instruct](https://www.runlocalai.co/models/llama-3.1-8b-instruct): Meta's small flagship. Strong general reasoning, 128K context, broad multilingual. The default first try for most local-AI use cases on consumer hardware. - [Llama 3.1 Nemotron 70B Instruct](https://www.runlocalai.co/models/llama-3.1-nemotron-70b-instruct): NVIDIA's HelpSteer2-tuned Llama 3.1 70B. Topped Arena Hard at release. The pre-Nemotron-3 NVIDIA reference open weights. - [Llama 3.1 Nemotron Nano 8B](https://www.runlocalai.co/models/llama-3.1-nemotron-nano-8b): Smallest of the Nemotron reasoning trio. NAS-optimized for inference efficiency on RTX hardware. - [Llama 3.1 Nemotron Ultra 253B](https://www.runlocalai.co/models/llama-3.1-nemotron-ultra-253b): NVIDIA's top open reasoning model in the Llama 3.1 lineage. Server-tier; trained for groundbreaking reasoning accuracy on agentic workloads. - [Llama 3.2 11B Vision](https://www.runlocalai.co/models/llama-3.2-11b-vision): Llama 3.2 multimodal at 11B. Consumer-tier multimodal predecessor to Llama 4 Scout. - [Llama 3.2 11B Vision Instruct](https://www.runlocalai.co/models/llama-3.2-11b-vision-instruct): First-party multimodal Llama. Accepts images alongside text for VQA, document understanding, and chart reading. Runs on 12GB+ VRAM. - [Llama 3.2 1B Instruct](https://www.runlocalai.co/models/llama-3.2-1b-instruct): True edge-tier Llama. Runs on a phone or Raspberry Pi. Useful for classification, simple summarization, and on-device agents. - [Llama 3.2 3B Instruct](https://www.runlocalai.co/models/llama-3.2-3b-instruct): Lightweight 3B model for edge and laptop deployment. Check the quantized artifact and context against available memory. The earlier Apple Silicon speed estimate had no linked meas… - [Llama 3.2 90B Vision](https://www.runlocalai.co/models/llama-3.2-90b-vision): Llama 3.2 multimodal at 90B. Datacenter-tier predecessor to Llama 4 Maverick. Strong visual reasoning. - [Llama 3.2 90B Vision Instruct](https://www.runlocalai.co/models/llama-3.2-90b-vision-instruct): The 90B vision Llama. Best-in-class first-party multimodal open weight at the time of release. Workstation-class only. - [Llama 3.3 70B Instruct](https://www.runlocalai.co/models/llama-3.3-70b-instruct): Late-2024 refresh of the 70B Llama line. Roughly matches Llama 3.1 405B on most benchmarks at one-fifth the parameter count. The default high-end model for serious local inference… - [Llama 4 405B](https://www.runlocalai.co/models/llama-4-405b): Meta's dense flagship in the Llama 4 line. 405B params; comparable footprint to Llama 3.1 405B with the Llama 4 reasoning improvements. - [Llama 4 70B](https://www.runlocalai.co/models/llama-4-70b): Llama 4 dense at 70B. Drop-in successor to Llama 3.3 70B; same hardware envelope, better on reasoning benchmarks. - [Llama 4 Maverick](https://www.runlocalai.co/models/llama-4-maverick): Meta's high-end Llama 4 sibling — 128 experts MoE built for performance over efficiency. Multilingual strength is its standout. Effectively a server-tier model; consumer hardware… - [Llama 4 Scout](https://www.runlocalai.co/models/llama-4-scout): Meta's 2026 flagship MoE model. 109B total parameters with only 17B active per forward pass and a record 10-million-token context window — unmatched in production at any tier. Bui… - [LLaVA 1.6 Mistral 7B](https://www.runlocalai.co/models/llava-1.6-mistral-7b): LLaVA 1.6 on Mistral 7B base. Apache 2.0 vision-language with strong OCR. - [LLaVA-OneVision 7B](https://www.runlocalai.co/models/llava-onevision-7b): LLaVA-OneVision unified single-image / multi-image / video VLM on Qwen 2 base. - [LongCat-2.0](https://www.runlocalai.co/models/longcat-2.0): LongCat-2.0 is Meituan's 1.6-trillion-parameter MoE language model, activating ~48B parameters per token, released under MIT — announced June 30, 2026, with weights and inference… - [Magistral 32B](https://www.runlocalai.co/models/magistral-32b): Mistral's reasoning-specialized fine-tune of a Mistral Small base. Reasoning-token emission similar to Qwen 3 / DeepSeek R1 in a smaller footprint. Research license — non-commerci… - [MedGemma 27B](https://www.runlocalai.co/models/medgemma-27b): Medical-specialist Gemma fine-tune. Trained on de-identified medical literature and imaging. Research use under HAI-DEF terms. - [MiniCPM 3 4B](https://www.runlocalai.co/models/minicpm-3-4b): OpenBMB's edge-optimized 4B. MIT license; designed for phone deployment. Strong reasoning per parameter. - [MiniCPM-V 2.6 8B](https://www.runlocalai.co/models/minicpm-v-2.6-8b): Multimodal MiniCPM at 8B. Vision + text; strong on document Q&A for the size class. - [MiniCPM-V 3 8B](https://www.runlocalai.co/models/minicpm-v-3-8b): MiniCPM-V successor. Multimodal at 8B with stronger document Q&A than 2.6. - [MiniMax-M3](https://www.runlocalai.co/models/minimax-m3): MiniMax-M3 is a native multimodal Mixture-of-Experts model from MiniMax (Hugging Face MiniMaxAI/MiniMax-M3, 2026-06), with ~428B total / ~23B active parameters per token. It accep… - [Ministral 3B Instruct](https://www.runlocalai.co/models/ministral-3b): Mistral edge model at 3B. Designed for on-device inference with extended 128k context. Research license only. - [Ministral 8B Instruct](https://www.runlocalai.co/models/ministral-8b): Mistral 8B with sliding-window attention and 128k context. Research license — Mistral 7B v0.3 is the commercial alternative. - [Mistral 7B Instruct v0.3](https://www.runlocalai.co/models/mistral-7b-instruct-v0.3): The reference 7B from Mistral. Apache 2.0 with native function calling. Mature ecosystem. - [Mistral Large 2 (123B)](https://www.runlocalai.co/models/mistral-large-2): Mistral's flagship dense model. Open weights but restricted commercial license — research and non-commercial only. - [Mistral Medium 3 24B (dense)](https://www.runlocalai.co/models/mistral-medium-3-24b): Dense variant in the Mistral Medium 3.5 family. Research license — non-commercial open. Same training data as the MoE flagship but in a smaller dense package. - [Mistral Medium 3.5 (675B MoE)](https://www.runlocalai.co/models/mistral-medium-3.5): Mistral's April 2026 frontier MoE. 675B total / 41B active. Strong European-multilingual lineage carries through; the new release competes head-to-head with DeepSeek V4-Pro on mos… - [Mistral Nemo 12B Instruct](https://www.runlocalai.co/models/mistral-nemo-12b): Joint Mistral/NVIDIA release with native 128K context and a new Tekken tokenizer. Strong multilingual; popular fine-tune base. - [Mistral Saba 24B](https://www.runlocalai.co/models/mistral-saba-24b): Mistral's Arabic and South Asian language specialist at 24B. Research license. - [Mistral Small 3 24B](https://www.runlocalai.co/models/mistral-small-3-24b): Re-release of Mistral Small under Apache 2.0. Competitive with Llama 3.3 70B at one-third the size for many tasks. - [Mistral Small 3.2 24B](https://www.runlocalai.co/models/mistral-small-3.2-24b): Iterative refresh of Mistral Small 3 24B. Same architecture; improved instruction following and tool-call reliability. Apache 2.0. - [Mixtral 8x22B Instruct](https://www.runlocalai.co/models/mixtral-8x22b-instruct): The bigger Mixtral. 141B total / 39B active. Strong general model, workstation-tier deployment. - [Mixtral 8x7B Instruct](https://www.runlocalai.co/models/mixtral-8x7b-instruct): The MoE model that introduced the 8-experts pattern to the open-weight world. 47B params total, 13B active. Still a viable workhorse on 36GB+ setups. - [Molmo 72B](https://www.runlocalai.co/models/molmo-72b): Molmo flagship. Apache 2.0 VLM rivaling proprietary models on UI pointing and visual reasoning. - [Molmo 7B-D](https://www.runlocalai.co/models/molmo-7b-d): AI2's fully-open VLM. Trained on PixMo dataset; pointing capability for UI grounding. - [Moondream 2](https://www.runlocalai.co/models/moondream-2): Tiny vision-language model. ~1.9B; designed for edge / embedded multimodal use cases. Apache 2.0. - [Muse Glimmer 30B](https://www.runlocalai.co/models/muse-glimmer): Muse Glimmer 30B is Meta's first open agentic model, released 2026-08-10 under the Apache-2.0 license — a genuinely permissive, commercial-friendly license that's unusual for a fr… - [Nemotron 3 Nano (30B-A3B)](https://www.runlocalai.co/models/nemotron-3-nano): NVIDIA's hybrid Mamba-2 + Transformer MoE for on-device agents. 30B total / 3B active. 1M-token context window with reasoning ON/OFF modes and 4× faster inference than the previou… - [Nemotron 3 Nano 9B](https://www.runlocalai.co/models/nemotron-3-nano-9b): NVIDIA's Nemotron 3 at 9B. Tuned for NVIDIA-stack deployment patterns; strong tool-calling reliability. - [Nemotron 3 Super (120B-A12B)](https://www.runlocalai.co/models/nemotron-3-super): Workstation-tier Nemotron 3. 120B total / 12B active. 5× higher throughput than the prior Super, 1M context, designed for multi-agent applications. - [Nemotron 3 Super 49B](https://www.runlocalai.co/models/nemotron-3-super-49b): Nemotron 3 mid-tier. 49B dense; fits 32GB cards with AWQ. NVIDIA stack alignment carries through. - [Nemotron 3 Ultra (550B-A55B)](https://www.runlocalai.co/models/nemotron-3-ultra): NVIDIA Nemotron 3 Ultra (550B-A55B) is a frontier-scale open-weight reasoning model from NVIDIA (Hugging Face nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16, 2026-06), with 550B to… - [Nemotron Mini 4B Instruct](https://www.runlocalai.co/models/nemotron-mini-4b): NVIDIA's edge-tier Nemotron. Distilled from Minitron lineage with role-play tuning. - [NV-Embed v2](https://www.runlocalai.co/models/nv-embed-v2): NVIDIA's research-grade embedding model. Mistral-7B base. Top of MTEB at release. - [OLMo 2 13B](https://www.runlocalai.co/models/olmo-2-13b): AI2's fully-open 13B. Apache 2.0; full training data + checkpoints + recipes published. The reproducibility-first model in the 13B class. - [OLMo 2 32B](https://www.runlocalai.co/models/olmo-2-32b): Fully-open OLMo 2. AI2 publishes the full training data, code, and weights — the most reproducible 32B model. - [OpenBioLLM Llama 3 70B](https://www.runlocalai.co/models/openbiollm-llama-3-70b): Medical / biomedical fine-tune of Llama 3 70B. Strong on USMLE and clinical-knowledge benchmarks; right pick when domain-specific medical depth matters more than general capabilit… - [OpenCoder 8B](https://www.runlocalai.co/models/opencoder-8b): Fully-open coding model — training data + recipes published. Apache 2.0 with verifiable open-data lineage. The right pick for academic / reproducibility-sensitive work. - [PaliGemma 2 10B](https://www.runlocalai.co/models/paligemma-2-10b): Mid-tier PaliGemma 2 fine-tuning base. Better baseline for complex vision tasks. - [PaliGemma 2 3B](https://www.runlocalai.co/models/paligemma-2-3b): PaliGemma 2 — Gemma 2 base + SigLIP vision encoder. Designed for fine-tuning on specific vision tasks. - [Phi-3.5 Mini Instruct](https://www.runlocalai.co/models/phi-3.5-mini-instruct): Compact 3.8B Phi for edge deployment. 128K context. Strong reasoning per parameter. - [Phi-3.5 Vision](https://www.runlocalai.co/models/phi-3.5-vision): Multimodal Phi 3.5. Document and chart understanding at edge size. MIT licensed. - [Phi-4 14B](https://www.runlocalai.co/models/phi-4-14b): Microsoft's Phi-4 14B trained on synthetic textbook-quality data. Punches above weight on reasoning and math; MIT licensed. - [Phi-4 Mini 4B](https://www.runlocalai.co/models/phi-4-mini-4b): Microsoft's edge-tier Phi-4 variant. 3.8B params; designed for phone / tablet / Pi deployment. Strong reasoning per parameter — Phi family's traditional advantage carries to the s… - [Phi-4 Multimodal](https://www.runlocalai.co/models/phi-4-multimodal): Multimodal variant of Phi-4 14B. Vision + text. Smaller than Llama 4 Scout but covers most image-Q&A workflows; right-sized for 16GB consumer cards. - [Phi-4 Reasoning 14B](https://www.runlocalai.co/models/phi-4-reasoning-14b): Reasoning-focused fine-tune of Phi-4. Visible chain-of-thought, competitive with much larger models on math and STEM benchmarks. - [Phi-4 Reasoning Mini 4B](https://www.runlocalai.co/models/phi-4-reasoning-mini-4b): Phi-4 reasoning at the edge tier. 3.8B with reasoning-token emission. The right pick when reasoning matters AND edge deployment is required. - [Phind CodeLlama 34B v2](https://www.runlocalai.co/models/phind-codellama-34b-v2): Phind's CodeLlama-derived coder at 34B. Older release; retained for historical / continuity value. Newer Qwen Coder lineage has surpassed it. - [Pixtral 12B](https://www.runlocalai.co/models/pixtral-12b): Mistral's multimodal entry. 12B parameters, vision + text, Apache 2.0. Good document and chart understanding. - [Qwen 2-VL 7B](https://www.runlocalai.co/models/qwen-2-vl-7b): Qwen 2 vision-language predecessor to Qwen 2.5-VL. Apache 2.0 with strong document Q&A. - [Qwen 2.5 0.5B Instruct](https://www.runlocalai.co/models/qwen-2.5-0.5b-instruct): Smallest Qwen 2.5. Apache 2.0; phone / Pi-class deployment target. - [Qwen 2.5 1.5B Instruct](https://www.runlocalai.co/models/qwen-2.5-1.5b-instruct): Compact Qwen 2.5. The 1.5B Apache-2.0 baseline. - [Qwen 2.5 14B Instruct](https://www.runlocalai.co/models/qwen-2.5-14b-instruct): 14B Qwen 2.5. Sweet spot for 16GB VRAM. Many production deployments still on this version. - [Qwen 2.5 32B Instruct](https://www.runlocalai.co/models/qwen-2.5-32b-instruct): Dense 32B Qwen 2.5. Strong daily-driver on 24GB cards prior to Qwen 3 32B. - [Qwen 2.5 3B Instruct](https://www.runlocalai.co/models/qwen-2.5-3b-instruct): Mid-edge Qwen 2.5. Note: 3B variant uses Qwen License (not Apache 2.0). - [Qwen 2.5 72B Instruct](https://www.runlocalai.co/models/qwen-2.5-72b-instruct): The flagship of Qwen 2.5. Workstation-tier; needs 48GB+ VRAM for usable inference. - [Qwen 2.5 7B Instruct](https://www.runlocalai.co/models/qwen-2.5-7b-instruct): The community-default small Qwen prior to Qwen 3. Still widely used because of mature ecosystem support. - [Qwen 2.5 Coder 1.5B](https://www.runlocalai.co/models/qwen-2.5-coder-1.5b): Smallest Qwen 2.5 Coder. Targets edge / autocomplete on integrated GPUs and Apple Silicon laptops. - [Qwen 2.5 Coder 14B Instruct](https://www.runlocalai.co/models/qwen-2.5-coder-14b-instruct): Coding-specialized Qwen 2.5 at 14B. The 16GB-VRAM tier coding model — fits comfortably with 8K context. - [Qwen 2.5 Coder 32B Instruct](https://www.runlocalai.co/models/qwen-2.5-coder-32b-instruct): Coding-specialist Qwen 2.5. Beats GPT-4o on HumanEval and matches Sonnet on many code-edit benchmarks. The default local-coding model on 24GB cards. - [Qwen 2.5 Coder 3B](https://www.runlocalai.co/models/qwen-2.5-coder-3b): Compact Qwen 2.5 Coder. Sweet spot for laptop autocomplete and small refactor agents. - [Qwen 2.5 Coder 7B Instruct](https://www.runlocalai.co/models/qwen-2.5-coder-7b-instruct): Coding-specialized Qwen 2.5 at 7B. The 8-12GB-VRAM coding model — entry-tier autocomplete + IDE assistant. Smaller sibling of the 14B / 32B Coder line. - [Qwen 2.5 Math 72B](https://www.runlocalai.co/models/qwen-2.5-math-72b): Largest Qwen 2.5 Math. Datacenter-tier math specialist; eclipsed by R1 distills for general reasoning. - [Qwen 2.5 Math 7B](https://www.runlocalai.co/models/qwen-2.5-math-7b): Qwen 2.5 fine-tuned for math problem-solving with chain-of-thought and tool-integrated reasoning. - [Qwen 2.5-VL 3B](https://www.runlocalai.co/models/qwen-2.5-vl-3b): Smallest Qwen 2.5-VL. Edge-deployable VLM with strong document Q&A. - [Qwen 2.5-VL 72B](https://www.runlocalai.co/models/qwen-2.5-vl-72b): Qwen 2.5 vision-language flagship at 72B. Strong on document understanding + multi-image queries. Apache 2.0. - [Qwen 2.5-VL 7B](https://www.runlocalai.co/models/qwen-2.5-vl-7b): Consumer-tier Qwen 2.5 VL. 7B + vision. Fits 8GB cards; the smallest practical multimodal Qwen. - [Qwen 3 14B](https://www.runlocalai.co/models/qwen-3-14b): 14B Qwen 3. Fits on 12GB cards at Q4. Strong default for users with a single mid-range GPU. - [Qwen 3 235B-A22B](https://www.runlocalai.co/models/qwen-3-235b-a22b): Qwen 3 flagship MoE. 235B total / 22B active per token, with built-in 'thinking' and 'non-thinking' modes that trade speed for reasoning depth at inference time. Best open-weight… - [Qwen 3 30B-A3B](https://www.runlocalai.co/models/qwen-3-30b-a3b): Mid-tier Qwen 3 MoE. 30B total / 3B active means 70B-class quality at 7B-class inference speed on a single 24GB card. The sweet spot of the Qwen 3 lineup for prosumer hardware. - [Qwen 3 32B](https://www.runlocalai.co/models/qwen-3-32b): Dense Qwen 3 32B. Best dense open-weight model in its size class at release; pairs nicely with a single RTX 5090 or 4090. - [Qwen 3 4B](https://www.runlocalai.co/models/qwen-3-4b): Compact Qwen 3 for edge and laptop deployment. Outperforms many 7B models from prior generations. - [Qwen 3.6 27B (MTP)](https://www.runlocalai.co/models/qwen-3-6-27b-mtp): Qwen 3.6 27B dense (not MoE) with Multi-Token Prediction. Sits between the 14B and 35B-A3B as a "single dense model with MTP throughput acceleration." Targets workloads where the… - [Qwen 3.6 35B-A3B (MTP)](https://www.runlocalai.co/models/qwen-3-6-35b-a3b-mtp): Qwen 3.6 35B-A3B with Multi-Token Prediction (MTP). The "A3B" suffix means ~3B activated parameters per token via Mixture-of-Experts — inference cost stays mid-tier while total pa… - [Qwen 3 7B](https://www.runlocalai.co/models/qwen-3-7b): Qwen 3 mid-tier. Same reasoning-mode toggle as Qwen 3 32B/14B/8B. Hits the consumer-laptop sweet spot. - [Qwen 3 8B](https://www.runlocalai.co/models/qwen-3-8b): Qwen 3 at the 8B scale. Direct head-to-head against Llama 3.1 8B on most benchmarks; usually wins on coding and structured output. - [Qwen 3 Coder 32B](https://www.runlocalai.co/models/qwen-3-coder-32b): Coding-specialized fine-tune of Qwen 3 32B. Curated coding corpus; outperforms Qwen 2.5 Coder 32B on SWE-Bench by ~6 points. Apache 2.0. - [Qwen 3 Embedding 8B](https://www.runlocalai.co/models/qwen-3-embedding-8b): Qwen 3 family embedding model. Apache 2.0 with strong multilingual coverage. - [Qwen 3.5 235B-A17B (MoE)](https://www.runlocalai.co/models/qwen-3.5-235b-a17b): Alibaba's May 2026 flagship. 397B total / 17B active MoE with hybrid thinking-mode toggle inherited from Qwen 3. Strongest open scientific reasoner per GPQA Diamond. The strongest… - [QwQ 32B Preview](https://www.runlocalai.co/models/qwq-32b): Qwen team's reasoning-focused experimental release. Visible chain-of-thought in