Describe your build — any GPUs, CPU, RAM, OS, runtime, use case. We'll compute effective VRAM honestly, recommend a runtime, and tell you which models fit comfortably, which are borderline, and which aren't practical.
Total VRAM ≠ pooled VRAM. We never sum VRAM unless the silicon truly pools (Apple unified memory). We always explain why effective is lower than total.
Calculations follow the RunLocalAI Will-It-Run Framework: effective VRAM, model working set, runtime constraints, fit tiers, and measured-vs-estimated evidence labels.
Add GPUs, set CPU/RAM/OS, optionally pick a runtime + use case. URL updates as you change fields — share a build by copying the URL.
Single NVIDIA GeForce RTX 3080 16GB (Mobile) — 16 GB VRAM minus ~1.8 GB runtime/driver overhead = ~14 GB usable for weights + KV cache + activations. The remaining uncertainty band covers OS display use and background CUDA allocations.
Publicly inspectable measured rows for the selected hardware slug(s). Exact measured rows calibrate the fit table instead of leaving it as pure VRAM estimation.
| Model | Evidence | Quant | Tok/s | Provenance |
|---|---|---|---|---|
| Turkcell LLM 7B v1 NVIDIA GeForce RTX 3080 16GB (Mobile) | 4K ctx ollama version is 0.24.0 Microsoft Windows [Version 10.0.26200.8457] Driver 571.96 | Q4_K_M | 85.77 | |
| RefinedNeuro RN TR R2 NVIDIA GeForce RTX 3080 16GB (Mobile) | 4K ctx ollama version is 0.24.0 Microsoft Windows [Version 10.0.26200.8457] Driver 571.96 | Q4_K_M | 79.27 | |
| RefinedNeuro RN TR R1 NVIDIA GeForce RTX 3080 16GB (Mobile) | 4K ctx ollama version is 0.24.0 Microsoft Windows [Version 10.0.26200.8457] Driver 571.96 | Q4_K_M | 79.89 | |
| Qwen 3 4B NVIDIA GeForce RTX 3080 16GB (Mobile) | 4K ctx ollama version is 0.24.0 Microsoft Windows [Version 10.0.26200.8457] Driver 571.96 | Q4_K_M | 103.7 | |
| Qwen 3 14B NVIDIA GeForce RTX 3080 16GB (Mobile) | 4K ctx ollama version is 0.24.0 Microsoft Windows [Version 10.0.26200.8457] Driver 571.96 | Q4_K_M | 38.26 | |
| Qwen 2.5 7B Instruct NVIDIA GeForce RTX 3080 16GB (Mobile) | 4K ctx ollama version is 0.24.0 Microsoft Windows [Version 10.0.26200.8457] Driver 571.96 | Q4_K_M | 80.42 | |
| Phi-4 Reasoning 14B NVIDIA GeForce RTX 3080 16GB (Mobile) | 4K ctx ollama version is 0.24.0 Microsoft Windows [Version 10.0.26200.8457] Driver 571.96 | Q4_K_M | 40.44 | |
| Phi-3.5 Mini Instruct NVIDIA GeForce RTX 3080 16GB (Mobile) | 4K ctx ollama version is 0.24.0 Microsoft Windows [Version 10.0.26200.8457] Driver 571.96 | Q4_K_M | 155.4 | |
| Mistral Nemo 12B Instruct NVIDIA GeForce RTX 3080 16GB (Mobile) | 4K ctx ollama version is 0.24.0 Microsoft Windows [Version 10.0.26200.8457] Driver 571.96 | Q4_K_M | 65.72 | |
| Mistral 7B Instruct v0.3 NVIDIA GeForce RTX 3080 16GB (Mobile) | 4K ctx ollama version is 0.24.0 Microsoft Windows [Version 10.0.26200.8457] Driver 571.96 | Q4_K_M | 89.64 | |
| Llama 3.2 11B Vision Instruct NVIDIA GeForce RTX 3080 16GB (Mobile) | 4K ctx ollama version is 0.24.0 Microsoft Windows [Version 10.0.26200.8457] Driver 571.96 | Q4_K_M | 67.00 | |
| Malhajar Mistral 7B Turkish NVIDIA GeForce RTX 3080 16GB (Mobile) | 4K ctx ollama version is 0.24.0 Microsoft Windows [Version 10.0.26200.8457] Driver 571.96 | Q4_K_M | 87.28 |
| Model | Benchmark | Setup | Score | Log |
|---|---|---|---|---|
| Ministral 3 14B | HumanEval+ | Q4_K_M / ollama-0.32 | 76.8 | gist |
| Qwen3.5 9B | HumanEval+ | Q4_K_M / ollama-0.32 | 42.7 | gist |
| Ornith 1.0 9B | HumanEval+ | Q4_K_M / ollama-0.32 | 73.2 | gist |
| Qwen 3 8B | HumanEval+ | Q4_K_M / ollama-0.24 | 2.4 | gist |
| Llama 3.1 8B Instruct | MBPP+ | Q4_K_M / ollama-0.24 | 39.2 | gist |
| Phi-4 14B | MBPP+ | Q4_K_M / ollama-0.24 | 60.3 | gist |
| Qwen 2.5 Coder 7B Instruct | MBPP+ | Q4_K_M / ollama-0.24 | 66.9 | gist |
| Llama 3.1 8B Instruct | HumanEval+ | Q4_K_M / ollama-0.24 | 56.1 | gist |
| Phi-4 14B | HumanEval+ | Q4_K_M / ollama-0.24 | 78.7 | gist |
| Qwen 2.5 Coder 7B Instruct | HumanEval+ | Q4_K_M / ollama-0.24 | 81.1 | gist |
| Llama 3.2 3B Instruct | TurkishMMLU (Generative) | Q4_K_M / ollama-0.24 | 11.4 | gist |
| Turkish Llama 8B Instruct v0.1 | TurkishMMLU (Generative) | Q4_K_M / ollama-0.24 | 11.0 | gist |
Best engine for this topology + skill level + use case.
345 models considered. Categorized by headroom at the recommended quant + a sensible context for your use case.
| Model | Params | Quant | VRAM est. | Context | Evidence | Note |
|---|---|---|---|---|---|---|
| OLMo 2 13B | 13B | Q4_K_M | 11.4 GB | 4,096 | No measured row yet | Fits cleanly at Q4_K_M + 4,096 ctx with 18% headroom. |
| OpenThaiGPT 1.0.0 Beta 13B Chat | 13B | Q4_K_M | 10.8 GB | 4,096 | No measured row yet | Fits cleanly at Q4_K_M + 4,096 ctx with 23% headroom. |
| mGPT 13B | 13B | Q4_K_M | 9.2 GB | 2,048 | No measured row yet | Fits cleanly at Q4_K_M + 2,048 ctx with 34% headroom. |
| Stable LM 2 12B | 12B | Q4_K_M | 10.6 GB | 4,096 | No measured row yet | Fits cleanly at Q4_K_M + 4,096 ctx with 25% headroom. |
| FLUX.1 [dev] | 12B | Q4_K_M | 6.9 GB | 0 | No measured row yet | Comfortable fit with 51% headroom — room to extend context or run alongside other workloads. |
| FLUX.1 [schnell] | 12B | Q4_K_M | 6.9 GB | 0 | No measured row yet | Comfortable fit with 51% headroom — room to extend context or run alongside other workloads. |
| Merlyn Education Safety 12B AWQ | 12B | Q4_K_M | 8.4 GB | 2,048 | No measured row yet | Fits cleanly at Q4_K_M + 2,048 ctx with 40% headroom. |
| Trendyol LLM Asure 12B | 12B | GGUF_UNKNOWN | 10.9 GB | 8,192 | No measured row yet | Fits cleanly at GGUF_UNKNOWN + 8,192 ctx with 22% headroom. |
| Bielik 11B v2.3 Instruct | 11B | Q4_K_M | 9.2 GB | 4,096 | No measured row yet | Fits cleanly at Q4_K_M + 4,096 ctx with 35% headroom. |
| Bielik 11B v2.3 Instruct | 11B | Q4_K_M | 9.2 GB | 4,096 | No measured row yet | Fits cleanly at Q4_K_M + 4,096 ctx with 35% headroom. |
| Bielik-11B v3.0 Instruct FP8 Dynamic | 11B | Q4_K_M | 9.2 GB | 4,096 | No measured row yet | Fits cleanly at Q4_K_M + 4,096 ctx with 35% headroom. |
| SOLAR 10.7B v1.0 | 11B | Q4_K_M | 8.9 GB | 4,096 | No measured row yet | Fits cleanly at Q4_K_M + 4,096 ctx with 37% headroom. |
| Falcon 3 10B | 10B | Q4_K_M | 11.3 GB | 8,192 | No measured row yet | Fits cleanly at Q4_K_M + 8,192 ctx with 19% headroom. |
| YTU Turkish Gemma 9B v0.1 | 9B | Q4_K_M | 10.7 GB | 8,192 | No measured row yet | Fits cleanly at Q4_K_M + 8,192 ctx with 24% headroom. |
| NVIDIA Nemotron Nano 9B v2 Japanese | 9B | Q4_K_M | 9.8 GB | 8,192 | No measured row yet | Fits cleanly at Q4_K_M + 8,192 ctx with 30% headroom. |
| Nemotron 3 Nano 9B | 9B | Q4_K_M | 10.1 GB | 8,192 | No measured row yet | Fits cleanly at Q4_K_M + 8,192 ctx with 28% headroom. |
| Gemma 2 9B Instruct | 9B | Q4_K_M | 10.6 GB | 8,192 | No measured row yet | Fits cleanly at Q4_K_M + 8,192 ctx with 24% headroom. |
| Turkish Gemma 9B T1 | 9B | Q4_K_M | 9.8 GB | 8,192 | No measured row yet | Fits cleanly at Q4_K_M + 8,192 ctx with 30% headroom. |
| Yi Coder 9B | 9B | Q4_K_M | 10.2 GB | 8,192 | No measured row yet | Fits cleanly at Q4_K_M + 8,192 ctx with 27% headroom. |
| GLM-4 9B | 9B | Q4_K_M | 10.3 GB | 8,192 | No measured row yet | Fits cleanly at Q4_K_M + 8,192 ctx with 27% headroom. |
| Qwen3.5 9B | 9B | Q4_K_M | 11.4 GB | 8,192 | No measured row yet | Fits cleanly at Q4_K_M + 8,192 ctx with 18% headroom. |
| Ornith 1.0 9B | 9B | Q4_K_M | 10.4 GB | 8,192 | No measured row yet | Fits cleanly at Q4_K_M + 8,192 ctx with 26% headroom. |
| Qwen3.5 9B Thai Law Base | 9B | Q4_K_M | 7.4 GB | 4,096 | No measured row yet | Comfortable fit with 47% headroom — room to extend context or run alongside other workloads. |
| Granite 4.1 8B Instruct | 9B | Q5_K_M | 10.7 GB | 8,192 | No measured row yet | Fits cleanly at Q5_K_M + 8,192 ctx with 23% headroom. |
| Model | Params | Quant | VRAM est. | Context | Evidence | Note |
|---|---|---|---|---|---|---|
| DeepSeek MoE 16B Base | 16B | Q4_K_M | 14 GB | 4,096 | No measured row yet | Tight fit at Q4_K_M — only 0% headroom. KV cache for longer context will OOM. Cap context tighter or drop one quant level. |
| Phi-4 Reasoning 14B | 14B | Q4_K_M | 15.8 GB | 8,192 | Measured on this hardware: 40.44 tok/s at Q4_K_M (4,096 measured ctx). The larger target context still needs validation before calling it comfortable. | |
| Qwen 3 14B | 14B | Q4_K_M | 15.8 GB | 8,192 | Measured on this hardware: 38.26 tok/s at Q4_K_M (4,096 measured ctx). The larger target context still needs validation before calling it comfortable. | |
| Pixtral 12B | 12B | Q4_K_M | 13.4 GB | 8,192 | No measured row yet | Tight fit at Q4_K_M — only 5% headroom. KV cache for longer context will OOM. Cap context tighter or drop one quant level. |
| Mistral Nemo 12B Instruct | 12B | Q4_K_M | 13.9 GB | 8,192 | Measured on this hardware: 65.72 tok/s at Q4_K_M (4,096 measured ctx). Tight fit at Q4_K_M — only 1% headroom. KV cache for longer context will OOM. Cap context tighter or drop one quant level. | |
| Gemma 3 12B | 12B | Q4_K_M | 13.7 GB | 8,192 | No measured row yet | Tight fit at Q4_K_M — only 2% headroom. KV cache for longer context will OOM. Cap context tighter or drop one quant level. |
| Gemma 4 12B | 12B | Q4_K_M | 14 GB | 8,192 | No measured row yet | Tight fit at Q4_K_M — only 0% headroom. KV cache for longer context will OOM. Cap context tighter or drop one quant level. |
| Bielik 11B v2.2 Instruct GGUF | 11B | Q4_K_M | 11.9 GB | 8,192 | No measured row yet | Tight fit at Q4_K_M — only 15% headroom. KV cache for longer context will OOM. Cap context tighter or drop one quant level. |
| Llama 3.2 11B Vision Instruct | 11B | Q4_K_M | 13.8 GB | 8,192 | Measured on this hardware: 67.00 tok/s at Q4_K_M (4,096 measured ctx). Tight fit at Q4_K_M — only 1% headroom. KV cache for longer context will OOM. Cap context tighter or drop one quant level. | |
| Llama 3.2 11B Vision | 11B | Q4_K_M | 12.3 GB | 8,192 | No measured row yet | Tight fit at Q4_K_M — only 12% headroom. KV cache for longer context will OOM. Cap context tighter or drop one quant level. |
| Bielik 11B v3.0 Instruct GGUF | 11B | Q4_K_M | 11.9 GB | 8,192 | No measured row yet | Tight fit at Q4_K_M — only 15% headroom. KV cache for longer context will OOM. Cap context tighter or drop one quant level. |
| Hermes 3 Llama 3.1 8B | 8B | Q8_0 | 12.9 GB | 8,192 | No measured row yet | Tight fit at Q8_0 — only 8% headroom. KV cache for longer context will OOM. Cap context tighter or drop one quant level. |
| Qwen 3 8B | 8B | Q8_0 | 12.6 GB | 8,192 | No measured row yet | Tight fit at Q8_0 — only 10% headroom. KV cache for longer context will OOM. Cap context tighter or drop one quant level. |
| DeepSeek R1 Distill Qwen 7B | 7B | Q8_0 | 12 GB | 8,192 | No measured row yet | Tight fit at Q8_0 — only 14% headroom. KV cache for longer context will OOM. Cap context tighter or drop one quant level. |
| Model | Params | Quant | VRAM est. | Context | Evidence | Note |
|---|---|---|---|---|---|---|
| NV-Embed v2 | 8B | FP16 | 19.7 GB | 8,192 | No measured row yet | ~19.7 GB needed at FP16 + 8,192 ctx — overshoots effective VRAM by 41%. Drop quant or move to a larger build. |
| Qwen 3 Embedding 8B | 8B | FP16 | 20.8 GB | 8,192 | No measured row yet | ~20.8 GB needed at FP16 + 8,192 ctx — overshoots effective VRAM by 49%. Drop quant or move to a larger build. |
| PaliGemma 2 10B | 10B | BF16 | 26 GB | 8,192 | No measured row yet | ~26.0 GB needed at BF16 + 8,192 ctx — overshoots effective VRAM by 86%. Drop quant or move to a larger build. |
| Mellum2 12B-A2.5B | 12B | Q4_K_M | 14.5 GB | 8,192 | No measured row yet | ~14.5 GB needed at Q4_K_M + 8,192 ctx — overshoots effective VRAM by 3%. Drop quant or move to a larger build. |
| Baichuan 4 13B | 13B | Q4_K_M | 14.7 GB | 8,192 | No measured row yet | ~14.7 GB needed at Q4_K_M + 8,192 ctx — overshoots effective VRAM by 5%. Drop quant or move to a larger build. |
| GLM-4V 9B | 14B | Q4_K_M | 15.9 GB | 8,192 | No measured row yet | ~15.9 GB needed at Q4_K_M + 8,192 ctx — overshoots effective VRAM by 13%. Drop quant or move to a larger build. |
| Phi-4 14B | 14B | Q4_K_M | 15.8 GB | 8,192 | No measured row yet | ~15.8 GB needed at Q4_K_M + 8,192 ctx — overshoots effective VRAM by 13%. Drop quant or move to a larger build. |
| Qwen 2.5 14B Instruct | 14B | Q4_K_M | 16.4 GB | 8,192 | No measured row yet | ~16.4 GB needed at Q4_K_M + 8,192 ctx — overshoots effective VRAM by 17%. Drop quant or move to a larger build. |
| Qwen 2.5 Coder 14B Instruct | 14B | Q4_K_M | 15.8 GB | 8,192 | No measured row yet | ~15.8 GB needed at Q4_K_M + 8,192 ctx — overshoots effective VRAM by 13%. Drop quant or move to a larger build. |
| DeepSeek R1 Distill Qwen 14B | 14B | Q4_K_M | 15.8 GB | 8,192 | No measured row yet | ~15.8 GB needed at Q4_K_M + 8,192 ctx — overshoots effective VRAM by 13%. Drop quant or move to a larger build. |
| Phi-4 Multimodal | 14B | Q4_K_M | 16.5 GB | 8,192 | No measured row yet | ~16.5 GB needed at Q4_K_M + 8,192 ctx — overshoots effective VRAM by 18%. Drop quant or move to a larger build. |
| Ministral 3 14B | 14B | Q4_K_M | 16.6 GB | 8,192 | No measured row yet | ~16.6 GB needed at Q4_K_M + 8,192 ctx — overshoots effective VRAM by 18%. Drop quant or move to a larger build. |
| StarCoder 2 15B | 15B | Q4_K_M | 17 GB | 8,192 | No measured row yet | ~17.0 GB needed at Q4_K_M + 8,192 ctx — overshoots effective VRAM by 21%. Drop quant or move to a larger build. |
| DeepSeek V2 Lite Chat | 16B | Q4_K_M | 16.9 GB | 8,192 | No measured row yet | ~16.9 GB needed at Q4_K_M + 8,192 ctx — overshoots effective VRAM by 21%. Drop quant or move to a larger build. |
| DeepSeek V3 Lite (16B MoE) | 16B | Q4_K_M | 18 GB | 8,192 | No measured row yet | ~18.0 GB needed at Q4_K_M + 8,192 ctx — overshoots effective VRAM by 28%. Drop quant or move to a larger build. |
| DeepSeek Coder V2 Lite (16B) | 16B | Q4_K_M | 18 GB | 8,192 | No measured row yet | ~18.0 GB needed at Q4_K_M + 8,192 ctx — overshoots effective VRAM by 28%. Drop quant or move to a larger build. |
NVLink vs PCIe, tensor- vs pipeline-parallel, mixed-card honesty.
Curated multi-GPU / cluster setups with effective-VRAM math.
OS + runtime install commands for your stack.
Runtime × OS × hardware support truth table.
If you're sizing a fresh AI build (not just a card to drop into an existing system), the build-budget walkthroughs cover the whole BOM honestly: AI PC build under $1,000 or AI PC build under $2,000 cover the realistic 2026 budget tiers.
Vertical-fit shopping? AI PC for students covers the budget + portability tradeoffs; AI PC for developers covers the coding workflow specifics; AI PC for small business covers the document-RAG / always-on machine.
Form-factor first? See best laptop for local AI, best Mac for local AI, best mini PC for local AI, or best used GPU for local AI.