RUNLOCALAIv38
->Will it run?Best GPUCompareTroubleshootStartLearnPulseModelsHardwareToolsBench
Run check
RUNLOCALAI

Independently operated catalog for local-AI hardware and software. Hand-written verdicts. Source-cited claims. Reproducible commands when we have them.

OP·Eruo Fredoline
DIR
  • Models
  • Hardware
  • Tools
  • Benchmarks
TOOLS
  • Will it run?
  • Compare hardware
  • Cost vs cloud
  • Choose my GPU
  • Prompting kits
  • Quick answers
REF
  • All buyer guides
  • Learn local AI
  • Methodology
  • Glossary
  • Errors KB
  • Trust
EDITOR
  • About
  • Author
  • How we make money
  • Editorial policy
  • Contact
LEGAL
  • Privacy
  • Terms
  • Sitemap
MAIL · MONTHLY DIGEST
Get monthly local AI changes
Monthly recap. No spam.
DISCLOSURE

Some links on this site are affiliate links (Amazon Associates and other first-class retailers). When you buy through them, we earn a small commission at no extra cost to you. Affiliate links do not influence our verdicts — there are cards we rate highly that we don't have affiliate relationships with, and cards that sell well that we refuse to recommend. Read more →

© 2026 runlocalai.coIndependently operated
RUNLOCALAI · v38
← Back to Will-it-run

Custom build engine

Describe your build — any GPUs, CPU, RAM, OS, runtime, use case. We'll compute effective VRAM honestly, recommend a runtime, and tell you which models fit comfortably, which are borderline, and which aren't practical.

Total VRAM ≠ pooled VRAM. We never sum VRAM unless the silicon truly pools (Apple unified memory). We always explain why effective is lower than total.

Calculations follow the RunLocalAI Will-It-Run Framework: effective VRAM, model working set, runtime constraints, fit tiers, and measured-vs-estimated evidence labels.

Describe your build

Add GPUs, set CPU/RAM/OS, optionally pick a runtime + use case. URL updates as you change fields — share a build by copying the URL.

Build summary

Total VRAM
16 GB
Effective VRAM
~14 GB
range 13-14 GB
Topology
single gpu
none
Setup difficulty
beginner
speed penalty ~0%
Why effective VRAM is lower than total

Single NVIDIA GeForce RTX 3080 16GB (Mobile) — 16 GB VRAM minus ~1.8 GB runtime/driver overhead = ~14 GB usable for weights + KV cache + activations. The remaining uncertainty band covers OS display use and background CUDA allocations.

Measured evidence on this hardware

Publicly inspectable measured rows for the selected hardware slug(s). Exact measured rows calibrate the fit table instead of leaving it as pure VRAM estimation.

Speed rows
12 measured
ModelEvidenceQuantTok/sProvenance
Turkcell LLM 7B v1
NVIDIA GeForce RTX 3080 16GB (Mobile)
4K ctx
ollama version is 0.24.0
Microsoft Windows [Version 10.0.26200.8457]
Driver 571.96
Q4_K_M85.77
Measured hererun logindependent reproduction pending
RefinedNeuro RN TR R2
NVIDIA GeForce RTX 3080 16GB (Mobile)
4K ctx
ollama version is 0.24.0
Microsoft Windows [Version 10.0.26200.8457]
Driver 571.96
Q4_K_M79.27
Measured hererun logindependent reproduction pending
RefinedNeuro RN TR R1
NVIDIA GeForce RTX 3080 16GB (Mobile)
4K ctx
ollama version is 0.24.0
Microsoft Windows [Version 10.0.26200.8457]
Driver 571.96
Q4_K_M79.89
Measured hererun logindependent reproduction pending
Qwen 3 4B
NVIDIA GeForce RTX 3080 16GB (Mobile)
4K ctx
ollama version is 0.24.0
Microsoft Windows [Version 10.0.26200.8457]
Driver 571.96
Q4_K_M103.7
Measured hererun logindependent reproduction pending
Qwen 3 14B
NVIDIA GeForce RTX 3080 16GB (Mobile)
4K ctx
ollama version is 0.24.0
Microsoft Windows [Version 10.0.26200.8457]
Driver 571.96
Q4_K_M38.26
Measured hererun logindependent reproduction pending
Qwen 2.5 7B Instruct
NVIDIA GeForce RTX 3080 16GB (Mobile)
4K ctx
ollama version is 0.24.0
Microsoft Windows [Version 10.0.26200.8457]
Driver 571.96
Q4_K_M80.42
Measured hererun logindependent reproduction pending
Phi-4 Reasoning 14B
NVIDIA GeForce RTX 3080 16GB (Mobile)
4K ctx
ollama version is 0.24.0
Microsoft Windows [Version 10.0.26200.8457]
Driver 571.96
Q4_K_M40.44
Measured hererun logindependent reproduction pending
Phi-3.5 Mini Instruct
NVIDIA GeForce RTX 3080 16GB (Mobile)
4K ctx
ollama version is 0.24.0
Microsoft Windows [Version 10.0.26200.8457]
Driver 571.96
Q4_K_M155.4
Measured hererun logindependent reproduction pending
Mistral Nemo 12B Instruct
NVIDIA GeForce RTX 3080 16GB (Mobile)
4K ctx
ollama version is 0.24.0
Microsoft Windows [Version 10.0.26200.8457]
Driver 571.96
Q4_K_M65.72
Measured hererun logindependent reproduction pending
Mistral 7B Instruct v0.3
NVIDIA GeForce RTX 3080 16GB (Mobile)
4K ctx
ollama version is 0.24.0
Microsoft Windows [Version 10.0.26200.8457]
Driver 571.96
Q4_K_M89.64
Measured hererun logindependent reproduction pending
Llama 3.2 11B Vision Instruct
NVIDIA GeForce RTX 3080 16GB (Mobile)
4K ctx
ollama version is 0.24.0
Microsoft Windows [Version 10.0.26200.8457]
Driver 571.96
Q4_K_M67.00
Measured hererun logindependent reproduction pending
Malhajar Mistral 7B Turkish
NVIDIA GeForce RTX 3080 16GB (Mobile)
4K ctx
ollama version is 0.24.0
Microsoft Windows [Version 10.0.26200.8457]
Driver 571.96
Q4_K_M87.28
Measured hererun logindependent reproduction pending
Quality rows
12 public
ModelBenchmarkSetupScoreLog
Ministral 3 14BHumanEval+Q4_K_M / ollama-0.3276.8gist
Qwen3.5 9BHumanEval+Q4_K_M / ollama-0.3242.7gist
Ornith 1.0 9BHumanEval+Q4_K_M / ollama-0.3273.2gist
Qwen 3 8BHumanEval+Q4_K_M / ollama-0.242.4gist
Llama 3.1 8B InstructMBPP+Q4_K_M / ollama-0.2439.2gist
Phi-4 14BMBPP+Q4_K_M / ollama-0.2460.3gist
Qwen 2.5 Coder 7B InstructMBPP+Q4_K_M / ollama-0.2466.9gist
Llama 3.1 8B InstructHumanEval+Q4_K_M / ollama-0.2456.1gist
Phi-4 14BHumanEval+Q4_K_M / ollama-0.2478.7gist
Qwen 2.5 Coder 7B InstructHumanEval+Q4_K_M / ollama-0.2481.1gist
Llama 3.2 3B InstructTurkishMMLU (Generative)Q4_K_M / ollama-0.2411.4gist
Turkish Llama 8B Instruct v0.1TurkishMMLU (Generative)Q4_K_M / ollama-0.2411.0gist

Recommended runtime

Best engine for this topology + skill level + use case.

ExLlamaV2
primary
involved

Highest single-stream throughput on consumer NVIDIA. EXL2 mixed-bit quants are the leading consumer-tier inference format.

llama.cpp
alternative
moderate

Cross-format flexibility — GGUF works everywhere; the engine that powers Ollama and LM Studio.

WORKLOAD PROFILE
FITS
OLMo 2 13B @ Q4_K_M, 4K context on NVIDIA GeForce RTX 3080 16GB (Mobile)
0 GB16 GBVRAM ceiling
Weights7.8 GB
KV cache3.3 GB
Activations0.4 GB
Runtime1.8 GB
Headroom2.8 GB
ESTIMATED DECODE RATE
53 tok/s
Bandwidth-derived estimate · efficiency 0.80. Real-world rates land within ±20% on well-tuned runtimes.
53 tokens per second02550100150

Models that fit your build

345 models considered. Categorized by headroom at the recommended quant + a sensible context for your use case.

Comfortable
24 models · ≥15% headroom
ModelParamsQuantVRAM est.ContextEvidenceNote
OLMo 2 13B13BQ4_K_M11.4 GB4,096No measured row yetFits cleanly at Q4_K_M + 4,096 ctx with 18% headroom.
OpenThaiGPT 1.0.0 Beta 13B Chat13BQ4_K_M10.8 GB4,096No measured row yetFits cleanly at Q4_K_M + 4,096 ctx with 23% headroom.
mGPT 13B13BQ4_K_M9.2 GB2,048No measured row yetFits cleanly at Q4_K_M + 2,048 ctx with 34% headroom.
Stable LM 2 12B12BQ4_K_M10.6 GB4,096No measured row yetFits cleanly at Q4_K_M + 4,096 ctx with 25% headroom.
FLUX.1 [dev]12BQ4_K_M6.9 GB0No measured row yetComfortable fit with 51% headroom — room to extend context or run alongside other workloads.
FLUX.1 [schnell]12BQ4_K_M6.9 GB0No measured row yetComfortable fit with 51% headroom — room to extend context or run alongside other workloads.
Merlyn Education Safety 12B AWQ12BQ4_K_M8.4 GB2,048No measured row yetFits cleanly at Q4_K_M + 2,048 ctx with 40% headroom.
Trendyol LLM Asure 12B12BGGUF_UNKNOWN10.9 GB8,192No measured row yetFits cleanly at GGUF_UNKNOWN + 8,192 ctx with 22% headroom.
Bielik 11B v2.3 Instruct11BQ4_K_M9.2 GB4,096No measured row yetFits cleanly at Q4_K_M + 4,096 ctx with 35% headroom.
Bielik 11B v2.3 Instruct11BQ4_K_M9.2 GB4,096No measured row yetFits cleanly at Q4_K_M + 4,096 ctx with 35% headroom.
Bielik-11B v3.0 Instruct FP8 Dynamic11BQ4_K_M9.2 GB4,096No measured row yetFits cleanly at Q4_K_M + 4,096 ctx with 35% headroom.
SOLAR 10.7B v1.011BQ4_K_M8.9 GB4,096No measured row yetFits cleanly at Q4_K_M + 4,096 ctx with 37% headroom.
Falcon 3 10B10BQ4_K_M11.3 GB8,192No measured row yetFits cleanly at Q4_K_M + 8,192 ctx with 19% headroom.
YTU Turkish Gemma 9B v0.19BQ4_K_M10.7 GB8,192No measured row yetFits cleanly at Q4_K_M + 8,192 ctx with 24% headroom.
NVIDIA Nemotron Nano 9B v2 Japanese9BQ4_K_M9.8 GB8,192No measured row yetFits cleanly at Q4_K_M + 8,192 ctx with 30% headroom.
Nemotron 3 Nano 9B9BQ4_K_M10.1 GB8,192No measured row yetFits cleanly at Q4_K_M + 8,192 ctx with 28% headroom.
Gemma 2 9B Instruct9BQ4_K_M10.6 GB8,192No measured row yetFits cleanly at Q4_K_M + 8,192 ctx with 24% headroom.
Turkish Gemma 9B T19BQ4_K_M9.8 GB8,192No measured row yetFits cleanly at Q4_K_M + 8,192 ctx with 30% headroom.
Yi Coder 9B9BQ4_K_M10.2 GB8,192No measured row yetFits cleanly at Q4_K_M + 8,192 ctx with 27% headroom.
GLM-4 9B9BQ4_K_M10.3 GB8,192No measured row yetFits cleanly at Q4_K_M + 8,192 ctx with 27% headroom.
Qwen3.5 9B9BQ4_K_M11.4 GB8,192No measured row yetFits cleanly at Q4_K_M + 8,192 ctx with 18% headroom.
Ornith 1.0 9B9BQ4_K_M10.4 GB8,192No measured row yetFits cleanly at Q4_K_M + 8,192 ctx with 26% headroom.
Qwen3.5 9B Thai Law Base9BQ4_K_M7.4 GB4,096No measured row yetComfortable fit with 47% headroom — room to extend context or run alongside other workloads.
Granite 4.1 8B Instruct9BQ5_K_M10.7 GB8,192No measured row yetFits cleanly at Q5_K_M + 8,192 ctx with 23% headroom.
Borderline
14 models · tight, may need quant downgrade
ModelParamsQuantVRAM est.ContextEvidenceNote
DeepSeek MoE 16B Base16BQ4_K_M14 GB4,096No measured row yetTight fit at Q4_K_M — only 0% headroom. KV cache for longer context will OOM. Cap context tighter or drop one quant level.
Phi-4 Reasoning 14B14BQ4_K_M15.8 GB8,192
40.44 tok/sMeasured hereNVIDIA GeForce RTX 3080 16GB (Mobile) / 4K ctx
Measured on this hardware: 40.44 tok/s at Q4_K_M (4,096 measured ctx). The larger target context still needs validation before calling it comfortable.
Qwen 3 14B14BQ4_K_M15.8 GB8,192
38.26 tok/sMeasured hereNVIDIA GeForce RTX 3080 16GB (Mobile) / 4K ctx
Measured on this hardware: 38.26 tok/s at Q4_K_M (4,096 measured ctx). The larger target context still needs validation before calling it comfortable.
Pixtral 12B12BQ4_K_M13.4 GB8,192No measured row yetTight fit at Q4_K_M — only 5% headroom. KV cache for longer context will OOM. Cap context tighter or drop one quant level.
Mistral Nemo 12B Instruct12BQ4_K_M13.9 GB8,192
65.72 tok/sMeasured hereNVIDIA GeForce RTX 3080 16GB (Mobile) / 4K ctx
Measured on this hardware: 65.72 tok/s at Q4_K_M (4,096 measured ctx). Tight fit at Q4_K_M — only 1% headroom. KV cache for longer context will OOM. Cap context tighter or drop one quant level.
Gemma 3 12B12BQ4_K_M13.7 GB8,192No measured row yetTight fit at Q4_K_M — only 2% headroom. KV cache for longer context will OOM. Cap context tighter or drop one quant level.
Gemma 4 12B12BQ4_K_M14 GB8,192No measured row yetTight fit at Q4_K_M — only 0% headroom. KV cache for longer context will OOM. Cap context tighter or drop one quant level.
Bielik 11B v2.2 Instruct GGUF11BQ4_K_M11.9 GB8,192No measured row yetTight fit at Q4_K_M — only 15% headroom. KV cache for longer context will OOM. Cap context tighter or drop one quant level.
Llama 3.2 11B Vision Instruct11BQ4_K_M13.8 GB8,192
67.00 tok/sMeasured hereNVIDIA GeForce RTX 3080 16GB (Mobile) / 4K ctx
Measured on this hardware: 67.00 tok/s at Q4_K_M (4,096 measured ctx). Tight fit at Q4_K_M — only 1% headroom. KV cache for longer context will OOM. Cap context tighter or drop one quant level.
Llama 3.2 11B Vision11BQ4_K_M12.3 GB8,192No measured row yetTight fit at Q4_K_M — only 12% headroom. KV cache for longer context will OOM. Cap context tighter or drop one quant level.
Bielik 11B v3.0 Instruct GGUF11BQ4_K_M11.9 GB8,192No measured row yetTight fit at Q4_K_M — only 15% headroom. KV cache for longer context will OOM. Cap context tighter or drop one quant level.
Hermes 3 Llama 3.1 8B8BQ8_012.9 GB8,192No measured row yetTight fit at Q8_0 — only 8% headroom. KV cache for longer context will OOM. Cap context tighter or drop one quant level.
Qwen 3 8B8BQ8_012.6 GB8,192No measured row yetTight fit at Q8_0 — only 10% headroom. KV cache for longer context will OOM. Cap context tighter or drop one quant level.
DeepSeek R1 Distill Qwen 7B7BQ8_012 GB8,192No measured row yetTight fit at Q8_0 — only 14% headroom. KV cache for longer context will OOM. Cap context tighter or drop one quant level.
Not practical
16 models · oversize for this build
ModelParamsQuantVRAM est.ContextEvidenceNote
NV-Embed v28BFP1619.7 GB8,192No measured row yet~19.7 GB needed at FP16 + 8,192 ctx — overshoots effective VRAM by 41%. Drop quant or move to a larger build.
Qwen 3 Embedding 8B8BFP1620.8 GB8,192No measured row yet~20.8 GB needed at FP16 + 8,192 ctx — overshoots effective VRAM by 49%. Drop quant or move to a larger build.
PaliGemma 2 10B10BBF1626 GB8,192No measured row yet~26.0 GB needed at BF16 + 8,192 ctx — overshoots effective VRAM by 86%. Drop quant or move to a larger build.
Mellum2 12B-A2.5B12BQ4_K_M14.5 GB8,192No measured row yet~14.5 GB needed at Q4_K_M + 8,192 ctx — overshoots effective VRAM by 3%. Drop quant or move to a larger build.
Baichuan 4 13B13BQ4_K_M14.7 GB8,192No measured row yet~14.7 GB needed at Q4_K_M + 8,192 ctx — overshoots effective VRAM by 5%. Drop quant or move to a larger build.
GLM-4V 9B14BQ4_K_M15.9 GB8,192No measured row yet~15.9 GB needed at Q4_K_M + 8,192 ctx — overshoots effective VRAM by 13%. Drop quant or move to a larger build.
Phi-4 14B14BQ4_K_M15.8 GB8,192No measured row yet~15.8 GB needed at Q4_K_M + 8,192 ctx — overshoots effective VRAM by 13%. Drop quant or move to a larger build.
Qwen 2.5 14B Instruct14BQ4_K_M16.4 GB8,192No measured row yet~16.4 GB needed at Q4_K_M + 8,192 ctx — overshoots effective VRAM by 17%. Drop quant or move to a larger build.
Qwen 2.5 Coder 14B Instruct14BQ4_K_M15.8 GB8,192No measured row yet~15.8 GB needed at Q4_K_M + 8,192 ctx — overshoots effective VRAM by 13%. Drop quant or move to a larger build.
DeepSeek R1 Distill Qwen 14B14BQ4_K_M15.8 GB8,192No measured row yet~15.8 GB needed at Q4_K_M + 8,192 ctx — overshoots effective VRAM by 13%. Drop quant or move to a larger build.
Phi-4 Multimodal14BQ4_K_M16.5 GB8,192No measured row yet~16.5 GB needed at Q4_K_M + 8,192 ctx — overshoots effective VRAM by 18%. Drop quant or move to a larger build.
Ministral 3 14B14BQ4_K_M16.6 GB8,192No measured row yet~16.6 GB needed at Q4_K_M + 8,192 ctx — overshoots effective VRAM by 18%. Drop quant or move to a larger build.
StarCoder 2 15B15BQ4_K_M17 GB8,192No measured row yet~17.0 GB needed at Q4_K_M + 8,192 ctx — overshoots effective VRAM by 21%. Drop quant or move to a larger build.
DeepSeek V2 Lite Chat16BQ4_K_M16.9 GB8,192No measured row yet~16.9 GB needed at Q4_K_M + 8,192 ctx — overshoots effective VRAM by 21%. Drop quant or move to a larger build.
DeepSeek V3 Lite (16B MoE)16BQ4_K_M18 GB8,192No measured row yet~18.0 GB needed at Q4_K_M + 8,192 ctx — overshoots effective VRAM by 28%. Drop quant or move to a larger build.
DeepSeek Coder V2 Lite (16B)16BQ4_K_M18 GB8,192No measured row yet~18.0 GB needed at Q4_K_M + 8,192 ctx — overshoots effective VRAM by 28%. Drop quant or move to a larger build.

Related

Multi-GPU buying guide →

NVLink vs PCIe, tensor- vs pipeline-parallel, mixed-card honesty.

Hardware combinations →

Curated multi-GPU / cluster setups with effective-VRAM math.

Setup path-finder →

OS + runtime install commands for your stack.

Compatibility matrix →

Runtime × OS × hardware support truth table.

Shopping a full build instead of a single card?

If you're sizing a fresh AI build (not just a card to drop into an existing system), the build-budget walkthroughs cover the whole BOM honestly: AI PC build under $1,000 or AI PC build under $2,000 cover the realistic 2026 budget tiers.

Vertical-fit shopping? AI PC for students covers the budget + portability tradeoffs; AI PC for developers covers the coding workflow specifics; AI PC for small business covers the document-RAG / always-on machine.

Form-factor first? See best laptop for local AI, best Mac for local AI, best mini PC for local AI, or best used GPU for local AI.

See something off?Submit a benchmark·Report outdated·Suggest a correctionWe read every submission. Editorial review takes 1-7 days.