Custom build engine
Describe your build — any GPUs, CPU, RAM, OS, runtime, use case. We'll compute effective VRAM honestly, recommend a runtime, and tell you which models fit comfortably, which are borderline, and which aren't practical.
Total VRAM ≠ pooled VRAM. We never sum VRAM unless the silicon truly pools (Apple unified memory). We always explain why effective is lower than total.
Calculations follow the RunLocalAI Will-It-Run Framework: effective VRAM, model working set, runtime constraints, fit tiers, and measured-vs-estimated evidence labels.
Describe your build
Add GPUs, set CPU/RAM/OS, optionally pick a runtime + use case. URL updates as you change fields — share a build by copying the URL.
Build summary
Single Intel Arc Pro B60 24GB — 24 GB VRAM minus ~1.8 GB runtime/driver overhead = ~22 GB usable for weights + KV cache + activations. The remaining uncertainty band covers OS display use and background CUDA allocations.
Measured evidence on this hardware
Publicly inspectable measured rows for the selected hardware slug(s). Exact measured rows calibrate the fit table instead of leaving it as pure VRAM estimation.
No publicly inspectable benchmark rows are attached to this exact hardware yet. The engine will still calculate fit and runtime, but speed rows will remain estimated.
Coding agents — agentic tool-call burst
Workload-specific bottleneck. Where this kind of work actually breaks first, and what to budget for.
Coding agents emit 5-15 tool calls per task. Each call carries the full agent system prompt + context. KV-cache budget for that prompt × concurrent requests is the limit. The decode side is well-served by any modern card; the prefill side bottlenecks first.
- •32K context with KV-cache room to spare (~3-4 GB on 4090 AWQ-INT4)
- •Prefix cache: prefer SGLang for >5 tool calls / task
- •Decode latency: aim for >40 tok/s sustained
Models that fit your build
85 models considered (filtered by coding). Categorized by headroom at the recommended quant + a sensible context for your use case.
| Model | Params | Quant | VRAM est. | Context | Evidence | Note |
|---|---|---|---|---|---|---|
| Muse Glimmer 30B | 30B | Q4_K_M | 18.6 GB | 8,192 | No measured row yet | Fits cleanly at Q4_K_M + 8,192 ctx with 16% headroom. |
| DeepSeek Coder V2 Lite (16B) | 16B | Q4_K_M | 18 GB | 8,192 | No measured row yet | Fits cleanly at Q4_K_M + 8,192 ctx with 18% headroom. |
| DeepSeek V3 Lite (16B MoE) | 16B | Q4_K_M | 18 GB | 8,192 | No measured row yet | Fits cleanly at Q4_K_M + 8,192 ctx with 18% headroom. |
| StarCoder 2 15B | 15B | Q4_K_M | 17 GB | 8,192 | No measured row yet | Fits cleanly at Q4_K_M + 8,192 ctx with 23% headroom. |
| Qwen 2.5 Coder 14B Instruct | 14B | Q4_K_M | 15.8 GB | 8,192 | No measured row yet | Fits cleanly at Q4_K_M + 8,192 ctx with 28% headroom. |
| Ministral 3 14B | 14B | Q4_K_M | 16.6 GB | 8,192 | No measured row yet | Fits cleanly at Q4_K_M + 8,192 ctx with 25% headroom. |
| Qwen 2.5 14B Instruct | 14B | Q5_K_M | 18 GB | 8,192 | No measured row yet | Fits cleanly at Q5_K_M + 8,192 ctx with 18% headroom. |
| Qwen 3 14B | 14B | Q4_K_M | 15.8 GB | 8,192 | No measured row yet | Fits cleanly at Q4_K_M + 8,192 ctx with 28% headroom. |
| Mellum2 12B-A2.5B | 12B | Q4_K_M | 14.5 GB | 8,192 | No measured row yet | Fits cleanly at Q4_K_M + 8,192 ctx with 34% headroom. |
| Gemma 4 12B | 12B | Q4_K_M | 14 GB | 8,192 | No measured row yet | Fits cleanly at Q4_K_M + 8,192 ctx with 36% headroom. |
| Yi Coder 9B | 9B | Q4_K_M | 10.2 GB | 8,192 | No measured row yet | Comfortable fit with 54% headroom — room to extend context or run alongside other workloads. |
| Qwen3.5 9B | 9B | Q4_K_M | 11.4 GB | 8,192 | No measured row yet | Comfortable fit with 48% headroom — room to extend context or run alongside other workloads. |
| Ornith 1.0 9B | 9B | Q4_K_M | 10.4 GB | 8,192 | No measured row yet | Comfortable fit with 53% headroom — room to extend context or run alongside other workloads. |
| Granite 4.1 8B Instruct | 9B | Q8_0 | 14.2 GB | 8,192 | No measured row yet | Fits cleanly at Q8_0 + 8,192 ctx with 35% headroom. |
| OpenCoder 8B | 8B | Q4_K_M | 13 GB | 16,384 | No measured row yet | Comfortable fit with 41% headroom — room to extend context or run alongside other workloads. |
| Gervásio 8B PTPT | 8B | Q4_K_M | 6.6 GB | 4,096 | No measured row yet | Comfortable fit with 70% headroom — room to extend context or run alongside other workloads. |
| Dolphin 3.0 8B | 8B | Q4_K_M | 13.3 GB | 16,384 | No measured row yet | Fits cleanly at Q4_K_M + 16,384 ctx with 40% headroom. |
| DeepSeek R1 Distill Llama 8B | 8B | Q4_K_M | 13 GB | 16,384 | No measured row yet | Comfortable fit with 41% headroom — room to extend context or run alongside other workloads. |
| Qwen 3 8B | 8B | Q8_0 | 16.6 GB | 16,384 | No measured row yet | Fits cleanly at Q8_0 + 16,384 ctx with 24% headroom. |
| EXAONE Deep 7.8B | 8B | Q4_K_M | 12.3 GB | 16,384 | No measured row yet | Comfortable fit with 44% headroom — room to extend context or run alongside other workloads. |
| Qwen 2.5 Coder 7B Instruct | 7B | Q6_K | 13.6 GB | 16,384 | No measured row yet | Fits cleanly at Q6_K + 16,384 ctx with 38% headroom. |
| CodeQwen 1.5 7B | 7B | Q4_K_M | 11.6 GB | 16,384 | No measured row yet | Comfortable fit with 47% headroom — room to extend context or run alongside other workloads. |
| StarCoder 2 7B | 7B | Q4_K_M | 11.6 GB | 16,384 | No measured row yet | Comfortable fit with 47% headroom — room to extend context or run alongside other workloads. |
| Salamandra 7B | 7B | Q4_K_M | 7.6 GB | 8,192 | No measured row yet | Comfortable fit with 65% headroom — room to extend context or run alongside other workloads. |
| Model | Params | Quant | VRAM est. | Context | Evidence | Note |
|---|---|---|---|---|---|---|
| Llama 3.1 8B Instruct | 8B | FP16 | 19.1 GB | 16,384 | No measured row yet | Tight fit at FP16 — only 13% headroom. KV cache for longer context will OOM. Cap context tighter or drop one quant level. |
| Model | Params | Quant | VRAM est. | Context | Evidence | Note |
|---|---|---|---|---|---|---|
| GPT-OSS 20B | 21B | MXFP4 | 25.2 GB | 8,192 | No measured row yet | ~25.2 GB needed at MXFP4 + 8,192 ctx — overshoots effective VRAM by 14%. Drop quant or move to a larger build. |
| Codestral 22B | 22B | Q4_K_M | 24.7 GB | 8,192 | No measured row yet | ~24.7 GB needed at Q4_K_M + 8,192 ctx — overshoots effective VRAM by 12%. Drop quant or move to a larger build. |
| Devstral Small 2 24B | 24B | Q4_K_M | 26.7 GB | 8,192 | No measured row yet | ~26.7 GB needed at Q4_K_M + 8,192 ctx — overshoots effective VRAM by 21%. Drop quant or move to a larger build. |
| Mistral Small 3 24B | 24B | Q4_K_M | 26.7 GB | 8,192 | No measured row yet | ~26.7 GB needed at Q4_K_M + 8,192 ctx — overshoots effective VRAM by 21%. Drop quant or move to a larger build. |
| Gemma 4 26B-A4B | 26B | Q4_K_M | 29.8 GB | 8,192 | No measured row yet | ~29.8 GB needed at Q4_K_M + 8,192 ctx — overshoots effective VRAM by 35%. Drop quant or move to a larger build. |
| Qwen 3.6 27B (MTP) | 27B | Q3_K_M | 27.3 GB | 8,192 | No measured row yet | ~27.3 GB needed at Q3_K_M + 8,192 ctx — overshoots effective VRAM by 24%. Drop quant or move to a larger build. |
| Qwen3.5 27B | 27B | Q4_K_M | 31.4 GB | 8,192 | No measured row yet | ~31.4 GB needed at Q4_K_M + 8,192 ctx — overshoots effective VRAM by 43%. Drop quant or move to a larger build. |
| Qwen3.6 27B | 27B | Q4_K_M | 31.4 GB | 8,192 | No measured row yet | ~31.4 GB needed at Q4_K_M + 8,192 ctx — overshoots effective VRAM by 43%. Drop quant or move to a larger build. |
| Granite 4.1 30B Instruct | 29B | Q3_K_M | 29.2 GB | 8,192 | No measured row yet | ~29.2 GB needed at Q3_K_M + 8,192 ctx — overshoots effective VRAM by 33%. Drop quant or move to a larger build. |
| Sarvam 30B | 30B | Q4_K_M | 24.8 GB | 4,096 | No measured row yet | ~24.8 GB needed at Q4_K_M + 4,096 ctx — overshoots effective VRAM by 13%. Drop quant or move to a larger build. |
| Granite 4.1 30B | 30B | Q4_K_M | 32.9 GB | 8,192 | No measured row yet | ~32.9 GB needed at Q4_K_M + 8,192 ctx — overshoots effective VRAM by 49%. Drop quant or move to a larger build. |
| Qwen3 Coder 30B-A3B | 30B | Q4_K_M | 35 GB | 8,192 | No measured row yet | ~35.0 GB needed at Q4_K_M + 8,192 ctx — overshoots effective VRAM by 59%. Drop quant or move to a larger build. |
| North Mini Code 1.0 | 30B | Q4_K_M | 35 GB | 8,192 | No measured row yet | ~35.0 GB needed at Q4_K_M + 8,192 ctx — overshoots effective VRAM by 59%. Drop quant or move to a larger build. |
| Qwen 3 30B-A3B | 30B | Q4_K_M | 33.9 GB | 8,192 | No measured row yet | ~33.9 GB needed at Q4_K_M + 8,192 ctx — overshoots effective VRAM by 54%. Drop quant or move to a larger build. |
| GLM-4.7-Flash | 31B | Q4_K_M | 35.5 GB | 8,192 | No measured row yet | ~35.5 GB needed at Q4_K_M + 8,192 ctx — overshoots effective VRAM by 61%. Drop quant or move to a larger build. |
| Gemma 4 31B Dense | 31B | Q4_K_M | 34.4 GB | 8,192 | No measured row yet | ~34.4 GB needed at Q4_K_M + 8,192 ctx — overshoots effective VRAM by 56%. Drop quant or move to a larger build. |
Related
NVLink vs PCIe, tensor- vs pipeline-parallel, mixed-card honesty.
Curated multi-GPU / cluster setups with effective-VRAM math.
OS + runtime install commands for your stack.
Runtime × OS × hardware support truth table.
Shopping a full build instead of a single card?
If you're sizing a fresh AI build (not just a card to drop into an existing system), the build-budget walkthroughs cover the whole BOM honestly: AI PC build under $1,000 or AI PC build under $2,000 cover the realistic 2026 budget tiers.
Vertical-fit shopping? AI PC for students covers the budget + portability tradeoffs; AI PC for developers covers the coding workflow specifics; AI PC for small business covers the document-RAG / always-on machine.
Form-factor first? See best laptop for local AI, best Mac for local AI, best mini PC for local AI, or best used GPU for local AI.