Custom build engine

Describe your build — any GPUs, CPU, RAM, OS, runtime, use case. We'll compute effective VRAM honestly, recommend a runtime, and tell you which models fit comfortably, which are borderline, and which aren't practical.

Total VRAM ≠ pooled VRAM. We never sum VRAM unless the silicon truly pools (Apple unified memory). We always explain why effective is lower than total.

Calculations follow the RunLocalAI Will-It-Run Framework: effective VRAM, model working set, runtime constraints, fit tiers, and measured-vs-estimated evidence labels.

Describe your build

Add GPUs, set CPU/RAM/OS, optionally pick a runtime + use case. URL updates as you change fields — share a build by copying the URL.

Build summary

Total VRAM
0 GB
Effective VRAM
~0 GB
range 0-0 GB
Topology
apple cluster
thunderbolt
Setup difficulty
advanced
speed penalty ~60%

Measured evidence on this hardware

Publicly inspectable measured rows for the selected hardware slug(s). Exact measured rows calibrate the fit table instead of leaving it as pure VRAM estimation.

No publicly inspectable benchmark rows are attached to this exact hardware yet. The engine will still calculate fit and runtime, but speed rows will remain estimated.

WORKLOAD PROFILE
OVERFLOW
gpt2-base-french @ Q4_K_M, 1K context on Apple M4 Pro
0 GB0.0 GBVRAM ceiling
Weights0.1 GB
KV cache0.0 GB
Activations0.0 GB
Runtime0.7 GB
Overflow0.8 GB
ESTIMATED DECODE RATE
1502 tok/s
Bandwidth-derived estimate · efficiency 0.55. Real-world rates land within ±20% on well-tuned runtimes.
1502 tokens per second02550100150

Models that fit your build

313 models considered. Categorized by headroom at the recommended quant + a sensible context for your use case.

Comfortable
0 models · ≥15% headroom

No model fits comfortably on this build.

Borderline
0 models · tight, may need quant downgrade

No borderline models — clean fit ladder.

Not practical
16 models · oversize for this build
ModelParamsQuantVRAM est.ContextEvidenceNote
gpt2-base-french0BQ4_K_M0.1 GB1,024No measured row yet~0.1 GB needed at Q4_K_M + 1,024 ctx — overshoots effective VRAM by Infinity%. Drop quant or move to a larger build.
GPT-2 Spanish0BQ4_K_M0.1 GB1,024No measured row yet~0.1 GB needed at Q4_K_M + 1,024 ctx — overshoots effective VRAM by Infinity%. Drop quant or move to a larger build.
SmolLM2 135M Instruct0BQ4_K_M0.2 GB8,192No measured row yet~0.2 GB needed at Q4_K_M + 8,192 ctx — overshoots effective VRAM by Infinity%. Drop quant or move to a larger build.
Dostoevsky Doesn't Write It GPT20BQ4_K_M0.1 GB1,024No measured row yet~0.1 GB needed at Q4_K_M + 1,024 ctx — overshoots effective VRAM by Infinity%. Drop quant or move to a larger build.
LFM2.5-230M0BQ3_K_M0.2 GB8,192No measured row yet~0.2 GB needed at Q3_K_M + 8,192 ctx — overshoots effective VRAM by Infinity%. Drop quant or move to a larger build.
Gemma 3 270M0BQ4_K_M0.3 GB8,192No measured row yet~0.3 GB needed at Q4_K_M + 8,192 ctx — overshoots effective VRAM by Infinity%. Drop quant or move to a larger build.
GPT-2 Spanish Medium0BQ4_K_M0.2 GB1,024No measured row yet~0.2 GB needed at Q4_K_M + 1,024 ctx — overshoots effective VRAM by Infinity%. Drop quant or move to a larger build.
SmolLM 2 360M Instruct0BQ4_K_M0.5 GB8,192No measured row yet~0.5 GB needed at Q4_K_M + 8,192 ctx — overshoots effective VRAM by Infinity%. Drop quant or move to a larger build.
SmolLM2 360M Instruct0BQ4_K_M0.4 GB8,192No measured row yet~0.4 GB needed at Q4_K_M + 8,192 ctx — overshoots effective VRAM by Infinity%. Drop quant or move to a larger build.
VBART Large (Turkish Summarization)0BQ4_K_M0.2 GB1,024No measured row yet~0.2 GB needed at Q4_K_M + 1,024 ctx — overshoots effective VRAM by Infinity%. Drop quant or move to a larger build.
SigLIP SO400M (patch14-384)0BQ4_K_M0.3 GB0No measured row yet~0.3 GB needed at Q4_K_M + 0 ctx — overshoots effective VRAM by Infinity%. Drop quant or move to a larger build.
Qwen 2.5 0.5B Instruct1BQ4_K_M0.7 GB8,192No measured row yet~0.7 GB needed at Q4_K_M + 8,192 ctx — overshoots effective VRAM by Infinity%. Drop quant or move to a larger build.
Vikhr Qwen 2.5 0.5B Instruct1BQ4_K_M0.4 GB4,096No measured row yet~0.4 GB needed at Q4_K_M + 4,096 ctx — overshoots effective VRAM by Infinity%. Drop quant or move to a larger build.
GOT-OCR 2.01BQ4_K_M0.3 GB0No measured row yet~0.3 GB needed at Q4_K_M + 0 ctx — overshoots effective VRAM by Infinity%. Drop quant or move to a larger build.
Qwen3 0.6B Hindi Instruct v1 GGUF1BQ4_K_M0.4 GB2,048No measured row yet~0.4 GB needed at Q4_K_M + 2,048 ctx — overshoots effective VRAM by Infinity%. Drop quant or move to a larger build.
Qwen 3 0.6B1BQ4_K_M0.6 GB8,192No measured row yet~0.6 GB needed at Q4_K_M + 8,192 ctx — overshoots effective VRAM by Infinity%. Drop quant or move to a larger build.

Shopping a full build instead of a single card?

If you're sizing a fresh AI build (not just a card to drop into an existing system), the build-budget walkthroughs cover the whole BOM honestly: AI PC build under $1,000 or AI PC build under $2,000 cover the realistic 2026 budget tiers.

Vertical-fit shopping? AI PC for students covers the budget + portability tradeoffs; AI PC for developers covers the coding workflow specifics; AI PC for small business covers the document-RAG / always-on machine.

Form-factor first? See best laptop for local AI, best Mac for local AI, best mini PC for local AI, or best used GPU for local AI.