Enter your GPU, RAM, and model target. The RunLocalAI Will-It-Run Framework shows what fits, how fast it should feel, when the number is measured, and what to upgrade next. Honest verdicts, auditable methodology, confidence labels on every benchmark.
Pick your hardware + a model, get a verdict in seconds.
Check fit →Every hardware unit ranked by RunLocalAI Score (0–1000).
See leaderboard →Tell us budget + OS + workload, get a GPU recommendation.
Get recommendation →Use case + budget + scale → GPU + runtime + models + install script. Three tiers side-by-side.
Open stack builder →Paste a prompt or code. See exactly what 11 cloud providers would charge vs your local rig. Break-even in months.
Open cost calculator →Q4_K_M vs Q5_K_M vs Q8 settled with math, not folklore. Quality curve + VRAM fit visualization.
Open quant advisor →Watch tokens stream at your hardware's actual speed, side-by-side with Claude / GPT-5 / Groq. Race mode.
Open stream visualizer →Paste a Claude Code .jsonl. See per-tool breakdown, cache savings, insights. Privacy-first.
Open decoder →Pick 2 models. Get a 10-row diff: params, context, license, TPS, fit, popularity. Use-case-weighted better-fit verdict.
Open battle card →Drag a tokens/day slider, watch the cloud vs local cost lines extend over your horizon. Crossover point drawn. One screenshot.
Open compounder →37 curated apps that plug into your local runtime: chat UIs, coding agents, RAG, voice, image, browser, mobile, editor plugins. Filter by VRAM + privacy.
Open apps directory →Every verdict has a named author and a real opinion. We say what to skip as much as what to buy. Some links are affiliate; the opinion comes first.
Benchmarks run with documented commands and open-weight models. Confidence tiers — measured / reproduced / inferred — visible on every result. /trust →
Privacy + cost + control beat cloud APIs for most everyday work. We'll tell you when local is the wrong answer too — /explained →
Most likely yes — the real question is which models. A GPU with 8-12 GB of VRAM, or any Apple Silicon Mac with 16 GB+ unified memory, runs 7B-14B-class models comfortably; larger models need more memory. Enter your GPU, RAM, and a target model for an exact verdict. the Will-It-Run checker →
Rule of thumb: a 7B model at 4-bit needs ~6 GB, a 32B ~18 GB, and a 70B ~40 GB. On unified-memory Macs, a chip comfortably loads models up to about (memory − 8) GB. Match your target model to your hardware before spending anything. the best-GPU guide →
Yes, for small models. 1-3B models run at conversational speed on an integrated GPU or a modern CPU; 7B is usable but slower. For 14B and up you want a discrete GPU or Apple Silicon with enough unified memory. the no-GPU path →
Often, for high-volume routine work — a one-time hardware cost replaces recurring per-token and subscription fees and pays back over months. Cloud APIs still win for frontier quality and day-one model access. Compare your own workload with the calculator. the cost calculator →
It depends on budget and the model you want to run. NVIDIA RTX cards give the best software support and speed per dollar; Apple Silicon wins on large unified memory for big models. Answer a few questions and get a specific pick. the GPU chooser →
Every benchmark row carries a confidence tier — measured, reproduced, or inferred — and the runs use documented commands and open-weight models. Nothing here is a marketing figure; you can see exactly how each number was produced. how we verify →