Will it run on your hardware?
Enter your GPU, RAM, and the model you want to run. Check the estimated memory fit, available performance evidence, and upgrade options. Use the methodology and benchmark confidence labels to distinguish measurements from estimates.
September 9 AI update · Markdown and agent access
Tell us what you're asking.
Pick your hardware + a model, get a verdict in seconds.
Check fit →Every hardware unit ranked by RunLocalAI Score (0–1000).
See leaderboard →Tell us budget + OS + workload, get a GPU recommendation.
Get recommendation →Use case + budget + scale → GPU + runtime + models + install script. Three tiers side-by-side.
Open stack builder →Paste a prompt or code. See exactly what 11 cloud providers would charge vs your local rig. Break-even in months.
Open cost calculator →Q4_K_M vs Q5_K_M vs Q8 settled with math, not folklore. Quality curve + VRAM fit visualization.
Open quant advisor →Watch tokens stream at your hardware's actual speed, side-by-side with Claude / GPT-5 / Groq. Race mode.
Open stream visualizer →Paste a Claude Code .jsonl. See per-tool breakdown, cache savings, insights. Privacy-first.
Open decoder →Pick 2 models. Get a 10-row diff: params, context, license, TPS, fit, popularity. Use-case-weighted better-fit verdict.
Open battle card →Drag a tokens/day slider, watch the cloud vs local cost lines extend over your horizon. Crossover point drawn. One screenshot.
Open compounder →37 curated apps that plug into your local runtime: chat UIs, coding agents, RAG, voice, image, browser, mobile, editor plugins. Filter by VRAM + privacy.
Open apps directory →Editor verdicts — recently updated.
Subroutines operators run.
Live stream of what changed.
Open-weight registry.
How this instrument operates.
Every verdict has a named author and a real opinion. We say what to skip as much as what to buy. Some links are affiliate; the opinion comes first.
Benchmarks run with documented commands and open-weight models. Confidence tiers — measured / reproduced / inferred — visible on every result. /trust →
Privacy + cost + control beat cloud APIs for most everyday work. We'll tell you when local is the wrong answer too — /explained →
Can your computer run local AI?
Can my computer run local AI?
Most likely yes — the real question is which models. A GPU with 8-12 GB of VRAM, or any Apple Silicon Mac with 16 GB+ unified memory, runs 7B-14B-class models comfortably; larger models need more memory. Enter your GPU, RAM, and a target model for an exact verdict. the Will-It-Run checker →
How much VRAM do I need to run a local LLM?
Rule of thumb: a 7B model at 4-bit needs ~6 GB, a 32B ~18 GB, and a 70B ~40 GB. On unified-memory Macs, a chip comfortably loads models up to about (memory − 8) GB. Match your target model to your hardware before spending anything. the best-GPU guide →
Can I run local AI without a dedicated GPU?
Yes, for small models. 1-3B models run at conversational speed on an integrated GPU or a modern CPU; 7B is usable but slower. For 14B and up you want a discrete GPU or Apple Silicon with enough unified memory. the no-GPU path →
Is running AI locally cheaper than cloud APIs?
It depends on your workload, hardware cost, electricity, maintenance and the cloud model you compare against. Use current provider prices and test whether local model quality is sufficient before estimating a payback period. the cost calculator →
What is the best GPU for local AI?
It depends on budget and the model you want to run. NVIDIA RTX cards give the best software support and speed per dollar; Apple Silicon wins on large unified memory for big models. Answer a few questions and get a specific pick. the GPU chooser →
How trustworthy are these benchmark numbers?
Check the source and confidence label on each benchmark. A measured run, a reproduced result and an estimate are different kinds of evidence. Compare hardware, runtime, model and quantization settings, and inspect the linked method or raw result before relying on a number. how we verify →