UNIT · NVIDIA · GPU
32 GB VRAMenthusiastReviewed June 2026

NVIDIA GeForce RTX 5090

RTX 5090 spec card — 32 GB VRAM, 1.79 TB/s bandwidth, 575 W; best for 70B Q4 + 8K context
diagram
Credit: RunLocalAI·License: CC-BY-4.0 (original illustration)·Source

Blackwell flagship with 32 GB GDDR7 and about 1.79 TB/s memory bandwidth. More context headroom than a 24 GB card for 32B Q4; 70B Q4 still requires offload or more GPU memory.

Released 2025·~$2499 editorial estimate · date unavailable·1792 GB/s memory bandwidth
▼ CHECK CURRENT PRICE· 1 retailer
NVIDIA GeForce RTX 5090

Affiliate disclosure: as an Amazon Associate and partner of other retailers, we earn from qualifying purchases. The verdict on this page is our editorial opinion; affiliate links never influence what we recommend.

RUNLOCALAI SCORE
See full leaderboard →
630/ 1000
BB-tier
Estimated
Throughput
500/ 500
VRAM-fit
170/ 200
Ecosystem
200/ 200
Efficiency
30/ 100

Sub-scores sum to 900 / 1000. Headline = 900 × 0.70 (Estimated-confidence discount) = 630. This is an algorithmic performance-tier score — distinct from, and often lower than, the editorial “Our verdict” below, which weighs value and real-world fit (especially for hardware we haven’t measured yet). How scoring works →

Extrapolated from 1792 GB/s bandwidth — 215.0 tok/s estimated. No measured benchmarks yet.

Plain-English: Comfortable at 32B and below — snappy enough for a coding agent; vision models supported.

7B chat
Comfortable
14B chat
Comfortable
32B chat
Comfortable
70B chat
Doesn't fit
Coding agent
Comfortable
Vision (≤8B VLM)
Comfortable
Long context (32K)
Comfortable
Comfortable — fits with headroom
~Tight — works, no slack
Marginal — needs aggressive quant
Doesn't fit usefully

Verdicts extrapolated from catalog VRAM + bandwidth + ecosystem flags. Hover any chip for the rationale. Want measured numbers? Submit your own run with runlocalai-bench --submit.

BLK · VERDICT

Our verdict

OP · Eruo Fredoline|VERIFIED JUN 12, 2026
9.6/10

What it does well

The RTX 5090 provides 32 GB GDDR7 at about 1.79 TB/s. It gives a 32B Q4 model more context headroom than a 24 GB GPU. It cannot fully hold 70B Q4_K_M weights (about 42.3 GB) or 32B FP16 weights (64 GB). Compare published benchmark configurations for speed; the earlier unlinked timing ranges have been removed.

Where it breaks

  • 575 W TGP is real. This is a 1000 W+ PSU card, not a 750 W card. Add headroom for CPU + drives + transient spikes; many operators end up at 1200 W. The 12V-2x6 connector replaces the controversial 4090-era 12VHPWR but the fitment + power-budget caution stays.
  • Supply + price are not normal yet. 2025-into-2026 retail is supply-constrained — MSRP $1,999 is rarely the price you actually pay. Scalper-adjacent pricing of $2,300–2,800 is the operator-grade reality.
  • 32B-class workloads are over-spec. If your daily target is Qwen 3 32B / Qwen 2.5 Coder 32B / QwQ 32B, the 5090 isn't doing more for you than a 4090 does — the workload fits 24 GB. You're paying the 5090 premium for headroom you don't need.
  • Multi-GPU economics are awkward. Two 5090s for ~$5,000 buy you 64 GB combined VRAM. Two used 3090s buy you 48 GB combined for ~$1,800. For homelab operators chasing $/VRAM, the calculus often favors the older silicon.

Ideal model range

  • Single-card starting point: 32B Q4 with more memory headroom than a 24 GB card.
  • 70B Q4: about 42.3 GB of weights before runtime and KV cache. A single RTX-5090 needs offload. Two 24 GB cards can be viable with a supported split and a constrained context budget.
  • FP16: 32B needs about 64 GB of weights; 70B needs about 140 GB. Reserve additional space before selecting a memory tier.
  • Concurrency and long context: measure the actual model, cache dtype and runtime. No fixed context or timing guarantee is made here.

Bad use cases

  • Genuine frontier-MoE workloads — DeepSeek V3 671B, Llama 4 Maverick / Behemoth — need workstation hardware (RTX 6000 Ada / RTX PRO 6000 Blackwell) or multi-GPU. 32 GB doesn't change that math.
  • Power-constrained builds — mini-ITX cases, 750 W PSUs, anyone running 24/7 inference and paying retail electricity. The 5090 is a thermal + power statement.
  • Maximum tok/s on small models — 7B at >~300 tok/s is throughput territory where smaller cards (RTX 5070 Ti, RTX 4070 Super) are better $/throughput. The 5090 is over-specced for sub-13B workloads.
  • Anyone betting on future supply normalization — if the 5090 is still scalper-priced when you check, the 4090 used market and the dual-3090 path are honest alternatives. Don't pay 30%+ premium for a card you can wait on.

Verdict

Consider this card if you have a verified 32B quantized workload that benefits from more than 24 GB, and the price, power and cooling requirements fit your build. Check the actual artifact and context in the hardware checker.

Choose another configuration if full-GPU 70B Q4 or 32B FP16 is the requirement. Consider more GPU memory or an explicitly supported multi-GPU split. A used dual-3090 setup has 48 GB nominal memory, with additional setup and power requirements.

How it compares

  • RTX 3090 / RTX 4090: both provide 24 GB. The 4090 has higher specified bandwidth; a matched benchmark is needed for a throughput comparison.
  • RTX 5090: 32 GB and about 1.79 TB/s. It adds context headroom for 32B Q4, but 70B Q4_K_M still exceeds its memory.
  • Two RTX 3090s: 48 GB nominal memory with a supported split. This can be viable for 70B Q4, with limited context room and additional system, power and cooling requirements.
  • 16 GB GPUs: start by checking 13B–14B Q4 models. A 32B Q4_K_M artifact is about 19.3 GB before overhead.
  • Apple M4 Max 128 GB: more total memory, shared with the OS. It does not hold 70B FP16 weights of about 140 GB. Consider Q4/Q8 and check available unified memory. A 192 GB or larger configuration provides a different FP16 budget.
  • Workstation GPUs: 48–96 GB on one card can simplify larger model allocations, at a different price and power level.
BLK · OVERVIEW

Overview

What the RTX 5090 actually is, in local-AI terms

The RTX 5090 has 32 GB of GDDR7 at about 1.79 TB/s and Blackwell hardware acceleration. Compared with the RTX 4090, it adds 8 GB of memory. That can provide more context room for a 32B quantized model, but it does not make a 70B Q4 model fully GPU-resident.

A 70B model needs at least 35 GB of raw 4-bit weights, and approximately 42.3 GB at Q4_K_M, before KV cache and runtime. Published benchmark configurations should establish speed comparisons; bandwidth alone is not a measurement.

Check current prices, available power and case clearance before choosing the 5090 over a 24 GB card.

Where it fits in the hardware ladder

In the consumer-NVIDIA tier:

Card VRAM BW Bin
RTX 4090 24 GB 1008 GB/s workstation default through 2025
RTX 5090 32 GB 1792 GB/s consumer flagship 2026
RTX Pro 6000 Blackwell 96 GB ~1.8 TB/s workstation tier above 5090

vs the datacenter ladder:

Card VRAM BW Notes
RTX 5090 32 GB 1.79 TB/s consumer; no NVLink
H100 SXM 80 GB 3.35 TB/s datacenter; NVLink
H200 141 GB 4.8 TB/s datacenter capacity tier

The 5090 has 32 GB on one card. This is useful headroom for 32B Q4, but both 70B INT4 and 405B Q4 exceed that capacity before context and runtime.

Best use cases

  • 32B Q4 coding and chat with more context headroom. Qwen 2.5 Coder 32B at Q4 is a candidate; FP16 weights need about 64 GB and cannot fit on this card.
  • FP4 experimentation. Use an engine that supports the model and Blackwell kernels. Faster arithmetic does not reduce the memory needed to store all model weights.
  • Training experiments. Check optimizer, activation and batch memory separately from inference.
  • Image pipelines. Check the combined footprint of the image model, resolution and any concurrent language model.

70B Q4 requires offload or a larger GPU memory configuration; a single 5090 has 32 GB.

What it can run

These are weight-budget estimates, not measured context limits. Use the hardware checker for a specific model and configuration.

Model class Quant Memory planning Notes
7B FP16 About 14 GB weights Add model-specific KV cache and runtime
13B–14B FP16 About 26–28 GB weights Limited remaining headroom; do not assume 64K context
32B Q4_K_M About 19.3 GB weights More headroom than 24 GB, but context still needs checking
32B FP16 / Q8 At least 64 / 32 GB weights Does not fit fully in 32 GB with runtime and context
70B Q4_K_M About 42.3 GB weights Requires offload or more GPU memory
70B 4-bit FP4 At least 35 GB weights before scales Hardware FP4 support does not make 70B fit in 32 GB
405B Q4_K_M About 245 GB weights Requires a substantially larger memory configuration

OS support

OS Quality
Linux (Ubuntu 24.04 LTS) excellent — reference
Windows 11 native excellent
Windows (WSL2) excellent
macOS unsupported

If your CUDA path is broken on WSL2, see /errors/wsl2-gpu-not-detected.

Software / runtime support

The 5090's Blackwell architecture is supported across the leading-edge inference engines, with the caveat that engine support for FP4 lags hardware availability through 2026:

  • Ollama / llama.cpp — full GGUF + CUDA; FP4 lands incrementally
  • vLLM — full AWQ / GPTQ / FP16 / FP8; FP4 maturing through 2026
  • SGLang — same coverage as vLLM
  • ExLlamaV2 — single-stream throughput king on this hardware via TabbyAPI
  • TensorRT-LLM — first-class; FP4 path the most mature here
  • LM Studio — full GUI path with CUDA acceleration
  • PyTorch — first-class CUDA target

What breaks first

  1. Power delivery. The 5090 pulls up to 575 W under sustained inference; the 12V-2x6 connector + cheap PSUs is a known fire-and-instability path. Pair with a Platinum-rated 1200 W+ PSU and high-quality cabling.
  2. Thermals in compact cases. 575 W of dissipation is a real cooling problem; small-form-factor builds throttle quickly without aggressive airflow.
  3. CUDA toolkit / driver lag for FP4. Engines are still catching up; expect a 6-12 month tail of "this engine doesn't yet use the 5090's FP4 path" through 2026.
  4. PCIe Gen5 x16 dependency. The 5090 wants Gen5 bandwidth for prefill on long contexts; older Gen4 boards still work but are bandwidth-limited.
  5. Multi-GPU absence of NVLink. Like the 4090, no NVLink — multi-card is PCIe only.

Alternatives by intent

If you want… Reach for
Cheaper, same-tier consumer RTX 5080 (16 GB) or used RTX 4090
Even more VRAM RTX Pro 6000 Blackwell (96 GB) or RTX A6000 (48 GB)
70B FP16 single-machine Apple M3 Ultra 192 GB unified memory
AMD path RX 9070 XT — much cheaper, ROCm tax applies
Datacenter throughput H100 SXM or H200

Best pairings

  • Ollama or llama.cpp + 32B Q4 — check artifact size and context before loading.
  • vLLM + a supported quantized model — reserve KV-cache space for concurrency.
  • TensorRT-LLM — verify model and FP4 kernel support in the installed version.
  • A compatible driver, CUDA and runtime combination — record versions when benchmarking.
  • A coding client connected to the chosen runtime — see the local coding-agent stack.

Who should avoid the RTX 5090

  • Operators happy with 24 GB. A 4090 is dramatically better dollars-per-token; the 5090 only wins when 32 GB or FP4 matter.
  • Anyone on a sub-1200 W PSU. Power delivery is non-negotiable.
  • Compact-case builders without aggressive cooling. 575 W is a real thermal problem.
  • Apple-ecosystem operators. Different stack entirely.
  • Workloads where 13B-class models suffice. A 16 GB card saves ~$2000 at the same tier of usefulness.

Related

Retailers we'd check:Amazon

Some links above are affiliate links. We may earn a commission at no extra cost to you. How we make money.

BLK · SPECS

Specs

VRAM32 GB
Power draw (peak)575 W
Released2025
MSRP$1999
Backends
CUDA
Vulkan

Models that fit

Open-weight models small enough to run on NVIDIA GeForce RTX 5090 with usable context.

Compare alternatives

Hardware worth comparing

The closest alternatives by price, memory bandwidth, and form factor, plus a step up and down — so you can frame the buying decision against real options.

Buyer guides where this card is the right answer

The 5090 only justifies its price for buyers who specifically need 32 GB on one card or are running production image/video gen. The guides below cover those workloads.

Frequently asked

What models can NVIDIA GeForce RTX 5090 run?

With 32GB VRAM, the NVIDIA GeForce RTX 5090 runs models up to ~32B in 4-bit, with room for context. See the model list below for tested combinations.

Does NVIDIA GeForce RTX 5090 support CUDA?

Yes — NVIDIA GeForce RTX 5090 is an NVIDIA card with full CUDA support, the most mature local-AI backend. llama.cpp, Ollama, vLLM, and ExLlamaV2 all run natively.

How much does NVIDIA GeForce RTX 5090 cost?

Editorial price estimate for NVIDIA GeForce RTX 5090: $2499 (MSRP $1999). Price check date unavailable. Verify current retailer price and availability.

Where next?

Reviewed by RunLocalAI Editorial. See our editorial policy for how we research and verify hardware specifications.