What can NVIDIA GeForce RTX 5090 run for vision?

Build: RTX 5090 + Ryzen 9 9950X + 64GB DDR5

Memory: 32 GB VRAM + 64 GB system RAM
Runner: llama.cpp / Ollama (CUDA)

Runs comfortably
29 models

Ranked by fit for vision use case + predicted speed. Click a row for VRAM breakdown.

#1Gemma 4 12B
12B
gemma
Commercial OK
Quant: Q4_K_MContext: 8,192VRAM: 15.8 GBHeadroom: 16.2 GBTTFT: fast
ollama run gemma4:12b
161
tok/s
Estimated
Weights
7.60 GB
KV cache
6.00 GB
Activations
0.39 GB
Runtime
1.80 GB
Time to first token (prefill, 512-token prompt): ~246 ms (fast)
#2Qwen3.5 9B
9B
qwen
Commercial OK
Quant: Q4_K_MContext: 8,192VRAM: 13.2 GBHeadroom: 18.8 GBTTFT: fast
ollama run qwen3.5:9b
214
tok/s
Estimated
Weights
6.60 GB
KV cache
4.50 GB
Activations
0.34 GB
Runtime
1.80 GB
Time to first token (prefill, 512-token prompt): ~184 ms (fast)
#3Gemma 4 E4B (Effective 4B)
4B
gemma
Commercial OK
Quant: Q8_0Context: 8,192VRAM: 8.4 GBHeadroom: 23.6 GBTTFT: instant
ollama run gemma4:e4b
274
tok/s
Estimated
Weights
4.40 GB
KV cache
2.00 GB
Activations
0.23 GB
Runtime
1.80 GB
Time to first token (prefill, 512-token prompt): ~82 ms (instant)
#4Ministral 3 14B
14B
mistral
Commercial OK
Quant: Q4_K_MContext: 8,192VRAM: 18.4 GBHeadroom: 13.6 GBTTFT: fast
ollama run ministral-3:14b
138
tok/s
Estimated
Weights
9.10 GB
KV cache
7.00 GB
Activations
0.46 GB
Runtime
1.80 GB
Time to first token (prefill, 512-token prompt): ~287 ms (fast)
#5Gemma 3 4B
4B
gemma
Commercial OK
Quant: Q8_0Context: 8,192VRAM: 8.4 GBHeadroom: 23.6 GBTTFT: instant
ollama run gemma3:4b
274
tok/s
Estimated
Weights
4.40 GB
KV cache
2.00 GB
Activations
0.23 GB
Runtime
1.80 GB
Time to first token (prefill, 512-token prompt): ~82 ms (instant)
#6Gemma 4 E2B (Effective 2B)
2B
gemma
Commercial OK
Quant: Q8_0Context: 8,192VRAM: 5.1 GBHeadroom: 26.9 GBTTFT: instant
ollama run gemma4:e2b
548
tok/s
Estimated
Weights
2.20 GB
KV cache
1.00 GB
Activations
0.12 GB
Runtime
1.80 GB
Time to first token (prefill, 512-token prompt): ~41 ms (instant)
#7Phi-3.5 Vision
4.2B
phi
Commercial OK
Quant: Q4_K_MContext: 8,192VRAM: 6.5 GBHeadroom: 25.5 GBTTFT: instant
459
tok/s
Estimated
Weights
2.50 GB
KV cache
2.10 GB
Activations
0.13 GB
Runtime
1.80 GB
Time to first token (prefill, 512-token prompt): ~86 ms (instant)
#8PaliGemma 2 3B
3B
gemma
Commercial OK
Quant: BF16Context: 8,192VRAM: 9.6 GBHeadroom: 22.4 GBTTFT: instant
194
tok/s
Estimated
Weights
6.00 GB
KV cache
1.50 GB
Activations
0.31 GB
Runtime
1.80 GB
Time to first token (prefill, 512-token prompt): ~61 ms (instant)
#9LLaVA 1.6 Mistral 7B
7B
other
Commercial OK
Quant: Q4_K_MContext: 8,192VRAM: 10.0 GBHeadroom: 22.0 GBTTFT: fast
276
tok/s
Estimated
Weights
4.50 GB
KV cache
3.50 GB
Activations
0.23 GB
Runtime
1.80 GB
Time to first token (prefill, 512-token prompt): ~143 ms (fast)
#10Phi-4 Multimodal
14B
phi
Commercial OK
Quant: Q4_K_MContext: 8,192VRAM: 18.3 GBHeadroom: 13.7 GBTTFT: fast
138
tok/s
Estimated
Weights
9.00 GB
KV cache
7.00 GB
Activations
0.46 GB
Runtime
1.80 GB
Time to first token (prefill, 512-token prompt): ~287 ms (fast)
#11LLaVA-OneVision 7B
7B
other
Commercial OK
Quant: Q4_K_MContext: 8,192VRAM: 10.0 GBHeadroom: 22.0 GBTTFT: fast
276
tok/s
Estimated
Weights
4.50 GB
KV cache
3.50 GB
Activations
0.23 GB
Runtime
1.80 GB
Time to first token (prefill, 512-token prompt): ~143 ms (fast)
#12Moondream 2
1.9B
other
Commercial OK
Quant: Q4_K_MContext: 2,048VRAM: 3.3 GBHeadroom: 28.7 GBTTFT: instant
1015
tok/s
Estimated
Weights
1.20 GB
KV cache
0.24 GB
Activations
0.06 GB
Runtime
1.80 GB
Time to first token (prefill, 512-token prompt): ~39 ms (instant)

Runs with tradeoffs
8 models

Tight VRAM, partial CPU offload, or context-limited.

Gemma 4 26B MoE
26B
gemma
Commercial OK
Quant: Q4_K_MContext: 8,192VRAM: 31.6 GBHeadroom: 0.4 GBTTFT: noticeable
  • Tight VRAM fit — only 0.4 GB headroom left for context growth
ollama run gemma4:26b-moe
74
tok/s
Estimated
Weights
16.00 GB
KV cache
13.00 GB
Activations
0.81 GB
Runtime
1.80 GB
Time to first token (prefill, 512-token prompt): ~532 ms (noticeable)
InternVL 2.5 26B
26B
other
Commercial OK
Quant: Q4_K_MContext: 8,192VRAM: 31.6 GBHeadroom: 0.4 GBTTFT: noticeable
  • Tight VRAM fit — only 0.4 GB headroom left for context growth
74
tok/s
Estimated
Weights
16.00 GB
KV cache
13.00 GB
Activations
0.81 GB
Runtime
1.80 GB
Time to first token (prefill, 512-token prompt): ~532 ms (noticeable)
Nemotron 3 Nano Omni 33B
33B
other
Commercial OK
Quant: Q4_K_MContext: 8,192VRAM: 47.7 GBHeadroom: 22.7 GBTTFT: noticeable
  • Partial CPU offload: ~33% of layers run on CPU
  • CPU is the bottleneck — upgrading RAM bandwidth helps more than VRAM here
ollama run nemotron3:33b
5
tok/s
Estimated
Weights
28.00 GB
KV cache
16.50 GB
Activations
1.41 GB
Runtime
1.80 GB
Time to first token (prefill, 512-token prompt): ~676 ms (noticeable)
Molmo 72B
72B
other
Commercial OK
Quant: Q4_K_MContext: 4,096VRAM: 62.9 GBHeadroom: 7.5 GBTTFT: noticeable
  • Partial CPU offload: ~49% of layers run on CPU
  • CPU is the bottleneck — upgrading RAM bandwidth helps more than VRAM here
2
tok/s
Estimated
Weights
41.00 GB
KV cache
18.00 GB
Activations
2.05 GB
Runtime
1.80 GB
Time to first token (prefill, 512-token prompt): ~1475 ms (noticeable)
InternVL 2.5 78B
78B
other
Commercial OK
Quant: Q4_K_MContext: 2,048VRAM: 58.8 GBHeadroom: 11.6 GBTTFT: noticeable
  • Partial CPU offload: ~46% of layers run on CPU
  • CPU is the bottleneck — upgrading RAM bandwidth helps more than VRAM here
2
tok/s
Estimated
Weights
45.00 GB
KV cache
9.75 GB
Activations
2.25 GB
Runtime
1.80 GB
Time to first token (prefill, 512-token prompt): ~1597 ms (noticeable)
Llama 3.2 90B Vision Instruct
90B
llama
Commercial OK
Quant: Q4_K_MContext: 2,048VRAM: 66.6 GBHeadroom: 3.8 GBTTFT: noticeable
  • Partial CPU offload: ~52% of layers run on CPU
  • CPU is the bottleneck — upgrading RAM bandwidth helps more than VRAM here
ollama run llama3.2-vision:90b
1
tok/s
Estimated
Weights
51.00 GB
KV cache
11.25 GB
Activations
2.55 GB
Runtime
1.80 GB
Time to first token (prefill, 512-token prompt): ~1843 ms (noticeable)
Qwen 2.5-VL 72B
72B
qwen
Commercial OK
Quant: AWQ-INT4Context: 2,048VRAM: 54.9 GBHeadroom: 15.5 GBTTFT: noticeable
  • Partial CPU offload: ~42% of layers run on CPU
  • CPU is the bottleneck — upgrading RAM bandwidth helps more than VRAM here
1
tok/s
Estimated
Weights
42.00 GB
KV cache
9.00 GB
Activations
2.10 GB
Runtime
1.80 GB
Time to first token (prefill, 512-token prompt): ~1475 ms (noticeable)
Llama 3.2 90B Vision
90B
llama
Commercial OK
Quant: AWQ-INT4Context: 2,048VRAM: 67.7 GBHeadroom: 2.7 GBTTFT: noticeable
  • Partial CPU offload: ~53% of layers run on CPU
  • CPU is the bottleneck — upgrading RAM bandwidth helps more than VRAM here
1
tok/s
Estimated
Weights
52.00 GB
KV cache
11.25 GB
Activations
2.60 GB
Runtime
1.80 GB
Time to first token (prefill, 512-token prompt): ~1843 ms (noticeable)

What if you upgraded?

Hypothetical scenarios. We re-ran the compatibility engine for each.

+32 GB system RAM

~$80–150

Doubles your CPU-offload working set. Helps when models don't quite fit in VRAM.

Unlocks: 187 new comfortable, 48 new tradeoff

  • Qwen 3 0.6B
  • Llama 3.1 8B Instruct
  • Qwen 3 30B-A3B
  • Qwen 2.5 Coder 32B Instruct

Upgrade to NVIDIA A100 40GB

see current pricing

40 GB VRAM (vs your 32 GB) plus a bandwidth jump from ~1792 GB/s to ~1555 GB/s.

Unlocks: 184 new comfortable

  • Qwen 3 0.6B
  • Llama 3.1 8B Instruct
  • Qwen 3 30B-A3B
  • Qwen 2.5 Coder 32B Instruct

Add a second NVIDIA GeForce RTX 5090

~$2499

Tensor parallelism splits the model across both cards, effectively doubling VRAM. Bandwidth doesn't double — runs ~1.5× the single-card speed in practice.

Unlocks: 226 new comfortable

  • Qwen 3 0.6B
  • Llama 3.1 8B Instruct
  • Qwen 3 30B-A3B
  • Qwen 2.5 Coder 32B Instruct

Some links above are affiliate links. We may earn a commission at no extra cost to you. How we make money.

Won't run
top 5 popular models

Need more memory than you have. Shown for orientation.

DeepSeek V4 Pro (1.6T MoE)
1600B
deepseek
Commercial OK

Even with CPU offload, needs more memory than your VRAM (32 GB) + 60% of system RAM (38 GB) combined.

Qwen 3.5 235B-A17B (MoE)
397B
qwen
Commercial OK

Even with CPU offload, needs more memory than your VRAM (32 GB) + 60% of system RAM (38 GB) combined.

Qwen 3 235B-A22B
235B
qwen
Commercial OK

Even with CPU offload, needs more memory than your VRAM (32 GB) + 60% of system RAM (38 GB) combined.

DeepSeek R1 (671B reasoning)
671B
deepseek
Commercial OK

Even with CPU offload, needs more memory than your VRAM (32 GB) + 60% of system RAM (38 GB) combined.

Llama 4 Scout
109B
llama
Commercial OK

Even with CPU offload, needs more memory than your VRAM (32 GB) + 60% of system RAM (38 GB) combined.

How to read these numbers

Measured here
Measured here - RunLocalAI ran this exact combo on owner hardware with public evidence.

Source-backed
Source-backed / community - a reproduced public source supports the speed, but it is not labeled as owner-measured.

Extrapolated
Extrapolated - predicted from a measured benchmark on similar-bandwidth hardware.

Estimated
Estimated - formula based on VRAM bandwidth and model architecture; not a benchmark row.

RunLocalAI Will-It-Run Framework →

Want a specific benchmark we don't have? Email Contact support and we'll prioritize it.