What can Apple Mac Studio (M3 Ultra) run?

Build: Apple Mac Studio (M3 Ultra) + — + 192 GB RAM (macos)

Memory: 192 GB unified memory
Runner: MLX-LM (Apple Metal)

Runs comfortably
277 models

Full-VRAM resident, with room for context. No compromises.

#1Qwen 3 235B-A22B
235B
qwen
Commercial OK
Quant: Q4_K_MContext: 2,048VRAM: 179.0 GBHeadroom: 5.0 GB
ollama run qwen3:235b
3
tok/s
Estimated
Weights
142.00 GB
KV cache
29.38 GB
Activations
7.10 GB
Runtime
0.50 GB
#2Qwen 3 0.6B
0.6B
qwen
Commercial OK
Quant: Q4_K_MContext: 8,192VRAM: 1.1 GBHeadroom: 182.9 GB
1215
tok/s
Estimated
Weights
0.30 GB
KV cache
0.30 GB
Activations
0.02 GB
Runtime
0.50 GB
#3Llama 4 Scout
109B
llama
Commercial OK
Quant: Q5_K_MContext: 8,192VRAM: 136.9 GBHeadroom: 47.1 GB
ollama run llama4:scout
6
tok/s
Estimated
Weights
78.00 GB
KV cache
54.50 GB
Activations
3.91 GB
Runtime
0.50 GB
#4Llama 3.1 8B Instruct
8B
llama
Commercial OK
Quant: FP16Context: 8,192VRAM: 18.5 GBHeadroom: 165.5 GB
ollama run llama3.1:8b
28
tok/s
Estimated
Weights
16.10 GB
KV cache
1.07 GB
Activations
0.81 GB
Runtime
0.50 GB
#5Qwen 3 30B-A3B
30B
qwen
Commercial OK
Quant: Q8_0Context: 8,192VRAM: 49.1 GBHeadroom: 134.9 GB
ollama run qwen3:30b
14
tok/s
Estimated
Weights
32.00 GB
KV cache
15.00 GB
Activations
1.61 GB
Runtime
0.50 GB
#6Qwen 2.5 Coder 32B Instruct
32B
qwen
Commercial OK
Quant: Q8_0Context: 8,192VRAM: 38.4 GBHeadroom: 145.6 GB
ollama run qwen2.5-coder:32b
13
tok/s
Estimated
Weights
34.00 GB
KV cache
2.15 GB
Activations
1.71 GB
Runtime
0.50 GB
#7Qwen3.6 27B
27B
qwen
Commercial OK
Quant: Q4_K_MContext: 8,192VRAM: 31.9 GBHeadroom: 152.1 GB
ollama run qwen3.6:27b
27
tok/s
Estimated
Weights
17.00 GB
KV cache
13.50 GB
Activations
0.86 GB
Runtime
0.50 GB
#8Llama 3.3 70B Instruct
70B
llama
Commercial OK
Quant: Q8_0Context: 8,192VRAM: 76.7 GBHeadroom: 107.3 GB
ollama run llama3.3:70b
6
tok/s
Estimated
Weights
70.00 GB
KV cache
2.68 GB
Activations
3.51 GB
Runtime
0.50 GB
#9Qwen 3 32B
32B
qwen
Commercial OK
Quant: Q8_0Context: 8,192VRAM: 52.2 GBHeadroom: 131.8 GB
ollama run qwen3:32b
13
tok/s
Estimated
Weights
34.00 GB
KV cache
16.00 GB
Activations
1.71 GB
Runtime
0.50 GB
#10Gemma 4 31B Dense
31B
gemma
Commercial OK
Quant: Q8_0Context: 8,192VRAM: 50.7 GBHeadroom: 133.3 GB
ollama run gemma4:31b
13
tok/s
Estimated
Weights
33.00 GB
KV cache
15.50 GB
Activations
1.66 GB
Runtime
0.50 GB
#11Qwen 3 8B
8B
qwen
Commercial OK
Quant: Q8_0Context: 8,192VRAM: 13.1 GBHeadroom: 170.9 GB
ollama run qwen3:8b
52
tok/s
Estimated
Weights
8.20 GB
KV cache
4.00 GB
Activations
0.42 GB
Runtime
0.50 GB
#12Gemma 4 12B
12B
gemma
Commercial OK
Quant: Q4_K_MContext: 8,192VRAM: 14.5 GBHeadroom: 169.5 GB
ollama run gemma4:12b
61
tok/s
Estimated
Weights
7.60 GB
KV cache
6.00 GB
Activations
0.39 GB
Runtime
0.50 GB

Runs with tradeoffs
1 models

Tight VRAM, partial CPU offload, or context-limited.

Llama 3.1 Nemotron Ultra 253B
253B
llama
Commercial OK
Quant: Q4_K_MContext: 2,048VRAM: 183.3 GBHeadroom: 0.7 GB
  • Tight VRAM fit — only 0.7 GB headroom left for context growth
3
tok/s
Estimated
Weights
144.00 GB
KV cache
31.63 GB
Activations
7.20 GB
Runtime
0.50 GB

What if you upgraded?

Hypothetical scenarios. We re-ran the compatibility engine for each.

Move up an Apple memory tier

~$200–400 over base

On Apple Silicon, more unified memory is the only path forward — VRAM and system RAM are the same pool.

Some links above are affiliate links. We may earn a commission at no extra cost to you. How we make money.

Won't run
top 5 popular models

Need more memory than you have. Shown for orientation.

DeepSeek V4 Pro (1.6T MoE)
1600B
deepseek
Commercial OK

Needs ~1024 GB unified memory minimum at smallest quant; you have 184 GB available after OS overhead.

Qwen 3.5 235B-A17B (MoE)
397B
qwen
Commercial OK

Needs ~256 GB unified memory minimum at smallest quant; you have 184 GB available after OS overhead.

DeepSeek R1 (671B reasoning)
671B
deepseek
Commercial OK

Needs ~420 GB unified memory minimum at smallest quant; you have 184 GB available after OS overhead.

DeepSeek V4 Flash (284B MoE)
284B
deepseek
Commercial OK

Needs ~192 GB unified memory minimum at smallest quant; you have 184 GB available after OS overhead.

GLM-5.2
753B
glm
Commercial OK

Needs ~528 GB unified memory minimum at smallest quant; you have 184 GB available after OS overhead.

How to read these numbers

Measured here
Measured here - RunLocalAI ran this exact combo on owner hardware with public evidence.

Source-backed
Source-backed / community - a reproduced public source supports the speed, but it is not labeled as owner-measured.

Extrapolated
Extrapolated - predicted from a measured benchmark on similar-bandwidth hardware.

Estimated
Estimated - formula based on VRAM bandwidth and model architecture; not a benchmark row.

RunLocalAI Will-It-Run Framework →

Want a specific benchmark we don't have? Email Contact support and we'll prioritize it.