What can Apple Mac Studio (M3 Ultra) run for chat?

Build: Apple Mac Studio (M3 Ultra) + — + 192 GB RAM (macos)

Memory: 192 GB unified memory
Runner: MLX-LM (Apple Metal)

Runs comfortably
268 models

Ranked by fit for chat use case + predicted speed. Click a row for VRAM breakdown.

#1Qwen 3 0.6B
0.6B
qwen
Commercial OK
Quant: Q4_K_MContext: 8,192VRAM: 1.1 GBHeadroom: 182.9 GB
1215
tok/s
Estimated
Weights
0.30 GB
KV cache
0.30 GB
Activations
0.02 GB
Runtime
0.50 GB
#2Llama 3.2 3B Instruct
3B
llama
Commercial OK
Quant: Q8_0Context: 8,192VRAM: 5.6 GBHeadroom: 178.4 GB
ollama run llama3.2:3b
138
tok/s
Estimated
Weights
3.40 GB
KV cache
1.50 GB
Activations
0.18 GB
Runtime
0.50 GB
#3Qwen 3 1.7B
1.7B
qwen
Commercial OK
Quant: Q4_K_MContext: 8,192VRAM: 2.3 GBHeadroom: 181.7 GB
429
tok/s
Estimated
Weights
0.90 GB
KV cache
0.85 GB
Activations
0.05 GB
Runtime
0.50 GB
#4TinyLlama 1.1B Chat v1.0
1.1B
llama
Commercial OK
Quant: Q4_K_MContext: 2,048VRAM: 1.3 GBHeadroom: 182.7 GB
663
tok/s
Estimated
Weights
0.60 GB
KV cache
0.14 GB
Activations
0.03 GB
Runtime
0.50 GB
#5Gemma 2 2B Instruct
2B
gemma
Commercial OK
Quant: Q4_K_MContext: 8,192VRAM: 2.7 GBHeadroom: 181.3 GB
364
tok/s
Estimated
Weights
1.10 GB
KV cache
1.00 GB
Activations
0.06 GB
Runtime
0.50 GB
#6DeepSeek V2 Lite Chat
15.7B
deepseek
Commercial OK
Quant: Q4_K_MContext: 8,192VRAM: 17.4 GBHeadroom: 166.6 GB
304
tok/s
Estimated
Weights
8.60 GB
KV cache
7.85 GB
Activations
0.44 GB
Runtime
0.50 GB
#7Kumru 2B
2.4B
mistral
Commercial OK
Quant: Q4_K_MContext: 8,192VRAM: 3.3 GBHeadroom: 180.7 GB
ollama run alibayram/kumru:latest
304
tok/s
Estimated
Weights
1.50 GB
KV cache
1.20 GB
Activations
0.08 GB
Runtime
0.50 GB
#8TinyLlama 1.1B Chat v0.3 AWQ
1.1B
other
Commercial OK
Quant: Q4_K_MContext: 2,048VRAM: 1.3 GBHeadroom: 182.7 GB
663
tok/s
Estimated
Weights
0.60 GB
KV cache
0.14 GB
Activations
0.03 GB
Runtime
0.50 GB
#9TinyLlama 1.1B Chat v0.3 GPTQ
1.1B
other
Commercial OK
Quant: Q4_K_MContext: 2,048VRAM: 1.3 GBHeadroom: 182.7 GB
663
tok/s
Estimated
Weights
0.60 GB
KV cache
0.14 GB
Activations
0.03 GB
Runtime
0.50 GB
Quant: Q4_K_MContext: 8,192VRAM: 3.1 GBHeadroom: 180.9 GB
304
tok/s
Estimated
Weights
1.30 GB
KV cache
1.20 GB
Activations
0.07 GB
Runtime
0.50 GB
#11Falcon 3 3B Instruct
3B
falcon
Commercial OK
Quant: Q4_K_MContext: 8,192VRAM: 3.8 GBHeadroom: 180.2 GB
243
tok/s
Estimated
Weights
1.70 GB
KV cache
1.50 GB
Activations
0.09 GB
Runtime
0.50 GB
#12PhoGPT 4B Chat
3.7B
other
Commercial OK
Quant: Q4_K_MContext: 8,192VRAM: 4.5 GBHeadroom: 179.5 GB
197
tok/s
Estimated
Weights
2.00 GB
KV cache
1.85 GB
Activations
0.11 GB
Runtime
0.50 GB

Runs with tradeoffs
1 models

Tight VRAM, partial CPU offload, or context-limited.

Llama 3.1 Nemotron Ultra 253B
253B
llama
Commercial OK
Quant: Q4_K_MContext: 2,048VRAM: 183.3 GBHeadroom: 0.7 GB
  • Tight VRAM fit — only 0.7 GB headroom left for context growth
3
tok/s
Estimated
Weights
144.00 GB
KV cache
31.63 GB
Activations
7.20 GB
Runtime
0.50 GB

What if you upgraded?

Hypothetical scenarios. We re-ran the compatibility engine for each.

Move up an Apple memory tier

~$200–400 over base

On Apple Silicon, more unified memory is the only path forward — VRAM and system RAM are the same pool.

Some links above are affiliate links. We may earn a commission at no extra cost to you. How we make money.

Won't run
top 5 popular models

Need more memory than you have. Shown for orientation.

DeepSeek V4 Pro (1.6T MoE)
1600B
deepseek
Commercial OK

Needs ~1024 GB unified memory minimum at smallest quant; you have 184 GB available after OS overhead.

Qwen 3.5 235B-A17B (MoE)
397B
qwen
Commercial OK

Needs ~256 GB unified memory minimum at smallest quant; you have 184 GB available after OS overhead.

DeepSeek R1 (671B reasoning)
671B
deepseek
Commercial OK

Needs ~420 GB unified memory minimum at smallest quant; you have 184 GB available after OS overhead.

DeepSeek V4 Flash (284B MoE)
284B
deepseek
Commercial OK

Needs ~192 GB unified memory minimum at smallest quant; you have 184 GB available after OS overhead.

GLM-5.2
753B
glm
Commercial OK

Needs ~528 GB unified memory minimum at smallest quant; you have 184 GB available after OS overhead.

How to read these numbers

Measured here
Measured here - RunLocalAI ran this exact combo on owner hardware with public evidence.

Source-backed
Source-backed / community - a reproduced public source supports the speed, but it is not labeled as owner-measured.

Extrapolated
Extrapolated - predicted from a measured benchmark on similar-bandwidth hardware.

Estimated
Estimated - formula based on VRAM bandwidth and model architecture; not a benchmark row.

RunLocalAI Will-It-Run Framework →

Want a specific benchmark we don't have? Email Contact support and we'll prioritize it.