What can MacBook Pro 16" M4 Max run for agents?

Build: MacBook Pro M4 Max 128GB

Memory: 128 GB unified memory
Runner: MLX-LM (Apple Metal)

Runs comfortably
207 models

Ranked by fit for agents use case + predicted speed. Click a row for VRAM breakdown.

#1Hermes 3 Llama 3.1 8B
8B
hermes
Commercial OK
Quant: Q8_0Context: 8,192VRAM: 13.4 GBHeadroom: 106.6 GB
ollama run hermes3:8b
35
tok/s
Estimated
Weights
8.50 GB
KV cache
4.00 GB
Activations
0.43 GB
Runtime
0.50 GB
#2Dolphin 3.0 Mistral 24B
24B
dolphin
Commercial OK
Quant: Q4_K_MContext: 8,192VRAM: 27.2 GBHeadroom: 92.8 GB
ollama run dolphin-mistral:24b
21
tok/s
Estimated
Weights
14.00 GB
KV cache
12.00 GB
Activations
0.71 GB
Runtime
0.50 GB
#3Qwen 2.5 7B Instruct
7B
qwen
Commercial OK
Quant: Q8_0Context: 8,192VRAM: 9.5 GBHeadroom: 110.5 GB
ollama run qwen2.5:7b
40
tok/s
Estimated
Weights
8.10 GB
KV cache
0.47 GB
Activations
0.41 GB
Runtime
0.50 GB
#4Qwen 2.5 14B Instruct
14B
qwen
Commercial OK
Quant: Q8_0Context: 8,192VRAM: 24.0 GBHeadroom: 96.0 GB
ollama run qwen2.5:14b
20
tok/s
Estimated
Weights
15.70 GB
KV cache
7.00 GB
Activations
0.79 GB
Runtime
0.50 GB
#5Qwen3 Coder 30B-A3B
30B
qwen
Commercial OK
Quant: Q4_K_MContext: 8,192VRAM: 35.5 GBHeadroom: 84.5 GB
ollama run qwen3-coder:30b
17
tok/s
Estimated
Weights
19.00 GB
KV cache
15.00 GB
Activations
0.96 GB
Runtime
0.50 GB
#6Mistral 7B Instruct v0.3
7B
mistral
Commercial OK
Quant: Q5_K_MContext: 8,192VRAM: 9.4 GBHeadroom: 110.6 GB
ollama run mistral:7b
62
tok/s
Estimated
Weights
5.10 GB
KV cache
3.50 GB
Activations
0.26 GB
Runtime
0.50 GB
#7Llama 3.1 Nemotron Nano 8B
8B
llama
Commercial OK
Quant: Q4_K_MContext: 8,192VRAM: 9.7 GBHeadroom: 110.3 GB
62
tok/s
Estimated
Weights
4.90 GB
KV cache
4.00 GB
Activations
0.25 GB
Runtime
0.50 GB
#8North Mini Code 1.0
30B
other
Commercial OK
Quant: Q4_K_MContext: 8,192VRAM: 35.5 GBHeadroom: 84.5 GB
ollama run north-mini-code-1.0
17
tok/s
Estimated
Weights
19.00 GB
KV cache
15.00 GB
Activations
0.96 GB
Runtime
0.50 GB
#9GLM-4.7-Flash
31B
glm
Commercial OK
Quant: Q4_K_MContext: 8,192VRAM: 36.0 GBHeadroom: 84.0 GB
ollama run glm-4.7-flash
16
tok/s
Estimated
Weights
19.00 GB
KV cache
15.50 GB
Activations
0.96 GB
Runtime
0.50 GB
#10Ornith 1.0 9B
9B
other
Commercial OK
Quant: Q4_K_MContext: 8,192VRAM: 10.9 GBHeadroom: 109.1 GB
ollama run ornith:9b
55
tok/s
Estimated
Weights
5.60 GB
KV cache
4.50 GB
Activations
0.29 GB
Runtime
0.50 GB
#11Laguna XS 2.1
33B
other
Commercial OK
Quant: Q4_K_MContext: 8,192VRAM: 38.0 GBHeadroom: 82.0 GB
ollama run laguna-xs-2.1
15
tok/s
Estimated
Weights
20.00 GB
KV cache
16.50 GB
Activations
1.01 GB
Runtime
0.50 GB
#12Ornith 1.0 35B
35B
other
Commercial OK
Quant: Q4_K_MContext: 8,192VRAM: 40.1 GBHeadroom: 79.9 GB
ollama run ornith:35b
14
tok/s
Estimated
Weights
21.00 GB
KV cache
17.50 GB
Activations
1.06 GB
Runtime
0.50 GB

What if you upgraded?

Hypothetical scenarios. We re-ran the compatibility engine for each.

Move up an Apple memory tier

~$200–400 over base

On Apple Silicon, more unified memory is the only path forward — VRAM and system RAM are the same pool.

Some links above are affiliate links. We may earn a commission at no extra cost to you. How we make money.

Won't run
top 5 popular models

Need more memory than you have. Shown for orientation.

DeepSeek V4 Pro (1.6T MoE)
1600B
deepseek
Commercial OK

Needs ~1024 GB unified memory minimum at smallest quant; you have 120 GB available after OS overhead.

Qwen 3.5 235B-A17B (MoE)
397B
qwen
Commercial OK

Needs ~256 GB unified memory minimum at smallest quant; you have 120 GB available after OS overhead.

Qwen 3 235B-A22B
235B
qwen
Commercial OK

Needs ~160 GB unified memory minimum at smallest quant; you have 120 GB available after OS overhead.

DeepSeek R1 (671B reasoning)
671B
deepseek
Commercial OK

Needs ~420 GB unified memory minimum at smallest quant; you have 120 GB available after OS overhead.

DeepSeek V4 Flash (284B MoE)
284B
deepseek
Commercial OK

Needs ~192 GB unified memory minimum at smallest quant; you have 120 GB available after OS overhead.

How to read these numbers

Measured here
Measured here - RunLocalAI ran this exact combo on owner hardware with public evidence.

Source-backed
Source-backed / community - a reproduced public source supports the speed, but it is not labeled as owner-measured.

Extrapolated
Extrapolated - predicted from a measured benchmark on similar-bandwidth hardware.

Estimated
Estimated - formula based on VRAM bandwidth and model architecture; not a benchmark row.

RunLocalAI Will-It-Run Framework →

Want a specific benchmark we don't have? Email Contact support and we'll prioritize it.