Apple unifiedUnified memorybeginner
Pooled memory

What runs on Mac Studio M3 Ultra 192GB?

Apple Silicon flagship with 192 GB unified memory. Genuinely pools — total VRAM ≈ effective VRAM. Trades NVIDIA throughput for the largest model envelope at any reasonable power budget.

At a glance
Effective VRAM
140 / 192 GB
Genuinely pooled
Speed penalty
~0%
vs ideal single-card
Recommended runtime
mlx-lm
none
Setup difficulty
beginner
~370W peak
24
Models fit
6
Borderline
8
Not practical
Deployment recipe
Apple Silicon AI

Single-Mac deployment recipe with MLX-LM + Ollama. The canonical Apple Silicon path.

Memory budget
Total VRAM
192 GB
Effective for inference
140 GB
73% of total
Genuinely pooled

Apple unified memory genuinely pools — there is no separate GPU VRAM. The CPU, GPU, and Neural Engine all share the same 192 GB pool with ~800 GB/s memory bandwidth. Effective ceiling for inference is ~140 GB because macOS reserves system memory and you need headroom for KV cache and activations. Concretely: a 200B-class model at Q4 (~110 GB weights) fits comfortably with 25-30 GB of context budget. This is the rare case where 'pooled VRAM' is genuine, not marketing. The tradeoff: 800 GB/s bandwidth is 25-30% of an RTX 4090, so tokens-per-second scale lower even though the model fits.

Why total VRAM is not the whole story

Genuinely pooled — 192 GB of unified memory shared between CPU/GPU. Effective ceiling is ~140 GB after OS reservations and KV-cache budget.

See the multi-GPU guide for topology tradeoffs, and the RunLocalAI Will-It-Run Framework for the citable fit-tier method.

Topology

Topology
apple-unified
Interconnect
unified-memory~800 GB/s
Component count
1 unit
Components
Recommended runtime
mlx-lm
Also: llama-cpp, ollama, lm-studio
Recommended split strategy
none
Setup difficulty
beginner
~370W peak

Models that fit comfortably (24)

Effective VRAM utilization ≤ 85% at the smallest production quant. Comfortable headroom for KV cache.

144B·AWQ-INT496 GB·69% of effective VRAM·~0% speed penalty vs ideal
141B·Q4_K_M96 GB·69% of effective VRAM·~0% speed penalty vs ideal
141B·Q4_K_M96 GB·69% of effective VRAM·~0% speed penalty vs ideal
132B·AWQ-INT496 GB·69% of effective VRAM·~0% speed penalty vs ideal
132B·Q4_K_M96 GB·69% of effective VRAM·~0% speed penalty vs ideal
123B·Q4_K_M88 GB·63% of effective VRAM·~0% speed penalty vs ideal
120B·Q4_K_M84 GB·60% of effective VRAM·~0% speed penalty vs ideal
109B·Q4_K_M80 GB·57% of effective VRAM·~0% speed penalty vs ideal
105B·Q4_K_M74 GB·53% of effective VRAM·~0% speed penalty vs ideal
105B·Q4_K_M74 GB·53% of effective VRAM·~0% speed penalty vs ideal
104B·Q4_K_M70 GB·50% of effective VRAM·~0% speed penalty vs ideal
104B·AWQ-INT472 GB·51% of effective VRAM·~0% speed penalty vs ideal
90B·Q4_K_M60 GB·43% of effective VRAM·~0% speed penalty vs ideal
90B·AWQ-INT464 GB·46% of effective VRAM·~0% speed penalty vs ideal
78B·Q4_K_M52 GB·37% of effective VRAM·~0% speed penalty vs ideal
72B·Q4_K_M48 GB·34% of effective VRAM·~0% speed penalty vs ideal
72B·Q4_K_M48 GB·34% of effective VRAM·~0% speed penalty vs ideal
72B·Q4_K_M48 GB·34% of effective VRAM·~0% speed penalty vs ideal
72B·AWQ-INT448 GB·34% of effective VRAM·~0% speed penalty vs ideal
70B·Q4_K_M48 GB·34% of effective VRAM·~0% speed penalty vs ideal
70B·AWQ-INT448 GB·34% of effective VRAM·~0% speed penalty vs ideal
70B·Q4_K_M48 GB·34% of effective VRAM·~0% speed penalty vs ideal
70B·Q4_K_M48 GB·34% of effective VRAM·~0% speed penalty vs ideal
70B·Q4_K_M48 GB·34% of effective VRAM·~0% speed penalty vs ideal

Borderline (6)

Fits but with little headroom. KV cache for long context may not fit; verify before deployment.

253B·Q4_K_M160 GB·114% of effective VRAM·~0% speed penalty vs ideal

Effective VRAM utilization >114% — KV cache for long context will not fit. Cap context at ~4-8K or move to a larger combo.

236B·Q4_K_M160 GB·114% of effective VRAM·~0% speed penalty vs ideal

Effective VRAM utilization >114% — KV cache for long context will not fit. Cap context at ~4-8K or move to a larger combo.

236B·Q4_K_M160 GB·114% of effective VRAM·~0% speed penalty vs ideal

Effective VRAM utilization >114% — KV cache for long context will not fit. Cap context at ~4-8K or move to a larger combo.

235B·Q4_K_M160 GB·114% of effective VRAM·~0% speed penalty vs ideal

Effective VRAM utilization >114% — KV cache for long context will not fit. Cap context at ~4-8K or move to a larger combo.

Kimi K1.5
Borderline
200B·AWQ-INT4140 GB·100% of effective VRAM·~0% speed penalty vs ideal

Effective VRAM utilization >100% — KV cache for long context will not fit. Cap context at ~4-8K or move to a larger combo.

GLM-5
Borderline
200B·Q4_K_M140 GB·100% of effective VRAM·~0% speed penalty vs ideal

Effective VRAM utilization >100% — KV cache for long context will not fit. Cap context at ~4-8K or move to a larger combo.

Not practical (8)

Model weights exceed effective combo VRAM. Even with the recommended split strategy, this configuration won't run cleanly. Drop to a smaller quant or move to a larger combo.

1600B·Q4_K_M1024 GB·731% of effective VRAM·~0% speed penalty vs ideal

Model weights exceed effective combo VRAM. Even with the recommended split strategy, this configuration won't run cleanly.

Step-3
Not practical
1000B·AWQ-INT4640 GB·457% of effective VRAM·~0% speed penalty vs ideal

Model weights exceed effective combo VRAM. Even with the recommended split strategy, this configuration won't run cleanly.

Kimi K2.6
Not practical
1000B·Q4_K_M700 GB·500% of effective VRAM·~0% speed penalty vs ideal

Model weights exceed effective combo VRAM. Even with the recommended split strategy, this configuration won't run cleanly.

DeepSeek V4
Not practical
745B·AWQ-INT4480 GB·343% of effective VRAM·~0% speed penalty vs ideal

Model weights exceed effective combo VRAM. Even with the recommended split strategy, this configuration won't run cleanly.

675B·Q4_K_M448 GB·320% of effective VRAM·~0% speed penalty vs ideal

Model weights exceed effective combo VRAM. Even with the recommended split strategy, this configuration won't run cleanly.

671B·Q4_K_M420 GB·300% of effective VRAM·~0% speed penalty vs ideal

Model weights exceed effective combo VRAM. Even with the recommended split strategy, this configuration won't run cleanly.

671B·Q4_K_M420 GB·300% of effective VRAM·~0% speed penalty vs ideal

Model weights exceed effective combo VRAM. Even with the recommended split strategy, this configuration won't run cleanly.

Llama 4 405B
Not practical
405B·AWQ-INT4280 GB·200% of effective VRAM·~0% speed penalty vs ideal

Model weights exceed effective combo VRAM. Even with the recommended split strategy, this configuration won't run cleanly.

Benchmark opportunities

estimates, not measurements

Pending benchmark targets for this combo. Once measured, results land in the catalog as benchmarks.

Mac Studio M3 Ultra 192GB + Qwen 3.5 235B-A17B (MLX-4bit)
pending
Estimate: 8-14 tok/s decode (single stream)

Apple Silicon at the frontier-MoE envelope. 17B-active makes this fit comfortably in 192GB unified memory. Bandwidth-bound; expect ~25-30% of NVIDIA tok/s but largest fittable model wins.

Going deeper