← /pulse/ollama-0-31-1-gemma4-mtp
ADVISORYRUNTIME UPDATE·2026-07-10

Ollama 0.31.1: Gemma 4 up to ~90% faster on Apple Silicon

▼ WHAT HAPPENED

Ollama v0.31.1 (2026-06-30) enables multi-token prediction (MTP) in its MLX engine for Gemma 4 on Apple Silicon, reporting up to ~90% faster average generation in a coding-agent benchmark. MTP is on by default and auto-tuned - Ollama decides how many tokens to draft at runtime, output is identical to standard decoding, and no configuration is needed. This lands on what is probably the most common local-AI pairing there is: a Gemma-class model on an M-series Mac through Ollama. The follow-up v0.31.2 also enabled flash attention on Pascal-era NVIDIA cards (GTX 10-series, P40) and iGPU offload for vision models, improving the budget and legacy-hardware story in the same week.

▼ OPERATOR ANGLE

If you run Gemma 4 on an M-series Mac through Ollama, update to 0.31.1+ and re-measure your tokens/sec - any benchmark you recorded before June 30 is now stale. Published Gemma-4-on-Apple-Silicon numbers across the web (including catalog sites) predate MTP; trust your own re-run.

▼ ENTITIES REFERENCED

[pulse item] · runlocalai.co/pulse/ollama-0-31-1-gemma4-mtp