ollama v0.32.4 adds Laguna support and Qwen3 fixes
▼ WHAT HAPPENED
ollama v0.32.4 introduces significant improvements for Apple GPU users by supporting the MLX engine on Laguna, enhancing performance and compatibility with Apple's hardware. Additionally, this release addresses a critical issue in Qwen3 MoE decoding, ensuring smoother operation across differently-quantized experts with up to 9% speedup on M5 Max GPUs. The update also includes quantization improvements for speculative-decoding drafts, which can enhance the efficiency of draft-model output heads during creation. Local AI operators running Apple Silicon or similar setups should consider this upgrade for better model performance and stability.
▼ OPERATOR ANGLE
Upgrade ollama to v0.32.4 if you are using Apple GPUs with Laguna support or encounter issues with Qwen3 MoE decoding on your system.