Updatify / Ollama | Release notes

Create your changelog

Update Jul 25, 2026 tracked by Updatify

v0.32.4

What’s Changed

  • Support Laguna on Apple GPUs via the MLX engine
  • Quantize draft-model output heads at the requested type when creating speculative-decoding drafts.
  • Fixed Qwen3 MoE decoding for differently-quantized experts, plus faster packed gate/up projection (~4–9% on M5 Max).

Full Changelog: https://github.com/ollama/ollama/compare/v0.32.3…v0.32.4

Read the original release on Ollama ↗
← All Ollama releases

Powered by Updatify