Update Aug 19, 2026 tracked by Updatify

v0.32.15

What’s Changed

  • New desktop onboarding flow on first launch
  • Caches resolved model metadata between requests, cutting time-to-first-token by roughly half (TTFT dropped from ~995 ms to ~524 ms in benchmarks)
  • Fixes a bug where chat and generate could wedge after a mid-stream parser error
  • Qwen 3.8 system messages are now normalized so non-leading system messages are handled consistently
  • MLX and llama.cpp dependency updates

New Contributors

Full Changelog: https://github.com/ollama/ollama/compare/v0.32.14…v0.32.15