Ollama v0.32.10-rc1

v0.32.10-rc1Pre-release

v0.32.10

Changed 2
  • Models that don't set a repeat_penalty now default to 1.0 (off) instead of 1.1, matching other engines and speeding up speculative decoding
  • Faster prefill on NVFP4 MLX models with a global scale, about 7–8% on Qwen3.6 and Muse Glimmer
Fixed 1
  • Fixed blob verification being skipped when an OCI manifest's config and layer share a digest
What's Changed
  • Models that don't set a repeat_penalty now default to 1.0 (off) instead of 1.1, matching other engines and speeding up speculative decoding; set a per-model parameter if an older model repeats itself.
  • Faster prefill on NVFP4 MLX models with a global scale, about 7–8% on Qwen3.6 and Muse Glimmer.
  • Fixed blob verification being skipped when an OCI manifest's config and layer share a digest.
New Contributors

Full Changelog: https://github.com/ollama/ollama/compare/v0.32.8...v0.32.10-rc1

View original

Upgraded? How did it go?

Discussion