unsloth v0.1.806-beta

v0.1.806-betaPre-release

2x Faster Qwen3.8-Flash + GLM-5.3-Flash MTP

Added 11
  • Enable MTP by default for Qwen3.8-Flash-Next and GLM-5.3-Flash to run up to 2x faster
  • Add support for MiniMax-Music3, Higgs, and MOSS audio models
  • Add live progress updates while audio is being generated
  • Add local media APIs for video, audio, and MLX-served models
  • Add ability to archive and manage audio clips
  • Add OpenAI-compatible Videos API for video generation
Changed 7
  • Enable MLX models to use their full context size and support much longer batched generation
  • Improve Qwen automatically to apply recommended settings for thinking and non-thinking modes
  • Improve multi-GPU training with automatic model placement across multiple GPUs
  • Improve AMD/ROCm detection, installation, and GPU compatibility with BF16 support on more GPUs
  • Improve MLX model loading to release GPU memory more cleanly between generation bursts and model switches
  • Improve model loading with fewer errors across local servers and connected providers
  • Improve chat safety to preserve tool cards, reply details, and conversation branches during edits
Fixed 2
  • Fix custom TTS playback and Whisper pairing checks for audio reliability
  • Fix MCP connection persistence for faster tool calls within each chat

From unsloth

Run Qwen3.8-Flash-Next and GLM-5.3-Flash up to 2x faster with MTP. MTP is enabled by default, you can still disable it. Also our new release includes 170+ training, chat, hardware, and performance improvements.

Highlights
  • Smoother model loading (less errors) across local servers and connected providers.
  • Safer chat edits that preserve tool cards, reply details, and conversation branches.
  • New local media APIs for video, audio, and MLX-served models.
  • New audio support with new models, progress tracking including: MiniMax-Music3, Higgs, MOSS and more!
  • Improved multi-GPU planning, memory fitting, and split-model training.
  • Strengthened AMD/ROCm detection, installation, and GPU compatibility.
  • Upgraded MCP, Deep Research, OAuth, and agent tool reliability.
Qwen3.8-Flash + GLM-5.3-Flash
  • Qwen and GLM now generate faster with MTP enabled by default.
  • Use GLM tools across longer, multi-turn chats.
  • Qwen automatically applies the recommended settings for thinking and non-thinking modes.

Download Qwen3.8-Flash-Next and GLM-5.3-Flash. See the Qwen guide and GLM guide for recommended settings and available GGUFs.

Faster MLX inference
  • Fine-tune both large MoE models with text or images on Apple Silicon using MLX.
  • Long Qwen chats now run much faster on Mac, with follow-up turns up to 30x faster.
  • MLX models now use their full context size and support much longer batched generation.
  • MLX releases GPU memory more cleanly between generation bursts and model switches.
  • Serve MLX models through Unsloth's OpenAI-compatible API.
Audio
  • Added support for MiniMax-Music3, Higgs, MOSS audio models.
  • Added live progress updates while audio is being generated.
  • Audio clips can now be archived and managed.
  • Improved reliability with custom TTS playback fixes, Whisper pairing checks, and stronger audio testing.
Chat + tools
  • Run several tool calls at once without mixing up their arguments.
  • Keep tools available when chatting with images.
  • Each chat keeps its MCP connection for faster tool calls.
  • Local models can edit code using Codex’s apply_patch tool.
  • Continue long chats with images and other media using Auto Compaction.
  • Review and approve Deep Research plans before research starts.
Training + hardware
  • Train larger models across multiple GPUs with automatic placement.
  • AMD installs choose the best build across Windows and Linux, with BF16 on more GPUs.
  • Export GLM-5.3 MLX fine-tunes to GGUF.
  • Choose custom GGUF shard sizes and save locations.
API + Desktop
  • Generate videos through the new OpenAI-compatible Videos API.
  • Updates download in the background and install when you restart.
  • Choose a custom port for LAN access.
  • Generate audio with Higgs, MOSS and MiniMax models.
  • Track audio generation progress and archive finished clips.
  • Model downloads show clearer progress and can switch from Xet to HTTP automatically.
Download Unsloth Desktop

Unsloth Desktop is free and open source. Download it for:

  • Windows
  • macOS
  • Linux

🦥 Download Unsloth Desktop

What's Changed
New Contributors

Full Changelog: https://github.com/unslothai/unsloth/compare/v0.1.804-beta...v0.1.806-beta

View original

Upgraded? How did it go?

Discussion