unsloth v0.1.805-beta

v0.1.805-betaPre-release

2x Faster Qwen3.8-Flash + GLM-5.3-Flash MTP

Added 10
  • New local media APIs for video, audio, and MLX-served models
  • Support for MiniMax-Music3, Higgs, and MOSS audio models
  • Live progress updates while audio is being generated
  • Audio archiving and management capabilities
  • Fine-tune large MoE models with text or images on Apple Silicon using MLX
  • Serve MLX models through Unsloth's OpenAI-compatible API
Changed 16
  • Enable MTP by default for Qwen3.8-Flash-Next and GLM-5.3-Flash, with option to disable it
  • Smoother model loading across local servers and connected providers with fewer errors
  • Safer chat edits that preserve tool cards, reply details, and conversation branches
  • GLM-5.3-Flash tools now work across longer, multi-turn chats
  • Qwen automatically applies recommended settings for thinking and non-thinking modes
  • Long Qwen chats run much faster on Mac with follow-up turns up to 30x faster
Fixed 5
  • Run multiple tool calls at once without mixing up their arguments
  • Ctrl+F Search now works
  • Studio stops offering transformers upgrade where it cannot load anything
  • Studio CPT no longer overwrites LFM2 all-linear LoRA targets
  • Studio accepts trailing slash for model discovery

From unsloth

Run Qwen3.8-Flash-Next and GLM-5.3-Flash up to 2x faster with MTP. MTP is enabled by default, you can still disable it. Also our new release includes 170+ training, chat, hardware, and performance improvements.

Highlights
  • Smoother model loading (less errors) across local servers and connected providers.
  • Safer chat edits that preserve tool cards, reply details, and conversation branches.
  • New local media APIs for video, audio, and MLX-served models.
  • New audio support with new models, progress tracking including: MiniMax-Music3, Higgs, MOSS and more!
  • Improved multi-GPU planning, memory fitting, and split-model training.
  • Ctrl+F Search now works.
  • Strengthened AMD/ROCm detection, installation, and GPU compatibility.
  • Upgraded MCP, Deep Research, OAuth, and agent tool reliability.
Qwen3.8-Flash + GLM-5.3-Flash
  • Qwen and GLM now generate faster with MTP enabled by default.
  • Use GLM tools across longer, multi-turn chats.
  • Qwen automatically applies the recommended settings for thinking and non-thinking modes.

Download Qwen3.8-Flash-Next and GLM-5.3-Flash. See the Qwen guide and GLM guide for recommended settings and available GGUFs.

Faster MLX inference
  • Fine-tune both large MoE models with text or images on Apple Silicon using MLX.
  • Long Qwen chats now run much faster on Mac, with follow-up turns up to 30x faster.
  • MLX models now use their full context size and support much longer batched generation.
  • MLX releases GPU memory more cleanly between generation bursts and model switches.
  • Serve MLX models through Unsloth's OpenAI-compatible API.
Audio
  • Added support for MiniMax-Music3, Higgs, MOSS audio models.
  • Added live progress updates while audio is being generated.
  • Audio clips can now be archived and managed.
  • Improved reliability with custom TTS playback fixes, Whisper pairing checks, and stronger audio testing.
Chat + tools
  • Run several tool calls at once without mixing up their arguments.
  • Keep tools available when chatting with images.
  • Each chat keeps its MCP connection for faster tool calls.
  • Local models can edit code using Codex’s apply_patch tool.
  • Continue long chats with images and other media using Auto Compaction.
  • Review and approve Deep Research plans before research starts.
Training + hardware
  • Train larger models across multiple GPUs with automatic placement.
  • AMD installs choose the best build across Windows and Linux, with BF16 on more GPUs.
  • Export GLM-5.3 MLX fine-tunes to GGUF.
  • Choose custom GGUF shard sizes and save locations.
API + Desktop
  • Generate videos through the new OpenAI-compatible Videos API.
  • Updates download in the background and install when you restart.
  • Choose a custom port for LAN access.
  • Generate audio with Higgs, MOSS and MiniMax models.
  • Track audio generation progress and archive finished clips.
  • Model downloads show clearer progress and can switch from Xet to HTTP automatically.
Download Unsloth Desktop

Unsloth Desktop is free and open source. Download it for:

  • Windows
  • macOS
  • Linux

🦥 Download Unsloth Desktop

What's Changed
New Contributors

Full Changelog: https://github.com/unslothai/unsloth/compare/v0.1.804-beta...v0.1.805-beta

View original

Upgraded? How did it go?

Discussion