# unsloth v0.1.806-beta — 2x Faster Qwen3.8-Flash + GLM-5.3-Flash MTP - Product: unsloth (https://whatsnew.fyi/product/unsloth) - Vendor: unslothai - Date: 2026-09-02 - Version: v0.1.806-beta - Original notes: https://github.com/unslothai/unsloth/releases/tag/v0.1.806-beta - Permalink: https://whatsnew.fyi/product/unsloth/releases/v0.1.806-beta - Labels: Pre-release What's New is an index, not a publisher: every entry below links to the vendor's own release notes, which are the authoritative source. Entries are labelled where they are hand-curated sample data, pre-releases, or drawn from a secondary source such as a developer blog. Reuse: the summaries, labels and curation here are © What's New. Quote freely with attribution and a link back; wholesale republication of the corpus is not permitted — terms: https://whatsnew.fyi/terms. The vendors' own release notes remain their publishers'. --- - **added** — Enable MTP by default for Qwen3.8-Flash-Next and GLM-5.3-Flash to run up to 2x faster - **added** — Add support for MiniMax-Music3, Higgs, and MOSS audio models - **added** — Add live progress updates while audio is being generated - **added** — Add local media APIs for video, audio, and MLX-served models - **added** — Add ability to archive and manage audio clips - **added** — Add OpenAI-compatible Videos API for video generation - **added** — Add support for running multiple tool calls simultaneously without mixing up their arguments - **added** — Add local code editing capability using Codex's apply_patch tool - **added** — Add setting that tells the model the current date - **added** — Add ability to export GLM-5.3 MLX fine-tunes to GGUF - **added** — Add option to choose custom GGUF shard sizes and save locations - **changed** — Enable MLX models to use their full context size and support much longer batched generation - **changed** — Improve Qwen automatically to apply recommended settings for thinking and non-thinking modes - **changed** — Improve multi-GPU training with automatic model placement across multiple GPUs - **changed** — Improve AMD/ROCm detection, installation, and GPU compatibility with BF16 support on more GPUs - **changed** — Improve MLX model loading to release GPU memory more cleanly between generation bursts and model switches - **changed** — Improve model loading with fewer errors across local servers and connected providers - **changed** — Improve chat safety to preserve tool cards, reply details, and conversation branches during edits - **fixed** — Fix custom TTS playback and Whisper pairing checks for audio reliability - **fixed** — Fix MCP connection persistence for faster tool calls within each chat Run Qwen3.8-Flash-Next and GLM-5.3-Flash up to 2x faster with MTP. MTP is enabled by default, you can still disable it. Also our new release includes 170+ training, chat, hardware, and performance improvements. ##### Highlights * **Smoother model loading** (less errors) across local servers and connected providers. * **Safer chat edits** that preserve tool cards, reply details, and conversation branches. * **New local media APIs** for video, audio, and MLX-served models. * New audio support with new models, progress tracking including: MiniMax-Music3, Higgs, MOSS and more! * Improved multi-GPU planning, memory fitting, and split-model training. * Strengthened AMD/ROCm detection, installation, and GPU compatibility. * Upgraded MCP, Deep Research, OAuth, and agent tool reliability. ##### Qwen3.8-Flash + GLM-5.3-Flash - Qwen and GLM now generate faster with MTP enabled by default. - Use GLM tools across longer, multi-turn chats. - Qwen automatically applies the recommended settings for thinking and non-thinking modes. Download [Qwen3.8-Flash-Next](https://huggingface.co/unsloth/Qwen3.8-Flash-Next) and [GLM-5.3-Flash](https://huggingface.co/unsloth/GLM-5.3-Flash). See the [Qwen guide](https://unsloth.ai/docs/models/qwen3.8-next) and [GLM guide](https://unsloth.ai/docs/models/glm-5.3-flash) for recommended settings and available GGUFs. ##### Faster MLX inference - Fine-tune both large MoE models with text or images on Apple Silicon using MLX. - Long Qwen chats now run much faster on Mac, with follow-up turns up to 30x faster. - MLX models now use their full context size and support much longer batched generation. - MLX releases GPU memory more cleanly between generation bursts and model switches. - Serve MLX models through Unsloth's OpenAI-compatible API. ##### Audio * Added support for **MiniMax-Music3, Higgs, MOSS audio models**. * Added **live progress updates** while audio is being generated. * Audio clips can now be **archived and managed**. * Improved reliability with **custom TTS playback fixes, Whisper pairing checks, and stronger audio testing**. ##### Chat + tools - Run several tool calls at once without mixing up their arguments. - Keep tools available when chatting with images. - Each chat keeps its MCP connection for faster tool calls. - Local models can edit code using Codex’s apply_patch tool. - Continue long chats with images and other media using Auto Compaction. - Review and approve Deep Research plans before research starts. ##### Training + hardware - Train larger models across multiple GPUs with automatic placement. - AMD installs choose the best build across Windows and Linux, with BF16 on more GPUs. - Export GLM-5.3 MLX fine-tunes to GGUF. - Choose custom GGUF shard sizes and save locations. ##### API + Desktop - Generate videos through the new OpenAI-compatible Videos API. - Updates download in the background and install when you restart. - Choose a custom port for LAN access. - Generate audio with Higgs, MOSS and MiniMax models. - Track audio generation progress and archive finished clips. - Model downloads show clearer progress and can switch from Xet to HTTP automatically. ##### Download Unsloth Desktop Unsloth Desktop is **free and open source**. Download it for: - **Windows** - **macOS** - **Linux** **[🦥 Download Unsloth Desktop](https://unsloth.ai/download)** ##### What's Changed * Put the smart offload planner back behind its flag by @danielhanchen in https://github.com/unslothai/unsloth/pull/9862 * Studio: stop the per-chunk autosave writing back messages the server owns by @danielhanchen in https://github.com/unslothai/unsloth/pull/9865 * Bump install.sh / install.ps1 pins to unsloth>=2026.8.22 by @danielhanchen in https://github.com/unslothai/unsloth/pull/9868 * Fix Studio hydrating synced GGUF files before selection by @milewski in https://github.com/unslothai/unsloth/pull/9539 * Fix datasets PyA _[Truncated at 4000 characters — full notes: https://github.com/unslothai/unsloth/releases/tag/v0.1.806-beta]_