# unsloth v0.1.804-beta — Qwen3.8-Flash-Next + GLM-5.3-Flash - Product: unsloth (https://whatsnew.fyi/product/unsloth) - Vendor: unslothai - Date: 2026-08-27 - Version: v0.1.804-beta - Original notes: https://github.com/unslothai/unsloth/releases/tag/v0.1.804-beta - Permalink: https://whatsnew.fyi/product/unsloth/releases/v0.1.804-beta - Labels: Pre-release What's New is an index, not a publisher: every entry below links to the vendor's own release notes, which are the authoritative source. Entries are labelled where they are hand-curated sample data, pre-releases, or drawn from a secondary source such as a developer blog. Reuse: the summaries, labels and curation here are © What's New. Quote freely with attribution and a link back; wholesale republication of the corpus is not permitted — terms: https://whatsnew.fyi/terms. The vendors' own release notes remain their publishers'. --- - **added** — Support for running Qwen3.8-Flash-Next locally on 75GB RAM - **added** — Support for running GLM-5.3-Flash locally on 102GB total RAM and VRAM - **added** — Local chats resume after a disconnect instead of losing the reply - **added** — Export chats as JSONL for backups or use in other tools - **added** — Search and download embedding models directly from Hugging Face - **changed** — Inference speed improved 5x for RAM offloading - **changed** — Large GGUFs automatically split across GPU and system RAM - **changed** — Smarter GPU and RAM offloading to run larger models with less setup - **changed** — Vision chats now handle multiple images properly - **changed** — Text-to-speech models only load when actually used - **changed** — Model settings stay saved when switching chats - **changed** — Collapse tool activity by default for cleaner agent chats - **changed** — Deep Research keeps going when a provider asks it to slow down - **changed** — Model reasoning defaults now followed in CLI start command - **fixed** — Infinite repeated compaction now works - **fixed** — Linux voice recording fixed - **fixed** — NVIDIA and Wayland interface freezes fixed - **fixed** — AMD model loading crashes fixed - **fixed** — llama.cpp models now load from Windows profiles with non-English characters - **fixed** — Non-English web links now work properly as chat sources Qwen3.8-Flash-Next and GLM-5.3-Flash can now run locally in Unsloth! * Run Qwen3.8-Flash on 75GB RAM, GLM-5.3-Flash on 102GB RAM+VRAM * 5x Faster inference for RAM offloading * "Infinite" repeated compaction now works * 100+ chat, reliability and performance improvements Qwen Guide: https://unsloth.ai/docs/models/qwen3.8-next Qwen GGUFs: https://huggingface.co/unsloth/Qwen3.8-Flash-Next-GGUF GLM Guide: https://unsloth.ai/docs/models/glm-5.3-flash GLM GGUFs: https://huggingface.co/unsloth/GLM-5.3-Flash-GGUF ##### Highlights - **Qwen3.8-Flash-Next on 75GB RAM** - **GLM-5.3-Flash on 102GB total memory** - **Smarter GPU + RAM offloading** - run larger models with less setup - **Chats recover after disconnects** instead of losing the reply - **See what fits before loading** with clearer memory estimates ##### Qwen3.8-Flash-Next Qwen3.8-Flash-Next is a new 125B multimodal reasoning model and an early preview of Qwen4's architecture. - The 1-bit Unsloth Dynamic GGUF runs on 75GB RAM or unified memory. - It's 79% smaller than BF16 while retaining 80% top-1 accuracy. - Chat with text and images using up to 262K context. - Switch between None, Low, Medium and Extra High reasoning. - Preserved Thinking keeps reasoning consistent across longer chats. ##### GLM-5.3-Flash GLM-5.3-Flash is Z.ai's new 320B multimodal model, with only 18B parameters active at a time. - Run the 1-bit model on 102GB of combined RAM + VRAM. - Chat with text, images and long documents using up to 1M context. - Switch between Low, High and Max reasoning. - Stronger coding, agent and vision performance than GLM-5.2. - Recommended settings are applied automatically in Unsloth. ##### Chat + tools - Local chats resume after a disconnect instead of losing the reply. - Deep Research keeps going when a provider asks it to slow down. - Vision chats now handle multiple images properly. - Images returned by MCP tools appear directly in chat. - Export chats as JSONL for backups or use in other tools. - Adjust Auto Compaction for longer chats, or turn it off. - Collapse tool activity by default for cleaner agent chats. ##### Models + performance - Large GGUFs automatically split across GPU and system RAM. - See estimated memory usage before loading a model. - View VRAM usage directly from your downloaded models. - Model settings stay saved when switching chats. - Search and download embedding models directly from Hugging Face. - Text-to-speech models only load when you actually use them. ##### Desktop + reliability - Linux voice recording fixed. - NVIDIA + Wayland interface freezes fixed. - AMD model loading crashes fixed. - llama.cpp models now load from Windows profiles with non-English characters. - Non-English web links now work properly as chat sources. - Desktop download links always point to the latest stable release. ##### What's Changed * Bump install.sh / install.ps1 pin to unsloth>=2026.8.21 by @danielhanchen in https://github.com/unslothai/unsloth/pull/9699 * fix(studio): resolve PowerShell by absolute path in the update gate by @yzxcj797 in https://github.com/unslothai/unsloth/pull/9452 * Prevent shared preview loads from evicting the active Studio model by @NilayYadav in https://github.com/unslothai/unsloth/pull/7104 * studio: harden lockfile audit followups for #5604 by @danielhanchen in https://github.com/unslothai/unsloth/pull/5695 * Fix Linux voice recording by capturing raw PCM where WebKitGTK's MediaRecorder produces no audio by @Fizza-Mukhtar in https://github.com/unslothai/unsloth/pull/9564 * Studio: take remend 1.3.1, which stops repairing markdown that is already complete by @danielhanchen in https://github.com/unslothai/unsloth/pull/9667 * CLI: follow model reasoning default in unsloth start by @shimmyshimmer in https://github.com/unslothai/unsloth/pull/9733 * Name the encoding when the poll probe writes its artifact by @danielhanchen in https://github.co _[Truncated at 4000 characters — full notes: https://github.com/unslothai/unsloth/releases/tag/v0.1.804-beta]_