U

Unsloth

U
Unsloth AI

Command Palette + Desktop UI/UX

This release brings faster navigation, shareable run settings, and clearer errors to Unsloth Desktop. It also 4x speeds up Laya decisions, expands hosted Decision API support, and keeps NVFP4, INT4, and MXFP4 checkpoints in 4-bit during LoRA training. Highlights Command palette on Cmd/Ctrl+P. Also, you can now share your GGUF model's run settings Laya decisions up to 4.1x faster, plus connections…

U
Unsloth AI v0.1.901-beta

v0.1.901-beta

perf(studio): reuse embeddings for identical files in linked folders …

U
Unsloth AI

Laya Decision Models + Library

We're adding support for decision models, a unified Library for docs and media, document viewer, many Apple Silicon improvements, creation of Skills, and ~4.5× faster image and video generation. Run and serve Decision Models like Laya (open-source Jev) locally Skills Editor to create, edit, and delete Skills directly in Desktop ModelScope model downloading is now here for users who can't use HF Li…

U
Unsloth AI v2.1

Qwen-Image-2.1 + Skills

You can now run Qwen-Image-2.1 locally with Unsloth! This release also includes custom Agent Skills, and easier chat/project management. It also brings 2x faster reasoning blocks (60 FPS vs 30 FPS), more reliable training, and improved Linux installs and updates. Qwen Image 2.1 Guide Highlights Fast FP8 Qwen-Image-2.1 + diffusers update + GGUF fixes Qwen-Image-2.1 works for image editing and image…

U
Unsloth AI v2.1

Qwen-Image-2.1 + Skills

You can now run Qwen-Image-2.1 locally with Unsloth! This release also includes custom Agent Skills, and easier chat/project management. It also brings 2x faster reasoning blocks (60 FPS vs 30 FPS), more reliable training, and improved Linux installs and updates. Qwen Image 2.1 Guide Highlights Qwen-Image-2.1 fixes - diffusers update + GGUF fix Qwen-Image-2.1 works for image editing and image gen!…

U
Unsloth AI v2.1

Qwen-Image-2.1 + Skills

You can now run Qwen-Image-2.1 locally with Unsloth! This release also includes custom Agent Skills, and easier chat/project management. It also brings 2x faster reasoning blocks (60 FPS vs 30 FPS), more reliable training, and improved Linux installs and updates. Qwen Image 2.1 Guide Highlights Qwen-Image-2.1 support for image gen + more (Another update coming today for fixes!) Add custom Agent Sk…

U
Unsloth AI v2.1

Qwen-Image-2.1 + Skills

We're releasing support for Qwen-Image-2.1, custom Agent Skills, and easier chat/project management. It also brings 2x faster reasoning blocks (60 FPS vs 30 FPS), more reliable training, and improved Linux installs and updates. Qwen Image 2.1 Guide Highlights Qwen-Image-2.1 support for image generation and more Add custom skills to guide models through specific tasks. 2x faster long reasoning bloc…

U
Unsloth AI

Docker + Multi User + AMD Support

We're releasing our new updated Docker image along with multi-user accounts, RDNA1+2, FP8/INT8 diffusion support, ARM64 CUDA Windows support and many training, GRPO and inference improvements. Also a Qwen3.8-Flash-Next 2x faster MTP hotfix. Highlights: Qwen3.8-Flash-Next MTP hotfix (2x faster) from v0.1.810-beta New Docker with NVIDIA & AMD support: Guide Multi user accounts with isolation Setting…

U
Unsloth AI

Docker + Multi User + AMD Support

We're releasing our new updated Docker image along with multi-user accounts, RDNA1+2, FP8/INT8 diffusion support, ARM64 CUDA Windows support and many training, GRPO and inference improvements Highlights: New Docker with NVIDIA & AMD support: Guide Multi user accounts with isolation Settings > Accounts INT8/FP8 Image Diffusion inference - 2x faster ARM64 Windows CUDA support for training, inference…

U
Unsloth AI

Windows ARM64 Binaries

Studio: keep the GPU order the user asked for instead of re-emitting … …it ascending (#11034) * Studio: keep the GPU order the user asked for instead of re-emitting it ascending A multi-GPU CUDA launch built the child's CUDA_VISIBLE_DEVICES from gpu_indices, which every producer sorts, so a parent CUDA_VISIBLE_DEVICES=1,0 reached llama-server as 0,1. A numeric mask carries enumeration ORDER as wel…

U
Unsloth AI

Large Performance Gains + Fixes

This is a large performance and reliability + bug fix release for Unsloth Highlights 1.2-1.7x faster diffusion. AMD 20% perf boost vs ROCM via Vulkan 2x faster updating, remove SAC + AV false positives for Windows Blender MCP, detect Hermes, AMD gibberish fixed (reported to AMD) Over 250+ bug fixes, 60% smaller binaries and performance improvements Strix iGPU BIOS popup - 3x faster inference if mo…

U
Unsloth AI

Large Perf Improvements + Fixes

This is a large performance and reliability + bug fix release for Unsloth Highlights AMD uses Vulkan by default - 20% perf boost for prefill, decoding vs ROCM Windows llama-server.exe is now signed, reducing false positives for SAC AMD gibberish issues in Strix, iGPUs fixed in (upstream - reported to AMD) Over 200+ bug fixes, 50% smaller binaries and performance improvements Updated PyTorch to 2.1…

U
Unsloth AI v5.3-Flash

2x Faster Qwen3.8-Flash + GLM-5.3-Flash MTP

Run Qwen3.8-Flash-Next and GLM-5.3-Flash up to 2x faster with MTP. MTP is enabled by default, you can still disable it. Also our new release includes 170+ training, chat, hardware, and performance improvements. Highlights Smoother model loading (less errors) across local servers and connected providers. Faster and less laggy UI with follow-up turns much faster for all chats. Safer chat edits that…

U
Unsloth AI v5.3-Flash

2x Faster Qwen3.8-Flash + GLM-5.3-Flash MTP

Run Qwen3.8-Flash-Next and GLM-5.3-Flash up to 2x faster with MTP. MTP is enabled by default, you can still disable it. Also our new release includes 170+ training, chat, hardware, and performance improvements. Highlights Smoother model loading (less errors) across local servers and connected providers. Safer chat edits that preserve tool cards, reply details, and conversation branches. New local…

U
Unsloth AI v5.3-Flash

Qwen3.8-Flash-Next + GLM-5.3-Flash

Qwen3.8-Flash-Next and GLM-5.3-Flash can now run locally in Unsloth! Run Qwen3.8-Flash on 75GB RAM, GLM-5.3-Flash on 102GB RAM+VRAM 5x Faster inference for RAM offloading "Infinite" repeated compaction now works 100+ chat, reliability and performance improvements Qwen Guide: https://unsloth.ai/docs/models/qwen3.8-next Qwen GGUFs: https://huggingface.co/unsloth/Qwen3.8-Flash-Next-GGUF GLM Guide: ht…

U
Unsloth AI

Bug Fixes + Auto compaction + LAN Remote Access

Thanks for the support for Qwen3.8-27B and Unsloth Desktop! This is a bug fix release with 170+ PRs. MLX fixed - Some MLX and Mac runtimes did not run correctly LAN API keyless / password-less + Keyboard shortcuts XET / HTTP download toggle - clearer download progress AMD bug fixes + 170 bug, reliability & performance fixes Features Auto Compaction (Experimental) for longer chats beyond context li…

U
Unsloth AI

Bug Fixes + Auto compaction + LAN Remote Access

Thanks for the support for Qwen3.8-27B and Unsloth Desktop! This is a bug fix release with 170+ PRs. MLX fixed - Some MLX and Mac runtimes did not run correctly LAN API keyless / password-less is now supported XET / HTTP download toggle - clearer download progress AMD bug fixes for Strix Halo, all RDNA GPUs + 170 bug fixes Features Auto Compaction (Experimental) for longer chats beyond context lim…

U
Unsloth AI

Auto compaction (preview) + LAN Remote Access

Thanks for the support for Qwen3.8-27B and Unsloth Desktop last week! For this release, we merged 200+ PRs to introduce many new features, fixes including: Auto Compaction (Experimental) for longer chats beyond context limits Remote & LAN Access (Preview) for easy network access without Cloudflare links Faster Chat - Improved streaming performance, reduced UI lag, and smoother long conversations.…

U
Unsloth AI

pre-split-full: Merge main into studio-mmproj-fit

Three conflicts, all small. The llama_cpp.py one is two independent additions at the same spot, the projector pin state and the Metal context refusal; both are kept. The mmproj-fallback test file is likewise both sides' tests, main's loadFallbackNotice coverage alongside the wording assertion here. image-input-support.ts takes main's line. The missing .ts on that runtime import is what made the fr…

U
Unsloth AI

Qwen3.8-27B

Qwen3.8-27B and Qwen3.8-2.4T can now be run locally in Unsloth! Run on 17GB RAM via Unsloth Dynamic GGUFs. You can also fine-tune Qwen3.8-27B in Unsloth. Qwen3.8-27B is by far the strongest model for its size. We also uploaded NVFP4 quants. Guide: https://unsloth.ai/docs/models/qwen3.8 GGUF: https://huggingface.co/unsloth/Qwen3.8-27B-GGUF See 1-bit Qwen3.8-2.4T GGUF running in Unsloth: Highlights…

U
Unsloth AI v0.1.71-beta

v0.1.71-beta

Offer the media pickers only what the host can run, and name the H3 s…

U
Unsloth AI v0.1.702-beta

v0.1.702-beta

Unsloth Desktop is here! The first desktop app to run and train AI models locally. Research, export and deploy from the same open-source app on Windows, macOS and Linux. v0.1.702-beta Update (August 13th) Added tool calling / web search & more for all external providers Fixed bypass permissions not working for sandboxing UI and UX fixes - VRAM usage is now tunable 10% faster inference + reduced VR…

U
Unsloth AI

Introducing Unsloth Desktop 🦥

Unsloth Desktop is here! The first desktop app to run and train AI models locally. Research, export and deploy from the same open-source app on Windows, macOS and Linux. 🦥 Download Unsloth Desktop for Linux, Windows, MacOS Here's what you can do with Unsloth Desktop: Get up to 50% more accurate tool calling with self-healing calls and sandboxed code execution. Run Muse Glimmer 30B, Kimi K3, Qwen3.…

U
Unsloth AI

Introducing Unsloth Desktop 🦥

Unsloth Desktop is here! The first desktop app to run and train AI models locally. Research, export and deploy from the same open-source app on Windows, macOS and Linux. 🦥 Download Unsloth Desktop for Linux, Windows, MacOS Here's what you can do with Unsloth Desktop: Get up to 50% more accurate tool calling with self-healing calls and sandboxed code execution. Run Muse Glimmer 30B, Kimi K3, Qwen3.…

U
Unsloth AI

Meta Muse Glimmer

Meta has released Muse Glimmer, the first open model from Meta Superintelligence Labs. Muse Glimmer is a dense, 30B parameter model by Meta, built for local agentic and coding workflows. Released under the Apache 2.0 license you can run Muse Glimmer 30B via Unsloth and Unsloth Dynamic quants for the best performance. Run Muse Glimmer Fine-tune Muse Glimmer Muse Glimmer 30B can run locally on 20GB…

U
Unsloth AI

Meta Muse Glimmer

Meta has released Muse Glimmer, the first open model from Meta Superintelligence Labs. Muse Glimmer is a dense, 30B parameter model by Meta, built for local agentic and coding workflows. Released under the Apache 2.0 license you can run Muse Glimmer 30B via Unsloth and Unsloth Dynamic quants for the best performance. Run Muse Glimmer Fine-tune Muse Glimmer Muse Glimmer 30B can run locally on 20GB…

U
Unsloth AI

DSpark + DeepSeek-V4 Flash 0731

Hey everyone! For folks who missed the news - Kimi K3 & DeepSeek v4 Flash can now run locally with Unsloth Dynamic GGUFs! We now added more efficient and faster downloading for Colab, low memory systems and also high memory and CPU systems - we auto fallback to HTTP as well if XET is stuck August 7th Update Many bug fixes + Mac fixes and smoother installations. DeepSeek V4 Flash 0731 + DSpark You…

U
Unsloth AI v0.1.527-beta

Unsloth v0.1.527-beta

What's Changed Studio: judge the cached-pipeline signals on the snapshot the row actually loads by @danielhanchen in #7851 Bump install.sh / install.ps1 pin to unsloth>=2026.8.3 by @danielhanchen in #7860 Studio: make the sidebar menu separators visible in dark mode by @shimmyshimmer in #7865 Studio: run sandbox matplotlib headless by @NilayYadav in #7789 Don't send desktop startup requests to an…