Unsloth v0.1.901-beta Released: Major LoRA Memory Reductions

At a Glance
| Item | Value |
|---|---|
| Repository | unslothai/unsloth |
| Version | v0.1.901-beta |
| Published | 2026-10-01 |
| License | Apache-2.0 |
| Source type | Primary source (the publisher itself) |
Values determined by this site’s code when the information was collected. Dates are JST.
Overview
The latest version “v0.1.901-beta" of the Python tool “unslothai/unsloth" (License: Apache-2.0) for running and fine-tuning LLMs and diffusion models locally has been released. This tool is characterized by its support for a wide range of models and formats, including GGUF, MLX, Qwen3.8, DeepSeek-V4, MiniMax-H3, Gemma 4, and FLUX.
The most significant changes in this update are the improved usability brought by the introduction of the command palette in the desktop version and a dramatic reduction in memory usage for 4-bit checkpoints (NVFP4, INT4, MXFP4) during LoRA training. This allows for more efficient fine-tuning of larger models even in environments with limited GPU memory.
Key Changes
Desktop UI and Usability Improvements
In the desktop version, a command palette that can be opened with the keyboard shortcut Cmd/Ctrl+P has been added, significantly speeding up in-app navigation. It is also now possible to share GGUF model execution settings as links and to save settings without loading the model.
Additionally, chat features have been enhanced with a “Continue response" feature to extend cut-off answers, and a “Fix with the model" feature that allows users to insert HTML canvas errors into the message box to ask the model for fixes. Features to display warnings when models only partially fit into GPU memory and to prioritize thumbnail loading in the image gallery have also been implemented, improving the user experience.
Significant Reduction in Memory Usage for LoRA Training
It is now possible to directly train pre-quantized 4-bit checkpoints (NVFP4, INT4, MXFP4) using their original published weights. Because these are loaded in a packed state during training, the loaded model consumes almost the same amount of memory as a 4-bit checkpoint.
Specifically, the following improvements have been reported:
– Qwen3.8-27B-NVFP4: Peak memory on a single RTX PRO 6000 was reduced by 45% from 72.9 GB to 40.2 GB, and step time was reduced by 11%.
– Qwen3-8B w4a16: On B200, peak memory was reduced from the traditional 18.8 GB to 8.5 GB.
– Kimi-K2.7-Code: Where a 16-bit copy required about 2 TB of memory, it can now be trained on 6 B200s with the INT4 experts packed.
– gpt-oss-120b: Training is now possible on a single B200 while keeping MXFP4 experts.
Note that when using gpt-oss for inference only, if you want to revert to native loading to improve prefill speed, set the environment variable UNSLOTH_MXFP4_KEEP_PACKED=0.
Laya Decision Model Speedup and Decision API Expansion
Processing speed for short requests using Laya decision models has been accelerated by up to 4.1x. Additionally, Laya’s load time has been reduced from 17 seconds to about 9 seconds, and peak host RAM usage has been cut by about half.
Furthermore, connection to hosted decision providers such as TypeSafe is now supported, allowing them to be added from “Connections" and models to be selected in “Settings > API". When the Decision API is enabled, enabling “Unsloth Decisions" in the chat MCP menu makes the decide tool available.
Image and Video Processing Optimization
When offloading image models, INT8 or FP8 precision is now maintained. This reduced Z-Image processing time from 38.4 seconds to 2.3 seconds on a B200 with memory limited to 10 GiB. In addition, rendering speed for 1536×1536 images using Qwen-Image-2.1 has been accelerated by up to 3.6x on Radeon 8060S.
Expansion of Training Features and Supported Models
Using FastModel, it is now possible to fine-tune text encoder-decoder models such as T5, T5Gemma, BART, and Marian. A new unsloth eval command has also been added, allowing evaluation tasks via lm-evaluation-harness to be run on checkpoints or LoRA adapters. Furthermore, training and batch generation for Gemma 4 31B now work correctly on multi-GPU split models.
Supported Models and Hardware
This update brings support for numerous new model architectures and optimizations for specific hardware environments.
Details of Newly Supported Models
In addition to the text encoder-decoder models and Gemma 4 mentioned earlier, the following models are now newly supported:
– Idefics3 (Granite Docling VLM): Support for the Idefics3 architecture has been added via PR #4241.
– Qwen2.5 Coder: Qwen2.5 Coder has been officially added to the model registry and chat templates (PR #4221).
– Kimi-K3: Support added for training while keeping MXFP4 experts packed and dequantizing them at runtime (PR #11750).
– MiniCPM3: Support added for averaging GRPO gradients across DDP ranks and maintaining MiniCPM3’s unique pre-head scaling (PR #12193).
– Qwen3.8-Flash-Next: This model can now be loaded by keeping parameters without placement specifications, such as n-gram tables, on the CPU (PR #12141).
– Custom Vision Projectors: Models equipped with custom vision projectors are now supported in Unsloth Studio and Desktop (PR #11352).
Hardware-Specific Optimizations and Updates
Key changes for users of various platforms include:
- AMD (Windows): Graphics cards from the RX 5000 to RX 9000 series have been improved to fetch PyTorch from AMD’s multi-architecture index. This ensures proper installation on environments such as the RX 5700 XT (RDNA 1) (PR #11755, PR #11846). Additionally, in host environments mixing NVIDIA and AMD, llama.cpp automatically selects the backend to match the installed PyTorch.
- Apple Silicon (Mac): MLX libraries have been updated to 0.32.3 and mlx-vlm to 0.7.4. A feature has been added to export LoRA adapters fine-tuned on Macs into PEFT or GGUF formats for use on other GPU machines (PR #7539). Furthermore, batched responses are now possible when loading vision models with quantized KV caches.
- NVIDIA GPU: On B200, optimizations were made to maintain INT8 or FP8 precision when offloading image models. This significantly speeds up Z-Image processing in environments with memory limited to 10 GiB. Windows 11-style buttons have also been introduced to the UI on Windows and Linux environments.
How to Update
Depending on your environment, follow the steps below to update to the latest version.
Unsloth Desktop (GUI)
If you are using the desktop application version, download and run the latest installer for your platform from the official releases page:
– Windows: Unsloth-Desktop-Windows.exe
– macOS: Unsloth-Desktop-MacOS.dmg
– Linux (Ubuntu/deb): Unsloth-Desktop-Ubuntu.deb
– Linux (AppImage): Unsloth-Desktop-Linux.AppImage
Unsloth Studio (CLI/Development Version)
If you installed from source, pull the repository and then re-run the installation script.
For macOS, Linux, and WSL:
cd unsloth && git pull
./install.sh --local
unsloth studio -p 8888
For Windows (PowerShell):
cd unsloth; git pull
.\install.ps1 --local
unsloth studio -p 8888
If you want to force a specific backend (such as Vulkan), set the environment variable before installation:
export UNSLOTH_LLAMA_CPP_BACKEND=vulkan
./install.sh --local
Related Articles
- unsloth Desktop v0.1.900-beta Released with Laya and Speedups
- Unsloth Update: Qwen-Image-2.1 Support and Agent Skills Added
- Unsloth v0.1.811-beta Released with AMD & NVIDIA Docker Images
- Unsloth v0.1.810-beta Released with Multi-User and AMD Support
What to Read Next
- Follow this tool → Unsloth overview and release history (73 releases tracked)
- Other quantization, model formats and fine-tuning → ggml / ExLlamaV3
Sources
- https://github.com/unslothai/unsloth/releases/tag/v0.1.901-beta
- https://github.com/unslothai/unsloth/pull/4241
- https://github.com/unslothai/unsloth/pull/4221
- https://github.com/unslothai/unsloth/pull/11750
- https://github.com/unslothai/unsloth/pull/12193
- https://github.com/unslothai/unsloth/pull/12141
- https://github.com/unslothai/unsloth/pull/11352
- https://github.com/unslothai/unsloth/pull/11755
- https://github.com/unslothai/unsloth/pull/11846
- https://github.com/unslothai/unsloth/pull/7539
