unsloth Desktop v0.1.900-beta Released with Laya and Speedups

unsloth Desktop v0.1.900-beta Released with Laya and Speedups

At a Glance

Item Value
Repository unslothai/unsloth
Version v0.1.900-beta
Published 2026-09-28
Source type Primary source (the publisher itself)

Values determined by this site’s code at collection time. Dates are JST.

Overview

unsloth (unslothai/unsloth), known as a fine-tuning and inference tool for local LLMs, has updated its Desktop app to v0.1.900-beta. The highlight of this release is the ability to locally run and serve Laya, one of the “Decision Models" based on the open-source Jev. In addition, it brings a “Library" feature to manage documents and media together, a PDF/Office file viewing feature, a Skills Editor to directly edit skills within the app, and massive speedups for image and video generation (approximately 4.5x faster for LTX-2.3 clips and 1.7 to 6.3x faster for VAE decoding). For those already using unsloth Desktop on a daily basis, the image and video generation speedups alone make this update worth taking.

Breaking Changes and Deprecations

This release notes no breaking changes or deprecations that would replace existing options, APIs, or settings.

Key Changes

Local Execution of Decision Models via Laya

You can now run Laya locally, which is one of the “Decision Models" that answers questions with probabilities for yes/no, multiple-choice, and scoring tasks. To use it, enable the Decision API via Settings > API and select the model to use along with whether to run it on the CPU or GPU. If using the TypeSafe SDK, it can be accessed via the Jev-compatible /v1/systemone endpoint. When GPU is selected, it is reported to run natively using MLX on Apple Silicon. Details can be found in the related PR (#11603). For those who have been replacing classification and judgment tasks with existing generative models, this is worth considering as a replacement in terms of accuracy and speed.

Addition of Skills Editor

A Skills Editor has been added, allowing you to directly create, edit, and delete Skills within the Desktop app. Workflows that previously required editing files externally and loading them can now be completed entirely within the app.

Support for Downloading Models from ModelScope

For users in environments unable to access Hugging Face, downloading models via ModelScope is now supported. This increases options for acquiring models for users who could not use HF due to regional restrictions or other reasons.

Library + Document Viewer

A Library / Document Viewer tab has been added to view PDF, Word, Excel, and PowerPoint files directly inside unsloth. Chats, images, videos, and more can also be centrally managed here, and attached file card displays have been improved. Users can now create their own sidebar sections and reorder them by dragging. This directly impacts users who employ workflows involving reading and interacting with documents.

Improvements for Apple Silicon

Several improvements for Apple Silicon environments are included, such as batched serving, structured outputs, and TurboQuant KV cache. Those serving models on a Mac should check for changes in memory efficiency and response speed.

Model Allocation Settings per GPU

In multi-GPU environments, you can now set which parts of the model are allocated to each GPU. This is relevant for those distributing models across multi-GPU configurations.

Speedup of Image and Video Generation

LTX-2.3 clip generation is reported to be approximately 4.5x faster, driven by distilled sampling, compilation-related fixes, and host-provided FP8 weights. Image and video VAE decoding is also 1.7 to 6.3x faster, which is said to correspondingly improve overall generation speed in workflows. Furthermore, initial rendering for MiniMax-H3 is reported to be up to 1 minute faster. Users generating images and videos with LTX-2.3 and MiniMax-H3 should be able to notice the improvement in perceived speed.

Addition of New Themes

Several new food-themed options have been added.

Specifications and Supported Hardware

Among Decision Models, Laya, which is based on the open-source Jev, can now be executed locally. When GPU is selected, it is reported to operate natively using MLX on Apple Silicon. Execution on CPU is also supported, and the execution method can be selected from Settings > API.

For image and video generation, LTX-2.3 clip generation is accelerated in combination with host-provided FP8 weights, implying that an environment capable of handling FP8 format models is assumed (detailed requirements are not explicitly stated in the release notes). The VAE decoding speedup applies generally to image and video generation, and reduced initial rendering times have also been reported for MiniMax-H3.

As a model distribution channel, downloading via ModelScope is now supported in addition to Hugging Face. Users who cannot access HF due to regional or environmental constraints can acquire models through this route.

For multi-GPU environments, a feature has been added to configure which parts of the model are assigned to each GPU, catering to setups where models are distributed across multiple GPUs.

How to Get It

Specific installation commands (such as pip install or git pull) are not mentioned in this material. If you are using the unsloth Desktop app, please get the latest version (v0.1.900-beta) through the in-app update function or from the releases page.

Note that newly added features such as the Decision API and Skills Editor need to be enabled and configured from the Settings screen after updating.

Releases Since Our Last Article

Compiled by Local Model Watch from the project’s GitHub releases: the versions between this release and the last one we covered, which did not get separate articles. Full history: release tracker.

Version Released Release notes
prebuilt-wheels-cu13 (Flash-Attention2, Causal-Conv1D, Mamba_SSM Binaries) 2026-09-27 GitHub
v0.1.815-beta (Qwen-Image-2.1 + Skills) 2026-09-23 GitHub
v0.1.814-beta (Qwen-Image-2.1 + Skills) 2026-09-23 GitHub
v0.1.813-beta (Qwen-Image-2.1 + Skills) 2026-09-23 GitHub

Related Articles

What to Read Next

Sources