SGLang v0.5.21 Released with Dynamic PD Role Switching

SGLang v0.5.21 Released with Dynamic PD Role Switching

At a Glance

Item Value
Repository sgl-project/sglang
Version v0.5.21
Published 2026-10-02
License Apache-2.0
Source type Primary source (the publisher itself)

Values determined by this site’s code when the information was collected. Dates are JST.

Overview

SGLang is a high-performance serving framework for large language models (LLMs) and multimodal models. Version 0.5.21 has been released in this update.

The most significant change in this release is that PD (Prefill/Decode) instances can now dynamically switch between Prefill and Decode states during execution without restarting. This is expected to significantly improve the utilization efficiency of inference resources.

Breaking Changes & Deprecations

This update removes features that were deprecated in past releases as well as experimental features. Caution is required in environments using these.

Change Old State / Setting New State
Removal of experimental C++ radix tree SGLANG_EXPERIMENTAL_CPP_RADIX_TREE environment variable Unavailable
Removal of specific Radix Cache implementations SWARadixCache / MambaRadixCache Integrated into UnifiedRadixCache
Removal of unused HiRadixCache HiRadixCache Unavailable
Removal of past deprecated endpoints APIs/environment variables/aliases deprecated for 2+ releases Unavailable

These primarily affect developers and users operating with specific experimental cache settings explicitly specified.

Key Changes

Dynamic Role Switching for PD Instances

In PD (Prefill/Decode) instances, Prefill and Decode roles can now be switched on-the-fly without restarting the instance. This enables flexible adaptation to changes in inference workloads.

Rust-Based Core for Prefix Cache

Prefix Cache now operates on a Rust-based core by default. This aims to improve cache management efficiency and performance.

Introduction of New APIs (Decisions API / Score API)

New APIs have been added to utilize LLMs and VLMs as low-latency classifiers or scorers.
– /v1/decisions (Decisions API): Functions the model as a low-latency classifier or scorer.
– /v1/score (Score API): Calculates scores for all candidates in a single request.

Performance Improvements for Specific Models

  • DeepSeek-V4.1: First Token generation on long prompts is 22% faster.
  • Kimi K3: Prefill throughput in PD serving improved by 20.6%.
  • LFM2-VL: Achieved 1.66x to 2.56x speedups at batch size 1 through the introduction of DSpark speculative decoding.

Parallelism and Communication Optimizations

Accuracy improvements and optimizations were made under PP (Pipeline Parallelism), DP (Data Parallelism), and CP (Context Parallelism) environments. Specifically, inter-layer communication is now directly controlled by SGLang, enabling more accurate results. Support for DeepEP v2 in MoE (Mixture-of-Experts) models and w4a8 MoE optimizations on H200 are also included.

Supported Models and Hardware

This update significantly expands support for a diverse range of model architectures, including the latest LLMs (Large Language Models), VLMs (Multimodal Models), and Diffusion models. This allows primary generative AI tasks—such as text generation, image understanding, and image generation—to be executed consistently on SGLang’s high-performance inference engine.

Newly Supported Models

The following models have been newly added and can now be inferred. Please refer to the official Cookbooks for specific usage instructions for each model.

LLM / VLM (Large Language Models and Multimodal Models)

Models for text generation and combined image-text understanding.
– DeepSeek-V4.1 Flash
– GigaChat 3.5
– IQuest-Q1
– MiMo-V2.6 / MiMo-V2.6-Pro
– Ling-3.0-flash-VL

Diffusion (Diffusion Models)

Diffusion models used for tasks such as image generation.
– DiffusionGemma
– Qwen-Image 2.1
– Anima Base v1.0
– Ming-Image 0.1 Design / Design-Layer
– FLUX 3 Action

Supported Hardware and Platforms

SGLang is designed to operate on diverse compute resources. This release provides optimized Docker images for the following platforms, allowing users to maximize hardware performance.

  • NVIDIA GPU: Supports operation in CUDA 13 environments.
  • AMD GPU: Supports both of AMD’s latest architectures, MI35x and MI30x, in ROCm 10 environments.
  • Intel GPU: Supports Intel XPU platforms.
  • Intel CPU: Supports CPU environments such as Intel Xeon processors.

How to Get It

Please update using one of the following methods according to your environment.

Updating via pip

If you are using uv for package management, you can install version 0.5.21 directly with pre-releases allowed by running the following command:

uv pip install --prerelease=allow sglang==0.5.21

Using Docker Images

If you are using a container-based environment, it is recommended to pull and use the optimized Docker images for each platform below:

  • NVIDIA (CUDA 13): lmsysorg/sglang:v0.5.21
  • AMD MI35x: lmsysorg/sglang:v0.5.21-rocm10-mi35x
  • AMD MI30x: lmsysorg/sglang:v0.5.21-rocm10-mi30x
  • Intel GPU: lmsysorg/sglang:v0.5.21-xpu
  • Intel CPU: lmsysorg/sglang:v0.5.21-xeon

Related Articles

What to Read Next

Sources