{"id":9028,"date":"2026-10-02T11:11:13","date_gmt":"2026-10-02T02:11:13","guid":{"rendered":"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/10\/02\/sglang-v0-5-21-released-2\/"},"modified":"2026-10-02T11:11:13","modified_gmt":"2026-10-02T02:11:13","slug":"sglang-v0-5-21-released-2","status":"publish","type":"post","link":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/10\/02\/sglang-v0-5-21-released-2\/","title":{"rendered":"SGLang v0.5.21 Released with Dynamic PD Role Switching"},"content":{"rendered":"<p><!-- lmw:facts --><\/p>\n<h2>At a Glance<\/h2>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Item<\/th>\n<th>Value<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>Repository<\/td>\n<td><a href=\"https:\/\/github.com\/sgl-project\/sglang\">sgl-project\/sglang<\/a><\/td>\n<\/tr>\n<tr>\n<td>Version<\/td>\n<td><a href=\"https:\/\/github.com\/sgl-project\/sglang\/releases\/tag\/v0.5.21\">v0.5.21<\/a><\/td>\n<\/tr>\n<tr>\n<td>Published<\/td>\n<td>2026-10-02<\/td>\n<\/tr>\n<tr>\n<td>License<\/td>\n<td>Apache-2.0<\/td>\n<\/tr>\n<tr>\n<td>Source type<\/td>\n<td>Primary source (the publisher itself)<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p><em>Values determined by this site&#8217;s code when the information was collected. Dates are JST.<\/em><\/p>\n<p><!-- \/lmw:facts --><\/p>\n<h2>Overview<\/h2>\n<p>SGLang is a high-performance serving framework for large language models (LLMs) and multimodal models. Version 0.5.21 has been released in this update.<\/p>\n<p>The most significant change in this release is that PD (Prefill\/Decode) instances can now dynamically switch between Prefill and Decode states during execution without restarting. This is expected to significantly improve the utilization efficiency of inference resources.<\/p>\n<h2>Breaking Changes &amp; Deprecations<\/h2>\n<p>This update removes features that were deprecated in past releases as well as experimental features. Caution is required in environments using these.<\/p>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th style=\"text-align: left;\">Change<\/th>\n<th style=\"text-align: left;\">Old State \/ Setting<\/th>\n<th style=\"text-align: left;\">New State<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td style=\"text-align: left;\">Removal of experimental C++ radix tree<\/td>\n<td style=\"text-align: left;\"><code>SGLANG_EXPERIMENTAL_CPP_RADIX_TREE<\/code> environment variable<\/td>\n<td style=\"text-align: left;\">Unavailable<\/td>\n<\/tr>\n<tr>\n<td style=\"text-align: left;\">Removal of specific Radix Cache implementations<\/td>\n<td style=\"text-align: left;\"><code>SWARadixCache<\/code> \/ <code>MambaRadixCache<\/code><\/td>\n<td style=\"text-align: left;\">Integrated into <code>UnifiedRadixCache<\/code><\/td>\n<\/tr>\n<tr>\n<td style=\"text-align: left;\">Removal of unused HiRadixCache<\/td>\n<td style=\"text-align: left;\"><code>HiRadixCache<\/code><\/td>\n<td style=\"text-align: left;\">Unavailable<\/td>\n<\/tr>\n<tr>\n<td style=\"text-align: left;\">Removal of past deprecated endpoints<\/td>\n<td style=\"text-align: left;\">APIs\/environment variables\/aliases deprecated for 2+ releases<\/td>\n<td style=\"text-align: left;\">Unavailable<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p>These primarily affect developers and users operating with specific experimental cache settings explicitly specified.<\/p>\n<h2>Key Changes<\/h2>\n<h3>Dynamic Role Switching for PD Instances<\/h3>\n<p>In PD (Prefill\/Decode) instances, Prefill and Decode roles can now be switched on-the-fly without restarting the instance. This enables flexible adaptation to changes in inference workloads.<\/p>\n<h3>Rust-Based Core for Prefix Cache<\/h3>\n<p>Prefix Cache now operates on a Rust-based core by default. This aims to improve cache management efficiency and performance.<\/p>\n<h3>Introduction of New APIs (Decisions API \/ Score API)<\/h3>\n<p>New APIs have been added to utilize LLMs and VLMs as low-latency classifiers or scorers.<br \/>\n&#8211; <code>\/v1\/decisions<\/code> (Decisions API): Functions the model as a low-latency classifier or scorer.<br \/>\n&#8211; <code>\/v1\/score<\/code> (Score API): Calculates scores for all candidates in a single request.<\/p>\n<h3>Performance Improvements for Specific Models<\/h3>\n<ul>\n<li>DeepSeek-V4.1: First Token generation on long prompts is 22% faster.<\/li>\n<li>Kimi K3: Prefill throughput in PD serving improved by 20.6%.<\/li>\n<li>LFM2-VL: Achieved 1.66x to 2.56x speedups at batch size 1 through the introduction of DSpark speculative decoding.<\/li>\n<\/ul>\n<h3>Parallelism and Communication Optimizations<\/h3>\n<p>Accuracy improvements and optimizations were made under PP (Pipeline Parallelism), DP (Data Parallelism), and CP (Context Parallelism) environments. Specifically, inter-layer communication is now directly controlled by SGLang, enabling more accurate results. Support for DeepEP v2 in MoE (Mixture-of-Experts) models and w4a8 MoE optimizations on H200 are also included.<\/p>\n<h2>Supported Models and Hardware<\/h2>\n<p>This update significantly expands support for a diverse range of model architectures, including the latest LLMs (Large Language Models), VLMs (Multimodal Models), and Diffusion models. This allows primary generative AI tasks\u2014such as text generation, image understanding, and image generation\u2014to be executed consistently on SGLang&#8217;s high-performance inference engine.<\/p>\n<h3>Newly Supported Models<\/h3>\n<p>The following models have been newly added and can now be inferred. Please refer to the official Cookbooks for specific usage instructions for each model.<\/p>\n<h4>LLM \/ VLM (Large Language Models and Multimodal Models)<\/h4>\n<p>Models for text generation and combined image-text understanding.<br \/>\n&#8211; DeepSeek-V4.1 Flash<br \/>\n&#8211; GigaChat 3.5<br \/>\n&#8211; IQuest-Q1<br \/>\n&#8211; MiMo-V2.6 \/ MiMo-V2.6-Pro<br \/>\n&#8211; Ling-3.0-flash-VL<\/p>\n<h4>Diffusion (Diffusion Models)<\/h4>\n<p>Diffusion models used for tasks such as image generation.<br \/>\n&#8211; DiffusionGemma<br \/>\n&#8211; Qwen-Image 2.1<br \/>\n&#8211; Anima Base v1.0<br \/>\n&#8211; Ming-Image 0.1 Design \/ Design-Layer<br \/>\n&#8211; FLUX 3 Action<\/p>\n<h3>Supported Hardware and Platforms<\/h3>\n<p>SGLang is designed to operate on diverse compute resources. This release provides optimized Docker images for the following platforms, allowing users to maximize hardware performance.<\/p>\n<ul>\n<li><strong>NVIDIA GPU<\/strong>: Supports operation in CUDA 13 environments.<\/li>\n<li><strong>AMD GPU<\/strong>: Supports both of AMD&#8217;s latest architectures, MI35x and MI30x, in ROCm 10 environments.<\/li>\n<li><strong>Intel GPU<\/strong>: Supports Intel XPU platforms.<\/li>\n<li><strong>Intel CPU<\/strong>: Supports CPU environments such as Intel Xeon processors.<\/li>\n<\/ul>\n<h2>How to Get It<\/h2>\n<p>Please update using one of the following methods according to your environment.<\/p>\n<h3>Updating via pip<\/h3>\n<p>If you are using <code>uv<\/code> for package management, you can install version 0.5.21 directly with pre-releases allowed by running the following command:<\/p>\n<pre><code class=\"language-bash\">uv pip install --prerelease=allow sglang==0.5.21\n<\/code><\/pre>\n<h3>Using Docker Images<\/h3>\n<p>If you are using a container-based environment, it is recommended to pull and use the optimized Docker images for each platform below:<\/p>\n<ul>\n<li><strong>NVIDIA (CUDA 13)<\/strong>: <code>lmsysorg\/sglang:v0.5.21<\/code><\/li>\n<li><strong>AMD MI35x<\/strong>: <code>lmsysorg\/sglang:v0.5.21-rocm10-mi35x<\/code><\/li>\n<li><strong>AMD MI30x<\/strong>: <code>lmsysorg\/sglang:v0.5.21-rocm10-mi30x<\/code><\/li>\n<li><strong>Intel GPU<\/strong>: <code>lmsysorg\/sglang:v0.5.21-xpu<\/code><\/li>\n<li><strong>Intel CPU<\/strong>: <code>lmsysorg\/sglang:v0.5.21-xeon<\/code><\/li>\n<\/ul>\n<p><!-- lmw:related --><\/p>\n<h2>Related Articles<\/h2>\n<ul>\n<li><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/19\/sglang-v0-5-20-released\/\">SGLang v0.5.20 Released: New Models and Optimizations<\/a><\/li>\n<\/ul>\n<p><!-- \/lmw:related --><\/p>\n<p><!-- lmw:next-steps --><\/p>\n<h2>What to Read Next<\/h2>\n<ul>\n<li><strong>Follow this tool<\/strong> \u2192 <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/engine-sglang-en\/\">SGLang overview and release history (61 releases tracked)<\/a><\/li>\n<li><strong>Other inference engines and runtimes<\/strong> \u2192 <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/engine-llama-cpp-en\/\">llama.cpp<\/a> \/ <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/engine-ollama-en\/\">Ollama<\/a> \/ <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/engine-vllm-en\/\">vLLM<\/a><\/li>\n<\/ul>\n<p><!-- \/lmw:next-steps --><\/p>\n<h2>Sources<\/h2>\n<ul>\n<li><a href=\"https:\/\/github.com\/sgl-project\/sglang\/releases\/tag\/v0.5.21\">https:\/\/github.com\/sgl-project\/sglang\/releases\/tag\/v0.5.21<\/a><\/li>\n<\/ul>\n","protected":false},"excerpt":{"rendered":"<p>SGLang v0.5.21 is out, bringing dynamic PD role switching, a Rust-based prefix cache core, new APIs, and expanded model\/hardware support.<\/p>\n","protected":false},"author":1,"featured_media":9027,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[105],"tags":[163,165,169,1547,1193],"class_list":["post-9028","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-engines-and-tools","tag-gguf-en","tag-moe-en","tag-sglang-en","tag-verified","tag--en"],"lang":"en","translations":{"en":9028,"ja":9026},"_links":{"self":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts\/9028","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/comments?post=9028"}],"version-history":[{"count":0,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts\/9028\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/media\/9027"}],"wp:attachment":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/media?parent=9028"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/categories?post=9028"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/tags?post=9028"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}