vLLM v0.31.0 Released: Fast Restart and Hardware Optimization

At a Glance
| Item | Value |
|---|---|
| Repository | vllm-project/vllm |
| Version | v0.31.0 |
| Published | 2026-10-05 |
| License | Apache-2.0 |
| Source type | Primary source (the publisher itself) |
Values determined by this site’s code when the information was collected. Dates are JST.
Overview
The latest version of vLLM, v0.31.0, an inference engine that enables high-speed inference and serving for LLMs (Large Language Models), has been released. vLLM is an Apache-2.0 licensed open-source project written in Python, characterized by high throughput and memory efficiency.
The most significant impact on users in this version is the introduction of the new “Fast restart" feature. The vllm preload command now allows starting a daemon that keeps quantized weights in GPU memory, significantly reducing model loading times when restarting the engine.
Breaking Changes and Deprecations
In this version, several configuration options and features have been changed or removed due to security enhancements and internal structural reorganization. Particular caution is required when using multimodal models, as per-request argument specification is now denied by default.
| Item | Old Setting / Feature | New Setting / Migration Target |
|---|---|---|
| Multimodal arguments | Allowed per-request mm_processor_kwargs, etc. |
Denied by default (allowed via --trust-request-mm-kwargs) |
| Tokenizer mode | tokenizer_mode="slow" |
Removed (use default “hf") |
| Mamba cache configuration | --enable-mamba-fine-grained-prefix-cache |
--enable-mamba-shared-prefix-checkpoint |
| FP8 quantization specification | quantization="fp8" |
fp8_per_tensor (shorthand) |
| Quark quantization | Implicit online quantization | Removed (use vLLM standard online quantization API) |
| Inference backend | AllSpark INT8 W8A16 | Removed |
| Eager mode | --enforce-eager |
Also disables JIT kernel warm-up |
| XPU configuration | VLLM_XPU_ENABLE_XPU_GRAPH |
Removed (enabled by default) |
For users of multimodal models, per-request argument specification is restricted for security reasons. If the traditional behavior is required in a trusted environment, append the --trust-request-mm-kwargs option when starting the server.
Key Changes
Introduction of Fast Restart Feature
The newly introduced vllm preload CLI makes available the “weight-cache daemon", which keeps quantized weights resident in GPU memory. This dramatically reduces weight reloading wait times in development environments where engine restarts are repeated, or in operational environments where models are frequently switched. This feature also supports data parallelism (DP) and MTP draft models, and status monitoring via the /health endpoint is available.
Optimization for DeepSeek-V4.1-Flash
Numerous optimizations have been introduced for the latest DeepSeek-V4.1-Flash model to maximize hardware performance. FlashMLA mega attention and NVFP4 compressed KV cache are enabled by default on NVIDIA Blackwell (SM100), and advanced kernel fusions such as sparse MQA logit calculations using DeepGEMM and “Mega-Gate" (which fuses gate GEMM and expert selection) have been implemented. This improves inference speed on latest-generation GPUs.
Improved Security and Accuracy of Prefix Caching
The hash calculation logic for prefix caching has been improved. Previously, LoRA names and cache_salt were mixed without being tagged, which could lead to cache collisions between different LoRA adapters or salt specifications. Starting from this version, each element is tagged by source, and in the case of LoRA, the adapter path is also included in the hash. This resolves the issue where old caches were incorrectly reused even when reloading different adapters with the same name.
Greater Flexibility in Scheduling Control
A new option --max-num-active-seqs has been added, allowing the number of running sequences to be restricted independently of max_num_seqs. In addition, the waiting queue logic has been revamped so that requests already holding KV blocks are scheduled preferentially. This suppresses throughput drops and deadlocks in resource-constrained situations.
Model Runner V2 and Enhanced Speculative Decoding
In the internal architecture Model Runner V2, speculative decoding using draft models and custom logit processors are now supported. Furthermore, features to reduce inference latency have been expanded, such as the addition of the new draft method LiLiCorr and support for asynchronous scheduling in DFlash.
Supported Models and Hardware
This version brings support for numerous new model architectures and quantization formats, alongside significant hardware-specific optimizations.
New Supported Models and Feature Extensions
Inference support and feature extensions have been added for a variety of latest models.
- MiMo V2 / Cohere2MoE: MXFP4 MoE, BF16 MoE router, and DFlash draft models are now supported in MiMo V2. Cohere2MoE supports auxiliary hidden states for EAGLE3 and DFlash drafters.
- GLM-5.3-Flash / GLM-5.2: GLM-5.3-Flash introduces FlashAttention and FlashMLA sparse backends for SM90 (opt-in), as well as NoPE sparse MLA backends for SM120. Memory consumption has been reduced by 3 GiB through indexer decode workspace optimization, and metadata operations are 1.6x to 4.8x faster. Additionally, GLM-5.2-MXFP4 is available in the ROCm environment’s DeepSeek-V3.2 execution path.
- DeepSeek Series: DeepSeek V4 now supports fused inverse RoPE and FP8 quantization via FlashInfer sparse MLA, FIM (Fill-in-the-Middle) completion using the
suffixargument, and adding tool definitions to existing system messages. AMD-Quark mixed-precision checkpoints, such as DeepSeek-V4.1-Flash-MXFP4 and GLM-5.3-Flash Quark MXFP4, are also supported. - Qwen3.8-Flash-Next (Qwen4Exp): FP8 main KV cache in the QSA path and FP8 tensor parallelism via FlashInfer TRTLLM MoE are supported. Optimizations to reduce weight loading processing on DGX Spark by about 25 seconds are also included.
- Kimi K3 / MiniMax-M3: Kimi-K3 adds GEMM conversion for vision patch embeddings and fused kernels for KimiViT QK RoPE (up to 29x faster). MiniMax-M3 includes optimizations for
Conv3dLayerpatch embeddings (approx. 62x faster) and encoder CUDA graph support. - DiffusionGemma / Granite 4.2: DiffusionGemma implements structured generation mode with bounded single-choice tokens, speeding up constrained decoding by approximately 25%. A built-in
granite_thinking_parseris provided for Granite 4.2.
Expansion of Multimodal Models and LoRA
For multimodal models, audio models such as Whisper and Qwen2-Audio now support processing audio clips exceeding 30 seconds. Qwen2.5-VL has been improved to respect video fps settings for temporal M-RoPE, and Gemma4 gains support for FP8 KV cache with a head dimension of 512.
The scope of LoRA adapter support has also expanded, enabling LoRA to be applied to Nemotron VL language models, ModernBert, VoyageQwen3 embedding models, and RoBERTa sequence classification models.
Advancements in Hardware Support
Optimizations have focused intensively on NVIDIA’s latest-generation architectures, SM100 and SM103 (Blackwell), implementing NVFP4 compressed KV caches, MXFP8-related fused kernels, and low-latency reduce-scatter processing utilizing multi-memory. Matrix multiplication kernel settings for SM120 have also been added.
In addition, XPU graphs are now enabled by default in Intel XPU environments, removing the need for explicit specification via environment variables. Compiled packages for ROCm 7.2.3 are also provided for AMD ROCm environments.
How to Get It
To update an existing environment to v0.31.0, upgrade using the recommended package manager uv or pip.
For a standard CUDA 13.0 environment, run the following command:
uv pip install vllm --torch-backend=auto
When using standard pip, run the following command:
pip install --upgrade vllm
To install builds for ROCm environments, specify the dedicated wheel index for installation:
pip install vllm --extra-index-url https://wheels.vllm.ai/rocm/0.31.0/rocm723
If using a container environment, you can pull and update the official Docker image:
docker pull vllm/vllm-openai:v0.31.0
Related Articles
- vLLM v0.30.0 Released: Fast Start Weight Caching and New Models
- vLLM v0.29.0 Released: Model Runner V2 Default and More
- DeepSeek-V4-Pro-0813 Text Generation Model: ~998GB Memory, GGUF Builds
- K2-Horizon-32B-NVFP4: Our Test Answers, 32GB+ VRAM
What to Read Next
- Follow this tool → vLLM overview and release history (98 releases tracked)
- Other inference engines and runtimes → llama.cpp / Ollama / SGLang
Sources
- vllm-project/vllm v0.31.0 Release Notes
- GitHub PR #58830: Gate per-request multimodal processor kwargs
- GitHub PR #51899: Tag prefix-cache extra keys by source
- GitHub PR #59335: Include the LoRA path in prefix-cache block hashes
- GitHub PR #57833: Prefer fresh multimodal payloads over a stale receiver cache
- GitHub PR #58545: Remove the slow tokenizer mode
- GitHub PR #57382: Rename –enable-mamba-fine-grained-prefix-cache
- GitHub PR #53585: Remove online quantization support in fp8.py
- GitHub PR #51800: Remove quark-specific silent online quantization

