vLLM v0.31.0 Released: Fast Restart and Hardware Optimization

vLLM v0.31.0 Released: Fast Restart and Hardware Optimization

At a Glance

Item Value
Repository vllm-project/vllm
Version v0.31.0
Published 2026-10-05
License Apache-2.0
Source type Primary source (the publisher itself)

Values determined by this site’s code when the information was collected. Dates are JST.

Overview

The latest version of vLLM, v0.31.0, an inference engine that enables high-speed inference and serving for LLMs (Large Language Models), has been released. vLLM is an Apache-2.0 licensed open-source project written in Python, characterized by high throughput and memory efficiency.

The most significant impact on users in this version is the introduction of the new “Fast restart" feature. The vllm preload command now allows starting a daemon that keeps quantized weights in GPU memory, significantly reducing model loading times when restarting the engine.

Breaking Changes and Deprecations

In this version, several configuration options and features have been changed or removed due to security enhancements and internal structural reorganization. Particular caution is required when using multimodal models, as per-request argument specification is now denied by default.

Item Old Setting / Feature New Setting / Migration Target
Multimodal arguments Allowed per-request mm_processor_kwargs, etc. Denied by default (allowed via --trust-request-mm-kwargs)
Tokenizer mode tokenizer_mode="slow" Removed (use default “hf")
Mamba cache configuration --enable-mamba-fine-grained-prefix-cache --enable-mamba-shared-prefix-checkpoint
FP8 quantization specification quantization="fp8" fp8_per_tensor (shorthand)
Quark quantization Implicit online quantization Removed (use vLLM standard online quantization API)
Inference backend AllSpark INT8 W8A16 Removed
Eager mode --enforce-eager Also disables JIT kernel warm-up
XPU configuration VLLM_XPU_ENABLE_XPU_GRAPH Removed (enabled by default)

For users of multimodal models, per-request argument specification is restricted for security reasons. If the traditional behavior is required in a trusted environment, append the --trust-request-mm-kwargs option when starting the server.

Key Changes

Introduction of Fast Restart Feature

The newly introduced vllm preload CLI makes available the “weight-cache daemon", which keeps quantized weights resident in GPU memory. This dramatically reduces weight reloading wait times in development environments where engine restarts are repeated, or in operational environments where models are frequently switched. This feature also supports data parallelism (DP) and MTP draft models, and status monitoring via the /health endpoint is available.

Optimization for DeepSeek-V4.1-Flash

Numerous optimizations have been introduced for the latest DeepSeek-V4.1-Flash model to maximize hardware performance. FlashMLA mega attention and NVFP4 compressed KV cache are enabled by default on NVIDIA Blackwell (SM100), and advanced kernel fusions such as sparse MQA logit calculations using DeepGEMM and “Mega-Gate" (which fuses gate GEMM and expert selection) have been implemented. This improves inference speed on latest-generation GPUs.

Improved Security and Accuracy of Prefix Caching

The hash calculation logic for prefix caching has been improved. Previously, LoRA names and cache_salt were mixed without being tagged, which could lead to cache collisions between different LoRA adapters or salt specifications. Starting from this version, each element is tagged by source, and in the case of LoRA, the adapter path is also included in the hash. This resolves the issue where old caches were incorrectly reused even when reloading different adapters with the same name.

Greater Flexibility in Scheduling Control

A new option --max-num-active-seqs has been added, allowing the number of running sequences to be restricted independently of max_num_seqs. In addition, the waiting queue logic has been revamped so that requests already holding KV blocks are scheduled preferentially. This suppresses throughput drops and deadlocks in resource-constrained situations.

Model Runner V2 and Enhanced Speculative Decoding

In the internal architecture Model Runner V2, speculative decoding using draft models and custom logit processors are now supported. Furthermore, features to reduce inference latency have been expanded, such as the addition of the new draft method LiLiCorr and support for asynchronous scheduling in DFlash.

Supported Models and Hardware

This version brings support for numerous new model architectures and quantization formats, alongside significant hardware-specific optimizations.

New Supported Models and Feature Extensions

Inference support and feature extensions have been added for a variety of latest models.

  • MiMo V2 / Cohere2MoE: MXFP4 MoE, BF16 MoE router, and DFlash draft models are now supported in MiMo V2. Cohere2MoE supports auxiliary hidden states for EAGLE3 and DFlash drafters.
  • GLM-5.3-Flash / GLM-5.2: GLM-5.3-Flash introduces FlashAttention and FlashMLA sparse backends for SM90 (opt-in), as well as NoPE sparse MLA backends for SM120. Memory consumption has been reduced by 3 GiB through indexer decode workspace optimization, and metadata operations are 1.6x to 4.8x faster. Additionally, GLM-5.2-MXFP4 is available in the ROCm environment’s DeepSeek-V3.2 execution path.
  • DeepSeek Series: DeepSeek V4 now supports fused inverse RoPE and FP8 quantization via FlashInfer sparse MLA, FIM (Fill-in-the-Middle) completion using the suffix argument, and adding tool definitions to existing system messages. AMD-Quark mixed-precision checkpoints, such as DeepSeek-V4.1-Flash-MXFP4 and GLM-5.3-Flash Quark MXFP4, are also supported.
  • Qwen3.8-Flash-Next (Qwen4Exp): FP8 main KV cache in the QSA path and FP8 tensor parallelism via FlashInfer TRTLLM MoE are supported. Optimizations to reduce weight loading processing on DGX Spark by about 25 seconds are also included.
  • Kimi K3 / MiniMax-M3: Kimi-K3 adds GEMM conversion for vision patch embeddings and fused kernels for KimiViT QK RoPE (up to 29x faster). MiniMax-M3 includes optimizations for Conv3dLayer patch embeddings (approx. 62x faster) and encoder CUDA graph support.
  • DiffusionGemma / Granite 4.2: DiffusionGemma implements structured generation mode with bounded single-choice tokens, speeding up constrained decoding by approximately 25%. A built-in granite_thinking_parser is provided for Granite 4.2.

Expansion of Multimodal Models and LoRA

For multimodal models, audio models such as Whisper and Qwen2-Audio now support processing audio clips exceeding 30 seconds. Qwen2.5-VL has been improved to respect video fps settings for temporal M-RoPE, and Gemma4 gains support for FP8 KV cache with a head dimension of 512.

The scope of LoRA adapter support has also expanded, enabling LoRA to be applied to Nemotron VL language models, ModernBert, VoyageQwen3 embedding models, and RoBERTa sequence classification models.

Advancements in Hardware Support

Optimizations have focused intensively on NVIDIA’s latest-generation architectures, SM100 and SM103 (Blackwell), implementing NVFP4 compressed KV caches, MXFP8-related fused kernels, and low-latency reduce-scatter processing utilizing multi-memory. Matrix multiplication kernel settings for SM120 have also been added.

In addition, XPU graphs are now enabled by default in Intel XPU environments, removing the need for explicit specification via environment variables. Compiled packages for ROCm 7.2.3 are also provided for AMD ROCm environments.

How to Get It

To update an existing environment to v0.31.0, upgrade using the recommended package manager uv or pip.

For a standard CUDA 13.0 environment, run the following command:

uv pip install vllm --torch-backend=auto

When using standard pip, run the following command:

pip install --upgrade vllm

To install builds for ROCm environments, specify the dedicated wheel index for installation:

pip install vllm --extra-index-url https://wheels.vllm.ai/rocm/0.31.0/rocm723

If using a container environment, you can pull and update the official Docker image:

docker pull vllm/vllm-openai:v0.31.0

Related Articles

What to Read Next

Sources