vLLM v0.31.0: Major Serving Improvements, Speculative Decoding, and Breaking Changes
vLLM v0.31.0 introduces speculative decoding, broader model support, significant serving performance upgrades, and major API changes including deprecations and new defaults. Developers using vLLM for large model inference or multimodal workloads must review breaking changes and new features.
What changed?
vLLM v0.31.0 is a major release with over 700 commits, introducing several serving performance upgrades, speculative decoding capabilities, extended model and multimodal support, and important breaking changes in configuration and API behavior.
- Speculative decoding support using draft models, including LiLiCorr drafter and custom logits processors.
- Fast restart support: post-quantized weights can remain on GPU across engine restarts via `vllm preload`.
- Large-scale serving enhancements: sharded weight transfer, all2all communication backend, improved admission scheduling, and parallelism features.
- Expanded model coverage: optimized paths for DeepSeek-V4.1-Flash, GLM-5.3-Flash, Qwen3.8-Flash-Next, Kimi K3, MiniMax-M3, and others.
- More multimodal types supported and processing improved with new kernels.
- Scheduler controls improved with new flags and better queue handling.
- BREAKING: Gating of per-request multimodal kwargs—now requires `--trust-request-mm-kwargs`.
- BREAKING: Removed features and renamed flags (e.g., `tokenizer_mode="slow"` removed, quantization shorthand, backend removals).
- BREAKING: XPU graphs are enabled by default and their environment variable removed.
