sglang v0.5.21 Brings New Models, Performance Gains, and Low-Latency APIs
sglang v0.5.21 introduces support for top open LLMs and diffusion models, major API updates for low-latency scoring/classification, performance boosts for popular models, and a new Rust prefix cache.
What changed?
sglang v0.5.21 is a significant release with: - Support for new models, including DeepSeek-V4.1 Flash, GigaChat 3.5, IQuest-Q1, MiMo-V2.6/Pro, Ling-3.0-flash-VL (LLMs/VLMs), plus major diffusion models like DiffusionGemma, Qwen-Image 2.1, and others. - New Decisions API (`/v1/decisions`) to turn any LLM/VLM into a low-latency classifier or scorer. - Score API (`/v1/score`) now allows scoring all candidates in a single request. - PD (prefill/decode) instances can now switch roles at runtime, no restart needed. - The prefix cache migrates to a Rust core, improving efficiency. - DeepSeek-V4.1 sees a 22% faster first-token latency on long prompts; Kimi K3 gets 20.6% faster prefill throughput. - Improvements in pipeline parallelism, expert parallelism and MoE support, especially for large scale deployment.
Why does it matter to an everyday developer?
This update matters for multiple reasons: - **Broader Model Access:** Out-of-the-box access to top recent LLMs/VLMs and diffusion models means faster iteration and richer application options. - **Low-Latency Scoring/Classifiers:** The new Decisions and Score APIs let you build fast, real-time classifiers or candidate ranking endpoints, enabling applications like rapid moderation, reranking, or content filtering. - **Performance Gains:** Whether serving DeepSeek-V4.1, Kimi K3, or complex pipeline/expert parallel setups, you'll see tangible throughput and latency improvements—this directly allows you to support more users or faster responses. - **Operational Flexibility:** Being able to switch PD roles at runtime and having a faster, more robust Rust cache reduces downtime and maximizes GPU utilization, which is important for reliability and scaling.
What can the developer do now?
Developers can now: - Start using or fine-tuning the new supported models with sglang to experiment with state-of-the-art text, vision-language, or image generation tasks. - Leverage the new Decisions API to turn any LLM/VLM into a low-latency scoring or classification service (see the [`/v1/decisions` docs](https://docs.sglang.io/docs/supported-models/decision_models)). - Use the Score API to efficiently evaluate/rerank candidates in bulk operations—particularly useful for search, recommendation, or retrieval applications. - Upgrade existing deployments for performance leaps, especially if using DeepSeek-V4.1 or Kimi K3 models. - Adopt or test the new Rust prefix cache for enhanced efficiency. To upgrade: install the new version via pip or use new Docker images matching your hardware (CUDA, ROCm, Intel GPU/CPU). Detailed upgrade instructions are in the [release notes](https://github.com/sgl-project/sglang/releases/tag/v0.5.21).
