V

vLLM

V
vLLM AI v0.31.0rc4

v0.31.0rc4

[Bugfix][HiSparse] Fix MTP acceptance collapse under FULL graphs with…

V
vLLM AI v0.31.0rc3

v0.31.0rc3: [Model Runner V2] Support randomized dummy inputs (#58411)

Signed-off-by: Robert Shaw robshaw@redhat.com Co-authored-by: Robert Shaw robshaw@redhat.com Co-authored-by: Claude Opus 5.5 noreply@anthropic.com Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com> Co-authored-by: Nick Hill nickhill123@gmail.com (cherry picked from commit e5e38ba)

V
vLLM AI v0.31.0rc2

v0.31.0rc2

[Bugfix][Mamba] Keep the prompt-end prefill checkpoint under sparse r…

V
vLLM AI v0.30.0

v0.30.0

v0.30.0 Highlights This release features 762 commits from 315 contributors (104 new)! New models: DeepSeek-V4.1-Flash (#56214, #56228, #56208) with the whole KV stored in MXFP8 through the FlashMLA V4.1 record on SM100 (#56893), DeepGEMM Mega-mHC (#56962), and async Engram prefetch with Engram DP sharding (#56512); DeepSeek-V4-Flash-Vision-Exp (#54566), also on ROCm (#55107) and with LoRA (#55897)…

V
vLLM AI v0.30.0rc2

v0.30.0rc2

[Bugfix][NIXL] Avoid receive reports for notification-only requests (…

V
vLLM AI v0.29.1rc0

v0.29.1rc0

[watermarking] Dual-key gumbel-max watermarking for speculative decod…

V
vLLM AI v0.29.0

v0.29.0

v0.29.0 Highlights This release features 594 commits from 277 contributors (91 new)! Model Runner V2 is now the default for all models (#53183), completing the rollout that began with pooling models (#48290). MRV2 also gained CUDA graph memory profiling for KV cache auto-sizing (#53306), batch-sharded sampling that cuts per-step logits memory by 1/TP (#50465), prompt embeds (#42963), extract_hidde…

V
vLLM AI v0.29.0rc6

v0.29.0rc6

[Bugfix][Core] Apply dense prefix cache default to hybrid models (#55…

V
vLLM AI v0.29.0rc5

v0.29.0rc5

[Core] Default prefix_cache_retention_interval to dense for Mamba + E…

V
vLLM AI v0.29.0rc3

v0.29.0rc3

[CI] Remove deleted nvidia/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-BF1…

V
vLLM AI v0.29.0rc2

v0.29.0rc2

[Bugfix][Multimodal] Handle prefix-covered items in SHM worker cache …

V
vLLM AI v0.28.1rc0

v0.28.1rc0

[Tools][Recipes] Improve sweep recommendations and short-alias parsin…

V
vLLM AI v0.28.0

v0.28.0

v0.28.0 Highlights This release features 584 commits from 270 contributors (76 new)! Kimi-K3 performance push: a major optimization effort for Kimi-K3 across the stack — Decode Context Parallel (DCP) support (#50484), fused FlashKDA decode and prefill kernels (#50654, #51311, #52458), SiTU activation support for MegaMoE (#50510), GEMM-RS for sequence parallelism (#52079), combined all-gathers with…

V
vLLM AI v0.28.0rc1

v0.28.0rc1

[Bugfix][Security] Guard _load_ov2_processor with resolve_trust_remot…

V
vLLM AI v0.27.2rc0

v0.27.2rc0: [Spec Decode] DSpark confidence-scheduled verification (#47808)

Signed-off-by: Lucas Wilkinson lwilkins@redhat.com Signed-off-by: Lucas Wilkinson LucasWilkinson@users.noreply.github.com Signed-off-by: Benjamin Chislett chislett.ben@gmail.com Signed-off-by: Lucas Wilkinson wilkinson.lucas@gmail.com Signed-off-by: Nick Hill nickhill123@gmail.com Co-authored-by: OpenAI Codex codex@openai.com Co-authored-by: Claude Opus 5 (1M context) noreply@anthropic.com Co-auth…

V
vLLM AI v0.27.1

v0.27.1

This is a patch release on top of v0.27.0. Support quantized DSpark Markov heads (#50424)

V
vLLM AI v0.27.0

v0.27.0

vLLM v0.27.0 Release Notes Highlights This release features 561 commits from 242 contributors (64 new)! Kimi K3 support with a full stack landing in one release: core model files and kernels (#50089, #50000), Python (#50093) and Rust (#50104) frontends, AttnRes kernels (#50090), DeepGEMM support (#50458), compressed-tensors quantized checkpoints (#50500), DSpark AR fusion (#50242), and an option t…