v0.31.0: [Misc] Add Transformers version upper bound in requirements (#59614)
Signed-off-by: Isotr0py Isotr0py@outlook.com Co-authored-by: Harry Mellor 19981378+hmellor@users.noreply.github.com (cherry picked from commit 58b3298)
Signed-off-by: Isotr0py Isotr0py@outlook.com Co-authored-by: Harry Mellor 19981378+hmellor@users.noreply.github.com (cherry picked from commit 58b3298)
Signed-off-by: Isotr0py Isotr0py@outlook.com Co-authored-by: Harry Mellor 19981378+hmellor@users.noreply.github.com (cherry picked from commit 58b3298)
[Bugfix][HiSparse] Fix MTP acceptance collapse under FULL graphs with…
Signed-off-by: Robert Shaw robshaw@redhat.com Co-authored-by: Robert Shaw robshaw@redhat.com Co-authored-by: Claude Opus 5.5 noreply@anthropic.com Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com> Co-authored-by: Nick Hill nickhill123@gmail.com (cherry picked from commit e5e38ba)
Release vllm-proto 0.4.0
[Bugfix][Mamba] Keep the prompt-end prefill checkpoint under sparse r…
Signed-off-by: khluu khluu000@gmail.com Co-authored-by: Claude Opus 5.5 noreply@anthropic.com (cherry picked from commit aedaba8)
Signed-off-by: Andreas Karatzas akaratza@amd.com Co-authored-by: OpenAI Codex noreply@openai.com
v0.30.0 Highlights This release features 762 commits from 315 contributors (104 new)! New models: DeepSeek-V4.1-Flash (#56214, #56228, #56208) with the whole KV stored in MXFP8 through the FlashMLA V4.1 record on SM100 (#56893), DeepGEMM Mega-mHC (#56962), and async Engram prefetch with Engram DP sharding (#56512); DeepSeek-V4-Flash-Vision-Exp (#54566), also on ROCm (#55107) and with LoRA (#55897)…
[Bugfix][NIXL] Avoid receive reports for notification-only requests (…
Signed-off-by: jiahanc 173873397+jiahanc@users.noreply.github.com Co-authored-by: OpenAI Codex codex@openai.com
Release vllm-proto 0.3.0
Validated by PR #56538 CI at fa2a26f.
[watermarking] Dual-key gumbel-max watermarking for speculative decod…
vllm-proto 0.1.0
v0.29.0 Highlights This release features 594 commits from 277 contributors (91 new)! Model Runner V2 is now the default for all models (#53183), completing the rollout that began with pooling models (#48290). MRV2 also gained CUDA graph memory profiling for KV cache auto-sizing (#53306), batch-sharded sampling that cuts per-step logits memory by 1/TP (#50465), prompt embeds (#42963), extract_hidde…
[Bugfix][Core] Apply dense prefix cache default to hybrid models (#55…
[Core] Default prefix_cache_retention_interval to dense for Mamba + E…
Generated-by: Codex codex@openai.com Signed-off-by: Codex codex@openai.com
[CI] Remove deleted nvidia/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-BF1…
[Bugfix][Multimodal] Handle prefix-covered items in SHM worker cache …
Signed-off-by: Kevin Luu 51931015+khluu@users.noreply.github.com Signed-off-by: Yongye Zhu zyy1102000@gmail.com Co-authored-by: Codex codex@openai.com Co-authored-by: Yongye Zhu zyy1102000@gmail.com
[Tools][Recipes] Improve sweep recommendations and short-alias parsin…
v0.28.0 Highlights This release features 584 commits from 270 contributors (76 new)! Kimi-K3 performance push: a major optimization effort for Kimi-K3 across the stack — Decode Context Parallel (DCP) support (#50484), fused FlashKDA decode and prefill kernels (#50654, #51311, #52458), SiTU activation support for MegaMoE (#50510), GEMM-RS for sequence parallelism (#52079), combined all-gathers with…
(cherry picked from commit b389ac2) Signed-off-by: khluu khluu000@gmail.com
[Bugfix][Security] Guard _load_ov2_processor with resolve_trust_remot…
Signed-off-by: Lucas Wilkinson lwilkins@redhat.com Signed-off-by: Lucas Wilkinson LucasWilkinson@users.noreply.github.com Signed-off-by: Benjamin Chislett chislett.ben@gmail.com Signed-off-by: Lucas Wilkinson wilkinson.lucas@gmail.com Signed-off-by: Nick Hill nickhill123@gmail.com Co-authored-by: OpenAI Codex codex@openai.com Co-authored-by: Claude Opus 5 (1M context) noreply@anthropic.com Co-auth…
This is a patch release on top of v0.27.0. Support quantized DSpark Markov heads (#50424)
vLLM v0.27.0 Release Notes Highlights This release features 561 commits from 242 contributors (64 new)! Kimi K3 support with a full stack landing in one release: core model files and kernels (#50089, #50000), Python (#50093) and Rust (#50104) frontends, AttnRes kernels (#50090), DeepGEMM support (#50458), compressed-tensors quantized checkpoints (#50500), DSpark AR fusion (#50242), and an option t…
v0.27.0rc2