fix(deps): update dependency vllm to >=0.26.0,<0.27 - #449
Open
aar-public-version-bump-bot[bot] wants to merge 1 commit into
Open
fix(deps): update dependency vllm to >=0.26.0,<0.27#449aar-public-version-bump-bot[bot] wants to merge 1 commit into
aar-public-version-bump-bot[bot] wants to merge 1 commit into
Conversation
Contributor
Author
|
aar-public-version-bump-bot
Bot
force-pushed
the
renovate/vllm-0.x
branch
3 times, most recently
from
July 15, 2026 00:26
7a43786 to
0068d56
Compare
aar-public-version-bump-bot
Bot
force-pushed
the
renovate/vllm-0.x
branch
2 times, most recently
from
July 17, 2026 00:31
bb10f8b to
10eefae
Compare
auto-merge was automatically disabled
July 18, 2026 00:29
Pull request was closed
aar-public-version-bump-bot
Bot
force-pushed
the
renovate/vllm-0.x
branch
3 times, most recently
from
July 23, 2026 03:05
712f4d3 to
2aecb40
Compare
aar-public-version-bump-bot
Bot
force-pushed
the
renovate/vllm-0.x
branch
12 times, most recently
from
July 27, 2026 03:06
bb88979 to
7938ccf
Compare
aar-public-version-bump-bot
Bot
force-pushed
the
renovate/vllm-0.x
branch
4 times, most recently
from
July 29, 2026 00:28
4fac9ea to
2157781
Compare
aar-public-version-bump-bot
Bot
force-pushed
the
renovate/vllm-0.x
branch
from
July 29, 2026 03:04
2157781 to
d35bf07
Compare
martinreinhardt01
force-pushed
the
renovate/vllm-0.x
branch
from
July 29, 2026 09:28
d35bf07 to
d45ea61
Compare
Contributor
Author
Edited/Blocked NotificationRenovate will not automatically rebase this PR, because it does not recognize the last commit author and assumes somebody else may have edited the PR. You can manually request rebase by checking the rebase/retry box above. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
This PR contains the following updates:
>=0.22,<0.23→>=0.26.0,<0.27Release Notes
vllm-project/vllm (vllm)
v0.26.0Compare Source
vLLM v0.26.0 Release Notes
Highlights
This release features 411 commits from 212 contributors (61 new)!
fused_topk_bias(1.5–2x kernel, #47463), and redundant repeat/copy removal (1.8% E2E TPOT, #48137), plus ROCm two-stage compressor for HCA prefill (#47718), sparse decode/prefill optimizations (#48519, #48788, #46275), and DSpark speculative decoding on AMD (#47419) and XPU (#47677).lm_headfor generation models viahead_dtype(#48390), extended to the LoRA path (#48525) and given a ROCmtorch.mmfast path (#48688), improving accuracy for generation heads.vllm-benchport (#48107).Model Support
lm_headon the LoRA path (#48525), optimizedTrtLlmLoRAExperts(#48759).torch.compile(#48901).Engine Core
lm_headfor generation models viahead_dtype(#48390); lower memory for capturing large CUDA graph sizes (#48483); opt-in persistence and reuse of the memory-profiling result across boots (#47388); improved InstantTensor loading (#46868).kv_cache_dtypeforspeculative_config(#48787).blocks_per_chunkconfig for heterogeneous KV groups (#48878), P2P default host/port env vars (#47636).new_block_ids(#44490), DSv3.2 + MTP + sequence-parallel accuracy (#48036).Hardware & Performance
fused_topk_bias1.5–2x (#47463), redundant repeat/copy removal (1.8% TPOT, #48137)._copy_mamba_state_blockto uint64 (#48110), stop upcasting logits to fp32 in the sampler (#48641).head_dtypetorch.mmfast path (#48688), DSv4 two-stage compressor kernel (#47718), sparse decode/prefill optimizations (#48519, #48788, #46275), DSv3.2 sparse MLA KV-split heuristic (#46832) and MTP CUDA-graph mode (#45149), MXFP8 GEMM for MiniMax-M3 (#46117), AITER sparse paged attention + spec decode for MiniMax-M3 (#47287, #47984), MiniMax-M2 fused QK-norm + all-reduce via AITER (#44849), HybridW4A16 linear kernel (#40977), Qwen3-30B-A3B QK-Norm+RoPE+KV runtime fusion (#42749).Large Scale Serving & Distributed
Quantization
nvfp4_per_tokenonline MoE quantization (#48538), CuTe-DSL FlashInfer MXFP4 quantization (#48417); bounded peak memory when repacking FP4 MoE weights for Marlin (#47851) and for NVFP4 MoE weight loading (#46276).kv_cache_dtype_skip_layerssupport (#47309).API & Frontend
vllm-benchport (#48107),continue_final_messagehandling with renderer sentinel (#47844).bad_wordsin/v1/completions(#46793), exposelogprob_token_idson Python OpenAI endpoints (#43463),include_reasoningparam for non-Harmony models (#44301), populatenum_cache_creation_tokenson Messages responses (#48535)./abort_requestson the RLHF dev API router (#47173); Deepstream video decoding backend (#42424); overlap preprocessing and computation for pooling models in offline inference (#47699).Security
Dependencies
Deprecations & Removals
New Contributors
Contributors
@mgoin, @yewentao256, @NickLucche, @njhill, @LucasWilkinson, @micah-wil, @khluu, @AndreasKaratzas, @WoosukKwon, @aoshen02, @BugenZhao, @vanshbhatia-amd, @jperezdealgaba, @jeejeelee, @hmellor, @tlrmchlsmth, @MatthewBonanni, @Change72, @chaojun-zhang, @gau-nernst, @gnovack, @ZJY0516, @reidliu41, @Yejing-Lai, @taneem-ibrahim, @Srinivasoo7, @Sunt-ing, @AmeenP, @zhenwei-intel, @LopezCastroRoberto, @matteso1, @yzong-rh, @stefankoncarevic, @Isotr0py, @djramic, @charlifu, @peizhang56, @giuseppegrossi, @drakosha, @NickCao, @benchislett, @Rohan138, @hickeyma, @muhammadfawaz1, @rasmith, @zixi-qi, @music-dino, @BWAAEEEK, @xianbaoqian, @wendyliu235, @atalman, @joerowell, @ErenAta16, @AlejandroParedesLT, @gcanlin, @liranschour, @omerpaz95, @Alex-ai-future, @tanpinsiang, @Fangzhou-Ai, @KKothuri, @zxd1997066, @akii96, @nemanjaudovic, @Etelis, @afierka-intel, @DaoyuanLi2816, @HDCharles, @tahsintunan, @xiaohongchen1991, @sagearc, @mikekg, @kliuae, @qli88, @arpera, @yushangdi, @edwinlim0919, @tjtanaa, @ariG23498, @lucifer1004, @netanel-haber, @kl527, @Rukhaiya2004, @ronensc, @guan404ming, @shaunkotek, @liulanze, @pierDipi, @eldarkurtic, @simon-mo, @robinguo23, @CienetStingLin, @RishabhSaini, @Sahil170595, @jasonlizhengjian, @jacklin78911-collab, @walterbm, @amd-ethany, @zqzten, @nicklasfrahm, @hongxiayang, @alexeldeib, @ManaEstras, @sungbin1015, @aoright, @Saddss, @vivek8123, @voipmonitor, @shawntsai, @almayne, @ilmarkov, @cleonard530, @kjiang249, @chaunceyjiang, @bigPYJ1151, @tsvikas, @deng451e, @tvirolai-amd, @zhewenl, @zihaomu, @ap9272, @staugust, @Yancey0623, @GongLei-HW, @albertoperdomo2, @guoriyue, @ViranjanPagar, @Functionhx, @XuZhou26, @MynameFelix, @larryli2-amd, @ashwing, @thisisjimmyfb, @robertgshaw2-redhat, @mayuyuace, @ibondarenko1, @zhejiangxiaomai, @vx120, @hugo-cen, @tanish-malekar, @zzt93, @guybd, @R3hankhan123, @mmangkad, @omera-nv, @yma11, @Gavin-Morris-04, @pavanimajety, @shanjiaz, @wenpengw-nv, @atalhens, @langzhao-netizen, @emricksini-h, @Zhenzhong1, @DanBlanaru, @mgehre-amd, @mwoodson, @wangxiyuan, @adsridhar, @hnt2601, @gangula-karthik, @tomerg-nvidia, @adhi29, @rishitdholakia13, @divakar-amd, @chaeminlim-mb, @joanvelja, @russellb, @janeyx99, @aarushjain29, @wjabbour, @mahadrehmann, @krishy91, @tzielinski-habana, @avalliappan-nvidia, @ruikangliu, @majian4work, @maxyanghu, @brijrajk, @lengrongfu, @Josephasafg, @elvircrn, @xiaguan, @ricky-chaoju, @iyastreb, @ovidiusm, @tuukkjs, @noooop, @samnordmann, @AndyDai-nv, @xiao-llm, @DiegoCao, @yuvalluria, @jhu960213, @woosebastian, @Debasish-87, @esmeetu, @hao-aaron, @zhangj1an, @wendadawen, @juliendenize, @passtoor-agi, @mosya415, @labAxiaoming, @devalshahamd, @wangxingda, @xuebwang-amd, @fuscof-ibm, @alexxu-roblox, @frida-andersson, @lishunyang12, @izhuhaoran
v0.25.1Compare Source
vLLM v0.25.1
Highlights
This release features 2 commits from 2 contributors (1 new)!
v0.25.1 is a patch release containing two targeted bug fixes on top of v0.25.0.
Bug Fixes
import torchcodecraised aRuntimeErrorat import time when system FFmpeg was missing, which blocked startup (e.g.vllm serve Qwen/Qwen3-VL-2B-Instruct) even when TorchCodec was not in use. The error is now deferred to runtime so it only surfaces if TorchCodec is actually needed.!!!!!tokens. A dtype-match guard now routes incompatible mixed-dtype graphs to the safe path, while same-dtype models retain the full allreduce + RMSNorm + quant fusion.Contributors
@Isotr0py, @hugo-cen
New Contributors
v0.25.0Compare Source
vLLM v0.25.0 Release Notes
Highlights
This release features 558 commits from 232 contributors (64 new)!
Model Support
tok_sparse_selectfrom MSA replacing Triton kernels (#47502).mm_token_type_idsfix (#46552), tied-embeddinglm_head.biasfix (#46835).architecturescrash fix (#46037), DeepSeek-V2 hidden-size and aux-hidden-state fixes (#46986, #46973).Engine Core
Configuration
📅 Schedule: (UTC)
🚦 Automerge: Enabled.
♻ Rebasing: Whenever PR is behind base branch, or you tick the rebase/retry checkbox.
🔕 Ignore: Close this PR and you won't be reminded about this update again.
This PR has been generated by Mend Renovate.