Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
1083 commits
Select commit Hold shift + click to select a range
0585b5b
Skip docs build if PR doesn't affect docs (#43972)
hmellor May 29, 2026
3f6f508
[Bugfix][CPU] Remove invalid extra deps (#43977)
bigPYJ1151 May 29, 2026
11dfa31
Add vLLM library info to Hugging Face Hub requests (#43857)
Wauplin May 29, 2026
f191d56
docs: clarify ITL acronym in optimization docs (#43922)
chunyang-wen May 29, 2026
5502c3b
[Misc] added unit tests for the core pooling methods (#43818)
taneem-ibrahim May 29, 2026
4ff865c
[Bugfix] Disable allreduce_rms_fusion when pipeline_parallel_size > 1…
zixi-qi May 29, 2026
84b2a8a
[MoE Refactor] WNA16 MoE backend selection into oracle module (#42553)
bnellnm May 29, 2026
4aaba00
[EPLB] Make async EPLB default (#43219)
ilmarkov May 29, 2026
d07ad06
[Bugfix] Use storage_block_size in KV cache reshape for compressed sp…
zixi-qi May 29, 2026
8b9deee
[Bugfix] Fix Ray placement group allocation with grouped nodes (#43998)
czhu-cohere May 29, 2026
739096a
[Bug] Fix torch device issue for MOE permute (#44005)
yewentao256 May 29, 2026
6aabe22
[CI] Make Model Executor test hangs fail fast with a traceback (#43971)
khluu May 29, 2026
6de08e8
[CI] Remove redundant test_chat_with_tool_reasoning.py (#44011)
sfeng33 May 29, 2026
acbc203
Add @khluu to CODEOWNERS (#44019)
khluu May 29, 2026
5dbf160
[Feature] SSL support for dp supervisor (#43688)
yewentao256 May 29, 2026
38b864d
[Metrics] Exclude KV transfer tokens from iteration_tokens_total (#43…
tlrmchlsmth May 29, 2026
46409fd
[Fronten] Clean up stop_token_ids override for Harmony (#44009)
yzong-rh May 29, 2026
106aa92
[MoE Refactor] Migrate MoeWNA16Method quantization to MK oracle (#42647)
bnellnm May 29, 2026
7b98f49
[MoE Refactor] Remove supports_expert_map (#43108)
bnellnm May 29, 2026
8c6daf6
[CI] Remove duplicate Harmony test coverage (#44023)
sfeng33 May 29, 2026
8fad266
[CI] Fix smoke test step key to bypass block gate (#43974)
khluu May 29, 2026
187457a
Revert "[MoE Refactor] Migrate MoeWNA16Method quantization to MK orac…
bnellnm May 29, 2026
559d671
[PERF]MiniMax-M2 gate kernel (#38445)
jeejeelee May 30, 2026
1e2ce5d
offload prompt_embeds decode in render_prompts_async to avoid blockin…
gagandhakrey May 30, 2026
1a096d8
[Refactor] Remove dead current_tool_name_sent assignments from tool p…
sfeng33 May 30, 2026
ef8840a
[ROCm][CI] Fix failure in the Phi3V pooling test (#44028)
AndreasKaratzas May 30, 2026
c0056b1
[ROCm] cmake: support PYTORCH_FOUND_HIP for torch 2.13 native HIP lan…
nemanjaudovic May 30, 2026
e949999
[BugFix][Platform] Fix import vllm.platforms.rocm error on non-CUDA t…
Liangliang-Ma May 30, 2026
124fac1
[Bugfix] Fix RMSNorm kernels to multiply in weight's native dtype (#4…
liulanze May 30, 2026
3becc5d
[ROCm] Add attention sink support to AITer flash attention backend (#…
sphinx07 May 30, 2026
50c80d7
[Governance] Add @BugenZhao as Rust frontend code owner (#44047)
BugenZhao May 30, 2026
e110506
[Bug] Fix gemma4 MTP IMA issue when TP>1, `CUDA error: an illegal mem…
yewentao256 May 30, 2026
27fa5aa
[MRV2] Support breakable CUDA graph (#44050)
WoosukKwon May 30, 2026
3fd9d2d
[CPU][Zen] Route W8A8 and W4A16 linear inference through zentorch on …
aadwived May 30, 2026
6bdabba
[CI/Build] Enable Step3p7ForConditionalGeneration testing (#43956)
jeejeelee May 31, 2026
8b8546d
docs: fix MLA attention docstring examples (#44118)
nightcityblade May 31, 2026
f46e6be
[Misc] Use VLLMValidationError consistently in chat completion and co…
umut-polat Jun 1, 2026
4721bb3
[MRV2] Remove Eagle's dedicated CUDA graph pool (#44078)
LucasWilkinson Jun 1, 2026
29d6933
[BugFix] Fix `_has_module` to verify native deps via trial import (#4…
jeffreywang88 Jun 1, 2026
1fd8bd0
[Docs] Replace broken video url in examples (#44159)
Isotr0py Jun 1, 2026
98f1279
[CPU][RISC-V] Add missing RVV cpu_types helpers for WNA16 (#42730)
Wcy1023n Jun 1, 2026
1f6048a
fix: glm5.1 pp model loading (#42944)
UranusSeven Jun 1, 2026
0910f7e
[Frontend] Resettle generative scoring entrypoint. (#44153)
noooop Jun 1, 2026
de21863
[Rust Frontend] Add InternLM2 tool parser (#43481)
willamhou Jun 1, 2026
8796838
[Bugfix] fix wrong partial_rotary_factor calculation for bailing_moe …
zzt93 Jun 1, 2026
bd0aecd
[XPU][CI] Fix test_audio_in_video flake by using module-scoped server…
chaojun-zhang Jun 1, 2026
985c97a
[Perf] Optimize cutlass fp8 scaled mm bypassing padding, 20% kernel p…
yewentao256 Jun 1, 2026
023808c
[Feature] Add support for JetBrains' Mellum v2 code generation model …
shadeMe Jun 1, 2026
0357335
[Kernel][DSv4] Optimize sparse FP8 compressor kernels (#44161)
zyongye Jun 1, 2026
fd9e91d
[ROCm][CI] Fix and stabilize EAGLE3 acceptance tests (#41294)
AndreasKaratzas Jun 1, 2026
182c67d
[Rust Frontend] Support streaming `generate` endpoint (#43779)
Xunzhuo Jun 1, 2026
266b9d9
[Frontend][Core] Add sparse NCCL weight transfer support for in-place…
bedeks Jun 1, 2026
6f8b40a
[BugFix][CI] Fix added `_has_module` tests (#44248)
njhill Jun 1, 2026
e4cbc43
[Test][BugFix] Fix double-BOS in PD+specdec acceptance test (#44234)
njhill Jun 1, 2026
8c3cc98
[DSV4] Remove unncessary classes & functions (#44246)
WoosukKwon Jun 1, 2026
48c0d13
[ROCm][CI] Skip unbacked dynamic shapes tests on PyTorch < 2.11 (#44256)
JartX Jun 2, 2026
517e74a
[DSV4] Refactor RoPE initialization (#44262)
WoosukKwon Jun 2, 2026
d68f0b2
[Bugfix][Mooncake] Release GPU pin on failed store in MooncakeStoreCo…
Dao007forever Jun 2, 2026
2588ec4
[ROCm] Upgrade AITER to v0.1.13.post1 (#44265)
micah-wil Jun 2, 2026
816cc73
[Bugfix][CI] Normalize NIXL connector CUDA wheel installs (#44266)
alec-flowers Jun 2, 2026
9affc17
[Refactor] Move unstreamed tool-arg flush from serving layer to parse…
sfeng33 Jun 2, 2026
54d0c36
[CI] Stabilize OpenAI schema fuzzing for malformed structural tags (#…
AndreasKaratzas Jun 2, 2026
279d25f
[BugFix] Fix TypeError in MiniCPM-O audio feature unpadding (#38053)
Krishnachaitanyakc Jun 2, 2026
480fada
[BugFix][kv_offload]: Prevent offloading stale sliding window blocks …
orozery Jun 2, 2026
a3a5a5e
[XPU][Bugfix] Fix per_token_group_fp8_quant missing dummy args on XPU…
chaojun-zhang Jun 2, 2026
a045c74
[MM][CG] Profile encoder CUDA graph pool memory (#41714)
BWAAEEEK Jun 2, 2026
f91fb2f
[Bugfix] Convert Gemma4-MM ViT linear layers to vllm native impl (#43…
Isotr0py Jun 2, 2026
8a9eb40
[Model Runner V2] Support zeroing freshly allocated KV blocks for hyb…
izhuhaoran Jun 2, 2026
1edfd09
[Model Runner V2] Use actual batch max_seq_len for attn metadata (#43…
izhuhaoran Jun 2, 2026
68dafcc
[Refactor] Unify reasoning + tool-call parsing behind Parser.parse() …
sfeng33 Jun 2, 2026
dcdfe66
[Perf] use triton moe backend on hopper by default (#44220)
ZJY0516 Jun 2, 2026
0b25cf4
[CPU][Perf] Enable fused kernels for GDN's gated delta rules (#43534)
fadara01 Jun 2, 2026
93da882
[kv_offload] Add `@override` decorators to subclass method implementa…
ronensc Jun 2, 2026
b817b23
[Rust Frontend] add --enable-request-id-headers flag support. (#43883)
cinnamonica02 Jun 2, 2026
7c37096
[Core][Refactor]: thread `scheduler_block_size` into KVCacheManager a…
ivanium Jun 2, 2026
d247a9d
[EC Connector] Non blocking EC Connector lookup (#41627)
omerpaz95 Jun 2, 2026
e303132
[Parser] Migrate `ResponsesParser` to unified `Parser` interface (#42…
albertoperdomo2 Jun 2, 2026
f8e9c56
[Multimodal] Automatically select registered video loader for VLM (#4…
Isotr0py Jun 2, 2026
689b0ee
[HARDWARE][POWER] Enable SHM communicator support for PowerPC (#43754)
Rukhaiya2004 Jun 2, 2026
2a2b5ca
[KV Offload] Add `on_schedule_end()` hook to separate step lifecycle …
ronensc Jun 2, 2026
f69ede4
[XPU][Mamba] Triton-based selective scan forward op for XPU (#43421)
mfylcek Jun 2, 2026
0eeba5e
Fix DFlash prefix cache corruption due to missing lookahead block (#4…
shreyas269 Jun 2, 2026
b623f7e
[Frontend] Consolidate dev entrypoints. (#44170)
noooop Jun 2, 2026
654bd2b
[Bugfix] Sync block_size from EngineCore to frontend for hybrid Mamba…
Gruner-atero Jun 2, 2026
2fd0e52
[Bugfix] Fix Gemma4 startup crash with recent transformers multimodal…
lucianommartins Jun 2, 2026
0cbc48c
Support ModelOpt MXFP8 non-gated MoE (#42958)
TomerBN-Nvidia Jun 2, 2026
0bdfd5e
[Bugfix] Vendor MiniCPMV/MiniCPMO processors to unblock Transformers …
wjinxu Jun 2, 2026
ea0d045
[FlashAttention] Sync FA with upstream (#44065)
MatthewBonanni Jun 2, 2026
c91a87f
[BugFix] [GDN] Read linear_key_head_dim from hf_text_config for multi…
IdoAtadTD Jun 2, 2026
6314de8
[XPU] [Bug] remove xpuw4a16 output size check (#44168)
zufangzhu Jun 2, 2026
880fc03
[Rust Frontend] Support recursive tool parameter conversion (#44299)
BugenZhao Jun 2, 2026
88f1721
[ROCm] Fix AITER RMSNormQuantFusion for Kimi-Linear (#44308)
pschlan-amd Jun 2, 2026
586201e
[Rust Frontend] Cover different thinking modes in roundtrip tests (#4…
BugenZhao Jun 2, 2026
4d93bc3
Migrate header files to torch stable abi (#44013)
cleonard530 Jun 2, 2026
53fa09d
[Misc] Support local image encoding in benchmarks (#43843)
xiaozcy Jun 2, 2026
774e552
[compressed-tensors] Asymmetric support for MoE WNA16 marlin (#44025)
brian-dellabetta Jun 2, 2026
cab5c9a
[Core] Move `max_concurrent_batches` to `VllmConfig` (#44274)
njhill Jun 2, 2026
478b49d
[Refactor] Remove dead code from parser infrastructure (#44279)
sfeng33 Jun 2, 2026
3f3e270
[XPU] Enable rms_norm/act quant fusions (#43963)
zhenwei-intel Jun 2, 2026
afcb580
[BugFix] Fix Humming MoE deploy error (#43100)
adotdad Jun 2, 2026
fe32e78
[Bugfix] flashinfer: fail fast when --kv-cache-dtype nvfp4 used on un…
Kartavyasonar Jun 2, 2026
2427094
[Feature] Support EPLB for DeepSeek v4 Mega Moe (#43339)
wzhao18 Jun 2, 2026
ed9a752
[Anthropic] Support system role messages inside messages array (#44283)
chaunceyjiang Jun 2, 2026
da107a5
[MRV2] Also enable MRV2 for Llama and Mistral dense models (#43458)
njhill Jun 2, 2026
b8b49e2
Bump actions/github-script from 8.0.0 to 9.0.0 (#39667)
dependabot[bot] Jun 2, 2026
e4a2e58
[MRV2] Remove assignment of graph_pool in cudagraph_utils (#44338)
WoosukKwon Jun 2, 2026
e9e08c4
[Bugfix] Cache the EAGLE/MTP lookahead block in the SWA prefix-cache …
ivanium Jun 2, 2026
5577811
[Misc] Remove stray empty file (#44350)
MatthewBonanni Jun 2, 2026
e15f202
[ModelRunnerV2] Avoid pipeline parallel bubbles (#42187)
njhill Jun 2, 2026
3099de3
[Kernel][MoE] Add GELU_TANH to CPU, CUTLASS, and WNA16 MoE backends (…
lesj0610 Jun 2, 2026
0917a00
Fix sparse NCCL weight transfer test construction (#44345)
bedeks Jun 2, 2026
8b3b71e
[CI/Build] Bump flashinfer to v0.6.12 (#44036)
vadiklyutiy Jun 2, 2026
a4ac746
[MoE/b12x] Accept W4A16 (kNvfp4Static, None) in FlashInferB12xExperts…
ECMGit Jun 2, 2026
bd98e97
[Misc] Remove dead VLLM_RPC_TIMEOUT env var and fix profiling doc tha…
DaoyuanLi2816 Jun 3, 2026
b254e04
[DSV4] Minor cleanup for DeepseekV4MegaMoEExperts (#44367)
WoosukKwon Jun 3, 2026
ca17b6b
[Perf] Apply single-pass min_larger finding and binary search in Trit…
cakeng Jun 3, 2026
969aec4
[Bugfix] Fix Deepseek v4 non-mega-moe model init error (#44356)
wzhao18 Jun 3, 2026
27a93cd
[docker] Stop using extra-index-url for flashinfer-jit-cache (#44366)
khluu Jun 3, 2026
02a0149
[Platform] Add is_cumem_allocator_available (#43838)
wangxiyuan Jun 3, 2026
4454a18
[ROCm][CI] Fix stale wvSplitK GEMM fallback test for N=5 (#44368)
JartX Jun 3, 2026
7b476c8
[ROCm][CI] Skip fp8 reload tests on gfx90a (MI250) (#44369)
JartX Jun 3, 2026
53b88d1
[CI] Reject out-of-vocabulary before they reach the GPU logprob path…
AndreasKaratzas Jun 3, 2026
e670638
[CI] Add missing vllm/parser/ CI trigger and fix test_parse.py (#44352)
sfeng33 Jun 3, 2026
3f0a91b
Nit Changes in Tiered KV Offload (#44293)
rshavitt Jun 3, 2026
597bc15
fix: resolve CUTLASS fmin compatibility for DeepSeek-V4 init (#44236)
Oxygen56 Jun 3, 2026
f020435
[Bugfix] fix crash in postprocess for null tool args (#43862)
william-rom Jun 3, 2026
e0081ef
[Benchmark] Enable reasoning-model (thinking) benchmarking via `--cha…
qiching Jun 3, 2026
71df063
Enable perf_token_group_quant/_C_stable_libtorch for ROCm (#42758)
charlifu Jun 3, 2026
87954eb
[ROCm][CI] Optimize ROCm Docker build: registry cache, DeepEP, and ci…
AndreasKaratzas Jun 3, 2026
9af53a3
[Perf] Add tuned selective_state_update configs for H200 and RTX PRO …
Majid-Taheri Jun 3, 2026
7268457
[KV Offloading] Enable HMA models for Tiering Offloading (#44287)
varun-sundar-rabindranath Jun 3, 2026
4aaed4c
[Rust Frontend] Add server router extension hook (#43774)
NolanHo Jun 3, 2026
6550ff1
[Rust Frontend] Add dynamic LoRA endpoints (#43778)
Xunzhuo Jun 3, 2026
449be4f
[Rust Frontend] Fix several hf chat template rendering issues (#44311)
BugenZhao Jun 3, 2026
0e2b131
[Doc] Update ViT CUDA graph interfaces (#44388)
shen-shanshan Jun 3, 2026
ace95c9
[Bugfix] Update TrtLLM MoE routing methods (#44347)
wzhao18 Jun 3, 2026
209709a
[Bugfix] Fix unstreamed tool call args dropped in Responses API strea…
sfeng33 Jun 3, 2026
02564b4
[XPU]fallback to TRITON_ATTN for vit attn on xpu when use float32 dty…
yma11 Jun 3, 2026
1fa9ea0
[Perf] Triton fast path for small CPU→GPU `swap_blocks_batch` in the …
Etelis Jun 3, 2026
95b1615
[Perf] Improve multimodal item handling from O(n) to O(log n) per ste…
andylolu2 Jun 3, 2026
823d271
[Attention][CPU] Standardize kv layout to blocks first (#44393)
bigPYJ1151 Jun 3, 2026
3d76f39
[SharedOffloadRegion] Align blocks to page-size (#43689)
varun-sundar-rabindranath Jun 3, 2026
309385a
[Rust Frontend] Add /server_info to Rust frontend (#43942)
Xunzhuo Jun 3, 2026
e523267
[XPU] Add XPU block-scaled W8A8 fp8 path (#39968)
xwu-intel Jun 3, 2026
e3e132d
[Refactor] Suppress SyntaxWarning from ast.literal_eval in tool parse…
sfeng33 Jun 3, 2026
27f1d34
[Frontend][Responses API] Move developer-to-system conversion into HF…
chaunceyjiang Jun 3, 2026
ec8d60b
[Model Runner V2] Use FlashInfer sampler (#42472)
njhill Jun 3, 2026
4d1fd13
[CI/Build] Fix LoRA testing (#44425)
jeejeelee Jun 3, 2026
df7252c
[CI] Align PD tests to HMA on by default (#44174)
NickLucche Jun 3, 2026
0c6631f
[KVCache] Support Pluggable KVCacheSpec (#37505)
MengqingCao Jun 3, 2026
51e0c57
fix(config): validate max_num_scheduled_tokens >= 0 on all paths (#44…
Oxygen56 Jun 3, 2026
0a5cbf6
Handle spinloop ext load failure gracefully (#43659)
pschlan-amd Jun 3, 2026
59d0236
[10b/n] Migrate custom all-reduce, DeepSeek V4 fused MLA, MiniMax red…
cleonard530 Jun 3, 2026
5b2a2be
[ROCm][CI] Move Model Executor test step from MI250 to MI300 (gfx942)…
JartX Jun 3, 2026
2b91012
[Refactor] Remove dead code fp quant (#44122)
yewentao256 Jun 3, 2026
271328e
[LoRA] Fix dedup for post-replacement module aliases (#44413)
linitra24 Jun 3, 2026
a248b45
[Model] Add Gemma4 Unified (encoder-free) support (#44429)
lucianommartins Jun 3, 2026
dad95e3
[Feature] Support batch invariant rms norm with residual (#42453)
yewentao256 Jun 3, 2026
2b237c7
[Bugfix] Honor tool_choice="none" in Chat Completions streaming (#42752)
hoobnn Jun 3, 2026
91945b6
[Bug Fix][Model Runner V2][Spec Decode] Warmup & capture with differe…
TheEpicDolphin Jun 3, 2026
6bad553
[Minor] Remove FlashInfer version check in topk_topp_sampler (#44442)
WoosukKwon Jun 3, 2026
bdbf08f
Bump actions/stale from 10.1.1 to 10.2.0 (#35078)
dependabot[bot] Jun 3, 2026
128adab
[Bugfix] Fix Gemma4 MTP block_table batch_size mismatch under concurr…
Dymasik Jun 4, 2026
0414d75
[XPU] skip unapplied UT in test_gpu_model_runner.py (#44289)
yma11 Jun 4, 2026
ceb0111
[Model Runner V2][Spec Decode] Add Gemma4 MTP support (#43241)
TheEpicDolphin Jun 4, 2026
0c1e6f6
[Bugfix] Fix VLLMNotFoundError when using LoRA adapter name in poolin…
wanghenshui Jun 4, 2026
b58e082
[KV Connector] Update lmcache kv_offloading_backend to use LMCacheMPC…
maobaolong Jun 4, 2026
f25952e
[MM][Perf][CG] Support ViT full CUDA graph for InternVL (#41759)
oguzhankir Jun 4, 2026
e6018c6
[Refactor] Remove dead code in tests and parallel_state (#41471)
yewentao256 Jun 4, 2026
f0cd590
optimize the compressor 128 split cutedsl kernel (#44230)
Jie-Fang Jun 4, 2026
4f423bd
[EPLB] Nixl communicator optimization. Zero-copy transfers (#41633)
ilmarkov Jun 4, 2026
5e2af28
[CI] Resolve release V2 docker build after ROCm CI wheels change (#44…
AndreasKaratzas Jun 4, 2026
b4b4aaa
[Inductor] Fast-path Inductor fallback for vllm::*/vllm_aiter::* cust…
okorzh-amd Jun 4, 2026
d01d0b4
[Frontend] Consolidate online serving utils. (#44479)
noooop Jun 4, 2026
22c2e87
[CI] Reverted gitignore changes (#44497)
AndreasKaratzas Jun 4, 2026
a618356
[Prefix Caching] DeepSeekv4 - Support selective prefix-cache retentio…
wzhao18 Jun 4, 2026
1bdc60e
Fix Kimi-K2.5 FlashInfer ViT metadata (#44493)
Kevin-XiongC Jun 4, 2026
d0975a4
[perf] Add gemma RMS AR fusion (#42646)
jiahanc Jun 4, 2026
9061935
[Attention] Mamba attention module refactor - LINEAR (#43556)
wangxiyuan Jun 4, 2026
4b87b3e
[Bugfix] fix EVS for qwen3-vl (#44205)
garrygale Jun 4, 2026
e68988a
Refactor CT NVFP4 linear to use a single class (#42443)
dsikka Jun 4, 2026
f35b557
Add GH token to docs build pre run check (#44534)
hmellor Jun 4, 2026
9354fb1
[Bugfix][Compile] Guard per_token_group_fp8_quant lookup on non-CUDA …
QiliangCui2023 Jun 4, 2026
68f5e56
[PD][Nixl] Mamba prefix caching mode support (#42554)
NickLucche Jun 4, 2026
0c96dd6
[ROCm] Bump fastsafetensors to v0.3.2 from PyPI, remove git source bu…
wjabbour Jun 4, 2026
6f68ca3
[ROCm][CI] Stabilize memory-release in the Hybrid model generation te…
AndreasKaratzas Jun 4, 2026
3e77036
[ROCm][CI] Specifying time outs for the lm eval models (#44255)
AndreasKaratzas Jun 4, 2026
b5235fc
[DSv4] Adding TRTLLM gen attention kernel (#43827)
zyongye Jun 4, 2026
06ee2d8
[Quant] Support compressed-tensors WNA8O8Int linears and WNInt embedd…
mgoin Jun 4, 2026
b21443e
Add model support for granite speech plus (#43519)
zvik Jun 4, 2026
3dbb4e0
[Bugfix] MiniCPM-V-4.6 video inference crash: placeholder count misma…
tc-mb Jun 4, 2026
4cc78c9
[Core] Freeze garbage collector in workers after model initialization…
tlrmchlsmth Jun 4, 2026
99ef652
[Bugfix] Reject non-positive values for ParallelConfig int knobs (#44…
jwzheng96 Jun 4, 2026
06f9463
[ROCm][CI] Add test for Aiter unified attn kernel (#44436)
divakar-amd Jun 4, 2026
3da29aa
[DOC] Add INT8 W4A8 docs and Arm's supported quantization schemes (#3…
fadara01 Jun 4, 2026
8d9536a
[Misc] Add unit tests for pooler head classes (#44471)
taneem-ibrahim Jun 4, 2026
439203d
[Bugfix] Fix test_cutlass_moe.py (#44380)
bnellnm Jun 4, 2026
a947f7a
[Kernel][Test] Extend lightning_attn and awq_triton kernel tests to X…
adobrzyn Jun 4, 2026
38fd240
use split_group for pytorch process group creation (#41980)
tushar00jain Jun 4, 2026
41a4829
[Logs Refactor] Optimize shutdown logs, easier to follow and consiste…
yewentao256 Jun 4, 2026
a55fccf
[mamba] unify KDA conv states into one cache to match 2-state SSM lay…
ZJY0516 Jun 4, 2026
b7c5baf
fix: keep DeepSeek V4 RoPE cache on inv_freq device (#43926)
galletas1712 Jun 4, 2026
62d6f06
[Rust Frontend] Skip loading multimodal processor if `--language-mode…
BugenZhao Jun 5, 2026
063ce98
[XPU][MoE] support block_fp8_moe on xpu (#42139)
zufangzhu Jun 5, 2026
56aff0d
[10/n] Migrate cuda_view and silu_and_mul_per_block_quant kernels to …
cleonard530 Jun 5, 2026
4efd6ff
[DSV4] Refactor DeepseekV4Attention (#44569)
WoosukKwon Jun 5, 2026
da1daf4
[Bugfix] Exclude vision embedder from quantization in Gemma4 Unified …
lucianommartins Jun 5, 2026
96229fa
[KVConnector][1/N] PP-aware handshake aggregation and intermediate-PP…
zixi-qi Jun 5, 2026
c505cd9
[CI/Build] Disable CPU-Compatibility Tests (#44605)
bigPYJ1151 Jun 5, 2026
165b786
[ROCM] [FEAT] Integrate Aiter hipBLASLt GEMM online tuning (#40426)
hanlin12-AMD Jun 5, 2026
b4a6f26
[ROCm][perf] Use workspace manager for sparse indexer allocations (#4…
tuukkjs Jun 5, 2026
ef3af56
Fix `LLM.wait_for_completion` output type docstring (#44617)
viiccwen Jun 5, 2026
ca73293
[Bugfix][Rust Frontend] Fix UTF-8 char-boundary panic in incremental …
Sunt-ing Jun 5, 2026
6542d48
[Bugfix] Fix test_invocations flaky failure with newer openai SDK (#4…
XuZhou26 Jun 5, 2026
d2f70da
fix: pad dummy run query_start_loc (#44603)
UranusSeven Jun 5, 2026
d61d856
[Bugfix] Update mistral tokenizer test for continue_final_message fix…
XuZhou26 Jun 5, 2026
e64237a
[Rust Frontend] Support include_reasoning=false (#44391)
ricky-chaoju Jun 5, 2026
d98b8f3
[NixlConnector] Initiate deprecation cycle for `kv_both` role (#43874)
NickLucche Jun 5, 2026
efc347f
docs: fix tokenizer optimization typo (#44066)
chunyang-wen Jun 5, 2026
8a83e6f
[Rust Frontend] Batch auto-abort requests by engine (#44591)
HueCodes Jun 5, 2026
7fe7800
[BUG] Fix FP64 Gumbel precision coverage (#43150)
tianyu-z Jun 5, 2026
62215e7
Remove KV cache scale boilerplate from model weight loading methods (…
hmellor Jun 5, 2026
bbb6c27
[Bugfix] Fix gemma4 crash on CPU: guard mem_get_info call (#44615)
adhithyamulticoreware Jun 5, 2026
02d2da0
[DSV4] Move more ops out of eager breakpoint (#44561)
WoosukKwon Jun 5, 2026
6a11d72
[Reasoning][Structured Outputs] Add Command A plus tags for structura…
rishitdholakia13 Jun 5, 2026
c66b198
[CI] Bump mistral-common (#44649)
hmellor Jun 5, 2026
a80af24
Speed up docs build (#44635)
hmellor Jun 5, 2026
ef0df7d
[CI] Bump mypy version `1.19.1` -> `1.20.2` (#44647)
hmellor Jun 5, 2026
7f003a1
Support MiniCPMV batched preprocessing (#44609)
yma11 Jun 5, 2026
6a89457
Add objectstore as a secondary tier to multi-tier kv cache offloading…
effi-ofer Jun 5, 2026
aa6fb8a
[Bugfix] [ROCm] [Critical] fallback to regular abi for ROCm (#44648)
tjtanaa Jun 5, 2026
91e17d4
Fix sarvam forward compatibility with transformers v5 (#38804)
Vikrantpalle Jun 5, 2026
b593396
Upgrade tpu-inference to v0.21.0 (#44621)
CienetStingLin Jun 5, 2026
703fb17
[Bugfix] GPT-OSS instruction rendering (#44330)
yzong-rh Jun 5, 2026
e28e369
Male Mergify comment less spammy (#44666)
hmellor Jun 5, 2026
c73b0d0
[Core][Engine] allow DP ray placement groups to be set on specific no…
walterbm Jun 5, 2026
4200f62
[ROCm][GPT-OSS] Fuse RoPE + static Q FP8 quant on fused RoPE+KV path …
akii96 Jun 5, 2026
f6a708a
[Doc] Add Llama-3.2-3B-Instruct to batch-invariance tested models (#4…
DaoyuanLi2816 Jun 5, 2026
a50e675
[Cohere] fix RoutingMethodType (#44021)
Terrencezzj Jun 5, 2026
1e36b67
fix-function-call-id-null
ankrovv Jun 5, 2026
9c1acc9
Fix Harmony tool descriptions for optional fields
shenoyvvarun Jun 5, 2026
7755735
fix(responses): support required Harmony function tools
ankrovv Jun 5, 2026
0739ad0
test(responses): trim required tool coverage
ankrovv Jun 5, 2026
91fc1a4
refactor(responses): tighten required tool shim
ankrovv Jun 5, 2026
0e51515
style(responses): simplify required tool checks
ankrovv Jun 5, 2026
46396bd
refactor(responses): clarify required function path
ankrovv Jun 5, 2026
77f6912
fix(responses): allow text format with required tools
ankrovv Jun 5, 2026
9dfbbdb
update-notimpl-message-for-required-tool-choice
ankrovv Jun 5, 2026
477587c
fix(responses): render harmony instructions as developer
Jun 6, 2026
3df92cf
add-json-constraint-chat
ankrovv Jun 18, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
The table of contents is too big for display.
Diff view
Diff view
  •  
  •  
  •  
1 change: 1 addition & 0 deletions .buildkite/ci_config.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -8,6 +8,7 @@ run_all_patterns:
- "CMakeLists.txt"
- "requirements/common.txt"
- "requirements/cuda.txt"
- "requirements/kv_connectors.txt"
- "requirements/build/cuda.txt"
- "requirements/test/cuda.txt"
- "setup.py"
Expand Down
23 changes: 23 additions & 0 deletions .buildkite/ci_config_rocm.yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,23 @@
name: vllm_rocm_ci
job_dirs:
- ".buildkite/hardware_tests"
run_all_patterns:
- "docker/Dockerfile.rocm"
- "docker/Dockerfile.rocm_base"
- "docker/ci-rocm.hcl"
- "docker/docker-bake-rocm.hcl"
- ".buildkite/hardware_tests/amd.yaml"
- ".buildkite/scripts/ci-bake-rocm.sh"
- ".buildkite/scripts/hardware_ci/run-amd-test.py"
- ".buildkite/scripts/hardware_ci/run-amd-test.sh"
- "CMakeLists.txt"
- "requirements/common.txt"
- "requirements/rocm.txt"
- "requirements/build/rocm.txt"
- "requirements/test/rocm.txt"
- "setup.py"
- "csrc/"
- "cmake/"
run_all_exclude_patterns:
- "csrc/cpu/"
- "cmake/cpu_extension.cmake"
81 changes: 66 additions & 15 deletions .buildkite/hardware_tests/amd.yaml
Original file line number Diff line number Diff line change
@@ -1,22 +1,73 @@
group: Hardware - AMD Build
group: Hardware - AMD Build
steps:
- label: "AMD: :docker: build image"
key: image-build-amd
# Ensure ci_base is up-to-date before building the test image.
# Compares a content hash of ci_base-affecting files against the remote
# image label. If hashes match the build is skipped (< 30 s); if they
# differ ci_base is rebuilt and pushed automatically.
- label: "AMD: :docker: ensure ci_base"
key: ensure-ci-base-amd
depends_on: []
device: amd_cpu
no_plugin: true
commands:
- >
docker build
--build-arg max_jobs=16
--build-arg REMOTE_VLLM=1
--build-arg ARG_PYTORCH_ROCM_ARCH='gfx90a;gfx942;gfx950'
--build-arg VLLM_BRANCH=$BUILDKITE_COMMIT
--tag "rocm/vllm-ci:${BUILDKITE_COMMIT}"
-f docker/Dockerfile.rocm
--target test
--no-cache
--progress plain .
- docker push "rocm/vllm-ci:${BUILDKITE_COMMIT}"
- bash .buildkite/scripts/ci-bake-rocm.sh ci-base-rocm-ci-with-deps
env:
DOCKER_BUILDKIT: "1"
VLLM_BAKE_FILE: "docker/docker-bake-rocm.hcl"
PYTORCH_ROCM_ARCH: "gfx90a;gfx942;gfx950"
REMOTE_VLLM: "1"
VLLM_BRANCH: "$BUILDKITE_COMMIT"
retry:
automatic:
- exit_status: -1 # Agent was lost
limit: 1
- exit_status: -10 # Agent was lost
limit: 1

- label: "AMD: :docker: build test image and artifacts"
key: image-build-amd
depends_on:
- ensure-ci-base-amd
device: amd_cpu
no_plugin: true
commands:
- |
if [[ "${ROCM_CI_ARTIFACT_ONLY:-0}" == "1" ]]; then
echo "ROCM_CI_ARTIFACT_ONLY=1; building ROCm wheel artifact only"
IMAGE_TAG="" bash .buildkite/scripts/ci-bake-rocm.sh test-rocm-ci-with-artifacts
else
bash .buildkite/scripts/ci-bake-rocm.sh test-rocm-ci-with-wheel
fi
- |
docker run --rm --network=none --entrypoint /bin/bash "rocm/vllm-ci:${BUILDKITE_COMMIT}" -ec '
if [ ! -d /vllm-workspace ]; then echo Missing directory: /vllm-workspace >&2; exit 1; fi
if [ ! -d /vllm-workspace/tests ]; then echo Missing directory: /vllm-workspace/tests >&2; exit 1; fi
if [ ! -d /vllm-workspace/src/vllm ]; then echo Missing directory: /vllm-workspace/src/vllm >&2; exit 1; fi
if [ ! -x /vllm-workspace/src/vllm/vllm-rs ]; then echo Missing executable: /vllm-workspace/src/vllm/vllm-rs >&2; exit 1; fi
command -v python3
command -v uv
command -v pytest
if ! command -v amd-smi >/dev/null 2>&1 && ! command -v rocminfo >/dev/null 2>&1; then
echo No ROCm CLI found in image >&2
exit 1
fi
python3 - <<PY
import torch, vllm
print(torch.__version__)
print(vllm.__version__)
PY
echo AMD image smoke OK
'
env:
DOCKER_BUILDKIT: "1"
VLLM_BAKE_FILE: "docker/docker-bake-rocm.hcl"
PYTORCH_ROCM_ARCH: "gfx90a;gfx942;gfx950"
IMAGE_TAG: "rocm/vllm-ci:$BUILDKITE_COMMIT"
REMOTE_VLLM: "1"
VLLM_BRANCH: "$BUILDKITE_COMMIT"
retry:
automatic:
- exit_status: -1 # Agent was lost
limit: 1
- exit_status: -10 # Agent was lost
limit: 1
64 changes: 45 additions & 19 deletions .buildkite/hardware_tests/cpu.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -12,28 +12,35 @@ steps:
- vllm/_custom_ops.py
- tests/kernels/attention/test_cpu_attn.py
- tests/kernels/moe/test_cpu_fused_moe.py
- tests/kernels/moe/test_cpu_quant_fused_moe.py
- tests/kernels/test_onednn.py
- tests/kernels/test_awq_int4_to_int8.py
- tests/kernels/quantization/test_cpu_fp8_scaled_mm.py
- tests/kernels/mamba/cpu/test_cpu_gdn_ops.py
commands:
- |
bash .buildkite/scripts/hardware_ci/run-cpu-test.sh 20m "
bash .buildkite/scripts/hardware_ci/run-cpu-test.sh 30m "
pytest -x -v -s tests/kernels/attention/test_cpu_attn.py
pytest -x -v -s tests/kernels/moe/test_cpu_fused_moe.py
pytest -x -v -s tests/kernels/moe/test_cpu_quant_fused_moe.py
pytest -x -v -s tests/kernels/test_onednn.py
pytest -x -v -s tests/kernels/test_awq_int4_to_int8.py"
pytest -x -v -s tests/kernels/test_awq_int4_to_int8.py
pytest -x -v -s tests/kernels/quantization/test_cpu_fp8_scaled_mm.py
pytest -x -v -s tests/kernels/mamba/cpu/test_cpu_gdn_ops.py"

- label: CPU-Compatibility Tests
depends_on: []
device: intel_cpu
no_plugin: true
source_file_dependencies:
- cmake/cpu_extension.cmake
- setup.py
- vllm/platforms/cpu.py
commands:
- |
bash .buildkite/scripts/hardware_ci/run-cpu-test.sh 20m "
bash .buildkite/scripts/hardware_ci/run-cpu-compatibility-test.sh"
# Note: SDE can't be downloaded from CI host because of AWS WAF
# - label: CPU-Compatibility Tests
# depends_on: []
# device: intel_cpu
# no_plugin: true
# source_file_dependencies:
# - cmake/cpu_extension.cmake
# - setup.py
# - vllm/platforms/cpu.py
# commands:
# - |
# bash .buildkite/scripts/hardware_ci/run-cpu-test.sh 20m "
# bash .buildkite/scripts/hardware_ci/run-cpu-compatibility-test.sh"

- label: CPU-Language Generation and Pooling Model Tests
depends_on: []
Expand All @@ -50,22 +57,41 @@ steps:
pytest -x -v -s tests/models/language/generation -m cpu_model
pytest -x -v -s tests/models/language/pooling -m cpu_model"

- label: CPU-ModelRunnerV2 Tests
depends_on: []
device: intel_cpu
no_plugin: true
soft_fail: true
source_file_dependencies:
- vllm/v1/worker/cpu/
- vllm/v1/worker/gpu/
- vllm/v1/sample/ops/topk_topp_triton.py
- vllm/v1/sample/ops/topk_topp_sampler.py
- tests/v1/sample/test_topk_topp_sampler.py
commands:
- |
bash .buildkite/scripts/hardware_ci/run-cpu-test.sh 45m "
uv pip install git+https://github.com/triton-lang/triton-cpu.git@270e696d
VLLM_USE_V2_MODEL_RUNNER=1 pytest -x -v -s tests/models/language/generation/test_granite.py -m cpu_model
# TODO: move to CPU-Kernel Tests once triton-cpu has a pre-built wheel
pytest -x -v -s tests/v1/sample/test_topk_topp_sampler.py::TestTritonTopkTopp"

- label: CPU-Quantization Model Tests
depends_on: []
device: intel_cpu
no_plugin: true
source_file_dependencies:
- csrc/cpu/
- vllm/model_executor/layers/quantization/cpu_wna16.py
- vllm/model_executor/layers/quantization/gptq_marlin.py
- vllm/model_executor/layers/quantization/auto_gptq.py
- vllm/model_executor/layers/quantization/compressed_tensors/schemes/compressed_tensors_w8a8_int8.py
- vllm/model_executor/layers/quantization/kernels/scaled_mm/cpu.py
- vllm/model_executor/layers/quantization/kernels/mixed_precision/cpu.py
- vllm/model_executor/kernels/linear/mixed_precision/cpu.py
- vllm/model_executor/kernels/linear/scaled_mm/cpu.py
- vllm/model_executor/layers/fused_moe/experts/cpu_moe.py
- tests/quantization/test_compressed_tensors.py
- tests/quantization/test_cpu_wna16.py
commands:
- |
bash .buildkite/scripts/hardware_ci/run-cpu-test.sh 20m "
bash .buildkite/scripts/hardware_ci/run-cpu-test.sh 30m "
pytest -x -v -s tests/quantization/test_compressed_tensors.py::test_compressed_tensors_w8a8_logprobs
pytest -x -v -s tests/quantization/test_cpu_wna16.py"

Expand Down
7 changes: 0 additions & 7 deletions .buildkite/hardware_tests/intel.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -8,10 +8,3 @@ steps:
commands:
- bash .buildkite/scripts/hardware_ci/run-hpu-test.sh

- label: "Intel GPU Test"
depends_on: []
soft_fail: true
device: intel_gpu
no_plugin: true
commands:
- bash .buildkite/scripts/hardware_ci/run-xpu-test.sh
72 changes: 72 additions & 0 deletions .buildkite/image_build/image_build.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -13,6 +13,60 @@ steps:
- exit_status: -10 # Agent was lost
limit: 2

- label: ":docker: :smoking: Non-root smoke tests"
key: image-build-smoke-test
depends_on:
- image-build
commands:
# Smoke 1: the default (root) image must still be importable
# under a non-root UID via `--user 2000:0`. Validates the `vllm` passwd
# entry + group-0-writable /home/vllm + uv path cleanup from #31959.
# Uses `import vllm` rather than `vllm serve --help` because the latter
# instantiates `VllmConfig` which requires a GPU attached to the
# container.
- docker run --rm --user 2000:0 --entrypoint python3 "$IMAGE_TAG" -c "import vllm; print(vllm.__version__)"
# Smoke 2: assert the non-root enabling invariants are baked
# into the image. Runs as UID 2000:0 via a shell so we can verify
# filesystem perms + passwd/group file state + wrapper presence without
# triggering vLLM's GPU-requiring config-init path. The opt-in
# `vllm-openai-nonroot` target adds only `USER vllm`, `WORKDIR
# /home/vllm`, and an `ENTRYPOINT` override on top of these invariants;
# its build correctness is reviewed at the Dockerfile level. Wrapper
# logic is covered separately by the pre-commit hook
# `test-nonroot-entrypoint` (see .pre-commit-config.yaml).
- |
docker run --rm --user 2000:0 --entrypoint /bin/sh "$IMAGE_TAG" -ec '
if ! getent passwd 2000 | grep -q ^vllm:; then
echo FAIL: UID 2000 != vllm
exit 1
fi
if ! id -gn 2>/dev/null | grep -qx root; then
echo FAIL: GID 0 not root group
exit 1
fi
touch /home/vllm/.smoke && rm /home/vllm/.smoke
touch /opt/uv/cache/.smoke && rm /opt/uv/cache/.smoke
if ! test -x /usr/local/bin/vllm-nonroot-entrypoint.sh; then
echo FAIL: wrapper missing
exit 1
fi
if ! test -w /etc/passwd; then
echo FAIL: /etc/passwd not group-writable
exit 1
fi
if ! test -w /etc/group; then
echo FAIL: /etc/group not group-writable
exit 1
fi
echo non-root invariants OK
'
retry:
automatic:
- exit_status: -1 # Agent was lost
limit: 2
- exit_status: -10 # Agent was lost
limit: 2

- label: ":docker: Build CPU image"
key: image-build-cpu
depends_on: []
Expand Down Expand Up @@ -56,3 +110,21 @@ steps:
limit: 2
- exit_status: -10 # Agent was lost
limit: 2

- label: ":docker: Build arm64 image"
key: arm64-image-build
depends_on: []
source_file_dependencies:
- ".buildkite/image_build/image_build.yaml"
- ".buildkite/image_build/image_build_arm64.sh"
- "docker/Dockerfile"
commands:
- .buildkite/image_build/image_build_arm64.sh $REGISTRY $REPO $BUILDKITE_COMMIT
env:
DOCKER_BUILDKIT: "1"
retry:
automatic:
- exit_status: -1 # Agent was lost
limit: 2
- exit_status: -10 # Agent was lost
limit: 2
37 changes: 37 additions & 0 deletions .buildkite/image_build/image_build_arm64.sh
Original file line number Diff line number Diff line change
@@ -0,0 +1,37 @@
#!/bin/bash
set -e

if [[ $# -lt 3 ]]; then
echo "Usage: $0 <registry> <repo> <commit>"
exit 1
fi

REGISTRY=$1
REPO=$2
BUILDKITE_COMMIT=$3

# authenticate with AWS ECR
aws ecr-public get-login-password --region us-east-1 | docker login --username AWS --password-stdin "$REGISTRY" || true

# skip build if image already exists
if [[ -z $(docker manifest inspect "$REGISTRY"/"$REPO":"$BUILDKITE_COMMIT"-arm64) ]]; then
echo "Image not found, proceeding with build..."
else
echo "Image found"
exit 0
fi

# build (Grace/GH200 is the arm64 GPU target; sm_90)
docker build --file docker/Dockerfile \
--platform linux/arm64 \
--build-arg max_jobs=16 \
--build-arg nvcc_threads=4 \
--build-arg torch_cuda_arch_list="9.0" \
--build-arg USE_SCCACHE=1 \
--build-arg buildkite_commit="$BUILDKITE_COMMIT" \
--tag "$REGISTRY"/"$REPO":"$BUILDKITE_COMMIT"-arm64 \
--target test \
--progress plain .

# push
docker push "$REGISTRY"/"$REPO":"$BUILDKITE_COMMIT"-arm64
2 changes: 1 addition & 1 deletion .buildkite/image_build/image_build_hpu.sh
Original file line number Diff line number Diff line change
Expand Up @@ -11,7 +11,7 @@ REPO=$2
BUILDKITE_COMMIT=$3

# authenticate with AWS ECR
aws ecr-public get-login-password --region us-east-1 | docker login --username AWS --password-stdin "$REGISTRY"
aws ecr-public get-login-password --region us-east-1 | docker login --username AWS --password-stdin "$REGISTRY" || true

# skip build if image already exists
if [[ -z $(docker manifest inspect "$REGISTRY"/"$REPO":"$BUILDKITE_COMMIT"-hpu) ]]; then
Expand Down
4 changes: 2 additions & 2 deletions .buildkite/image_build/image_build_xpu.sh
Original file line number Diff line number Diff line change
Expand Up @@ -11,8 +11,8 @@ REPO=$2
BUILDKITE_COMMIT=$3

# authenticate with AWS ECR
aws ecr-public get-login-password --region us-east-1 | docker login --username AWS --password-stdin "$REGISTRY"
aws ecr get-login-password --region us-east-1 | docker login --username AWS --password-stdin 936637512419.dkr.ecr.us-east-1.amazonaws.com
aws ecr-public get-login-password --region us-east-1 | docker login --username AWS --password-stdin "$REGISTRY" || true
aws ecr get-login-password --region us-east-1 | docker login --username AWS --password-stdin 936637512419.dkr.ecr.us-east-1.amazonaws.com || true

# skip build if image already exists
if ! docker manifest inspect "$REGISTRY"/"$REPO":"$BUILDKITE_COMMIT"-xpu &> /dev/null; then
Expand Down
Loading
Loading