Skip to content

Pull requests: vllm-project/vllm

Author
Filter by author
Loading
Label
Filter by label
Loading
Use alt + click/return to exclude labels
or + click/return for logical OR
Projects
Filter by project
Loading
Milestones
Filter by milestone
Loading
Reviews
Assignee
Filter by who’s assigned
Assigned to nobody Loading
Sort

Pull requests list

[Bugfix] Rank NVFP4 CuteDSL W4A16 kernel below the native W4A4 kernels bug Something isn't working
#55405 opened Sep 4, 2026 by iharshlalakiya Loading…
3 of 4 tasks
[Docs] Remove dead snippet includes left behind by refactors documentation Improvements or additions to documentation
#55402 opened Sep 4, 2026 by simpleqt Loading…
2 tasks done
[Bugfix][Parser] Hide incomplete Muse Glimmer ATEM markers bug Something isn't working tool-calling
#55396 opened Sep 4, 2026 by lming2001 Draft
3 of 4 tasks
[Bugfix] do not treat weight filenames as custom proposer paths bug Something isn't working
#55393 opened Sep 4, 2026 by Chessing234 Loading…
1 task done
Revert "[CI] Remove deleted nvidia/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-BF16 and its arch aliases" nvidia ready ONLY add when PR is ready to merge/full CI is needed
#55392 opened Sep 4, 2026 by mgoin Member Loading…
[Bugfix] reject empty prompt lists in CompletionRequest bug Something isn't working frontend
#55391 opened Sep 4, 2026 by Chessing234 Loading…
1 task done
[Bugfix] Annotate MTP draft KV cache groups positionally on the hybrid grouping path bug Something isn't working kv-cache-manager verified Run pre-commit for new contributors without triggering other tests
#55390 opened Sep 4, 2026 by Navjot10 Loading…
[Model] Extend device-side mm normalization to GLM4V/GLM5Next glm
#55389 opened Sep 4, 2026 by cjackal Contributor Loading…
4 tasks done
[Bugfix] error cleanly when vllm serve --model has no value bug Something isn't working
#55388 opened Sep 4, 2026 by Chessing234 Loading…
1 task done
[CI/Build] Move the default CUDA build from 13.0 to 13.2 ci/build documentation Improvements or additions to documentation nvidia
#55387 opened Sep 4, 2026 by atalman Contributor Loading…
[Bugfix][Compile] Fold LoRA wrap state into the AOT cache key bug Something isn't working torch.compile
#55386 opened Sep 4, 2026 by he-yufeng Contributor Loading…
[Bugfix] Share resolved KV cache layout across config copies bug Something isn't working kv-cache-manager kv-connector nvidia ready ONLY add when PR is ready to merge/full CI is needed
#55384 opened Sep 4, 2026 by LucasWilkinson Collaborator Loading…
[Turing] Restore flashinfer support for SM75 nvidia
#55380 opened Sep 4, 2026 by ir1ka Contributor Loading…
6 of 7 tasks
[CI] Disable Ascend NPU test ci/build
#55379 opened Sep 4, 2026 by khluu Member Loading…
[ROCm] Hold the router weight in fp32 when fp32 routing is requested deepseek Related to DeepSeek models rocm Related to AMD ROCm
#55378 opened Sep 4, 2026 by stefanskiasan Loading…
[Bugfix] Autotune FlashInfer deferred MoE decode kernels before CUDA graph capture bug Something isn't working nvidia
#55377 opened Sep 4, 2026 by mingg26 Contributor Loading…
[Build] Remove obsolete TPU Dockerfile ci/build ready ONLY add when PR is ready to merge/full CI is needed
#55376 opened Sep 4, 2026 by WoosukKwon Collaborator Loading…
4 tasks done
[Bugfix][Qwen4Exp] fix state index strides in fused PLE conv bug Something isn't working qwen Related to Qwen models
#55375 opened Sep 4, 2026 by peakcrosser7 Contributor Loading…
3 of 4 tasks
[KV Connector] Support piecewise prefix loading with NIXL connectors kv-connector
#55374 opened Sep 4, 2026 by z-zanez Contributor Loading…
3 tasks done
ProTip! Add no:assignee to see everything that’s not assigned.