-
-
Notifications
You must be signed in to change notification settings - Fork 21.7k
Pull requests: vllm-project/vllm
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
[Bugfix] Rank NVFP4 CuteDSL W4A16 kernel below the native W4A4 kernels
bug
Something isn't working
#55405
opened Sep 4, 2026 by
iharshlalakiya
Loading…
3 of 4 tasks
[Docs] Remove dead snippet includes left behind by refactors
documentation
Improvements or additions to documentation
#55402
opened Sep 4, 2026 by
simpleqt
Loading…
2 tasks done
[Bugfix][Core] Fix logging calls whose arguments do not match their placeholders
bug
Something isn't working
kv-connector
#55401
opened Sep 4, 2026 by
simpleqt
Loading…
[Feature][Core][Rust Frontend] Add a Rejected finish reason for engine-stated request rejections
rust
#55399
opened Sep 4, 2026 by
FeathBow
Contributor
Loading…
[5/N] HiSparse: support direct GPU landing for P/D transfers
kv-cache-manager
kv-connector
scheduler
#55398
opened Sep 4, 2026 by
MatthewBonanni
Member
•
4/6
•
Draft
[Bugfix] do not treat weight filenames as custom proposer paths
bug
Something isn't working
#55393
opened Sep 4, 2026 by
Chessing234
Loading…
1 task done
Revert "[CI] Remove deleted nvidia/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-BF16 and its arch aliases"
nvidia
ready
ONLY add when PR is ready to merge/full CI is needed
#55392
opened Sep 4, 2026 by
mgoin
Member
Loading…
[Bugfix] reject empty prompt lists in CompletionRequest
bug
Something isn't working
frontend
#55391
opened Sep 4, 2026 by
Chessing234
Loading…
1 task done
[Bugfix] Annotate MTP draft KV cache groups positionally on the hybrid grouping path
bug
Something isn't working
kv-cache-manager
verified
Run pre-commit for new contributors without triggering other tests
#55390
opened Sep 4, 2026 by
Navjot10
Loading…
[Model] Extend device-side mm normalization to GLM4V/GLM5Next
glm
#55389
opened Sep 4, 2026 by
cjackal
Contributor
Loading…
4 tasks done
[Bugfix] error cleanly when vllm serve --model has no value
bug
Something isn't working
#55388
opened Sep 4, 2026 by
Chessing234
Loading…
1 task done
[CI/Build] Move the default CUDA build from 13.0 to 13.2
ci/build
documentation
Improvements or additions to documentation
nvidia
#55387
opened Sep 4, 2026 by
atalman
Contributor
Loading…
[Bugfix][Compile] Fold LoRA wrap state into the AOT cache key
bug
Something isn't working
torch.compile
#55386
opened Sep 4, 2026 by
he-yufeng
Contributor
Loading…
[perf] wire FA and FlashMLA for sm90 GLM5Next NoPE SparseMLA
glm
#55385
opened Sep 4, 2026 by
JaredforReal
Contributor
•
Draft
[Bugfix] Share resolved KV cache layout across config copies
bug
Something isn't working
kv-cache-manager
kv-connector
nvidia
ready
ONLY add when PR is ready to merge/full CI is needed
#55384
opened Sep 4, 2026 by
LucasWilkinson
Collaborator
Loading…
[Turing] Restore flashinfer support for SM75
nvidia
#55380
opened Sep 4, 2026 by
ir1ka
Contributor
Loading…
6 of 7 tasks
[ROCm] Hold the router weight in fp32 when fp32 routing is requested
deepseek
Related to DeepSeek models
rocm
Related to AMD ROCm
#55378
opened Sep 4, 2026 by
stefanskiasan
Loading…
[Bugfix] Autotune FlashInfer deferred MoE decode kernels before CUDA graph capture
bug
Something isn't working
nvidia
#55377
opened Sep 4, 2026 by
mingg26
Contributor
Loading…
[Build] Remove obsolete TPU Dockerfile
ci/build
ready
ONLY add when PR is ready to merge/full CI is needed
#55376
opened Sep 4, 2026 by
WoosukKwon
Collaborator
Loading…
4 tasks done
[Bugfix][Qwen4Exp] fix state index strides in fused PLE conv
bug
Something isn't working
qwen
Related to Qwen models
#55375
opened Sep 4, 2026 by
peakcrosser7
Contributor
Loading…
3 of 4 tasks
[KV Connector] Support piecewise prefix loading with NIXL connectors
kv-connector
#55374
opened Sep 4, 2026 by
z-zanez
Contributor
Loading…
3 tasks done
Previous Next
ProTip!
Add no:assignee to see everything that’s not assigned.