-
Notifications
You must be signed in to change notification settings - Fork 2.7k
Pull requests: NVIDIA/TensorRT-LLM
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
[https://nvbugs/6566891][fix] Use FlashInfer FA2 for Gemma4 on SM120 and SM121
#17557
opened Aug 12, 2026 by
lfr-0531
Collaborator
Loading…
1 task done
[TRTLLM-15309][feat] VisualGen MGMN mode: run multi-node under trtllm-llmapi-launch
VisualGen
#17556
opened Aug 12, 2026 by
zhenhuaw-me
Member
•
Draft
1 task done
[None][chore] Drop the skip_* sampling flags from one-model spec metadata
#17554
opened Aug 12, 2026 by
zhaoyangwang-nvidia
Collaborator
Loading…
5 tasks done
[None][fix] Forward reasoning_effort to the chat template
#17553
opened Aug 12, 2026 by
joerowell
Contributor
Loading…
[None][feat] Address a speculative draft by subfolder of its repo
#17552
opened Aug 12, 2026 by
joerowell
Contributor
Loading…
[None][fix] Honour producer ignored layers for fused MoE and any quant format
#17551
opened Aug 12, 2026 by
joerowell
Contributor
Loading…
[None][fix] GVR indexer top-K: repair the non-converged threshold search
#17550
opened Aug 12, 2026 by
longcheng-nv
Collaborator
Loading…
1 task done
[None][test] Replace disaggregated DWDP accuracy tests with aggregated coverage
ci: full pre-merge approved
#17546
opened Aug 12, 2026 by
tianyuz-nv
Collaborator
Loading…
2 of 3 tasks
feat: Add adaptive speculative decoding engine and benchmark results
#17545
opened Aug 12, 2026 by
tolani007
Loading…
1 task done
[TRTLLM-13215][perf] Skip redundant one-model sampling-param refills
#17544
opened Aug 12, 2026 by
zhaoyangwang-nvidia
Collaborator
Loading…
6 tasks done
[None][infra] Disable RTXPro6000D multi gpu stages due to some nodes are offline
#17543
opened Aug 12, 2026 by
yiqingy0
Collaborator
Loading…
1 task
[TRTLLM-14834][chore] Introduce observability/ for logging and profiling
api-compatible
Accepted LLM API contract change that is backwards-compatible
VisualGen
#17541
opened Aug 12, 2026 by
YihuiLu512
Collaborator
•
Draft
1 task done
[None][fix] Report Server-Timing on the video tensor route
VisualGen
#17540
opened Aug 12, 2026 by
karljang
Collaborator
Loading…
[#17522][fix] Restore conv_state ordering barriers in the Triton conv kernels
#17539
opened Aug 12, 2026 by
wilyan09007
Loading…
1 task done
[None][feat] Don't review PR stacks to support new model
#17537
opened Aug 12, 2026 by
Wanli-Jiang
Collaborator
•
Draft
[TRTLLM-14832][chore] Introduce executor/params/ for public request contracts
#17536
opened Aug 12, 2026 by
lori-ren
Contributor
Loading…
1 task done
[None][chore] Stop forcing TRTLLM_DISABLE_KV_CACHE_TRANSFER_OVERLAP=1 in disagg gen_only benchmark
#17535
opened Aug 12, 2026 by
dc3671
Collaborator
Loading…
3 of 4 tasks
[None][fix] Clamp conversation-affinity ADP routing to per-rank slot
#17534
opened Aug 12, 2026 by
lancelly
Collaborator
Loading…
[None][infra] Report Blossom trigger failures
#17533
opened Aug 12, 2026 by
Mgluhovskoi
Collaborator
Loading…
[TRTLLM-14956][refactor] make MoE implementation selection reproducible
#17532
opened Aug 12, 2026 by
xxi-nv
Collaborator
Loading…
3 of 4 tasks
[None][perf] Cut per-iteration executor bookkeeping in the hang detector and profiler
#17531
opened Aug 12, 2026 by
pranav-nvidia
Contributor
•
Draft
Previous Next
ProTip!
Follow long discussions with comments:>50.