-
Notifications
You must be signed in to change notification settings - Fork 825
Pull requests: NVIDIA/TransformerEngine
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
[PyTorch] Extend no-load-balance CP to the a2a comm type
community-contribution
PRs from external contributor outside the core maintainers, representing community-driven work.
#3530
opened Sep 16, 2026 by
Rudin6
Loading…
13 tasks
[PyTorch] Allow F16 GDN inside FP8 autocast
community-contribution
PRs from external contributor outside the core maintainers, representing community-driven work.
#3529
opened Sep 16, 2026 by
layalir
Contributor
Loading…
4 of 5 tasks
feat(attention): cuDNN FROST attention backend for head_dim in (256, 512], with context parallelism
community-contribution
PRs from external contributor outside the core maintainers, representing community-driven work.
#3527
opened Sep 16, 2026 by
nvegesna-netizen
Contributor
Loading…
[Pytorch][NCCL EP] Create tokens_per_expert on pinned CPU memory when running eager
#3526
opened Sep 16, 2026 by
YangFei1990
Collaborator
Loading…
8 of 13 tasks
[JAX] WAR for XLA associative scan rewriter crash
2.20
#3525
opened Sep 16, 2026 by
KshitijLakhani
Collaborator
Loading…
2 of 13 tasks
[JAX] Remove duplicate score-mod test run
2.20
#3524
opened Sep 16, 2026 by
KshitijLakhani
Collaborator
Loading…
1 of 13 tasks
Remove the view() calls from Linear/LNLinear/LNMLP modules
#3523
opened Sep 15, 2026 by
ptrendx
Member
Loading…
1 of 13 tasks
[JAX] Support single-process-multi-device EP via XLA-borrowed communicator
#3522
opened Sep 15, 2026 by
phu0ngng
Collaborator
Loading…
8 of 13 tasks
[PyTorch] GDN2 support and linear attention refactor
2.20
#3521
opened Sep 15, 2026 by
ksivaman
Member
Loading…
10 of 13 tasks
Accept compact MXFP8 scales in GEMM preparation
community-contribution
PRs from external contributor outside the core maintainers, representing community-driven work.
[PyTorch] Support distributed weights in GroupedLinear's grouped-tensor path
org-contribution
#3517
opened Sep 15, 2026 by
fanshiqing
Member
Loading…
7 of 8 tasks
[JAX] Align fused attention backward output gradient sharding
2.20
attention
bug
Something isn't working
#3516
opened Sep 14, 2026 by
KshitijLakhani
Collaborator
•
Draft
1 of 13 tasks
Support building docs for publishing via GHA
#3515
opened Sep 14, 2026 by
fheinecke
Collaborator
Loading…
1 of 13 tasks
Abstract CUDA hardcodes into configurable te_device_type / te_platform
community-contribution
PRs from external contributor outside the core maintainers, representing community-driven work.
#3513
opened Sep 14, 2026 by
lxd-cumt
Contributor
Loading…
[PyTorch] Add head-parallel FA4 backward for cuDNN CP attention
attention
community-contribution
PRs from external contributor outside the core maintainers, representing community-driven work.
#3510
opened Sep 11, 2026 by
bzantium
Loading…
8 of 13 tasks
[PyTorch] Fuse MoE chunk sorting and padding
community-contribution
PRs from external contributor outside the core maintainers, representing community-driven work.
#3509
opened Sep 11, 2026 by
bzantium
Loading…
7 of 13 tasks
[Draft] Port cuDNN frontend attention to Python API
#3508
opened Sep 11, 2026 by
vcherepanov-nv
Collaborator
•
Draft
13 tasks
[PyTorch] Avoid temporary state casts when loading FusedAdam checkpoints
community-contribution
PRs from external contributor outside the core maintainers, representing community-driven work.
Fix THD P2P pad detection and tail-zero gating
community-contribution
PRs from external contributor outside the core maintainers, representing community-driven work.
#3506
opened Sep 10, 2026 by
RPalmr
Loading…
[PyTorch] Support FP8 weight caching in compiled Linear
#3505
opened Sep 10, 2026 by
pggPL
Collaborator
Loading…
6 of 8 tasks
[Pytorch][Attention] Fix FP8 THD backward workspace scaling
#3504
opened Sep 10, 2026 by
sudhakarsingh27
Member
Loading…
4 of 13 tasks
MOE Sequential Block with Dispatch and Combine as Basic Ops
2.20
#3503
opened Sep 10, 2026 by
vthumbe1503
Collaborator
Loading…
13 tasks
[PyTorch] Add timestep-conditioned AdaptiveLayerNorm to op fuser
community-contribution
PRs from external contributor outside the core maintainers, representing community-driven work.
#3501
opened Sep 10, 2026 by
zupengwang
Loading…
8 of 9 tasks
Fix hybrid extra state size mismatch
#3499
opened Sep 9, 2026 by
negvet
Collaborator
Loading…
13 tasks
[JAX] Fix undefined sr_rng_state that breaks --dry-run in two encoder examples
community-contribution
PRs from external contributor outside the core maintainers, representing community-driven work.
#3498
opened Sep 9, 2026 by
Anai-Guo
Contributor
Loading…
Previous Next
ProTip!
Mix and match filters to narrow down what you’re looking for.