Skip to content

Pull requests: NVIDIA/TransformerEngine

Author
Filter by author
Loading
Label
Filter by label
Loading
Use alt + click/return to exclude labels
or + click/return for logical OR
Projects
Filter by project
Loading
Milestones
Filter by milestone
Loading
Reviews
Assignee
Filter by who’s assigned
Assigned to nobody Loading
Sort

Pull requests list

[PyTorch] Extend no-load-balance CP to the a2a comm type community-contribution PRs from external contributor outside the core maintainers, representing community-driven work.
#3530 opened Sep 16, 2026 by Rudin6 Loading…
13 tasks
[PyTorch] Allow F16 GDN inside FP8 autocast community-contribution PRs from external contributor outside the core maintainers, representing community-driven work.
#3529 opened Sep 16, 2026 by layalir Contributor Loading…
4 of 5 tasks
feat(attention): cuDNN FROST attention backend for head_dim in (256, 512], with context parallelism community-contribution PRs from external contributor outside the core maintainers, representing community-driven work.
#3527 opened Sep 16, 2026 by nvegesna-netizen Contributor Loading…
[Pytorch][NCCL EP] Create tokens_per_expert on pinned CPU memory when running eager
#3526 opened Sep 16, 2026 by YangFei1990 Collaborator Loading…
8 of 13 tasks
[JAX] WAR for XLA associative scan rewriter crash 2.20
#3525 opened Sep 16, 2026 by KshitijLakhani Collaborator Loading…
2 of 13 tasks
[JAX] Remove duplicate score-mod test run 2.20
#3524 opened Sep 16, 2026 by KshitijLakhani Collaborator Loading…
1 of 13 tasks
Remove the view() calls from Linear/LNLinear/LNMLP modules
#3523 opened Sep 15, 2026 by ptrendx Member Loading…
1 of 13 tasks
[JAX] Support single-process-multi-device EP via XLA-borrowed communicator
#3522 opened Sep 15, 2026 by phu0ngng Collaborator Loading…
8 of 13 tasks
[PyTorch] GDN2 support and linear attention refactor 2.20
#3521 opened Sep 15, 2026 by ksivaman Member Loading…
10 of 13 tasks
Accept compact MXFP8 scales in GEMM preparation community-contribution PRs from external contributor outside the core maintainers, representing community-driven work.
#3519 opened Sep 15, 2026 by wujingyue Contributor Draft
[JAX] Align fused attention backward output gradient sharding 2.20 attention bug Something isn't working
#3516 opened Sep 14, 2026 by KshitijLakhani Collaborator Draft
1 of 13 tasks
Support building docs for publishing via GHA
#3515 opened Sep 14, 2026 by fheinecke Collaborator Loading…
1 of 13 tasks
Abstract CUDA hardcodes into configurable te_device_type / te_platform community-contribution PRs from external contributor outside the core maintainers, representing community-driven work.
#3513 opened Sep 14, 2026 by lxd-cumt Contributor Loading…
[PyTorch] Add head-parallel FA4 backward for cuDNN CP attention attention community-contribution PRs from external contributor outside the core maintainers, representing community-driven work.
#3510 opened Sep 11, 2026 by bzantium Loading…
8 of 13 tasks
[PyTorch] Fuse MoE chunk sorting and padding community-contribution PRs from external contributor outside the core maintainers, representing community-driven work.
#3509 opened Sep 11, 2026 by bzantium Loading…
7 of 13 tasks
[Draft] Port cuDNN frontend attention to Python API
#3508 opened Sep 11, 2026 by vcherepanov-nv Collaborator Draft
13 tasks
[PyTorch] Avoid temporary state casts when loading FusedAdam checkpoints community-contribution PRs from external contributor outside the core maintainers, representing community-driven work.
#3507 opened Sep 11, 2026 by bzantium Draft
5 of 13 tasks
Fix THD P2P pad detection and tail-zero gating community-contribution PRs from external contributor outside the core maintainers, representing community-driven work.
#3506 opened Sep 10, 2026 by RPalmr Loading…
[PyTorch] Support FP8 weight caching in compiled Linear
#3505 opened Sep 10, 2026 by pggPL Collaborator Loading…
6 of 8 tasks
[Pytorch][Attention] Fix FP8 THD backward workspace scaling
#3504 opened Sep 10, 2026 by sudhakarsingh27 Member Loading…
4 of 13 tasks
MOE Sequential Block with Dispatch and Combine as Basic Ops 2.20
#3503 opened Sep 10, 2026 by vthumbe1503 Collaborator Loading…
13 tasks
[PyTorch] Add timestep-conditioned AdaptiveLayerNorm to op fuser community-contribution PRs from external contributor outside the core maintainers, representing community-driven work.
#3501 opened Sep 10, 2026 by zupengwang Loading…
8 of 9 tasks
Fix hybrid extra state size mismatch
#3499 opened Sep 9, 2026 by negvet Collaborator Loading…
13 tasks
[JAX] Fix undefined sr_rng_state that breaks --dry-run in two encoder examples community-contribution PRs from external contributor outside the core maintainers, representing community-driven work.
#3498 opened Sep 9, 2026 by Anai-Guo Contributor Loading…
ProTip! Mix and match filters to narrow down what you’re looking for.