-
Notifications
You must be signed in to change notification settings - Fork 5k
Pull requests: deepspeedai/DeepSpeed
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
fix(inference): break circular import in ops.transformer.inference
#8465
opened Sep 9, 2026 by
chakshu-dhannawat
Contributor
Loading…
Fix Muon optimizer under ZeRO CPU offload and bound gather buffers
#8464
opened Sep 8, 2026 by
jinyouzhi
Contributor
Loading…
Checkpoint the RNG the curriculum sampler actually draws from
#8460
opened Sep 8, 2026 by
alanhuangyoo
Contributor
Loading…
Say when ZenFlow will never re-select its important columns
#8456
opened Sep 8, 2026 by
alanhuangyoo
Contributor
Loading…
Rebuild the full argument list when partitioning activations
#8455
opened Sep 8, 2026 by
vineethsaivs
Contributor
Loading…
fix(inference): reject non-positive max_out_tokens at config validation
#8454
opened Sep 8, 2026 by
chakshu-dhannawat
Contributor
Loading…
[Phase 2] Add NEON SIMD path for CPU Adam on AArch64
#8453
opened Sep 8, 2026 by
PKUWZP
Collaborator
Loading…
Sort checkpoint shard files by numeric rank
#8452
opened Sep 8, 2026 by
ebarkhordar
Contributor
Loading…
docs(runtime): document max_norm and use_graph in clip_tensors_by_global_norm
#8450
opened Sep 7, 2026 by
simpleqt
Loading…
docs: point CONTRIBUTING's installation link at the README section
#8449
opened Sep 7, 2026 by
simpleqt
Loading…
docs(sparse_attention): fix copy-pasted BigBird return line in sliding-window layout
#8448
opened Sep 7, 2026 by
simpleqt
Loading…
fix(checkpointing): remove identical if/else arms and leftover debug print in WriterFactory
#8446
opened Sep 6, 2026 by
simpleqt
Loading…
2 tasks done
fix(zenflow): restore assert and remove duplicated identical branch in gradient copy
#8445
opened Sep 6, 2026 by
simpleqt
Loading…
DeepSpeed cannot start without mpi4py on a machine with no launcher
#8444
opened Sep 6, 2026 by
alanhuangyoo
Contributor
Loading…
Muon silently discards the param groups it is given
#8440
opened Sep 6, 2026 by
alanhuangyoo
Contributor
Loading…
Muon is silently disabled under ZeRO-3 when the model is built with zero.Init
#8438
opened Sep 6, 2026 by
alanhuangyoo
Contributor
Loading…
[muon] Per-head Muon for linear-attention layers: read head geometry from the owning module
#8436
opened Sep 6, 2026 by
alanhuangyoo
Contributor
Loading…
[muon] Keep the momentum out of steps the loss scaler discards
#8435
opened Sep 6, 2026 by
alanhuangyoo
Contributor
Loading…
Fix SequenceTiledCompute backward for empty trailing shards
#8434
opened Sep 6, 2026 by
taking-lying-flat
Loading…
[muon] Reconcile the momentum dtype when a checkpoint is restored
#8433
opened Sep 6, 2026 by
alanhuangyoo
Contributor
Loading…
ZeRO-3: adaptive prefetch bucket size
#8431
opened Sep 6, 2026 by
promptsmith1990
Contributor
Loading…
Guard steps_per_print() None check in PipelineEngine._exec_optimizer_step
#8430
opened Sep 6, 2026 by
promptsmith1990
Contributor
Loading…
Previous Next
ProTip!
Mix and match filters to narrow down what you’re looking for.