3D Parallelism (TP/PP/DP), Context Parallelism (CP), PyTorch FSDP, DeepSpeed ZeRO memory sharding, long-horizon fault tolerance, and async checkpoint recovery.
Megatron-LM TP/PP communication scheduling, Context/Sequence Parallelism (CP/SP) for long context, PyTorch FSDP, DeepSpeed ZeRO-1/2/3 parameter/optimizer sharding, and activation checkpointing memory trade-offs.
10K+ GPU async checkpoint persistence and rapid recovery, straggler and silent hang detection, FP8/BF16 mixed-precision numerical stability, automated loss spike remediation, and W&B telemetry.