Megatron FSDP Skill
For stable background and recommendation level, see:
- @docs/training/megatron-fsdp.md
- @skills/nemo-mbridge-perf-megatron-fsdp/card.yaml
Enablement
Minimal Megatron FSDP override in Bridge:
Example recipe fixup:
Performance harness note:
Code Anchors
Bridge config definition:
Bridge validation:
Runtime wrapper selection:
Perf harness overrides:
Pitfalls
- Public recipes often expose
use_megatron_fsdpbut still default tockpt_format="torch_dist". If save/load is enabled, switch tofsdp_dtensor. use_torch_fsdp2exists, but on the validated branch Bridge still fails before training because_ddp_wrappassespg_collection.- CPU offloading is only valid when
pipeline_model_parallel_size == 1and activation recomputation is disabled. - Upstream warns that FSDP and TP/CP can want different
CUDA_DEVICE_MAX_CONNECTIONSsettings on Hopper and earlier. - Megatron FSDP and FSDP2 are mutually exclusive.
Verification
Use the existing 2-GPU functional smoke test:
Success criteria:
- Pytest reports
1 passed - The log shows finite loss at the last iteration
- The run finishes without a checkpoint format assertion


