Nemo Mbridge Perf Megatron Fsdp

nvidia/skills/skills/nemo-mbridge-perf-megatron-fsdp

作者 nvidiacf5224d14250Apache-2.03.5K 個星標收錄於 2026年10月8日更新於 2026年10月8日儲存庫今天更新

Operational guide for enabling Megatron FSDP in Megatron-Bridge, including config knobs, code anchors, pitfalls, and verification.

僅含說明DevOps & Cloud
AI 產生的概覽

在 Megatron-Bridge 中啟用 Megatron FSDP 的操作指南,涵蓋設定選項、程式碼錨點、常見陷阱與驗證方式。

功能
此技能提供在 Megatron-Bridge 中開啟 Megatron FSDP 的操作指南。它列出所需的設定覆寫項目,指向 Bridge 設定與執行時封裝中的相關程式碼錨點,說明常見陷阱,並提供使用現有 2 張 GPU 冒煙測試的驗證流程。它產出的是指引與設定片段,而非可執行的成品。
適用情境
適用於從 DDP 切換到以 FSDP 為基礎的資料平行時,或是在追查由 FSDP 設定變更所引起的記憶體不足或效能退步時。也適用於處理 use_megatron_fsdp、data_parallel_sharding_strategy 或分片資料平行等相關設定。
執行需求
需要存取 Megatron-Bridge 程式碼庫及其文件。驗證過程使用 pytest、torch.distributed 與 uv,並需要至少兩張 GPU。此技能未附帶指令碼,僅為說明性內容。

Megatron FSDP Skill

For stable background and recommendation level, see:

  • @docs/training/megatron-fsdp.md
  • @skills/nemo-mbridge-perf-megatron-fsdp/card.yaml

Enablement

Minimal Megatron FSDP override in Bridge:

python
cfg.dist.use_megatron_fsdp = Truecfg.ddp.use_megatron_fsdp = Truecfg.ddp.data_parallel_sharding_strategy = "optim_grads_params"cfg.ddp.average_in_collective = Falsecfg.checkpoint.ckpt_format = "fsdp_dtensor"

Example recipe fixup:

python
cfg = llama3_8b_pretrain_config()cfg.dist.use_megatron_fsdp = Truecfg.ddp.use_megatron_fsdp = Truecfg.ddp.data_parallel_sharding_strategy = "optim_grads_params"cfg.ddp.average_in_collective = Falsecfg.checkpoint.ckpt_format = "fsdp_dtensor"cfg.checkpoint.save = "/tmp/fsdp_ckpts"cfg.checkpoint.load = None

Performance harness note:

bash
python scripts/performance/launch.py --use_megatron_fsdp true

Code Anchors

Bridge config definition:

148:154:src/megatron/bridge/training/config.py
use_megatron_fsdp: bool = False"""Use Megatron's Fully Sharded Data Parallel. Cannot be used together with use_torch_fsdp2."""
use_torch_fsdp2: bool = False"""Use the torch FSDP2 implementation. FSDP2 is not currently working with Pipeline Parallel.It is still not in a stable release stage, and may therefore contain bugs or otherpotential issues."""

Bridge validation:

1533:1578:src/megatron/bridge/training/config.py
if self.dist.use_megatron_fsdp and self.dist.use_torch_fsdp2:    raise ValueError(...)...assert not self.dist.use_tp_pp_dp_mapping, "use_tp_pp_dp_mapping is not supported with Megatron FSDP"...assert self.checkpoint.ckpt_format == "fsdp_dtensor", (    "Megatron FSDP only supports fsdp_dtensor checkpoint format")

Runtime wrapper selection:

217:243:src/megatron/bridge/models/common/unimodal.py
if use_megatron_fsdp:    DP = FullyShardedDataParallelelif use_torch_fsdp2:    DP = TorchFullyShardedDataParallelelse:    DP = DistributedDataParallel...DP(    config=get_model_config(model_chunk),    ddp_config=ddp_config,    module=model_chunk,    ...    pg_collection=pg_collection,)

Perf harness overrides:

74:98:scripts/performance/utils/overrides.py
recipe.ddp.use_megatron_fsdp = Truerecipe.ddp.data_parallel_sharding_strategy = "optim_grads_params"recipe.ddp.keep_fp8_transpose_cache = Falserecipe.ddp.average_in_collective = False...recipe.checkpoint.load = None

Pitfalls

  1. Public recipes often expose use_megatron_fsdp but still default to ckpt_format="torch_dist". If save/load is enabled, switch to fsdp_dtensor.
  2. use_torch_fsdp2 exists, but on the validated branch Bridge still fails before training because _ddp_wrap passes pg_collection.
  3. CPU offloading is only valid when pipeline_model_parallel_size == 1 and activation recomputation is disabled.
  4. Upstream warns that FSDP and TP/CP can want different CUDA_DEVICE_MAX_CONNECTIONS settings on Hopper and earlier.
  5. Megatron FSDP and FSDP2 are mutually exclusive.

Verification

Use the existing 2-GPU functional smoke test:

bash
CUDA_VISIBLE_DEVICES=0,1 uv run python -m torch.distributed.run --nproc_per_node=2 \  -m pytest tests/functional_tests/training/test_megatron_fsdp.py::TestMegatronFSDP::test_fsdp_pretrain_basic -v -s

Success criteria:

  • Pytest reports 1 passed
  • The log shows finite loss at the last iteration
  • The run finishes without a checkpoint format assertion

來源與署名

來源:nvidia/skills位於skills/nemo-mbridge-perf-megatron-fsdp提交cf5224d

授權條款: Apache-2.0

內容歸原作者所有。SourceWeft 從公開儲存庫中收錄這些內容。

檢舉或申請下架