Nemo Mbridge Perf Tp Dp Comm Overlap

nvidia/skills/skills/nemo-mbridge-perf-tp-dp-comm-overlap

作者 nvidiacf5224d14250Apache-2.03.5K 個星標收錄於 2026年10月8日更新於 2026年10月8日儲存庫今天更新

Operational guide for enabling TP, DP, and PP communication overlap in Megatron-Bridge, including config knobs, code anchors, pitfalls, and verification.

僅含說明DevOps & Cloud
AI 產生的概覽

在 Megatron-Bridge 中啟用 TP、DP 與 PP 通訊重疊的操作指南,包含設定選項、程式碼位置、注意事項與驗證方式。

功能
此技能是一份操作指南,用於在 Megatron-Bridge 中開啟張量平行、資料平行與管線平行的通訊重疊。它提供最小的 Bridge 設定覆寫範例、可選的 TP 預設集、混合精度相關設定,以及控制重疊行為的程式碼位置。它也列出常見注意事項,並說明如何使用已簽入的單元測試進行驗證。
適用情境
當需要在 Megatron-Bridge 中啟用 TP、DP 或 PP 通訊重疊,或需要將輸送量回退追溯到通訊重疊設定變更時使用。它也適用於處理 overlap_param_gather、overlap_grad_reduce、序列平行重疊或 TP/DP 重疊等設定。
執行需求
不附帶指令碼,僅為說明性內容。需要具備 Megatron-Bridge 環境及所引用的原始碼檔案,可選依賴 Transformer Engine 與 nemo_run,並能夠透過 uv 執行 pytest。

TP / DP / PP Communication Overlap Skill

For stable background and recommendation level, see:

  • @docs/training/communication-overlap.md

Enablement

Minimal Bridge override:

python
from megatron.bridge.training.comm_overlap import CommOverlapConfig
cfg.model.tensor_model_parallel_size = 4cfg.model.sequence_parallel = Truecfg.model.pipeline_model_parallel_size = 4cfg.model.virtual_pipeline_model_parallel_size = 2
cfg.comm_overlap = CommOverlapConfig(    tp_comm_overlap=True,)
cfg.ddp.use_distributed_optimizer = Truecfg.ddp.overlap_grad_reduce = Truecfg.ddp.overlap_param_gather = True

Optional TP preset:

python
from megatron.bridge.training.comm_overlap import userbuffers_bf16_h100_h12288_tp4_mbs1_seqlen2048
cfg.comm_overlap.tp_comm_overlap_cfg = userbuffers_bf16_h100_h12288_tp4_mbs1_seqlen2048

Precision knobs belong to mixed precision:

python
cfg.mixed_precision.grad_reduce_in_fp32 = Falsecfg.mixed_precision.fp8_param_gather = False

Code Anchors

Bridge overlap gating:

439:449:src/megatron/bridge/training/comm_overlap.py
if self.user_comm_overlap_cfg.tp_comm_overlap is True:    if model_cfg.tensor_model_parallel_size < 2:        ...    elif not model_cfg.sequence_parallel:        ...    elif not HAVE_TE:        ...

PP overlap selection:

451:458:src/megatron/bridge/training/comm_overlap.py
if model_cfg.pipeline_model_parallel_size > 1:    if vp_size > 1:        comm_overlap_cfg.overlap_p2p_comm = True        comm_overlap_cfg.batch_p2p_comm = False    else:        comm_overlap_cfg.overlap_p2p_comm = False        comm_overlap_cfg.batch_p2p_comm = True

DP overlap defaults:

572:579:src/megatron/bridge/training/comm_overlap.py
if self.data_parallel_size > 1:    comm_overlap_cfg.bucket_size = 128 * 1024 * 1024    comm_overlap_cfg.overlap_grad_reduce = True    comm_overlap_cfg.overlap_param_gather = True

Launch-time env tuning:

570:609:src/megatron/bridge/recipes/run_plugins.py
executor.env_vars["CUDA_DEVICE_MAX_CONNECTIONS"] = str(cuda_device_max_connections)...executor.env_vars["NVTE_FWD_LAYERNORM_SM_MARGIN"] = str(self.layernorm_sm_margin)executor.env_vars["NVTE_BWD_LAYERNORM_SM_MARGIN"] = str(self.layernorm_sm_margin)

Pitfalls

  1. TP overlap silently disables itself if sequence_parallel=False or Transformer Engine is unavailable.
  2. PP overlap is not enabled for all PP cases. Bridge only auto-selects overlap_p2p_comm=True when PP > 1 and VPP > 1.
  3. bucket_size is a parameter-count knob, not a byte-size knob.
  4. grad_reduce_in_fp32 and fp8_param_gather should be set through mixed precision, not as standalone DDP tuning first.
  5. CUDA_DEVICE_MAX_CONNECTIONS and LayerNorm SM margin are launch-time plugin settings, not CommOverlapConfig fields.

Verification

Use the checked-in overlap unit coverage first:

bash
uv run python -m pytest tests/unit_tests/training/test_comm_overlap.py -q

Optional second check if nemo_run is available:

bash
uv run python -m pytest tests/unit_tests/recipes/test_run_plugins.py -q

Success criteria:

  • first command reports 26 passed
  • second command validates plugin-owned env wiring when not skipped

來源與署名

來源:nvidia/skills位於skills/nemo-mbridge-perf-tp-dp-comm-overlap提交cf5224d

授權條款: Apache-2.0

內容歸原作者所有。SourceWeft 從公開儲存庫中收錄這些內容。

檢舉或申請下架