Nemo Mbridge Perf Tp Dp Comm Overlap

nvidia/skills/skills/nemo-mbridge-perf-tp-dp-comm-overlap

作者 nvidiacf5224d14250Apache-2.03.5K 个星标收录于 2026年10月8日更新于 2026年10月8日仓库今天更新

Operational guide for enabling TP, DP, and PP communication overlap in Megatron-Bridge, including config knobs, code anchors, pitfalls, and verification.

仅含说明DevOps & Cloud
AI 生成的概览

在 Megatron-Bridge 中启用 TP、DP 和 PP 通信重叠的操作指南,包含配置项、代码位置、注意事项和验证方法。

功能
该技能是一份操作指南,用于在 Megatron-Bridge 中开启张量并行、数据并行和流水线并行的通信重叠。它给出最小的 Bridge 配置覆盖示例、可选的 TP 预设、混合精度相关配置项,以及控制重叠行为的代码位置。它还列出常见注意事项,并说明如何使用已签入的单元测试进行验证。
适用场景
当需要在 Megatron-Bridge 中启用 TP、DP 或 PP 通信重叠,或需要将吞吐量回退追溯到通信重叠配置变更时使用。它也适用于处理 overlap_param_gather、overlap_grad_reduce、序列并行重叠或 TP/DP 重叠等设置。
运行要求
不附带脚本,仅为说明性内容。需要具备 Megatron-Bridge 环境及所引用的源文件,可选依赖 Transformer Engine 和 nemo_run,并能够通过 uv 运行 pytest。

TP / DP / PP Communication Overlap Skill

For stable background and recommendation level, see:

  • @docs/training/communication-overlap.md

Enablement

Minimal Bridge override:

python
from megatron.bridge.training.comm_overlap import CommOverlapConfig
cfg.model.tensor_model_parallel_size = 4cfg.model.sequence_parallel = Truecfg.model.pipeline_model_parallel_size = 4cfg.model.virtual_pipeline_model_parallel_size = 2
cfg.comm_overlap = CommOverlapConfig(    tp_comm_overlap=True,)
cfg.ddp.use_distributed_optimizer = Truecfg.ddp.overlap_grad_reduce = Truecfg.ddp.overlap_param_gather = True

Optional TP preset:

python
from megatron.bridge.training.comm_overlap import userbuffers_bf16_h100_h12288_tp4_mbs1_seqlen2048
cfg.comm_overlap.tp_comm_overlap_cfg = userbuffers_bf16_h100_h12288_tp4_mbs1_seqlen2048

Precision knobs belong to mixed precision:

python
cfg.mixed_precision.grad_reduce_in_fp32 = Falsecfg.mixed_precision.fp8_param_gather = False

Code Anchors

Bridge overlap gating:

439:449:src/megatron/bridge/training/comm_overlap.py
if self.user_comm_overlap_cfg.tp_comm_overlap is True:    if model_cfg.tensor_model_parallel_size < 2:        ...    elif not model_cfg.sequence_parallel:        ...    elif not HAVE_TE:        ...

PP overlap selection:

451:458:src/megatron/bridge/training/comm_overlap.py
if model_cfg.pipeline_model_parallel_size > 1:    if vp_size > 1:        comm_overlap_cfg.overlap_p2p_comm = True        comm_overlap_cfg.batch_p2p_comm = False    else:        comm_overlap_cfg.overlap_p2p_comm = False        comm_overlap_cfg.batch_p2p_comm = True

DP overlap defaults:

572:579:src/megatron/bridge/training/comm_overlap.py
if self.data_parallel_size > 1:    comm_overlap_cfg.bucket_size = 128 * 1024 * 1024    comm_overlap_cfg.overlap_grad_reduce = True    comm_overlap_cfg.overlap_param_gather = True

Launch-time env tuning:

570:609:src/megatron/bridge/recipes/run_plugins.py
executor.env_vars["CUDA_DEVICE_MAX_CONNECTIONS"] = str(cuda_device_max_connections)...executor.env_vars["NVTE_FWD_LAYERNORM_SM_MARGIN"] = str(self.layernorm_sm_margin)executor.env_vars["NVTE_BWD_LAYERNORM_SM_MARGIN"] = str(self.layernorm_sm_margin)

Pitfalls

  1. TP overlap silently disables itself if sequence_parallel=False or Transformer Engine is unavailable.
  2. PP overlap is not enabled for all PP cases. Bridge only auto-selects overlap_p2p_comm=True when PP > 1 and VPP > 1.
  3. bucket_size is a parameter-count knob, not a byte-size knob.
  4. grad_reduce_in_fp32 and fp8_param_gather should be set through mixed precision, not as standalone DDP tuning first.
  5. CUDA_DEVICE_MAX_CONNECTIONS and LayerNorm SM margin are launch-time plugin settings, not CommOverlapConfig fields.

Verification

Use the checked-in overlap unit coverage first:

bash
uv run python -m pytest tests/unit_tests/training/test_comm_overlap.py -q

Optional second check if nemo_run is available:

bash
uv run python -m pytest tests/unit_tests/recipes/test_run_plugins.py -q

Success criteria:

  • first command reports 26 passed
  • second command validates plugin-owned env wiring when not skipped

来源与署名

来源:nvidia/skills位于skills/nemo-mbridge-perf-tp-dp-comm-overlap提交cf5224d

许可证: Apache-2.0

内容归原作者所有。SourceWeft 从公开仓库中收录这些内容。

举报或申请下架