Nemo Mbridge Perf Tp Dp Comm Overlap

nvidia/skills/skills/nemo-mbridge-perf-tp-dp-comm-overlap

by nvidiacf5224d14250Apache-2.03.5K starsListed Oct 8, 2026Updated Oct 8, 2026Repository updated today

Operational guide for enabling TP, DP, and PP communication overlap in Megatron-Bridge, including config knobs, code anchors, pitfalls, and verification.

Instructions onlyDevOps & Cloud
AI-generated overview

Guide for enabling TP, DP, and PP communication overlap in Megatron-Bridge, with config knobs, code anchors, pitfalls, and verification.

What it does
This skill is an operational guide for turning on tensor-parallel, data-parallel, and pipeline-parallel communication overlap in Megatron-Bridge. It shows the minimal Bridge configuration override, optional TP presets, mixed-precision knobs, and the code locations that gate overlap behavior. It also lists common pitfalls and verification steps using checked-in unit tests.
When to use it
Use it when enabling TP, DP, or PP communication overlap in Megatron-Bridge, or when tracing a throughput regression back to a communication overlap configuration change. It is also relevant when working with settings such as overlap_param_gather, overlap_grad_reduce, sequence-parallel overlap, or TP/DP overlap.
Requirements
No scripts are shipped; it is instructions only. It assumes a Megatron-Bridge environment with the referenced source files, optional Transformer Engine and nemo_run availability, and the ability to run pytest via uv.

TP / DP / PP Communication Overlap Skill

For stable background and recommendation level, see:

  • @docs/training/communication-overlap.md

Enablement

Minimal Bridge override:

python
from megatron.bridge.training.comm_overlap import CommOverlapConfig
cfg.model.tensor_model_parallel_size = 4cfg.model.sequence_parallel = Truecfg.model.pipeline_model_parallel_size = 4cfg.model.virtual_pipeline_model_parallel_size = 2
cfg.comm_overlap = CommOverlapConfig(    tp_comm_overlap=True,)
cfg.ddp.use_distributed_optimizer = Truecfg.ddp.overlap_grad_reduce = Truecfg.ddp.overlap_param_gather = True

Optional TP preset:

python
from megatron.bridge.training.comm_overlap import userbuffers_bf16_h100_h12288_tp4_mbs1_seqlen2048
cfg.comm_overlap.tp_comm_overlap_cfg = userbuffers_bf16_h100_h12288_tp4_mbs1_seqlen2048

Precision knobs belong to mixed precision:

python
cfg.mixed_precision.grad_reduce_in_fp32 = Falsecfg.mixed_precision.fp8_param_gather = False

Code Anchors

Bridge overlap gating:

439:449:src/megatron/bridge/training/comm_overlap.py
if self.user_comm_overlap_cfg.tp_comm_overlap is True:    if model_cfg.tensor_model_parallel_size < 2:        ...    elif not model_cfg.sequence_parallel:        ...    elif not HAVE_TE:        ...

PP overlap selection:

451:458:src/megatron/bridge/training/comm_overlap.py
if model_cfg.pipeline_model_parallel_size > 1:    if vp_size > 1:        comm_overlap_cfg.overlap_p2p_comm = True        comm_overlap_cfg.batch_p2p_comm = False    else:        comm_overlap_cfg.overlap_p2p_comm = False        comm_overlap_cfg.batch_p2p_comm = True

DP overlap defaults:

572:579:src/megatron/bridge/training/comm_overlap.py
if self.data_parallel_size > 1:    comm_overlap_cfg.bucket_size = 128 * 1024 * 1024    comm_overlap_cfg.overlap_grad_reduce = True    comm_overlap_cfg.overlap_param_gather = True

Launch-time env tuning:

570:609:src/megatron/bridge/recipes/run_plugins.py
executor.env_vars["CUDA_DEVICE_MAX_CONNECTIONS"] = str(cuda_device_max_connections)...executor.env_vars["NVTE_FWD_LAYERNORM_SM_MARGIN"] = str(self.layernorm_sm_margin)executor.env_vars["NVTE_BWD_LAYERNORM_SM_MARGIN"] = str(self.layernorm_sm_margin)

Pitfalls

  1. TP overlap silently disables itself if sequence_parallel=False or Transformer Engine is unavailable.
  2. PP overlap is not enabled for all PP cases. Bridge only auto-selects overlap_p2p_comm=True when PP > 1 and VPP > 1.
  3. bucket_size is a parameter-count knob, not a byte-size knob.
  4. grad_reduce_in_fp32 and fp8_param_gather should be set through mixed precision, not as standalone DDP tuning first.
  5. CUDA_DEVICE_MAX_CONNECTIONS and LayerNorm SM margin are launch-time plugin settings, not CommOverlapConfig fields.

Verification

Use the checked-in overlap unit coverage first:

bash
uv run python -m pytest tests/unit_tests/training/test_comm_overlap.py -q

Optional second check if nemo_run is available:

bash
uv run python -m pytest tests/unit_tests/recipes/test_run_plugins.py -q

Success criteria:

  • first command reports 26 passed
  • second command validates plugin-owned env wiring when not skipped

Source and attribution

Source:nvidia/skillsinskills/nemo-mbridge-perf-tp-dp-comm-overlapat commitcf5224d

License: Apache-2.0

Content belongs to its original authors. SourceWeft indexes it from a public repository.

Report or request removal