MoE Expert-Parallel Overlap Skill
References
- Stable docs: @docs/training/communication-overlap.md
- Structured metadata: @skills/nemo-mbridge-perf-expert-parallel-overlap/card.yaml
What It Is
Expert-parallel (EP) overlap hides the cost of token dispatch/combine all-to-all
communication by running it concurrently with expert FFN compute. Optionally,
delayed expert weight-gradient computation (delay_wgrad_compute) provides
additional overlap by deferring wgrad to overlap with the next layer's forward.
Bridge supports two dispatcher paths:
Quick Decision
Use EP overlap when:
- the model is MoE with
EP > 1 - expert dispatch/combine communication is a meaningful part of step time
- you have memory headroom and are tuning for throughput
Prefer:
alltoalldispatcher for the first rollout (broader compatibility)flex+ DeepEP/HybridEP when running on supported GPUs and seeking additional gains
Avoid EP overlap when:
- full activation recompute is enabled
moe_shared_expert_overlapis enabled- the run is still being brought up for correctness
- PyTorch < 2.6.0
Expected outcome:
- if all-to-all dispatch is a clear profile bottleneck, overlap can produce a modest to meaningful speedup
- if the run is tiny, communication-light, or dominated by another wall, the gain may be negligible
Correctness-First alltoall Benchmark
For the plain EP-overlap isolation benchmark, keep flex dispatch and delayed
wgrad disabled. The measured shape was Qwen3 MoE 30B-A3B SFT on 16 H100 GPUs:
EP=16, alltoall, BF16, global batch size 1024, CUDA graphs disabled,
moe_permute_fusion=false, measured over iterations 3-8.
Use these overrides for the plain-overlap case:
Do not use --moe_a2a_overlap true for this isolation test: the performance
harness helper enables both overlap_moe_expert_parallel_comm and
delay_wgrad_compute, so it does not isolate plain EP overlap.
Steady-window timing from that benchmark:
This is evidence for enabling plain EP overlap on this inter-node all-to-all shape. It does not show a meaningful independent win from delayed wgrad, and it does not validate fused MoE permutation because that path was disabled for the runtime stack.
HybridEP Production-Shape Benchmark
A 2026-07-25 controlled Qwen3 30B-A3B pretraining comparison validated plain EP overlap with the production HybridEP path:
The independent run reduced step time by 15.059% and raised throughput by 17.729% over the reproduced baseline. Loss was finite, skipped and NaN iterations remained zero, and rank-0 peak allocated memory was 62.166 GiB.
A matched Nsight Systems comparison captured the same 463,348 rank-0 kernels per case. Enabling overlap increased communication concurrent with GEMM and attention from 9.079ms (0.11% of communication time) to 3,958.997ms (36.55%). GPU-active interval union fell from 22.821s to 21.221s.
Use this as evidence for the mechanism, not as a universal speedup promise. The dispatcher, graph scopes, routing, parallelism, batch shape, and runtime were held fixed while only plain EP overlap changed.
Enablement
alltoall dispatcher
Enable delay_wgrad_compute=True only after the plain overlap path is known to
work and its extra compatibility constraints have been checked.
flex dispatcher (DeepEP or HybridEP)
Benchmark plain EP overlap first. Enable delay_wgrad_compute=True only as a
separate follow-up A/B after its CUDA-graph and TE compatibility constraints
are satisfied.
Compatibility And Constraints
expert_model_parallel_size > 1num_moe_experts > 1moe_token_dispatcher_typemust be"alltoall"or"flex"moe_shared_expert_overlap = False- Base precision is BF16 or FP16
- PyTorch
>= 2.6.0 - If
PP > 1,virtual_pipeline_model_parallel_sizemust be set recompute_granularity != "full",recompute_method = None,recompute_num_layers = Nonemtp_num_layersmust beNoneor1delay_wgrad_computerequiresoverlap_moe_expert_parallel_commas a prerequisitedelay_wgrad_computewithoverlap_grad_reducerequires TE >= 2.7.0delay_wgrad_computewithgradient_accumulation_fusionrequires TE >= 2.7.0- CUDA graph
attnscope +delay_wgrad_computerequires TE >= 2.12.0,gradient_accumulation_fusion = True, and no attention bias - DeepEP: Ampere, Hopper, B200, B300 GPUs only
- HybridEP: Ampere, Hopper, B200, B300, GB200/GB300 with NVL72
Minimal Working Config
Use this as the correctness-first starting point. Add delayed wgrad, flex dispatch, and CUDA-graph interactions only after the plain overlap path is known to work.
Minimal Runnable Command
Performance harness example inside a Slurm allocation. Keep the model, parallelism, dispatcher, and runtime fixed, and vary only the two overlap overrides:
Do not use --moe_a2a_overlap true when separating plain EP overlap from
delayed wgrad: the performance harness helper enables both
overlap_moe_expert_parallel_comm and delay_wgrad_compute.
Unit test verification:
Verification
Unit tests
Log checks
After a successful run with EP overlap:
- Confirm no assertion errors during
CommOverlapConfigfinalization - Confirm
overlap_moe_expert_parallel_commappears asTruein the logged config - If using flex dispatcher, confirm
moe_token_dispatcher_type = "flex"and the correct backend in logs
Success criteria
- Config validation passes for the selected dispatcher and overlap settings
- Training runs complete without hangs or assertion failures
- Throughput improves or at least does not regress for the target workload
- Loss trajectory matches baseline (overlap should not affect convergence)
Profile interpretation
Use an unprofiled steady window for the throughput acceptance result. Use a matched profile to explain the mechanism:
- Keep the dispatcher, routing, graph scopes, batch shape, parallel layout, and runtime fixed.
- Capture the same rank and steady iteration while toggling only plain EP overlap.
- Build interval unions for communication and compute kernels, then measure their intersection.
- Do not use summed kernel duration as wall time. Concurrent kernels can run longer under SM or bandwidth contention even when exposed time decreases.
- Corroborate interval results with dispatch/combine NVTX ranges, final step time, loss finiteness, skipped/NaN counts, and peak memory.
Code Anchors
Bridge overlap validation
Delayed wgrad validation
Flex-dispatcher activation
Perf harness override
Tests
Failure Diagnosis
Known Limitations
- Setting
moe_flex_dispatcher_backendalone does not activate flex dispatch — you must callapply_flex_dispatcher_backend(...). - Public recipes are often conservative and leave MoE overlap disabled by default.
- Controlled end-to-end and profile evidence exists for one Qwen3 30B-A3B HybridEP H100 shape; repeat the matched A/B before generalizing it to another model, dispatcher, topology, precision, or batch shape.
- MoE overlap and shared-expert overlap are mutually exclusive.
- CUDA graph plus delayed wgrad is a multi-constraint path that requires careful TE version and scope validation.
Last signature refresh: 2026-08-03.


