Hierarchical Context Parallel Skill
This skill covers hierarchical context parallelism: nested context-parallel process
groups used by cp_comm_type="a2a+p2p" and configured with
hierarchical_context_parallel_sizes.
For what hierarchical CP is, when to use it, and the decision tree
(a2a+p2p vs pure a2a vs p2p), see:
- @docs/training/hierarchical-context-parallel.md
- @skills/nemo-mbridge-perf-hierarchical-context-parallel/card.yaml
Enablement
Minimal Bridge override:
Required constraints:
prod(hierarchical_context_parallel_sizes) == context_parallel_sizeseq_length % (2 * context_parallel_size) == 0- Transformer Engine
>= 1.12.0
Code Anchors
Upstream config and validation:
Bridge MPU path:
Bridge decentralized-PG path:
Implementation Map
The code anchors above show the config declarations and argument validation.
Validation (MCore)
TransformerConfig.__post_init__ enforces that a2a+p2p requires HCP sizes and the product matches CP.
Process group creation
parallel_state.initialize_model_parallel creates hierarchical CP sub-groups
when HCP sizes are provided via create_hierarchical_groups. Bridge currently
gets those groups through the MPU-backed ProcessGroupCollection.
TE integration
TEDotProductAttention passes the hierarchical groups to Transformer Engine
when a2a+p2p is used. Requires Transformer Engine >= 1.12.0.
Pitfalls
- Bridge HCP is MPU-only today: If
use_decentralized_pg=True, Bridge initializes flat CP groups and leaves HCP unset. - No checked-in Bridge recipe currently exercises HCP directly.
- Single-GPU load helpers clear
hierarchical_context_parallel_sizes. - Silent broken training on old stacks: If you use
a2a+p2pwithout settinghierarchical_context_parallel_sizes, MCore now asserts. Older versions would silently disable CP communication, so each rank attended only to its local chunk and produced artificially high throughput with broken gradients. - Product must match:
prod(hierarchical_context_parallel_sizes)must exactly equalcontext_parallel_size. A mismatch triggers an assertion. - Verify in logs: Look for the process group initialization output. You should see
HIERARCHICAL_CONTEXT_PARALLEL_GROUPSbeing created. If you only seeCONTEXT_PARALLEL_GROUP, HCP is not active.
Verification
No dedicated Bridge end-to-end test exists yet for HCP (see @skills/nemo-mbridge-perf-hierarchical-context-parallel/card.yaml
follow_up_validation). Use the existing unit tests and log inspection instead.
Run the decentralized-PG unit test to confirm the flat-CP behavior is preserved:
For a manual smoke check, launch a 4-GPU run with a small recipe and
cp_comm_type=a2a+p2p plus hierarchical_context_parallel_sizes=[2,2]:
Success criteria:
- Logs show
HIERARCHICAL_CONTEXT_PARALLEL_GROUPSbeing created - Training completes at least one step without error
- If you only see
CONTEXT_PARALLEL_GROUP, HCP is not active


