MLM vs Bridge Training
For how they differ, the arg mapping tables, gotchas, and translation script, see:
- @docs/megatron-lm-to-megatron-bridge.md
First Answer Checklist
For MLM-vs-Bridge correlation questions, always name these items up front:
- Bridge recipe:
vanilla_gpt_pretrain_config. - Bridge entry point:
scripts/training/run_recipe.py. - MLM entry point:
3rdparty/Megatron-LM/pretrain_gpt.py. - Launch wrapper for both:
uv run python -m torch.distributed.run. - Fresh-run cleanup:
rm -rf nemo_experimentsbefore the Bridge run.
Also state that MLM needs
PYTHONPATH=3rdparty/Megatron-LM:$PYTHONPATH, matched Bridge and MLM losses
should agree within BF16 rounding, and files under 3rdparty/Megatron-LM/
should not be modified from this repo.
Correlation Testing
Use vanilla_gpt_pretrain_config for loss-correlation testing. This recipe uses
bare GPTModelProvider defaults (LayerNorm, GeLU, learned_absolute position
embeddings, vocab_size inherited from tokenizer) — matching MLM
pretrain_gpt.py defaults with no args.
MLM Correlation Run (2L/256H, 1 GPU)
Bridge Correlation Run (same config, 1 GPU)
Verification
With matched parameters the LM losses should be nearly identical at each
iteration. Compare lm loss values from both logs — they should agree to
within BF16 rounding.
Multi-GPU Examples
MLM 2-GPU with TP=2
Bridge 2-GPU with TP=2
Available Recipes
Common recipes (use with --recipe):
vanilla_gpt_pretrain_config— Minimal GPT (bare GPTModelProvider defaults, ideal for correlation testing and custom configs)llama32_1b_pretrain_config— Llama 3.2 1B (16L, 2048H, GBS=512, seq=8192)llama3_8b_pretrain_config— Llama 3 8Bqwen3_8b_pretrain_config— Qwen3 8Bdeepseek_v2_lite_pretrain_config— DeepSeek-V2-Lite 16B MoE
SFT/PEFT variants use _sft_config / _peft_config suffix.
Megatron-Core Submodule
For what the submodule is and why two versions exist, see @docs/megatron-lm-to-megatron-bridge.md.
Check current version
Switch to dev for testing newer MCore features
Switch back to main
After pulling latest main
When you pull the latest Bridge main branch, the submodule pointer may have been updated. Re-sync the submodule:
Pitfalls
-
Always
rm -rf nemo_experimentsbefore a fresh correlation run. Bridge auto-resumes from stale checkpoints silently. -
uv runrequired: Always useuv run python -m torch.distributed.run(not baretorchrunorpython). -
MLM PYTHONPATH: Must include
3rdparty/Megatron-LMsogpt_builders.pyis importable. -
Scheduler overrides: When overriding
train.train_itersto a small value, also setscheduler.lr_warmup_itersandscheduler.lr_decay_itersor you get an assertion error. -
Use
dataset.seq_lengthin CLI overrides for both pretraining and fine-tuning datasets. -
MoE OOM: Large MoE models require full activation recomputation and typically multi-node EP. TP does NOT reduce per-GPU expert memory.
-
uv sync --lockedfails after switching to dev: The lockfile is generated against the main MCore commit. Useuv sync(without--locked) when on dev.
