Resiliency
Stable docs: @docs/training/resiliency.md, @docs/training/checkpointing.md Card: @skills/nemo-mbridge-resiliency/card.yaml
Enablement
Fault tolerance (Slurm only)
Option 1: NeMo Run plugin (recommended)
Option 2: Direct config + ft_launcher
Launch with ft_launcher (not torchrun):
Section-based timeout monitoring covers setup, training steps, checkpointing,
and out-of-section time independently. Timeouts are saved to ft_state.json
for subsequent runs when calc_ft_timeouts=True.
NVRx straggler detection
Preemption
Plugin (Slurm)
Direct config
Re-run state machine (experimental)
Exit codes: 16 = resume to disambiguate, 17 = failed validation.
In-process restart (experimental)
Required environment variables:
The PyTorch NCCL watchdog timeout must exceed hard_timeout. NeMo-Run's
Slurm Executor is not supported; launch directly with srun --kill-on-bad-exit=0.
Async checkpoint save
Local checkpointing (NVRx)
Code Anchors
Fault tolerance
- Config:
src/megatron/bridge/training/config.py—FaultToleranceConfig - Runtime:
src/megatron/bridge/training/fault_tolerance.py - Plugin:
src/megatron/bridge/recipes/run_plugins.py—FaultTolerancePlugin - Perf plugin:
scripts/performance/nemo-mbridge-resiliency_plugins.py - Tests:
tests/unit_tests/training/test_fault_tolerance.py - Example:
examples/training_features/nemo-mbridge-resiliency/fault_tolerance/
Straggler detection
- Config:
src/megatron/bridge/training/config.py—NVRxStragglerDetectionConfig - Runtime:
src/megatron/bridge/training/nvrx_straggler.py - Train loop:
src/megatron/bridge/training/train.py—check_nvrx_straggler_detection - Tests:
tests/unit_tests/training/test_nvrx_straggler.py,tests/functional_tests/training/test_nvrx_straggler.py - Example:
examples/training_features/nemo-mbridge-resiliency/straggler_detection/
In-process restart
- Config:
src/megatron/bridge/training/config.py—InProcessRestartConfig - Runtime:
src/megatron/bridge/training/inprocess_restart.py - Entry point:
src/megatron/bridge/training/pretrain.py—maybe_wrap_for_inprocess_restart - Tests:
tests/unit_tests/training/test_inprocess_restart.py,tests/functional_tests/training/test_inprocess_restart.py
Preemption
- Plugin:
src/megatron/bridge/recipes/run_plugins.py—PreemptionPlugin - Signal handler:
src/megatron/bridge/training/utils/sig_utils.py - Tests:
tests/unit_tests/recipes/test_run_plugins.py
Re-run state machine
- Config:
src/megatron/bridge/training/config.py—RerunStateMachineConfig - Init:
src/megatron/bridge/training/initialize.py—init_rerun_state
Checkpointing
- Async save:
src/megatron/bridge/training/checkpointing.py—schedule_async_save - Local ckpt:
src/megatron/bridge/training/checkpointing.py—LocalCheckpointManager - Tests:
tests/functional_tests/training/test_local_checkpointing.py
Pitfalls
-
ft_launcher, not torchrun: Direct
FaultToleranceConfigrequiresft_launcher. Usingtorchrunsilently disables FT. For non-Slurm, setGROUP_RANK=0. -
Async save requires torch_dist:
async_save=Trueonly works withckpt_format="torch_dist". Other formats silently fail or error. -
IPR + NeMo-Run: In-process restart is not compatible with NeMo-Run or Slurm preemption plugins. Requires specific PyTorch/NCCL versions and env vars.
-
NVRx vs legacy straggler: Two detectors exist. Use NVRx (
nvrx_straggler); do not enable both. -
stop_if_detected default: NVRx logs but does not stop training by default. Set
stop_if_detected=Truefor automatic termination. -
NCCL watchdog vs hard_timeout: For IPR, NCCL watchdog timeout must exceed
hard_timeoutor PyTorch kills the process before recovery. -
Rerun state machine is alpha: Use
check_for_nan_in_loss=Truefor NaN detection, but don't rely on full rerun workflows yet.
Verification
Fault tolerance
Look for [FaultTolerance] / [RankMonitorServer] log lines with section
timeouts. Simulated fault should trigger restart from checkpoint.
Straggler detection
Look for GPU relative performance and GPU individual performance reports
with per-rank scores.
Async checkpoint
Look for Scheduling async checkpoint save in logs. Training iterations
should continue while checkpoint files are being written.
In-process restart
Requires compatible PyTorch/NCCL versions.


