CPU Offloading
References
- Stable docs: @docs/training/cpu-offloading.md
- Structured metadata: @skills/nemo-mbridge-perf-cpu-offloading/card.yaml
What It Is
Two independent mechanisms to move data from GPU to CPU memory:
Quick Decision
Enablement
Optimizer CPU offloading (recommended for large models)
CLI overrides:
Activation CPU offloading (small/medium models only)
Config Parameter Reference
Optimizer offloading
Activation offloading
Compatibility And Constraints
Activation offloading
pipeline_model_parallel_sizemust be 1recompute_granularitymust beNone- Cannot combine with
fine_grained_activation_offloading - Cannot combine with CUDA graphs
cpu_offloading_num_layersmust be in[0, num_layers-1)
Optimizer offloading
- Requires
use_distributed_optimizer = True(default in most recipes) - No PP, recompute, or CUDA graph restrictions
optimizer_offload_fractionmust be in[0.0, 1.0]
Practical: large MoE models
Activation offloading is blocked for Qwen3-30B-A3B and similar large MoE models. The PP=1 constraint means each GPU holds all 48 layers; model weights + optimizer states alone (~70 GB) exceed H100 80 GB capacity.
Minimal Runnable Command
Verification
Unit tests
Success criteria
- Config validation passes for the selected offloading mode
- Training completes without OOM or NCCL errors
- Loss matches the non-offloaded baseline (max delta < 0.001)
- Memory usage drops proportionally to offload fraction
Code Anchors
MCore activation offload constraints
MCore CUDA graph incompatibility
MCore fine-grained offloading mutual exclusion
MCore HybridDeviceOptimizer instantiation
Bridge CUDA graph guard
Bridge activation offloading in PEFT
Failure Diagnosis
Known Limitations
- Activation offloading requires PP=1, making it impractical for large models (30B+ MoE) that need pipeline parallelism.
- Optimizer offloading throughput penalty scales linearly (~1.9x at 25%, ~4.2x at 100% for Qwen3-30B-A3B).
- D2H/H2D overlap provides only ~7% speedup because CPU Adam compute is the dominant bottleneck.
fine_grained_activation_offloadingis a separate module-level approach that works with PP > 1 but cannot be combined with layer-levelcpu_offloading.


