← ClaudeAtlas

nemo-mbridge-perf-memory-tuninglisted

Techniques for reducing peak GPU memory in Megatron Bridge — expandable segments, parallelism resizing, activation recompute, CPU offloading constraints, and common OOM fixes.
yangwhale/CloseCrab · ★ 4 · AI & Automation · score 80
Install: claude install-skill yangwhale/CloseCrab
# Memory Tuning Stable docs: @docs/parallelisms.md Card: @skills/nemo-mbridge-perf-memory-tuning/card.yaml ## What It Is GPU OOM failures during training often stem from memory **fragmentation** rather than raw capacity. PyTorch's default CUDA allocator can leave unusable gaps between allocations. The single most effective fix is: ```bash export PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True ``` This tells PyTorch to use expandable (non-fixed-size) memory segments, which dramatically reduces fragmentation and often eliminates borderline OOM without any model or parallelism changes. Beyond fragmentation, actual peak memory is determined by: - **Parameter + optimizer state memory** — controlled by TP, PP, DP sharding (distributed optimizer, FSDP) - **Activation memory** — controlled by activation recompute, sequence length, micro-batch size - **Temporary / workspace memory** — CUDA kernels, NCCL buffers, CUDA graphs For configuration planning, use the Bridge theoretical estimator before launching large jobs: ```python from megatron.bridge.training.utils.theoretical_memory_utils import estimate_training_memory estimate = estimate_training_memory(cfg, num_microbatches=num_microbatches) ``` The estimator reports the most-loaded GPU shard and separates dense/embedding, routed MoE expert, and activation components. It does not include allocator fragmentation, CUDA/NCCL workspace, CUDA graph buffers, token imbalance, or dispatcher workspace, so validate final confi