nemo-mbridge-perf-memory-tuninglisted
Install: claude install-skill yangwhale/CloseCrab
# Memory Tuning
Stable docs: @docs/parallelisms.md
Card: @skills/nemo-mbridge-perf-memory-tuning/card.yaml
## What It Is
GPU OOM failures during training often stem from memory **fragmentation** rather
than raw capacity. PyTorch's default CUDA allocator can leave unusable gaps
between allocations. The single most effective fix is:
```bash
export PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True
```
This tells PyTorch to use expandable (non-fixed-size) memory segments, which
dramatically reduces fragmentation and often eliminates borderline OOM without
any model or parallelism changes.
Beyond fragmentation, actual peak memory is determined by:
- **Parameter + optimizer state memory** — controlled by TP, PP, DP sharding
(distributed optimizer, FSDP)
- **Activation memory** — controlled by activation recompute, sequence length,
micro-batch size
- **Temporary / workspace memory** — CUDA kernels, NCCL buffers, CUDA graphs
For configuration planning, use the Bridge theoretical estimator before launching
large jobs:
```python
from megatron.bridge.training.utils.theoretical_memory_utils import estimate_training_memory
estimate = estimate_training_memory(cfg, num_microbatches=num_microbatches)
```
The estimator reports the most-loaded GPU shard and separates dense/embedding,
routed MoE expert, and activation components. It does not include allocator
fragmentation, CUDA/NCCL workspace, CUDA graph buffers, token imbalance, or
dispatcher workspace, so validate final confi