← ClaudeAtlas

nemo-mbridge-perf-megatron-fsdplisted

Operational guide for enabling Megatron FSDP in Megatron-Bridge, including config knobs, code anchors, pitfalls, and verification.
yangwhale/CloseCrab · ★ 4 · AI & Automation · score 80
Install: claude install-skill yangwhale/CloseCrab
# Megatron FSDP Skill For stable background and recommendation level, see: - @docs/training/megatron-fsdp.md - @skills/nemo-mbridge-perf-megatron-fsdp/card.yaml ## Enablement Minimal Megatron FSDP override in Bridge: ```python cfg.dist.use_megatron_fsdp = True cfg.ddp.use_megatron_fsdp = True cfg.ddp.data_parallel_sharding_strategy = "optim_grads_params" cfg.ddp.average_in_collective = False cfg.checkpoint.ckpt_format = "fsdp_dtensor" ``` Example recipe fixup: ```python cfg = llama3_8b_pretrain_config() cfg.dist.use_megatron_fsdp = True cfg.ddp.use_megatron_fsdp = True cfg.ddp.data_parallel_sharding_strategy = "optim_grads_params" cfg.ddp.average_in_collective = False cfg.checkpoint.ckpt_format = "fsdp_dtensor" cfg.checkpoint.save = "/tmp/fsdp_ckpts" cfg.checkpoint.load = None ``` Performance harness note: ```bash python scripts/performance/launch.py --use_megatron_fsdp true ``` ## Code Anchors Bridge config definition: ```148:154:src/megatron/bridge/training/config.py use_megatron_fsdp: bool = False """Use Megatron's Fully Sharded Data Parallel. Cannot be used together with use_torch_fsdp2.""" use_torch_fsdp2: bool = False """Use the torch FSDP2 implementation. FSDP2 is not currently working with Pipeline Parallel. It is still not in a stable release stage, and may therefore contain bugs or other potential issues.""" ``` Bridge validation: ```1533:1578:src/megatron/bridge/training/config.py if self.dist.use_megatron_fsdp and self.dist.use_torch_fsdp2: rais