← ClaudeAtlas

nemo-mbridge-perf-moe-optimization-workflowlisted

Systematic workflow for MoE training optimization in Megatron Bridge, based on the Megatron-Core MoE paper. Covers the Three Walls framework, parallel folding, recompute strategy, dispatcher choice, and CUDA-graph bring-up.
yangwhale/CloseCrab · ★ 4 · AI & Automation · score 80
Install: claude install-skill yangwhale/CloseCrab
# MoE Training Optimization Workflow Stable docs: @docs/training/moe-optimization.md Card: @skills/nemo-mbridge-perf-moe-optimization-workflow/card.yaml Source: [Scalable Training of MoE Models with Megatron Core](https://arxiv.org/abs/2603.07685) ## Quick Reference Think in terms of the paper's Three Walls: - memory wall - communication wall - compute and host-overhead wall MoE tuning is iterative. Fixing one wall usually exposes the next one, so the best workflow is: fit first, scale second, profile third, then retune. ## First Answer Checklist For MoE optimization workflow prompts, present the response in this order: 1. **Fit**: make the model memory-feasible first. Use the smallest model parallelism that fits, prefer selective recompute before full recompute, add offloading only after recompute and parallelism are insufficient, and use `--fake-init-process-group` to sanity-check large layouts. 2. **Scale**: maximize DP after the model fits, keep hot communication inside the fastest interconnect, use PP plus VPP for multi-node scaling, prefer EP over extra TP for expert layers, and add CP when long context makes attention memory dominant. 3. **Profile**: identify the dominant wall: memory, communication, host overhead, or compute. 4. **Retune**: change dispatcher, overlap, FP8 mode, CUDA graphs, or recompute based on the profiled bottleneck. 5. Include the exact Parallel Folding meshes: `Attention: TP x CP x DP x PP` and `MoE: ETP x EP x