← ClaudeAtlas

nemo-mbridge-perf-expert-parallel-overlaplisted

Validate and use MoE expert-parallel communication overlap in Megatron-Bridge, including overlap_moe_expert_parallel_comm, delay_wgrad_compute, and flex dispatcher backends such as DeepEP and HybridEP.
yangwhale/CloseCrab · ★ 4 · AI & Automation · score 80
Install: claude install-skill yangwhale/CloseCrab
# MoE Expert-Parallel Overlap Skill ## References - Stable docs: @docs/training/communication-overlap.md - Structured metadata: @skills/nemo-mbridge-perf-expert-parallel-overlap/card.yaml ## What It Is Expert-parallel (EP) overlap hides the cost of token dispatch/combine all-to-all communication by running it concurrently with expert FFN compute. Optionally, delayed expert weight-gradient computation (`delay_wgrad_compute`) provides additional overlap by deferring wgrad to overlap with the next layer's forward. Bridge supports two dispatcher paths: | Dispatcher | Backend | When to use | |---|---|---| | `alltoall` | Standard MoE all-to-all | Default, broadest compatibility | | `flex` | DeepEP or HybridEP | Higher overlap on Ampere/Hopper/Blackwell | ## Quick Decision Use EP overlap when: - the model is MoE with `EP > 1` - expert dispatch/combine communication is a meaningful part of step time - you have memory headroom and are tuning for throughput Prefer: - `alltoall` dispatcher for the first rollout (broader compatibility) - `flex` + DeepEP/HybridEP when running on supported GPUs and seeking additional gains Avoid EP overlap when: - full activation recompute is enabled - `moe_shared_expert_overlap` is enabled - the run is still being brought up for correctness - PyTorch < 2.6.0 Expected outcome: - if all-to-all dispatch is a clear profile bottleneck, overlap can produce a modest to meaningful speedup - if the run is tiny, communication-light, or dominated