← ClaudeAtlas

nemo-mbridge-perf-moe-comm-overlaplisted

MoE expert-parallel communication overlap in Megatron Bridge. Covers dispatch/combine overlap, flex dispatcher backends, and expert wgrad scheduling.
yangwhale/CloseCrab · ★ 4 · AI & Automation · score 80
Install: claude install-skill yangwhale/CloseCrab
# MoE Communication Overlap For the higher-level overview, see: - @docs/training/communication-overlap.md - @skills/nemo-mbridge-perf-moe-comm-overlap/card.yaml ## Quick Decision Use MoE communication overlap when: - `EP > 1` - token dispatch or combine time is visible in the profile - the run is already correct and you are now tuning throughput Avoid turning it on as an early bring-up step. It is easier to validate after the dispatcher, routing mode, and recompute plan are already stable. ## Enablement ```python cfg.comm_overlap.overlap_moe_expert_parallel_comm = True # Optional: delayed wgrad for additional overlap cfg.comm_overlap.delay_wgrad_compute = True # IMPORTANT: disable shared expert overlap when using dispatch overlap cfg.model.moe_shared_expert_overlap = False ``` ### Prerequisites - `expert_model_parallel_size > 1` - `num_moe_experts > 1` - `moe_token_dispatcher_type` must be `"alltoall"` or `"flex"` - Precision: BF16 or FP16 - If PP is used, VPP (`virtual_pipeline_model_parallel_size`) must be set (non-`None`) ### Flex dispatcher activation Setting `moe_flex_dispatcher_backend` alone does **not** activate flex dispatch. You must also set `moe_token_dispatcher_type = "flex"`. ## Recompute And CUDA Graph Interaction - Full recompute is not a good companion for the overlap path. - `delay_wgrad_compute` adds further constraints if CUDA-graph scopes include attention or MoE-router work. - In practice, selective recompute is the safer pairing when o