nemo-mbridge-perf-moe-comm-overlaplisted
Install: claude install-skill yangwhale/CloseCrab
# MoE Communication Overlap
For the higher-level overview, see:
- @docs/training/communication-overlap.md
- @skills/nemo-mbridge-perf-moe-comm-overlap/card.yaml
## Quick Decision
Use MoE communication overlap when:
- `EP > 1`
- token dispatch or combine time is visible in the profile
- the run is already correct and you are now tuning throughput
Avoid turning it on as an early bring-up step. It is easier to validate after
the dispatcher, routing mode, and recompute plan are already stable.
## Enablement
```python
cfg.comm_overlap.overlap_moe_expert_parallel_comm = True
# Optional: delayed wgrad for additional overlap
cfg.comm_overlap.delay_wgrad_compute = True
# IMPORTANT: disable shared expert overlap when using dispatch overlap
cfg.model.moe_shared_expert_overlap = False
```
### Prerequisites
- `expert_model_parallel_size > 1`
- `num_moe_experts > 1`
- `moe_token_dispatcher_type` must be `"alltoall"` or `"flex"`
- Precision: BF16 or FP16
- If PP is used, VPP (`virtual_pipeline_model_parallel_size`) must be set (non-`None`)
### Flex dispatcher activation
Setting `moe_flex_dispatcher_backend` alone does **not** activate flex dispatch.
You must also set `moe_token_dispatcher_type = "flex"`.
## Recompute And CUDA Graph Interaction
- Full recompute is not a good companion for the overlap path.
- `delay_wgrad_compute` adds further constraints if CUDA-graph scopes include
attention or MoE-router work.
- In practice, selective recompute is the safer pairing when o