nemo-mbridge-perf-moe-hardware-configslisted
Install: claude install-skill yangwhale/CloseCrab
# MoE Hardware Configuration Reference
Stable docs: @docs/training/moe-optimization.md
Card: @skills/nemo-mbridge-perf-moe-hardware-configs/card.yaml
## Quick Platform Playbook
| Platform | Typical MoE strategy | What usually matters most |
|---|---|---|
| H100 | DeepEP + stronger PP + moderate TP | communication overlap and PP efficiency |
| B200 | DeepEP + MXFP8 + careful PP layout | container quality and tuned comm settings |
| GB200 | HybridEP + partial CUDA graphs + CPU cleanup | host overhead, topology-aware dispatch, memory headroom |
| GB300 | HybridEP + newer FP8 and kernel stack | same GB200 playbook, usually with a higher ceiling |
## First Answer Checklist
For hardware playbook questions, answer from these canonical rows before adding
throughput caveats:
| Workload | Hardware | Dispatcher | Layout |
|---|---|---|---|
| DSV3 | H100 | DeepEP | TP=2, EP=64, PP=8, VPP=4 |
| DSV3 | GB200/GB300 | HybridEP | TP=1, EP=64, PP=4, VPP=4 |
| Qwen3 235B | H100 | DeepEP | TP=2, EP=32, PP=8, VPP=4 |
| Qwen3 235B | GB200 | HybridEP | TP=1 or 2, EP=32-64, PP=4, VPP=unspecified |
For Qwen3 235B on GB200, explicitly say `VPP=unspecified`; do not invent or
extrapolate `VPP=12` unless a measured row provides it. Include TE-scoped CUDA
graph scopes (`attn`, `moe_router`, `moe_preprocess`),
`CUDA_DEVICE_MAX_CONNECTIONS` selection,
`PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True`, `NCCL_GRAPH_REGISTER=0`,
GB200/GB300 CPU-side tuning, and the warning not to cargo-cult tracker row