← ClaudeAtlas

nemo-mbridge-perf-moe-hardware-configslisted

Representative MoE training playbooks by hardware platform and model family. Summarizes rounded throughput bands, parallelism patterns, and common tuning stacks.
yangwhale/CloseCrab · ★ 4 · AI & Automation · score 80
Install: claude install-skill yangwhale/CloseCrab
# MoE Hardware Configuration Reference Stable docs: @docs/training/moe-optimization.md Card: @skills/nemo-mbridge-perf-moe-hardware-configs/card.yaml ## Quick Platform Playbook | Platform | Typical MoE strategy | What usually matters most | |---|---|---| | H100 | DeepEP + stronger PP + moderate TP | communication overlap and PP efficiency | | B200 | DeepEP + MXFP8 + careful PP layout | container quality and tuned comm settings | | GB200 | HybridEP + partial CUDA graphs + CPU cleanup | host overhead, topology-aware dispatch, memory headroom | | GB300 | HybridEP + newer FP8 and kernel stack | same GB200 playbook, usually with a higher ceiling | ## First Answer Checklist For hardware playbook questions, answer from these canonical rows before adding throughput caveats: | Workload | Hardware | Dispatcher | Layout | |---|---|---|---| | DSV3 | H100 | DeepEP | TP=2, EP=64, PP=8, VPP=4 | | DSV3 | GB200/GB300 | HybridEP | TP=1, EP=64, PP=4, VPP=4 | | Qwen3 235B | H100 | DeepEP | TP=2, EP=32, PP=8, VPP=4 | | Qwen3 235B | GB200 | HybridEP | TP=1 or 2, EP=32-64, PP=4, VPP=unspecified | For Qwen3 235B on GB200, explicitly say `VPP=unspecified`; do not invent or extrapolate `VPP=12` unless a measured row provides it. Include TE-scoped CUDA graph scopes (`attn`, `moe_router`, `moe_preprocess`), `CUDA_DEVICE_MAX_CONNECTIONS` selection, `PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True`, `NCCL_GRAPH_REGISTER=0`, GB200/GB300 CPU-side tuning, and the warning not to cargo-cult tracker row