← ClaudeAtlas

nemo-mbridge-perf-parallelism-strategieslisted

Operational guide for choosing and combining parallelism strategies in Megatron Bridge, including sizing rules, hardware topology mapping, and combined parallelism configuration.
yangwhale/CloseCrab · ★ 4 · AI & Automation · score 80
Install: claude install-skill yangwhale/CloseCrab
# Parallelism Strategy Selection Skill For stable background on each parallelism type, see: - @docs/parallelisms.md - @skills/nemo-mbridge-perf-parallelism-strategies/card.yaml ## Decision by Model Size ### Dense models | Model size | GPUs | Recommended starting point | |---|---|---| | < 1B | 1-8 | DP only | | 1-10B | 8-16 | TP=2-4 + DP | | 10-70B | 16-64 | TP=4-8 + PP=2-4 + DP | | 70-175B | 64-256 | TP=8 + PP=4-8 + DP | | 175-500B | 256-1024 | TP=8 + PP=8-16 + CP=2 + DP | ### MoE models MoE parallelism differs from dense models. Because only a fraction of parameters are active per token, TP can often stay at 1 or 2 — the active parameter shard already fits on a single GPU. EP is the primary scaling dimension, with PP handling cross-node layer distribution. | Model (total / active) | TP | PP | EP | Notes | |---|---|---|---|---| | OLMoE 7B / 1B | 1 | 1 | 8 | EP only, fits single node | | Moonlight 16B / 3B | 2 | 1 | 8 | small TP for shared layers | | DeepSeek-V2 236B / 21B | 1 | 4 | 32 | no TP at all | | GLM-4.5 Air 106B / 12B | 1 | 4 | 8 | no TP at all | | Qwen3 30B-A3B | 4 | 2 | 4 | | | GLM-4.5 355B / 32B | 2 | 8 | 16 | | | Qwen3 235B-A22B | 4 | 16 | 8 | CP=2 for pretrain | | DeepSeek-V3 671B / 37B | 2 | 16 | 64 | TP=2, not 8 | | Kimi-K2 1T | 2 | 16 | 32 | | Key patterns: - TP is sized by **active** params, not total params. A 671B MoE with 37B active needs far less TP than a 70B dense model. - EP scales with expert count. Common: EP = num_experts or num_expert