nemo-automodel-distributed-training

Solid

Guide for selecting and configuring distributed training strategies in NeMo AutoModel, including FSDP2, Megatron FSDP, DDP, and parallelism settings.

AI & Automation 4 stars 0 forks Updated today Apache-2.0

Install

View on GitHub

Quality Score: 83/100

Stars 20%
23
Recency 20%
100
Frontmatter 20%
70
Documentation 15%
100
Issue Health 10%
80
License 10%
100
Description 5%
100

Skill Content

# Distributed Training in NeMo AutoModel ## Purpose NeMo AutoModel uses PyTorch-native distributed training. All parallelism is orchestrated through a single `MeshContext` object that holds device meshes, strategy configs, and axis names. ## Instructions For conceptual distributed-training questions, answer directly from the quick patterns in this skill without inspecting the repository. Start with the strategy choice, then list only the YAML fields and constraints relevant to the question. Use direct action verbs in the final answer: recommend the strategy, show the minimal YAML, state the sizing constraint, and name the unsupported strategies. Do not discuss model onboarding, recipes, Slurm, SkyPilot, or checkpointing unless the user asks. ## Examples ### TP plus PP for a large multi-node model Recommend `strategy: fsdp2`. Mention `tp_size`, `pp_size`, `cp_size`, `ep_size`, and the `pipeline` sub-config. State that `dp_size` is inferred from `world_size / (tp_size * pp_size * cp_size)`. ```yaml distributed: strategy: fsdp2 tp_size: 8 pp_size: 4 cp_size: 1 ep_size: 1 pipeline: pp_schedule: interleaved1f1b pp_microbatch_size: 1 ``` ### MoE expert parallelism Recommend `strategy: fsdp2` with `ep_size > 1`. Say this creates a separate `moe_mesh`; include the `moe` sub-config when relevant; state that `ep_size` must divide `dp_size * cp_size`. Do not recommend `megatron_fsdp` or `ddp`. ```yaml distributed: strategy: fsdp2 ep_size: 8 moe: ...

Details

Author
yangwhale
Repository
yangwhale/CloseCrab
Created
6 months ago
Last Updated
today
Language
Python
License
Apache-2.0

Similar Skills

Semantically similar based on skill content — not just same category