← ClaudeAtlas

nemo-mbridge-multi-node-slurmlisted

Convert single-node scripts to multi-node Slurm sbatch jobs and debug common multi-node failures. Covers srun-native vs uv run torch.distributed approaches, container setup, NCCL timeouts, OOM sizing for MoE models, and interactive allocation.
yangwhale/CloseCrab · ★ 4 · AI & Automation · score 80
Install: claude install-skill yangwhale/CloseCrab
# Multi-Node Slurm Convert single-node `uv run python -m torch.distributed.run` commands into multi-node Slurm sbatch scripts with Enroot container support, and debug common multi-node failures. ## First Answer Checklist When converting or debugging Bridge multi-node jobs, answer in this order: 1. Prefer the **srun-native** launch shape for Bridge scripts that reach `initialize.py`: `#SBATCH --ntasks-per-node=8` and a direct `srun ... uv run python <script> ...` launch. Do not wrap these jobs in `python -m torch.distributed.run`. 2. State that Bridge derives `RANK`, `WORLD_SIZE`, `LOCAL_RANK`, `MASTER_ADDR`, and `MASTER_PORT` from SLURM variables during `initialize.py` distributed init. 3. Require shared paths and matching container mounts for the repo, data, logs, `HF_HOME`, `UV_CACHE_DIR`, and `NEMO_HOME`. 4. For NCCL timeout reports, do these first-log checks before speculating: - grep for real errors while filtering warning/frame noise - inspect `Failures:` to find the first failed rank and node - grep for `ncclUniqueId`, `timeout`, or `crash on rank 0` ## Two Approaches: srun-native vs uv run torch.distributed | Approach | `ntasks-per-node` | Process spawning | Best for | |---|---|---|---| | **srun-native** (preferred) | 8 | Slurm spawns 8 tasks/node | Conversion, inference, Bridge scripts | | **uv run torch.distributed** (legacy) | 1 | `uv run python -m torch.distributed.run` spawns 8 procs/node | MLM pretrain_gpt.py | **Prefer srun-nat