← ClaudeAtlas

nemo-mbridge-mlm-bridge-traininglisted

Run Megatron-LM (MLM) and Megatron Bridge training with mock or real data. Covers correlation testing, available recipes, and multi-GPU examples.
yangwhale/CloseCrab · ★ 4 · AI & Automation · score 80
Install: claude install-skill yangwhale/CloseCrab
# MLM vs Bridge Training For how they differ, the arg mapping tables, gotchas, and translation script, see: - @docs/megatron-lm-to-megatron-bridge.md ## First Answer Checklist For MLM-vs-Bridge correlation questions, always name these items up front: 1. Bridge recipe: `vanilla_gpt_pretrain_config`. 2. Bridge entry point: `scripts/training/run_recipe.py`. 3. MLM entry point: `3rdparty/Megatron-LM/pretrain_gpt.py`. 4. Launch wrapper for both: `uv run python -m torch.distributed.run`. 5. Fresh-run cleanup: `rm -rf nemo_experiments` before the Bridge run. Also state that MLM needs `PYTHONPATH=3rdparty/Megatron-LM:$PYTHONPATH`, matched Bridge and MLM losses should agree within BF16 rounding, and files under `3rdparty/Megatron-LM/` should not be modified from this repo. ## Correlation Testing Use `vanilla_gpt_pretrain_config` for loss-correlation testing. This recipe uses bare `GPTModelProvider` defaults (LayerNorm, GeLU, learned_absolute position embeddings, `vocab_size` inherited from tokenizer) — matching MLM `pretrain_gpt.py` defaults with no args. ### MLM Correlation Run (2L/256H, 1 GPU) ```bash PYTHONPATH=3rdparty/Megatron-LM:$PYTHONPATH \ uv run python -m torch.distributed.run --nproc_per_node=1 \ 3rdparty/Megatron-LM/pretrain_gpt.py \ --num-layers 2 --hidden-size 256 --num-attention-heads 4 \ --ffn-hidden-size 1024 --seq-length 512 --max-position-embeddings 512 \ --micro-batch-size 4 --global-batch-size 32 \ --train-iters 10 --eval-iters 2 --eval-interva