nemo-mbridge-mlm-bridge-traininglisted
Install: claude install-skill yangwhale/CloseCrab
# MLM vs Bridge Training
For how they differ, the arg mapping tables, gotchas, and translation script, see:
- @docs/megatron-lm-to-megatron-bridge.md
## First Answer Checklist
For MLM-vs-Bridge correlation questions, always name these items up front:
1. Bridge recipe: `vanilla_gpt_pretrain_config`.
2. Bridge entry point: `scripts/training/run_recipe.py`.
3. MLM entry point: `3rdparty/Megatron-LM/pretrain_gpt.py`.
4. Launch wrapper for both: `uv run python -m torch.distributed.run`.
5. Fresh-run cleanup: `rm -rf nemo_experiments` before the Bridge run.
Also state that MLM needs
`PYTHONPATH=3rdparty/Megatron-LM:$PYTHONPATH`, matched Bridge and MLM losses
should agree within BF16 rounding, and files under `3rdparty/Megatron-LM/`
should not be modified from this repo.
## Correlation Testing
Use `vanilla_gpt_pretrain_config` for loss-correlation testing. This recipe uses
bare `GPTModelProvider` defaults (LayerNorm, GeLU, learned_absolute position
embeddings, `vocab_size` inherited from tokenizer) — matching MLM
`pretrain_gpt.py` defaults with no args.
### MLM Correlation Run (2L/256H, 1 GPU)
```bash
PYTHONPATH=3rdparty/Megatron-LM:$PYTHONPATH \
uv run python -m torch.distributed.run --nproc_per_node=1 \
3rdparty/Megatron-LM/pretrain_gpt.py \
--num-layers 2 --hidden-size 256 --num-attention-heads 4 \
--ffn-hidden-size 1024 --seq-length 512 --max-position-embeddings 512 \
--micro-batch-size 4 --global-batch-size 32 \
--train-iters 10 --eval-iters 2 --eval-interva