← ClaudeAtlas

nemo-mbridge-perf-sequence-packinglisted

Validate and use packed sequences and long-context training in Megatron-Bridge, distinguishing offline packed SFT for LLMs from in-batch packing for VLMs, and applying the right CP constraints.
yangwhale/CloseCrab · ★ 4 · AI & Automation · score 80
Install: claude install-skill yangwhale/CloseCrab
# Sequence Packing Skill For stable background and recommendation level, see: - @docs/training/packed-sequences.md - @skills/nemo-mbridge-perf-sequence-packing/card.yaml ## Enablement Offline packed SFT for LLM finetuning: ```python from megatron.bridge.data.datasets.packed_sequence import PackedSequenceSpecs cfg.train.micro_batch_size = 1 cfg.dataset.seq_length = 4096 cfg.model.seq_length = 4096 cfg.dataset.dataset_kwargs = {"pad_to_max_length": True} cfg.dataset.packed_sequence_specs = PackedSequenceSpecs( packed_sequence_size=4096, pad_seq_to_mult=1, ) ``` If CP is enabled: ```python cfg.model.context_parallel_size = 2 cfg.model.calculate_per_token_loss = True cfg.ddp.average_in_collective = False cfg.dataset.packed_sequence_specs.pad_seq_to_mult = cfg.model.context_parallel_size * 2 # If sequence_parallel is also enabled, use lcm(2*CP, CP*TP): # import math # cfg.dataset.packed_sequence_specs.pad_seq_to_mult = math.lcm(2 * CP, CP * TP) # See src/megatron/bridge/training/vlm_step.py for reference logic. ``` If CUDA graphs are enabled for this packed path: ```python cfg.dataset.packed_sequence_specs.pad_cu_seqlens = True cfg.dataset.dataset_kwargs["pad_to_max_length"] = True ``` **Note:** `pad_cu_seqlens = True` also requires a metadata JSON file alongside the packed dataset (asserted in `src/megatron/bridge/data/datasets/sft.py`). Custom packed datasets that omit the metadata file will hit an assertion at dataset initialization. In-batch packing for VL