← ClaudeAtlas

nemo-mbridge-perf-cuda-graphslisted

Validate and use CUDA graph capture in Megatron Bridge, including local full-iteration graphs and Transformer Engine scoped graphs for attention, MLP, and MoE modules.
yangwhale/CloseCrab · ★ 4 · AI & Automation · score 80
Install: claude install-skill yangwhale/CloseCrab
# CUDA Graphs Stable documentation: @docs/training/cuda-graphs.md Card: @skills/nemo-mbridge-perf-cuda-graphs/card.yaml ## What It Is CUDA graphs capture GPU operations once and replay them with minimal host-driver overhead. Bridge supports two implementations: | `cuda_graph_impl` | Mechanism | Scope support | |---|---|---| | `"local"` | MCore `FullCudaGraphWrapper` wrapping entire fwd+bwd | `full_iteration` | | `"transformer_engine"` | TE `make_graphed_callables()` per layer | `attn`, `mlp`, `moe`, `moe_router`, `moe_preprocess`, `mamba` | ## Quick Decision Start with TE-scoped graphs for most training workloads, then verify replay timing against eager on the same dispatcher, layout, and container: - dense models: `attn`, then optionally `mlp` - dropless MoE: `attn moe_router moe_preprocess` - VLMs: the same dropless-MoE scope, but only after the real-data path is stable Use `local` + `full_iteration` only when you specifically want full-iteration capture and can satisfy the tighter constraints. For recompute-heavy workloads: - TE-scoped graphs pair naturally with selective recompute - full recompute usually pushes you toward `local` full-iteration graphs or away from graphs entirely Related docs: - @docs/training/cuda-graphs.md - @docs/training/activation-recomputation.md ## Enablement ### Local full-iteration graph ```python cfg.model.cuda_graph_impl = "local" cfg.model.cuda_graph_scope = ["full_iteration"] cfg.model.cuda_graph_warmup_steps = 3 cfg.model.u