ai-pretraining

Featured

Builds a transformer/GPT and BPE tokenizer from scratch. Use when implementing autograd, self-attention, a nanoGPT-style pretraining loop, or a byte-level tokenizer.

AI & Automation 89 stars 19 forks Updated 3 days ago MIT

Install

View on GitHub

Quality Score: 91/100

Stars 20%
65
Recency 20%
100
Frontmatter 20%
70
Documentation 15%
100
Issue Health 10%
80
License 10%
100
Description 5%
100

Skill Content

# Pretraining From Scratch **Domain**: building a transformer/GPT and a BPE tokenizer from first principles — the from-first-principles training-layer competency. Does NOT cover applications-layer fine-tuning, RLHF, or inference optimization; those belong to sibling skills. Canonical teachers: Karpathy "Neural Networks: Zero to Hero" (micrograd → makemore → "Let's build GPT" → "Let's build the GPT Tokenizer" → "Let's reproduce GPT-2"), Karpathy nanochat (full-stack from-scratch successor to nanoGPT, 2025), Raschka "Build a Large Language Model From Scratch", nanoGPT, minbpe, "Attention Is All You Need". GPT-2 is the pedagogical spine here — the right thing to build *first*. The 2026 from-scratch baseline then swaps four components onto that spine (RoPE, RMSNorm, SwiGLU, GQA) and runs attention through FlashAttention/SDPA; see [Modern Architecture Deltas](references/modern-architecture-deltas.md). ## ASCII Flow ```text Raw text corpus | v BPE Tokenizer (byte-level merges, vocab, encode/decode) | v Token IDs -> Embedding table (vocab_size x n_embd) | v + Positional Embedding (learned, shape: block_size x n_embd) | v Transformer Block x N ├── LayerNorm (pre-norm placement in GPT-2 style) ├── Multi-Head Self-Attention (causal mask, k/q/v projections) ├── Residual connection ├── LayerNorm ├── FFN (Linear -> GELU -> Linear, 4x expansion) └── Residual connection | v Final LayerNorm | v LM Head (Linear, n_embd -> vocab_size, weight-tied to emb...

Details

Author
vasilyu1983
Repository
vasilyu1983/AI-Agents-public
Created
10 months ago
Last Updated
3 days ago
Language
Python
License
MIT

Integrates with

Similar Skills

Semantically similar based on skill content — not just same category