← ClaudeAtlas

incremental-model-patternslisted

Build reliable incremental dbt models: choose the right unique_key and strategy (append, merge, delete+insert), handle late-arriving data and out-of-order events, write a safe is_incremental filter, and design the full-refresh fallback — so the model is idempotent from day one.
mcorbett51090/RavenClaude · ★ 7 · AI & Automation · score 65
Install: claude install-skill mcorbett51090/RavenClaude
# Skill: incremental-model-patterns **Purpose:** Produce incremental dbt models that are correct under rerun, late-data, and full-refresh scenarios. Used by `analytics-engineer` (primary) and `data-quality-testing-engineer` (testing incremental correctness). ## When to use - A fact table is large enough that a full rebuild is expensive or violates SLA (general rule: > 10M rows or > 30 minutes for a full table scan). - A source grows mostly via append — new events, new orders, new log lines. - The model must remain idempotent: running it twice must produce the same result as running it once. --- ## Step 1: Choose the incremental strategy | Strategy | When to use | Key constraint | |---|---|---| | `append_only` (no deduplication) | Source is truly append-only and no row is ever updated | Reruns will create duplicate rows — only safe if you NEVER rerun on overlapping data | | `merge` (default for most warehouses) | Source rows can be updated (e.g. order status changes), or you need upsert behaviour | Requires a `unique_key`; most expensive at scale | | `delete+insert` (BigQuery, Snowflake) | Bulk-replace a time-partition or date-range in one operation | Requires a `partition_by` clause; excellent for date-partitioned fact tables with late updates | | `insert_overwrite` (Spark/Databricks) | Partition-level replacement on Spark | Partition column must be in the `partition_by` block | **Recommendation for most fact tables:** `merge` with a `unique_key` on the natural event g