← ClaudeAtlas

test-data-engineeringlisted

Use when generating synthetic/realistic test data, seeding fixtures, masking PII for a test environment, or designing deterministic/reusable test datasets. Underpins reliable automation across all the other testing skills.
mejbaurbahar/fagun · ★ 1 · Testing & QA · score 72
Install: claude install-skill mejbaurbahar/fagun
# Test Data Engineering ## Why this is its own discipline Flaky and unreliable tests are very often a test-data problem, not a test-logic problem: shared mutable state, non-deterministic random values, or production data with unpredictable content. Good test data design is what makes [[test-automation]] and [[regression-testing]] actually reliable. ## Principles - **Deterministic where it matters**: seed random generators explicitly so a failing test's exact input is reproducible — "it failed with some random string" is not debuggable; "it failed with seed 12345, input was X" is. - **Isolated per test**: each test creates and owns its data (via API/factory, not shared fixtures mutated across tests) so tests can run in any order or in parallel without interfering. - **Realistic shape, safe content**: production-like distributions (field lengths, null rates, unicode presence) without using real user data. - **Boundary-aware**: generated data should include boundary/edge values deliberately (empty string, max length, unicode/emoji, negative numbers) — see [[manual-testing]]'s boundary value analysis — not just "random valid-looking" values, which tend to cluster in the safe middle of the input space and miss real bugs. ## Synthetic data generation - Use a factory/builder pattern in code (e.g. Faker-style libraries) rather than hardcoded fixture files, so tests can override just the field under test and get sensible defaults for everything else. - For load/performance testing,