test-data-engineeringlisted
Install: claude install-skill mejbaurbahar/fagun
# Test Data Engineering
## Why this is its own discipline
Flaky and unreliable tests are very often a test-data problem, not a test-logic problem: shared mutable state, non-deterministic random values, or production data with unpredictable content. Good test data design is what makes [[test-automation]] and [[regression-testing]] actually reliable.
## Principles
- **Deterministic where it matters**: seed random generators explicitly so a failing test's exact input is reproducible — "it failed with some random string" is not debuggable; "it failed with seed 12345, input was X" is.
- **Isolated per test**: each test creates and owns its data (via API/factory, not shared fixtures mutated across tests) so tests can run in any order or in parallel without interfering.
- **Realistic shape, safe content**: production-like distributions (field lengths, null rates, unicode presence) without using real user data.
- **Boundary-aware**: generated data should include boundary/edge values deliberately (empty string, max length, unicode/emoji, negative numbers) — see [[manual-testing]]'s boundary value analysis — not just "random valid-looking" values, which tend to cluster in the safe middle of the input space and miss real bugs.
## Synthetic data generation
- Use a factory/builder pattern in code (e.g. Faker-style libraries) rather than hardcoded fixture files, so tests can override just the field under test and get sensible defaults for everything else.
- For load/performance testing,