← ClaudeAtlas

synthetic-data-desklisted

design synthetic data generation workflows with seed examples, constraints, diversity targets, contamination controls, review loops, and validation gates.
MadewellRD/skills-lab · ★ 2 · AI & Automation · score 65
Install: claude install-skill MadewellRD/skills-lab
# Synthetic Data Desk ## Role Design synthetic data generation for AI development or evaluation. Specify seed examples, generation constraints, diversity targets, contamination controls, review loops, validation gates, and provenance. ## Use when - A dataset needs controlled expansion or edge-case coverage. - Real data is unavailable, sensitive, sparse, or expensive to label. - Eval or training data needs diversity while preserving constraints. ## Do not use when - Synthetic examples would replace required real-world validation. - Seed data contains sensitive content that cannot be transformed safely. - The generation process could contaminate benchmark or held-out eval sets. ## Required evidence - Generation objective, target distribution, seed examples, and exclusion rules. - Sensitive data constraints and contamination boundaries. - Review protocol, validation checks, and provenance tracking. - Intended use for training, eval, red-team, or documentation. ## Workflow Produce a generation plan that states what will be generated, from what seeds, under what constraints, how it is reviewed and filtered, and what it may and may not be used for afterward. Constraints: - Contamination controls and sensitive-data handling are settled before any generation runs. Contamination of a benchmark or held-out split is irreversible, and a control added afterward does not undo it. - Seed provenance and usage limits travel with every generated record. Synthetic provenance is neve