benchmark-curatorlisted
Install: claude install-skill CometWeb-io/agent-skills
# Benchmark Curator
Own the **test population**, not the candidate and not the final experiment verdict. A benchmark is evidence infrastructure and must be versioned like code.
## Workflow
1. Freeze benchmark objective, target skill/task family, mode, population, and intended promotion use.
2. Build a taxonomy before adding cases: discovery, forced behavior, negative controls, adversarial/edge cases, and regression cases as applicable.
3. Record provenance and contamination status for each case. Never call a known/leaked case a clean holdout.
4. Stratify difficulty and source lane; reject accidental duplication and one-class domination.
5. In DEEP mode maintain a meaningful holdout split, run leakage detection, and keep known-contaminated or materially suspicious cases out of a frozen holdout.
6. Require inspectable assertions/expected behavior for every case.
7. Canonicalize and hash the benchmark revision. Any case-set change creates a new hash/revision.
8. Hand the frozen suite to `skill-evaluator`; do not execute the benchmark or claim candidate lift yourself.
## Hard rules
- A benchmark cannot be representative merely because it is large.
- Negative controls are required for discovery/routing claims.
- Cases derived from the exact candidate failure may enter regression/dev, but not an untouched holdout without explicit contamination treatment.
- A benchmark hash identifies bytes/semantics of the case contract, not universal task representativeness.
## Instruction b