eval-geniuslisted
Install: claude install-skill alexgreensh/eval-genius
# Eval Genius
An eval is a claim you are willing to defend under hostile audit. You measure to earn the right to say "this is better" and have it hold when someone sharp pushes back.
Behave like a measurement engineer: state the promise, fix the bar before looking, hold
everything else constant, distrust the instrument first, report the number that hurts.
---
## Step 0: Does this need an eval, and where does it go?
Three questions decide it (`references/00-start-here.md`): does the output vary? will it
change again, and would a quiet regression cost something? is a decision or public claim
coming? No to all: a spot check, stop. Yes to any: an eval, sized to the project's stage.
Two checks sit in front. Preflight: a domain expert can verify the output without
redoing the work, else CANNOT-MEASURE (`references/01-foundation.md`). Triage: a
known, frequent defect is fixed now, not measured (FIX_NOW / MEASURE / CANNOT-MEASURE).
| Stage the user is at | Instrument | Smallest useful version |
|---|---|---|
| Exploring prompts and models | Spot check | 10 inputs, eyeball |
| First working version | Smoke eval | 20 to 50 real inputs, code-checked; **this run is the baseline** |
| Changing one thing | Paired eval vs baseline | Same items both arms, per-item diff, bar written first |
| Merging or shipping | CI gate | Held-out items, three-way outcome, a known-bad item that must fail |
| Comparing or claiming publicly | Benchmark | Versioned dataset and harness, intervals, report |