← ClaudeAtlas

eval-geniuslisted

Decide whether an AI/LLM/agent/retrieval system needs an eval, where it fits in the dev process, which one to run, and how to read the result; then design, gate, judge, and defend it. Not ordinary unit tests.
alexgreensh/eval-genius · ★ 14 · AI & Automation · score 80
Install: claude install-skill alexgreensh/eval-genius
# Eval Genius An eval is a claim you are willing to defend under hostile audit. You measure to earn the right to say "this is better" and have it hold when someone sharp pushes back. Behave like a measurement engineer: state the promise, fix the bar before looking, hold everything else constant, distrust the instrument first, report the number that hurts. --- ## Step 0: Does this need an eval, and where does it go? Three questions decide it (`references/00-start-here.md`): does the output vary? will it change again, and would a quiet regression cost something? is a decision or public claim coming? No to all: a spot check, stop. Yes to any: an eval, sized to the project's stage. Two checks sit in front. Preflight: a domain expert can verify the output without redoing the work, else CANNOT-MEASURE (`references/01-foundation.md`). Triage: a known, frequent defect is fixed now, not measured (FIX_NOW / MEASURE / CANNOT-MEASURE). | Stage the user is at | Instrument | Smallest useful version | |---|---|---| | Exploring prompts and models | Spot check | 10 inputs, eyeball | | First working version | Smoke eval | 20 to 50 real inputs, code-checked; **this run is the baseline** | | Changing one thing | Paired eval vs baseline | Same items both arms, per-item diff, bar written first | | Merging or shipping | CI gate | Held-out items, three-way outcome, a known-bad item that must fail | | Comparing or claiming publicly | Benchmark | Versioned dataset and harness, intervals, report |