skill-evaluatorlisted
Install: claude install-skill CometWeb-io/agent-skills
# Skill Evaluator
Measure **behavioral lift**, not prose quality. A skill can be structurally excellent and still fail to trigger or improve outcomes; a weak-looking package can sometimes work. Keep those questions separate.
## Workflow
1. Freeze candidate skill identity, baseline (`NO_SKILL` or prior version), **rubric hash**, **benchmark hash**, eval suite hash, host/model/harness configuration, reasoning effort, and budgets.
2. Separate discovery tests from forced-invocation behavior tests.
3. Include negative trigger controls so improved recall cannot hide precision collapse.
4. Use identical cases and grading conditions for candidate and baseline. Predeclare exclusions before execution.
5. Prefer deterministic graders for observable state; use model grading only for assertions that cannot be checked mechanically.
6. Repeat stochastic real-host cases. STANDARD requires at least 3 runs/case for promotion claims; DEEP requires at least 5.
7. Record skipped/ungradeable assertions as skipped, never silently failed or passed.
8. For paired candidate/baseline cases, use paired significance rather than independent-rate intuition; DEEP real-host promotion also requires stable repeated-run evidence.
9. If using early stopping, predeclare checkpoints and alpha spending before execution.
10. Compare pass rate, trigger precision/recall, invariant regressions, token/cost usage, and wall-clock time.
11. Report `IMPROVED`, `NO_MATERIAL_CHANGE`, `TRADEOFF`, `REGRESSION`, `INSUFFICIENT