skill-eval-looplisted
Install: claude install-skill novahiz/novahiz
# skill-eval-loop
Skill quality is a measured delta, not a feeling. Define done, run with and without the skill, score, and keep the score across versions.
## What an eval is
`prompt → captured run (trace + artifacts) → checks → score you can compare over time`
Three layers, cheapest first:
1. **Deterministic checks**: files exist, commands ran in order, JSONL events, exit codes.
2. **Rubric / judge**: qualitative structure and style; schema-constrained JSON output; blind to which condition produced the text when comparing A/B.
3. **Regression diff**: baseline snapshot vs current; fail CI when score or trigger rate drops past a threshold.
## Prompt set
Start with 10-20 rows, grow from real failures.
| Column | Meaning |
|---|---|
| id | stable key |
| prompt | user-style request |
| expect_trigger | yes / no |
| must_steps | commands or artifacts that prove success (optional) |
| must_not | failures that count as regression |
Include negative prompts (nearby topics that must not load the skill). Without negatives, a broader description always looks “better” until production noise appears.
## Trigger testing
- Explicit load: invoke the skill directly; fix the **body** if output is wrong.
- Implicit load: paraphrase the user ask without naming the skill; fix the **description** if it never fires.
- Over-trigger: unrelated prompts must not load it; narrow description or add a do-not-use clause.
Description problems look like body problems until you separate these two