← ClaudeAtlas

skill-eval-looplisted

Measure whether an agent skill actually helps: prompt sets with should/should-not trigger, with/without skill runs, deterministic checks, rubric grading, CI regressions. Use when hardening a SKILL.md, deciding if a rewrite improved quality, or adding a quality gate for skills.
novahiz/novahiz · ★ 1 · AI & Automation · score 70
Install: claude install-skill novahiz/novahiz
# skill-eval-loop Skill quality is a measured delta, not a feeling. Define done, run with and without the skill, score, and keep the score across versions. ## What an eval is `prompt → captured run (trace + artifacts) → checks → score you can compare over time` Three layers, cheapest first: 1. **Deterministic checks**: files exist, commands ran in order, JSONL events, exit codes. 2. **Rubric / judge**: qualitative structure and style; schema-constrained JSON output; blind to which condition produced the text when comparing A/B. 3. **Regression diff**: baseline snapshot vs current; fail CI when score or trigger rate drops past a threshold. ## Prompt set Start with 10-20 rows, grow from real failures. | Column | Meaning | |---|---| | id | stable key | | prompt | user-style request | | expect_trigger | yes / no | | must_steps | commands or artifacts that prove success (optional) | | must_not | failures that count as regression | Include negative prompts (nearby topics that must not load the skill). Without negatives, a broader description always looks “better” until production noise appears. ## Trigger testing - Explicit load: invoke the skill directly; fix the **body** if output is wrong. - Implicit load: paraphrase the user ask without naming the skill; fix the **description** if it never fires. - Over-trigger: unrelated prompts must not load it; narrow description or add a do-not-use clause. Description problems look like body problems until you separate these two