agent-evallisted
Install: claude install-skill Hefrock/agent-skills
# Agent Eval
Turns "does this actually work" into a repeatable, evidence-based answer instead of a gut feeling.
## How this works
1. **Figure out what "good" means first.** Before writing any eval, get a concrete definition of success from the user, or infer it from context and confirm it back to them. What does a correct/good output look like? What does a clearly bad one look like? Is there a reference answer, or is this judgment-based?
2. **Pick the eval type** — don't default to one without considering the fit:
- **Reference-based**: there's a known correct answer (exact or fuzzy/semantic match). Cheapest and most reliable, but only works when "correct" is well-defined.
- **Rubric-based (LLM-as-judge)**: quality is graded against explicit criteria (e.g. "factually grounded," "follows the required format," "appropriately concise"). Use `references/llm-judge-prompt.md` as the starting template — don't write a judge prompt from scratch each time.
- **Pairwise comparison**: judging which of two outputs is better, not scoring each in isolation. More reliable than absolute scoring for subjective quality, but watch for position bias — always run both orderings and average. Use `scripts/run_pairwise.py` for this — it runs both orderings itself and reconciles them (see `references/pairwise-comparison.md`), rather than leaving "remember to run it twice" as a step to repeat by hand each time.
- **Programmatic/structural**: format compliance, schema validation, code th