← ClaudeAtlas

agent-evallisted

Designs and runs evaluations for LLM or agent outputs — builds rubrics, sets up LLM-as-judge scoring, creates regression test sets, and reports pass rates with concrete failure examples. Use this skill whenever the user wants to evaluate, test, grade, or score an agent's or LLM's outputs; asks "how do I know if this is working," "is this any good," "set up an eval," or "did this prompt change make things worse"; needs a rubric for judging quality; wants to compare two prompts, models, or outputs; or wants to catch regressions before shipping a change. Also trigger when reviewing agent trajectories specifically — did the agent pick the right tool, the right arguments, the right sequence — not just the final output.
Hefrock/agent-skills · ★ 0 · AI & Automation · score 68
Install: claude install-skill Hefrock/agent-skills
# Agent Eval Turns "does this actually work" into a repeatable, evidence-based answer instead of a gut feeling. ## How this works 1. **Figure out what "good" means first.** Before writing any eval, get a concrete definition of success from the user, or infer it from context and confirm it back to them. What does a correct/good output look like? What does a clearly bad one look like? Is there a reference answer, or is this judgment-based? 2. **Pick the eval type** — don't default to one without considering the fit: - **Reference-based**: there's a known correct answer (exact or fuzzy/semantic match). Cheapest and most reliable, but only works when "correct" is well-defined. - **Rubric-based (LLM-as-judge)**: quality is graded against explicit criteria (e.g. "factually grounded," "follows the required format," "appropriately concise"). Use `references/llm-judge-prompt.md` as the starting template — don't write a judge prompt from scratch each time. - **Pairwise comparison**: judging which of two outputs is better, not scoring each in isolation. More reliable than absolute scoring for subjective quality, but watch for position bias — always run both orderings and average. Use `scripts/run_pairwise.py` for this — it runs both orderings itself and reconciles them (see `references/pairwise-comparison.md`), rather than leaving "remember to run it twice" as a step to repeat by hand each time. - **Programmatic/structural**: format compliance, schema validation, code th