agent-evaluationlisted
Install: claude install-skill aicodedecode/awesome-muse-skills
# Agent Evaluation
Demos lie; evals tell the truth. Agent evaluation is the practice of measuring what an agent does
on representative tasks with clear success criteria — so improvements are real and regressions
get caught.
## Overview
An eval is a task plus a grader. Tasks should mirror real usage: same tools, same ambiguity, same
edge cases. Graders range from exact-match checks to model-based judges to human review. Track
success rate, cost, latency, and failure modes together — an agent that's 95% successful at 10×
the cost may be worse than the 90% one. Run evals on every meaningful change; that's what makes
them a safety net instead of a ceremony.
## When to use
- Before shipping any agent: establishing a baseline of real performance.
- Comparing approaches: prompts, models, tools, architectures — decide by numbers.
- After any change: catching regressions before users do.
- Debugging: a failing eval case is a reproducible bug report.
## Core concepts
- **Task suites**: 20–100 representative tasks covering happy paths, edge cases, and adversarial
inputs. Small enough to run often, diverse enough to matter.
- **Graders**: exact match for deterministic outputs; structured rubrics for open-ended work; model
judges for scale (with spot-checked human agreement); humans for the highest-stakes cases.
- **Success criteria**: defined per task before running — what counts as done, including partial
credit rules. Vague criteria produce arguable results.
- **Failure