← ClaudeAtlas

agent-evaluationlisted

Evaluate AI agents rigorously — task benchmarks, success criteria, failure taxonomy, cost/latency tracking, and regression testing. Use when you need to know whether an agent actually works, not whether it demos well.
aicodedecode/awesome-muse-skills · ★ 0 · AI & Automation · score 75
Install: claude install-skill aicodedecode/awesome-muse-skills
# Agent Evaluation Demos lie; evals tell the truth. Agent evaluation is the practice of measuring what an agent does on representative tasks with clear success criteria — so improvements are real and regressions get caught. ## Overview An eval is a task plus a grader. Tasks should mirror real usage: same tools, same ambiguity, same edge cases. Graders range from exact-match checks to model-based judges to human review. Track success rate, cost, latency, and failure modes together — an agent that's 95% successful at 10× the cost may be worse than the 90% one. Run evals on every meaningful change; that's what makes them a safety net instead of a ceremony. ## When to use - Before shipping any agent: establishing a baseline of real performance. - Comparing approaches: prompts, models, tools, architectures — decide by numbers. - After any change: catching regressions before users do. - Debugging: a failing eval case is a reproducible bug report. ## Core concepts - **Task suites**: 20–100 representative tasks covering happy paths, edge cases, and adversarial inputs. Small enough to run often, diverse enough to matter. - **Graders**: exact match for deterministic outputs; structured rubrics for open-ended work; model judges for scale (with spot-checked human agreement); humans for the highest-stakes cases. - **Success criteria**: defined per task before running — what counts as done, including partial credit rules. Vague criteria produce arguable results. - **Failure