ai-eval-reviewlisted
Install: claude install-skill sorawit-w/agent-skills
# AI Eval Review
Audit the eval layer of an AI product or feature against a structured set of
design-completeness questions. Each of the seven blocks is a *gap detector* —
the value of this skill is whether the elicitation surfaces eval decisions
the builder hasn't made yet, not whether the artifact looks complete.
The job is **eval-design-completeness**: *have we designed how we'll know if
this works?* Covers offline criteria, ground-truth quality, online signal,
cohort breakdowns and disparate impact, adversarial / robustness coverage,
and drift detection. Regulatory rigor (EU AI Act, FDA SaMD, FTC) is a
cross-cutting lens applied across blocks, not a separate block.
This skill is the eval-side companion to `ai-ux-review`. Same shape, same
elicitation pattern, different subject — `ai-ux-review` audits the
human-AI design surface (was the experience intentionally designed?);
this skill audits the measurement layer behind it (do we have signal for
whether the design works?).
This is not an implementation tool. It does not write eval code, label
datasets, set up monitoring dashboards, or compute metrics. It names the
gaps; you take them to your tools or your team to close.
## What this skill produces
Always produced under the resolved review root (default
`docs/ai-ux/` — same folder as `ai-ux-review` for clean composition):
1. **`ai-eval-review.md`** — canonical, editable Markdown with one top-level
section per block (seven blocks total), plus a `## Gap Summary` sect