← ClaudeAtlas

agent-evaluationlisted

Use when evaluating an AI agent — task completion, tool-use correctness, trajectory scoring, automation rate, and human-in-the-loop review. Triggers on "agent evaluation", "agent eval", "task completion rate", "tool-use accuracy", "trajectory", "automation rate".
noctua84/nescio-ai · ★ 0 · AI & Automation · score 73
Install: claude install-skill noctua84/nescio-ai
# Agent Evaluation ## Purpose Analyze and produce an agent evaluation that delivers actionable, measurable results. **Category**: AI & Automation ## Inputs ### Required - **Objective**: What you want to achieve with this deliverable - **Context**: Relevant background information ### Optional - **Constraints**: Any limitations or requirements to consider - **Existing Work**: Previous documents or data to build on ## Context Before starting, read the repo's `CLAUDE.md` and any relevant notes under `memory/` (e.g. `memory/repo/<repo>/`, `memory/feedback/`) for prior decisions and constraints. ## Process ### Step 1: Context & Research - Review any existing agent evaluation documents in the project - Identify key stakeholders and their requirements - Select the most appropriate framework: AI Readiness Assessment, Automation ROI Calculator, Human-in-the-Loop Design ### Step 2: Analysis & Framework Application - Apply the selected framework to structure the agent evaluation - Identify gaps, opportunities, and risks - Define success metrics: Time Saved Per Task, Automation Rate, Error Reduction %, Cost Per AI Operation - Document assumptions and dependencies - Validate approach against industry best practices ### Step 3: Build the Deliverable - Structure the agent evaluation using the output format below - Include specific, actionable recommendations — not generic advice - Add concrete numbers, timelines, and benchmarks where applicable - Cross-reference with existing pro