← ClaudeAtlas

agent-quality-gradinglisted

Use when evaluating how well an agent completed real tasks, used tools, communicated, or produced assets; grade each dimension from traceable evidence and keep quality separate from reliability.
sandbaseai/workbuddy-skill · ★ 2 · AI & Automation · score 81
Install: claude install-skill sandbaseai/workbuddy-skill
# Agent Quality Grading Use this skill to measure how well an agent served a user, independently of whether the underlying system was reliable. A polished message does not prove the task happened, and a successful tool call does not prove the result was useful. Grade only the evidence available for the requested scope. ## Establish scope and privacy Define the time window, channels, agents, task classes, asset types, comparison baseline, and scoring audience. Obtain authorization for every trace and artifact. Prefer redacted or synthetic data; never copy credentials, personal data, private prompts, tokens, or raw secret-like content into the report. Record missing logs, clock assumptions, unavailable assets, and sampling limits. Reconstruct each conversation or task from inbound request, agent turns, tool calls/results, generated artifacts, and final user-visible outcome in chronological order. Preserve IDs or hashes only when needed to join evidence. A claim such as “sent,” “updated,” or “deployed” requires a matching successful operation in the same trace; otherwise grade task completion as unverified or fabricated according to the repository policy. ## Grade four independent axes Give each relevant task or channel a grade from A–F, with a short rationale and exact evidence reference: - **TASK** — did the agent satisfy the requested outcome fully, partially, not at all, or claim work without a matching operation? Check scope, correctness, side effects, constraints, a