← ClaudeAtlas

design-skill-evalslisted

Design falsifiable skill evals: routing cases, oracles, failures.
SylphxAI/skills · ★ 1 · Web & Frontend · score 74
Install: claude install-skill SylphxAI/skills
# design-skill-evals # Design Skill Evals Produce a **Skill Evaluation Program** that can disprove whether an exact skill bundle improves realistic work while preserving routing boundaries, safety, cost, and evidence integrity. ## When to use - An exact skill bundle needs falsifiable routing/behavior evals with holdouts and replay - Promotion claims need candidate-bound digests, controls, oracles, and provenance - Not for the skill procedure itself (`author-skill`) or portfolio decisions (`curate-skill-repository`) ## Atomic boundary Own the eval contract: candidate digests, routing/behavior tasks, controls, rubrics/oracles, artifact checks, model/provider matrix, hidden-data handling, metrics/uncertainty, provenance, replay, attestation, expiry, and regression triage. Do not own the skill procedure, repository portfolio decision, runtime capability, or adoption/outcome telemetry. ## Workflow 1. Freeze the question: exact skill name/description/body/references/scripts, intended jobs/artifacts, nearest neighbours, risks, output budget, supported runtimes/tools, and claim that the eval may falsify. 2. Read `references/skill-eval-systems.md`. Compute an injection-contract digest over name+description and behavior digest over the complete ordered bundle. Bind candidate commit, catalog digest, task/rubric/runner/policy/model-registry digests, parameters, seed, tool availability, retries, and expiry. For repeated or continuously monitored evaluations, also