evals-ops
SolidBuild and run evals for LLM and agent systems: golden datasets, LLM-as-a-judge with bias control, trajectory/step/outcome scoring, adversarial refuters, and CI regression gates. Triggers on: eval, evals, eval harness, golden dataset, golden set, LLM-as-a-judge, judge rubric, judge bias, judge calibration, Cohen kappa, agreement with human labels, pass@k, pass^k, trajectory eval, step-level eval, tool-call accuracy, regression gate, eval CI, blocking vs advisory eval, faithfulness score, DeepEval, Braintrust, Opik, Langfuse, AgentOps, did my prompt change make it worse, is my agent getting better.
Install
Quality Score: 84/100
Skill Content
Details
- Author
- 0xDarkMatter
- Repository
- 0xDarkMatter/claude-mods
- Created
- 10 months ago
- Last Updated
- 1 weeks ago
- Language
- Shell
- License
- MIT
Integrates with
Bundled in these plugins
Similar Skills
Semantically similar based on skill content — not just same category
agent-eval
Use when measuring whether an LLM or agent system actually got better and gating merges on it: golden sets, fixing an inflated LLM-as-judge, scoring RAG (faithfulness, contextual recall) or agent trajectories (tool correctness, completion), or picking an eval framework. NOT building the agent loop, tools or RAG plumbing (that is `building-agents`).
eval-genius
Decide whether an AI/LLM/agent/retrieval system needs an eval, where it fits in the dev process, which one to run, and how to read the result; then design, gate, judge, and defend it. Not ordinary unit tests.
ai-evals
Use when testing an LLM-backed feature, prompt, tool loop or multi-step agent, where the same input can produce different outputs and a prompt or model change can regress behaviour with no code diff. The behaviour spec that precedes the prompt, scenario datasets including adversarial and degradation classes, deterministic assertions over OpenTelemetry traces, calibrated LLM-as-judge, CI gates with baselines, human-in-the-loop, and the production scoring loop that turns incidents into scenarios.