evals-ops

Solid

Build and run evals for LLM and agent systems: golden datasets, LLM-as-a-judge with bias control, trajectory/step/outcome scoring, adversarial refuters, and CI regression gates. Triggers on: eval, evals, eval harness, golden dataset, golden set, LLM-as-a-judge, judge rubric, judge bias, judge calibration, Cohen kappa, agreement with human labels, pass@k, pass^k, trajectory eval, step-level eval, tool-call accuracy, regression gate, eval CI, blocking vs advisory eval, faithfulness score, DeepEval, Braintrust, Opik, Langfuse, AgentOps, did my prompt change make it worse, is my agent getting better.

AI & Automation 42 stars 7 forks Updated 1 weeks ago MIT

Install

View on GitHub

Quality Score: 84/100

Stars 20%
54
Recency 20%
90
Frontmatter 20%
70
Documentation 15%
100
Issue Health 10%
80
License 10%
100
Description 5%
100

Skill Content

# Evals Ops **Evals are the prerequisite, not the polish.** You cannot tune a prompt, a retriever, a compaction strategy or a memory layer without a harness that says whether the change made things better. Teams that skip this ship vibes and learn about regressions from users. This skill is the operational layer: what to measure, how to build the dataset, how to make a judge trustworthy, and how to gate CI on it without teaching everyone to ignore red. ## Route first | The ask | Go to | |---|---| | "What should I even measure?" | [Three levels](#three-levels-of-agent-eval) → `references/eval-taxonomy.md` | | "Where do the test cases come from?" | [Golden set](#the-golden-set) → `references/golden-datasets.md` | | "My judge disagrees with me / is it any good?" | [Judges](#llm-as-a-judge) → `references/llm-judge.md` | | "Verify a finding is real, not plausible" | [Refuters](#adversarial-verification) → `references/adversarial-verification.md` | | "Is my RAG retrieval any good?" | [Retrieval](#retrieval) → `references/retrieval-eval.md` | | "Where do the human labels come from?" | `references/annotation-workflow.md` | | "Should this block the merge?" | [Gating](#regression-gating) → `references/regression-gating.md` | | "Did this change really make it worse?" | [Is the drop real](#is-the-drop-real) | | "Optimise against the eval / run it overnight" | [Hillclimbing](#hillclimbing) → `references/hillclimbing.md` | | "Which platform should we use?" | `references/tooling-landsca...

Details

Author
0xDarkMatter
Repository
0xDarkMatter/claude-mods
Created
10 months ago
Last Updated
1 weeks ago
Language
Shell
License
MIT

Integrates with

Bundled in these plugins

Similar Skills

Semantically similar based on skill content — not just same category