agent-eval
FeaturedUse when measuring whether an LLM or agent system actually got better and gating merges on it: golden sets, fixing an inflated LLM-as-judge, scoring RAG (faithfulness, contextual recall) or agent trajectories (tool correctness, completion), or picking an eval framework. NOT building the agent loop, tools or RAG plumbing (that is `building-agents`).
Install
Quality Score: 92/100
Skill Content
Details
- Author
- ericrisco
- Repository
- ericrisco/rsc-harness
- Created
- 3 months ago
- Last Updated
- 2 days ago
- Language
- JavaScript
- License
- MIT
Integrates with
Similar Skills
Semantically similar based on skill content — not just same category
agent-eval
Designs and runs evaluations for LLM or agent outputs — builds rubrics, sets up LLM-as-judge scoring, creates regression test sets, and reports pass rates with concrete failure examples. Use this skill whenever the user wants to evaluate, test, grade, or score an agent's or LLM's outputs; asks "how do I know if this is working," "is this any good," "set up an eval," or "did this prompt change make things worse"; needs a rubric for judging quality; wants to compare two prompts, models, or outputs; or wants to catch regressions before shipping a change. Also trigger when reviewing agent trajectories specifically — did the agent pick the right tool, the right arguments, the right sequence — not just the final output.
agent-quality
Use when evaluating a coding-agent product, gating a release, or when the user mentions evals, Agent Quality, 评测, 模块测试, 整体测试, trajectory, LLM judge, regression fixture, or independent verification of agent behavior. Use after an implementer claims done. Not for ordinary app unit tests with no agent loop.
evals-ops
Build and run evals for LLM and agent systems: golden datasets, LLM-as-a-judge with bias control, trajectory/step/outcome scoring, adversarial refuters, and CI regression gates. Triggers on: eval, evals, eval harness, golden dataset, golden set, LLM-as-a-judge, judge rubric, judge bias, judge calibration, Cohen kappa, agreement with human labels, pass@k, pass^k, trajectory eval, step-level eval, tool-call accuracy, regression gate, eval CI, blocking vs advisory eval, faithfulness score, DeepEval, Braintrust, Opik, Langfuse, AgentOps, did my prompt change make it worse, is my agent getting better.