agent-eval

Featured

Use when measuring whether an LLM or agent system actually got better and gating merges on it: golden sets, fixing an inflated LLM-as-judge, scoring RAG (faithfulness, contextual recall) or agent trajectories (tool correctness, completion), or picking an eval framework. NOT building the agent loop, tools or RAG plumbing (that is `building-agents`).

AI & Automation 116 stars 9 forks Updated 2 days ago MIT

Install

View on GitHub

Quality Score: 92/100

Stars 20%
69
Recency 20%
100
Frontmatter 20%
70
Documentation 15%
100
Issue Health 10%
80
License 10%
100
Description 5%
100

Skill Content

# Measure agent quality you can defend and gate on Turn "the agent feels better" into a number you can put in a PR check. You own the eval dataset, the scorer mix, the LLM-as-judge calibration, and the block-on-regression CI gate — framework-neutral, provider-neutral. ## Do NOT use — route instead | The ask | Route to | Why it is not this skill | | --- | --- | --- | | Build the agent loop, tools, RAG plumbing | `building-agents` | It builds the system; you score it. They cross-link. | | "Make the answers shorter / rewrite the prompt" | `prompt-engineering` | Evals say it is worse; that skill changes the words. You never edit the prompt. | | pytest/jest on deterministic functions | `testing-py` / `testing-web` | Assert-equals on pure code, not stochastic outputs scored by a judge. | | Dashboards / tracing of live production traffic | `observability` | Online monitoring; you are offline + pre-merge. | | Red-team, jailbreak, prompt injection | `agent-safety` | Adversarial coverage, not quality measurement. | | Per-token cost budgets and accounting | `cost-tracking` | You report cost-per-task as one metric; the discipline lives there. | | A/B stats on product/funnel metrics | `ab-testing` | Web experiments, not offline model comparison on a fixed set. | ## The eval anatomy Every framework instantiates the same five-stage pipeline. Learn it once; the tool is a detail. ```text dataset ──▶ runner ──▶ scorers ──▶ metrics ──▶ gate (JSONL (calls the (det / judge (aggregate ...

Details

Author
ericrisco
Repository
ericrisco/rsc-harness
Created
3 months ago
Last Updated
2 days ago
Language
JavaScript
License
MIT

Integrates with

Similar Skills

Semantically similar based on skill content — not just same category

AI & Automation Listed

agent-eval

Designs and runs evaluations for LLM or agent outputs — builds rubrics, sets up LLM-as-judge scoring, creates regression test sets, and reports pass rates with concrete failure examples. Use this skill whenever the user wants to evaluate, test, grade, or score an agent's or LLM's outputs; asks "how do I know if this is working," "is this any good," "set up an eval," or "did this prompt change make things worse"; needs a rubric for judging quality; wants to compare two prompts, models, or outputs; or wants to catch regressions before shipping a change. Also trigger when reviewing agent trajectories specifically — did the agent pick the right tool, the right arguments, the right sequence — not just the final output.

0 Updated today
Hefrock
AI & Automation Listed

agent-quality

Use when evaluating a coding-agent product, gating a release, or when the user mentions evals, Agent Quality, 评测, 模块测试, 整体测试, trajectory, LLM judge, regression fixture, or independent verification of agent behavior. Use after an implementer claims done. Not for ordinary app unit tests with no agent loop.

1 Updated 3 weeks ago
xin-yi33
AI & Automation Solid

evals-ops

Build and run evals for LLM and agent systems: golden datasets, LLM-as-a-judge with bias control, trajectory/step/outcome scoring, adversarial refuters, and CI regression gates. Triggers on: eval, evals, eval harness, golden dataset, golden set, LLM-as-a-judge, judge rubric, judge bias, judge calibration, Cohen kappa, agreement with human labels, pass@k, pass^k, trajectory eval, step-level eval, tool-call accuracy, regression gate, eval CI, blocking vs advisory eval, faithfulness score, DeepEval, Braintrust, Opik, Langfuse, AgentOps, did my prompt change make it worse, is my agent getting better.

42 Updated 1 weeks ago
0xDarkMatter