ai-evals
FeaturedDesigns trustworthy LLM, agent, responsible-AI, and multimodal evaluations. Use when measuring quality, fairness, privacy, grounding, safety, or judge reliability.
Install
Quality Score: 91/100
Skill Content
Details
- Author
- vasilyu1983
- Repository
- vasilyu1983/AI-Agents-public
- Created
- 10 months ago
- Last Updated
- 3 days ago
- Language
- Python
- License
- MIT
Integrates with
Similar Skills
Semantically similar based on skill content — not just same category
llm-eval-testing
When the user wants to design, build, or operate evaluations (evals) for LLM-powered products — chatbots, RAG systems, agents, classification, summarization, structured output. Use when the user mentions "LLM evals," "evals," "RAG evaluation," "RAGAS," "DeepEval," "LangSmith," "LangFuse," "PromptLayer," "OpenAI evals," "judge model," "rubric eval," "LLM-as-judge," "Inspect AI," "AnthropicEvals," "Vertex evals," "Braintrust," or "regression tests for prompts." For AI testing tools see ai-augmented-testing. For chaos see chaos-engineering. For production monitoring see production-testing.
ai-evals
Use when testing an LLM-backed feature, prompt, tool loop or multi-step agent, where the same input can produce different outputs and a prompt or model change can regress behaviour with no code diff. The behaviour spec that precedes the prompt, scenario datasets including adversarial and degradation classes, deterministic assertions over OpenTelemetry traces, calibrated LLM-as-judge, CI gates with baselines, human-in-the-loop, and the production scoring loop that turns incidents into scenarios.
evals-ops
Build and run evals for LLM and agent systems: golden datasets, LLM-as-a-judge with bias control, trajectory/step/outcome scoring, adversarial refuters, and CI regression gates. Triggers on: eval, evals, eval harness, golden dataset, golden set, LLM-as-a-judge, judge rubric, judge bias, judge calibration, Cohen kappa, agreement with human labels, pass@k, pass^k, trajectory eval, step-level eval, tool-call accuracy, regression gate, eval CI, blocking vs advisory eval, faithfulness score, DeepEval, Braintrust, Opik, Langfuse, AgentOps, did my prompt change make it worse, is my agent getting better.