agent-evaluation-framework-builder
SolidDesigns an eval suite for an LLM agent or pipeline including success metrics, trajectory scoring, LLM-as-judge setup, and regression test cases.
Install
Quality Score: 85/100
Skill Content
Details
- Author
- Notysoty
- Repository
- Notysoty/openagentskills
- Created
- 5 months ago
- Last Updated
- 6 days ago
- Language
- JavaScript
- License
- MIT
Similar Skills
Semantically similar based on skill content — not just same category
agent-evaluation-rubrics
This skill should be used when building agent evaluation systems: deterministic checks, regression suites, multi-dimensional rubrics, quality gates, production monitoring, baseline comparison, and outcome measurement for agent pipelines.
evaluation
Build evaluation frameworks for agent systems. Use when testing agent performance, validating context engineering choices, or measuring improvements over time.
agent-ux-eval
Evaluate any AI agent experience against the Agent UX Evaluation Framework. Use this skill whenever asked to evaluate, assess, audit, score, or review an AI agent, chatbot, copilot, or agentic workflow. Also trigger when the user mentions "agent evaluation framework", "experience evaluation", "agent UX audit", or "score this agent". Supports screenshot-based, Chrome plugin, and URL-fetch evaluation methods.