ai-evaluation
SolidUse when setting up quality assurance for AI features — defining evaluation criteria, measuring output quality, using AI-as-judge, monitoring production AI, detecting drift, and building user feedback loops
Install
Quality Score: 85/100
Skill Content
Details
- Author
- peterbamuhigire
- Repository
- peterbamuhigire/chwezi-dev-engine
- Created
- 7 months ago
- Last Updated
- 3 days ago
- Language
- HTML
- License
- MIT
Integrates with
Similar Skills
Semantically similar based on skill content — not just same category
ai-agent-observability-evaluation
Use when measuring, evaluating, replaying, evidencing, or tracking success for AI agent tasks, steps, traces, and outcomes.
evaluate-an-ai-system
Evaluate an AI feature against versioned cases, a baseline, failure slices, and human judgment before deciding to release it. Measure usefulness, safety, latency, and cost without hiding trade-offs in one score. Use when prompts, models, tools, or retrieval changes need repeatable evidence.
ai-evaluation-red-team
Design and execute repeatable evaluations and authorized red-team tests for prompts, models, agents, tools, and RAG systems across quality, safety, security, reliability, cost, and latency. Use before releasing or materially changing AI behavior.