← ClaudeAtlas

ai-evaluation-red-teamlisted

Design and execute repeatable evaluations and authorized red-team tests for prompts, models, agents, tools, and RAG systems across quality, safety, security, reliability, cost, and latency. Use before releasing or materially changing AI behavior.
Yaz-inc/yazinc-ai-toolkit · ★ 0 · AI & Automation · score 60
Install: claude install-skill Yaz-inc/yazinc-ai-toolkit
# AI Evaluation Red Team ## Objective Replace subjective AI approval with versioned datasets, measurable assertions, adversarial scenarios, and explicit release thresholds. ## Workflow 1. Define the AI task, users, model and tool boundaries, failure impact, prohibited behavior, data sensitivity, and release criteria. 2. Build representative and edge-case datasets with expected outcomes, rubrics, deterministic assertions, and human review samples. 3. Measure task quality, groundedness, refusal behavior, format compliance, latency, cost, and consistency across candidate configurations. 4. Create authorized adversarial cases for prompt injection, data extraction, tool misuse, privilege escalation, unsafe output handling, and denial of service. 5. Separate model, prompt, retrieval, tool, policy, and infrastructure failures so remediation targets the correct layer. 6. Run evaluations with fixed versions and controlled randomness, then compare against an approved baseline. 7. Publish failed cases, severity, reproducibility, residual risk, and the release decision without exposing sensitive prompts or data. ## Safety and authorization - Use synthetic or explicitly approved datasets and providers. Treat sending prompts to a provider as data publication. - Do not probe third-party systems, accounts, or models beyond authorized terms and scope. - Do not use evaluator-model scores as the only evidence for high-impact safety or compliance claims. ## Evidence and completion - Reco