ai-agent-safety-and-red-team
SolidUse when red-teaming agent tool and data perimeters for indirect prompt injection, action escalation, tenant exfiltration, unsafe tool chains, self-modification, or containment regressions.
Install
Quality Score: 85/100
Skill Content
Details
- Author
- peterbamuhigire
- Repository
- peterbamuhigire/chwezi-dev-engine
- Created
- 7 months ago
- Last Updated
- 3 days ago
- Language
- HTML
- License
- MIT
Integrates with
Similar Skills
Semantically similar based on skill content — not just same category
agent-red-teaming
Plan, execute, document, and retest authorized security assessments of AI agents and multi-agent workflows using safe adversarial cases, synthetic identities, canaries, and evidence-based findings. Use when defining red-team rules of engagement, assessing prompt injection or excessive agency, testing tool and identity boundaries, evaluating memory or cross-agent attacks, scoring a campaign, or verifying remediation in an approved environment.
agent-red-teaming
Use when planning or documenting an authorized security assessment of an AI agent or multi-agent workflow with safe adversarial cases, synthetic identities, canaries, and evidence-based retesting.
ai-evaluation-red-team
Design and execute repeatable evaluations and authorized red-team tests for prompts, models, agents, tools, and RAG systems across quality, safety, security, reliability, cost, and latency. Use before releasing or materially changing AI behavior.