agent-red-teaming

Solid

Plan, execute, document, and retest authorized security assessments of AI agents and multi-agent workflows using safe adversarial cases, synthetic identities, canaries, and evidence-based findings. Use when defining red-team rules of engagement, assessing prompt injection or excessive agency, testing tool and identity boundaries, evaluating memory or cross-agent attacks, scoring a campaign, or verifying remediation in an approved environment.

AI & Automation 161 stars 32 forks Updated 1 weeks ago MIT

Install

View on GitHub

Quality Score: 84/100

Stars 20%
74
Recency 20%
90
Frontmatter 20%
70
Documentation 15%
100
Issue Health 10%
50
License 10%
100
Description 5%
100

Skill Content

# Agent Red Teaming Find exploitable control failures without creating uncontrolled harm. Treat written authorization and rules of engagement as prerequisites for execution, not paperwork to complete afterward. ## Inputs Collect: - Named target owner and explicit authorization for the exact systems to be tested - Target identifiers, environment, accounts, endpoints, models, versions, and a reproducible configuration digest - Start/end time, tester identities, source addresses, rate and cost limits, and emergency contact - In-scope objectives and out-of-scope systems, tenants, data, techniques, and effects - Agent architecture, tools, privileges, memory, retrieval, handoffs, identities, and external integrations - Protected assets, security requirements, prior incidents, existing controls, and expected benign tasks - Approved synthetic data, canary values, test destinations, cleanup plan, and evidence-handling rules If target-specific authorization or scope is missing, stop at a non-executable assessment plan. Do not probe a live target to infer scope. ## Output contract Deliver: 1. Signed-off or explicitly pending rules of engagement with scope, constraints, stop conditions, contacts, and cleanup duties 2. A system and privilege map plus prioritized threat hypotheses 3. A machine-readable, owner-approved campaign plan with unique case IDs, targets, environment, configuration digest, tester subjects, authorization reference, time window, limits, stop conditions, cleanu...

Details

Author
seb1n
Repository
seb1n/awesome-ai-agent-skills
Created
6 months ago
Last Updated
1 weeks ago
Language
Python
License
MIT

Integrates with

Similar Skills

Semantically similar based on skill content — not just same category