← ClaudeAtlas

agent-red-teaminglisted

Use when planning or documenting an authorized security assessment of an AI agent or multi-agent workflow with safe adversarial cases, synthetic identities, canaries, and evidence-based retesting.
sandbaseai/workbuddy-skill · ★ 2 · AI & Automation · score 81
Install: claude install-skill sandbaseai/workbuddy-skill
# Agent Red Teaming Find exploitable control failures without creating uncontrolled harm. Written authorization and rules of engagement are prerequisites for execution, not paperwork to complete afterward. Public reachability is not permission. ## Authorization and scope Before execution, record the named owner, authorization reference, exact targets, environment, models and configuration digest, tester identity, time window, in/out-of-scope systems and tenants, allowed techniques, rate/cost limits, stop conditions, emergency contact, cleanup owner, and evidence policy. If target-specific authorization or scope is missing, produce only a non-executable plan and threat model. Never probe a live target to infer scope. Use synthetic accounts, inert destinations, non-destructive tools, and benign canary values that grant no access. Confirm how to disable test tools, revoke credentials, restore fixtures, and report an unexpected effect before testing. ## Assessment contract Deliver: 1. Approved or explicitly pending rules of engagement. 2. A system and privilege map plus prioritized threat hypotheses. 3. A case plan with unique case IDs, target, environment, configuration digest, authorization reference, limits, stop conditions, safe oracle, and cleanup duty. 4. Execution records tied to approved cases and unique test IDs, including timestamps, observed limits, structured redacted evidence, and cleanup. 5. Deduplicated findings with reproduction, impact, likelihood