ai-redteamlisted
Install: claude install-skill Garyson26/Trident-SecOps-Skills
# AI Red Team
## Authorization Boundary
- Only target models, agents, and applications the user owns or has written authorization to test.
- Do not produce working malware, CSAM, weapons synthesis, or content that bypasses model safety for harmful real-world outcomes.
- Treat findings as defensive: every successful attack must ship with a detection and a mitigation.
## Attack Taxonomy
- `Direct injection`: instructions in user input override system policy.
- `Indirect injection`: hostile content in retrieved docs, web pages, emails, files, image alt-text, or tool output.
- `Jailbreak`: roleplay, hypothetical, encoding, obfuscation, multi-turn drift, persona pinning.
- `Tool abuse`: forcing browse, shell, code, or MCP tools to act outside scope (SSRF, file read, command exec).
- `Data exfiltration`: leaking system prompt, secrets, embeddings, training data, or per-user memory.
- `Agent hijack`: rewriting plans, looping, escalating permissions, calling unintended sub-agents.
- `RAG poisoning`: index pollution, ranking manipulation, embedding collisions.
- `Resource exhaustion`: token bombs, recursive tool calls, prompt-bomb DoS.
## Evaluation Workflow
1. Map the AI surface: model, system prompt, tools, retrievers, memory, output sinks, and downstream callers.
2. Build a threat model with abuse cases per surface.
3. Generate a versioned attack corpus (`attacks.jsonl`) with `{id, technique, payload, expected_block, severity}`.
4. Run the harness; record `{pass, fail, partia