ai-agent-redteam

Featured

Use when red-teaming an agentic AI / LLM application — indirect & zero-click prompt injection, MCP tool poisoning, persistent memory poisoning, excessive-agency tool abuse, multi-turn jailbreaks, PyRIT/Garak/Promptfoo harnesses

AI & Automation 382 stars 66 forks Updated 5 days ago MIT

Install

View on GitHub

Quality Score: 95/100

Stars 20%
86
Recency 20%
100
Frontmatter 20%
70
Documentation 15%
100
Issue Health 10%
80
License 10%
100
Description 5%
100

Skill Content

# AI Agent Red Teaming Offensive testing of **autonomous LLM agents** — systems that combine model reasoning with tools, memory, retrieval, and multi-step planning. This is distinct from model-level testing (see `ai-security`): the attack surface here is the *agentic pipeline* — untrusted data channels, tool/MCP integrations, persistent memory, and delegated authority. Assumes authorized engagement. ## When to Activate - Pentesting an LLM agent with tool/function-calling, an MCP client, or a code interpreter - Testing RAG / email / browser assistants for indirect or zero-click prompt injection - Auditing MCP server integrations for tool poisoning, rug-pull, or line-jumping - Assessing persistent memory / long-term context for poisoning and belief drift - Evaluating excessive agency: confused-deputy, SSRF/RCE-via-tool, over-privileged actions - Running automated jailbreak campaigns (PAIR/TAP/Crescendo/Best-of-N) and measuring ASR - Standing up a repeatable PyRIT/Garak/Promptfoo harness mapped to OWASP Agentic Top 10 / ATLAS ## Technique Map | Technique | ATT&CK | CWE | Reference | Script | |-----------|--------|-----|-----------|--------| | Indirect / zero-click prompt injection (EchoLeak-class) | T1566.002 / AML.T0051.001 | CWE-1427 | references/indirect-prompt-injection.md | scripts/indirect_injection_forge.py | | RAG corpus poisoning & markdown/image exfiltration | T1567 / AML.T0070 | CWE-1426 | references/indirect-prompt-injection.md | scripts/indirect_injection_forge...

Details

Author
hypnguyen1209
Repository
hypnguyen1209/offensive-claude
Created
4 months ago
Last Updated
5 days ago
Language
Python
License
MIT

Integrates with

Bundled in these plugins

Similar Skills

Semantically similar based on skill content — not just same category