ai-security

Featured

Use when attacking an AI/ML system or model — prompt injection & jailbreaks (Crescendo, Skeleton Key, Best-of-N), RAG/vector poisoning, agentic/MCP exploitation (CVE-2025-54136), ML supply-chain RCE (pickle CVE-2025-32434), model extraction / membership inference / adversarial suffixes (GCG)

AI & Automation 382 stars 66 forks Updated 5 days ago MIT

Install

View on GitHub

Quality Score: 95/100

Stars 20%
86
Recency 20%
100
Frontmatter 20%
70
Documentation 15%
100
Issue Health 10%
80
License 10%
100
Description 5%
100

Skill Content

# AI/ML Security ## When to Activate - Red-teaming an LLM/chatbot/copilot for direct & indirect prompt injection and multi-turn jailbreaks. - Testing a RAG pipeline for document/embedding poisoning, embedding inversion, and cross-tenant retrieval leakage. - Auditing an AI agent / MCP server for tool poisoning, excessive agency, and command injection (RCE). - Scanning a model artifact (HuggingFace, `.pt/.pkl/.bin/.gguf`) for deserialization payloads before loading it. - Assessing a model API for extraction/distillation, membership inference, and adversarial-suffix robustness. - Mapping findings to OWASP LLM Top-10 (2025) + MITRE ATLAS for a report. ## Technique Map | Technique | ATT&CK | CWE | Reference | Script | |-----------|--------|-----|-----------|--------| | Direct prompt injection / system-prompt leak (LLM01/LLM07) | T1059.006, T1606 | CWE-1427 | references/prompt-injection-jailbreak.md | scripts/promptinject_harness.py | | Multi-turn jailbreak: Crescendo / Skeleton Key | T1059.006 | CWE-1427 | references/prompt-injection-jailbreak.md | scripts/promptinject_harness.py | | Best-of-N / many-shot / token-smuggling jailbreak | T1059.006, T1027 | CWE-1427 | references/prompt-injection-jailbreak.md | scripts/promptinject_harness.py | | Indirect injection via ingested content (EchoLeak CVE-2025-32711) | T1190, T1059.006 | CWE-74 | references/prompt-injection-jailbreak.md | scripts/promptinject_harness.py | | RAG knowledge-base poisoning (PoisonedRAG, 5 docs) | T1195, T156...

Details

Author
hypnguyen1209
Repository
hypnguyen1209/offensive-claude
Created
4 months ago
Last Updated
5 days ago
Language
Python
License
MIT

Integrates with

Bundled in these plugins

Similar Skills

Semantically similar based on skill content — not just same category