AI/LLM defensive security reference: prompt-injection defense, OWASP LLM Top 10 defensive mapping, MCP and agentic tool-call hardening, training-data poisoning detection, model-output validation and guardrails, MITRE ATLAS defensive correlation, and NIST AI RMF governance. Agent-extending skill that amplifies backend, security, and AI-application engineering with production-grade defensive patterns for LLM-backed systems. NOT for: offensive techniques (jailbreak authoring, attack-payload crafting, red-team exploitation), model training or fine-tuning methodology, prompt optimization for capability, web-app OWASP Top 10 (see moai-ref-owasp-checklist), or general API design (see moai-ref-api-patterns).
AI & Automation 5 stars
4 forksUpdated 2 weeks agoApache-2.0
# LLM / AI Defensive Security Reference
Defensive practitioner reference for hardening LLM-backed applications and agents.
Every section is framed as defense, hardening, detection, or verification — it
describes the misconfiguration, how to detect it, and how to prevent it, never how
to exploit it. Cross-domain web-app vulnerabilities live in
`moai-ref-owasp-checklist`; API design lives in `moai-ref-api-patterns`.
## Target Use
Apply when reviewing or building an LLM-backed system — a chat product, a
retrieval-augmented application, an autonomous agent, or an MCP server. The
material assumes an untrusted-input threat model: any text the model reads (user
turns, retrieved documents, tool results, file contents) may carry adversarial
instructions, and any text the model emits may be acted on downstream.
## Trust Boundaries in an LLM System
The core defensive insight: an LLM does not distinguish "data" from "instructions"
the way a parser does. Treat every text channel that reaches the model as a
boundary where adversarial instructions can enter.
| Channel | Entry risk | Primary defense |
|---------|-----------|-----------------|
| End-user prompt | Direct prompt injection | Instruction-hierarchy enforcement, input screening |
| Retrieved documents (RAG) | Indirect prompt injection | Provenance tagging, content isolation, retrieval allowlist |
| Tool / function results | Injected instructions in tool output | Treat tool output as untrusted data, re-validate before re-promp...