agent-safety

Featured

Use when bounding an LLM agent that already runs — scoping its task domain, gating tools to least privilege, defending against prompt injection in untrusted web/email/RAG text, requiring human approval on irreversible actions, capping runtime and cost, or triaging what it already did. NOT building the loop, tools, or RAG (that is `building-agents`).

AI & Automation 116 stars 9 forks Updated 2 days ago MIT

Install

View on GitHub

Quality Score: 92/100

Stars 20%
69
Recency 20%
100
Frontmatter 20%
70
Documentation 15%
100
Issue Health 10%
80
License 10%
100
Description 5%
100

Skill Content

# Agent safety You are the security review for an agent's **agency**, not for its code. The loop works, tools are wired, memory persists — your job is to make that autonomy *bounded*. If you want to review ordinary endpoints, auth, or secrets handling, that is `../secure-coding/SKILL.md` — this skill is the *Agentic* Top 10, the risks that exist only because a model has tools and autonomy. If the loop or tools do not exist yet, that is `../building-agents/SKILL.md`. You arrive *after* both. `references/threat-model.md` carries the OWASP Agentic Top 10 2026 risks mapped to the controls below, the pre-ship guardrail checklist, and the incident-triage flow for "the agent did X" — open it when you are reviewing before ship or reconstructing an incident. ## The ownership split Agent security splits into four layers — **Model · Harness · Tools · Environment**. The model provider owns only the Model layer (alignment, refusals). Everything else is yours: the Harness (loop, memory, context assembly), the Tools (what the agent can *do*), and the Environment (creds, network, blast radius). Do not outsource a layer you own to "the model is aligned." Three excesses cause almost every agentic incident. Cut all three: - **Excessive functionality** — tools the task never needs. - **Excessive permissions** — broader scopes/creds than the tool needs. - **Excessive autonomy** — acting without checking back when it should. The operating principle is **least agency**: autonomy is earned pe...

Details

Author
ericrisco
Repository
ericrisco/rsc-harness
Created
3 months ago
Last Updated
2 days ago
Language
JavaScript
License
MIT

Similar Skills

Semantically similar based on skill content — not just same category

AI & Automation Listed

agent-threat-model

Threat-model an AI agent deployment against the lethal trifecta — private data, untrusted content, and an exfiltration vector — producing a per-capability matrix, a named architectural fix for every unsafe path, and a pre-launch checklist. Use before shipping an agent, when reviewing MCP server or tool permissions, when the user asks about prompt injection or data exfiltration risk, or when deciding whether an agent's capability surface is safe to expose.

2 Updated yesterday
sananthanarayan
AI & Automation Listed

ai-agent-guardrails

Apply safety controls when an LLM agent has authority to act on real systems. Covers blast-radius classification, dry-run-first patterns, out-of-band approval gates, scope locking, idempotency, kill switches, and rollback strategies. Invoke when designing an autonomous agent, when granting an LLM write access to production, or after an agent makes an unexpected change.

19 Updated 1 months ago
GoldenWing-360
AI & Automation Listed

agent-builder

Build, implement, review, or harden a SINGLE production-grade agent — its contract, schemas, tools and permission tiers, durable state, guardrails, traces, and evals. Use when the user wants to create one agent (not a multi-agent product), implement an agent runner, add tools/memory/evals to an existing agent, or review whether one agent is production-ready. For multi-agent products, orchestration, or framework selection, use the agentic-product-architect skill instead. The full operational standard this skill applies is AGENT_STANDARD.md (bundled with this skill); copy-paste artifacts are in templates/.

14 Updated 1 months ago
Moai-Team-LLC