ai-safetylisted
Install: claude install-skill aicodedecode/awesome-muse-skills
# AI Safety
AI safety is the practice of building systems that do what we intend, don't cause unintended harm,
and remain under meaningful human control. It spans technical work (alignment, evaluation,
robustness) and governance (policies, oversight, accountability) — this skill covers both at a
practitioner level.
## Overview
Think in layers. Technical: train models to follow intent faithfully (alignment), test for
dangerous capabilities and failure modes (evaluation), harden against adversaries (robustness), and
keep humans meaningfully in control of consequential actions (oversight). Organizational: risk
assessments before deployment, clear ownership, incident response plans, and honest communication
about limitations. Safety isn't a feature added at the end — it's a property of how you build.
## When to use
- Planning any AI deployment: assessing risks before building.
- Designing safeguards for agents with real-world actions.
- Evaluating models for dangerous capabilities or problematic behaviors.
- Establishing team practices: review processes, incident response, accountability.
## Core concepts
- **Risk assessment**: identify hazards (what could go wrong?), assess likelihood and severity, and
decide mitigations before deployment. Proportional to capability and autonomy — more powerful
systems need more rigor.
- **Alignment**: the technical problem of models pursuing intended goals — instruction following,
honesty, and not optimizing proxies in harmful