← ClaudeAtlas

ai-safetylisted

Practice AI safety across the lifecycle — risk assessment, alignment concepts, evaluation, deployment safeguards, and governance. Use when building or deploying AI systems responsibly.
aicodedecode/awesome-muse-skills · ★ 0 · AI & Automation · score 75
Install: claude install-skill aicodedecode/awesome-muse-skills
# AI Safety AI safety is the practice of building systems that do what we intend, don't cause unintended harm, and remain under meaningful human control. It spans technical work (alignment, evaluation, robustness) and governance (policies, oversight, accountability) — this skill covers both at a practitioner level. ## Overview Think in layers. Technical: train models to follow intent faithfully (alignment), test for dangerous capabilities and failure modes (evaluation), harden against adversaries (robustness), and keep humans meaningfully in control of consequential actions (oversight). Organizational: risk assessments before deployment, clear ownership, incident response plans, and honest communication about limitations. Safety isn't a feature added at the end — it's a property of how you build. ## When to use - Planning any AI deployment: assessing risks before building. - Designing safeguards for agents with real-world actions. - Evaluating models for dangerous capabilities or problematic behaviors. - Establishing team practices: review processes, incident response, accountability. ## Core concepts - **Risk assessment**: identify hazards (what could go wrong?), assess likelihood and severity, and decide mitigations before deployment. Proportional to capability and autonomy — more powerful systems need more rigor. - **Alignment**: the technical problem of models pursuing intended goals — instruction following, honesty, and not optimizing proxies in harmful