← ClaudeAtlas

framework-red-teamlisted

Use when an agent-control framework needs adversarial testing for fake productivity, loops, recursion, evidence gaps, or vague wording.
ihabkhaled/AI-Psychiatry · ★ 3 · AI & Automation · score 74
Install: claude install-skill ihabkhaled/AI-Psychiatry
# Framework Red Team ## Core principle Semantic compliance is stronger than literal compliance. Use observable evidence and causal history; never collect or demand private chain-of-thought. The goal is correct, safe delivery with sufficient reasoning, followed by termination. ## Procedure 1. Lock the primary objective, mandatory requirements, Definition of Done, and current evidence before changing any classification or budget. 2. Identify the specific observable signal. Do not infer a violation merely from time, token use, discomfort, or a label. 3. Act as an adversarial executor, identify concrete exploits, then switch roles and patch each exploit. Compare the current outcome with the previous outcome and preserve causal history across renames, handoffs, replans, and compression. 4. Produce the compact record: `exploit, observable trace, violated intent, patch, regression`. Mark unsupported claims `not confirmed`; do not convert confidence into proof. 5. Apply one bounded corrective action with an explicit attempt or time limit and exit condition. If a default limit prevents required correctness evidence, use `$executive-override` with `reason, evidence, exact limit, narrow scope, exit condition` rather than resetting a counter. 6. Revalidate only the affected requirement or policy. Report `a closed exploit list with a regression for each patch` and return to productive work. ## Repository runtime Apply this procedure inside the installed `.ai/` framework. Record obse