sycophancy-guard
SolidAlways-on background filter preventing sycophantic belief drift. Based on Chandra et al. 2026 proving even ideal Bayesian users spiral from sycophantic AI, including factual sycophants cherry-picking true info. ALWAYS active on every response like anti-ai-slop. Explicit triggers include "drift check", "am I spiraling", "check my assumptions", "red team this", "belief audit", "sycophancy check", "are you just agreeing with me", "challenge this", "what am I wrong about". Monitors confirmation bias, selective evidence, escalating agreement without new evidence, unverifiable strategic claims without pushback. Auto-triggers a Decision Council on detected drift. Logs belief positions to your knowledge graph. The immune system against RLHF sycophancy over extended single-user interactions.
Install
Quality Score: 79/100
Skill Content
Details
- Author
- 0xUrsanomics
- Repository
- 0xUrsanomics/utopia-os
- Created
- 5 days ago
- Last Updated
- today
- Language
- Python
- License
- MIT
Similar Skills
Semantically similar based on skill content — not just same category
confirmation-bias
Activate when: user says 'we keep finding evidence that supports our view,' 'the team is all aligned on this,' 'I've done the research and it checks out,' or a decision moves forward with only supporting evidence cited. Do NOT activate when: context is explicit advocacy (legal brief, pitch deck) where one-sided argument is the design; or stakes are too low to justify structured disconfirmation. More: deciqai.com/s/confirmation-bias
anti-sycophancy
Strip validating, hedging, and flattering language from responses. Disagree when warranted. State the answer.
decision-council
Forces 5 AI advisor personas to argue about a decision, anonymously peer-review each other, and synthesize a verdict with anti-false-consensus guardrails. Auto-detects high-stakes decisions from context. Trigger on: pricing, deal terms, partnerships, market entry, investments, exit strategies, go/no-go calls, trading decisions, DeFi/protocol entry, market entry/exit, futures positions, training program changes, competition prep, injury decisions, career moves, big purchases, financial planning, macro economic events (central-bank decisions, FX, inflation), geopolitical developments affecting markets, major AI/Web3/tech releases that change capability or competitive landscape, or any question where being wrong costs money, health, or reputation. Also trigger on "run council", "stress test", "what am I missing", "devil's advocate", "sanity check", or when the user leans toward an answer already. Do NOT trigger for brainstorming, writing, code, factual lookups, routine daily choices, or low-stakes ops.