← ClaudeAtlas

evidence-tieringlisted

Shared credibility framework for weighing any AI-related claim — a benchmark result, a model/tool announcement, a research paper, a demo, or a viral thread — by the quality of its source rather than how confidently it's stated. Use this any time you're about to repeat a claim about AI capabilities, performance, or news without having checked where it actually came from. Other skills (ai-release-triage, ai-pulse) call into this for their source-vetting step, but use it standalone too whenever a claim about AI needs a credibility check before you pass it on.
GadDev/skills · ★ 0 · Code & Development · score 72
Install: claude install-skill GadDev/skills
# Evidence Tiering A rubric for one recurring problem: claims about AI (a benchmark number, a "outperforms X" line, a viral capability demo) get repeated at face value far more than their evidence quality justifies. This skill is the check that runs before repeating any such claim. ## When to use this - Before stating a benchmark result, performance comparison, or capability claim about any AI model or tool - Before summarizing a launch announcement, research paper, or news item - When a claim arrived via a summary, thread, or secondhand write-up rather than its original source - Any time `ai-release-triage` or `ai-pulse` calls for a source-vetting pass ## Credibility tiers (highest to lowest) 1. **Primary technical documentation** — model/system cards, technical reports, papers with a methodology section, official API docs. This is authoritative on what's *claimed*, not automatically on whether the claim holds up under scrutiny. 2. **Independent reproductions** — a third party running the same benchmark or task themselves, ideally with published methodology or code. The strongest evidence a capability claim is real. 3. **Hands-on technical journalism** — outlets that tested the thing themselves and describe specific results or failure cases, not just relaying a press release. 4. **Official launch blog / press release** — useful for what changed nominally (price, availability, headline numbers); treat performance claims here as marketing