evidence-tieringlisted
Install: claude install-skill GadDev/skills
# Evidence Tiering
A rubric for one recurring problem: claims about AI (a benchmark number, a
"outperforms X" line, a viral capability demo) get repeated at face value far
more than their evidence quality justifies. This skill is the check that
runs before repeating any such claim.
## When to use this
- Before stating a benchmark result, performance comparison, or capability
claim about any AI model or tool
- Before summarizing a launch announcement, research paper, or news item
- When a claim arrived via a summary, thread, or secondhand write-up rather
than its original source
- Any time `ai-release-triage` or `ai-pulse` calls for a source-vetting
pass
## Credibility tiers (highest to lowest)
1. **Primary technical documentation** — model/system cards, technical
reports, papers with a methodology section, official API docs. This is
authoritative on what's *claimed*, not automatically on whether the
claim holds up under scrutiny.
2. **Independent reproductions** — a third party running the same
benchmark or task themselves, ideally with published methodology or
code. The strongest evidence a capability claim is real.
3. **Hands-on technical journalism** — outlets that tested the thing
themselves and describe specific results or failure cases, not just
relaying a press release.
4. **Official launch blog / press release** — useful for what changed
nominally (price, availability, headline numbers); treat performance
claims here as marketing