deep-researchlisted
Install: claude install-skill akuroglo/claude-code-setup
# Deep Research
One entrance, two collection modes, the same quality bar on both.
Three things below are measured, not assumed. They are why this pipeline is shaped the way it is:
- **Self-reported confidence is decorative.** Calibration error above 0.377; on a 1–5 scale RLHF models answer "5" in 96.5% of cases. Never ask the model how sure it is — count independent confirmations instead.
- **Agents that debate each other collapse into agreement.** Three independent 2025 studies found multi-agent debate degrading into sycophancy and consensus collapse, often below single-agent performance. So: parallel independent agents, reconciled by source quality, never by vote.
- **Guessing is rational unless you make it costly.** Most benchmarks penalise "I don't know" exactly as much as a wrong answer, so models learn to guess. Two models on one task: 52% abstention / 26% error versus 1% abstention / 75% error.
**The overriding rule: an admitted gap beats a plausible invention.** Every phase has a way to say "not established". Use it.
---
## Phase 0 — Contract
Before any search, ask the user three questions. Do not skip them and do not guess the answers.
1. **Depth.** Quick scan (one round per aspect) · standard (until saturation) · exhaustive (saturation plus an adversarial pass).
2. **Source classes.** Docs and articles only, or also native material — video transcripts, forum threads, practitioner posts, incident reports? Native material is where practice lives; it is off by