experiment-analysislisted
Install: claude install-skill mcorbett51090/RavenClaude
# Skill: experiment-analysis
> **Invoked by:** `applied-statistician` (primary). Pairs with [`../power-and-sample-size/SKILL.md`](../power-and-sample-size/SKILL.md) (run first, at design time) and [`../choose-statistical-test/SKILL.md`](../choose-statistical-test/SKILL.md) (names the underlying test).
>
> **When to invoke:** "is this A/B winner real?"; "the experiment finished — what does it say?"; "can we ship variant B?"
>
> **Output:** a verdict (ship / don't ship / inconclusive) backed by the primary-metric effect size + CI, the guardrail check, the multiplicity correction, and an explicit pitfall screen.
## Procedure
1. **Recover the analysis plan.** Was there a pre-registered primary metric, test, and stopping rule ([`../../templates/analysis-plan.md`](../../templates/analysis-plan.md))? If analysis decisions were made *after* seeing data, flag the p-hacking exposure (pitfall #1) and treat findings as exploratory.
2. **Check the stopping rule.** Was the test stopped at the planned sample, or stopped early when it "looked significant"? Early stopping on a fixed-horizon test = peeking (pitfall #3) → the nominal p-value is not valid; downgrade confidence.
3. **Run the primary-metric test** (via `choose-statistical-test`). Report the **effect size + confidence interval** as the headline — the p-value is secondary (pitfall #6). "Significant but the CI includes trivially small effects" is not a ship signal.
4. **Check guardrail metrics.** A primary-metric win that degrades