← ClaudeAtlas

experiment-analysislisted

Analyze a completed A/B test or experiment defensibly — check it against the pre-registered plan, run the primary-metric test, report effect size + CI (not just p), check guardrail metrics, apply a multiple-comparison correction across metrics/segments, and screen for the peeking/p-hacking pitfalls before declaring a winner. Used by `applied-statistician` (primary).
mcorbett51090/RavenClaude · ★ 7 · AI & Automation · score 65
Install: claude install-skill mcorbett51090/RavenClaude
# Skill: experiment-analysis > **Invoked by:** `applied-statistician` (primary). Pairs with [`../power-and-sample-size/SKILL.md`](../power-and-sample-size/SKILL.md) (run first, at design time) and [`../choose-statistical-test/SKILL.md`](../choose-statistical-test/SKILL.md) (names the underlying test). > > **When to invoke:** "is this A/B winner real?"; "the experiment finished — what does it say?"; "can we ship variant B?" > > **Output:** a verdict (ship / don't ship / inconclusive) backed by the primary-metric effect size + CI, the guardrail check, the multiplicity correction, and an explicit pitfall screen. ## Procedure 1. **Recover the analysis plan.** Was there a pre-registered primary metric, test, and stopping rule ([`../../templates/analysis-plan.md`](../../templates/analysis-plan.md))? If analysis decisions were made *after* seeing data, flag the p-hacking exposure (pitfall #1) and treat findings as exploratory. 2. **Check the stopping rule.** Was the test stopped at the planned sample, or stopped early when it "looked significant"? Early stopping on a fixed-horizon test = peeking (pitfall #3) → the nominal p-value is not valid; downgrade confidence. 3. **Run the primary-metric test** (via `choose-statistical-test`). Report the **effect size + confidence interval** as the headline — the p-value is secondary (pitfall #6). "Significant but the CI includes trivially small effects" is not a ship signal. 4. **Check guardrail metrics.** A primary-metric win that degrades