ab-testing

Featured

Use when designing or analyzing a controlled experiment — falsifiable hypothesis, sample size from an MDE, reading significance/CI/power, CUPED, or rescuing tests that won't go significant. NOT recurring metric tracking (that is `analytics`), NOT north-star/KPI trees (that is `kpi-framework`), NOT projecting metrics forward (that is `forecasting`).

Testing & QA 116 stars 9 forks Updated 2 days ago MIT

Install

View on GitHub

Quality Score: 92/100

Stars 20%
69
Recency 20%
100
Frontmatter 20%
70
Documentation 15%
100
Issue Health 10%
80
License 10%
100
Description 5%
100

Skill Content

# A/B testing — design and read a defensible experiment An experiment without a pre-committed sample size and a single primary metric is not an experiment. It is a dashboard you stare at until it tells you what you wanted to hear. The discipline lives almost entirely *before* traffic ships: a falsifiable hypothesis, one primary metric, a sample size derived from the smallest effect worth detecting, and a stop rule you cannot renegotiate at 2pm on day four. ## Pre-test checklist — every line true before any traffic Each one is a place experiments die silently. - [ ] A **falsifiable hypothesis** — names the change, the direction, and the metric it moves. - [ ] Exactly **ONE primary metric**. More than one primary = multiple comparisons = inflated false positives. - [ ] **Guardrail metrics** — what you refuse to harm (latency, refunds, unsubscribes) even for a win. - [ ] The **randomization unit = the analysis unit** (usually the user). Mixing them is pseudoreplication. - [ ] An **MDE** — the smallest lift that would change a decision. Not "any difference." - [ ] A **computed sample size** and the **duration** it implies at your real daily eligible traffic. - [ ] A **fixed stop rule** — a date or an n you commit to before launch. No "we'll see how it looks." ## Step 1 — Hypothesis and metrics State a null you can reject. "The new checkout button changes purchase conversion" with H0: conversion equal across arms, H1: it differs. Vague aspirations ("improve the funnel") have...

Details

Author
ericrisco
Repository
ericrisco/rsc-harness
Created
3 months ago
Last Updated
2 days ago
Language
JavaScript
License
MIT

Similar Skills

Semantically similar based on skill content — not just same category

AI & Automation Solid

experiment-analysis

Designs and reads A/B tests end to end — a sharp hypothesis, a primary metric plus guardrails, sample size and power, and a frequentist read (significance, confidence interval, practical effect) leading to a clear ship/kill call. Use when you say "design this A/B test," "is this result significant," "how many users do I need," "can we ship this," or "did the experiment win?"

6 Updated 3 weeks ago
Sidsaladi9
Testing & QA Listed

ab-testing

When the user wants to plan, design, or implement an A/B test or experiment, or build a growth experimentation program. Also use when the user mentions "A/B test," "split test," "experiment," "test this change," "variant copy," "multivariate test," "hypothesis," "should I test this," "which version is better," "test two versions," "statistical significance," "how long should I run this test," "growth experiments," "experiment velocity," "experiment backlog," "ICE score," "experimentation program," or "experiment playbook." Use this whenever someone is comparing two approaches and wants to measure which performs better, or when they want to build a systematic experimentation practice. For tracking implementation, see analytics. For page-level conversion optimization, see cro.

7 Updated yesterday
TommyBez
Testing & QA Listed

ab-testing

When the user wants to plan, design, or implement an A/B test or experiment, or build a growth experimentation program. Also use when the user mentions "A/B test," "split test," "experiment," "test this change," "variant copy," "multivariate test," "hypothesis," "should I test this," "which version is better," "test two versions," "statistical significance," "how long should I run this test," "growth experiments," "experiment velocity," "experiment backlog," "ICE score," "experimentation program," or "experiment playbook." Use this whenever someone is comparing two approaches and wants to measure which performs better, or when they want to build a systematic experimentation practice. For tracking implementation, see analytics. For page-level conversion optimization, see cro.

0 Updated 3 weeks ago
gaznilmuzammil12-prog