measurement-experimentation-opslisted
Install: claude install-skill sergeyizmailov/knowledge-delta-skills
# Measurement & Experimentation Ops
The other skills act on measured differences; this decides whether a difference
is real or noise before they do.
## Pick the testing mode by the decision at stake
Three modes, different evidence bars — match to the cost of being wrong:
1. **Causal**: estimates incrementality (no design *proves* causality without
assumptions). Two sub-modes — a randomized experiment (one treatment vs a
non-overlapping control/holdout, pre-sized) is the strongest; quasi-experimental
causal estimation (GeoLift synthetic control, pre/post) is the fallback when you
can't randomize. Use for expensive, hard-to-reverse bets: offer, funnel, landing
page, "does this channel even lift sales." Cost: volume + discipline + often a
Meta rep.
2. **Screening** (directional): many concepts in one ad set / parallel ABO cells;
delivery is UNEQUAL by design, so a "winner" is a hypothesis, not a proof.
Use for high-throughput creative hunting where being fast beats being certain.
Never present a screen result as validated.
3. **Infrastructure** (isolate infra variance): hold the CREATIVE fixed, vary one
infra axis (domain / proxy cluster / account batch) across a balanced set to
attribute delivery/ban/CPM differences to infra, not creative. The grey
inversion of a normal test — see meta-grey-ops/06.
## Feasibility gate (grey reality — check BEFORE promising a clean test)
Causal measurement often isn't available on grey/small-account buys