model-bakeofflisted
Install: claude install-skill ityaadiii/skills-that-say-i-dont-know
# Comparing models without manufacturing a winner
A sweep of 5 models across 4 workloads and 5 metrics is 100 simultaneous tests. At
p < 0.05 that produces about 5 "significant" results from pure noise, and those are the
ones that end up in the summary.
## The refusal
**Do not name a winner that does not survive correction.** If nothing survives, the
finding is "these models are indistinguishable on this evidence", and that is a real,
useful, publishable result.
Also refuse to compare unpaired when the models saw the same items. Throwing the
pairing away discards most of the available power.
## Procedure
1. **Same items, every model.** If they saw different items, stop. That is a different
and much weaker study, and it should be labelled as one.
2. **Count the comparisons before you run them.** Models times workloads times metrics.
Write the number down. It goes in the output.
3. **Use McNemar on the discordant pairs** for accuracy-style outcomes. Items every
model got right, or every model got wrong, carry no information about which is better.
4. **Correct with Holm** across the full sweep, not per workload. See `holm()` in
`lib/stats.ts`.
5. **Report the effect size next to the p-value.** Significant and tiny is a real
category and it usually means "do not switch".
6. **Check the practical gates separately.** Latency, cost, and failure modes disqualify
models regardless of accuracy. A model that wins by 2 points at 30x the latency has
not won. Say