forge-calibratelisted
Install: claude install-skill yerros/context-forge
# forge-calibrate
A reviewer you have never measured is a reviewer you are trusting on vibes. This
skill turns "the review feels inconsistent" into two numbers — **recall** against
planted findings and **agreement** between repeated runs — and a list of what the
reviewer never sees. Run it before and after changing a lens, an agent prompt, a
rule card, or the confidence gate; without the before/after pair, a prompt change
is a guess.
## Argument
- `--runs N` — how many times each case is reviewed (default 3; 5 for a decision
you'll act on).
- `--case <name>` — only that golden case.
- `--focus=<lenses>` — passed through to the review.
- `seed` — copy the bundled golden set into the project (see Golden set) and stop.
## Golden set
A case is a directory with `before/` and `after/` trees (the diff is
`diff -ruN before after`), an `expected.txt` of finding keys
(`<lens>:<file>:<tag>`, one per line, `#` comments allowed), and optionally a
`context/` with the spec / standards / `rules.txt` the case depends on.
- Bundled: `${CLAUDE_PLUGIN_ROOT}/skills/forge-calibrate/golden/` — eight cases, one
per failure class: swallowed error, scope creep + unrequested config, hollow
test, silent breakage of an untouched caller, rule-card violation that step 0
must catch by ID; and three **gate cases** for `forge-gatekeeper` — hardcoded
live token (step 0 must catch it), missing object-level authorization (IDOR),
irreversible migration shipped with its code change.
- Project: `<