calibratelisted
Install: claude install-skill mehboobali98/rmine-skills
# Calibrate
`/estimate` prices against the anchor table in
`skills/estimate/references/rubric.md`, which was derived from past
**estimates** — so it reproduces how those were made, including any systematic
error. This skill is the only thing that can tell you whether they were any
good. Now that the rubric is shared, a correction here moves every estimate the
team produces, which raises the bar for making one.
```sh
python3 "${CLAUDE_PLUGIN_ROOT}/skills/calibrate/scripts/actuals.py" <path>...
```
Point it at the product repo's committed `estimates/` directory — that's the
corpus `/estimate` builds up, one file per ticket. It also accepts individual
`estimate-<id>.md` files, or any single file holding many estimates. It parses
each estimate's issue reference and `Total` line, pulls matching time entries
via `rmine time list`, and reports a ratio per ticket plus aggregates.
**If the directory is thin, stop there.** The bar below needs 20+ tickets, and
a corpus that small means the answer is "not yet" regardless of what the ratios
say. That is a finding worth reporting, not a failure.
The script does the arithmetic. **Interpreting it is the hard part, and doing it
wrong is worse than not running it** — a confidently wrong multiplier applied
across a team's estimates is a lot of damage from one bad inference.
## Read the data quality before reading the numbers
Logged time is not ground truth. It is what people remembered to type in. Check
all of these before drawing any co