← ClaudeAtlas

tune-metriclisted

Use when the user invokes it to push one measured metric toward a target in an unattended keep-or-revert loop. Not for one decided change, which build-change builds, an unproven failure, which find-cause owns, or choosing what to improve, which audit-architecture owns.
blauwtje/exo · ★ 1 · AI & Automation · score 70
Install: claude install-skill blauwtje/exo
# Tune metric A number counts only when a frozen harness measured it and the regression checks stayed green around it. The enemy is the run that edits its own yardstick, keeps changes it never measured, and calls a lucky sample or a plateau the finish. The overcorrection is a loop so cautious it stops at the first reject or asks the user about every reversible fix, while nobody is there to answer. ## When to use - The user invokes it with a metric, its direction and a target: latency, bundle size, memory, a score. - Not for one decided change: `build-change` builds it and needs no loop. - Not for a failure with an unproven cause: `find-cause` proves it first. - Not for finding what to improve: `audit-architecture` ranks the candidates. ## The loop 1. **Frame.** Write the metric, its direction and a stop predicate pairing a target with an attempt floor, such as "p95 at most 60% of baseline and at least 10 attempts", so one lucky sample cannot end the run. A budget the user names is the only other end. Locate the hot path with `exo:locate-code`, and pick a case that reproduces the complaint; with none, building one comes first. 2. **Prove the harness, then freeze it.** One command reports the median of at least five runs and their spread. Run it on two workloads that must differ and confirm it separates them beyond the spread, because a harness blind to a known difference is blind to yours. Commit the harness and its inputs; from here nothing edits them or the regression c