calibrate
SolidCalibrate skills/role cards for leaks/gaps with recall, precision, and confidence-accuracy checks.
Install
Quality Score: 81/100
Skill Content
Details
- Author
- Borda
- Repository
- Borda/AI-Rig
- Created
- 5 months ago
- Last Updated
- today
- Language
- Python
- License
- Apache-2.0
Similar Skills
Semantically similar based on skill content — not just same category
calibrate
Reflective end-of-session self-improvement. Scans the current Claude Code session for corrections, preferences, repeated patterns, errors, success patterns, and voice violations, then proposes numbered concrete patches to memory, settings, ceo-only skills, and ceo-only rules. Corporate files route to a separate review queue and are NEVER auto-applied. Use at end of every working session. Light mode for low-token state or quick sweeps. CEO-only - never propagates to execs.
audit
Audit Codex configuration/workflow drift; emit ranked gaps and measurable gates.
nasde-benchmark-calibration
Calibrate assessment rubrics by reviewing agent work in GitHub/GitLab PRs and feeding human comments back into the rubric. Use this skill when the user wants to: - Calibrate, tune, or sanity-check assessment criteria / dimensions of a benchmark - Review trial diffs alongside the LLM-as-a-Judge scores in a PR/MR - Investigate why judge scores feel off, too harsh, too lenient, or misaligned with how a human would grade the code - Pull review comments back from PRs/MRs and turn them into concrete rubric edits Even if the user doesn't say "calibrate" — if they're worried the LLM judge's scores diverge from human judgment, or want to align scores with a real developer's opinion before freezing a benchmark, this skill applies.