ceph-troubleshootinglisted
Install: claude install-skill air-gapped/skills
# ceph-troubleshooting
Defaults below were read from `src/common/options/*.yaml.in` at v18.2.8,
v19.2.6 and v20.2.4 on **2026-09-23**; procedures from docs.ceph.com and
Rook docs at v1.20.7. On a live cluster, `ceph config show osd.N <opt>` beats
any default written here.
## First moves
1. `ceph health detail` — work from the codes, not from `ceph status`.
2. Under Rook, run Ceph commands through the toolbox or
`kubectl rook-ceph ceph <args>` (plugin v0.9.6, 2026-03-31). If the
toolbox fails with `handle_auth_bad_method` / errno 13 right after a
Ceph upgrade, it is a key rotation, not an outage — see
rook-ceph-best-practices.
3. `AUTH_INSECURE_*` codes on 19.2.6+ / 20.2.4+ belong to the CVE-2025-30156
rotation (rook-ceph-best-practices), not to this skill.
## Recovery and backfill are slow — mClock
`osd_op_queue` defaults to `mclock_scheduler` on Quincy through Tentacle.
Under mClock, `osd_max_backfills` and `osd_recovery_max_active*` are
**locked**: `ceph config set` reports success and the value is reverted to
the profile's built-in. This is the usual "raised backfills, nothing
changed" report.
| Goal | Do |
|---|---|
| Faster recovery, accept client latency | `ceph config set osd osd_mclock_profile high_recovery_ops` (revert to `balanced` afterwards) |
| Use the classic knobs anyway | `ceph config set osd osd_mclock_override_recovery_settings true`, then set `osd_max_backfills` / `osd_recovery_max_active` |
| Check what is really in force | `ceph confi