ceph-performancelisted
Install: claude install-skill air-gapped/skills
# ceph-performance
Defaults read from `src/common/options/*.yaml.in` at v19.2.6 and v20.2.4,
docs.ceph.com (tentacle pages) and Rook docs at v1.20.7, on **2026-09-23**.
Measure before and after every change; `ceph config show osd.N <opt>` is
the value actually in force.
## mClock: the NVMe capacity trap
mClock sizes every reservation from each OSD's measured IOPS capacity. At
OSD start an `osd bench` sets `osd_mclock_max_capacity_iops_{hdd,ssd}`.
If the result exceeds the sanity threshold — **500 (HDD), 80 000 (SSD)** —
it is discarded and the OSD falls back to the default capacity
(**21 500** SSD) with a cluster-log warning. Modern NVMe routinely exceeds
80 000, so mClock under-drives it.
1. `ceph config show osd.N osd_mclock_max_capacity_iops_ssd` on every OSD;
`21500` means the fallback hit.
2. Measure the device with fio at 4 KiB random write, then
`ceph config set osd.N osd_mclock_max_capacity_iops_ssd <iops>`.
3. Re-check after replacing a device — the value is per OSD.
Profiles: `balanced` is the default from 17.2.7 and in Reef, Squid and
Tentacle (17.2.0–17.2.6 shipped `high_client_ops`). Use `high_client_ops`
for latency-sensitive steady state, `high_recovery_ops` only for a
maintenance window.
## Memory
- Under Rook, Ceph reads `POD_MEMORY_REQUEST` and uses it 1:1 as
`osd_memory_target`; with only a limit, target = limit ×
`osd_memory_target_cgroup_limit_ratio` (0.8). Set requests and limits
explicitly on `spec.resources.osd`; the limit needs head