← ClaudeAtlas

ceph-performancelisted

Tune and benchmark a healthy Ceph cluster, with Rook-on-Kubernetes specifics: mClock profiles and the NVMe capacity-fallback trap (osd_mclock_max_capacity_iops), osd_memory_target from pod requests, BlueStore allocation and cache defaults, PG autoscaler targets (mon_target_pg_per_osd, bulk, target_size_ratio), public/cluster network and Multus choices, msgr2 modes, Tentacle Fast EC (allow_ec_optimizations, stripe_unit), krbd vs rbd-nbd in ceph-csi, librbd cache, and benchmark method (rados bench, rbd bench, fio). Squid 19.2 and Tentacle 20.2.
air-gapped/skills · ★ 5 · AI & Automation · score 80
Install: claude install-skill air-gapped/skills
# ceph-performance Defaults read from `src/common/options/*.yaml.in` at v19.2.6 and v20.2.4, docs.ceph.com (tentacle pages) and Rook docs at v1.20.7, on **2026-09-23**. Measure before and after every change; `ceph config show osd.N <opt>` is the value actually in force. ## mClock: the NVMe capacity trap mClock sizes every reservation from each OSD's measured IOPS capacity. At OSD start an `osd bench` sets `osd_mclock_max_capacity_iops_{hdd,ssd}`. If the result exceeds the sanity threshold — **500 (HDD), 80 000 (SSD)** — it is discarded and the OSD falls back to the default capacity (**21 500** SSD) with a cluster-log warning. Modern NVMe routinely exceeds 80 000, so mClock under-drives it. 1. `ceph config show osd.N osd_mclock_max_capacity_iops_ssd` on every OSD; `21500` means the fallback hit. 2. Measure the device with fio at 4 KiB random write, then `ceph config set osd.N osd_mclock_max_capacity_iops_ssd <iops>`. 3. Re-check after replacing a device — the value is per OSD. Profiles: `balanced` is the default from 17.2.7 and in Reef, Squid and Tentacle (17.2.0–17.2.6 shipped `high_client_ops`). Use `high_client_ops` for latency-sensitive steady state, `high_recovery_ops` only for a maintenance window. ## Memory - Under Rook, Ceph reads `POD_MEMORY_REQUEST` and uses it 1:1 as `osd_memory_target`; with only a limit, target = limit × `osd_memory_target_cgroup_limit_ratio` (0.8). Set requests and limits explicitly on `spec.resources.osd`; the limit needs head