health-checkslisted
Install: claude install-skill Amey-Thakur/AI-SKILLS
# Health checks
Three different questions, three different answers: should you restart me
(liveness), should you send me traffic (readiness), am I still booting
(startup). Conflating them turns partial outages into total ones.
## Method
1. **Liveness checks only the process.** Event loop responsive, not
deadlocked: return 200 if the handler runs at all. Never include
dependencies: if the database blips and liveness checks it, the
orchestrator restart-loops your entire healthy fleet during the one
moment it needs stability.
2. **Readiness checks ability to serve.** Required dependencies
(DB pool has a connection, config loaded, migrations current) with
short per-check timeouts, cached for a few seconds. Unready is
recoverable and expected: during startup, shutdown drain (see
graceful-shutdown), and dependency outages.
3. **Distinguish required from degradable dependencies.** The database
is required; the recommendation service is not. Degradable
dependencies never fail readiness; they flip feature flags and show
up in metrics. Otherwise one optional system's outage drains every
pod that could have served 90% of traffic.
4. **Startup probe covers slow boots.** Cache warming, model loading,
migration waits: a startup probe with a generous budget keeps
liveness (tight thresholds) from killing pods mid-boot. Without it
you either boot-loop or loosen liveness for everyone.
5. **Fail readiness on saturation, carefully.** Rejecting at