← ClaudeAtlas

logging-and-observabilitylisted

Instrument a service so production problems are diagnosable: structured JSON logs with consistent fields and levels, correlation ids across requests and jobs, the RED and USE metrics, distributed tracing with OpenTelemetry, health and readiness endpoints, alerts on symptoms with runbooks, and dashboards per service; with the setup for Node, Python, and Go and the cost and privacy rules that keep it sustainable. Use when a service goes to production, when an incident could not be traced, when logs are noisy or expensive, or when alerts fire without meaning.
KhaledSaeed18/dotclaude · ★ 5 · DevOps & Infrastructure · score 80
Install: claude install-skill KhaledSaeed18/dotclaude
Observability is the ability to ask a new question about the system's behaviour without shipping new code. Logs, metrics, and traces each answer different questions; a service needs all three, wired so that one request can be followed across them. ## Logs - **Structured** (JSON, one object per line) to stdout; the platform ships them. Never format strings for humans in production logs; dashboards and queries need fields. - **Standard fields on every line**: `timestamp` (ISO 8601 UTC), `level`, `message`, `service`, `version`, `env`, `request_id` or `trace_id`, and the domain ids in play (`user_id`, `order_id`). A logger bound with context (`logger.child({ requestId })`) so handlers do not repeat them. - **Levels**: `error` for failures needing attention (paged or triaged), `warn` for degraded but handled, `info` for business events and request summaries (one line per request with method, path, status, duration), `debug` off in production. A log at `error` that nobody would act on is noise that hides the one that matters. - **One request summary line** at the boundary rather than a line per step; steps at `debug`. - **Never log**: secrets, tokens, passwords, full card numbers, full request bodies, personal data beyond the ids needed. Redaction in the logger config (pino `redact`, structlog processors), not by remembering. - **Sampling** for high-volume `info` and for `debug` in production when needed; errors are never sampled. - **Retention and cost**: logs cost by volume; a