observability-and-instrumentation
SolidDesigns or reviews logs, metrics, traces, alerts, correlation, dashboards, and diagnostic instrumentation so failures and performance changes can be detected and explained at component boundaries. Use for observability work, incident readiness, instrumentation, or production diagnostics. Not for fixing a specific incident before reproducing it or for adding noisy logging without an operational question.
Install
Quality Score: 84/100
Skill Content
Details
- Author
- thiientv
- Repository
- thiientv/godmode
- Created
- 1 weeks ago
- Last Updated
- today
- Language
- Python
- License
- MIT
Similar Skills
Semantically similar based on skill content — not just same category
observability
Instrument software so production questions get answered from signals, not guesses. Use when adding logging, metrics, tracing, or alerts, or when a system is hard to debug in production.
observability-incident-response
Design and assess application observability and incident response across logs, metrics, traces, correlation, alerting, runbooks, recovery, and post-incident learning. Use for production readiness, reliability improvement, or active incident diagnosis.
design-observability
Design the signals needed to tell whether users are succeeding and to explain failures quickly. Connect service indicators, structured events, traces, alerts, and runbooks to real journeys. Use when a live system is hard to operate or before a new production service launches.