qa-observability

Solid

Implement OpenTelemetry logs/metrics/traces, SLI/SLO gates, burn-rate alerts, and APM integrations. Use when adding or validating observability.

AI & Automation 89 stars 19 forks Updated 3 days ago MIT

Install

View on GitHub

Quality Score: 85/100

Stars 20%
65
Recency 20%
100
Frontmatter 20%
70
Documentation 15%
100
Issue Health 10%
80
License 10%
100
Description 5%
100

Skill Content

# QA Observability and Performance Engineering Use telemetry (logs, metrics, traces, profiles) as a QA signal and a debugging substrate. Core references (see `data/sources.json`): OpenTelemetry, W3C Trace Context, and SLO practices (Google SRE). ## Quick Start (Default) If key context is missing, ask for: critical user journeys, service/dependency inventory, environments (local/staging/prod), current telemetry stack, and current SLO/SLA commitments (if any). 1. Establish the minimum bar: correlation IDs + structured logs + traces + golden metrics (latency, traffic, errors, saturation). 2. Verify propagation: confirm `traceparent` (and your request ID) flow across boundaries end-to-end. 3. Make failures diagnosable: every test failure captures a trace link (or trace ID) plus the correlated logs. 4. Define SLIs/SLOs and error budget policy; wire burn-rate alerts (prefer multi-window burn rates). 5. Produce artifacts: a readiness checklist plus an SLO definition and alert rules (use `assets/checklists/template-observability-readiness-checklist.md` and `assets/monitoring/slo/*`). ## Default QA stance - Treat telemetry as part of acceptance criteria (especially for integration/E2E tests). - Require correlation: request_id + trace_id (traceparent) across boundaries. - Prefer SLO-based release gating and burn-rate alerting over raw infra thresholds. - Budget overhead: sampling, cardinality, retention, and cost are quality constraints. - Redact PII/secrets by default (logs and...

Details

Author
vasilyu1983
Repository
vasilyu1983/AI-Agents-public
Created
10 months ago
Last Updated
3 days ago
Language
Python
License
MIT

Integrates with

Similar Skills

Semantically similar based on skill content — not just same category

AI & Automation Listed

observability

Use when instrumenting a feature for production — adding or improving logging, metrics, tracing, or alerting — or when production issues are reported and the current telemetry cannot explain them. Common triggers: observability, instrumentation, add logging, structured logging, log levels, add metrics, adding metrics, metrics dashboard, set up metrics, set up tracing, distributed tracing, opentelemetry, set up alerting, alerting on, alert rule, runbook, telemetry setup, app telemetry, monitor this feature, how do we observe, what is working in production, monitoring alerts, instrument this, production visibility.

1 Updated yesterday
MarkBovee
AI & Automation Listed

observability

Use for production logs, metrics, traces, health, alerts, SLOs, and performance measurement; not debug prints.

1 Updated today
kreek
AI & Automation Listed

sota-observability

State-of-the-art observability and reliability engineering (2026). Use when instrumenting code (structured logging, metrics, distributed tracing with OpenTelemetry, SLOs, alerting, health endpoints) or auditing an existing codebase's observability posture (can on-call answer "why is this request slow?" and "what broke at 3am?"). Not for security detections, SIEM, or threat hunting — use sota-detection-engineering. Triggers: logging, metrics, tracing, monitoring, alerting, SLO, SLI, error budget, OpenTelemetry, OTel, Prometheus, Grafana, debugging production, incident, on-call, telemetry, instrumentation, health check, runbook, Sentry, crash reporting, profiling.

23 Updated today
martinholovsky