Datadog
MonitoringCommonly used with
Skills using Datadog (139)
datadog-automation
Automate Datadog tasks via Rube MCP (Composio): query metrics, search logs, manage monitors/dashboards, create events and downtimes. Always search tools first for current schemas.
evidence-hygiene
Evidence-capture and PoC-redaction discipline for bug-bounty submissions: cookie redaction protocol (which fields to mask, Preview annotation / Burp panel hiding / DevTools workflow), PII black-bar discipline (what to mask in other-user data — names, emails, phones, faces — vs what is safe to leave — usernames, trace IDs, request bodies), HAR file sanitization (jq filters for Cookie/Set-Cookie/Authorization headers), Burp Repeater/Intruder screenshot hygiene (hide request body, show only Results table for rate-limit attacks), Chrome DevTools Console PoC patterns (credentials include so cookies are not echoed, labeled console.log), screenshot capture order, filename conventions, post-submission rotation hygiene. Use BEFORE any PoC screenshot, BEFORE attaching a HAR, or whenever preparing evidence with session cookies or other-user PII. Pairs with bugcrowd-reporting and report-writing.
agent-observability-spec
Specify the tracing, metrics, and alerting for an AI agent or LLM feature in production. Use when asked what to log for an LLM app, design agent tracing or spans, define quality and cost monitors, or answer 'how do we know if the agent is misbehaving?'. Produces an observability spec with a trace schema, metric definitions with owners and alert thresholds, sampling and retention policy, and a privacy note for logged content.
boss
OPS on-demand: This skill should be used when the user asks to "/ops:boss", "boss mode", or "what needs…
flow
OPS on-demand: This skill should be used when the user asks to "/ops:flow", "dev lifecycle", or "where…
ops-accounts
OPS on-demand: This skill should be used when the user asks to "rotate accounts", "switch…
ops-ar
OPS on-demand: This skill should be used when the user asks to "A&R this track", "demo verdict", or…
ops-aws-audit
OPS on-demand: This skill should be used when the user asks to "audit AWS", "unused AWS resources", or…
ops-comms
OPS on-demand: This skill should be used when the user asks to "send a message"…
ops-competitors
OPS on-demand: This skill should be used when the user asks to "competitor intel", "what did they…
ops-credentials
OPS on-demand: This skill should be used when the user asks to "which keys are set", "missing…
ops-daemon
OPS on-demand: This skill should be used when the user asks to "daemon health", "background services…
ops-dash
OPS on-demand: This skill should be used when the user asks to "ops dashboard", "pixel HQ", or…
ops-deploy
OPS on-demand: This skill should be used when the user asks to "deploy status", "what is in…
ops-desktop
OPS on-demand: This skill should be used when the user asks to "control the desktop", "click on…
ops-doctor
OPS on-demand: This skill should be used when the user asks to "plugin broken", "ops doctor", or…
ops-ecom
OPS on-demand: This skill should be used when the user asks to "shopify", "orders inventory", or…
ops-feature-dev
OPS on-demand: This skill should be used when the user asks to "guided feature", "feature-dev", or…
ops-fires
OPS on-demand: This skill should be used when the user asks to "production fires", "what is on fire"…
ops-fleet
OPS on-demand: This skill should be used when the user asks to "fleet dashboard", "CLIProxy sessions"…
ops-go
OPS on-demand: This skill should be used when the user asks to "morning briefing", "/ops:ops-go", or…
ops-gtm
OPS on-demand: This skill should be used when the user asks to "go to market", "GTM plan", or…
ops-home
OPS on-demand: This skill should be used when the user asks to "homey", "smart home", or…
ops-humanizer
OPS on-demand: This skill should be used when the user asks to "humanize this", "this reads like…
ops-inbox
OPS on-demand: This skill should be used when the user asks to "check inbox", "inbox zero"…
ops-integrate
OPS on-demand: This skill should be used when the user asks to "add an API", "integrate SaaS", or…
ops-leadgen
OPS on-demand: This skill should be used when the user asks to "leadgen drafts", "cold email approve"…
ops-linear
OPS on-demand: This skill should be used when the user asks to "linear sprint", "create a ticket", or…
ops-mac
OPS on-demand: This skill should be used when the user asks to "mac is slow", "macos fix", or…
ops-marketing
OPS on-demand: This skill should be used when the user asks to "klaviyo", "ads spend", or…
ops-mcp
OPS on-demand: This skill should be used when the user asks to "MCP down", "reconnect MCP", or…
ops-merge
OPS on-demand: This skill should be used when the user asks to "merge PRs", "salvage branches", or…
ops-monitor
OPS on-demand: This skill should be used when the user asks to "datadog", "APM alerts", or…
ops-next
OPS on-demand: This skill should be used when the user asks to "what should I do next", "priority…
ops-orchestrate
OPS on-demand: This skill should be used when the user asks to "orchestrate projects", "dispatch…
ops-package
OPS on-demand: This skill should be used when the user asks to "ship a parcel", "print a label", or…
ops-pocket
OPS on-demand: This skill should be used when the user asks to "pocket memos", "voice memo pipeline"…
ops-projects
OPS on-demand: This skill should be used when the user asks to "portfolio dashboard", "gsd projects"…
ops-recap
OPS on-demand: This skill should be used when the user asks to "recap daemon", "tmux marquee", or…
ops-release
OPS on-demand: This skill should be used when the user asks to "release the plugin", "publish ops…
ops-revenue
OPS on-demand: This skill should be used when the user asks to "burn rate", "runway", or…
ops-secret-sync
OPS on-demand: This skill should be used when the user asks to "doppler github secrets", "secret…
ops-settings
OPS on-demand: This skill should be used when the user asks to "update credentials", "ops settings", or…
ops-ship
OPS on-demand: This skill should be used when the user asks to "ship ops plugin", "merge all PRs and…
ops-social-planner
OPS on-demand: This skill should be used when the user asks to "content calendar", "what is scheduled"…
ops-socials
OPS on-demand: This skill should be used when the user asks to "tweet", "post to linkedin", or…
ops-speedup
Cross-platform, hardware-adaptive system optimizer. Auto-detects macOS / Linux / WSL / Windows (MINGW/Cygwin/MSYS2) and CPU/RAM/disk/GPU profile, then picks the right cleanup strategy. Scans reclaimable disk space, memory pressure, runaway processes, startup bloat, network issues. CleanMyMac built into Claude Code.
ops-status
Lightweight green/red status panel for every configured integration. No gather, no actions.
mcp-oauth-remote-gateway
Manual OAuth for remote MCP servers on headless gateways.
datadog-cli
Datadog CLI for searching logs, querying metrics, tracing requests, and managing dashboards. Use this when debugging production issues or working with Datadog observability.
deployment-pipeline-design
Design multi-stage CI/CD pipelines with approval gates, security checks, and deployment orchestration. Use this skill when designing zero-downtime deployment pipelines, implementing canary rollout strategies, setting up multi-environment promotion workflows, or debugging failed deployment gates in CI/CD.
ops-rotate
OPS on-demand: This skill should be used when the user asks to "rotate Claude", "max seats", or…
ops-rotate-setup
OPS on-demand: This skill should be used when the user asks to "rotate setup", "enroll Claude seat", or…
ops-deploy-fix
OPS on-demand: This skill should be used when the user asks to "deploy auto-fix", "failed deploy", or…
ops-desk
OPS on-demand: This skill should be used when the user asks to "desk sweep", "open decisions", or…
ops-resume
OPS on-demand: This skill should be used when the user asks to "resume sessions", "reopen ghostty…
ops-rules
OPS on-demand: This skill should be used when running any ops skill, or when the user asks to "ops…
http-mcp-headers
Implement secret-safe HTTP headers for MCP transport in gh-aw.
otel-queries
Analyze gh-aw OpenTelemetry traces from JSONL mirrors or OTLP backends.
the-homepage
Structure a dev-tool homepage that converts developers into champions. Use when the landing page is written for the buyer instead of the developer, reads as salesy, buries what the product does, or makes it hard to start. Pairs with the ShipReady homepage audit.
ledger
OPS on-demand: This skill should be used when the user asks to "ops ledger", "what did we handle", or…
review-logging-patterns
Review code for logging patterns and suggest evlog adoption. Guides setup on Nuxt, Next.js, SvelteKit, Nitro, TanStack Start, React Router, NestJS, Express, Hono, Fastify, Elysia, Cloudflare Workers, and standalone TypeScript. Detects console.log spam, unstructured errors, and missing context. Covers wide events, structured errors, drain adapters (Axiom, OTLP, HyperDX, PostHog, Sentry, Better Stack, Datadog), sampling, enrichers, and AI SDK integration (token usage, tool calls, streaming metrics, telemetry integration, cost estimation, embedding metadata).
company-research
Create a company research brief with executive quotes, product strategy, and org context. Use when preparing for interviews, competitive analysis, partnerships, or market-entry work.
positioning-workshop
Run a positioning workshop that surfaces target customer, unmet need, category, benefits, and differentiation. Use when your product messaging feels fuzzy, generic, or misaligned.
golang-samber-slog
Structured logging extensions for Golang using samber/slog-**** packages — multi-handler pipelines (slog-multi), log sampling (slog-sampling), attribute formatting (slog-formatter), HTTP middleware (slog-fiber, slog-gin, slog-chi, slog-echo), and backend routing (slog-datadog, slog-sentry, slog-loki, slog-syslog, slog-logstash, slog-graylog...). Apply when using or adopting slog, or when the codebase already imports any github.com/samber/slog-* package.
scaffold-cost-check
Measure Mycelium's own scaffold token cost (CLAUDE.md + engine + harness + canvas + memory) and surface a structured estimate. One-shot audit; pair with /framework-health for trend tracking.
rum-tracking
Guides product analytics and RUM (Real User Monitoring) event tracking in web (React/Next.js) and mobile (React Native/Expo) apps. Decides what user interactions are valuable to capture, what's noise, what's PII to avoid, and how to implement, audit, update, and remove tracking code cleanly. Covers event naming, property schemas, tracking plans, GDPR/CCPA/DPDPA compliance, OpenTelemetry semantic conventions for browser and mobile RUM, and platforms (PostHog, Segment, Mixpanel, Amplitude, Datadog RUM, Sentry, OTel, Dash0). Modes: guide (default), implement, audit, remove, plan. Triggers on "track this event", "add analytics", "what should I track", "is this PII", "tracking plan", "remove tracking", "audit analytics", "/rum-tracking".
experiment
Designing A/B tests: hypothesis docs, sample size, feature flags, significance analysis, CUPED, SRM detection, switchback experiments. Use when hypothesis validation is needed.
observability-checklist
Reviews a service or codebase against a full observability checklist — logs, metrics, traces, and alerting gaps.
canvas
Create a live React canvas beside chat for standalone analytical artifacts that benefit from visual layout: quantitative/security/ architecture analyses, data-heavy timelines/charts/tables, interactive explorations, or repeatable tools. Do not use for simple prose or code changes.
platform-skills
Use when troubleshooting, implementing, reviewing, or auditing platform infrastructure as a system — where Kubernetes, GitOps, CI/CD, and security concerns intersect. Provides structured diagnosis with blast radius, validation steps, and rollback plan for: Kubernetes, Flux CD, Argo CD, Terraform, GitHub Actions (composite actions, OIDC, SHA pinning), AWS, Azure, GKE, Linkerd, KEDA, Karpenter, supply chain security (Cosign, SBOM, SLSA), Falco, Chaos Engineering, DORA metrics, Datadog/Dynatrace/LLM observability, SOC 2, and PR review.
backend-dev-guidelines
Comprehensive backend development guide for Langfuse's Next.js 14/tRPC/Express/TypeScript monorepo. Use when creating tRPC routers, public API endpoints, BullMQ queue processors, services, or working with tRPC procedures, Next.js API routes, Prisma database access, ClickHouse analytics queries, Redis queues, OpenTelemetry instrumentation, Zod v4 validation, env.mjs configuration, tenant isolation patterns, or async patterns. Covers layered architecture (tRPC procedures → services, queue processors → services), dual database system (PostgreSQL + ClickHouse), projectId filtering for multi-tenant isolation, traceException error handling, observability patterns, and testing strategies (Jest for web, vitest for worker).
k8s-rightsize
Right-size Kubernetes pod CPU/memory requests from real Prometheus usage data (per-pod p95 CPU, p98 memory over 7 days), compute how many nodes the cluster needs, and tune KEDA queue-based autoscaling. Use when the user asks to rightsize workloads, fix over/under-provisioned requests, cut node count or cloud cost, size Celery or background workers, tune KEDA listLength or HPA settings, or answer "how many nodes do we need". Works for any namespace with kube-prometheus and cAdvisor metrics.
golang-samber-oops
Structured error handling in Golang with samber/oops — error builders, stack traces, error codes, error context, error wrapping, error attributes, user-facing vs developer messages, panic recovery, and logger integration. Apply when using or adopting samber/oops, or when the codebase already imports github.com/samber/oops.
golang-samber-slog
Structured logging extensions for Golang using samber/slog-**** packages — multi-handler pipelines (slog-multi), log sampling (slog-sampling), attribute formatting (slog-formatter), HTTP middleware (slog-fiber, slog-gin, slog-chi, slog-echo), and backend routing (slog-datadog, slog-sentry, slog-loki, slog-syslog, slog-logstash, slog-graylog...). Apply when using or adopting slog, or when the codebase already imports any github.com/samber/slog-* package.
monitoring-observability
Monitoring, metrics, alerting, and observability standards. TRIGGER when: implementing health checks, collecting metrics, or defining alert rules. SKIP: log statement formatting (use logging-standards); CI pipeline setup (use github-actions-template).
flutter-observability
Design, implement, or repair production observability in Flutter using operational logs, error and crash capture, breadcrumbs, traces, release context, privacy controls, and verification. Use for diagnosability and incident signals; not product analytics, profiling-only optimization, or a general code review.
datadog-cli
Datadog CLI for searching logs, querying metrics, and tracing requests. Use when user mentions "Datadog", "DD logs", "production debugging", "check metrics", "trace request", or needs to investigate production incidents. Triggers on "datadog", "DD", "로그 검색", "메트릭 조회", "장애 조사".
claude-jobs
Find job openings at tech companies. Use when user asks about jobs, careers, openings, positions, roles, or salaries - either at specific companies or general tech job queries.
observability-slo
Observability and reliability engineering — structured logging, metrics, distributed tracing, correlation ids, error tracking, dashboards, SLIs/SLOs and error budgets, alerting that pages on symptoms rather than causes, on-call practice, incident response and blameless postmortems. Use when the user says "logging", "monitoring", "observability", "metrics", "tracing", "Prometheus", "Grafana", "Datadog", "Sentry", "OpenTelemetry", "SLO", "SLA", "uptime", "alerting", "on-call", "incident", "postmortem", "how do we know if it breaks", "it broke and we didn't notice" or "debugging production"; and as a pass in any project audit. By Devleck.
alerting-rules-tuner
Cut alert noise and make every page mean something — rewrite alerting rules to fire on user-felt symptoms (error rate, latency SLO burn, failed requests) instead of causes (high CPU, full disk), with duration windows and severity routing so only urgent, actionable conditions reach a human. Use when on-call is fatigued by low-value pages, when real incidents get missed in the noise, or when alerts fire on causes rather than impact.
observability
Production observability done right — structured logs, distributed traces, metrics, alerting, SLO/SLI. Use when adding logging to a new service, designing dashboards, choosing between OpenTelemetry / Datadog / Grafana stack, defining SLOs for a feature, writing alert rules, or untangling a noisy alert channel. Stack-agnostic; recipes target OpenTelemetry as the canonical instrumentation, Prometheus + Grafana / Datadog as the canonical backends. Pairs with performance (perf budgets), security-web (audit logs), and incident-response (alert → runbook).
performance
Application performance for backend (Go+Fiber, Python+FastAPI), web frontends, and React Native mobile clients. Use when measuring or improving Web Vitals (LCP / INP / CLS), mobile FPS / TTI / memory, image/font/code-split optimization, profiling, or interpreting Lighthouse / Reanimated / Hermes / Flipper traces. Pairs with frontend-fundamentals + react-native-patterns + db-design.
mcp-oauth-remote-gateway
Manual OAuth for remote MCP servers on headless gateways.
ecs-observability
Advise on Amazon ECS observability architecture — select the logs/metrics/traces stack (CloudWatch, Container Insights, X-Ray, ADOT/OpenTelemetry, Managed Prometheus/Grafana, FireLens to third-party) by compliance needs, existing tooling, scale, budget, and launch types (EC2, Fargate, Managed Instances, ECS Anywhere). Use for "how should we monitor our ECS services", "Container Insights or Prometheus for ECS", "are we losing ECS container logs", "set up tracing on Fargate", "ECS logging best practices", "Datadog vs CloudWatch for ECS", "GPU metrics for ECS tasks", or "plan live-debug access to an ECS task". Any ECS logging, metrics, tracing, or alerting design question qualifies even if "observability" is never said. Skip for EKS/Kubernetes (eks-* skills), deployment mechanics/CI-CD/deploy-failure diagnosis (ecs-devops; deploy-failure alerting stays here), security posture beyond observability audit logging (ecs-security), live-estate audits (ecs-operation-review), and FinOps audits of observability spend.
instrumentation-setup
Wire instrumentation.ts in Next.js 16 — Sentry, pino logger, OpenTelemetry traces. Edge vs Node runtime split. The single place observability is bootstrapped.
agent-guard
Scan AI agent skills, plugins, and MCP servers for malicious code BEFORE installation — catches prompt injection, credential theft, data exfiltration, and backdoors. Skills and the static MCP source scan use NVIDIA SkillSpector (static patterns + taint tracking + YARA + live OSV.dev CVE lookup + LLM semantic analysis, which runs by default through the user's own claude / codex / gemini CLI login — no API key — or any hosted provider with a key); the optional live MCP runtime check uses cisco-ai-mcp-scanner with separate MCP_SCANNER_LLM_* settings and any LiteLLM-supported provider. Skills follow the open SKILL.md standard (agentskills.io) and MCP is an open protocol, so one scan covers every agent: repos are downloaded as commit-pinned ZIP snapshots (never git clone before a verdict), and the exact scanned commit is installed via the bundled universal installer into Claude Code, Claude Desktop, Codex, Antigravity/Gemini, Hermes, and OpenClaw at once — or a chosen subset via --tools. Scan once, install everywh
doc-this-tracer
Use as an optional Discovery agent that resolves 🔴 gaps via dynamic analysis when static analysis falls short. STRICTLY DESCRIPTIVE — cites log lines / span IDs / samples; never proposes fixes or labels behavior as wrong. Read-only — never executes mutating code. Sources: log files, distributed traces (OTLP, Jaeger, Datadog), anonymized production samples, error-tracker exports (Sentry, Bugsnag, Rollbar). Resolves gaps like actual state machines, caller payloads, dead endpoints, error rates. Also runs the 🟢 corroboration sweep: stamps Evidence provenance (static → static + runtime, artifact-cited) on telemetry-matched scenarios — hard-advisory when the legacy system cannot be run live. Updates .doc-this-sdd/dynamic.md; promotes 🔴→🟢 ONLY on a cited specific runtime artifact — otherwise the gap stays 🔴. No 🟡. Triggers: '/doc-this-tracer', 'analyze logs', 'mine traces'. NOT for live system probing or load testing. NOT for static-only analysis (doc-this-code-analyst/detective).
golang-samber-oops
Structured error handling in Golang with samber/oops — error builders, stack traces, error codes, error context, error wrapping, error attributes, user-facing vs developer messages, panic recovery, and logger integration. Apply when using or adopting samber/oops, or when the codebase already imports github.com/samber/oops.
golang-samber-slog
Structured logging extensions for Golang using samber/slog-**** packages — multi-handler pipelines (slog-multi), log sampling (slog-sampling), attribute formatting (slog-formatter), HTTP middleware (slog-fiber, slog-gin, slog-chi, slog-echo), and backend routing (slog-datadog, slog-sentry, slog-loki, slog-syslog, slog-logstash, slog-graylog...). Apply when using or adopting slog, or when the codebase already imports any github.com/samber/slog-* package.
deployment-monitor
Monitor deployments with read-only evidence gathering, anomaly detection, cadence summaries, and local alerts. Use when watching a deploy, validating staging or production after release, comparing Datadog/database/runtime evidence, investigating new errors or skipped work, or linking post-deploy anomalies back to GitHub code, commits, PRs, or issues.
aio-dashboard-design
Design and review SaaS analytics dashboards — chart selection, anti-patterns, WCAG 2.2 a11y, color palettes, and storytelling.
integration-specialist-agent
Resilient third-party integration patterns — API clients, webhooks, MCP server wiring, retries/circuit-breakers, and idempotency. Use when connecting to an external service (Stripe/GitHub/Slack/S3/etc.), building or verifying a webhook receiver, adding retry/backoff or circuit-breaker resilience, registering an MCP server, or reviewing an integration for signature-verification and rate-limit gaps.
observability-setup
Instruments applications with structured logging, metrics, and distributed traces, then derives SLIs, SLOs, error budgets, and alerts that page only on user-facing pain. Use this skill when the user asks to "add observability", "instrument my app", "set up OpenTelemetry/OTel", "add structured logging", "expose Prometheus metrics", "add tracing", "define SLOs/SLIs", "set up alerting", "reduce alert noise", "build a Grafana dashboard", or wants the three pillars (logs/metrics/traces) wired into a service.
logging-design-patterns
Structured logging best practices - Pino JSON output, log levels, correlation IDs, PII redaction, sampling, async context, canonical log lines
kora-telemetry-logging
Structured logging for Kora — SLF4J+Logback (LogbackModule), KoraAsyncAppender, StructuredArgument, Kora MDC, @Log/@Mdc. Use when wiring logback.xml, per-package log levels, or JSON logs for ELK/Datadog/Splunk.
devops-automator
Automates infrastructure and cloud operations: provisioning, configuration management and operational tooling. Use for day-to-day DevOps. For pipelines, use ci-cd-pipeline-builder.
artillery
When the user wants to design, implement, debug, or operate Artillery load tests. Use when the user mentions "Artillery," "artillery.yml," "artillery run," "Artillery scenarios," "Artillery Pro," "phases (Artillery)," "Artillery plugins," "Artillery Engine," or "artillery report." For JS-tested perf with thresholds see k6. For JVM perf see gatling. For Python see locust. For JMeter see jmeter.
k6
When the user wants to design, implement, debug, or operate k6 load tests. Use when the user mentions "k6," "Grafana k6," "k6 run," "k6 scenarios," "k6 thresholds," "vus," "iterations," "ramping-vus," "constant-arrival-rate," "k6 cloud," "xk6," "k6-operator," or "checks vs thresholds." For JMeter see jmeter. For Gatling see gatling. For Locust see locust. For Artillery see artillery. For overall perf testing strategy see test-strategy.
test-reporting
When the user wants to design, integrate, or improve test reporting — JUnit XML aggregation, HTML dashboards, Allure, Cucumber Reports, screenshots / videos / traces, failure triage, and historical analytics. Use when the user mentions "test reports," "JUnit XML," "Allure," "mochawesome," "Cucumber Reports," "Datadog CI Visibility," "Buildkite Test Analytics," "test result aggregation," "trace / artifacts on failure," "test dashboard," or "trend analysis." For CI orchestration see ci-test-orchestration. For flake analytics see flaky-test-management.
Showing top 100 of 139 skills using Datadog by quality score.
See all 139 skills via search →Integration detected automatically from skill content. Some results may be false positives.