Topic · DevOps & Cloud
Best observability skills for Claude Code, Codex and other agents.
- skills
- 562
- official
- 82
Observability skills, ranked
Ranked by score. Sort bymost stars,trending,newest,recently updated
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 1 | Runs a metrics-first audit of a deployed Vercel project, gating investigations on real signals to produce ranked, citation-backed cost and performance recommendations. | vercel-labs/ | 32k | 8 repos | ~4.3k | Automated safety check: Pass | No licence | 1 mo ago |
| 2 | Enforces that every error in a service is captured to Sentry v8, with patterns for controllers, routes, cron jobs and database performance spans instead of console logging alone. | diet103/ | 10k | 2 repos | ~2.3k | Automated safety check: Notes | MIT | 2 mo ago |
| 3 | Installs and configures Kubeshark on a Kubernetes cluster, choosing between the quick CLI path and a Helm install with custom values. | kubeshark/ | 12k | — | ~3.6k | Automated safety check: Notes | Apache-2.0 | 6 days ago |
| 4 | Syntax reference for KFL2, the CEL-based display filter language used to search Kubernetes network traffic captured by Kubeshark, loaded before any filter is written. | kubeshark/ | 12k | — | ~3.6k | Automated safety check: Pass | Apache-2.0 | 6 days ago |
| 5 | Installs, checks and troubleshoots the KubeSphere ServiceMesh extension (Istio, Kiali, Jaeger), including grayscale release, sidecar injection, topology and tracing issues. | kubesphere/ | 17k | — | ~2.4k | Automated safety check: Pass | Unknown | 2 mo ago |
| 6 | Procedure for adding or changing YugabyteDB Active Session History wait states in TServer and DocDB C++ code, including the macro to use for sync and async paths. | yugabyte/ | 11k | — | ~4.5k | Automated safety check: Pass | Unknown | today |
| 7 | Helps instrument a custom Python or TypeScript agent to record events for Failproof AI, verify what gets written, and run an evaluator worker that scores the runs. | FailproofAI/ | 5.3k | — | ~6k | Automated safety check: Pass | Unknown | yesterday |
| 8 | Build or review Langfuse backend code. An agent skill from langfuse/langfuse. | langfuse/ | 35k | — | ~1.9k | Automated safety check: Pass | Unknown | today |
| 9 | A skill your agent uses when designing, generating an MVP blueprint for, auditing, troubleshooting, refactoring, or explaining an agentic harness for any domain. | DenisSergeevitch/ | 2.4k | — | ~7.4k | Automated safety check: Pass | MIT | 2 days ago |
| 10 | 10.Looper Scaffold a well-designed agent loop with best-practice coaching and a cross-model review council. | ksimback/ | 710 | — | ~2.7k | Automated safety check: Notes | MIT | 1 mo ago |
| 11 | Investigates past Kubernetes incidents from Kubeshark traffic snapshots: takes captures, dissects API calls, extracts PCAPs and compares traffic over time. | kubeshark/ | 12k | — | ~5.3k | Automated safety check: Pass | Apache-2.0 | 6 days ago |
| 12 | Installs and runs Crabwalk, a real-time monitor that shows an OpenClaw agent's activity as a graph on a web page you can open from another machine. | crabwise-ai/ | 874 | 1 repo | ~1.4k | Automated safety check: Notes | MIT | 2 mo ago |
| 13 | Routes every LLM call in a repository through the Caveman Cloud gateway in record mode, so requests and costs are measured without changing behavior. | JuliusBrussee/ | 110k | 1 repo | ~2.6k | Automated safety check: Warn | Apache-2.0 | today |
| 14 | Designs and deploys Axiom dashboards through the API, choosing chart types and writing APL or metrics queries, with templates and migration notes for Splunk and Grafana. | openclaw/ | 9.5k | — | ~4.9k | Automated safety check: Pass | MIT | today |
| 15 | 15.Motel Debug Debug applications with motel, a local OpenTelemetry ingest and query server. | kitlangton/ | 298 | — | ~2.2k | Automated safety check: Pass | MIT | 1 mo ago |
| 16 | A skill your agent uses when composing, adapting, or validating an Deep Researcher Agent workflow YAML under configs/ — selecting a shipped profile, enabling tools and datasourceregistry sources… | NVIDIA-AI-Blueprints/ | 883 | — | ~1.1k | Automated safety check: Notes | Apache-2.0 | today |
| 17 | Queries live Prometheus metrics and manages Grafana dashboards as code for Happy's infrastructure, using the grafanactl CLI and the Grafana datasource proxy API. | slopus/ | 24k | — | ~2k | Automated safety check: Notes | MIT | today |
| 18 | Finds unused data in Axiom by analyzing query patterns, then deploys a cost dashboard and ingest monitors to keep spend under the contract limit. | openclaw/ | 9.5k | — | ~1.7k | Automated safety check: Pass | MIT | today |
| 19 | 19.Temps Manage, deploy, operate, and instrument applications with Temps. | gotempsh/ | 822 | — | ~1.9k | Automated safety check: Pass | Apache-2.0 | today |
| 20 | Walks through adding a new built-in evlog drain adapter for an observability platform: source, build config, exports, tests, docs and PR scope. | evloghq/ | 1.9k | — | ~2.9k | Automated safety check: Pass | MIT | today |
| 21 | Installs and configures the WizTelemetry Platform Service extension for KubeSphere, the shared API server behind its observability extensions. | kubesphere/ | 17k | — | ~1.8k | Automated safety check: Pass | Unknown | 2 mo ago |
| 22 | 22.Diagnose Opt-in evidence-first causal diagnosis for bugs, browser or app/device failures, flaky behavior, and performance regressions. | CherryHQ/ | 4k | — | ~1.8k | Automated safety check: Pass | AGPL-3.0 | 4 days ago |
| 23 | Full Sentry SDK setup for Mini Programs — error monitoring, tracing, offline cache, source maps. | lizhiyao/ | 686 | — | ~5k | Automated safety check: Pass | MIT | today |
| 24 | Add and verify lightweight macOS runtime telemetry. An agent skill from robinebers/openusage. | robinebers/ | 4.3k | — | ~934 | Automated safety check: Pass | MIT | yesterday |
| 25 | Explores and queries OpenTelemetry metrics in Axiom MetricsDB, listing datasets, metrics and tags first and picking the right aggregation for each metric's type. | openclaw/ | 9.5k | — | ~2.6k | Automated safety check: Pass | MIT | today |
| 26 | Checks status, tails logs, queries the task queue, and starts, stops or restarts the Pilot daemon running on a specific AWS EC2 instance over SSM. | qf-studio/ | 722 | — | ~3k | Automated safety check: Notes | Unknown | yesterday |
| 27 | Investigates a service incident to its root cause by querying a UModel object graph alongside metrics, logs, topology and recent deployments. | alibaba/ | 412 | — | ~1.9k | Automated safety check: Pass | Unknown | 13 days ago |
| 28 | Write a new library instrumentation end-to-end. An agent skill from DataDog/dd-trace-java. | DataDog/ | 736 | — | ~3.7k | Automated safety check: Notes | Apache-2.0 | today |
| 29 | Guides adding a new built-in enricher to the evlog package, covering the source, tests, docs, README, a related skill and a changeset. | evloghq/ | 1.9k | — | ~1.7k | Automated safety check: Pass | MIT | today |
| 30 | AWS Bedrock AgentCore comprehensive expert for deploying and managing AI agents at scale. | zxkane/ | 367 | 1 repo | ~2.5k | Automated safety check: Pass | MIT | 3 mo ago |
| 31 | Implement distributed tracing with Jaeger and Tempo to track requests across microservices and identify performance bottlenecks. | wshobson/ | 40k | 11 repos | ~527 | Automated safety check: Pass | MIT | 2 days ago |
| 32 | Triage OpenSEO production errors in Cloudflare Workers Observability — verified query recipes, counting gotchas, and a known-noise filter list applied automatically. | every-app/ | 23k | — | ~1.7k | Automated safety check: Pass | MIT | today |
| 33 | This skill should be used when the user asks to "add tracking", "add a PostHog event", "change telemetry consent", "instrument onboarding", "debug analytics", or changes telemetry.ts… | OpenHands/ | 90k | — | ~305 | Automated safety check: Pass | MIT | today |
| 34 | Agentforce session tracing extraction and analysis. An agent skill from Jaganpro/sf-skills. | Jaganpro/ | 424 | — | ~1.8k | Automated safety check: Pass | MIT | 5 mo ago |
| 35 | Give the human an off switch and a cost meter for the coding agents on this machine, using ClawMetry. | vivekchand/ | 424 | — | ~1.1k | Automated safety check: Pass | MIT | today |
| 36 | Build production-ready AI agent backends using the CloudBase Agent Python SDK — create agents with LangGraph/CrewAI/LlamaIndex, serve them via FastAPI with AG-UI protocol streaming +… | TencentCloudBase/ | 1.1k | 2 repos | ~2.9k | Automated safety check: Notes | MIT | yesterday |
| 37 | Correlates a Claude Code session's local transcript with a model router's production cloud logs to explain why a specific response rendered the way it did. | weave-os/ | 5.6k | — | ~2.9k | Automated safety check: Pass | Apache-2.0 | today |
| 38 | 38.Agentmeasure Check whether agent telemetry preserves measurement semantics. | roy-tong/ | 218 | — | ~753 | Automated safety check: Pass | MIT | 2 days ago |
| 39 | Adds telemetry events to the Warp codebase through its trait-based system, after agreeing with you what to track and why. | warpdotdev/ | 65k | 1 repo | ~1.3k | Automated safety check: Pass | AGPL-3.0 | today |
| 40 | 40.Effect TS Write idiomatic Effect v4 TypeScript following official best practices from effect-solutions and the Effect source. | mattiacerutti/ | 187 | — | ~2.8k | Automated safety check: Pass | MIT | today |
| 41 | Checks the health of an OmniRoute gateway: provider circuit breakers, latency percentiles, budget guard alerts, connection cooldowns and model lockouts. | diegosouzapw/ | 74k | 1 repo | ~1.2k | Automated safety check: Pass | MIT | today |
| 42 | Analyze and instrument repositories for Traceway observability. | tracewayapp/ | 1.6k | — | ~17k | Automated safety check: Notes | MIT | yesterday |
| 43 | Add Temps analytics to React applications with comprehensive tracking capabilities including page views, custom events, scroll tracking, engagement monitoring, session recording, and Web Vitals… | gotempsh/ | 822 | — | ~2.7k | Automated safety check: Pass | Apache-2.0 | today |
| 44 | Review code for correct logging and error handling patterns. | getsentry/ | 917 | — | ~2.6k | Automated safety check: Pass | Unknown | today |
| 45 | This skill should be used when the user wants to "set up tracing", "monitor my ADK agent", "configure logging", "add observability", "debug production traffic", or needs guidance on monitoring… | pifferologo/ | 129 | 1 repo | ~2.5k | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 46 | PostHog AI Observability integration for LangChain (Python). An agent skill from Jwuthri/Tracely-ai. | Jwuthri/ | 1.5k | — | ~2.2k | Automated safety check: Pass | MIT | yesterday |
| 47 | A skill your agent uses when adding Spring AI-specific model observations, token usage, latency, externally configured cost attribution, advisor telemetry, or protected prompt and completion logging. | rrezartprebreza/ | 296 | 1 repo | ~1.6k | Automated safety check: Pass | MIT | 16 days ago |
| 48 | AWS cost optimization, monitoring, and operational excellence expert. | zxkane/ | 367 | 1 repo | ~2.4k | Automated safety check: Pass | MIT | 3 mo ago |
Questions, answered from the data.
What is the best observability skill?
Vercel Optimize Audit (official) from vercel-labs/agent-skills ranks first of the 562 observability skills listed here, with the highest score: its repository has 32k GitHub stars, 8 other GitHub owners carry a copy, its SKILL.md loads about 4.3k tokens and it passes the automated safety check with no findings. Next come Sentry v8 Error Tracking and Kubeshark Installer.
Which observability skills are official?
82 of the 562 observability skills are official, published by the vendor's own GitHub organization: Vercel Optimize Audit, Apm Integrations, Logging Observability, Agent Platform Alert Configuration, Redis Observability and 77 more.
How are these skills ranked?
By Skill Navigator score, which combines the GitHub stars of the skill's repository (shared across that repo's skills and discounted for large collections), how many other GitHub owners carry a copy of the skill, and automated SKILL.md quality checks, minus penalties for safety-check warnings and for each further skill from the same repository. Skills that fail the safety check are not listed.
Explore related skills
More topics in DevOps & Cloud
- Deployment1,152
- CI/CD921
- Containers723
- Container orchestration519
- Infrastructure as code351
- Monitoring and alerting333
- Secrets management318
- Runbooks and postmortems277
- Incident response271
- Cloud networking230
- Backup and disaster recovery170
- Site reliability engineering150
- Cloud architecture114
- MLOps101
- Cloud cost optimization93
- GitOps87
- Linux administration75
- Platform engineering51
- Chaos engineering28