Topic · DevOps & Cloud
Best observability skills, page 4
Observability skills, ranked
Ranked by score. Sort bymost stars,trending,newest,recently updated
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 145 | Performs a comprehensive Amazon ECS operations review across the 6 review pillars (Resiliency & HA, Observability, Security, Operations, Performance, Additional Analysis) using read-only AWS APIs… | aws/ | 102 | — | ~4.8k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 146 | Product analytics with posthog. An agent skill from langfuse/langfuse. | langfuse/ | 36k | — | ~3.2k | Automated safety check: Pass | Unknown | today |
| 147 | Decide whether and how errors report to Sentry. An agent skill from langfuse/langfuse. | langfuse/ | 36k | — | ~3k | Automated safety check: Pass | Unknown | today |
| 148 | Run any Python LLM agent as an Agent Observability experiment using the public agento11y.experiments package: define a test suite, run an existing agent through typed trials, bind or record… | grafana/ | 127 | — | ~2k | Automated safety check: Pass | Apache-2.0 | today |
| 149 | Add or change OpenTelemetry tracing, metrics, and structured logging in the .NET API (apps/api) — OTLP exporter setup, custom ActivitySource spans, IMeterFactory metrics, log/trace correlation. | atherio-danp/ | 109 | — | ~1.3k | Automated safety check: Notes | No licence | 2 mo ago |
| 150 | Investigate agent activity and JSONL/gzip evidence. An agent skill from QuentinCody/interlinked-cli. | QuentinCody/ | 178 | — | ~10k | Automated safety check: Pass | MIT | 6 days ago |
| 151 | Finds the metrics and labels behind a Prometheus series explosion using a connected Prometheus MCP server, then proposes relabeling, dropping or recording-rule fixes with measured impact. | prometheus/ | 120 | — | ~547 | Automated safety check: Pass | Apache-2.0 | 4 days ago |
| 152 | Build safe, version-pinned telemetrygen commands for synthetic OTLP traces, metrics, and logs. | ollygarden/ | 106 | — | ~2.2k | Automated safety check: Pass | Apache-2.0 | today |
| 153 | LiteLLM-RS Observability Architecture. An agent skill from majiayu000/litellm-rs. | majiayu000/ | 117 | — | ~1.3k | Automated safety check: Pass | MIT | today |
| 154 | Guides Docker, CI/CD pipelines, deployment strategies, infrastructure as code, and observability setup. | CloudAI-X/ | 1.4k | — | ~2.7k | Automated safety check: Notes | MIT | 2 days ago |
| 155 | Troubleshooting guide for NanoClaw's containerized agents: where the logs are, how the two session databases show message flow, and how to raise the log level. | nanocoai/ | 31k | — | ~3.8k | Automated safety check: Notes | MIT | 2 days ago |
| 156 | 156.Error Handler Design error handling, structured logging, and observability with OpenTelemetry (traces, metrics, logs), error classification, recovery patterns (retry with jitter, circuit breaker, bulkhead… | EliasOulkadi/ | 114 | — | ~3.6k | Automated safety check: Notes | MIT | 4 days ago |
| 157 | 157.Self Ops Runtime self-diagnosis knowledge: database layout, observability APIs, timeout hierarchy, restart procedure, failure replay. | spytensor/ | 452 | — | ~767 | Automated safety check: Pass | MIT | 2 mo ago |
| 158 | 158.Dotnet Devops Configures .NET CI/CD pipelines (GitHub Actions with setup-dotnet, NuGet cache, reusable workflows; Azure DevOps with DotNetCoreCLI, templates, multi-stage), containerization (multi-stage… | novotnyllc/ | 233 | — | ~1.1k | Automated safety check: Pass | MIT | yesterday |
| 159 | Enhancement-overlay skill — the DECISION + WIRING layer for LM observability that the single-backend skills [[langsmith]], [[phoenix]], [[mlflow]] do NOT cover. | agentsope/ | 466 | — | ~4.4k | Automated safety check: Pass | MIT | today |
| 160 | Implement and quality-check OpenTelemetry metric instrumentation in Kibana code that uses @kbn/metrics. | elastic/ | 21k | — | ~5.1k | Automated safety check: Pass | Unknown | today |
| 161 | Profiles a slow command or runtime you own with the right profiler for its language, using uninstrumented baseline timings and keeping raw profiling evidence local and private. | loopx-project/ | 6.2k | — | ~880 | Automated safety check: Pass | Apache-2.0 | today |
| 162 | Build production-ready monitoring, logging, and tracing systems. | davila7/ | 32k | 8 repos | ~3.2k | Automated safety check: Pass | MIT | today |
| 163 | Installs and configures the OpenSearch extension that stores KubeSphere logs, events, audit data and notification history, with its dashboard and curator options. | kubesphere/ | 17k | — | ~2.8k | Automated safety check: Pass | Unknown | 2 mo ago |
| 164 | Installs and configures KubeSphere's WizTelemetry Data Pipeline, built on Vector, which collects and routes logs, audit events, events and notifications to OpenSearch. | kubesphere/ | 17k | — | ~2.6k | Automated safety check: Pass | Unknown | 2 mo ago |
| 165 | Installs and configures the WizTelemetry Auditing extension for KubeSphere, which collects and stores Kubernetes audit events, and covers the audit query API. | kubesphere/ | 17k | — | ~2.2k | Automated safety check: Pass | Unknown | 2 mo ago |
| 166 | Installs and configures the WizTelemetry Events extension for KubeSphere, which exports Kubernetes events for storage, with dependency checks and the event query API. | kubesphere/ | 17k | — | ~1.8k | Automated safety check: Pass | Unknown | 2 mo ago |
| 167 | Installs and configures WizTelemetry Logging for KubeSphere, with container log and optional disk log collection, dependency checks and the log query API. | kubesphere/ | 17k | — | ~2.3k | Automated safety check: Pass | Unknown | 2 mo ago |
| 168 | Installs and configures WizTelemetry Tracing for KubeSphere, an OpenTelemetry-based tracing extension, and covers its components, query API and auto-instrumentation. | kubesphere/ | 17k | — | ~4.6k | Automated safety check: Pass | Unknown | 2 mo ago |
| 169 | Add purposeful debug logging to improve observability without changing behavior. | dmitriiweb/ | 111 | — | ~426 | Automated safety check: Pass | MIT | 9 mo ago |
| 170 | Generates valid PromQL queries for Cloud Monitoring metrics from metric descriptors and resource parameters, with a validator script and error-recovery notes. | google/ | 21k | — | ~2.6k | Automated safety check: Pass | Apache-2.0 | today |
| 171 | Add Pydantic Logfire observability to application code — traces, logs, metrics, and AI/agent spans. | pydantic/ | 140 | — | ~6.1k | Automated safety check: Pass | MIT | 7 days ago |
| 172 | 172.Dt Obs Problems DAVIS problem analysis including root cause identification, impact assessment, and correlation with other telemetry. | Dynatrace/ | 162 | — | ~4.6k | Automated safety check: Pass | Apache-2.0 | 7 days ago |
| 173 | 173.Auto Harness Diagnose and strengthen a repository's harness layer: AGENTS.md rules, knowledge layout, architecture boundaries, lint and type gates, API and generated-client contracts, test scaffolding… | PacificStudio/ | 268 | — | ~1.2k | Automated safety check: Pass | Apache-2.0 | 2 mo ago |
| 174 | Designs distributed systems: bounded-context service boundaries, sync and async communication, data ownership, resilience, tracing and rollout strategy. | Jeffallan/ | 12k | — | ~1.8k | Automated safety check: Pass | MIT | 5 days ago |
| 175 | Sets up application monitoring: structured logs, Prometheus metrics, OpenTelemetry tracing, Grafana dashboards, alert rules and load tests with k6 or Artillery. | Jeffallan/ | 12k | — | ~1.6k | Automated safety check: Pass | MIT | 5 days ago |
| 176 | Audits Prometheus recording and alerting rules through a connected Prometheus MCP server, finds gaps and noisy alerts, and drafts improved rule-group YAML. | prometheus/ | 120 | — | ~765 | Automated safety check: Pass | Apache-2.0 | 4 days ago |
| 177 | 177.Otel Upgrade Assess OpenTelemetry package and Collector upgrades across ecosystems. | ollygarden/ | 106 | — | ~1.1k | Automated safety check: Pass | Apache-2.0 | today |
| 178 | 178.Error Handling Implements error handling patterns, structured logging, retry strategies, circuit breakers, and graceful degradation. | CloudAI-X/ | 1.4k | — | ~3.4k | Automated safety check: Pass | MIT | 2 days ago |
| 179 | Traces an opaque production error in an OpenWork build to its cause using local server logs and Sentry, names the regressing PR and files a report. | different-ai/ | 24k | — | ~803 | Automated safety check: Pass | Unknown | today |
| 180 | Operational controls for long-lived or cloud-hosted agent systems — runtime lifecycle (start, pause, stop, restart), observability (logs, metrics, traces), least-privilege safety scopes and kill… | affaan-m/ | 276k | 4 repos | ~384 | Automated safety check: Pass | MIT | 4 days ago |
| 181 | Installs clidash, a dependency-free web dashboard that turns the JSON resource listings of a CLI such as NanoClaw's ncl into read-only tabs and tables. | nanocoai/ | 31k | — | ~1.6k | Automated safety check: Pass | MIT | 2 days ago |
| 182 | Adds a local monitoring dashboard to NanoClaw by installing its npm package and a pusher module that sends periodic JSON snapshots of agent activity. | nanocoai/ | 31k | — | ~1.3k | Automated safety check: Notes | MIT | 2 days ago |
| 183 | Opt the current project into local Claude Code usage telemetry, or verify its setup. | Eigenwise/ | 278 | — | ~2.9k | Automated safety check: Warn | MIT | today |
| 184 | Installs and configures the WizTelemetry Notification extension for KubeSphere: channel setup, alert routing by tenant labels, silences and troubleshooting. | kubesphere/ | 17k | — | ~6.1k | Automated safety check: Pass | Unknown | 2 mo ago |
| 185 | Installs the WizTelemetry Ruler extension for KubeSphere and manages event, audit and log alerting rules as RuleGroup and ClusterRuleGroup resources. | kubesphere/ | 17k | — | ~6.5k | Automated safety check: Pass | Unknown | 2 mo ago |
| 186 | 186.Logging Observability for .NET 10 applications. An agent skill from Resgrid/Core. | Resgrid/ | 229 | — | ~1.4k | Automated safety check: Pass | Apache-2.0 | today |
| 187 | 187.Debug Mastery Systematic debugging methodology with 4-phase process, root cause tracing, and elite observability standards. | xenitV1/ | 229 | — | ~2.5k | Automated safety check: Notes | MIT | 8 mo ago |
| 188 | A skill your agent uses when instrumenting or inspecting TRL training runs with Trackio, run names, metric schemas, dashboards, logs, grep or ripgrep, SFTP, Hugging Face Job logs, remote artifacts… | burtenshaw/ | 153 | — | ~351 | Automated safety check: Pass | Apache-2.0 | 26 days ago |
| 189 | 189.Neon Overview of Neon, a complete set of cloud backend primitives for apps and agents, spanning Lakebase Postgres, Auth, the Data API, Object Storage, Compute Functions, and the AI Gateway. | smontlouis/ | 171 | — | ~7.1k | Automated safety check: Notes | GPL-3.0 | today |
| 190 | Use early in an AI-agent project — before ship, before real traffic — to decide which evaluations to set up and to scaffold a starter experiment. | grafana/ | 127 | — | ~6.4k | Automated safety check: Pass | Apache-2.0 | today |
| 191 | Answers questions about agent spend, token use, traces, events, errors and tool usage by running read-only SQL against a local TMA1 observability store. | tma1-ai/ | 119 | — | ~5.1k | Automated safety check: Notes | Apache-2.0 | today |
| 192 | Use at review and ship, automatically for feature or architecture changes and on request for error handling, logging, crash reporting, or observability; verify errors reach production monitoring. | KbWen/ | 206 | — | ~938 | Automated safety check: Pass | MIT | 2 days ago |
Explore related skills
Category
More topics in DevOps & Cloud
- Deployment1,266
- CI/CD978
- Containers729
- Container orchestration546
- Infrastructure as code379
- Monitoring and alerting352
- Secrets management334
- Runbooks and postmortems324
- Incident response302
- Cloud networking227
- Backup and disaster recovery185
- Site reliability engineering150
- Cloud architecture118
- MLOps93
- Cloud cost optimization91
- GitOps88
- Linux administration73
- Platform engineering43
- Chaos engineering24