Topic · DevOps & Cloud
Best monitoring and alerting skills, page 3
Monitoring and alerting skills, ranked
Ranked by score. Sort bymost stars,trending,newest,recently updated
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 97 | Installs and configures the WizTelemetry Events extension for KubeSphere, which exports Kubernetes events for storage, with dependency checks and the event query API. | kubesphere/ | 17k | — | ~1.8k | Automated safety check: Pass | Unknown | 2 mo ago |
| 98 | Installs and configures WizTelemetry Logging for KubeSphere, with container log and optional disk log collection, dependency checks and the log query API. | kubesphere/ | 17k | — | ~2.3k | Automated safety check: Pass | Unknown | 2 mo ago |
| 99 | 99.Expert Ops 基础设施运维专家入口。用于 Codex CLI 的 $expert-ops 调用. An agent skill from ReJeCtAll/ExpertTeam-Codex. | ReJeCtAll/ | 113 | — | ~625 | Automated safety check: Pass | MIT | 3 mo ago |
| 100 | Improve and validate the Kaniop Grafana dashboard against repository metrics and the grigri live cluster. | pando85/ | 130 | — | ~987 | Automated safety check: Pass | AGPL-3.0 | today |
| 101 | 101.Playbook Lookup Query past incident resolutions from the knowledge base. An agent skill from papadopouloskyriakos/agentic-chatops. | papadopouloskyriakos/ | 107 | — | ~363 | Automated safety check: Notes | No licence | 2 days ago |
| 102 | Generates valid PromQL queries for Cloud Monitoring metrics from metric descriptors and resource parameters, with a validator script and error-recovery notes. | google/ | 21k | — | ~2.6k | Automated safety check: Pass | Apache-2.0 | today |
| 103 | Set up Prometheus and Grafana monitoring for AQUA vLLM model deployments on OCI. | oracle/ | 125 | — | ~1.5k | Automated safety check: Pass | UPL-1.0 | 1 mo ago |
| 104 | Sets up application monitoring: structured logs, Prometheus metrics, OpenTelemetry tracing, Grafana dashboards, alert rules and load tests with k6 or Artillery. | Jeffallan/ | 12k | — | ~1.6k | Automated safety check: Pass | MIT | 5 days ago |
| 105 | 105.SRE Engineer Defines SLIs, SLOs and error budgets, and sets up golden-signal monitoring, blameless postmortems, toil automation and chaos experiments for production systems. | Jeffallan/ | 12k | — | ~1.7k | Automated safety check: Pass | MIT | 5 days ago |
| 106 | Audits Prometheus recording and alerting rules through a connected Prometheus MCP server, finds gaps and noisy alerts, and drafts improved rule-group YAML. | prometheus/ | 118 | — | ~765 | Automated safety check: Pass | Apache-2.0 | 4 days ago |
| 107 | 107.Analyze Run Read and compare finished harness runs - frame-time tails, utilization, chunk latency, JFR waits, load phases - and look them up in Grafana. | xD3I/ | 114 | — | ~536 | Automated safety check: Pass | No licence | today |
| 108 | 108.Apm Usage Reference for APM (Agent Package Manager) — apm.yml syntax, install / uninstall / update commands, target detection, lockfile workflow. | mizchi/ | 356 | — | ~1.9k | Automated safety check: Pass | No licence | 6 days ago |
| 109 | Manage Grafana Cloud accounts — organizations, stacks, RBAC roles and assignments, SSO/SAML/OAuth/GitHub auth, service accounts for CI/CD, user invites, team membership, and API-driven provisioning. | grafana/ | 279 | — | ~1.5k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 110 | Adds a local monitoring dashboard to NanoClaw by installing its npm package and a pusher module that sends periodic JSON snapshots of agent activity. | nanocoai/ | 31k | — | ~1.3k | Automated safety check: Notes | MIT | 2 days ago |
| 111 | Installs and configures the WizTelemetry Notification extension for KubeSphere: channel setup, alert routing by tenant labels, silences and troubleshooting. | kubesphere/ | 17k | — | ~6.1k | Automated safety check: Pass | Unknown | 2 mo ago |
| 112 | Installs the WizTelemetry Ruler extension for KubeSphere and manages event, audit and log alerting rules as RuleGroup and ClusterRuleGroup resources. | kubesphere/ | 17k | — | ~6.5k | Automated safety check: Pass | Unknown | 2 mo ago |
| 113 | 113.Neon Overview of Neon, a complete set of cloud backend primitives for apps and agents, spanning Lakebase Postgres, Auth, the Data API, Object Storage, Compute Functions, and the AI Gateway. | smontlouis/ | 171 | — | ~7.1k | Automated safety check: Notes | GPL-3.0 | today |
| 114 | A skill your agent uses whenever a pull request is opened, reopened, or synchronized in microsoft/apm to assess whether and how the documentation corpus must change to stay truthful with the… | microsoft/ | 4k | — | ~3k | Automated safety check: Pass | MIT | today |
| 115 | 115.Logql Generator Generate LogQL queries, log stream selectors, metric queries, and alerting rules for Grafana Loki. | akin-ozer/ | 319 | — | ~3.8k | Automated safety check: Pass | Apache-2.0 | 2 mo ago |
| 116 | 116.Live Debug Debug the running local stack with traces, logs, and a shared headless browser. | macro-inc/ | 4.6k | — | ~2.4k | Automated safety check: Notes | AGPL-3.0 | today |
| 117 | A skill your agent uses for any question involving telemetry data: "investigate an issue", "debug a problem", "find out why something is slow", "check error rates", "analyze user behavior"… | coralogix/ | 121 | — | ~2.6k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 118 | Azure Application Insights SDK for .NET. An agent skill from microsoft/skills. | microsoft/ | 3.1k | 6 repos | ~4.6k | Automated safety check: Pass | MIT | 2 days ago |
| 119 | Finds where a Prometheus metric stops existing, whether at the target, the scrape, relabeling or the query, using the tools of a connected Prometheus MCP server. | prometheus/ | 118 | — | ~587 | Automated safety check: Pass | Apache-2.0 | 4 days ago |
| 120 | Monitor robot fleet telemetry via Azure IoT Operations, drift detection, Grafana dashboards, and Fabric analytics | microsoft/ | 123 | — | ~598 | Automated safety check: Pass | MIT | yesterday |
| 121 | Hands-on playbook for Windows 11 disk cleanup, dev-machine optimization, and proactive health alerting. | CodeAlive-AI/ | 157 | — | ~4k | Automated safety check: Pass | MIT | 2 days ago |
| 122 | A skill your agent uses when the user asks to "write a validator", "add validation", "implement admission control", "write a mutating webhook", "add a mutation handler", "validate incoming… | grafana/ | 279 | — | ~1.7k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 123 | Watches a live app after a deploy for console errors, performance regressions and page failures, comparing periodic screenshots against pre-deploy baselines. | garrytan/ | 136k | — | ~13k | Automated safety check: Notes | MIT | today |
| 124 | Create and manage Kibana connectors for Slack, PagerDuty, Jira, webhooks, and more via REST API or Terraform. | aspectrr/ | 405 | — | ~2k | Automated safety check: Pass | MIT | 5 mo ago |
| 125 | 125.Cloud Devops Cloud infrastructure and DevOps workflow covering AWS, Azure, GCP, Kubernetes, Terraform, CI/CD, monitoring, and cloud-native development. | davila7/ | 32k | 4 repos | ~1.4k | Automated safety check: Pass | MIT | today |
| 126 | Automate PagerDuty tasks via Rube MCP (Composio): manage incidents, services, schedules, escalation policies, and on-call rotations. | davepoon/ | 3.6k | 7 repos | ~2.6k | Automated safety check: Pass | MIT | 2 days ago |
| 127 | 127.Datasource Check 检查 Prometheus 数据源的连通性、数据延迟和指标采集健康度。 | kubehan/ | 125 | — | ~229 | Automated safety check: Pass | No licence | 20 days ago |
| 128 | Activate when code touches token management, credential resolution, git auth flows, GITHUBAPMPAT, ADOAPMPAT, AuthResolver, HostInfo, AuthContext, or any remote host authentication -- even if 'auth'… | microsoft/ | 4k | — | ~756 | Automated safety check: Pass | MIT | today |
| 129 | A skill your agent uses to post or patch ONE GitHub comment for a microsoft/apm autopilot run. | microsoft/ | 4k | — | ~883 | Automated safety check: Pass | MIT | today |
| 130 | A skill your agent uses when editing or creating CLI output, logging, warnings, error messages, progress indicators, or diagnostic summaries in the APM codebase. | microsoft/ | 4k | — | ~3.9k | Automated safety check: Pass | MIT | today |
| 131 | A skill your agent uses when the docs-impact-classifier returns a structural verdict, signalling that the documentation TOC must change to accommodate the PR. | microsoft/ | 4k | — | ~1.5k | Automated safety check: Pass | MIT | today |
| 132 | A skill your agent uses to classify the documentation impact of a pull request diff, returning one of three verdicts -- no-change, in-place edit, or structural change -- with bounded LLM cost. | microsoft/ | 4k | — | ~1.8k | Automated safety check: Pass | MIT | today |
| 133 | Review and tune Prometheus configuration and performance. An agent skill from prometheus/prometheus-mcp. | prometheus/ | 118 | — | ~724 | Automated safety check: Pass | Apache-2.0 | 4 days ago |
| 134 | Guides Qdrant monitoring setup including Prometheus scraping, health probes, Hybrid Cloud metrics, alerting, and log centralization. | qdrant/ | 253 | 2 repos | ~874 | Automated safety check: Pass | Apache-2.0 | yesterday |
| 135 | Build a unified telemetry pipeline with Grafana Alloy — one OpenTelemetry-compatible binary that collects metrics, logs, traces, and profiles and ships to Grafana Cloud / Prometheus / Loki / Tempo /… | grafana/ | 279 | — | ~1.3k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 136 | A skill your agent uses to triage ONE microsoft/apm issue already selected by autopilot-issue-triage-scheduler. | microsoft/ | 4k | — | ~6.4k | Automated safety check: Pass | MIT | today |
| 137 | A skill your agent uses to run a multi-persona expert advisory review on a labelled pull request in microsoft/apm. | microsoft/ | 4k | — | ~9.1k | Automated safety check: Pass | MIT | today |
| 138 | 138.Prometheus Prometheus monitoring expert for PromQL, alerting rules, Grafana dashboards, and observability | RightNow-AI/ | 18k | — | ~738 | Automated safety check: Pass | Apache-2.0 | 3 mo ago |
| 139 | Go observability — always-on production signals: slog logging, Prometheus metrics, OpenTelemetry tracing, pprof profiling, alerting, Grafana. | context-labs/ | 1.1k | 1 repo | ~3.3k | Automated safety check: Pass | MIT | 3 days ago |
| 140 | Read what FastLLM has been doing — usage records, time-series aggregates, the configuration audit trail, Prometheus metrics, control-plane health, and per-replica fleet status. | azrtydxb/ | 108 | — | ~916 | Automated safety check: Pass | Apache-2.0 | 3 days ago |
| 141 | Monitor AI agent health, detect anomalies, set up alerting, and maintain observability dashboards for production multi-agent systems. | cosmicstack-labs/ | 476 | — | ~2.7k | Automated safety check: Pass | MIT | 1 mo ago |
| 142 | 142.Dd Apm APM - install, onboard, instrument, enable, set up, configure, traces, services, dependencies, performance analysis, Data Streams Monitoring (DSM), queue lag, pipeline latency. | datadog-labs/ | 177 | — | ~2k | Automated safety check: Pass | MIT | today |
| 143 | 143.Agent Install Install the Datadog Agent on Kubernetes using the Datadog Operator — required before enabling Single Step Instrumentation (SSI), which automatically instruments applications for APM without code… | datadog-labs/ | 177 | — | ~2.1k | Automated safety check: Warn | MIT | today |
| 144 | A skill your agent uses for read-only incident, occupancy, speed, place, and analytics-sensor questions through the project-local VSS CLI. | NVIDIA-AI-Blueprints/ | 1.9k | — | ~1.5k | Automated safety check: Pass | Apache-2.0 | today |
Explore related skills
Category
More topics in DevOps & Cloud
- Deployment1,264
- CI/CD977
- Containers731
- Observability636
- Container orchestration542
- Infrastructure as code377
- Secrets management335
- Runbooks and postmortems324
- Incident response302
- Cloud networking225
- Backup and disaster recovery183
- Site reliability engineering152
- Cloud architecture118
- MLOps100
- Cloud cost optimization92
- GitOps89
- Linux administration70
- Platform engineering45
- Chaos engineering25