Search
Monitoring and alerting
Skills
Sort:BestMost starsTrending todayTrending this weekTrending this monthNewestRecently updatedName
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 97 | Checks OmniRoute server health, per-component status, and circuit breakers from the command line, with a live watch dashboard. | diegosouzapw/ | 75k | — | ~327 | Automated safety check: Pass | MIT | today |
| 98 | Runs a phased health audit of a Cisco ACI fabric through MCP tools: node status, links, tenant and policy review, faults and endpoint learning. | automateyournetwork/ | 677 | — | ~2.9k | Automated safety check: Pass | Apache-2.0 | today |
| 99 | A skill your agent uses to triage ONE microsoft/apm pull request already selected by autopilot-pr-triage-scheduler. | microsoft/ | 4k | — | ~1.6k | Automated safety check: Pass | MIT | yesterday |
| 100 | Installs and configures the WizTelemetry Events extension for KubeSphere, which exports Kubernetes events for storage, with dependency checks and the event query API. | kubesphere/ | 17k | — | ~1.8k | Automated safety check: Pass | Unknown | 2 mo ago |
| 101 | Installs and configures WizTelemetry Logging for KubeSphere, with container log and optional disk log collection, dependency checks and the log query API. | kubesphere/ | 17k | — | ~2.3k | Automated safety check: Pass | Unknown | 2 mo ago |
| 102 | Improve and validate the Kaniop Grafana dashboard against repository metrics and the grigri live cluster. | pando85/ | 132 | — | ~987 | Automated safety check: Pass | AGPL-3.0 | today |
| 103 | 103.Expert Ops 基础设施运维专家入口。用于 Codex CLI 的 $expert-ops 调用. An agent skill from ReJeCtAll/ExpertTeam-Codex. | ReJeCtAll/ | 113 | — | ~625 | Automated safety check: Pass | MIT | 3 mo ago |
| 104 | 104.Playbook Lookup Query past incident resolutions from the knowledge base. An agent skill from papadopouloskyriakos/agentic-chatops. | papadopouloskyriakos/ | 107 | — | ~363 | Automated safety check: Notes | No licence | 5 days ago |
| 105 | Generates valid PromQL queries for Cloud Monitoring metrics from metric descriptors and resource parameters, with a validator script and error-recovery notes. | google/ | 21k | — | ~2.6k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 106 | Set up Prometheus and Grafana monitoring for AQUA vLLM model deployments on OCI. | oracle/ | 125 | — | ~1.5k | Automated safety check: Pass | UPL-1.0 | 1 mo ago |
| 107 | Sets up application monitoring: structured logs, Prometheus metrics, OpenTelemetry tracing, Grafana dashboards, alert rules and load tests with k6 or Artillery. | Jeffallan/ | 12k | — | ~1.6k | Automated safety check: Pass | MIT | 8 days ago |
| 108 | 108.SRE Engineer Defines SLIs, SLOs and error budgets, and sets up golden-signal monitoring, blameless postmortems, toil automation and chaos experiments for production systems. | Jeffallan/ | 12k | — | ~1.7k | Automated safety check: Pass | MIT | 8 days ago |
| 109 | Audits Prometheus recording and alerting rules through a connected Prometheus MCP server, finds gaps and noisy alerts, and drafts improved rule-group YAML. | prometheus/ | 121 | — | ~765 | Automated safety check: Pass | Apache-2.0 | yesterday |
| 110 | 110.Apm Usage Reference for APM (Agent Package Manager) — apm.yml syntax, install / uninstall / update commands, target detection, lockfile workflow. | mizchi/ | 360 | — | ~1.9k | Automated safety check: Pass | No licence | 9 days ago |
| 111 | Manage Grafana Cloud accounts — organizations, stacks, RBAC roles and assignments, SSO/SAML/OAuth/GitHub auth, service accounts for CI/CD, user invites, team membership, and API-driven provisioning. | grafana/ | 282 | — | ~1.5k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 112 | Installs clidash, a dependency-free web dashboard that turns the JSON resource listings of a CLI such as NanoClaw's ncl into read-only tabs and tables. | nanocoai/ | 31k | — | ~1.6k | Automated safety check: Pass | MIT | yesterday |
| 113 | Adds a local monitoring dashboard to NanoClaw by installing its npm package and a pusher module that sends periodic JSON snapshots of agent activity. | nanocoai/ | 31k | — | ~1.3k | Automated safety check: Notes | MIT | yesterday |
| 114 | Installs and configures the WizTelemetry Notification extension for KubeSphere: channel setup, alert routing by tenant labels, silences and troubleshooting. | kubesphere/ | 17k | — | ~6.1k | Automated safety check: Pass | Unknown | 2 mo ago |
| 115 | Installs the WizTelemetry Ruler extension for KubeSphere and manages event, audit and log alerting rules as RuleGroup and ClusterRuleGroup resources. | kubesphere/ | 17k | — | ~6.5k | Automated safety check: Pass | Unknown | 2 mo ago |
| 116 | 116.Neon Overview of Neon, a complete set of cloud backend primitives for apps and agents, spanning Lakebase Postgres, Auth, the Data API, Object Storage, Compute Functions, and the AI Gateway. | smontlouis/ | 172 | — | ~7.1k | Automated safety check: Notes | GPL-3.0 | today |
| 117 | A skill your agent uses whenever a pull request is opened, reopened, or synchronized in microsoft/apm to assess whether and how the documentation corpus must change to stay truthful with the… | microsoft/ | 4k | — | ~3k | Automated safety check: Pass | MIT | yesterday |
| 118 | 118.Logql Generator Generate LogQL queries, log stream selectors, metric queries, and alerting rules for Grafana Loki. | akin-ozer/ | 320 | — | ~3.8k | Automated safety check: Pass | Apache-2.0 | 2 mo ago |
| 119 | 119.Live Debug Debug the running local stack with traces, logs, and a shared headless browser. | macro-inc/ | 4.6k | — | ~2.4k | Automated safety check: Notes | AGPL-3.0 | today |
| 120 | 120.Analyze Runtime Find out what a running Mendix app actually does — logs, Prometheus metrics, OpenTelemetry traces and the model catalog, joined across sources. | mendixlabs/ | 129 | — | ~2.8k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 121 | A skill your agent uses for any question involving telemetry data: "investigate an issue", "debug a problem", "find out why something is slow", "check error rates", "analyze user behavior"… | coralogix/ | 121 | — | ~2.6k | Automated safety check: Pass | Apache-2.0 | 4 days ago |
| 122 | Finds where a Prometheus metric stops existing, whether at the target, the scrape, relabeling or the query, using the tools of a connected Prometheus MCP server. | prometheus/ | 121 | — | ~587 | Automated safety check: Pass | Apache-2.0 | yesterday |
| 123 | Monitor robot fleet telemetry via Azure IoT Operations, drift detection, Grafana dashboards, and Fabric analytics | microsoft/ | 126 | — | ~598 | Automated safety check: Pass | MIT | yesterday |
| 124 | Hands-on playbook for Windows 11 disk cleanup, dev-machine optimization, and proactive health alerting. | CodeAlive-AI/ | 159 | — | ~4k | Automated safety check: Pass | MIT | 3 days ago |
| 125 | 125.Analyze Run Read and compare finished harness runs - frame-time tails, utilization, chunk latency, JFR waits, load phases - and look them up in Grafana. | xD3I/ | 116 | — | ~536 | Automated safety check: Pass | No licence | today |
| 126 | A skill your agent uses when the user asks to "write a validator", "add validation", "implement admission control", "write a mutating webhook", "add a mutation handler", "validate incoming… | grafana/ | 282 | — | ~1.7k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 127 | Create and manage Kibana connectors for Slack, PagerDuty, Jira, webhooks, and more via REST API or Terraform. | aspectrr/ | 405 | — | ~2k | Automated safety check: Pass | MIT | 5 mo ago |
| 128 | Watches a live app after a deploy for console errors, performance regressions and page failures, comparing periodic screenshots against pre-deploy baselines. | garrytan/ | 136k | — | ~13k | Automated safety check: Notes | MIT | today |
| 129 | Azure Application Insights SDK for .NET. An agent skill from microsoft/skills. | microsoft/ | 3.1k | 5 repos | ~4.6k | Automated safety check: Pass | MIT | yesterday |
| 130 | 130.Cloud Devops Cloud infrastructure and DevOps workflow covering AWS, Azure, GCP, Kubernetes, Terraform, CI/CD, monitoring, and cloud-native development. | davila7/ | 33k | 4 repos | ~1.4k | Automated safety check: Pass | MIT | today |
| 131 | Automate PagerDuty tasks via Rube MCP (Composio): manage incidents, services, schedules, escalation policies, and on-call rotations. | davepoon/ | 3.6k | 7 repos | ~2.6k | Automated safety check: Pass | MIT | yesterday |
| 132 | 132.Datasource Check 检查 Prometheus 数据源的连通性、数据延迟和指标采集健康度。 | kubehan/ | 125 | — | ~229 | Automated safety check: Pass | No licence | 23 days ago |
| 133 | Activate when code touches token management, credential resolution, git auth flows, GITHUBAPMPAT, ADOAPMPAT, AuthResolver, HostInfo, AuthContext, or any remote host authentication -- even if 'auth'… | microsoft/ | 4k | — | ~756 | Automated safety check: Pass | MIT | yesterday |
| 134 | A skill your agent uses to post or patch ONE GitHub comment for a microsoft/apm autopilot run. | microsoft/ | 4k | — | ~883 | Automated safety check: Pass | MIT | yesterday |
| 135 | A skill your agent uses when editing or creating CLI output, logging, warnings, error messages, progress indicators, or diagnostic summaries in the APM codebase. | microsoft/ | 4k | — | ~3.9k | Automated safety check: Pass | MIT | yesterday |
| 136 | A skill your agent uses when the docs-impact-classifier returns a structural verdict, signalling that the documentation TOC must change to accommodate the PR. | microsoft/ | 4k | — | ~1.5k | Automated safety check: Pass | MIT | yesterday |
| 137 | A skill your agent uses to classify the documentation impact of a pull request diff, returning one of three verdicts -- no-change, in-place edit, or structural change -- with bounded LLM cost. | microsoft/ | 4k | — | ~1.8k | Automated safety check: Pass | MIT | yesterday |
| 138 | Review and tune Prometheus configuration and performance. An agent skill from prometheus/prometheus-mcp. | prometheus/ | 121 | — | ~724 | Automated safety check: Pass | Apache-2.0 | yesterday |
| 139 | Build a unified telemetry pipeline with Grafana Alloy — one OpenTelemetry-compatible binary that collects metrics, logs, traces, and profiles and ships to Grafana Cloud / Prometheus / Loki / Tempo /… | grafana/ | 282 | — | ~1.3k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 140 | Guides Qdrant monitoring setup including Prometheus scraping, health probes, Hybrid Cloud metrics, alerting, and log centralization. | qdrant/ | 254 | 2 repos | ~874 | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 141 | A skill your agent uses to triage ONE microsoft/apm issue already selected by autopilot-issue-triage-scheduler. | microsoft/ | 4k | — | ~6.4k | Automated safety check: Pass | MIT | yesterday |
| 142 | A skill your agent uses to run a multi-persona expert advisory review on a labelled pull request in microsoft/apm. | microsoft/ | 4k | — | ~9.1k | Automated safety check: Pass | MIT | yesterday |
| 143 | 143.Prometheus Prometheus monitoring expert for PromQL, alerting rules, Grafana dashboards, and observability | RightNow-AI/ | 18k | — | ~738 | Automated safety check: Pass | Apache-2.0 | 3 mo ago |
| 144 | Go observability — always-on production signals: slog logging, Prometheus metrics, OpenTelemetry tracing, pprof profiling, alerting, Grafana. | context-labs/ | 1.1k | 1 repo | ~3.3k | Automated safety check: Pass | MIT | 6 days ago |