Topic · DevOps & Cloud
Best monitoring and alerting skills, page 2
Monitoring and alerting skills, ranked
Ranked by score. Sort bymost stars,trending,newest,recently updated
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 49 | 49.Loki Logs A skill your agent uses when investigating production behaviour from logs — a 500/error in prod, a failing or stuck Celery task, tracing one request/traceid/user across services, or confirming a… | letsrevel/ | 109 | — | ~918 | Automated safety check: Notes | MIT | yesterday |
| 50 | 对远程多实例MySQL数据库执行全方位深度巡检,覆盖基础健康、连接负载、性能慢查询、索引冗余、主从复制、容量空间、账号安全、配置风险八大维度,全自动完成巡检扫描、风险识别、问题定级、优化建议、报告归档与飞书推送,适用于生产/测试所有运行中MySQL实例常态化合规巡检。适用场景:用户要求进行 MySQL 全链路健康检查、MySQL 综合巡检、MySQL 风险扫描、MySQL 性能审计、MySQL… | openocta/ | 166 | — | ~2k | Automated safety check: Pass | MIT | 3 mo ago |
| 51 | Diagnose, troubleshoot, and advise on any Qdrant deployment by loading the latest official Qdrant skills live from skills.qdrant.tech. | qdrant/ | 253 | — | ~1.7k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 52 | Build, modify, and ship Grafana dashboards as JSON via the HTTP API — panel types (timeseries / stat / gauge / table / heatmap / logs / traces / node-graph), gridPos 24-column layout, units… | grafana/ | 279 | 1 repo | ~1.4k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 53 | Bootstrap, create, connect to, operate, secure, scale, upgrade, troubleshoot, inspect, and tear down Alibaba Cloud Container Compute Service (ACS) Agent Sandbox environments. | cinience/ | 397 | — | ~2.7k | Automated safety check: Pass | MIT | 1 mo ago |
| 54 | 当需要为 funboost 创建 Consumer 或 Publisher 的 Mixin 扩展类时使用。触发场景:添加监控、熔断、限流、链路追踪等横切关注点,编写自定义前置/后置处理钩子。关键词:mixin, consumeroverridecls, publisheroverridecls, ConsumerMixin, 自定义消费者, hook, 拦截器, 熔断器, 监控… | ydf0509/ | 892 | — | ~2.1k | Automated safety check: Pass | No licence | 1 mo ago |
| 55 | 55.Oryxos Init 初始化 OryxOS(或同类 JDK 21 + Spring Boot 3.x 企业级单体)的工程地基:Maven 多模块骨架、 结构化日志、Actuator + Prometheus 监控、Spring MVC + 虚拟线程、springdoc OpenAPI、 统一响应体与全局异常/错误码、Google 格式 + 阿里编码规约(Spotless + 阿里 P3C +… | oryx-labs/ | 186 | — | ~1.6k | Automated safety check: Pass | Apache-2.0 | today |
| 56 | A skill your agent uses to run a four-panel adversarial advisory review on any pull request that touches the OpenAPM specification artifact (docs/src/content/docs/specs/openapm-.md), its inline /… | microsoft/ | 4k | — | ~4.9k | Automated safety check: Pass | MIT | today |
| 57 | Help choose, configure, and test local agento11y guard packs for coding-agent tool calls. | grafana/ | 127 | — | ~2.6k | Automated safety check: Notes | Apache-2.0 | today |
| 58 | 58.Graft This repo is indexed by graft/. An agent skill from m4r1k/Eneru. | m4r1k/ | 149 | 1 repo | ~2.3k | Automated safety check: Pass | MIT | today |
| 59 | Write, validate, and optimize PromQL for Prometheus / Grafana Mimir / Grafana Cloud Metrics. | grafana/ | 279 | 1 repo | ~1.1k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 60 | Mandatory pre-release deep review for minor/major releases (X.Y.0 / X.0.0). | m4r1k/ | 149 | — | ~1.9k | Automated safety check: Pass | MIT | today |
| 61 | MUST USE when investigating performance issues on a ClickHouse-managed Postgres instance. | ClickHouse/ | 544 | — | ~1.2k | Automated safety check: Pass | Apache-2.0 | 9 days ago |
| 62 | Implements backend modules from proto definitions for goddess, marksman, and rabbit apps. | aide-family/ | 253 | — | ~4.1k | Automated safety check: Pass | No licence | 3 mo ago |
| 63 | Builds a picture of whether Prometheus itself is healthy and successfully monitoring its targets, covering readiness, firing alerts, target health and TSDB load. | prometheus/ | 118 | — | ~584 | Automated safety check: Pass | Apache-2.0 | 3 days ago |
| 64 | Checks the health of an OmniRoute gateway: provider circuit breakers, latency percentiles, budget guard alerts, connection cooldowns and model lockouts. | diegosouzapw/ | 74k | 1 repo | ~1.2k | Automated safety check: Pass | MIT | today |
| 65 | Set up tracing, metrics and dashboards for Istio, Linkerd and other service meshes, with golden-signal alerts, SLOs and guidance on sampling and cardinality. | wshobson/ | 40k | 9 repos | ~607 | Automated safety check: Pass | MIT | 3 days ago |
| 66 | Sets up Arize Phoenix to trace, evaluate and monitor LLM applications, with instrumentation for OpenAI, LangChain and LlamaIndex and a self-hosted server. | Orchestra-Research/ | 13k | 2 repos | ~2.9k | Automated safety check: Pass | MIT | 3 mo ago |
| 67 | Provides patterns to configure Spring Boot Actuator for production-grade monitoring, health probes, secured management endpoints, and Micrometer metrics across JVM services. | giuseppe-trisciuoglio/ | 355 | — | ~2.2k | Automated safety check: Notes | MIT | 27 days ago |
| 68 | A skill your agent uses to queue maintainer-accepted microsoft/apm issues (status/accepted) and fan them out through an isolated pool (default 2) of autopilot-issue-delivery-worker sessions. | microsoft/ | 4k | — | ~1.7k | Automated safety check: Pass | MIT | today |
| 69 | A skill your agent uses when changing Archestra tracing, metrics, OpenTelemetry, Tempo, Grafana, Prometheus, LLM/MCP spans, observability labels, or local observability setup. | archestra-ai/ | 4.4k | — | ~1.2k | Automated safety check: Pass | Unknown | today |
| 70 | Checks OmniRoute server health, per-component status, and circuit breakers from the command line, with a live watch dashboard. | diegosouzapw/ | 74k | 1 repo | ~327 | Automated safety check: Pass | MIT | today |
| 71 | Find the top CPU, memory, or disk consumers. Use for capacity reviews, noisy-neighbor hunts, and top-N questions about which jobs, pods, or… | prometheus/ | 118 | — | ~569 | Automated safety check: Pass | Apache-2.0 | 3 days ago |
| 72 | 72.Docs Site A skill your agent uses when editing, building, previewing, or deploying the prometheus-proxy documentation site under website/prometheus-proxy — covers the Zensical config, code-snippet resolution… | pambrose/ | 157 | — | ~250 | Automated safety check: Pass | Apache-2.0 | today |
| 73 | Does a read-only first pass on a possible production incident: gathers deploy, Sentry and health-check signals, proposes a severity and status message, then stops for human approval. | superset-sh/ | 15k | — | ~1k | Automated safety check: Pass | Unknown | today |
| 74 | A skill your agent uses when adding tracing, structured logging, metrics, or monitoring to a FrontMCP server. | agentfront/ | 146 | — | ~4.6k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 75 | Diagnose a live, running lean devnet from its Prometheus and container logs, then report findings with proposed fixes and stop for approval. | geanlabs/ | 177 | — | ~3.8k | Automated safety check: Warn | No licence | today |
| 76 | Overview of Neon, a complete set of cloud backend primitives around Lakebase Postgres: Auth, Object Storage, Functions, and the AI Gateway. | neondatabase/ | 100 | — | ~8.7k | Automated safety check: Notes | Apache-2.0 | yesterday |
| 77 | Checks whether the Telegram daemon is alive, shows recent logs, finds the remote control URL and diagnoses MCP server problems. | k1p1l0/ | 131 | — | ~746 | Automated safety check: Pass | Apache-2.0 | 12 days ago |
| 78 | Monitoring and observability strategy, implementation, and troubleshooting. | ahmedasmar/ | 203 | — | ~3.9k | Automated safety check: Pass | No licence | 5 mo ago |
| 79 | Queue open microsoft/apm issues and fan them out through an isolated pool (default 2) of autopilot-issue-triage-worker sessions. | microsoft/ | 4k | — | ~1.6k | Automated safety check: Pass | MIT | today |
| 80 | Create and manage production-ready Grafana dashboards for comprehensive system observability. | davila7/ | 32k | 14 repos | ~2.1k | Automated safety check: Pass | MIT | today |
| 81 | Set up metrics collection and visualization with Prometheus and Grafana. | BagelHole/ | 1.1k | — | ~2.5k | Automated safety check: Pass | MIT | 4 mo ago |
| 82 | APM - traces, services, dependencies, performance analysis. An agent skill from DataDog/pup. | DataDog/ | 1k | — | ~1.5k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 83 | Installs clidash, a dependency-free web dashboard that turns the JSON resource listings of a CLI such as NanoClaw's ncl into read-only tabs and tables. | nanocoai/ | 31k | 1 repo | ~1.6k | Automated safety check: Pass | MIT | yesterday |
| 84 | Build or refactor a Tool-channel provider (PagerDuty, Opsgenie, future incident/alerting tools) to be endpoint-routed: per-subscriber secrets encrypted on the channel endpoint resource, a stateless… | novuhq/ | 40k | — | ~2.8k | Automated safety check: Pass | Unknown | today |
| 85 | Quantifies elevated error rates with PromQL, compares them to a baseline, and isolates which jobs or instances an error spike is concentrated in. | prometheus/ | 118 | — | ~592 | Automated safety check: Pass | Apache-2.0 | 3 days ago |
| 86 | Complete guide to Prometheus setup, metric collection, scrape configuration, and recording rules. | davila7/ | 32k | 12 repos | ~2.6k | Automated safety check: Pass | MIT | today |
| 87 | In-memory caching in Golang using samber/hot — eviction algorithms (LRU, LFU, TinyLFU, W-TinyLFU, S3FIFO, ARC, TwoQueue, SIEVE, FIFO), TTL, cache loaders, sharding, stale-while-revalidate, missing… | samber/ | 3.4k | — | ~2k | Automated safety check: Pass | MIT | 6 days ago |
| 88 | Queue open microsoft/apm pull requests (community fixes, PRs without an issue, or PRs that close an issue) and fan them out through an isolated pool (default 2) of autopilot-pr-triage-worker sessions. | microsoft/ | 4k | — | ~1.6k | Automated safety check: Pass | MIT | today |
| 89 | 89.Ops Monitor OPS on-demand: This skill should be used when the user asks to "datadog", "APM alerts", or… | Lifecycle-Innovations-Limited/ | 540 | 1 repo | ~1.5k | Automated safety check: Notes | MIT | today |
| 90 | Run any Python LLM agent as an Agent Observability experiment using the public agento11y.experiments package: define a test suite, run an existing agent through typed trials, bind or record… | grafana/ | 127 | — | ~2k | Automated safety check: Pass | Apache-2.0 | today |
| 91 | Finds the metrics and labels behind a Prometheus series explosion using a connected Prometheus MCP server, then proposes relabeling, dropping or recording-rule fixes with measured impact. | prometheus/ | 118 | — | ~547 | Automated safety check: Pass | Apache-2.0 | 3 days ago |
| 92 | LiteLLM-RS Observability Architecture. An agent skill from majiayu000/litellm-rs. | majiayu000/ | 117 | — | ~1.3k | Automated safety check: Pass | MIT | today |
| 93 | Author, modify, or review Netdata collectors across Go, IBM, C, Rust and external plugins. | netdata/ | 81k | — | ~1.9k | Automated safety check: Pass | GPL-3.0 | today |
| 94 | Cut Grafana Cloud Metrics cost by shrinking active-series count with Adaptive Metrics aggregation rules — auto-recommendations from query history, custom exact/regex rules, label-drop config… | grafana/ | 279 | — | ~1.3k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 95 | Runs a phased health audit of a Cisco ACI fabric through MCP tools: node status, links, tenant and policy review, faults and endpoint learning. | automateyournetwork/ | 675 | — | ~2.9k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 96 | A skill your agent uses to triage ONE microsoft/apm pull request already selected by autopilot-pr-triage-scheduler. | microsoft/ | 4k | — | ~1.6k | Automated safety check: Pass | MIT | today |
Explore related skills
Category
More topics in DevOps & Cloud
- Deployment1,264
- CI/CD977
- Containers731
- Observability636
- Container orchestration542
- Infrastructure as code377
- Secrets management335
- Runbooks and postmortems324
- Incident response302
- Cloud networking225
- Backup and disaster recovery183
- Site reliability engineering152
- Cloud architecture118
- MLOps100
- Cloud cost optimization92
- GitOps89
- Linux administration70
- Platform engineering45
- Chaos engineering25