Topic · DevOps & Cloud
Best monitoring and alerting skills, page 6
Monitoring and alerting skills, ranked
Ranked by score. Sort bymost stars,trending,newest,recently updated
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 241 | Query live Pyroscope profiles with profilecli, analyze them with pprof, and correlate hot functions with checked-out source code. | grafana/ | 279 | — | ~1.8k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 242 | Prometheus and Grafana Cloud Metrics overview including PromQL query language, Metrics Drilldown, alerting, recording rules, and integration patterns. | grafana/ | 279 | — | ~1.2k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 243 | Diagnostic guide for active Prometheus cardinality problems — slow queries, OOMing Prometheus, high Grafana Cloud Active Series or DPM bills, "too many samples" ingest errors, series churn, or rapid… | grafana/ | 279 | — | ~4.6k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 244 | Sending telemetry data to Grafana Cloud — metrics via Prometheus remote write or OTLP, logs via Loki push or Alloy, traces via OTLP to Tempo, profiles via Pyroscope. | grafana/ | 279 | — | ~1.4k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 245 | Author Grafana Cloud Synthetic Monitoring checks, with deep coverage of k6 scripted and browser checks: SM's single-VU/single-iteration execution model, assertions that actually fail probesuccess… | grafana/ | 279 | — | ~5.7k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 246 | 246.QA Dashboard Build and visualize QA dashboards and reports with Allure Report, Grafana, and ReportPortal. | petrkindlmann/ | 165 | — | ~4.7k | Automated safety check: Pass | MIT | 3 mo ago |
| 247 | 247.Oss Alternatives Find actively maintained open source alternatives to any paid SaaS tool or commercial API. | tinyfish-io/ | 2.2k | — | ~2.2k | Automated safety check: Pass | MIT | 6 days ago |
| 248 | A skill your agent uses to deploy and operate a CollectX (clx) based DOCA telemetry collector on a host or BlueField — wiring providers / counters into the collector, running the collection daemon… | NVIDIA/ | 3.5k | — | ~2.8k | Automated safety check: Pass | Apache-2.0 | today |
| 249 | 249.Oh My Opencode Multi-agent orchestration plugin for OpenCode. An agent skill from LeoYeAI/openclaw-master-skills. | LeoYeAI/ | 2.2k | — | ~5.1k | Automated safety check: Notes | MIT | 2 mo ago |
| 250 | Design, audit, and troubleshoot production monitoring and observability using user-impact checks, layered telemetry, USE/RED, SLI/SLO/SLA, error budgets, cardinality controls, actionable alerting… | AnastasiyaW/ | 154 | — | ~4.1k | Automated safety check: Pass | MIT | today |
| 251 | 251.Skill Selector Meta-skill for picking project skills via APM. An agent skill from mizchi/skills. | mizchi/ | 356 | — | ~2.8k | Automated safety check: Pass | No licence | 6 days ago |
| 252 | Hunt down CoreWeave GPU cost leaks — idle reserved capacity, wrong-GPU-type right-sizing waste, allocated-but-idle instances, and on-demand spend that should be committed — then produce a… | jeremylongshore/ | 2.8k | — | ~3.5k | Automated safety check: Pass | MIT | today |
| 253 | Probe a target for accidentally-public admin / debug / introspection endpoints — Spring Boot Actuator, Apache server-status, Prometheus metrics, GraphQL playground, Swagger UI, phpMyAdmin… | jeremylongshore/ | 2.8k | — | ~2k | Automated safety check: Pass | MIT | today |
| 254 | A skill your agent uses when you need to implement or improve Java logging and observability — including selecting SLF4J with Logback/Log4j2, applying proper log levels (ERROR, WARN, INFO, DEBUG… | jabrena/ | 446 | — | ~804 | Automated safety check: Pass | Apache-2.0 | today |
| 255 | A skill your agent uses when you need to implement or improve Java metrics observability with Micrometer — including meter design, naming/tag conventions, cardinality control… | jabrena/ | 446 | — | ~868 | Automated safety check: Pass | Apache-2.0 | today |
| 256 | [OMX] Clean-room interview-driven planner: Metis clarifies, Momus challenges, Oracle synthesizes, then hands off to $ultragoal/$team. | yangyuan-zhen/ | 315 | — | ~4.6k | Automated safety check: Pass | AGPL-3.0 | 17 days ago |
| 257 | Observability and SRE expert. An agent skill from majiayu000/spellbook. | majiayu000/ | 286 | — | ~3.3k | Automated safety check: Pass | MIT | today |
| 258 | 258.Promql Generator Generate/create/write PromQL queries, metric expressions, alerting rules, recording rules, Prometheus dashboards. | akin-ozer/ | 319 | — | ~9.2k | Automated safety check: Pass | Apache-2.0 | 2 mo ago |
| 259 | 259.Uptime Monitor Monitor website uptime - check availability, response times, and status | gooseworks-ai/ | 1.2k | 1 repo | ~1.3k | Automated safety check: Pass | MIT | today |
| 260 | 260.Datadog Query and analyze Datadog logs, metrics, APM traces, and monitors using the Datadog API. | OpenHands/ | 158 | — | ~716 | Automated safety check: Pass | MIT | today |
| 261 | Observability patterns for Python applications. An agent skill from aiskillstore/marketplace. | aiskillstore/ | 430 | 1 repo | ~1.3k | Automated safety check: Pass | No licence | today |
| 262 | An on-call veteran SRE interviewer focused on monitoring and alerting. | PrepLabsAI/ | 112 | — | ~3.9k | Automated safety check: Pass | MIT | yesterday |
| 263 | A skill your agent uses when the user asks to "set up monitoring", "configure observability", "onboard new service", "create saved view", "set up notifications", "configure webhook", "set up Slack… | coralogix/ | 121 | — | ~2.3k | Automated safety check: Pass | Apache-2.0 | today |
| 264 | Query Coralogix's Service Catalog (APM v2 entities) with the cx service-catalog CLI — discover entity types, list known entities, check their schema, and pull aggregated or timeseries data for… | coralogix/ | 121 | — | ~2.1k | Automated safety check: Pass | Apache-2.0 | today |
| 265 | Create and manage APM service remapping rules — rewrite service names at ingestion time to collapse noisy inferred entities, clean up auto-generated names, handle org renames, or normalize naming… | datadog-labs/ | 177 | — | ~4.7k | Automated safety check: Pass | MIT | today |
| 266 | This skill provides AWS cost optimization, monitoring, and operational best practices with integrated MCP servers for billing analysis, cost estimation, observability, and security assessment. | Microck/ | 403 | 1 repo | ~2.5k | Automated safety check: Pass | Unknown | 1 mo ago |
| 267 | 267.Datapages Server Configure the Datapages server entry point: NewServer type arguments, the message broker, server options, static assets, TLS and Prometheus metrics. | romshark/ | 113 | — | ~2.2k | Automated safety check: Pass | MIT | today |
| 268 | Expert evaluator for Prometheus label strategy on Grafana Cloud. | grafana/ | 279 | — | ~6.4k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 269 | Set up ongoing competitor monitoring — captures per-competitor baselines across content, pricing, ads, social, SEO, and SERP features, saves them via competitor-tracker.py, configures per-dimension… | indranilbanerjee/ | 855 | 1 repo | ~3.2k | Automated safety check: Pass | MIT | 4 days ago |
| 270 | 270.QA Metrics Define, track, and act on QA metrics: test coverage percentage, flakiness rate, defect escape rate, MTTR, test execution time trends, automation ROI, quality gates, and SLAs for test suites. | petrkindlmann/ | 165 | — | ~5.3k | Automated safety check: Pass | MIT | 3 mo ago |
| 271 | Scheduled probes that run CONTINUOUSLY after release. An agent skill from petrkindlmann/qa-skills. | petrkindlmann/ | 165 | — | ~5.8k | Automated safety check: Pass | MIT | 3 mo ago |
| 272 | Monitoring and observability patterns for Prometheus metrics, Grafana dashboards, Langfuse v4 LLM tracing (astype, scorecurrentspan, shouldexportspan, LangfuseMedia), and drift detection. | yonatangross/ | 289 | — | ~2.2k | Automated safety check: Pass | MIT | today |
| 273 | 273.Golang Benchmark Golang benchmarking, profiling, and performance measurement. | aiskillstore/ | 430 | 1 repo | ~3.4k | Automated safety check: Pass | MIT | today |
| 274 | 274.Gorm Expert GORM v2 最佳实践与性能优化。适用于:代码审查、慢查询优化、N+1、连接池、 事务管理、分库分表、Prometheus/OTel监控、Session安全、Clause/Upsert、 缓存集成、BaseModel脚手架、SQL→struct生成、多租户隔离。 | LeoYeAI/ | 2.2k | — | ~3.4k | Automated safety check: Pass | MIT | 2 mo ago |
| 275 | 275.Alerting Oncall Set up alerting rules, configure on-call rotations, and manage incident response workflows. | BagelHole/ | 1.1k | — | ~3k | Automated safety check: Pass | MIT | 4 mo ago |
| 276 | 276.Loki Logging Configure Grafana Loki for log aggregation and analysis. An agent skill from BagelHole/DevOps-Security-Agent-Skills. | BagelHole/ | 1.1k | — | ~2.4k | Automated safety check: Pass | MIT | 4 mo ago |
| 277 | 277.Observability OpenTelemetry, distributed tracing, structured logging, metrics (Prometheus, Grafana, Datadog). | TheBeardedBearSAS/ | 107 | — | ~547 | Automated safety check: Pass | MIT | 23 days ago |
| 278 | Design observability strategies: SLI/SLO frameworks, alerting, and dashboards. | borghei/ | 881 | — | ~1.8k | Automated safety check: Pass | MIT | today |
| 279 | Monitoring and observability with OpenTelemetry, Prometheus, Grafana dashboards, and structured logging | rohitg00/ | 2.7k | — | ~1.4k | Automated safety check: Pass | Apache-2.0 | 4 mo ago |
| 280 | A skill your agent uses when designing a GEO Signal Monitor, AI answer monitoring system, citation tracking plan, brand-fact correction loop, GEO monthly report, alert rules, dashboard fields, or… | yaojingang/ | 868 | — | ~799 | Automated safety check: Pass | MIT | 7 days ago |
| 281 | Observability: structured logs, metrics (RED/USE), tracing, SLO/SLI. | softspark/ | 179 | — | ~2.2k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 282 | 282.Enable Dsm Enable Data Streams Monitoring (DSM) on services already instrumented with APM, for end-to-end latency, throughput, and consumer lag across Kafka, RabbitMQ, SQS, SNS, Kinesis, Pub/Sub, IBM MQ, Azure… | datadog-labs/ | 177 | — | ~5.7k | Automated safety check: Notes | MIT | today |
| 283 | Set up Apollo.io monitoring and observability. An agent skill from jeremylongshore/tons-of-skills-marketplace. | jeremylongshore/ | 2.8k | — | ~2.3k | Automated safety check: Pass | MIT | today |
| 284 | Monitor ClickHouse with Prometheus metrics, Grafana dashboards, system table queries, and alerting for query performance, merge health, and resource usage. | jeremylongshore/ | 2.8k | — | ~1.4k | Automated safety check: Pass | MIT | today |
| 285 | Set up Customer.io monitoring and observability. An agent skill from jeremylongshore/tons-of-skills-marketplace. | jeremylongshore/ | 2.8k | — | ~2.8k | Automated safety check: Pass | MIT | today |
| 286 | Set up comprehensive observability for Deepgram integrations. | jeremylongshore/ | 2.8k | — | ~3k | Automated safety check: Pass | MIT | today |
| 287 | Execute Deepgram production deployment checklist. An agent skill from jeremylongshore/tons-of-skills-marketplace. | jeremylongshore/ | 2.8k | — | ~2.3k | Automated safety check: Pass | MIT | today |
| 288 | Implement monitoring, logging, and tracing for Documenso integrations. | jeremylongshore/ | 2.8k | — | ~2.2k | Automated safety check: Pass | MIT | today |
Explore related skills
Category
More topics in DevOps & Cloud
- Deployment1,264
- CI/CD977
- Containers731
- Observability636
- Container orchestration542
- Infrastructure as code377
- Secrets management335
- Runbooks and postmortems324
- Incident response302
- Cloud networking225
- Backup and disaster recovery183
- Site reliability engineering152
- Cloud architecture118
- MLOps100
- Cloud cost optimization92
- GitOps89
- Linux administration70
- Platform engineering45
- Chaos engineering25