Topic · DevOps & Cloud
Best monitoring and alerting skills, page 7
Monitoring and alerting skills, ranked
Ranked by score. Sort bymost stars,trending,newest,recently updated
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 289 | Implement observability for Evernote integrations. An agent skill from jeremylongshore/tons-of-skills-marketplace. | jeremylongshore/ | 2.8k | — | ~1.2k | Automated safety check: Pass | MIT | today |
| 290 | Set up monitoring, metrics, and alerting for Figma API integrations. | jeremylongshore/ | 2.8k | — | ~2k | Automated safety check: Pass | MIT | today |
| 291 | Implement comprehensive observability for Gamma integrations. | jeremylongshore/ | 2.8k | — | ~2.1k | Automated safety check: Pass | MIT | today |
| 292 | Set up observability for Groq integrations: latency histograms, token throughput, rate limit gauges, cost tracking, and Prometheus alerts. | jeremylongshore/ | 2.8k | — | ~1.6k | Automated safety check: Pass | MIT | today |
| 293 | A skill your agent uses when you need production monitoring for an Intercom integration — instrumenting API calls with metrics and traces, standing up dashboards, or wiring alerts for error rate… | jeremylongshore/ | 2.8k | — | ~1.7k | Automated safety check: Pass | MIT | today |
| 294 | Set up observability for Klaviyo integrations with metrics, traces, and alerts. | jeremylongshore/ | 2.8k | — | ~1.5k | Automated safety check: Pass | MIT | today |
| 295 | Set up comprehensive observability for Langfuse with metrics, dashboards, and alerts. | jeremylongshore/ | 2.8k | — | ~2.2k | Automated safety check: Pass | MIT | today |
| 296 | Implement comprehensive observability for MaintainX integrations. | jeremylongshore/ | 2.8k | — | ~2.1k | Automated safety check: Pass | MIT | today |
| 297 | 297.Monitoring APIs Build real-time API monitoring dashboards with metrics, alerts, and health checks. | jeremylongshore/ | 2.8k | — | ~1.4k | Automated safety check: Pass | MIT | today |
| 298 | Observe a PostHog integration through application-side delivery metrics, ingestion warnings, billing volume, destination logs, and status evidence. | jeremylongshore/ | 2.8k | — | ~2.4k | Automated safety check: Pass | MIT | today |
| 299 | Migrate to Sentry from other error tracking tools like Rollbar, Bugsnag, or New Relic. | jeremylongshore/ | 2.8k | — | ~2.6k | Automated safety check: Notes | MIT | today |
| 300 | Configure Sentry across development, staging, and production environments with separate DSNs, environment-specific sample rates, per-environment alert rules, and dashboard filtering. | jeremylongshore/ | 2.8k | — | ~4.1k | Automated safety check: Notes | MIT | today |
| 301 | Integrate Sentry with your observability stack — logging, metrics, APM, and dashboards. | jeremylongshore/ | 2.8k | — | ~3.7k | Automated safety check: Pass | MIT | today |
| 302 | Set up performance monitoring and distributed tracing with Sentry. | jeremylongshore/ | 2.8k | — | ~4.1k | Automated safety check: Pass | MIT | today |
| 303 | Set up observability for Shopify app integrations with query cost tracking, rate limit monitoring, webhook delivery metrics, and structured logging. | jeremylongshore/ | 2.8k | — | ~1k | Automated safety check: Pass | MIT | today |
| 304 | 304.Uptime Kuma Interact with Uptime Kuma monitoring server. An agent skill from sundial-org/awesome-openclaw-skills. | sundial-org/ | 663 | — | ~575 | Automated safety check: Pass | No licence | 7 mo ago |
| 305 | 305.Grafana Helper Use Grafana's GCX CLI for dashboards, datasources, Prometheus metrics, Loki logs, Tempo traces, alert rules, and Grafana resource operations. | shepherdjerred/ | 112 | — | ~1.1k | Automated safety check: Pass | GPL-3.0 | today |
| 306 | 306.Monitoring A skill your agent uses when setting up uptime and health monitoring, alerts, or on-call basics for a service already in production, so you learn it is down before customers do — health and… | ericrisco/ | 167 | — | ~3.1k | Automated safety check: Pass | MIT | today |
| 307 | Generate a live Single Step Instrumentation (SSI) onboarding confirmation report — verifies APM instrumentation is working end-to-end with deep links into the Datadog UI. | datadog-labs/ | 177 | — | ~1.1k | Automated safety check: Pass | MIT | today |
| 308 | Generate a live Single Step Instrumentation (SSI) onboarding confirmation report for Linux hosts — verifies APM instrumentation is working end-to-end with deep links into the Datadog UI. | datadog-labs/ | 177 | — | ~1.2k | Automated safety check: Notes | MIT | today |
| 309 | 309.AI Operations Configure Harness AI-powered operations (AIDA) via MCP. An agent skill from majiayu000/claude-skill-registry. | majiayu000/ | 666 | 1 repo | ~1.4k | Automated safety check: Pass | Apache-2.0 | today |
| 310 | Query and analyze Claude Code observability data (metrics, logs, traces). | majiayu000/ | 666 | 1 repo | ~1.6k | Automated safety check: Pass | MIT | today |
| 311 | 311.Ops Report Generate a 24-hour operational health report for a PostHog service by querying Grafana dashboards and Prometheus metrics. | haacked/ | 134 | — | ~8k | Automated safety check: Notes | No licence | today |
| 312 | 312.Telemetry Operate the observability stack that deploys as one unit: Prometheus scrape configuration, recording and alerting rules, relabeling, retention, and high availability; OpenTelemetry Collector… | magnus919/ | 113 | — | ~3.9k | Automated safety check: Pass | MIT | yesterday |
| 313 | Observability and monitoring validation patterns for dashboards, alerting, log aggregation, APM traces, and SLA/SLO verification. | proffesor-for-testing/ | 494 | — | ~8.3k | Automated safety check: Pass | MIT | 3 days ago |
| 314 | Observability patterns for logging, monitoring, alerting, and distributed tracing. | TheSoftwareHouse/ | 284 | — | ~2k | Automated safety check: Pass | MIT | 2 days ago |
| 315 | Set up observability for Claude API integrations with metrics, logging, and alerting for latency, cost, errors, and token usage. | jeremylongshore/ | 2.8k | — | ~1.8k | Automated safety check: Pass | MIT | today |
| 316 | Monitor Clay enrichment pipeline health, credit consumption, and data quality metrics. | jeremylongshore/ | 2.8k | — | ~2.2k | Automated safety check: Pass | MIT | today |
| 317 | Set up GPU monitoring and observability for CoreWeave workloads. | jeremylongshore/ | 2.8k | — | ~1.4k | Automated safety check: Pass | MIT | today |
| 318 | Create grafana dashboard creator operations. An agent skill from jeremylongshore/tons-of-skills-marketplace. | jeremylongshore/ | 2.8k | — | ~582 | Automated safety check: Pass | MIT | today |
| 319 | Set up comprehensive observability for Lokalise integrations with metrics, traces, and alerts. | jeremylongshore/ | 2.8k | — | ~2.9k | Automated safety check: Pass | MIT | today |
| 320 | Generate prometheus config generator operations. An agent skill from jeremylongshore/tons-of-skills-marketplace. | jeremylongshore/ | 2.8k | — | ~585 | Automated safety check: Pass | MIT | today |
| 321 | 321.Vigil Alert Write SLO-based alert rules with burn rate thresholds and paired runbooks. | jeremylongshore/ | 2.8k | — | ~2.7k | Automated safety check: Notes | MIT | today |
| 322 | 322.Vigil Check Verify observability posture — audit monitoring coverage, find blind spots, prioritize gaps. | jeremylongshore/ | 2.8k | — | ~1k | Automated safety check: Notes | MIT | today |
| 323 | 323.Vigil Recon Observability reconnaissance — inventory what monitoring exists, map coverage, highlight blind spots. | jeremylongshore/ | 2.8k | — | ~1.1k | Automated safety check: Notes | MIT | today |
| 324 | Audit a JavaScript/TypeScript repo's npm, yarn, or pnpm configuration for supply-chain hardening: tool version, lifecycle scripts, unsafe dependency protocols, and minimum release age ≥3 days. | grafana/ | 279 | — | ~1.3k | Automated safety check: Warn | Apache-2.0 | yesterday |
| 325 | Grafana observability platform — dashboards, Prometheus PromQL, Loki LogQL, alerting, incidents, OnCall schedules, annotations, datasource queries, panel rendering (75+ tools). | automateyournetwork/ | 675 | — | ~2.7k | Automated safety check: Notes | Apache-2.0 | 2 days ago |
| 326 | Prometheus monitoring — PromQL instant/range queries, metric discovery, metadata, scrape target health, system health checks (6 tools). | automateyournetwork/ | 675 | — | ~2k | Automated safety check: Notes | Apache-2.0 | 2 days ago |
| 327 | A skill your agent uses when analyzing disruption from a DisruptionRegression alert, a Grafana disruption dashboard URL, or Prow CI job runs by examining interval data, audit logs, pod logs, and CPU… | openshift-eng/ | 120 | — | ~2.8k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 328 | Guides Qdrant monitoring and observability setup. An agent skill from qdrant/skills. | qdrant/ | 253 | — | ~368 | Automated safety check: Pass | Apache-2.0 | today |
| 329 | 329.Promql CLI CLI for querying Prometheus and PromQL-compatible engines (Thanos, Cortex, VictoriaMetrics, Grafana Mimir, Grafana Tempo...) — instant queries, range queries, metric discovery (metrics/labels/meta… | samber/ | 228 | — | ~1.9k | Automated safety check: Pass | MIT | 6 days ago |
| 330 | 330.Enable Ssi Enable Single Step Instrumentation (SSI) on Kubernetes — automatically instruments applications for APM without code changes. | datadog-labs/ | 177 | — | ~3.2k | Automated safety check: Warn | MIT | today |
| 331 | Reduce Sentry alert fatigue by surgically tuning issue grouping, fingerprint rules, severity mapping, sample rates, before-send filters, sourcemap pipelines, and release-health gates. | LeoYeAI/ | 2.2k | — | ~7.3k | Automated safety check: Pass | MIT | 2 mo ago |
| 332 | Find the Django ORM code and request path responsible for slow SQL by using APM traces, slow query logs, Django Debug Toolbar, query logging, and local reproduction. | hashgraph-online/ | 1.2k | — | ~644 | Automated safety check: Pass | MIT | today |
| 333 | 333.Magic Mouth Magic Mouth is trigger → message. An agent skill from Hmbown/Wizards-of-the-Ghosts. | Hmbown/ | 109 | — | ~836 | Automated safety check: Pass | CC0-1.0 | 6 mo ago |
| 334 | 334.Cloud Monitoring Monitor cloud infrastructure and applications using metrics, logs, and traces to provide real-time observability into performance, health, and reliability. | seb1n/ | 206 | — | ~2.8k | Automated safety check: Pass | MIT | 1 mo ago |
| 335 | Stand up a complete, ready-to-run computer-vision analytics stack on Intel hardware with one Docker Compose command — point it at your video sources and an OpenVINO/ONNX model to get live annotated… | open-edge-platform/ | 140 | — | ~4.4k | Automated safety check: Notes | Apache-2.0 | today |
| 336 | Reduces JavaScript dependency footprint with pnpm while preserving lockfile, workspace layout, and dependency range style. | grafana/ | 279 | — | ~3.6k | Automated safety check: Warn | Apache-2.0 | yesterday |
Explore related skills
More topics in DevOps & Cloud
- Deployment1,264
- CI/CD977
- Containers731
- Observability636
- Container orchestration542
- Infrastructure as code377
- Secrets management335
- Runbooks and postmortems324
- Incident response302
- Cloud networking225
- Backup and disaster recovery183
- Site reliability engineering152
- Cloud architecture118
- MLOps100
- Cloud cost optimization92
- GitOps89
- Linux administration70
- Platform engineering45
- Chaos engineering25