Search
DevOps & Cloud · Prometheus
Skills
Sort:BestMost starsTrending todayTrending this weekTrending this monthNewestRecently updatedName
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 97 | Validate, lint, audit, or fix PromQL queries and alerting rules; detects anti-patterns. | akin-ozer/ | 320 | — | ~4k | Automated safety check: Pass | Apache-2.0 | 2 mo ago |
| 98 | Canvas/A2UI inline network visualizations — topology maps, health dashboards, alert cards, change timelines, config diffs, path traces, and health scorecards rendered directly in the OpenClaw chat… | automateyournetwork/ | 676 | — | ~941 | Automated safety check: Pass | Apache-2.0 | yesterday |
| 99 | Reusable investigation patterns for AWS CloudWatch: Logs Insights query templates, alarm-to-deployment correlation, blast-radius narrowing decision tree, and PromQL-style metric query patterns for… | github/ | 40k | — | ~2.6k | Automated safety check: Pass | MIT | 2 days ago |
| 100 | Collect comprehensive infrastructure performance metrics across compute, storage, network, containers, load balancers, and databases. | jeremylongshore/ | 2.8k | — | ~1.2k | Automated safety check: Pass | MIT | yesterday |
| 101 | Monitor use when deploying monitoring stacks including Prometheus, Grafana, and Datadog. | jeremylongshore/ | 2.8k | — | ~1.2k | Automated safety check: Pass | MIT | yesterday |
| 102 | Prometheus and Grafana Cloud Metrics overview including PromQL query language, Metrics Drilldown, alerting, recording rules, and integration patterns. | grafana/ | 282 | — | ~1.2k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 103 | Diagnostic guide for active Prometheus cardinality problems — slow queries, OOMing Prometheus, high Grafana Cloud Active Series or DPM bills, "too many samples" ingest errors, series churn, or rapid… | grafana/ | 282 | — | ~4.6k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 104 | Sending telemetry data to Grafana Cloud — metrics via Prometheus remote write or OTLP, logs via Loki push or Alloy, traces via OTLP to Tempo, profiles via Pyroscope. | grafana/ | 282 | — | ~1.4k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 105 | A skill your agent uses to deploy and operate a CollectX (clx) based DOCA telemetry collector on a host or BlueField — wiring providers / counters into the collector, running the collection daemon… | NVIDIA/ | 3.6k | — | ~2.8k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 106 | 106.Oh My Opencode Multi-agent orchestration plugin for OpenCode. An agent skill from LeoYeAI/openclaw-master-skills. | LeoYeAI/ | 2.2k | — | ~5.1k | Automated safety check: Notes | MIT | 2 mo ago |
| 107 | Design, audit, and troubleshoot production monitoring and observability using user-impact checks, layered telemetry, USE/RED, SLI/SLO/SLA, error budgets, cardinality controls, actionable alerting… | AnastasiyaW/ | 154 | — | ~4.1k | Automated safety check: Pass | MIT | yesterday |
| 108 | Builds, configures, debugs, and optimizes AWS observability - operator-symptom questions and detecting Omni vs classic CloudWatch. | aws/ | 2.8k | — | ~7.2k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 109 | Hunt down CoreWeave GPU cost leaks — idle reserved capacity, wrong-GPU-type right-sizing waste, allocated-but-idle instances, and on-demand spend that should be committed — then produce a… | jeremylongshore/ | 2.8k | — | ~3.5k | Automated safety check: Pass | MIT | yesterday |
| 110 | Probe a target for accidentally-public admin / debug / introspection endpoints — Spring Boot Actuator, Apache server-status, Prometheus metrics, GraphQL playground, Swagger UI, phpMyAdmin… | jeremylongshore/ | 2.8k | — | ~2k | Automated safety check: Pass | MIT | yesterday |
| 111 | A skill your agent uses when you need to implement or improve Java metrics observability with Micrometer — including meter design, naming/tag conventions, cardinality control… | jabrena/ | 447 | — | ~868 | Automated safety check: Pass | Apache-2.0 | 3 days ago |
| 112 | [OMX] Clean-room interview-driven planner: Metis clarifies, Momus challenges, Oracle synthesizes, then hands off to $ultragoal/$team. | yangyuan-zhen/ | 316 | — | ~4.6k | Automated safety check: Pass | AGPL-3.0 | 20 days ago |
| 113 | Observability and SRE expert. An agent skill from majiayu000/spellbook. | majiayu000/ | 287 | — | ~3.3k | Automated safety check: Pass | MIT | 2 days ago |
| 114 | 114.Promql Generator Generate/create/write PromQL queries, metric expressions, alerting rules, recording rules, Prometheus dashboards. | akin-ozer/ | 320 | — | ~9.2k | Automated safety check: Pass | Apache-2.0 | 2 mo ago |
| 115 | Observability patterns for Python applications. An agent skill from aiskillstore/marketplace. | aiskillstore/ | 433 | 1 repo | ~1.3k | Automated safety check: Pass | No licence | yesterday |
| 116 | This skill provides AWS cost optimization, monitoring, and operational best practices with integrated MCP servers for billing analysis, cost estimation, observability, and security assessment. | Microck/ | 404 | 1 repo | ~2.5k | Automated safety check: Pass | Unknown | 1 mo ago |
| 117 | Expert evaluator for Prometheus label strategy on Grafana Cloud. | grafana/ | 282 | — | ~6.4k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 118 | 118.Datapages Server Configure the Datapages server entry point: NewServer type arguments, the message broker, server options, static assets, TLS and Prometheus metrics. | romshark/ | 113 | — | ~2.2k | Automated safety check: Pass | MIT | 3 days ago |
| 119 | Monitoring and observability patterns for Prometheus metrics, Grafana dashboards, Langfuse v4 LLM tracing (astype, scorecurrentspan, shouldexportspan, LangfuseMedia), and drift detection. | yonatangross/ | 292 | — | ~2.2k | Automated safety check: Pass | MIT | yesterday |
| 120 | 120.Golang Benchmark Golang benchmarking, profiling, and performance measurement. | aiskillstore/ | 433 | 1 repo | ~3.4k | Automated safety check: Pass | MIT | yesterday |
| 121 | 121.Gorm Expert GORM v2 最佳实践与性能优化。适用于:代码审查、慢查询优化、N+1、连接池、 事务管理、分库分表、Prometheus/OTel监控、Session安全、Clause/Upsert、 缓存集成、BaseModel脚手架、SQL→struct生成、多租户隔离。 | LeoYeAI/ | 2.2k | — | ~3.4k | Automated safety check: Pass | MIT | 2 mo ago |
| 122 | 122.Alerting Oncall Set up alerting rules, configure on-call rotations, and manage incident response workflows. | BagelHole/ | 1.2k | — | ~3k | Automated safety check: Pass | MIT | 4 mo ago |
| 123 | 123.Observability OpenTelemetry, distributed tracing, structured logging, metrics (Prometheus, Grafana, Datadog). | TheBeardedBearSAS/ | 107 | — | ~547 | Automated safety check: Pass | MIT | 26 days ago |
| 124 | Monitoring and observability with OpenTelemetry, Prometheus, Grafana dashboards, and structured logging | rohitg00/ | 2.7k | — | ~1.4k | Automated safety check: Pass | Apache-2.0 | 5 mo ago |
| 125 | Observability: structured logs, metrics (RED/USE), tracing, SLO/SLI. | softspark/ | 179 | — | ~2.2k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 126 | Set up Apollo.io monitoring and observability. An agent skill from jeremylongshore/tons-of-skills-marketplace. | jeremylongshore/ | 2.8k | — | ~2.3k | Automated safety check: Pass | MIT | yesterday |
| 127 | Monitor ClickHouse with Prometheus metrics, Grafana dashboards, system table queries, and alerting for query performance, merge health, and resource usage. | jeremylongshore/ | 2.8k | — | ~1.4k | Automated safety check: Pass | MIT | yesterday |
| 128 | Set up Customer.io monitoring and observability. An agent skill from jeremylongshore/tons-of-skills-marketplace. | jeremylongshore/ | 2.8k | — | ~2.8k | Automated safety check: Pass | MIT | yesterday |
| 129 | Set up comprehensive observability for Deepgram integrations. | jeremylongshore/ | 2.8k | — | ~3k | Automated safety check: Pass | MIT | yesterday |
| 130 | Execute Deepgram production deployment checklist. An agent skill from jeremylongshore/tons-of-skills-marketplace. | jeremylongshore/ | 2.8k | — | ~2.3k | Automated safety check: Pass | MIT | yesterday |
| 131 | Implement observability for Evernote integrations. An agent skill from jeremylongshore/tons-of-skills-marketplace. | jeremylongshore/ | 2.8k | — | ~1.2k | Automated safety check: Pass | MIT | yesterday |
| 132 | Set up monitoring, metrics, and alerting for Figma API integrations. | jeremylongshore/ | 2.8k | — | ~2k | Automated safety check: Pass | MIT | yesterday |
| 133 | Implement comprehensive observability for Gamma integrations. | jeremylongshore/ | 2.8k | — | ~2.1k | Automated safety check: Pass | MIT | yesterday |
| 134 | Set up observability for Groq integrations: latency histograms, token throughput, rate limit gauges, cost tracking, and Prometheus alerts. | jeremylongshore/ | 2.8k | — | ~1.6k | Automated safety check: Pass | MIT | yesterday |
| 135 | A skill your agent uses when you need production monitoring for an Intercom integration — instrumenting API calls with metrics and traces, standing up dashboards, or wiring alerts for error rate… | jeremylongshore/ | 2.8k | — | ~1.7k | Automated safety check: Pass | MIT | yesterday |
| 136 | Set up observability for Klaviyo integrations with metrics, traces, and alerts. | jeremylongshore/ | 2.8k | — | ~1.5k | Automated safety check: Pass | MIT | yesterday |
| 137 | Set up comprehensive observability for Langfuse with metrics, dashboards, and alerts. | jeremylongshore/ | 2.8k | — | ~2.2k | Automated safety check: Pass | MIT | yesterday |
| 138 | Implement comprehensive observability for MaintainX integrations. | jeremylongshore/ | 2.8k | — | ~2.1k | Automated safety check: Pass | MIT | yesterday |
| 139 | 139.Monitoring APIs Build real-time API monitoring dashboards with metrics, alerts, and health checks. | jeremylongshore/ | 2.8k | — | ~1.4k | Automated safety check: Pass | MIT | yesterday |
| 140 | Observe a PostHog integration through application-side delivery metrics, ingestion warnings, billing volume, destination logs, and status evidence. | jeremylongshore/ | 2.8k | — | ~2.4k | Automated safety check: Pass | MIT | yesterday |
| 141 | Set up observability for Shopify app integrations with query cost tracking, rate limit monitoring, webhook delivery metrics, and structured logging. | jeremylongshore/ | 2.8k | — | ~1k | Automated safety check: Pass | MIT | yesterday |
| 142 | Cost guardrail for AWS DevOps Agent that covers ALL AWS services and native agent tools. | aws/ | 103 | — | ~4.5k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 143 | 143.Grafana Helper Use Grafana's GCX CLI for dashboards, datasources, Prometheus metrics, Loki logs, Tempo traces, alert rules, and Grafana resource operations. | shepherdjerred/ | 112 | — | ~1.1k | Automated safety check: Pass | GPL-3.0 | yesterday |
| 144 | 144.Ops Report Generate a 24-hour operational health report for a PostHog service by querying Grafana dashboards and Prometheus metrics. | haacked/ | 134 | — | ~8k | Automated safety check: Notes | No licence | yesterday |