Search
Prometheus · For devops and sre engineers
Skills
Sort:BestMost starsTrending todayTrending this weekTrending this monthNewestRecently updatedName
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 97 | A skill your agent uses to deploy and operate a CollectX (clx) based DOCA telemetry collector on a host or BlueField — wiring providers / counters into the collector, running the collection daemon… | NVIDIA/ | 3.6k | — | ~2.8k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 98 | Design, audit, and troubleshoot production monitoring and observability using user-impact checks, layered telemetry, USE/RED, SLI/SLO/SLA, error budgets, cardinality controls, actionable alerting… | AnastasiyaW/ | 154 | — | ~4.1k | Automated safety check: Pass | MIT | yesterday |
| 99 | Builds, configures, debugs, and optimizes AWS observability - operator-symptom questions and detecting Omni vs classic CloudWatch. | aws/ | 2.8k | — | ~7.2k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 100 | Hunt down CoreWeave GPU cost leaks — idle reserved capacity, wrong-GPU-type right-sizing waste, allocated-but-idle instances, and on-demand spend that should be committed — then produce a… | jeremylongshore/ | 2.8k | — | ~3.5k | Automated safety check: Pass | MIT | yesterday |
| 101 | A skill your agent uses when you need to implement or improve Java metrics observability with Micrometer — including meter design, naming/tag conventions, cardinality control… | jabrena/ | 447 | — | ~868 | Automated safety check: Pass | Apache-2.0 | 3 days ago |
| 102 | [OMX] Clean-room interview-driven planner: Metis clarifies, Momus challenges, Oracle synthesizes, then hands off to $ultragoal/$team. | yangyuan-zhen/ | 316 | — | ~4.6k | Automated safety check: Pass | AGPL-3.0 | 20 days ago |
| 103 | Observability and SRE expert. An agent skill from majiayu000/spellbook. | majiayu000/ | 287 | — | ~3.3k | Automated safety check: Pass | MIT | 2 days ago |
| 104 | 104.Promql Generator Generate/create/write PromQL queries, metric expressions, alerting rules, recording rules, Prometheus dashboards. | akin-ozer/ | 320 | — | ~9.2k | Automated safety check: Pass | Apache-2.0 | 2 mo ago |
| 105 | Observability patterns for Python applications. An agent skill from aiskillstore/marketplace. | aiskillstore/ | 433 | 1 repo | ~1.3k | Automated safety check: Pass | No licence | yesterday |
| 106 | This skill provides AWS cost optimization, monitoring, and operational best practices with integrated MCP servers for billing analysis, cost estimation, observability, and security assessment. | Microck/ | 404 | 1 repo | ~2.5k | Automated safety check: Pass | Unknown | 1 mo ago |
| 107 | Expert evaluator for Prometheus label strategy on Grafana Cloud. | grafana/ | 282 | — | ~6.4k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 108 | 108.Datapages Server Configure the Datapages server entry point: NewServer type arguments, the message broker, server options, static assets, TLS and Prometheus metrics. | romshark/ | 113 | — | ~2.2k | Automated safety check: Pass | MIT | 3 days ago |
| 109 | Monitoring and observability patterns for Prometheus metrics, Grafana dashboards, Langfuse v4 LLM tracing (astype, scorecurrentspan, shouldexportspan, LangfuseMedia), and drift detection. | yonatangross/ | 292 | — | ~2.2k | Automated safety check: Pass | MIT | yesterday |
| 110 | 110.Golang Benchmark Golang benchmarking, profiling, and performance measurement. | aiskillstore/ | 433 | 1 repo | ~3.4k | Automated safety check: Pass | MIT | yesterday |
| 111 | 111.Gorm Expert GORM v2 最佳实践与性能优化。适用于:代码审查、慢查询优化、N+1、连接池、 事务管理、分库分表、Prometheus/OTel监控、Session安全、Clause/Upsert、 缓存集成、BaseModel脚手架、SQL→struct生成、多租户隔离。 | LeoYeAI/ | 2.2k | — | ~3.4k | Automated safety check: Pass | MIT | 2 mo ago |
| 112 | 112.Alerting Oncall Set up alerting rules, configure on-call rotations, and manage incident response workflows. | BagelHole/ | 1.2k | — | ~3k | Automated safety check: Pass | MIT | 4 mo ago |
| 113 | 113.Observability OpenTelemetry, distributed tracing, structured logging, metrics (Prometheus, Grafana, Datadog). | TheBeardedBearSAS/ | 107 | — | ~547 | Automated safety check: Pass | MIT | 27 days ago |
| 114 | Monitoring and observability with OpenTelemetry, Prometheus, Grafana dashboards, and structured logging | rohitg00/ | 2.7k | — | ~1.4k | Automated safety check: Pass | Apache-2.0 | 5 mo ago |
| 115 | Observability: structured logs, metrics (RED/USE), tracing, SLO/SLI. | softspark/ | 179 | — | ~2.2k | Automated safety check: Pass | Apache-2.0 | 3 days ago |
| 116 | Set up Apollo.io monitoring and observability. An agent skill from jeremylongshore/tons-of-skills-marketplace. | jeremylongshore/ | 2.8k | — | ~2.3k | Automated safety check: Pass | MIT | yesterday |
| 117 | Monitor ClickHouse with Prometheus metrics, Grafana dashboards, system table queries, and alerting for query performance, merge health, and resource usage. | jeremylongshore/ | 2.8k | — | ~1.4k | Automated safety check: Pass | MIT | yesterday |
| 118 | Set up Customer.io monitoring and observability. An agent skill from jeremylongshore/tons-of-skills-marketplace. | jeremylongshore/ | 2.8k | — | ~2.8k | Automated safety check: Pass | MIT | yesterday |
| 119 | Set up comprehensive observability for Deepgram integrations. | jeremylongshore/ | 2.8k | — | ~3k | Automated safety check: Pass | MIT | yesterday |
| 120 | Execute Deepgram production deployment checklist. An agent skill from jeremylongshore/tons-of-skills-marketplace. | jeremylongshore/ | 2.8k | — | ~2.3k | Automated safety check: Pass | MIT | yesterday |
| 121 | Implement observability for Evernote integrations. An agent skill from jeremylongshore/tons-of-skills-marketplace. | jeremylongshore/ | 2.8k | — | ~1.2k | Automated safety check: Pass | MIT | yesterday |
| 122 | Set up monitoring, metrics, and alerting for Figma API integrations. | jeremylongshore/ | 2.8k | — | ~2k | Automated safety check: Pass | MIT | yesterday |
| 123 | Implement comprehensive observability for Gamma integrations. | jeremylongshore/ | 2.8k | — | ~2.1k | Automated safety check: Pass | MIT | yesterday |
| 124 | Set up observability for Groq integrations: latency histograms, token throughput, rate limit gauges, cost tracking, and Prometheus alerts. | jeremylongshore/ | 2.8k | — | ~1.6k | Automated safety check: Pass | MIT | yesterday |
| 125 | A skill your agent uses when you need production monitoring for an Intercom integration — instrumenting API calls with metrics and traces, standing up dashboards, or wiring alerts for error rate… | jeremylongshore/ | 2.8k | — | ~1.7k | Automated safety check: Pass | MIT | yesterday |
| 126 | Set up observability for Klaviyo integrations with metrics, traces, and alerts. | jeremylongshore/ | 2.8k | — | ~1.5k | Automated safety check: Pass | MIT | yesterday |
| 127 | Set up comprehensive observability for Langfuse with metrics, dashboards, and alerts. | jeremylongshore/ | 2.8k | — | ~2.2k | Automated safety check: Pass | MIT | yesterday |
| 128 | Implement comprehensive observability for MaintainX integrations. | jeremylongshore/ | 2.8k | — | ~2.1k | Automated safety check: Pass | MIT | yesterday |
| 129 | 129.Monitoring APIs Build real-time API monitoring dashboards with metrics, alerts, and health checks. | jeremylongshore/ | 2.8k | — | ~1.4k | Automated safety check: Pass | MIT | yesterday |
| 130 | Observe a PostHog integration through application-side delivery metrics, ingestion warnings, billing volume, destination logs, and status evidence. | jeremylongshore/ | 2.8k | — | ~2.4k | Automated safety check: Pass | MIT | yesterday |
| 131 | Set up observability for Shopify app integrations with query cost tracking, rate limit monitoring, webhook delivery metrics, and structured logging. | jeremylongshore/ | 2.8k | — | ~1k | Automated safety check: Pass | MIT | yesterday |
| 132 | 132.Grafana Helper Use Grafana's GCX CLI for dashboards, datasources, Prometheus metrics, Loki logs, Tempo traces, alert rules, and Grafana resource operations. | shepherdjerred/ | 112 | — | ~1.1k | Automated safety check: Pass | GPL-3.0 | yesterday |
| 133 | 133.Ops Report Generate a 24-hour operational health report for a PostHog service by querying Grafana dashboards and Prometheus metrics. | haacked/ | 134 | — | ~8k | Automated safety check: Notes | No licence | yesterday |
| 134 | 134.Telemetry Operate the observability stack that deploys as one unit: Prometheus scrape configuration, recording and alerting rules, relabeling, retention, and high availability; OpenTelemetry Collector… | magnus919/ | 115 | — | ~3.9k | Automated safety check: Pass | MIT | yesterday |
| 135 | Observability patterns for logging, monitoring, alerting, and distributed tracing. | TheSoftwareHouse/ | 284 | — | ~2k | Automated safety check: Pass | MIT | 5 days ago |
| 136 | Manage alertmanager rules config operations. An agent skill from jeremylongshore/tons-of-skills-marketplace. | jeremylongshore/ | 2.8k | — | ~579 | Automated safety check: Pass | MIT | yesterday |
| 137 | Set up observability for Claude API integrations with metrics, logging, and alerting for latency, cost, errors, and token usage. | jeremylongshore/ | 2.8k | — | ~1.8k | Automated safety check: Pass | MIT | yesterday |
| 138 | Monitor Clay enrichment pipeline health, credit consumption, and data quality metrics. | jeremylongshore/ | 2.8k | — | ~2.2k | Automated safety check: Pass | MIT | yesterday |
| 139 | Set up comprehensive observability for Lokalise integrations with metrics, traces, and alerts. | jeremylongshore/ | 2.8k | — | ~2.9k | Automated safety check: Pass | MIT | yesterday |
| 140 | Generate prometheus config generator operations. An agent skill from jeremylongshore/tons-of-skills-marketplace. | jeremylongshore/ | 2.8k | — | ~585 | Automated safety check: Pass | MIT | yesterday |
| 141 | 141.Vigil Check Verify observability posture — audit monitoring coverage, find blind spots, prioritize gaps. | jeremylongshore/ | 2.8k | — | ~1k | Automated safety check: Notes | MIT | yesterday |
| 142 | Grafana observability platform — dashboards, Prometheus PromQL, Loki LogQL, alerting, incidents, OnCall schedules, annotations, datasource queries, panel rendering (75+ tools). | automateyournetwork/ | 676 | — | ~2.7k | Automated safety check: Notes | Apache-2.0 | yesterday |
| 143 | Prometheus monitoring — PromQL instant/range queries, metric discovery, metadata, scrape target health, system health checks (6 tools). | automateyournetwork/ | 676 | — | ~2k | Automated safety check: Notes | Apache-2.0 | yesterday |
| 144 | Guides Qdrant monitoring and observability setup. An agent skill from qdrant/skills. | qdrant/ | 254 | — | ~368 | Automated safety check: Pass | Apache-2.0 | yesterday |