Search
Prometheus · Site reliability engineering
Skills
Sort:BestMost starsTrending todayTrending this weekTrending this monthNewestRecently updatedName
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 1 | Configure Grafana Alerting, Incident Response Management (IRM), and SLOs end-to-end — provisions Grafana-managed and data-source-managed alert rules, contact points (Slack/PagerDuty/email/webhook)… | grafana/ | 282 | 1 repo | ~1.9k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 2 | Write, validate, and optimize PromQL for Prometheus / Grafana Mimir / Grafana Cloud Metrics. | grafana/ | 282 | 1 repo | ~1.1k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 3 | Monitoring and observability strategy, implementation, and troubleshooting. | ahmedasmar/ | 203 | — | ~3.9k | Automated safety check: Pass | No licence | 6 mo ago |
| 4 | Quantifies elevated error rates with PromQL, compares them to a baseline, and isolates which jobs or instances an error spike is concentrated in. | prometheus/ | 121 | — | ~592 | Automated safety check: Pass | Apache-2.0 | yesterday |
| 5 | 基础设施运维专家入口。用于 Codex CLI 的 $expert-ops 调用. An agent skill from ReJeCtAll/ExpertTeam-Codex. | ReJeCtAll/ | 113 | — | ~625 | Automated safety check: Pass | MIT | 3 mo ago |
| 6 | Defines SLIs, SLOs and error budgets, and sets up golden-signal monitoring, blameless postmortems, toil automation and chaos experiments for production systems. | Jeffallan/ | 12k | — | ~1.7k | Automated safety check: Pass | MIT | 7 days ago |
| 7 | Build AI-focused SRE incident response practices for LLM outages, degraded quality, runaway cost events, and safety regressions. | sickn33/ | 47k | 2 repos | ~3k | Automated safety check: Pass | MIT | 2 days ago |
| 8 | Configures alerting policies in Terraform for Google Kubernetes Engine (GKE) clusters, workloads, and services using PromQL and Google Cloud Managed Service for Prometheus. | google/ | 21k | — | ~5.3k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 9 | Configures best-practice, high-signal alerting policies for Cloud Run resources on Google Cloud (services, jobs, and worker pools) based on seasoned SRE practices. | google/ | 21k | — | ~1.7k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 10 | Configures Cloud Monitoring PromQL-based Service Level Objective (SLO) alerting policies on Google Cloud for resources registered in App Hub or individually specified. | google/ | 21k | — | ~3.1k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 11 | Design, audit, and troubleshoot production monitoring and observability using user-impact checks, layered telemetry, USE/RED, SLI/SLO/SLA, error budgets, cardinality controls, actionable alerting… | AnastasiyaW/ | 154 | — | ~4.1k | Automated safety check: Pass | MIT | yesterday |
| 12 | Observability and SRE expert. An agent skill from majiayu000/spellbook. | majiayu000/ | 287 | — | ~3.3k | Automated safety check: Pass | MIT | 2 days ago |
| 13 | Observability: structured logs, metrics (RED/USE), tracing, SLO/SLI. | softspark/ | 179 | — | ~2.2k | Automated safety check: Pass | Apache-2.0 | 3 days ago |
| 14 | 14.Telemetry Operate the observability stack that deploys as one unit: Prometheus scrape configuration, recording and alerting rules, relabeling, retention, and high availability; OpenTelemetry Collector… | magnus919/ | 115 | — | ~3.9k | Automated safety check: Pass | MIT | yesterday |
| 15 | Monitor cloud infrastructure and applications using metrics, logs, and traces to provide real-time observability into performance, health, and reliability. | seb1n/ | 206 | — | ~2.8k | Automated safety check: Pass | MIT | 2 mo ago |