Search
Grafana · Site reliability engineering
Skills
Sort:BestMost starsTrending todayTrending this weekTrending this monthNewestRecently updatedName
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 1 | Configure Grafana Alerting, Incident Response Management (IRM), and SLOs end-to-end — provisions Grafana-managed and data-source-managed alert rules, contact points (Slack/PagerDuty/email/webhook)… | grafana/ | 282 | 1 repo | ~1.9k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 2 | Write, validate, and optimize PromQL for Prometheus / Grafana Mimir / Grafana Cloud Metrics. | grafana/ | 282 | 1 repo | ~1.1k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 3 | Monitoring and observability strategy, implementation, and troubleshooting. | ahmedasmar/ | 203 | — | ~3.9k | Automated safety check: Pass | No licence | 6 mo ago |
| 4 | 基础设施运维专家入口。用于 Codex CLI 的 $expert-ops 调用. An agent skill from ReJeCtAll/ExpertTeam-Codex. | ReJeCtAll/ | 113 | — | ~625 | Automated safety check: Pass | MIT | 3 mo ago |
| 5 | Design, audit, and troubleshoot production monitoring and observability using user-impact checks, layered telemetry, USE/RED, SLI/SLO/SLA, error budgets, cardinality controls, actionable alerting… | AnastasiyaW/ | 154 | — | ~4.1k | Automated safety check: Pass | MIT | yesterday |
| 6 | Observability and SRE expert. An agent skill from majiayu000/spellbook. | majiayu000/ | 287 | — | ~3.3k | Automated safety check: Pass | MIT | 2 days ago |
| 7 | An on-call veteran SRE interviewer focused on monitoring and alerting. | PrepLabsAI/ | 112 | — | ~3.9k | Automated safety check: Pass | MIT | 4 days ago |
| 8 | Design observability strategies: SLI/SLO frameworks, alerting, and dashboards. | borghei/ | 891 | — | ~1.8k | Automated safety check: Pass | MIT | 4 days ago |
| 9 | Observability: structured logs, metrics (RED/USE), tracing, SLO/SLI. | softspark/ | 179 | — | ~2.2k | Automated safety check: Pass | Apache-2.0 | 3 days ago |
| 10 | Wire LangChain 1.0 / LangGraph 1.0 traces into an OpenTelemetry-native backend (Jaeger, Honeycomb, Grafana Tempo, Datadog) with LLM-specific SLOs, safe prompt-content policy, and subgraph-aware span… | jeremylongshore/ | 2.8k | — | ~3.6k | Automated safety check: Pass | MIT | yesterday |
| 11 | 11.Telemetry Operate the observability stack that deploys as one unit: Prometheus scrape configuration, recording and alerting rules, relabeling, retention, and high availability; OpenTelemetry Collector… | magnus919/ | 115 | — | ~3.9k | Automated safety check: Pass | MIT | yesterday |
| 12 | 12.Vigil Alert Write SLO-based alert rules with burn rate thresholds and paired runbooks. | jeremylongshore/ | 2.8k | — | ~2.7k | Automated safety check: Notes | MIT | yesterday |
| 13 | Monitor cloud infrastructure and applications using metrics, logs, and traces to provide real-time observability into performance, health, and reliability. | seb1n/ | 206 | — | ~2.8k | Automated safety check: Pass | MIT | 2 mo ago |