Search

Prometheus · Site reliability engineering

15 skills found.
Search results
#SkillRepositoryStarsUsed inTokensAuto-checkLicenceUpdated
1
1.Alerting IrmOfficial

Configure Grafana Alerting, Incident Response Management (IRM), and SLOs end-to-end — provisions Grafana-managed and data-source-managed alert rules, contact points (Slack/PagerDuty/email/webhook)…

grafana/skills2821 repo~1.9kAutomated safety check: PassApache-2.02 days ago
2
2.PromqlOfficial

Write, validate, and optimize PromQL for Prometheus / Grafana Mimir / Grafana Cloud Metrics.

grafana/skills2821 repo~1.1kAutomated safety check: PassApache-2.02 days ago
3

Monitoring and observability strategy, implementation, and troubleshooting.

ahmedasmar/devops-claude-skills203—~3.9kAutomated safety check: PassNo licence6 mo ago
4

Quantifies elevated error rates with PromQL, compares them to a baseline, and isolates which jobs or instances an error spike is concentrated in.

prometheus/prometheus-mcp121—~592Automated safety check: PassApache-2.0yesterday
5

基础设施运维专家入口。用于 Codex CLI 的 $expert-ops 调用. An agent skill from ReJeCtAll/ExpertTeam-Codex.

ReJeCtAll/ExpertTeam-Codex113—~625Automated safety check: PassMIT3 mo ago
6

Defines SLIs, SLOs and error budgets, and sets up golden-signal monitoring, blameless postmortems, toil automation and chaos experiments for production systems.

Jeffallan/claude-skills12k—~1.7kAutomated safety check: PassMIT7 days ago
7

Build AI-focused SRE incident response practices for LLM outages, degraded quality, runaway cost events, and safety regressions.

sickn33/agentic-awesome-skills47k2 repos~3kAutomated safety check: PassMIT2 days ago
8

Configures alerting policies in Terraform for Google Kubernetes Engine (GKE) clusters, workloads, and services using PromQL and Google Cloud Managed Service for Prometheus.

google/skills21k—~5.3kAutomated safety check: PassApache-2.0yesterday
9

Configures best-practice, high-signal alerting policies for Cloud Run resources on Google Cloud (services, jobs, and worker pools) based on seasoned SRE practices.

google/skills21k—~1.7kAutomated safety check: PassApache-2.0yesterday
10

Configures Cloud Monitoring PromQL-based Service Level Objective (SLO) alerting policies on Google Cloud for resources registered in App Hub or individually specified.

google/skills21k—~3.1kAutomated safety check: PassApache-2.0yesterday
11

Design, audit, and troubleshoot production monitoring and observability using user-impact checks, layered telemetry, USE/RED, SLI/SLO/SLA, error budgets, cardinality controls, actionable alerting…

AnastasiyaW/codex-claude-code-config154—~4.1kAutomated safety check: PassMITyesterday
12

Observability and SRE expert. An agent skill from majiayu000/spellbook.

majiayu000/spellbook287—~3.3kAutomated safety check: PassMIT2 days ago
13

Observability: structured logs, metrics (RED/USE), tracing, SLO/SLI.

softspark/ai-toolkit179—~2.2kAutomated safety check: PassApache-2.03 days ago
14

Operate the observability stack that deploys as one unit: Prometheus scrape configuration, recording and alerting rules, relabeling, retention, and high availability; OpenTelemetry Collector…

magnus919/agent-skills115—~3.9kAutomated safety check: PassMITyesterday
15

Monitor cloud infrastructure and applications using metrics, logs, and traces to provide real-time observability into performance, health, and reliability.

seb1n/awesome-ai-agent-skills206—~2.8kAutomated safety check: PassMIT2 mo ago