Topic · DevOps & Cloud
Best site reliability engineering skills for Claude Code, Codex and other agents.
- skills
- 150
- official
- 11
Site reliability engineering skills, ranked
Ranked by score. Sort bymost stars,trending,newest,recently updated
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 1 | Audible code-quality reactions for code reading, review, refactoring, debugging, and repository exploration. | AndrewVos/ | 243 | — | ~1.3k | Automated safety check: Pass | No licence | 5 mo ago |
| 2 | Analyze, benchmark, diagnose, and optimize large-model inference deployments from hardware inventory, model details, workload traces, and latency or throughput SLOs. | rednote-machine-learning/ | 142 | — | ~4.5k | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 3 | A skill your agent uses when running a previously designed distributed-systems test plan against a real or simulated cluster — driving fault injection, workload, chaos scenarios, linearizability /… | shenli/ | 231 | — | ~5.1k | Automated safety check: Notes | MIT | 2 mo ago |
| 4 | Diagnose failed, hung, slow, OOM, NaN, illegal-memory-access, NCCL, compilation, rollout, or training runs in AReno. | inclusionAI/ | 323 | — | ~486 | Automated safety check: Pass | Apache-2.0 | 13 days ago |
| 5 | Configure Grafana Alerting, Incident Response Management (IRM), and SLOs end-to-end — provisions Grafana-managed and data-source-managed alert rules, contact points (Slack/PagerDuty/email/webhook)… | grafana/ | 278 | 1 repo | ~1.9k | Automated safety check: Pass | Apache-2.0 | today |
| 6 | Data Cloud 360° view of a single Agentforce session. An agent skill from forcedotcom/sf-skills. | forcedotcom/ | 1.1k | — | ~2.9k | Automated safety check: Pass | Apache-2.0 | 4 days ago |
| 7 | Write, validate, and optimize PromQL for Prometheus / Grafana Mimir / Grafana Cloud Metrics. | grafana/ | 278 | 1 repo | ~1.1k | Automated safety check: Pass | Apache-2.0 | today |
| 8 | Define and implement Service Level Indicators (SLIs) and Service Level Objectives (SLOs) with error budgets and alerting. | wshobson/ | 40k | 10 repos | ~1.7k | Automated safety check: Pass | MIT | 2 days ago |
| 9 | Set up tracing, metrics and dashboards for Istio, Linkerd and other service meshes, with golden-signal alerts, SLOs and guidance on sampling and cardinality. | wshobson/ | 40k | 8 repos | ~607 | Automated safety check: Pass | MIT | 2 days ago |
| 10 | Observability best practices. An agent skill from 2SSK/dot-files. | 2SSK/ | 247 | — | ~676 | Automated safety check: Pass | MIT | 25 days ago |
| 11 | Monitoring and observability strategy, implementation, and troubleshooting. | ahmedasmar/ | 203 | — | ~3.9k | Automated safety check: Pass | No licence | 5 mo ago |
| 12 | Annual and quarterly revenue plan construction, top-down vs bottoms-up reconciliation, plan versioning, stretch goal handling, and FP&A-RevOps collaboration for B2B revenue teams. | swan-gtm/ | 165 | — | ~7.2k | Automated safety check: Pass | MIT | 3 days ago |
| 13 | Quantifies elevated error rates with PromQL, compares them to a baseline, and isolates which jobs or instances an error spike is concentrated in. | prometheus/ | 117 | — | ~592 | Automated safety check: Pass | Apache-2.0 | 3 days ago |
| 14 | Command reference for omniroute's quota and resilience commands: inspect circuit breakers, cooldowns, lockouts and quota pools, reset stuck providers and tune thresholds. | diegosouzapw/ | 74k | — | ~624 | Automated safety check: Pass | MIT | today |
| 15 | Design error handling, structured logging, and observability with OpenTelemetry (traces, metrics, logs), error classification, recovery patterns (retry with jitter, circuit breaker, bulkhead… | EliasOulkadi/ | 114 | — | ~3.6k | Automated safety check: Notes | MIT | 2 days ago |
| 16 | 16.Expert Ops 基础设施运维专家入口。用于 Codex CLI 的 $expert-ops 调用. An agent skill from ReJeCtAll/ExpertTeam-Codex. | ReJeCtAll/ | 113 | — | ~625 | Automated safety check: Pass | MIT | 3 mo ago |
| 17 | An on-call SRE interviewer who just got paged about a broken checkout API. | PrepLabsAI/ | 112 | — | ~2.6k | Automated safety check: Pass | MIT | today |
| 18 | A skill your agent uses to set up a CI gate that fails the build when Compose stability silently regresses, using the skydoves/compose-stability-analyzer Gradle plugin (primary) or the… | rosuH/ | 1.9k | 1 repo | ~4k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 19 | Build production-ready monitoring, logging, and tracing systems. | davila7/ | 32k | 7 repos | ~3.2k | Automated safety check: Pass | MIT | today |
| 20 | Designs chaos experiments, failure injection and game days for distributed systems, with blast radius limits, rollback plans and written learnings. | Jeffallan/ | 12k | — | ~1.8k | Automated safety check: Pass | MIT | 4 days ago |
| 21 | 21.SRE Engineer Defines SLIs, SLOs and error budgets, and sets up golden-signal monitoring, blameless postmortems, toil automation and chaos experiments for production systems. | Jeffallan/ | 12k | — | ~1.7k | Automated safety check: Pass | MIT | 4 days ago |
| 22 | Expert SRE incident responder specializing in rapid problem resolution, modern observability, and comprehensive incident management. | davila7/ | 32k | 7 repos | ~2.6k | Automated safety check: Pass | MIT | today |
| 23 | Maps architectural components in a codebase and measures their size to identify what should be extracted first. | tech-leads-club/ | 7k | — | ~3.4k | Automated safety check: Pass | Unknown | 17 days ago |
| 24 | 24.Sre [production-grade internal] Makes systems reliable in production — SLOs, monitoring, alerting, chaos engineering, incident runbooks, capacity planning. | nagisanzenin/ | 181 | — | ~3.2k | Automated safety check: Pass | No licence | 1 mo ago |
| 25 | Headcount and delivery-capacity planning — effective capacity from raw headcount, hire/contract/defer scenarios, and capacity-vs-commitment gap reports. | borghei/ | 874 | — | ~3k | Automated safety check: Pass | MIT | today |
| 26 | Comprehensive Scrum Master assistant for sprint planning, backlog grooming, retrospectives, capacity planning, and daily standups with intelligent context-aware reporting | alirezarezvani/ | 879 | — | ~3.3k | Automated safety check: Pass | MIT | 10 mo ago |
| 27 | Run stepped HTTP load tests with ab/wrk, ramping concurrency levels to collect p50/p90/p99 latency, detect performance inflection points, and recommend optimal concurrency. | zebbern/ | 4.6k | — | ~1.6k | Automated safety check: Notes | MIT | today |
| 28 | Validates and normalizes raw AIPerf outputs, then evaluates valid results against target SLOs and comparable prior candidates. | ai-dynamo/ | 8.2k | — | ~2.7k | Automated safety check: Pass | Apache-2.0 | today |
| 29 | 29.Mint Enroll SRE runbook for enrolling new GitHub repos into the fullsend token mint service using go run ./cmd/fullsend from this checkout. | fullsend-ai/ | 147 | — | ~3.6k | Automated safety check: Notes | Apache-2.0 | today |
| 30 | 30.PR Review Review a GitHub pull request using multiple expert personas. | agentic-community/ | 962 | — | ~4.7k | Automated safety check: Notes | Apache-2.0 | yesterday |
| 31 | Guidance for Azure DDoS Protection — Network Protection (per-VNet) and IP Protection (per public IP) tiers built on the same always-on Microsoft platform. | vinayaklatthe/ | 175 | — | ~2k | Automated safety check: Pass | MIT | 3 mo ago |
| 32 | 32.Expert Team 专家团总路由器。用于 Codex CLI 的 $expert-team 调用. An agent skill from ReJeCtAll/ExpertTeam-Codex. | ReJeCtAll/ | 113 | — | ~667 | Automated safety check: Pass | MIT | 3 mo ago |
| 33 | Build a cloud, SLO, and incident-readiness register after intake. | sickn33/ | 47k | 1 repo | ~6.5k | Automated safety check: Pass | MIT | yesterday |
| 34 | Search and read official Coralogix platform documentation using cx docs search and cx docs fetch. | coralogix/ | 121 | — | ~1.2k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 35 | A skill your agent uses when a Head of Ops, Knowledge Manager, or TPM-Internal needs to author, validate, or clean up company SOPs and internal runbooks (procurement intake, vendor offboarding… | alirezarezvani/ | 28k | — | ~4.1k | Automated safety check: Pass | MIT | 1 mo ago |
| 36 | Design production-ready observability strategies combining metrics, logs, and traces. | alirezarezvani/ | 28k | — | ~3.5k | Automated safety check: Pass | MIT | 1 mo ago |
| 37 | A skill your agent uses when a BizOps lead, COO, or process-improvement owner needs to document an end-to-end business process (procurement, employee onboarding, incident handoff… | alirezarezvani/ | 28k | — | ~2.2k | Automated safety check: Pass | MIT | 1 mo ago |
| 38 | A skill your agent uses when defining, reviewing, or operating SLOs/SLIs/error budgets. | alirezarezvani/ | 28k | — | ~2.6k | Automated safety check: Pass | MIT | 1 mo ago |
| 39 | Configures alerting policies in Terraform for Google Kubernetes Engine (GKE) clusters, workloads, and services using PromQL and Google Cloud Managed Service for Prometheus. | google/ | 21k | — | ~5.3k | Automated safety check: Pass | Apache-2.0 | today |
| 40 | Intent-based observability + traceability router across layers, boundaries, and signals. | first-fluke/ | 1.3k | — | ~4.9k | Automated safety check: Pass | MIT | yesterday |
| 41 | Implements infrastructure as code using Terraform, Kubernetes, and cloud platforms. | davila7/ | 32k | 1 repo | ~1.6k | Automated safety check: Pass | MIT | today |
| 42 | Manages IT infrastructure, monitoring, incident response, and service reliability. | davila7/ | 32k | 1 repo | ~3.7k | Automated safety check: Pass | MIT | today |
| 43 | Reference document for monopoly scale-benchmarks. An agent skill from sickn33/agentic-awesome-skills. | sickn33/ | 47k | 1 repo | ~1.4k | Automated safety check: Pass | MIT | yesterday |
| 44 | Design and operationalize SRE dashboards that surface reliability, latency, error, saturation, and capacity signals across services. | sickn33/ | 47k | 1 repo | ~1.1k | Automated safety check: Pass | MIT | yesterday |
| 45 | Configures best-practice, high-signal alerting policies for Google Cloud Run resources (services, jobs, and worker pools) based on seasoned SRE practices. | google/ | 21k | — | ~1.8k | Automated safety check: Pass | Apache-2.0 | today |
| 46 | Configures PromQL-based Service Level Objective (SLO) alerting policies for Google Cloud resources registered in App Hub or individually specified. | google/ | 21k | — | ~3k | Automated safety check: Pass | Apache-2.0 | today |
| 47 | Think and work like an expert Precision Engineering Specialist. | K-Dense-AI/ | 200 | — | ~4.2k | Automated safety check: Pass | MIT | 5 days ago |
| 48 | Design and operate service reliability targets in Elastic Observability: choose an SLI type and a defensible target, pick a time window and budgeting method, create and maintain SLOs through the… | elastic/ | 592 | — | ~9.2k | Automated safety check: Pass | Apache-2.0 | 5 days ago |
Questions, answered from the data.
What is the best site reliability engineering skill?
Endless Toil from AndrewVos/endless-toil ranks first of the 150 site reliability engineering skills listed here, with the highest score: its repository has 243 GitHub stars, its SKILL.md loads about 1.3k tokens and it passes the automated safety check with no findings. Next come Inference Autopilot and Executing Distributed System Tests.
Which site reliability engineering skills are official?
11 of the 150 site reliability engineering skills are official, published by the vendor's own GitHub organization: Alerting Irm, Promql, Gke Alert Configuration, Cloud Run Alert Configuration, Google Cloud Slo Alert Configuration and 6 more.
How are these skills ranked?
By Skill Navigator score, which combines the GitHub stars of the skill's repository (shared across that repo's skills and discounted for large collections), how many other GitHub owners carry a copy of the skill, and automated SKILL.md quality checks, minus penalties for safety-check warnings and for each further skill from the same repository. Skills that fail the safety check are not listed.
Explore related skills
Category
More topics in DevOps & Cloud
- Deployment1,152
- CI/CD921
- Containers723
- Observability562
- Container orchestration519
- Infrastructure as code351
- Monitoring and alerting333
- Secrets management318
- Runbooks and postmortems277
- Incident response271
- Cloud networking230
- Backup and disaster recovery170
- Cloud architecture114
- MLOps101
- Cloud cost optimization93
- GitOps87
- Linux administration75
- Platform engineering51
- Chaos engineering28