Topic · DevOps & Cloud
Best site reliability engineering skills, page 2
Site reliability engineering skills, ranked
Ranked by score. Sort bymost stars,trending,newest,recently updated
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 49 | Think and work like an expert Precision Engineering Specialist. | K-Dense-AI/ | 200 | — | ~4.2k | Automated safety check: Pass | MIT | 5 days ago |
| 50 | Design and operate service reliability targets in Elastic Observability: choose an SLI type and a defensible target, pick a time window and budgeting method, create and maintain SLOs through the… | elastic/ | 592 | — | ~9.2k | Automated safety check: Pass | Apache-2.0 | today |
| 51 | Triage a degraded or suspect service end to end: read SLO status and burn rate, check active alerting rules and ML anomalies, measure throughput, latency, and error rate, assess dependency health… | elastic/ | 592 | — | ~7.4k | Automated safety check: Pass | Apache-2.0 | today |
| 52 | Expert performance engineer specializing in modern observability. | Dokhacgiakhoa/ | 507 | — | ~743 | Automated safety check: Pass | Unknown | 3 mo ago |
| 53 | Assess APM service health using SLOs, alerts, ML, throughput, latency, error rate, and dependencies. | aspectrr/ | 405 | — | ~1.2k | Automated safety check: Pass | MIT | 5 mo ago |
| 54 | Work on Supercheck AI SRE agents, incidents, services, connectors, private agents, tool execution, resilience, memory/performance, environment configuration, CI/CD, or release management. | supercheck-io/ | 215 | — | ~1.3k | Automated safety check: Pass | AGPL-3.0 | today |
| 55 | When validating system performance under load, identifying bottlenecks through profiling, or optimizing application responsiveness. | ancoleman/ | 526 | — | ~2.9k | Automated safety check: Notes | MIT | 10 mo ago |
| 56 | Design or troubleshoot telemetry, SLOs, and incident diagnostics. | first-fluke/ | 1.3k | — | ~3.7k | Automated safety check: Pass | MIT | today |
| 57 | Design and run a monitoring system for a website or web app. | rampstackco/ | 940 | — | ~2.6k | Automated safety check: Pass | MIT | yesterday |
| 58 | Build a multi-quarter roadmap from a backlog of ideas, requests, and ongoing initiatives. | rampstackco/ | 940 | — | ~2.7k | Automated safety check: Pass | MIT | yesterday |
| 59 | 59.Cto Review /cs:cto-review <plan — Architecture and scaling interrogation. | alirezarezvani/ | 28k | — | ~848 | Automated safety check: Pass | MIT | 1 mo ago |
| 60 | Index of 37 advanced engineering agent skills for Claude Code, Codex, Gemini CLI, Cursor, OpenClaw. | alirezarezvani/ | 28k | — | ~1.1k | Automated safety check: Pass | MIT | 1 mo ago |
| 61 | 61.Slo Check Define and check simple service-level objectives for Claude Code from Agent Monitor data — session completion rate, tool success rate (PostToolUse/PreToolUse), and error rate (APIError/total) — then… | hoangsonww/ | 1.1k | — | ~769 | Automated safety check: Pass | MIT | yesterday |
| 62 | You are an SLO (Service Level Objective) expert specializing in implementing reliability standards and error budget-based engineering practices. | aiskillstore/ | 430 | 7 repos | ~566 | Automated safety check: Pass | No licence | today |
| 63 | Test or validate Supercheck changes using package checks, Jest, Playwright UI/API E2E, recorder browser tests, contract tests, AI SRE acceptance, release evidence, and commit/merge readiness gates. | supercheck-io/ | 215 | — | ~942 | Automated safety check: Pass | AGPL-3.0 | today |
| 64 | Define monitoring strategy, metrics collection, and alerting thresholds during PRD v0.8 Deployment & Ops. | mattgierhart/ | 180 | — | ~3.8k | Automated safety check: Notes | MIT | 1 mo ago |
| 65 | OpenDesign's incident retro: the daemon-restart data bug, the root cause, the fix, and the systemic follow-ups. | nexu-io/ | 100k | — | ~1.4k | Automated safety check: Pass | Apache-2.0 | today |
| 66 | Design, audit, or implement application structured logging architecture from repository evidence. | majiayu000/ | 286 | — | ~1.7k | Automated safety check: Pass | MIT | today |
| 67 | 67.Eng Runbook An engineering runbook — service overview, alerts table, dashboards links, common procedures with copy-pasteable commands, on-call rotation, and an incident-response checklist. | sanqiufong/ | 132 | 1 repo | ~380 | Automated safety check: Pass | Apache-2.0 | yesterday |
| 68 | Generates guidance for reliability, resilience, availability, redundancy, fault-tolerance, and disaster recovery (DR) for Google Cloud workloads based on the design principles and recommendations in… | google/ | 21k | — | ~2k | Automated safety check: Pass | Apache-2.0 | today |
| 69 | Authors portable, strongly-typed bioinformatics pipelines in the Common Workflow Language (CWL v1.2) as CommandLineTool/Workflow/ExpressionTool documents, validated with cwltool and run at scale on… | GPTomics/ | 1.2k | 1 repo | ~4.6k | Automated safety check: Pass | MIT | 1 mo ago |
| 70 | Execute this skill enables AI assistant to analyze capacity requirements and plan for future growth. | jeremylongshore/ | 2.8k | — | ~947 | Automated safety check: Pass | MIT | today |
| 71 | Define and track SLAs, SLIs, and SLOs for service reliability including availability, latency, and error rates. | jeremylongshore/ | 2.8k | — | ~1k | Automated safety check: Pass | MIT | today |
| 72 | Generates reliability-focused guidance for Google Cloud workloads based on the Google Cloud Well-Architected Framework. | davila7/ | 32k | — | ~1.8k | Automated safety check: Pass | MIT | today |
| 73 | Design scalable distributed systems using structured approaches for load balancing, caching, database scaling, and message queues. | ninehills/ | 281 | — | ~4.7k | Automated safety check: Pass | MIT | 3 mo ago |
| 74 | Design or improve observability for application and delivery flows: logs, metrics, traces, correlation, alerts, and operational diagnostics. | managedcode/ | 138 | — | ~976 | Automated safety check: Pass | MIT | yesterday |
| 75 | Design, audit, and troubleshoot production monitoring and observability using user-impact checks, layered telemetry, USE/RED, SLI/SLO/SLA, error budgets, cardinality controls, actionable alerting… | AnastasiyaW/ | 154 | — | ~4.1k | Automated safety check: Pass | MIT | today |
| 76 | 76.Vpe Advisor VP of Engineering advisor on org design, productivity, quality, delivery, and capacity planning. | borghei/ | 881 | — | ~2.3k | Automated safety check: Pass | MIT | today |
| 77 | Designs and runs performance and load tests: workload modelling from real traffic, thresholds tied to SLOs, warmup and ramp shapes, percentile-based analysis, and lightweight CI perf checks with k6… | jaktestowac/ | 116 | — | ~2.8k | Automated safety check: Pass | MIT | 1 mo ago |
| 78 | Route full software-development architecture work from product intent through design, implementation, testing, release, and operations. | majiayu000/ | 286 | — | ~848 | Automated safety check: Pass | MIT | today |
| 79 | Observability and SRE expert. An agent skill from majiayu000/spellbook. | majiayu000/ | 286 | — | ~3.3k | Automated safety check: Pass | MIT | today |
| 80 | [omh] Postmortem for an outage or SLO miss: postmortems, SLOs, error budgets, incident follow-ups, and service reliability evidence. | rlaope/ | 3.2k | — | ~1.6k | Automated safety check: Pass | MIT | today |
| 81 | An on-call veteran SRE interviewer focused on monitoring and alerting. | PrepLabsAI/ | 112 | — | ~3.9k | Automated safety check: Pass | MIT | yesterday |
| 82 | A Principal SRE interviewer focused on fault tolerance and monitoring. | PrepLabsAI/ | 112 | — | ~2.4k | Automated safety check: Pass | MIT | yesterday |
| 83 | 83.Cx Slos Manage Coralogix SLO (Service Level Objective) definitions with the cx slos CLI — list and inspect SLOs, check whether targets and error budgets are healthy, and create, update, or delete SLO… | coralogix/ | 121 | — | ~1k | Automated safety check: Pass | Apache-2.0 | today |
| 84 | Build AI-focused SRE incident response practices for LLM outages, degraded quality, runaway cost events, and safety regressions. | majiayu000/ | 666 | 3 repos | ~2.8k | Automated safety check: Pass | MIT | today |
| 85 | Think and work like an expert Experimental Physicist. An agent skill from K-Dense-AI/scientific-agents. | K-Dense-AI/ | 200 | — | ~7.7k | Automated safety check: Pass | MIT | 5 days ago |
| 86 | Scheduled probes that run CONTINUOUSLY after release. An agent skill from petrkindlmann/qa-skills. | petrkindlmann/ | 165 | — | ~5.8k | Automated safety check: Pass | MIT | 3 mo ago |
| 87 | Safe-release techniques DURING rollout: feature flags, progressive rollouts, canary analysis, guardrail metrics, production smoke tests, and synthetic users. | petrkindlmann/ | 165 | — | ~5.3k | Automated safety check: Pass | MIT | 3 mo ago |
| 88 | 88.Release It Build production-ready systems with stability patterns: circuit breakers, bulkheads, timeouts, and retry logic. | wondelai/ | 2.4k | — | ~4k | Automated safety check: Pass | MIT | 27 days ago |
| 89 | Design scalable distributed systems using structured approaches for load balancing, caching, database scaling, and message queues. | wondelai/ | 2.4k | — | ~4k | Automated safety check: Pass | MIT | 27 days ago |
| 90 | 生产可用性 / SRE 专家 Owner — 当任务涉及发布风险、运行稳定性、可观测性、容量、资源生命周期、内存泄漏、回滚、故障恢复、运行手册、长连接、队列、缓存或生产验收时使用;要求把实现映射到可运行、可监控、可恢复、可回滚。 | devcodex-labs/ | 439 | — | ~575 | Automated safety check: Pass | AGPL-3.0 | 21 days ago |
| 91 | Design and operationalize SRE dashboards that surface reliability, latency, error, saturation, and capacity signals across services. | BagelHole/ | 1.1k | — | ~903 | Automated safety check: Pass | MIT | 4 mo ago |
| 92 | Sets up AWS Resilience Hub v2 from scratch: creates resilience policies with SLO targets, registers systems and user journeys, onboards services with input sources, and runs a first failure mode… | aws/ | 2.8k | — | ~968 | Automated safety check: Pass | Apache-2.0 | today |
| 93 | Production incident response. An agent skill from borghei/Claude-Skills. | borghei/ | 881 | — | ~1.8k | Automated safety check: Pass | MIT | today |
| 94 | Design observability strategies: SLI/SLO frameworks, alerting, and dashboards. | borghei/ | 881 | — | ~1.8k | Automated safety check: Pass | MIT | today |
| 95 | 95.Scrum Master Data-driven Scrum Master for sprint health scoring, Monte Carlo velocity forecasting, retrospective analysis, capacity planning, and Tuckman team coaching. | borghei/ | 881 | — | ~2.5k | Automated safety check: Pass | MIT | today |
| 96 | Create or audit SLOs, SLIs, alert rules, incident response steps, escalation paths, postmortems, operational runbooks, and customer-impact communication. | majiayu000/ | 286 | — | ~460 | Automated safety check: Pass | MIT | today |
Explore related skills
Category
More topics in DevOps & Cloud
- Deployment1,264
- CI/CD977
- Containers731
- Observability636
- Container orchestration542
- Infrastructure as code377
- Monitoring and alerting351
- Secrets management335
- Runbooks and postmortems324
- Incident response302
- Cloud networking225
- Backup and disaster recovery183
- Cloud architecture118
- MLOps100
- Cloud cost optimization92
- GitOps89
- Linux administration70
- Platform engineering45
- Chaos engineering25