Search
Site reliability engineering
Skills
Sort:BestMost starsTrending todayTrending this weekTrending this monthNewestRecently updatedName
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 49 | Configures best-practice, high-signal alerting policies for Cloud Run resources on Google Cloud (services, jobs, and worker pools) based on seasoned SRE practices. | google/ | 21k | — | ~1.7k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 50 | Configures Cloud Monitoring PromQL-based Service Level Objective (SLO) alerting policies on Google Cloud for resources registered in App Hub or individually specified. | google/ | 21k | — | ~3.1k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 51 | Think and work like an expert Precision Engineering Specialist. | K-Dense-AI/ | 200 | — | ~4.2k | Automated safety check: Pass | MIT | 8 days ago |
| 52 | Design and operate service reliability targets in Elastic Observability: choose an SLI type and a defensible target, pick a time window and budgeting method, create and maintain SLOs through the… | elastic/ | 592 | — | ~9.2k | Automated safety check: Pass | Apache-2.0 | 3 days ago |
| 53 | Triage a degraded or suspect service end to end: read SLO status and burn rate, check active alerting rules and ML anomalies, measure throughput, latency, and error rate, assess dependency health… | elastic/ | 592 | — | ~7.4k | Automated safety check: Pass | Apache-2.0 | 3 days ago |
| 54 | Expert performance engineer specializing in modern observability. | Dokhacgiakhoa/ | 508 | — | ~743 | Automated safety check: Pass | Unknown | 4 mo ago |
| 55 | Assess APM service health using SLOs, alerts, ML, throughput, latency, error rate, and dependencies. | aspectrr/ | 405 | — | ~1.2k | Automated safety check: Pass | MIT | 5 mo ago |
| 56 | Work on Supercheck AI SRE agents, incidents, services, connectors, private agents, tool execution, resilience, memory/performance, environment configuration, CI/CD, or release management. | supercheck-io/ | 215 | — | ~1.3k | Automated safety check: Pass | AGPL-3.0 | 2 days ago |
| 57 | When validating system performance under load, identifying bottlenecks through profiling, or optimizing application responsiveness. | ancoleman/ | 525 | — | ~2.9k | Automated safety check: Notes | MIT | 10 mo ago |
| 58 | Design or troubleshoot telemetry, SLOs, and incident diagnostics. | first-fluke/ | 1.3k | — | ~3.7k | Automated safety check: Pass | MIT | yesterday |
| 59 | Design and run a monitoring system for a website or web app. | rampstackco/ | 945 | — | ~2.6k | Automated safety check: Pass | MIT | 4 days ago |
| 60 | Build a multi-quarter roadmap from a backlog of ideas, requests, and ongoing initiatives. | rampstackco/ | 945 | — | ~2.7k | Automated safety check: Pass | MIT | 4 days ago |
| 61 | 61.Cto Review /cs:cto-review <plan — Architecture and scaling interrogation. | alirezarezvani/ | 28k | — | ~848 | Automated safety check: Pass | MIT | 1 mo ago |
| 62 | Index of 37 advanced engineering agent skills for Claude Code, Codex, Gemini CLI, Cursor, OpenClaw. | alirezarezvani/ | 28k | — | ~1.1k | Automated safety check: Pass | MIT | 1 mo ago |
| 63 | 63.Slo Check Define and check simple service-level objectives for Claude Code from Agent Monitor data — session completion rate, tool success rate (PostToolUse/PreToolUse), and error rate (APIError/total) — then… | hoangsonww/ | 1.1k | — | ~769 | Automated safety check: Pass | MIT | yesterday |
| 64 | You are an SLO (Service Level Objective) expert specializing in implementing reliability standards and error budget-based engineering practices. | aiskillstore/ | 433 | 7 repos | ~566 | Automated safety check: Pass | No licence | yesterday |
| 65 | Test or validate Supercheck changes using package checks, Jest, Playwright UI/API E2E, recorder browser tests, contract tests, AI SRE acceptance, release evidence, and commit/merge readiness gates. | supercheck-io/ | 215 | — | ~942 | Automated safety check: Pass | AGPL-3.0 | 2 days ago |
| 66 | Define monitoring strategy, metrics collection, and alerting thresholds during PRD v0.8 Deployment & Ops. | mattgierhart/ | 180 | — | ~3.8k | Automated safety check: Notes | MIT | 1 mo ago |
| 67 | Design, audit, or implement application structured logging architecture from repository evidence. | majiayu000/ | 287 | — | ~1.7k | Automated safety check: Pass | MIT | 2 days ago |
| 68 | OpenDesign's incident retro: the daemon-restart data bug, the root cause, the fix, and the systemic follow-ups. | nexu-io/ | 100k | — | ~1.4k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 69 | 69.Eng Runbook An engineering runbook — service overview, alerts table, dashboards links, common procedures with copy-pasteable commands, on-call rotation, and an incident-response checklist. | sanqiufong/ | 132 | 1 repo | ~380 | Automated safety check: Pass | Apache-2.0 | yesterday |
| 70 | Generates guidance for reliability, resilience, availability, redundancy, fault-tolerance, and disaster recovery (DR) for Google Cloud workloads based on the design principles and recommendations in… | google/ | 21k | — | ~2k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 71 | Authors portable, strongly-typed bioinformatics pipelines in the Common Workflow Language (CWL v1.2) as CommandLineTool/Workflow/ExpressionTool documents, validated with cwltool and run at scale on… | GPTomics/ | 1.2k | 1 repo | ~4.6k | Automated safety check: Pass | MIT | 1 mo ago |
| 72 | Define and track SLAs, SLIs, and SLOs for service reliability including availability, latency, and error rates. | jeremylongshore/ | 2.8k | — | ~1k | Automated safety check: Pass | MIT | yesterday |
| 73 | Generates reliability-focused guidance for Google Cloud workloads based on the Google Cloud Well-Architected Framework. | davila7/ | 33k | — | ~1.8k | Automated safety check: Pass | MIT | yesterday |
| 74 | Design scalable distributed systems using structured approaches for load balancing, caching, database scaling, and message queues. | ninehills/ | 280 | — | ~4.7k | Automated safety check: Pass | MIT | 3 mo ago |
| 75 | Design or improve observability for application and delivery flows: logs, metrics, traces, correlation, alerts, and operational diagnostics. | managedcode/ | 138 | — | ~976 | Automated safety check: Pass | MIT | 4 days ago |
| 76 | Design, audit, and troubleshoot production monitoring and observability using user-impact checks, layered telemetry, USE/RED, SLI/SLO/SLA, error budgets, cardinality controls, actionable alerting… | AnastasiyaW/ | 154 | — | ~4.1k | Automated safety check: Pass | MIT | yesterday |
| 77 | 77.Vpe Advisor VP of Engineering advisor on org design, productivity, quality, delivery, and capacity planning. | borghei/ | 891 | — | ~2.3k | Automated safety check: Pass | MIT | 4 days ago |
| 78 | Designs and runs performance and load tests: workload modelling from real traffic, thresholds tied to SLOs, warmup and ramp shapes, percentile-based analysis, and lightweight CI perf checks with k6… | jaktestowac/ | 116 | — | ~2.8k | Automated safety check: Pass | MIT | 1 mo ago |
| 79 | Route full software-development architecture work from product intent through design, implementation, testing, release, and operations. | majiayu000/ | 287 | — | ~848 | Automated safety check: Pass | MIT | 2 days ago |
| 80 | Observability and SRE expert. An agent skill from majiayu000/spellbook. | majiayu000/ | 287 | — | ~3.3k | Automated safety check: Pass | MIT | 2 days ago |
| 81 | [omh] Postmortem for an outage or SLO miss: postmortems, SLOs, error budgets, incident follow-ups, and service reliability evidence. | rlaope/ | 3.2k | — | ~1.6k | Automated safety check: Pass | MIT | yesterday |
| 82 | An on-call veteran SRE interviewer focused on monitoring and alerting. | PrepLabsAI/ | 112 | — | ~3.9k | Automated safety check: Pass | MIT | 4 days ago |
| 83 | A Principal SRE interviewer focused on fault tolerance and monitoring. | PrepLabsAI/ | 112 | — | ~2.4k | Automated safety check: Pass | MIT | 4 days ago |
| 84 | 84.Cx Slos Manage Coralogix SLO (Service Level Objective) definitions with the cx slos CLI — list and inspect SLOs, check whether targets and error budgets are healthy, and create, update, or delete SLO… | coralogix/ | 121 | — | ~1k | Automated safety check: Pass | Apache-2.0 | 4 days ago |
| 85 | Scheduled probes that run CONTINUOUSLY after release. An agent skill from petrkindlmann/qa-skills. | petrkindlmann/ | 170 | — | ~5.8k | Automated safety check: Pass | MIT | 4 mo ago |
| 86 | Safe-release techniques DURING rollout: feature flags, progressive rollouts, canary analysis, guardrail metrics, production smoke tests, and synthetic users. | petrkindlmann/ | 170 | — | ~5.3k | Automated safety check: Pass | MIT | 4 mo ago |
| 87 | Think and work like an expert Experimental Physicist. An agent skill from K-Dense-AI/scientific-agents. | K-Dense-AI/ | 200 | — | ~7.7k | Automated safety check: Pass | MIT | 8 days ago |
| 88 | 88.Release It Build production-ready systems with stability patterns: circuit breakers, bulkheads, timeouts, and retry logic. | wondelai/ | 2.4k | — | ~4k | Automated safety check: Pass | MIT | 1 mo ago |
| 89 | Design scalable distributed systems using structured approaches for load balancing, caching, database scaling, and message queues. | wondelai/ | 2.4k | — | ~4k | Automated safety check: Pass | MIT | 1 mo ago |
| 90 | 生产可用性 / SRE 专家 Owner — 当任务涉及发布风险、运行稳定性、可观测性、容量、资源生命周期、内存泄漏、回滚、故障恢复、运行手册、长连接、队列、缓存或生产验收时使用;要求把实现映射到可运行、可监控、可恢复、可回滚。 | devcodex-labs/ | 439 | — | ~575 | Automated safety check: Pass | AGPL-3.0 | 24 days ago |
| 91 | Build autonomous AI agents with Claude Agent SDK. An agent skill from LeoYeAI/openclaw-master-skills. | LeoYeAI/ | 2.2k | — | ~7.3k | Automated safety check: Notes | MIT | 2 mo ago |
| 92 | Sets up AWS Resilience Hub v2 from scratch: creates resilience policies with SLO targets, registers systems and user journeys, onboards services with input sources, and runs a first failure mode… | aws/ | 2.8k | — | ~968 | Automated safety check: Pass | Apache-2.0 | yesterday |
| 93 | Design and operationalize SRE dashboards that surface reliability, latency, error, saturation, and capacity signals across services. | BagelHole/ | 1.2k | — | ~903 | Automated safety check: Pass | MIT | 4 mo ago |
| 94 | Production incident response. An agent skill from borghei/Claude-Skills. | borghei/ | 891 | — | ~1.8k | Automated safety check: Pass | MIT | 4 days ago |
| 95 | Design observability strategies: SLI/SLO frameworks, alerting, and dashboards. | borghei/ | 891 | — | ~1.8k | Automated safety check: Pass | MIT | 4 days ago |
| 96 | 96.Scrum Master Data-driven Scrum Master for sprint health scoring, Monte Carlo velocity forecasting, retrospective analysis, capacity planning, and Tuckman team coaching. | borghei/ | 891 | — | ~2.5k | Automated safety check: Pass | MIT | 4 days ago |