Topic · DevOps & Cloud
Best site reliability engineering skills, page 3
Site reliability engineering skills, ranked
Ranked by score. Sort bymost stars,trending,newest,recently updated
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 97 | Handle Venice API errors correctly. An agent skill from veniceai/skills. | veniceai/ | 143 | — | ~4.9k | Automated safety check: Pass | MIT | 2 days ago |
| 98 | A skill your agent uses when designing, operating, reviewing, or improving production reliability with SLOs, incident command, observability, error budgets, and operational excellence. | magnus919/ | 278 | — | ~1.3k | Automated safety check: Pass | MIT | 3 mo ago |
| 99 | Configure Harness AI-powered operations (AIDA) via MCP. An agent skill from majiayu000/claude-skill-registry. | majiayu000/ | 666 | 1 repo | ~1.4k | Automated safety check: Pass | Apache-2.0 | today |
| 100 | A skill your agent uses when planning capacity, redundancy, graceful degradation, or recovery for systems that must keep working under stress — performance budgets, dependency failure handling… | hashgraph-online/ | 1.2k | — | ~1.9k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 101 | 101.Monitoring A skill your agent uses when setting up uptime and health monitoring, alerts, or on-call basics for a service already in production, so you learn it is down before customers do — health and… | ericrisco/ | 156 | — | ~3.1k | Automated safety check: Pass | MIT | yesterday |
| 102 | Observability and monitoring validation patterns for dashboards, alerting, log aggregation, APM traces, and SLA/SLO verification. | proffesor-for-testing/ | 494 | — | ~8.3k | Automated safety check: Pass | MIT | 3 days ago |
| 103 | 103.Telemetry Operate the observability stack that deploys as one unit: Prometheus scrape configuration, recording and alerting rules, relabeling, retention, and high availability; OpenTelemetry Collector… | magnus919/ | 111 | — | ~3.9k | Automated safety check: Pass | MIT | yesterday |
| 104 | 104.Vllm Bench Serve Interactive online benchmark orchestrator for vLLM inference services using vllm bench serve. | ascend-ai-coding/ | 174 | — | ~5.6k | Automated safety check: Pass | No licence | today |
| 105 | SRE patterns for production service reliability: SLOs, error budgets, postmortems, and incident response. | aiskillstore/ | 430 | — | ~1.2k | Automated safety check: Pass | No licence | today |
| 106 | Design production observability strategies covering SLI/SLOs, metrics, logs, traces, dashboards, and alert quality. | aAAaqwq/ | 105 | 1 repo | ~3.3k | Automated safety check: Pass | MIT | 10 days ago |
| 107 | Calculation models and business impact matrices for quantitatively assessing incident impact based on SLA/SLO. | revfactory/ | 1.3k | — | ~1.2k | Automated safety check: Pass | Apache-2.0 | 6 mo ago |
| 108 | 108.Operations Operations excellence expertise for supply chain optimization, process improvement (Lean, Six Sigma), capacity planning, vendor management, quality assurance, and operational efficiency. | travisjneuman/ | 101 | — | ~3.4k | Automated safety check: Pass | MIT | yesterday |
| 109 | Your software-engineering brain trust. An agent skill from coco-research/coco. | coco-research/ | 473 | — | ~3.4k | Automated safety check: Pass | Unknown | today |
| 110 | 110.Claude Agent SDK Build autonomous AI agents with Claude Agent SDK. An agent skill from majiayu000/claude-skill-registry. | majiayu000/ | 666 | 1 repo | ~7.3k | Automated safety check: Notes | MIT | today |
| 111 | Produce a capacity planning document for a service covering traffic forecasts, resource requirements, and scaling strategy. | mohitagw15856/ | 1.4k | — | ~4.1k | Automated safety check: Pass | MIT | yesterday |
| 112 | 112.Slo Error Budget Define Service Level Objectives (SLOs) and an error budget policy for a service. | mohitagw15856/ | 1.4k | — | ~3k | Automated safety check: Pass | MIT | yesterday |
| 113 | Design production plans using MPS (Master Production Schedule), MRP (Material Requirements Planning), and capacity planning. | asgard-ai-platform/ | 241 | — | ~1.3k | Automated safety check: Pass | MIT | 4 mo ago |
| 114 | Design, automate, and operate end-to-end software releases: release process models and pipelines (trunk-based development, CD stages, release trains), progressive delivery and feature flags… | magnus919/ | 111 | — | ~3.9k | Automated safety check: Pass | MIT | yesterday |
| 115 | Unified platform operations guidance for CI/CD pipeline design, deployment strategies, observability, SLI/SLOs, and incident-ready rollouts. | rsmdt/ | 536 | — | ~991 | Automated safety check: Pass | MIT | 2 mo ago |
| 116 | 116.Kdd Workflow A skill your agent uses when planning a KDD project calendar across the venue's two submission cycles per year, including cycle choice, abstract-then-paper deadline pairs a week apart, rebuttal… | brycewang-stanford/ | 1.2k | — | ~1.7k | Automated safety check: Pass | MIT | 10 days ago |
| 117 | A skill your agent uses when load testing a service before launch or after a significant traffic change — writing k6 or Locust scripts, setting SLO-based pass/fail thresholds, diagnosing bottlenecks… | kid-sid/ | 189 | — | ~2.5k | Automated safety check: Pass | MIT | 2 mo ago |
| 118 | Complete performance engineering system — profiling, optimization, load testing, capacity planning, and performance culture. | LeoYeAI/ | 2.2k | — | ~7.1k | Automated safety check: Pass | MIT | 2 mo ago |
| 119 | Reduce Sentry alert fatigue by surgically tuning issue grouping, fingerprint rules, severity mapping, sample rates, before-send filters, sourcemap pipelines, and release-health gates. | LeoYeAI/ | 2.2k | — | ~7.3k | Automated safety check: Pass | MIT | 2 mo ago |
| 120 | 120.Yatta Personal productivity system for task and capacity management. | LeoYeAI/ | 2.2k | — | ~6.4k | Automated safety check: Pass | MIT | 2 mo ago |
| 121 | 121.Cloud Monitoring Monitor cloud infrastructure and applications using metrics, logs, and traces to provide real-time observability into performance, health, and reliability. | seb1n/ | 206 | — | ~2.8k | Automated safety check: Pass | MIT | 1 mo ago |
| 122 | 122.Task Automation Automate repetitive tasks and workflows using scripting, file watchers, scheduled jobs, CI triggers, and API polling to eliminate manual toil. | seb1n/ | 206 | — | ~2.4k | Automated safety check: Pass | MIT | 1 mo ago |
| 123 | Drives an interactive system design session: classifies depth, elicits scale/SLO/consistency inputs, computes capacity, then reveals components one by one, each justified by a constraint. | HoangNguyen0403/ | 570 | — | ~1.8k | Automated safety check: Pass | MIT | yesterday |
| 124 | Model technical capacity, unit cost, and budget constraints connected to demand, performance, and reliability decisions. | magnus919/ | 111 | — | ~4.7k | Automated safety check: Pass | MIT | yesterday |
| 125 | Compare intended product outcomes against observed results to close the launch-to-learning loop: collect post-launch evidence, distinguish expected from observed from uncertain from inferred claims… | magnus919/ | 111 | — | ~4.7k | Automated safety check: Pass | MIT | yesterday |
| 126 | Design, operate, and improve reliable production systems with SLOs, incident command, observability, error budgets, and operational practices. | magnus919/ | 111 | — | ~2.7k | Automated safety check: Pass | MIT | yesterday |
| 127 | Selects and implements NestJS runtime features, error and API contracts, security, testing, DevOps, performance, and safe scale. | aiskillstore/ | 430 | — | ~3.8k | Automated safety check: Pass | MIT | today |
| 128 | Expert knowledge for Azure Reliability development including best practices, decision making, architecture & design patterns, limits & quotas, and deployment. | MicrosoftDocs/ | 775 | — | ~2.2k | Automated safety check: Pass | CC-BY-4.0 | 2 days ago |
| 129 | Expert knowledge for Azure Sre Agent development including troubleshooting, best practices, decision making, architecture & design patterns, security, configuration, integrations & coding patterns… | MicrosoftDocs/ | 775 | — | ~2.7k | Automated safety check: Pass | CC-BY-4.0 | 2 days ago |
| 130 | Infrastructure as Code patterns (Terraform, Kubernetes), observability design (SLOs, metrics, alerting, dashboards), and pipeline security stages. | nWave-ai/ | 617 | — | ~1.5k | Automated safety check: Pass | MIT | 21 days ago |
| 131 | Foundational platform engineering knowledge from key references -- Continuous Delivery, SRE, Accelerate, Team Topologies, Chaos Engineering, and Secure Delivery. | nWave-ai/ | 617 | — | ~1.1k | Automated safety check: Pass | MIT | 21 days ago |
| 132 | SLA/SLO 기반 장애 영향도를 정량적으로 산정하는 계산 모델과 비즈니스 영향 매트릭스. An agent skill from revfactory/harness-100. | revfactory/ | 1.3k | — | ~765 | Automated safety check: Pass | Apache-2.0 | 6 mo ago |
| 133 | Corrects the wrong defaults a model has when building Datadog dashboards, verified against Datadog's docs in July 2026. | pproenca/ | 214 | — | ~2.6k | Automated safety check: Pass | MIT | 1 mo ago |
| 134 | 134.Pagerduty Expert Invoke when: User needs help with PagerDuty alerting policies, on-call scheduling, incident workflows, or SRE practices. | theneoai/ | 183 | — | ~3.4k | Automated safety check: Pass | MIT | 4 mo ago |
| 135 | Elite Site Reliability Engineer skill with expertise in SLO/SLI definition, incident management, chaos engineering, observability (Prometheus, Grafana, Datadog), and building self-healing systems. | theneoai/ | 183 | — | ~2.3k | Automated safety check: Pass | MIT | 4 mo ago |
| 136 | 136.System Architect Expert System Architect with 20+ years designing distributed systems at scale. | theneoai/ | 183 | — | ~3.2k | Automated safety check: Pass | MIT | 4 mo ago |
| 137 | Incident response and analysis via Harness MCP. An agent skill from harness/harness-skills. | harness/ | 115 | — | ~1.7k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 138 | 138.Manage Slos Assist with Harness Service Reliability Management (SRM) tasks that the MCP server currently supports: pulling recent deployments for incident correlation, generating on-call handover reports from… | harness/ | 115 | — | ~1.5k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 139 | 139.Sei Analytics Advanced engineering analytics via Harness Software Engineering Insights (SEI) MCP. | harness/ | 115 | — | ~1.7k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 140 | k6 script templates, load profiles, response time thresholds, SLO validation, and performance testing strategies. | vibeeval/ | 531 | — | ~1.6k | Automated safety check: Pass | MIT | 2 mo ago |
| 141 | PromQL queries, alerting rules, recording rules, Grafana dashboard JSON, SLO | vibeeval/ | 531 | — | ~1k | Automated safety check: Pass | MIT | 2 mo ago |
| 142 | 142.Capacity Planner Capacity planning expertise covering load testing methodologies, autoscaling policies, resource forecasting, performance budgets, cost-capacity curves, bottleneck identification, queue theory… | FerroxLabs/ | 608 | — | ~4.4k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 143 | Infrastructure cost modeling and capacity planning expert covering cloud cost estimation, resource right-sizing, growth projection modeling, reserved vs on-demand analysis, cost-per-request… | FerroxLabs/ | 608 | — | ~4.5k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 144 | Observability and monitoring. An agent skill from FerroxLabs/wayland. | FerroxLabs/ | 608 | — | ~3.9k | Automated safety check: Pass | Apache-2.0 | yesterday |
Explore related skills
Category
More topics in DevOps & Cloud
- Deployment1,152
- CI/CD921
- Containers723
- Observability562
- Container orchestration519
- Infrastructure as code351
- Monitoring and alerting333
- Secrets management318
- Runbooks and postmortems277
- Incident response271
- Cloud networking230
- Backup and disaster recovery170
- Cloud architecture114
- MLOps101
- Cloud cost optimization93
- GitOps87
- Linux administration75
- Platform engineering51
- Chaos engineering28