Search
Site reliability engineering
Skills
Sort:BestMost starsTrending todayTrending this weekTrending this monthNewestRecently updatedName
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 1 | Audible code-quality reactions for code reading, review, refactoring, debugging, and repository exploration. | AndrewVos/ | 243 | — | ~1.3k | Automated safety check: Pass | No licence | 5 mo ago |
| 2 | Analyze, benchmark, diagnose, and optimize large-model inference deployments from hardware inventory, model details, workload traces, and latency or throughput SLOs. | rednote-machine-learning/ | 144 | — | ~4.5k | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 3 | A skill your agent uses when running a previously designed distributed-systems test plan against a real or simulated cluster — driving fault injection, workload, chaos scenarios, linearizability /… | shenli/ | 231 | — | ~5.1k | Automated safety check: Notes | MIT | 2 mo ago |
| 4 | Diagnose failed, hung, slow, OOM, NaN, illegal-memory-access, NCCL, compilation, rollout, or training runs in AReno. | inclusionAI/ | 323 | — | ~486 | Automated safety check: Pass | Apache-2.0 | today |
| 5 | Configure Grafana Alerting, Incident Response Management (IRM), and SLOs end-to-end — provisions Grafana-managed and data-source-managed alert rules, contact points (Slack/PagerDuty/email/webhook)… | grafana/ | 281 | 1 repo | ~1.9k | Automated safety check: Pass | Apache-2.0 | today |
| 6 | Define and implement Service Level Indicators (SLIs) and Service Level Objectives (SLOs) with error budgets and alerting. | wshobson/ | 40k | 11 repos | ~1.7k | Automated safety check: Pass | MIT | 4 days ago |
| 7 | Data Cloud 360° view of a single Agentforce session. An agent skill from forcedotcom/sf-skills. | forcedotcom/ | 1.1k | — | ~3.4k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 8 | Write, validate, and optimize PromQL for Prometheus / Grafana Mimir / Grafana Cloud Metrics. | grafana/ | 281 | 1 repo | ~1.1k | Automated safety check: Pass | Apache-2.0 | today |
| 9 | Set up tracing, metrics and dashboards for Istio, Linkerd and other service meshes, with golden-signal alerts, SLOs and guidance on sampling and cardinality. | wshobson/ | 40k | 9 repos | ~607 | Automated safety check: Pass | MIT | 4 days ago |
| 10 | Observability best practices. An agent skill from 2SSK/dot-files. | 2SSK/ | 247 | — | ~676 | Automated safety check: Pass | MIT | today |
| 11 | Monitoring and observability strategy, implementation, and troubleshooting. | ahmedasmar/ | 203 | — | ~3.9k | Automated safety check: Pass | No licence | 6 mo ago |
| 12 | Annual and quarterly revenue plan construction, top-down vs bottoms-up reconciliation, plan versioning, stretch goal handling, and FP&A-RevOps collaboration for B2B revenue teams. | swan-gtm/ | 171 | — | ~7.2k | Automated safety check: Pass | MIT | yesterday |
| 13 | Designs chaos experiments, failure injection and game days for distributed systems, with blast radius limits, rollback plans and written learnings. | Jeffallan/ | 12k | — | ~1.8k | Automated safety check: Pass | MIT | 5 days ago |
| 14 | Quantifies elevated error rates with PromQL, compares them to a baseline, and isolates which jobs or instances an error spike is concentrated in. | prometheus/ | 120 | — | ~592 | Automated safety check: Pass | Apache-2.0 | 4 days ago |
| 15 | Command reference for omniroute's quota and resilience commands: inspect circuit breakers, cooldowns, lockouts and quota pools, reset stuck providers and tune thresholds. | diegosouzapw/ | 74k | — | ~624 | Automated safety check: Pass | MIT | today |
| 16 | Execute this skill enables AI assistant to analyze capacity requirements and plan for future growth. | jeremylongshore/ | 2.8k | — | ~947 | Automated safety check: Pass | MIT | today |
| 17 | Design error handling, structured logging, and observability with OpenTelemetry (traces, metrics, logs), error classification, recovery patterns (retry with jitter, circuit breaker, bulkhead… | EliasOulkadi/ | 114 | — | ~3.6k | Automated safety check: Notes | MIT | 4 days ago |
| 18 | Build production-ready monitoring, logging, and tracing systems. | davila7/ | 32k | 8 repos | ~3.2k | Automated safety check: Pass | MIT | today |
| 19 | 19.Expert Ops 基础设施运维专家入口。用于 Codex CLI 的 $expert-ops 调用. An agent skill from ReJeCtAll/ExpertTeam-Codex. | ReJeCtAll/ | 113 | — | ~625 | Automated safety check: Pass | MIT | 3 mo ago |
| 20 | An on-call SRE interviewer who just got paged about a broken checkout API. | PrepLabsAI/ | 112 | — | ~2.6k | Automated safety check: Pass | MIT | 2 days ago |
| 21 | A skill your agent uses to set up a CI gate that fails the build when Compose stability silently regresses, using the skydoves/compose-stability-analyzer Gradle plugin (primary) or the… | rosuH/ | 1.9k | 1 repo | ~4k | Automated safety check: Pass | Apache-2.0 | 3 days ago |
| 22 | 22.SRE Engineer Defines SLIs, SLOs and error budgets, and sets up golden-signal monitoring, blameless postmortems, toil automation and chaos experiments for production systems. | Jeffallan/ | 12k | — | ~1.7k | Automated safety check: Pass | MIT | 5 days ago |
| 23 | Expert SRE incident responder specializing in rapid problem resolution, modern observability, and comprehensive incident management. | davila7/ | 32k | 7 repos | ~2.6k | Automated safety check: Pass | MIT | today |
| 24 | Use this skill during any incident investigation, capacity planning, or operational troubleshooting when the issue may be caused by hitting AWS service limits. | aws/ | 102 | — | ~3.4k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 25 | Maps architectural components in a codebase and measures their size to identify what should be extracted first. | tech-leads-club/ | 7k | — | ~3.4k | Automated safety check: Pass | Unknown | yesterday |
| 26 | 26.Sre [production-grade internal] Makes systems reliable in production — SLOs, monitoring, alerting, chaos engineering, incident runbooks, capacity planning. | nagisanzenin/ | 181 | — | ~3.2k | Automated safety check: Pass | No licence | 1 mo ago |
| 27 | Comprehensive Scrum Master assistant for sprint planning, backlog grooming, retrospectives, capacity planning, and daily standups with intelligent context-aware reporting | alirezarezvani/ | 880 | — | ~3.3k | Automated safety check: Pass | MIT | 11 mo ago |
| 28 | Expert SRE incident responder specializing in rapid problem resolution. | Dokhacgiakhoa/ | 508 | — | ~706 | Automated safety check: Pass | Unknown | 4 mo ago |
| 29 | Run stepped HTTP load tests with ab/wrk, ramping concurrency levels to collect p50/p90/p99 latency, detect performance inflection points, and recommend optimal concurrency. | zebbern/ | 4.7k | — | ~1.6k | Automated safety check: Notes | MIT | today |
| 30 | Headcount and delivery-capacity planning — effective capacity from raw headcount, hire/contract/defer scenarios, and capacity-vs-commitment gap reports. | borghei/ | 886 | — | ~3k | Automated safety check: Pass | MIT | 2 days ago |
| 31 | Validates and normalizes raw AIPerf outputs, then evaluates valid results against target SLOs and comparable prior candidates. | ai-dynamo/ | 8.3k | — | ~2.7k | Automated safety check: Pass | Apache-2.0 | today |
| 32 | 32.Mint Enroll SRE runbook for enrolling new GitHub repos into the fullsend token mint service using go run ./cmd/fullsend from this checkout. | fullsend-ai/ | 149 | — | ~3.6k | Automated safety check: Notes | Apache-2.0 | today |
| 33 | 33.PR Review Review a GitHub pull request using multiple expert personas. | agentic-community/ | 967 | — | ~4.7k | Automated safety check: Notes | Apache-2.0 | 3 days ago |
| 34 | Guidance for Azure DDoS Protection — Network Protection (per-VNet) and IP Protection (per public IP) tiers built on the same always-on Microsoft platform. | vinayaklatthe/ | 175 | — | ~2k | Automated safety check: Pass | MIT | 3 mo ago |
| 35 | Build AI-focused SRE incident response practices for LLM outages, degraded quality, runaway cost events, and safety regressions. | sickn33/ | 47k | 2 repos | ~3k | Automated safety check: Pass | MIT | today |
| 36 | 36.Expert Team 专家团总路由器。用于 Codex CLI 的 $expert-team 调用. An agent skill from ReJeCtAll/ExpertTeam-Codex. | ReJeCtAll/ | 113 | — | ~667 | Automated safety check: Pass | MIT | 3 mo ago |
| 37 | Build a cloud, SLO, and incident-readiness register after intake. | sickn33/ | 47k | 1 repo | ~6.5k | Automated safety check: Pass | MIT | today |
| 38 | Search and read official Coralogix platform documentation using cx docs search and cx docs fetch. | coralogix/ | 121 | — | ~1.2k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 39 | A skill your agent uses when a Head of Ops, Knowledge Manager, or TPM-Internal needs to author, validate, or clean up company SOPs and internal runbooks (procurement intake, vendor offboarding… | alirezarezvani/ | 28k | — | ~4.1k | Automated safety check: Pass | MIT | 1 mo ago |
| 40 | Design production-ready observability strategies combining metrics, logs, and traces. | alirezarezvani/ | 28k | — | ~3.5k | Automated safety check: Pass | MIT | 1 mo ago |
| 41 | A skill your agent uses when a BizOps lead, COO, or process-improvement owner needs to document an end-to-end business process (procurement, employee onboarding, incident handoff… | alirezarezvani/ | 28k | — | ~2.2k | Automated safety check: Pass | MIT | 1 mo ago |
| 42 | A skill your agent uses when defining, reviewing, or operating SLOs/SLIs/error budgets. | alirezarezvani/ | 28k | — | ~2.6k | Automated safety check: Pass | MIT | 1 mo ago |
| 43 | Configures alerting policies in Terraform for Google Kubernetes Engine (GKE) clusters, workloads, and services using PromQL and Google Cloud Managed Service for Prometheus. | google/ | 21k | — | ~5.3k | Automated safety check: Pass | Apache-2.0 | today |
| 44 | Intent-based observability + traceability router across layers, boundaries, and signals. | first-fluke/ | 1.3k | — | ~4.9k | Automated safety check: Pass | MIT | today |
| 45 | Implements infrastructure as code using Terraform, Kubernetes, and cloud platforms. | davila7/ | 32k | 1 repo | ~1.6k | Automated safety check: Pass | MIT | today |
| 46 | Manages IT infrastructure, monitoring, incident response, and service reliability. | davila7/ | 32k | 1 repo | ~3.7k | Automated safety check: Pass | MIT | today |
| 47 | Reference document for monopoly scale-benchmarks. An agent skill from sickn33/agentic-awesome-skills. | sickn33/ | 47k | 1 repo | ~1.4k | Automated safety check: Pass | MIT | today |
| 48 | Design and operationalize SRE dashboards that surface reliability, latency, error, saturation, and capacity signals across services. | sickn33/ | 47k | 1 repo | ~1.1k | Automated safety check: Pass | MIT | today |