Search
Site reliability engineering
Skills
Sort:BestMost starsTrending todayTrending this weekTrending this monthNewestRecently updatedName
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 97 | Create or audit SLOs, SLIs, alert rules, incident response steps, escalation paths, postmortems, operational runbooks, and customer-impact communication. | majiayu000/ | 287 | — | ~460 | Automated safety check: Pass | MIT | 3 days ago |
| 98 | A skill your agent uses when a user asks to assess, plan, or validate an Amazon EKS cluster upgrade. | aws/ | 103 | — | ~7.7k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 99 | Helps engineering managers plan roadmaps, prioritize work, and communicate priorities effectively — produces the 20% tech debt framework (and its 5 traps), a phased release pressure-test, a… | manager-dot-dev/ | 114 | — | ~4.9k | Automated safety check: Pass | MIT | 5 mo ago |
| 100 | 100.Shadow Work Helps engineering managers identify, quantify, and reduce hidden capacity drains that make teams miss commitments even when everyone is busy. | manager-dot-dev/ | 114 | — | ~2.3k | Automated safety check: Pass | MIT | 5 mo ago |
| 101 | Observability: structured logs, metrics (RED/USE), tracing, SLO/SLI. | softspark/ | 179 | — | ~2.2k | Automated safety check: Pass | Apache-2.0 | 3 days ago |
| 102 | Tune Alchemy-backed reads with measured latency, cache semantics, batching, concurrency, and freshness SLOs. | jeremylongshore/ | 2.8k | — | ~1.2k | Automated safety check: Pass | MIT | yesterday |
| 103 | Reduce the operational cost of a BambooHR connector by removing redundant traffic, oversized data retention, retry waste, and support toil without claiming undocumented API prices. | jeremylongshore/ | 2.8k | — | ~1.1k | Automated safety check: Pass | MIT | yesterday |
| 104 | Tune CAST AI node and workload autoscaling against application SLOs, scheduling constraints, and recommendation confidence. | jeremylongshore/ | 2.8k | — | ~1.2k | Automated safety check: Pass | MIT | yesterday |
| 105 | Optimize Clari integration latency without violating concurrency, quota, correctness, or duplicate-work controls. | jeremylongshore/ | 2.8k | — | ~1k | Automated safety check: Pass | MIT | yesterday |
| 106 | 106.Clay Load Scale Scale Clay enrichment pipelines for high-volume processing (10K-100K+ leads/month). | jeremylongshore/ | 2.8k | — | ~2.4k | Automated safety check: Pass | MIT | yesterday |
| 107 | Instrument Firecrawl v2 requests, async jobs, queue pressure, credits, origin status, webhooks, output quality, and downstream delivery without logging content. | jeremylongshore/ | 2.8k | — | ~1.1k | Automated safety check: Pass | MIT | yesterday |
| 108 | Operate a Guidewire Cloud API integration in production — define SLIs/SLOs for token availability, bind success rate, FNOL p99 latency; route alerts so the on-call gets paged for real outages and… | jeremylongshore/ | 2.8k | — | ~3.1k | Automated safety check: Pass | MIT | yesterday |
| 109 | Instrument Ideogram request, queue, async, safety, storage, latency, and spend signals without logging image content. | jeremylongshore/ | 2.8k | — | ~1.1k | Automated safety check: Pass | MIT | yesterday |
| 110 | Triage LangChain 1.0 / LangGraph 1.0 production incidents — LLM-specific SLOs, provider outage runbook, latency spike decision tree, cost-overrun response, agent loop containment. | jeremylongshore/ | 2.8k | — | ~3.8k | Automated safety check: Pass | MIT | yesterday |
| 111 | Wire LangChain 1.0 / LangGraph 1.0 traces into an OpenTelemetry-native backend (Jaeger, Honeycomb, Grafana Tempo, Datadog) with LLM-specific SLOs, safe prompt-content policy, and subgraph-aware span… | jeremylongshore/ | 2.8k | — | ~3.6k | Automated safety check: Pass | MIT | yesterday |
| 112 | Instrument Mistral requests, streams, tools, batch, and workflows without logging sensitive content. | jeremylongshore/ | 2.8k | — | ~901 | Automated safety check: Pass | MIT | yesterday |
| 113 | Load test and scale Vercel deployments with concurrency tuning and capacity planning. | jeremylongshore/ | 2.8k | — | ~2k | Automated safety check: Pass | MIT | yesterday |
| 114 | Makes systems debuggable and reliably operable — instrumentation, alerting that is worth waking for, service objectives, and learning from failure. | cbrock84/ | 2k | — | ~931 | Automated safety check: Pass | MIT | 23 days ago |
| 115 | 115.Delivery Manager Expert delivery management for release planning, deployment strategy, incident response, change management, SLA/error-budget tracking, and DORA metrics across continuous delivery pipelines. | borghei/ | 891 | — | ~2.2k | Automated safety check: Pass | MIT | 4 days ago |
| 116 | 116.Venice Errors Handle Venice API errors correctly. An agent skill from veniceai/skills. | veniceai/ | 144 | — | ~4.9k | Automated safety check: Pass | MIT | 5 days ago |
| 117 | A skill your agent uses when designing, operating, reviewing, or improving production reliability with SLOs, incident command, observability, error budgets, and operational excellence. | magnus919/ | 289 | — | ~1.3k | Automated safety check: Pass | MIT | 3 mo ago |
| 118 | 118.Monitoring A skill your agent uses when setting up uptime and health monitoring, alerts, or on-call basics for a service already in production, so you learn it is down before customers do — health and… | ericrisco/ | 180 | — | ~3.1k | Automated safety check: Pass | MIT | 2 days ago |
| 119 | A skill your agent uses when planning capacity, redundancy, graceful degradation, or recovery for systems that must keep working under stress — performance budgets, dependency failure handling… | hashgraph-online/ | 1.3k | — | ~1.9k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 120 | Observability and monitoring validation patterns for dashboards, alerting, log aggregation, APM traces, and SLA/SLO verification. | proffesor-for-testing/ | 495 | — | ~8.3k | Automated safety check: Pass | MIT | 2 days ago |
| 121 | 121.Telemetry Operate the observability stack that deploys as one unit: Prometheus scrape configuration, recording and alerting rules, relabeling, retention, and high availability; OpenTelemetry Collector… | magnus919/ | 115 | — | ~3.9k | Automated safety check: Pass | MIT | yesterday |
| 122 | 122.Vllm Bench Serve Interactive online benchmark orchestrator for vLLM inference services using vllm bench serve. | ascend-ai-coding/ | 174 | — | ~5.6k | Automated safety check: Pass | No licence | yesterday |
| 123 | SRE patterns for production service reliability: SLOs, error budgets, postmortems, and incident response. | aiskillstore/ | 433 | — | ~1.2k | Automated safety check: Pass | No licence | yesterday |
| 124 | Design production observability strategies covering SLI/SLOs, metrics, logs, traces, dashboards, and alert quality. | aAAaqwq/ | 105 | 1 repo | ~3.3k | Automated safety check: Pass | MIT | 3 days ago |
| 125 | Calculation models and business impact matrices for quantitatively assessing incident impact based on SLA/SLO. | revfactory/ | 1.3k | — | ~1.2k | Automated safety check: Pass | Apache-2.0 | 6 mo ago |
| 126 | 126.Anth Load Scale Implement load testing, auto-scaling, and capacity planning for Claude API. | jeremylongshore/ | 2.8k | — | ~1.8k | Automated safety check: Pass | MIT | yesterday |
| 127 | Measure and improve SerpAPI latency, payload size, connection reuse, caching, and concurrency without breaking freshness or allowance controls. | jeremylongshore/ | 2.8k | — | ~1k | Automated safety check: Pass | MIT | yesterday |
| 128 | 128.Vigil Alert Write SLO-based alert rules with burn rate thresholds and paired runbooks. | jeremylongshore/ | 2.8k | — | ~2.7k | Automated safety check: Notes | MIT | yesterday |
| 129 | 129.Operations Operations excellence expertise for supply chain optimization, process improvement (Lean, Six Sigma), capacity planning, vendor management, quality assurance, and operational efficiency. | travisjneuman/ | 100 | — | ~3.4k | Automated safety check: Pass | MIT | yesterday |
| 130 | Your software-engineering brain trust. An agent skill from coco-research/coco. | coco-research/ | 513 | — | ~3.4k | Automated safety check: Pass | Unknown | yesterday |
| 131 | Produce a capacity planning document for a service covering traffic forecasts, resource requirements, and scaling strategy. | mohitagw15856/ | 1.4k | — | ~4.1k | Automated safety check: Pass | MIT | 2 days ago |
| 132 | 132.Slo Error Budget Define Service Level Objectives (SLOs) and an error budget policy for a service. | mohitagw15856/ | 1.4k | — | ~3k | Automated safety check: Pass | MIT | 2 days ago |
| 133 | Design production plans using MPS (Master Production Schedule), MRP (Material Requirements Planning), and capacity planning. | asgard-ai-platform/ | 242 | — | ~1.3k | Automated safety check: Pass | MIT | 4 mo ago |
| 134 | Design, automate, and operate end-to-end software releases: release process models and pipelines (trunk-based development, CD stages, release trains), progressive delivery and feature flags… | magnus919/ | 115 | — | ~3.9k | Automated safety check: Pass | MIT | yesterday |
| 135 | Unified platform operations guidance for CI/CD pipeline design, deployment strategies, observability, SLI/SLOs, and incident-ready rollouts. | rsmdt/ | 560 | — | ~991 | Automated safety check: Pass | MIT | 2 mo ago |
| 136 | 136.Kdd Workflow A skill your agent uses when planning a KDD project calendar across the venue's two submission cycles per year, including cycle choice, abstract-then-paper deadline pairs a week apart, rebuttal… | brycewang-stanford/ | 1.2k | — | ~1.7k | Automated safety check: Pass | MIT | 14 days ago |
| 137 | A skill your agent uses when load testing a service before launch or after a significant traffic change — writing k6 or Locust scripts, setting SLO-based pass/fail thresholds, diagnosing bottlenecks… | kid-sid/ | 190 | — | ~2.5k | Automated safety check: Pass | MIT | 2 mo ago |
| 138 | Complete performance engineering system — profiling, optimization, load testing, capacity planning, and performance culture. | LeoYeAI/ | 2.2k | — | ~7.1k | Automated safety check: Pass | MIT | 2 mo ago |
| 139 | Reduce Sentry alert fatigue by surgically tuning issue grouping, fingerprint rules, severity mapping, sample rates, before-send filters, sourcemap pipelines, and release-health gates. | LeoYeAI/ | 2.2k | — | ~7.3k | Automated safety check: Pass | MIT | 2 mo ago |
| 140 | 140.Yatta Personal productivity system for task and capacity management. | LeoYeAI/ | 2.2k | — | ~6.4k | Automated safety check: Pass | MIT | 2 mo ago |
| 141 | 141.Cloud Monitoring Monitor cloud infrastructure and applications using metrics, logs, and traces to provide real-time observability into performance, health, and reliability. | seb1n/ | 206 | — | ~2.8k | Automated safety check: Pass | MIT | 2 mo ago |
| 142 | 142.Task Automation Automate repetitive tasks and workflows using scripting, file watchers, scheduled jobs, CI triggers, and API polling to eliminate manual toil. | seb1n/ | 206 | — | ~2.4k | Automated safety check: Pass | MIT | 2 mo ago |
| 143 | Drives an interactive system design session: classifies depth, elicits scale/SLO/consistency inputs, computes capacity, then reveals components one by one, each justified by a constraint. | HoangNguyen0403/ | 572 | — | ~1.8k | Automated safety check: Pass | MIT | 2 days ago |
| 144 | Model technical capacity, unit cost, and budget constraints connected to demand, performance, and reliability decisions. | magnus919/ | 115 | — | ~4.7k | Automated safety check: Pass | MIT | yesterday |