Search

Kubernetes · Site reliability engineering

7 skills found.
Search results
#SkillRepositoryStarsUsed inTokensAuto-checkLicenceUpdated
1

Designs chaos experiments, failure injection and game days for distributed systems, with blast radius limits, rollback plans and written learnings.

Jeffallan/claude-skills12k—~1.8kAutomated safety check: PassMIT7 days ago
2

Design error handling, structured logging, and observability with OpenTelemetry (traces, metrics, logs), error classification, recovery patterns (retry with jitter, circuit breaker, bulkhead…

EliasOulkadi/shokunin114—~3.6kAutomated safety check: NotesMIT6 days ago
3

A skill your agent uses when defining, reviewing, or operating SLOs/SLIs/error budgets.

alirezarezvani/claude-skills28k—~2.6kAutomated safety check: PassMIT1 mo ago
4

Implements infrastructure as code using Terraform, Kubernetes, and cloud platforms.

davila7/claude-code-templates33k1 repo~1.6kAutomated safety check: PassMITyesterday
5

Triage a degraded or suspect service end to end: read SLO status and burn rate, check active alerting rules and ML anomalies, measure throughput, latency, and error rate, assess dependency health…

elastic/agent-skills592—~7.4kAutomated safety check: PassApache-2.03 days ago
6

A skill your agent uses when a user asks to assess, plan, or validate an Amazon EKS cluster upgrade.

aws/tools-for-devops-agent103—~7.7kAutomated safety check: PassApache-2.0yesterday
7

Selects and implements NestJS runtime features, error and API contracts, security, testing, DevOps, performance, and safe scale.

aiskillstore/marketplace433—~3.8kAutomated safety check: PassMITyesterday