Agent skill

SRE Engineer

by Jeffallan in Jeffallan/claude-skills

Defines SLIs, SLOs and error budgets, and sets up golden-signal monitoring, blameless postmortems, toil automation and chaos experiments for production systems.

MITAuto-check passedDevOps & Cloud

Install SRE Engineer

skills CLI
$ npx skills add Jeffallan/claude-skills --skill sre-engineer -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Jeffallan/claude-skills sre-engineer --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Jeffallan/claude-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/sre-engineer .claude/skills/sre-engineer && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
sre-engineer
GitHub stars
12k
Token cost
~1.7k tokens
SKILL.md length
284 words
Files
6 (incl. references)
Skills in repo
58
Repo updated
First seen
Licence
MIT

At a glance

Defines SLIs, SLOs and error budgets, and sets up golden-signal monitoring, blameless postmortems, toil automation and chaos experiments for production systems.

  • Works in 6 steps: Assess reliability - Review… → Define SLOs - Identify meaningful SLIs… → Verify alignment - Confirm SLO targets… → …
  • Defining SLIs and SLOs for a new service
  • SKILL.md covers Core Workflow, Reference Guide, Constraints and Output Templates, plus 1 more section
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

The agent reviews architecture, SLOs, incidents and toil, defines meaningful SLIs with targets, and checks those targets against user expectations before building golden signal dashboards and alerts. It then automates repetitive tasks and designs chaos experiments, verifying end to end that recovery meets RTO and RPO targets before an experiment counts as complete.

The rules call for quantitative SLOs such as 99.9% availability, error budgets derived from the targets, blameless postmortems for every incident, measured toil and runbooks behind alerts. They warn against SLOs with no user impact reasoning, toil above 50% with no automation plan, and ignoring error budget exhaustion. Outputs are SLO definitions, monitoring configuration such as Prometheus, automation scripts, runbooks and a short note on reliability impact. A worked example computes the downtime a 99.9% SLO allows over 30 days.

When your agent uses it

  • Defining SLIs and SLOs for a new service
  • Calculating error budgets and burn-rate policies
  • Designing golden signal dashboards and actionable alerts
  • Planning chaos experiments and checking recovery targets
  • Writing blameless postmortems and automating repetitive operations

Example prompts

  • “Define SLOs for our checkout API and work out the monthly error budget.”
  • “Design alerts based on the four golden signals, each linked to a runbook.”
  • “Write a blameless postmortem template and fill it in for last week's database failover.”
  • “Plan a chaos experiment that kills one worker node and verify recovery against our RTO.”

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Assess reliability - Review architecture, SLOs, incidents, toil levels
  2. Define SLOs - Identify meaningful SLIs and set appropriate targets
  3. Verify alignment - Confirm SLO targets reflect user expectations before proceeding
  4. Implement monitoring - Build golden signal dashboards and alerting
  5. Automate toil - Identify repetitive tasks and build automation
  6. Test resilience - Design and execute chaos experiments; verify recovery meets RTO/RPO targets before marking the experiment complete…

What it can do on your machine

Read from SKILL.md and the folder at commit 1be15d8. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are yaml, promql and python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • github.com
    • synergetic.solutions
    • jeffallan.github.io

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

SRE Engineer loads about 1.7k tokens when it runs, and up to ~17k if it reads all its reference files. Until then it costs about 99 tokens; SKILL.md has 284 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~99
When it runs · the whole SKILL.md, loaded when a task matches
~1.7k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~17k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Jeffallan/claude-skills at commit 1be15d8, republished under its MIT licence (© Jeffallan). 284 words, ~1,730 tokens.

Download SKILL.mdSave it as .claude/skills/sre-engineer/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.
name
sre-engineer
description
Defines service level objectives, creates error budget policies, designs incident response procedures, develops capacity models, and produces monitoring configurations and automation scripts for production systems. Use when defining SLIs/SLOs, managing error budgets, building reliable systems at scale, incident management, chaos engineering, toil reduction, or capacity planning.
license
MIT
metadata.author
https://github.com/Jeffallan
metadata.company
https://synergetic.solutions
metadata.version
1.1.0
metadata.domain
devops
metadata.triggers
SRE, site reliability, SLO, SLI, error budget, incident management, chaos engineering, toil reduction, on-call, MTTR
metadata.role
specialist
metadata.scope
implementation
metadata.output-format
code
metadata.related-skills
devops-engineer, cloud-architect, kubernetes-specialist

SRE Engineer

Core Workflow

  1. Assess reliability - Review architecture, SLOs, incidents, toil levels
  2. Define SLOs - Identify meaningful SLIs and set appropriate targets
  3. Verify alignment - Confirm SLO targets reflect user expectations before proceeding
  4. Implement monitoring - Build golden signal dashboards and alerting
  5. Automate toil - Identify repetitive tasks and build automation
  6. Test resilience - Design and execute chaos experiments; verify recovery meets RTO/RPO targets before marking the experiment complete; validate recovery behavior end-to-end

Reference Guide

Load detailed guidance based on context:

TopicReferenceLoad When
SLO/SLIreferences/slo-sli-management.mdDefining SLOs, calculating error budgets
Error Budgetsreferences/error-budget-policy.mdManaging budgets, burn rates, policies
Monitoringreferences/monitoring-alerting.mdGolden signals, alert design, dashboards
Automationreferences/automation-toil.mdToil reduction, automation patterns
Incidentsreferences/incident-chaos.mdIncident response, chaos engineering

Constraints

MUST DO
  • Define quantitative SLOs (e.g., 99.9% availability)
  • Calculate error budgets from SLO targets
  • Monitor golden signals (latency, traffic, errors, saturation)
  • Write blameless postmortems for all incidents
  • Measure toil and track reduction progress
  • Automate repetitive operational tasks
  • Test failure scenarios with chaos engineering
  • Balance reliability with feature velocity
MUST NOT DO
  • Set SLOs without user impact justification
  • Alert on symptoms without actionable runbooks
  • Tolerate >50% toil without automation plan
  • Skip postmortems or assign blame
  • Implement manual processes for recurring tasks
  • Deploy without capacity planning
  • Ignore error budget exhaustion
  • Build systems that can't degrade gracefully

Output Templates

When implementing SRE practices, provide:

  1. SLO definitions with SLI measurements and targets
  2. Monitoring/alerting configuration (Prometheus, etc.)
  3. Automation scripts (Python, Go, Terraform)
  4. Runbooks with clear remediation steps
  5. Brief explanation of reliability impact

Concrete Examples

SLO Definition & Error Budget Calculation
# 99.9% availability SLO over a 30-day window
# Allowed downtime: (1 - 0.999) * 30 * 24 * 60 = 43.2 minutes/month
# Error budget (request-based): 0.001 * total_requests

# Example: 10M requests/month → 10,000 error budget requests
# If 5,000 errors consumed in week 1 → 50% budget burned in 25% of window
# → Trigger error budget policy: freeze non-critical releases
Prometheus SLO Alerting Rule (Multiwindow Burn Rate)
yaml
groups:
  - name: slo_availability
    rules:
      # Fast burn: 2% budget in 1h (14.4x burn rate)
      - alert: HighErrorBudgetBurn
        expr: |
          (
            sum(rate(http_requests_total{status=~"5.."}[1h]))
            /
            sum(rate(http_requests_total[1h]))
          ) > 0.014400
          and
          (
            sum(rate(http_requests_total{status=~"5.."}[5m]))
            /
            sum(rate(http_requests_total[5m]))
          ) > 0.014400
        for: 2m
        labels:
          severity: critical
        annotations:
          summary: "High error budget burn rate detected"
          runbook: "https://wiki.internal/runbooks/high-error-burn"

      # Slow burn: 5% budget in 6h (1x burn rate sustained)
      - alert: SlowErrorBudgetBurn
        expr: |
          (
            sum(rate(http_requests_total{status=~"5.."}[6h]))
            /
            sum(rate(http_requests_total[6h]))
          ) > 0.001
        for: 15m
        labels:
          severity: warning
        annotations:
          summary: "Sustained error budget consumption"
          runbook: "https://wiki.internal/runbooks/slow-error-burn"
PromQL Golden Signal Queries
promql
# Latency — 99th percentile request duration
histogram_quantile(0.99, sum(rate(http_request_duration_seconds_bucket[5m])) by (le, service))

# Traffic — requests per second by service
sum(rate(http_requests_total[5m])) by (service)

# Errors — error rate ratio
sum(rate(http_requests_total{status=~"5.."}[5m])) by (service)
  /
sum(rate(http_requests_total[5m])) by (service)

# Saturation — CPU throttling ratio
sum(rate(container_cpu_cfs_throttled_seconds_total[5m])) by (pod)
  /
sum(rate(container_cpu_cfs_periods_total[5m])) by (pod)
Toil Automation Script (Python)
python
#!/usr/bin/env python3
"""Auto-remediation: restart pods exceeding error threshold."""
import subprocess, sys, json

ERROR_THRESHOLD = 0.05  # 5% error rate triggers restart

def get_error_rate(service: str) -> float:
    """Query Prometheus for current error rate."""
    import urllib.request
    query = f'sum(rate(http_requests_total{{status=~"5..",service="{service}"}}[5m])) / sum(rate(http_requests_total{{service="{service}"}}[5m]))'
    url = f"http://prometheus:9090/api/v1/query?query={urllib.request.quote(query)}"
    with urllib.request.urlopen(url) as resp:
        data = json.load(resp)
    results = data["data"]["result"]
    return float(results[0]["value"][1]) if results else 0.0

def restart_deployment(namespace: str, deployment: str) -> None:
    subprocess.run(
        ["kubectl", "rollout", "restart", f"deployment/{deployment}", "-n", namespace],
        check=True
    )
    print(f"Restarted {namespace}/{deployment}")

if __name__ == "__main__":
    service, namespace, deployment = sys.argv[1], sys.argv[2], sys.argv[3]
    rate = get_error_rate(service)
    print(f"Error rate for {service}: {rate:.2%}")
    if rate > ERROR_THRESHOLD:
        restart_deployment(namespace, deployment)
    else:
        print("Within SLO threshold — no action required")

Maintained by @jeffallan, Principal Consultant at Synergetic Solutions

Documentation

© Jeffallan, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 5 other files (references) in skills/sre-engineer of Jeffallan/claude-skills.

  • SKILL.md
  • references/automation-toil.md
  • references/error-budget-policy.md
  • references/incident-chaos.md
  • references/monitoring-alerting.md
  • references/slo-sli-management.md

Open the folder on GitHubat commit 1be15d8

Compare with similar skills

SRE Engineer next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

SRE Engineer compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
SRE Engineer this skillJeffallan/claude-skills12k—~1.7kAutomated safety check: PassMIT
Alerting Irmgrafana/skills2781 repos~1.9kAutomated safety check: PassApache-2.0
Pagerduty Experttheneoai/awesome-skills183—~3.4kAutomated safety check: PassMIT
Site Reliability Engineertheneoai/awesome-skills183—~2.3kAutomated safety check: PassMIT
Monitoring EngineerFerroxLabs/wayland608—~3.9kAutomated safety check: PassApache-2.0
Promqlgrafana/skills2781 repos~1.1kAutomated safety check: PassApache-2.0

Similar skills

  • Alerting Irm

    grafana/skills

    Official

    Configure Grafana Alerting, Incident Response Management (IRM), and SLOs end-to-end — provisions Grafana-managed and data-source-managed alert rules, contact points (Slack/PagerDuty/email/webhook)…

    278 GitHub starsUsed in 1 repo~1.9k tokens
    DevOps & CloudAuto-check passed
  • Pagerduty Expert

    theneoai/awesome-skills

    Invoke when: User needs help with PagerDuty alerting policies, on-call scheduling, incident workflows, or SRE practices.

    183 GitHub stars~3.4k tokensUpdated 4 mo ago
    DevOps & CloudAuto-check passed
  • Site Reliability Engineer

    theneoai/awesome-skills

    Elite Site Reliability Engineer skill with expertise in SLO/SLI definition, incident management, chaos engineering, observability (Prometheus, Grafana, Datadog), and building self-healing systems.

    183 GitHub stars~2.3k tokensUpdated 4 mo ago
    DevOps & CloudAuto-check passed
  • Monitoring Engineer

    FerroxLabs/wayland

    Observability and monitoring. An agent skill from FerroxLabs/wayland.

    608 GitHub stars~3.9k tokensUpdated yesterday
    DevOps & CloudAuto-check passed
  • Promql

    grafana/skills

    Official

    Write, validate, and optimize PromQL for Prometheus / Grafana Mimir / Grafana Cloud Metrics.

    278 GitHub starsUsed in 1 repo~1.1k tokens
    DevOps & CloudAuto-check passed
  • Superset Incident Triage

    superset-sh/superset

    Does a read-only first pass on a possible production incident: gathers deploy, Sentry and health-check signals, proposes a severity and status message, then stops for human approval.

    15k GitHub stars~1k tokensUpdated today
    DevOps & CloudAuto-check passed

More from Jeffallan/claude-skills

All 58 skills in this repo
  • API Designer

    Jeffallan/claude-skills

    Designs REST and GraphQL APIs from resource modeling to an OpenAPI 3.1 contract, with versioning, pagination and RFC 7807 error handling.

    12k GitHub starsUsed in 2 repos~2k tokens
    Auto-check passed
  • CLI Developer

    Jeffallan/claude-skills

    Walks through designing, building and polishing a command-line tool: user workflow and command hierarchy, implementation in commander, click, typer or cobra, completions and cross-platform testing.

    12k GitHub starsUsed in 1 repo~1.2k tokens
    Auto-check passed
  • Fine-Tuning Expert

    Jeffallan/claude-skills

    Guides LLM fine-tuning with LoRA and QLoRA through Hugging Face PEFT, from dataset validation and training checks to adapter merging, quantization and deployment.

    12k GitHub starsUsed in 1 repo~1.7k tokens
    Auto-check passed
  • GraphQL Architect

    Jeffallan/claude-skills

    Designs GraphQL schemas and Apollo Federation graphs, with DataLoader resolvers, subscriptions, query complexity limits and caching.

    12k GitHub starsUsed in 1 repo~1.3k tokens
    Auto-check passed
  • Kubernetes Specialist

    Jeffallan/claude-skills

    Creates and checks Kubernetes manifests, Helm charts, RBAC and network policies, and helps debug pod problems, with kubectl checks and rollback steps.

    12k GitHub starsUsed in 1 repo~2.1k tokens
    Auto-check passed
  • Laravel Specialist

    Jeffallan/claude-skills

    Builds Laravel 10+ applications with Eloquent models, Sanctum authentication, Horizon queues, API resources and Livewire components, tested with Pest or PHPUnit.

    12k GitHub starsUsed in 1 repo~2.1k tokens
    Auto-check passed

Works with

Categories

Questions about SRE Engineer

What does SRE Engineer do?

Defines SLIs, SLOs and error budgets, and sets up golden-signal monitoring, blameless postmortems, toil automation and chaos experiments for production systems. The agent reviews architecture, SLOs, incidents and toil, defines meaningful SLIs with targets, and checks those targets against user expectations before building golden signal dashboards and alerts. It then automates repetitive tasks and designs chaos experiments, verifying end to end that recovery meets RTO and RPO targets before an experiment counts as complete.

When should I use SRE Engineer?

SRE Engineer fits situations like: defining SLIs and SLOs for a new service; calculating error budgets and burn-rate policies; designing golden signal dashboards and actionable alerts; planning chaos experiments and checking recovery targets.

How do I install SRE Engineer in Claude Code?

Run `npx skills add Jeffallan/claude-skills --skill sre-engineer -a claude-code`. Or copy the skill folder (skills/sre-engineer in Jeffallan/claude-skills) into .claude/skills/sre-engineer in your project. Claude Code loads it when a task matches its description.

How do I install SRE Engineer in Codex?

Run `npx skills add Jeffallan/claude-skills --skill sre-engineer -a codex`. Or copy the skill folder (skills/sre-engineer in Jeffallan/claude-skills) into .agents/skills/sre-engineer in your project. Codex loads it when a task matches its description.

Can I use SRE Engineer in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Jeffallan/claude-skills --skill sre-engineer -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/sre-engineer, .gemini/skills/sre-engineer, .github/skills/sre-engineer and .opencode/skills/sre-engineer in your project.

What does SRE Engineer need to run?

SKILL.md names no scripts, command-line tools or credentials: SRE Engineer is instructions for the agent only.

Does SRE Engineer access the network?

SKILL.md names 3 domains. As links in the text: github.com, synergetic.solutions and jeffallan.github.io. This is read from the text; nothing was executed.

Is SRE Engineer safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does SRE Engineer use?

SRE Engineer is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does SRE Engineer use?

About 1.7k tokens (SKILL.md is roughly 6.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 15k tokens, read only when the agent opens those files.

What are the alternatives to SRE Engineer?

Skills that share tags, products or a category with SRE Engineer: Alerting Irm (grafana/skills, 278 stars), Pagerduty Expert (theneoai/awesome-skills, 183 stars), Site Reliability Engineer (theneoai/awesome-skills, 183 stars) and Monitoring Engineer (FerroxLabs/wayland, 608 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains SRE Engineer?

Jeffallan (a GitHub user) maintains it in Jeffallan/claude-skills, which has 11,754 GitHub stars. The repository holds 58 skills in this directory. The repository was last updated on October 3, 2026.

Source: Jeffallan/claude-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.