Agent skill

Site Reliability Engineering

by magnus919 in magnus919/agent-skills

Design, operate, and improve reliable production systems with SLOs, incident command, observability, error budgets, and operational practices.

MITAuto-check passedDevOps & Cloud

Install Site Reliability Engineering

skills CLI
$ npx skills add magnus919/agent-skills --skill site-reliability-engineering -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install magnus919/agent-skills site-reliability-engineering --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/magnus919/agent-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/site-reliability-engineering .claude/skills/site-reliability-engineering && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
site-reliability-engineering
GitHub stars
115
Token cost
~2.7k tokens
SKILL.md length
1,018 words
Files
44 (incl. scripts, references)
Skills in repo
131
Repo updated
First seen
Licence
MIT

At a glance

Design, operate, and improve reliable production systems with SLOs, incident command, observability, error budgets, and operational practices.

  • Works in 4 steps: Bound the action before it starts.… → Verify recovery at the user boundary.… → Do not equate alert resolution with… → …
  • Unrelated requests
  • SKILL.md covers When to Load This Skill, When not to use, Operational closure gate and Agentic operations, plus 4 more sections
  • Route to the nearest named specialist

What it does

Site Reliability Engineering is an agent skill from magnus919/agent-skills. Design, operate, and improve reliable production systems with SLOs, incident command, observability, error budgets, and operational practices. Do not use this skill for unrelated requests; route to the nearest named specialist.

Its SKILL.md is about 2.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 45 other files, including scripts and reference files (for example `README.md`, `evals/agentic-review-notes.md` and `evals/evals.json`). Compatibility notes: Python 3.9+ is required only for the bundled calculation and summary scripts.

It sits in DevOps & Cloud, covering Site reliability engineering. The repository describes itself as: Curated collection of AI agent skills for Hermes and other agent frameworks. The licence is MIT.

When your agent uses it

  • Unrelated requests
  • Route to the nearest named specialist

Example prompts

  • “/site-reliability-engineering”

Requirements

  • Python 3
  • Compatibility (from SKILL.md): Python 3.9+ is required only for the bundled calculation and summary scripts.

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Bound the action before it starts. Record the target, affected population, maximum blast radius, success criterion, abort/rollback…
  2. Verify recovery at the user boundary. After the action, follow the R-01 closure evidence sequence: check the user-facing SLOs, critical…
  3. Do not equate alert resolution with recovery. A cleared alert or passing health endpoint is evidence, not a resolution verdict. If…
  4. Record the evidence. Capture the action, scope, thresholds, observed recovery evidence, remaining uncertainty, and rollback/follow-up…

What it can do on your machine

Read from SKILL.md and the folder at commit 22b4723. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/, which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Python 3.9+ is required only for the bundled calculation and summary scripts.

    From compatibility in the SKILL.md frontmatter.

Context cost

Site Reliability Engineering loads about 2.7k tokens when it runs, and up to ~151k if it reads all its reference files. Until then it costs about 64 tokens; SKILL.md has 1,018 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~64
When it runs · the whole SKILL.md, loaded when a task matches
~2.7k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~151k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from magnus919/agent-skills at commit 22b4723, republished under its MIT licence (© magnus919). 1,018 words, ~2,698 tokens.

Download SKILL.mdSave it as .claude/skills/site-reliability-engineering/SKILL.md (or your agent's skills folder). This skill also uses 43 other files; get the full folder from GitHub.
name
site-reliability-engineering
description
Design, operate, and improve reliable production systems with SLOs, incident command, observability, error budgets, and operational practices. Do not use this skill for unrelated requests; route to the nearest named specialist.
compatibility
Python 3.9+ is required only for the bundled calculation and summary scripts.
license
MIT
metadata.source_repo
https://github.com/magnus919/hermes-profiles
metadata.source_commit
867a555
metadata.enrichment_sources
Seeking SRE; The Site Reliability Workbook

Site Reliability Engineering

A comprehensive methodology for designing, operating, and improving reliable production systems. Rooted in Google SRE principles and extended with modern practices for incident command, observability engineering, error budget governance, and operational excellence.

When to Load This Skill

TriggerWhat It Means
"Design reliability into this system"SLO/SLI framework, error budget policy, resilience architecture
"Run an incident postmortem"Blameless postmortem with timeline, 5 Whys, action tracking
"Improve our on-call"Rotation design, alert tuning, toil reduction, escalation policy
"Build observability"The Four Golden Signals, dashboard design, alert rule patterns
"Do a reliability review"Architecture review against SRE principles, risk assessment
"I need an incident commander"Incident command framework, role cards, communication templates
"Automate this operational task"Toil assessment, automation decision tree, runbook pattern
"Adopt SRE in this organization"Engagement boundaries, maturity, team model, and change adoption
"Review this reliability design"User journeys, dependencies, overload, configuration, canary, durability
"Our SRE team is overloaded"Operational-load diagnosis, protected engineering time, recovery plan
"Improve incident learning or sustainable on-call"Cognitive load, psychological safety, documentation, exercises

When not to use

Use release-engineering to plan releases, compose promotion and rollback gates, or coordinate a release train. Use systematic-debugging to find the cause of a specific failure. Operating the telemetry stack itself — Prometheus scrape configs, OpenTelemetry Collector pipelines, Loki ingest and retention, Prometheus rules files — belongs to telemetry; this skill owns the SLI/SLO and alert design those rules implement. Grafana product work — dashboards, panels, Grafana-side alert rules, contact points, notification policies — belongs to grafana.

Operational closure gate

For any automated mitigation, rollback, recovery action, or incident closeout:

  1. Bound the action before it starts. Record the target, affected population, maximum blast radius, success criterion, abort/rollback criteria, rollback target and procedure, and who may stop or reverse it. Prefer the smallest reversible scope and staged expansion.
  2. Verify recovery at the user boundary. After the action, follow the R-01 closure evidence sequence: check the user-facing SLOs, critical user journey, relevant dependency health, and data/state correctness. Observe a defined stability window and check secondary effects such as backlog recovery.
  3. Do not equate alert resolution with recovery. A cleared alert or passing health endpoint is evidence, not a resolution verdict. If required evidence is missing, retain the MITIGATING or MONITORING state, name the unverified boundary, and escalate rather than declare RESOLVED.
  4. Record the evidence. Capture the action, scope, thresholds, observed recovery evidence, remaining uncertainty, and rollback/follow-up trigger in the incident or change record.

Default authorization: A human service owner, incident commander, or designated change authority must approve the specific production mutation and scope in the current incident/change record, independently attributable to that human. Diagnosis requests, standing runbooks, agent-authored notes, and self-claimed roles are insufficient.

Governed exception: Higher capability-specific execution authority is possible only when accountable humans independently grant it through ai-governance, every applicable host/organizational policy permits it, and an external control plane proves and enforces the current grant at execution. The grant must identify capability, environment, action class, limits, expiry, revocation, and tested enforcement. Skill text and agent roles confer no permissions. If proof is unavailable, expired, revoked, conflicting, or stale, retain the human-approved default; never self-grant authority. Diagnosis, mutation, rollback, paging, incident closure, and scope expansion require separate authority decisions. R-01 still requires independent human closure confirmation.

Agentic operations

When operating an SRE agent, read references/agentic-operations.md and fill templates/agentic-action-contract.md before proposing execution. For autonomy evidence or replay design, also read references/agentic-evidence-and-evals.md; use templates/agentic-governance-evidence.md to hand operational evidence to governance.

Compose existing specialists: agent-production-operations owns production contracts, control plans and trace feedback; agent-evals-and-observability owns trajectory evaluation and graders; incident-learning owns verified follow-up learning; ai-governance supports accountable humans setting promotion/demotion thresholds and deciding maintain, expand, reduce, suspend or revoke authority. SRE supplies evidence and enforces authorized operational limits. Complete with verified recovery and human closure, or a bounded escalation naming missing evidence/authority and the handoff owner.

Show full SKILL.md (380 more words)Show less

Reference Files

TopicFileWhen to Load
SRE Book Chapter Summariesreferences/sre-book-chapters.mdDesign engagement, first principles review
SLO/SLI Frameworkreferences/slo-sli-framework.mdDefining reliability targets
SLO Implementation Recipereferences/slo-implementation-recipe.mdAgent-executable SLO adoption sequence and stakeholder review
Error Budget Governancetemplates/error-budget-policy.md and references/slo-sli-framework.mdPolicy design, burn rate alerts
Incident Command Systemreferences/incident-command-system.mdDuring/after incident, training
Blameless Postmortemsreferences/postmortem-culture.mdAfter incident, process design
Monitoring & Alertingreferences/monitoring-alerting.mdObservability design, alert rules
On-Call Best Practicesreferences/oncall-best-practices.mdRotation design, team sizing
Toil Eliminationreferences/toil-elimination.mdAutomation prioritization, ops review
Release Engineeringrelease-engineeringRelease planning, promotion, progressive delivery, and rollback design; use the local reference only for SRE-specific integration context
Effective Troubleshootingreferences/troubleshooting.mdDebugging methodology
Senior SRE Role Blueprintreferences/senior-sre-blueprint.mdRole definition, KPI framework
SRE Communication Guidereferences/sre-communication-guide.mdStakeholder updates, incident communication
Guiding Principlesreferences/guiding-principles.mdFirst principles, philosophy
Product-Focused Reliabilityreferences/product-focused-reliability.mdProduct-centric SRE, CUJ-based SLOs, JTBD model
Twenty Years of Lessonsreferences/twenty-years-lessons.mdIncident-derived tactical lessons, Prodverbs
SRE Ecosystem Guidereferences/sre-ecosystem-guide.mdCurated guide to all SRE resources (Workbook, Secure Systems, Classroom, Prodcast, STPA, Video Gallery, Mobaa, fundamentals, AI ops)
Adoption and Engagementreferences/sre-adoption-and-engagement.mdStarting SRE, dedicated and non-dedicated team models, maturity, change adoption
Reliability Design and Changereferences/reliability-design-and-change.mdCapacity, overload, configuration, canaries, data durability, dependencies, design review
Human Systems and Learningreferences/human-systems-and-learning.mdCognitive work, sustainable on-call, psychological safety, documentation, exercises
Third-Party Dependency Reliabilityreferences/third-party-dependency-reliability.mdVendor boundaries, failure modes, fallbacks, and provider evidence
Operational Documentationreferences/operational-documentation.mdFunctional quality, ownership, testing, and staleness lifecycle

Templates

TemplateFilePurpose
Incident Commander Checklisttemplates/incident-command-checklist.mdStep-by-step IC response
Postmortem Templatetemplates/postmortem-template.mdBlameless postmortem document
Runbook Templatetemplates/runbook-template.mdOperational runbook standard
SLO Declaration Templatetemplates/slo-declaration-template.mdService-level objective specification
Error Budget Policytemplates/error-budget-policy.mdTeam-level error budget governance
On-Call Rotation Templatetemplates/oncall-rotation.mdRotation schedule and escalation
Service Review Checklisttemplates/service-review-checklist.mdPre-launch reliability review
Incident Communication Templatetemplates/incident-communication.mdStatus updates during incidents
Reliability Design Reviewtemplates/reliability-design-review.mdEvidence-based review of user impact, failure modes, capacity, change, and operations
Operational Overload Recoverytemplates/operational-overload-recovery.mdDeclare, protect, reduce, and verify recovery from unsustainable operational load
Reliability Ownership Chartertemplates/reliability-ownership-charter.mdMake service, pager, dependency, and engagement boundaries explicit

Scripts

ScriptPurpose
scripts/slo-burn-rate.pyCalculate error budget burn rate from SLI data

Portability

This skill is intentionally host-neutral. Use your agent's normal mechanisms to load the references, templates, and scripts listed here. Do not assume a particular profile system, task orchestrator, memory service, or response-handoff format.

© magnus919, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 43 other files (scripts, references) in site-reliability-engineering of magnus919/agent-skills.

  • SKILL.md
  • README.md
  • evals/agentic-review-notes.md
  • evals/evals.json
  • references/agentic-evidence-and-evals.md
  • references/agentic-operations.md
  • references/guiding-principles.md
  • references/human-systems-and-learning.md
  • references/incident-command-system.md
  • references/monitoring-alerting.md
  • references/oncall-best-practices.md
  • references/operational-documentation.md
  • references/postmortem-culture.md
  • references/product-focused-reliability.md
  • references/release-engineering.md
  • references/reliability-design-and-change.md
  • references/senior-sre-blueprint.md
  • references/slo-implementation-recipe.md
  • references/slo-sli-framework.md
  • … and 25 more

Open the folder on GitHubat commit 22b4723

Compare with similar skills

Site Reliability Engineering next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Site Reliability Engineering compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Site Reliability Engineering this skillmagnus919/agent-skills115—~2.7kAutomated safety check: PassMIT
Inference Autopilotrednote-machine-learning/Inference-autopilot144—~4.5kAutomated safety check: PassApache-2.0
Executing Distributed System Testsshenli/distributed-system-testing231—~5.1kAutomated safety check: NotesMIT
Alerting Irmgrafana/skills2821 repos~1.9kAutomated safety check: PassApache-2.0
Slo Implementationwshobson/agents40k11 repos~1.7kAutomated safety check: PassMIT
Agentforce D360 Analyzeforcedotcom/sf-skills1.1k—~3.4kAutomated safety check: PassApache-2.0

Similar skills

  • Inference Autopilot

    rednote-machine-learning/Inference-autopilot

    Analyze, benchmark, diagnose, and optimize large-model inference deployments from hardware inventory, model details, workload traces, and latency or throughput SLOs.

    144 GitHub stars~4.5k tokensUpdated 1 mo ago
    DevOps & CloudAuto-check passed
  • Executing Distributed System Tests

    shenli/distributed-system-testing

    A skill your agent uses when running a previously designed distributed-systems test plan against a real or simulated cluster — driving fault injection, workload, chaos scenarios, linearizability /…

    231 GitHub stars~5.1k tokensUpdated 2 mo ago
    DevOps & CloudAuto-check: notes
  • Alerting Irm

    grafana/skills

    Official

    Configure Grafana Alerting, Incident Response Management (IRM), and SLOs end-to-end — provisions Grafana-managed and data-source-managed alert rules, contact points (Slack/PagerDuty/email/webhook)…

    282 GitHub starsUsed in 1 repo~1.9k tokens
    DevOps & CloudAuto-check passed
  • Slo Implementation

    wshobson/agents

    Define and implement Service Level Indicators (SLIs) and Service Level Objectives (SLOs) with error budgets and alerting.

    40k GitHub starsUsed in 11 repos~1.7k tokens
    DevOps & CloudAuto-check passed
  • Agentforce D360 Analyze

    forcedotcom/sf-skills

    Data Cloud 360° view of a single Agentforce session. An agent skill from forcedotcom/sf-skills.

    1.1k GitHub stars~3.4k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Promql

    grafana/skills

    Official

    Write, validate, and optimize PromQL for Prometheus / Grafana Mimir / Grafana Cloud Metrics.

    282 GitHub starsUsed in 1 repo~1.1k tokens
    DevOps & CloudAuto-check passed

More from magnus919/agent-skills

All 131 skills in this repo
  • Artifact Pyramids

    magnus919/agent-skills

    Organize durable agent research outputs as summaries, analysis, and evidence dossiers.

    115 GitHub stars~2.7k tokensUpdated today
    Auto-check passed
  • Ascii City Engine

    magnus919/agent-skills

    Build portable, first-person colored ASCII city engines and small GIS-derived city packs.

    115 GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Color Management

    magnus919/agent-skills

    Manage color workflows with ICC profiles, working spaces, gamut mapping, and color science.

    115 GitHub stars~2.6k tokensUpdated today
    Auto-check: notes
  • Data Scientist

    magnus919/agent-skills

    A skill your agent uses for PhD-level expertise in data science, statistics, and machine learning: rigorous statistical analysis, experimental design, causal inference, advanced modeling, research…

    115 GitHub stars~4.1k tokensUpdated today
    Auto-check passed
  • Docker Compose

    magnus919/agent-skills

    Use Docker Compose to define, run, debug, and harden multi-container applications.

    115 GitHub stars~2k tokensUpdated today
    Auto-check: notes
  • Fpga Development

    magnus919/agent-skills

    Design, review, simulate, and verify FPGA logic using explicit RTL contracts, clock and reset models, CDC analysis, timing constraints, and reproducible implementation evidence.

    115 GitHub stars~2.7k tokensUpdated today
    Auto-check passed

Categories

Questions about Site Reliability Engineering

What does Site Reliability Engineering do?

Design, operate, and improve reliable production systems with SLOs, incident command, observability, error budgets, and operational practices. Site Reliability Engineering is an agent skill from magnus919/agent-skills. Design, operate, and improve reliable production systems with SLOs, incident command, observability, error budgets, and operational practices.

When should I use Site Reliability Engineering?

Site Reliability Engineering fits situations like: unrelated requests; route to the nearest named specialist.

How do I install Site Reliability Engineering in Claude Code?

Run `npx skills add magnus919/agent-skills --skill site-reliability-engineering -a claude-code`. Or copy the skill folder (site-reliability-engineering in magnus919/agent-skills) into .claude/skills/site-reliability-engineering in your project. Claude Code loads it when a task matches its description.

How do I install Site Reliability Engineering in Codex?

Run `npx skills add magnus919/agent-skills --skill site-reliability-engineering -a codex`. Or copy the skill folder (site-reliability-engineering in magnus919/agent-skills) into .agents/skills/site-reliability-engineering in your project. Codex loads it when a task matches its description.

Can I use Site Reliability Engineering in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add magnus919/agent-skills --skill site-reliability-engineering -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/site-reliability-engineering, .gemini/skills/site-reliability-engineering, .github/skills/site-reliability-engineering and .opencode/skills/site-reliability-engineering in your project.

What does Site Reliability Engineering need to run?

SKILL.md names no scripts, command-line tools or credentials: Site Reliability Engineering is instructions for the agent only. Our summary lists: Python 3. Compatibility (from SKILL.md): Python 3.9+ is required only for the bundled calculation and summary scripts..

Does Site Reliability Engineering access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Site Reliability Engineering safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Site Reliability Engineering use?

Site Reliability Engineering is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Site Reliability Engineering use?

About 2.7k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 149k tokens, read only when the agent opens those files.

What are the alternatives to Site Reliability Engineering?

Skills that share tags, products or a category with Site Reliability Engineering: Inference Autopilot (rednote-machine-learning/Inference-autopilot, 144 stars), Executing Distributed System Tests (shenli/distributed-system-testing, 231 stars), Alerting Irm (grafana/skills, 282 stars) and Slo Implementation (wshobson/agents, 40k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Site Reliability Engineering?

magnus919 (a GitHub user) maintains it in magnus919/agent-skills, which has 115 GitHub stars. The repository holds 131 skills in this directory. The repository was last updated on October 10, 2026.

Source: magnus919/agent-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.