Agent skill

Site Reliability Engineering

by magnus919 in magnus919/hermes-profiles

A skill your agent uses when designing, operating, reviewing, or improving production reliability with SLOs, incident command, observability, error budgets, and operational excellence.

MITAuto-check passedDevOps & Cloud

Install Site Reliability Engineering

skills CLI
$ npx skills add magnus919/hermes-profiles --skill site-reliability-engineering -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install magnus919/hermes-profiles site-reliability-engineering --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/magnus919/hermes-profiles.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/site-reliability-engineering .claude/skills/site-reliability-engineering && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
site-reliability-engineering
GitHub stars
289
Token cost
~1.3k tokens
SKILL.md length
439 words
Files
25 (incl. scripts, references)
Skills in repo
31
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when designing, operating, reviewing, or improving production reliability with SLOs, incident command, observability, error budgets, and operational excellence.

  • Works in 2 steps: skill_view('site-reliability-engineering'… → Load references on demand by topic (see…
  • Improving production reliability with SLOs
  • SKILL.md covers When to Load This Skill, Loading Order, Reference Files and Templates, plus 3 more sections
  • Runs Python scripts from its folder

What it does

Site Reliability Engineering is an agent skill from magnus919/hermes-profiles. Use when designing, operating, reviewing, or improving production reliability with SLOs, incident command, observability, error budgets, and operational excellence.

Its SKILL.md is about 1.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 27 other files, including scripts and reference files (for example `references/guiding-principles.md`, `references/incident-command-system.md` and `references/monitoring-alerting.md`).

It sits in DevOps & Cloud, covering Site reliability engineering. The repository describes itself as: Curated Hermes Agent profiles for specialist swarms — opinionated, Hermes-optimized, artifact-pyramid native. The licence is MIT.

When your agent uses it

  • Improving production reliability with SLOs
  • Incident command
  • Operational excellence

Example prompts

  • “/site-reliability-engineering”

Requirements

  • Python 3

Workflow steps

2 steps, taken from the first numbered list in SKILL.md.

  1. skill_view('site-reliability-engineering') — this file, methodology index
  2. Load references on demand by topic (see below)

What it can do on your machine

Read from SKILL.md and the folder at commit 867a555. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python, from the files we listed), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Site Reliability Engineering loads about 1.3k tokens when it runs, and up to ~137k if it reads all its reference files. Until then it costs about 48 tokens; SKILL.md has 439 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~48
When it runs · the whole SKILL.md, loaded when a task matches
~1.3k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~137k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from magnus919/hermes-profiles at commit 867a555, republished under its MIT licence (© magnus919). 439 words, ~1,283 tokens.

Download SKILL.mdSave it as .claude/skills/site-reliability-engineering/SKILL.md (or your agent's skills folder). This skill also uses 24 other files; get the full folder from GitHub.
name
site-reliability-engineering
description
Use when designing, operating, reviewing, or improving production reliability with SLOs, incident command, observability, error budgets, and operational excellence.
title
Site Reliability Engineering — Methodology
type
skill
subjects
SRE, Reliability Engineering, Incident Management, Observability, Platform Engineering

Site Reliability Engineering

A comprehensive methodology for designing, operating, and improving reliable production systems. Rooted in Google SRE principles and extended with modern practices for incident command, observability engineering, error budget governance, and operational excellence.

When to Load This Skill

TriggerWhat It Means
"Design reliability into this system"SLO/SLI framework, error budget policy, resilience architecture
"Run an incident postmortem"Blameless postmortem with timeline, 5 Whys, action tracking
"Improve our on-call"Rotation design, alert tuning, toil reduction, escalation policy
"Build observability"The Four Golden Signals, dashboard design, alert rule patterns
"Do a reliability review"Architecture review against SRE principles, risk assessment
"I need an incident commander"Incident command framework, role cards, communication templates
"Automate this operational task"Toil assessment, automation decision tree, runbook pattern

Loading Order

  1. skill_view('site-reliability-engineering') — this file, methodology index
  2. Load references on demand by topic (see below)

Reference Files

TopicFileWhen to Load
SRE Book Chapter Summariesreferences/sre-book-chapters.mdDesign engagement, first principles review
SLO/SLI Frameworkreferences/slo-sli-framework.mdDefining reliability targets
Error Budget Governancereferences/error-budget-governance.mdPolicy design, burn rate alerts
Incident Command Systemreferences/incident-command-system.mdDuring/after incident, training
Blameless Postmortemsreferences/postmortem-culture.mdAfter incident, process design
Monitoring & Alertingreferences/monitoring-alerting.mdObservability design, alert rules
On-Call Best Practicesreferences/oncall-best-practices.mdRotation design, team sizing
Toil Eliminationreferences/toil-elimination.mdAutomation prioritization, ops review
Release Engineeringreferences/release-engineering.mdDeployment pipeline design
Effective Troubleshootingreferences/troubleshooting.mdDebugging methodology
Senior SRE Role Blueprintreferences/senior-sre-blueprint.mdRole definition, KPI framework
SRE Communication Guidereferences/sre-communication-guide.mdStakeholder updates, incident communication
Guiding Principlesreferences/guiding-principles.mdFirst principles, philosophy
Product-Focused Reliabilityreferences/product-focused-reliability.mdProduct-centric SRE, CUJ-based SLOs, JTBD model
Twenty Years of Lessonsreferences/twenty-years-lessons.mdIncident-derived tactical lessons, Prodverbs
SRE Ecosystem Guidereferences/sre-ecosystem-guide.mdCurated guide to all SRE resources (Workbook, Secure Systems, Classroom, Prodcast, STPA, Video Gallery, Mobaa, fundamentals, AI ops)
Show full SKILL.md (156 more words)Show less

Templates

TemplateFilePurpose
Incident Commander Checklisttemplates/incident-command-checklist.mdStep-by-step IC response
Postmortem Templatetemplates/postmortem-template.mdBlameless postmortem document
Runbook Templatetemplates/runbook-template.mdOperational runbook standard
SLO Declaration Templatetemplates/slo-declaration-template.mdService-level objective specification
Error Budget Policytemplates/error-budget-policy.mdTeam-level error budget governance
On-Call Rotation Templatetemplates/oncall-rotation.mdRotation schedule and escalation
Service Review Checklisttemplates/service-review-checklist.mdPre-launch reliability review
Incident Communication Templatetemplates/incident-communication.mdStatus updates during incidents

Scripts

ScriptPurpose
scripts/slo-burn-rate.pyCalculate error budget burn rate from SLI data
scripts/postmortem-summary.pyGenerate a postmortem summary from structured data

Output Contract

All output follows the artifact pyramid convention (three-layer progressive disclosure). The response to any caller is the absolute path to 00-index.md. Not a summary. Not a conversation. A path.

This skill is primarily used by the site-reliability-engineer profile. It works alongside:

  • technical-architect — receives reliability constraints, provides architecture context
  • data-architect — coordinates on observability pipeline and data durability
  • orchestrator — routes incident response tasks and reliability initiatives
  • implementation-planner — consumes reliability requirements for build plans

© magnus919, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 24 other files (scripts, references) in skills/site-reliability-engineering of magnus919/hermes-profiles.

  • SKILL.md
  • references/guiding-principles.md
  • references/incident-command-system.md
  • references/monitoring-alerting.md
  • references/oncall-best-practices.md
  • references/postmortem-culture.md
  • references/product-focused-reliability.md
  • references/release-engineering.md
  • references/senior-sre-blueprint.md
  • references/slo-sli-framework.md
  • references/sre-book-chapters.md
  • references/sre-communication-guide.md
  • references/sre-ecosystem-guide.md
  • references/toil-elimination.md
  • references/troubleshooting.md
  • references/twenty-years-lessons.md
  • scripts/slo-burn-rate.py
  • templates/error-budget-policy.md
  • … and 7 more

Open the folder on GitHubat commit 867a555

Compare with similar skills

Site Reliability Engineering next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Site Reliability Engineering compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Site Reliability Engineering this skillmagnus919/hermes-profiles289—~1.3kAutomated safety check: PassMIT
Inference Autopilotrednote-machine-learning/Inference-autopilot144—~4.5kAutomated safety check: PassApache-2.0
Executing Distributed System Testsshenli/distributed-system-testing231—~5.1kAutomated safety check: NotesMIT
Alerting Irmgrafana/skills2821 repos~1.9kAutomated safety check: PassApache-2.0
Slo Implementationwshobson/agents40k11 repos~1.7kAutomated safety check: PassMIT
Agentforce D360 Analyzeforcedotcom/sf-skills1.1k—~3.4kAutomated safety check: PassApache-2.0

Similar skills

  • Inference Autopilot

    rednote-machine-learning/Inference-autopilot

    Analyze, benchmark, diagnose, and optimize large-model inference deployments from hardware inventory, model details, workload traces, and latency or throughput SLOs.

    144 GitHub stars~4.5k tokensUpdated 1 mo ago
    DevOps & CloudAuto-check passed
  • Executing Distributed System Tests

    shenli/distributed-system-testing

    A skill your agent uses when running a previously designed distributed-systems test plan against a real or simulated cluster — driving fault injection, workload, chaos scenarios, linearizability /…

    231 GitHub stars~5.1k tokensUpdated 2 mo ago
    DevOps & CloudAuto-check: notes
  • Alerting Irm

    grafana/skills

    Official

    Configure Grafana Alerting, Incident Response Management (IRM), and SLOs end-to-end — provisions Grafana-managed and data-source-managed alert rules, contact points (Slack/PagerDuty/email/webhook)…

    282 GitHub starsUsed in 1 repo~1.9k tokens
    DevOps & CloudAuto-check passed
  • Slo Implementation

    wshobson/agents

    Define and implement Service Level Indicators (SLIs) and Service Level Objectives (SLOs) with error budgets and alerting.

    40k GitHub starsUsed in 11 repos~1.7k tokens
    DevOps & CloudAuto-check passed
  • Agentforce D360 Analyze

    forcedotcom/sf-skills

    Data Cloud 360° view of a single Agentforce session. An agent skill from forcedotcom/sf-skills.

    1.1k GitHub stars~3.4k tokensUpdated yesterday
    DevOps & CloudAuto-check passed
  • Promql

    grafana/skills

    Official

    Write, validate, and optimize PromQL for Prometheus / Grafana Mimir / Grafana Cloud Metrics.

    282 GitHub starsUsed in 1 repo~1.1k tokens
    DevOps & CloudAuto-check passed

More from magnus919/hermes-profiles

All 31 skills in this repo
  • Data Scientist

    magnus919/hermes-profiles

    PhD-level expertise in data science, statistics, and machine learning.

    289 GitHub stars~3.3k tokensUpdated 3 mo ago
    Auto-check passed
  • Brand Designer

    magnus919/hermes-profiles

    Create comprehensive brand identity documentation for any brand.

    289 GitHub stars~2.3k tokensUpdated 3 mo ago
    Auto-check passed
  • Sdd Authoring

    magnus919/hermes-profiles

    Specification authoring for AI-native SDD — writes formal specifications in structured formats (Gherkin, user stories, acceptance criteria), enforces spec quality gates, and produces…

    289 GitHub stars~1.4k tokensUpdated 3 mo ago
    Auto-check passed
  • Sdd Verification

    magnus919/hermes-profiles

    SDD acceptance criteria verification — maps specification acceptance criteria to tests, validates implementation output against spec requirements, and produces artifact-pyramid-compliant…

    289 GitHub stars~1.8k tokensUpdated 3 mo ago
    Auto-check passed
  • Sdd Work Decomposition

    magnus919/hermes-profiles

    SDD work decomposition — translates formal specifications into dependency-aware task plans with per-task acceptance criteria.

    289 GitHub stars~1.4k tokensUpdated 3 mo ago
    Auto-check passed
  • Artifact Pyramids

    magnus919/hermes-profiles

    Progressive disclosure for what AI agents produce. An agent skill from magnus919/hermes-profiles.

    289 GitHub stars~2.5k tokensUpdated 3 mo ago
    Auto-check passed

Categories

Questions about Site Reliability Engineering

What does Site Reliability Engineering do?

A skill your agent uses when designing, operating, reviewing, or improving production reliability with SLOs, incident command, observability, error budgets, and operational excellence. Site Reliability Engineering is an agent skill from magnus919/hermes-profiles. Use when designing, operating, reviewing, or improving production reliability with SLOs, incident command, observability, error budgets, and operational excellence.

When should I use Site Reliability Engineering?

Site Reliability Engineering fits situations like: improving production reliability with SLOs; incident command; operational excellence.

How do I install Site Reliability Engineering in Claude Code?

Run `npx skills add magnus919/hermes-profiles --skill site-reliability-engineering -a claude-code`. Or copy the skill folder (skills/site-reliability-engineering in magnus919/hermes-profiles) into .claude/skills/site-reliability-engineering in your project. Claude Code loads it when a task matches its description.

How do I install Site Reliability Engineering in Codex?

Run `npx skills add magnus919/hermes-profiles --skill site-reliability-engineering -a codex`. Or copy the skill folder (skills/site-reliability-engineering in magnus919/hermes-profiles) into .agents/skills/site-reliability-engineering in your project. Codex loads it when a task matches its description.

Can I use Site Reliability Engineering in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add magnus919/hermes-profiles --skill site-reliability-engineering -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/site-reliability-engineering, .gemini/skills/site-reliability-engineering, .github/skills/site-reliability-engineering and .opencode/skills/site-reliability-engineering in your project.

What does Site Reliability Engineering need to run?

Going by SKILL.md and its folder, Site Reliability Engineering needs Python for the scripts in its folder. Our summary lists: Python 3.

Does Site Reliability Engineering access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Site Reliability Engineering safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Site Reliability Engineering use?

Site Reliability Engineering is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Site Reliability Engineering use?

About 1.3k tokens (SKILL.md is roughly 5.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 136k tokens, read only when the agent opens those files.

What are the alternatives to Site Reliability Engineering?

Skills that share tags, products or a category with Site Reliability Engineering: Inference Autopilot (rednote-machine-learning/Inference-autopilot, 144 stars), Executing Distributed System Tests (shenli/distributed-system-testing, 231 stars), Alerting Irm (grafana/skills, 282 stars) and Slo Implementation (wshobson/agents, 40k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Site Reliability Engineering?

magnus919 (a GitHub user) maintains it in magnus919/hermes-profiles, which has 289 GitHub stars. The repository holds 31 skills in this directory. The repository was last updated on June 27, 2026.

Source: magnus919/hermes-profiles on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.