[production-grade internal] Makes systems reliable in production — SLOs, monitoring, alerting, chaos engineering, incident runbooks, capacity planning.

No licenceAuto-check passedDevOps & Cloud

Install Sre

skills CLI
$ npx skills add nagisanzenin/production-grade --skill sre -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install nagisanzenin/production-grade sre --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/nagisanzenin/production-grade.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/sre .claude/skills/sre && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
sre
GitHub stars
181
Token cost
~3.2k tokens
SKILL.md length
1,170 words
Files
6
Skills in repo
13
Repo updated
First seen
Licence
None found

At a glance

[production-grade internal] Makes systems reliable in production — SLOs, monitoring, alerting, chaos engineering, incident runbooks, capacity planning.

  • Works in 3 steps: Phase 1: Readiness Review (sequential —… → Phase 2: SLO Definition (sequential —… → Phases 3-5: Chaos + Incidents + Capacity…
  • Tasks that involve Site reliability engineering
  • SKILL.md covers Preprocessing, Brownfield Awareness, Engagement Mode and Progress Output, plus 11 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Sre is an agent skill from nagisanzenin/production-grade. [production-grade internal] Makes systems reliable in production — SLOs, monitoring, alerting, chaos engineering, incident runbooks, capacity planning. Routed via the production-grade orchestrator.

Its SKILL.md is about 3.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files (for example `phases/01-readiness-review.md`, `phases/02-slo-definition.md` and `phases/03-chaos-engineering.md`).

It sits in DevOps & Cloud, covering Site reliability engineering. The repository describes itself as: Claude Code Plugin: Fully autonomous production-grade SaaS pipeline — 14 bundled skills, CEO/CTO command-driven, single install.

When your agent uses it

  • Tasks that involve Site reliability engineering

Example prompts

  • “/sre”

Requirements

  • Python 3

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. Phase 1: Readiness Review (sequential — foundational assessment)
  2. Phase 2: SLO Definition (sequential — all other phases reference SLOs)
  3. Phases 3-5: Chaos + Incidents + Capacity (PARALLEL)

What it can do on your machine

Read from SKILL.md and the folder at commit 4b2f13f. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Sre loads about 3.2k tokens when it runs. Until then it costs about 50 tokens; SKILL.md has 1,170 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~50
When it runs · the whole SKILL.md, loaded when a task matches
~3.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

Without a licence we can't republish the file, so here is its outline and opening line. It has 1,170 words (~3,206 tokens).

“!cat Claude-Production-Grade-Suite/.protocols/ux-protocol.md 2>/dev/null || true !cat Claude-Production-Grade-Suite/.protocols/input-validation.md 2>/dev/null || true !cat Claude-Production-Grade-Suite/.protocols/tool-efficiency.md 2>/dev/null || true !cat Claude-Production-Grade-Suite/.protocols/visual-identity.md 2>/dev/null || true !cat Claude-Production-Grade-Suite/.protocols/freshness-protocol.md 2>/dev/null || true !cat Claude-Production-Grade-Suite/.protocols/receipt-protocol.md 2>/dev/null || true !cat Claude-Production-Grade-Suite/.protocols/boundary-safety.md 2>/dev/null || true !cat Claude-Production-Grade-Suite/.protocols/loop-protocol.md 2>/dev/null || true…”

— opening of SKILL.md by nagisanzenin
name
sre

Read the full SKILL.md on GitHub

Files

SKILL.md and 5 other files in skills/sre of nagisanzenin/production-grade.

  • SKILL.md
  • phases/01-readiness-review.md
  • phases/02-slo-definition.md
  • phases/03-chaos-engineering.md
  • phases/04-incident-management.md
  • phases/05-capacity-planning.md

Open the folder on GitHubat commit 4b2f13f

Compare with similar skills

Sre next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Sre compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Sre this skillnagisanzenin/production-grade181—~3.2kAutomated safety check: PassNone
Inference Autopilotrednote-machine-learning/Inference-autopilot144—~4.5kAutomated safety check: PassApache-2.0
Executing Distributed System Testsshenli/distributed-system-testing231—~5.1kAutomated safety check: NotesMIT
Alerting Irmgrafana/skills2821 repos~1.9kAutomated safety check: PassApache-2.0
Slo Implementationwshobson/agents40k11 repos~1.7kAutomated safety check: PassMIT
Agentforce D360 Analyzeforcedotcom/sf-skills1.1k—~3.4kAutomated safety check: PassApache-2.0

Similar skills

  • Inference Autopilot

    rednote-machine-learning/Inference-autopilot

    Analyze, benchmark, diagnose, and optimize large-model inference deployments from hardware inventory, model details, workload traces, and latency or throughput SLOs.

    144 GitHub stars~4.5k tokensUpdated 1 mo ago
    DevOps & CloudAuto-check passed
  • Executing Distributed System Tests

    shenli/distributed-system-testing

    A skill your agent uses when running a previously designed distributed-systems test plan against a real or simulated cluster — driving fault injection, workload, chaos scenarios, linearizability /…

    231 GitHub stars~5.1k tokensUpdated 2 mo ago
    DevOps & CloudAuto-check: notes
  • Alerting Irm

    grafana/skills

    Official

    Configure Grafana Alerting, Incident Response Management (IRM), and SLOs end-to-end — provisions Grafana-managed and data-source-managed alert rules, contact points (Slack/PagerDuty/email/webhook)…

    282 GitHub starsUsed in 1 repo~1.9k tokens
    DevOps & CloudAuto-check passed
  • Slo Implementation

    wshobson/agents

    Define and implement Service Level Indicators (SLIs) and Service Level Objectives (SLOs) with error budgets and alerting.

    40k GitHub starsUsed in 11 repos~1.7k tokens
    DevOps & CloudAuto-check passed
  • Agentforce D360 Analyze

    forcedotcom/sf-skills

    Data Cloud 360° view of a single Agentforce session. An agent skill from forcedotcom/sf-skills.

    1.1k GitHub stars~3.4k tokensUpdated yesterday
    DevOps & CloudAuto-check passed
  • Promql

    grafana/skills

    Official

    Write, validate, and optimize PromQL for Prometheus / Grafana Mimir / Grafana Cloud Metrics.

    282 GitHub starsUsed in 1 repo~1.1k tokens
    DevOps & CloudAuto-check passed

More from nagisanzenin/production-grade

All 13 skills in this repo
  • Production Grade

    nagisanzenin/production-grade

    A skill your agent uses when the user wants to build, create, or develop anything — websites, apps, APIs, services, platforms.

    181 GitHub stars~16k tokensUpdated 1 mo ago
    Auto-check: notes
  • Data Scientist

    nagisanzenin/production-grade

    [production-grade internal] Optimizes AI/ML/LLM usage when you need model selection, prompt engineering, cost reduction, or experiment design.

    181 GitHub stars~3k tokensUpdated 1 mo ago
    Auto-check passed
  • Product Manager

    nagisanzenin/production-grade

    [production-grade internal] Turns product ideas and business goals into formal requirements — BRD, user stories, acceptance criteria, prioritization.

    181 GitHub stars~3.7k tokensUpdated 1 mo ago
    Auto-check passed
  • Skill Maker

    nagisanzenin/production-grade

    [production-grade internal] Creates reusable Claude Code skills and plugins when you want to automate repeatable workflows into shareable tools.

    181 GitHub stars~2.2k tokensUpdated 1 mo ago
    Auto-check passed
  • Software Engineer

    nagisanzenin/production-grade

    [production-grade internal] Implements backend services, APIs, and business logic — builds features, fixes bugs, refactors code from specs.

    181 GitHub stars~3.6k tokensUpdated 1 mo ago
    Auto-check: notes
  • Technical Writer

    nagisanzenin/production-grade

    [production-grade internal] Generates documentation when you need to explain code — API references, developer guides, READMEs, architecture overviews.

    181 GitHub stars~3k tokensUpdated 1 mo ago
    Auto-check passed

Categories

Questions about Sre

What does Sre do?

[production-grade internal] Makes systems reliable in production — SLOs, monitoring, alerting, chaos engineering, incident runbooks, capacity planning. Sre is an agent skill from nagisanzenin/production-grade. [production-grade internal] Makes systems reliable in production — SLOs, monitoring, alerting, chaos engineering, incident runbooks, capacity planning.

When should I use Sre?

Sre fits situations like: tasks that involve Site reliability engineering.

How do I install Sre in Claude Code?

Run `npx skills add nagisanzenin/production-grade --skill sre -a claude-code`. Or copy the skill folder (skills/sre in nagisanzenin/production-grade) into .claude/skills/sre in your project. Claude Code loads it when a task matches its description.

How do I install Sre in Codex?

Run `npx skills add nagisanzenin/production-grade --skill sre -a codex`. Or copy the skill folder (skills/sre in nagisanzenin/production-grade) into .agents/skills/sre in your project. Codex loads it when a task matches its description.

Can I use Sre in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add nagisanzenin/production-grade --skill sre -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/sre, .gemini/skills/sre, .github/skills/sre and .opencode/skills/sre in your project.

What does Sre need to run?

SKILL.md names no scripts, command-line tools or credentials: Sre is instructions for the agent only. Our summary lists: Python 3.

Does Sre access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Sre safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Sre use?

No licence was found for Sre or its repository. Without one, default copyright applies: ask the author before reusing or redistributing it.

How many tokens does Sre use?

About 3.2k tokens (SKILL.md is roughly 13k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Sre?

Skills that share tags, products or a category with Sre: Inference Autopilot (rednote-machine-learning/Inference-autopilot, 144 stars), Executing Distributed System Tests (shenli/distributed-system-testing, 231 stars), Alerting Irm (grafana/skills, 282 stars) and Slo Implementation (wshobson/agents, 40k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Sre?

nagisanzenin (a GitHub user) maintains it in nagisanzenin/production-grade, which has 181 GitHub stars. The repository holds 13 skills in this directory. The repository was last updated on August 19, 2026.

Source: nagisanzenin/production-grade on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.