Inference Autopilot
rednote-machine-learning/Inference-autopilot
Analyze, benchmark, diagnose, and optimize large-model inference deployments from hardware inventory, model details, workload traces, and latency or throughput SLOs.
A skill your agent uses when designing, operating, reviewing, or improving production reliability with SLOs, incident command, observability, error budgets, and operational excellence.
$ npx skills add magnus919/hermes-profiles --skill site-reliability-engineering -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install magnus919/hermes-profiles site-reliability-engineering --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/magnus919/hermes-profiles.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/site-reliability-engineering .claude/skills/site-reliability-engineering && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "site-reliability-engineering" agent skill from https://github.com/magnus919/hermes-profiles/tree/main/skills/site-reliability-engineering into .claude/skills/site-reliability-engineering/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "site-reliability-engineering", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/magnus919/hermes-profiles/tree/main/skills/site-reliability-engineeringType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add magnus919/hermes-profiles --skill site-reliability-engineering -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install magnus919/hermes-profiles site-reliability-engineering --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/magnus919/hermes-profiles.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/site-reliability-engineering .agents/skills/site-reliability-engineering && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "site-reliability-engineering" agent skill from https://github.com/magnus919/hermes-profiles/tree/main/skills/site-reliability-engineering into .agents/skills/site-reliability-engineering/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "site-reliability-engineering", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add magnus919/hermes-profiles --skill site-reliability-engineering -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install magnus919/hermes-profiles site-reliability-engineering --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/magnus919/hermes-profiles.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/site-reliability-engineering .cursor/skills/site-reliability-engineering && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "site-reliability-engineering" agent skill from https://github.com/magnus919/hermes-profiles/tree/main/skills/site-reliability-engineering into .cursor/skills/site-reliability-engineering/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "site-reliability-engineering", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/magnus919/hermes-profiles.git --path skills/site-reliability-engineering--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add magnus919/hermes-profiles --skill site-reliability-engineering -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install magnus919/hermes-profiles site-reliability-engineering --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/magnus919/hermes-profiles.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/site-reliability-engineering .gemini/skills/site-reliability-engineering && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "site-reliability-engineering" agent skill from https://github.com/magnus919/hermes-profiles/tree/main/skills/site-reliability-engineering into .gemini/skills/site-reliability-engineering/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "site-reliability-engineering", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install magnus919/hermes-profiles site-reliability-engineeringInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add magnus919/hermes-profiles --skill site-reliability-engineering -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/magnus919/hermes-profiles.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/site-reliability-engineering .github/skills/site-reliability-engineering && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "site-reliability-engineering" agent skill from https://github.com/magnus919/hermes-profiles/tree/main/skills/site-reliability-engineering into .github/skills/site-reliability-engineering/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "site-reliability-engineering", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add magnus919/hermes-profiles --skill site-reliability-engineering -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install magnus919/hermes-profiles site-reliability-engineering --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/magnus919/hermes-profiles.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/site-reliability-engineering .opencode/skills/site-reliability-engineering && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "site-reliability-engineering" agent skill from https://github.com/magnus919/hermes-profiles/tree/main/skills/site-reliability-engineering into .opencode/skills/site-reliability-engineering/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "site-reliability-engineering", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
site-reliability-engineeringA skill your agent uses when designing, operating, reviewing, or improving production reliability with SLOs, incident command, observability, error budgets, and operational excellence.
Site Reliability Engineering is an agent skill from magnus919/hermes-profiles. Use when designing, operating, reviewing, or improving production reliability with SLOs, incident command, observability, error budgets, and operational excellence.
Its SKILL.md is about 1.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 27 other files, including scripts and reference files (for example `references/guiding-principles.md`, `references/incident-command-system.md` and `references/monitoring-alerting.md`).
It sits in DevOps & Cloud, covering Site reliability engineering. The repository describes itself as: Curated Hermes Agent profiles for specialist swarms — opinionated, Hermes-optimized, artifact-pyramid native. The licence is MIT.
2 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 867a555. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 1 file in scripts/ (Python, from the files we listed), which the agent can run.
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Site Reliability Engineering loads about 1.3k tokens when it runs, and up to ~137k if it reads all its reference files. Until then it costs about 48 tokens; SKILL.md has 439 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from magnus919/hermes-profiles at commit 867a555, republished under its MIT licence (© magnus919). 439 words, ~1,283 tokens.
.claude/skills/site-reliability-engineering/SKILL.md (or your agent's skills folder). This skill also uses 24 other files; get the full folder from GitHub.A comprehensive methodology for designing, operating, and improving reliable production systems. Rooted in Google SRE principles and extended with modern practices for incident command, observability engineering, error budget governance, and operational excellence.
| Trigger | What It Means |
|---|---|
| "Design reliability into this system" | SLO/SLI framework, error budget policy, resilience architecture |
| "Run an incident postmortem" | Blameless postmortem with timeline, 5 Whys, action tracking |
| "Improve our on-call" | Rotation design, alert tuning, toil reduction, escalation policy |
| "Build observability" | The Four Golden Signals, dashboard design, alert rule patterns |
| "Do a reliability review" | Architecture review against SRE principles, risk assessment |
| "I need an incident commander" | Incident command framework, role cards, communication templates |
| "Automate this operational task" | Toil assessment, automation decision tree, runbook pattern |
skill_view('site-reliability-engineering') — this file, methodology index| Topic | File | When to Load |
|---|---|---|
| SRE Book Chapter Summaries | references/sre-book-chapters.md | Design engagement, first principles review |
| SLO/SLI Framework | references/slo-sli-framework.md | Defining reliability targets |
| Error Budget Governance | references/error-budget-governance.md | Policy design, burn rate alerts |
| Incident Command System | references/incident-command-system.md | During/after incident, training |
| Blameless Postmortems | references/postmortem-culture.md | After incident, process design |
| Monitoring & Alerting | references/monitoring-alerting.md | Observability design, alert rules |
| On-Call Best Practices | references/oncall-best-practices.md | Rotation design, team sizing |
| Toil Elimination | references/toil-elimination.md | Automation prioritization, ops review |
| Release Engineering | references/release-engineering.md | Deployment pipeline design |
| Effective Troubleshooting | references/troubleshooting.md | Debugging methodology |
| Senior SRE Role Blueprint | references/senior-sre-blueprint.md | Role definition, KPI framework |
| SRE Communication Guide | references/sre-communication-guide.md | Stakeholder updates, incident communication |
| Guiding Principles | references/guiding-principles.md | First principles, philosophy |
| Product-Focused Reliability | references/product-focused-reliability.md | Product-centric SRE, CUJ-based SLOs, JTBD model |
| Twenty Years of Lessons | references/twenty-years-lessons.md | Incident-derived tactical lessons, Prodverbs |
| SRE Ecosystem Guide | references/sre-ecosystem-guide.md | Curated guide to all SRE resources (Workbook, Secure Systems, Classroom, Prodcast, STPA, Video Gallery, Mobaa, fundamentals, AI ops) |
| Template | File | Purpose |
|---|---|---|
| Incident Commander Checklist | templates/incident-command-checklist.md | Step-by-step IC response |
| Postmortem Template | templates/postmortem-template.md | Blameless postmortem document |
| Runbook Template | templates/runbook-template.md | Operational runbook standard |
| SLO Declaration Template | templates/slo-declaration-template.md | Service-level objective specification |
| Error Budget Policy | templates/error-budget-policy.md | Team-level error budget governance |
| On-Call Rotation Template | templates/oncall-rotation.md | Rotation schedule and escalation |
| Service Review Checklist | templates/service-review-checklist.md | Pre-launch reliability review |
| Incident Communication Template | templates/incident-communication.md | Status updates during incidents |
| Script | Purpose |
|---|---|
scripts/slo-burn-rate.py | Calculate error budget burn rate from SLI data |
scripts/postmortem-summary.py | Generate a postmortem summary from structured data |
All output follows the artifact pyramid convention (three-layer progressive disclosure). The response to any caller is the absolute path to 00-index.md. Not a summary. Not a conversation. A path.
This skill is primarily used by the site-reliability-engineer profile. It works alongside:
© magnus919, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 24 other files (scripts, references) in skills/site-reliability-engineering of magnus919/hermes-profiles.
Open the folder on GitHubat commit 867a555
Site Reliability Engineering next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Site Reliability Engineering this skillmagnus919/hermes-profiles | 289 | — | ~1.3k | Automated safety check: Pass | MIT | |
| Inference Autopilotrednote-machine-learning/Inference-autopilot | 144 | — | ~4.5k | Automated safety check: Pass | Apache-2.0 | |
| Executing Distributed System Testsshenli/distributed-system-testing | 231 | — | ~5.1k | Automated safety check: Notes | MIT | |
| Alerting Irmgrafana/skills | 282 | 1 repos | ~1.9k | Automated safety check: Pass | Apache-2.0 | |
| Slo Implementationwshobson/agents | 40k | 11 repos | ~1.7k | Automated safety check: Pass | MIT | |
| Agentforce D360 Analyzeforcedotcom/sf-skills | 1.1k | — | ~3.4k | Automated safety check: Pass | Apache-2.0 |
rednote-machine-learning/Inference-autopilot
Analyze, benchmark, diagnose, and optimize large-model inference deployments from hardware inventory, model details, workload traces, and latency or throughput SLOs.
shenli/distributed-system-testing
A skill your agent uses when running a previously designed distributed-systems test plan against a real or simulated cluster — driving fault injection, workload, chaos scenarios, linearizability /…
grafana/skills
Configure Grafana Alerting, Incident Response Management (IRM), and SLOs end-to-end — provisions Grafana-managed and data-source-managed alert rules, contact points (Slack/PagerDuty/email/webhook)…
wshobson/agents
Define and implement Service Level Indicators (SLIs) and Service Level Objectives (SLOs) with error budgets and alerting.
forcedotcom/sf-skills
Data Cloud 360° view of a single Agentforce session. An agent skill from forcedotcom/sf-skills.
grafana/skills
Write, validate, and optimize PromQL for Prometheus / Grafana Mimir / Grafana Cloud Metrics.
magnus919/hermes-profiles
PhD-level expertise in data science, statistics, and machine learning.
magnus919/hermes-profiles
Create comprehensive brand identity documentation for any brand.
magnus919/hermes-profiles
Specification authoring for AI-native SDD — writes formal specifications in structured formats (Gherkin, user stories, acceptance criteria), enforces spec quality gates, and produces…
magnus919/hermes-profiles
SDD acceptance criteria verification — maps specification acceptance criteria to tests, validates implementation output against spec requirements, and produces artifact-pyramid-compliant…
magnus919/hermes-profiles
SDD work decomposition — translates formal specifications into dependency-aware task plans with per-task acceptance criteria.
magnus919/hermes-profiles
Progressive disclosure for what AI agents produce. An agent skill from magnus919/hermes-profiles.
Categories
A skill your agent uses when designing, operating, reviewing, or improving production reliability with SLOs, incident command, observability, error budgets, and operational excellence. Site Reliability Engineering is an agent skill from magnus919/hermes-profiles. Use when designing, operating, reviewing, or improving production reliability with SLOs, incident command, observability, error budgets, and operational excellence.
Site Reliability Engineering fits situations like: improving production reliability with SLOs; incident command; operational excellence.
Run `npx skills add magnus919/hermes-profiles --skill site-reliability-engineering -a claude-code`. Or copy the skill folder (skills/site-reliability-engineering in magnus919/hermes-profiles) into .claude/skills/site-reliability-engineering in your project. Claude Code loads it when a task matches its description.
Run `npx skills add magnus919/hermes-profiles --skill site-reliability-engineering -a codex`. Or copy the skill folder (skills/site-reliability-engineering in magnus919/hermes-profiles) into .agents/skills/site-reliability-engineering in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add magnus919/hermes-profiles --skill site-reliability-engineering -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/site-reliability-engineering, .gemini/skills/site-reliability-engineering, .github/skills/site-reliability-engineering and .opencode/skills/site-reliability-engineering in your project.
Going by SKILL.md and its folder, Site Reliability Engineering needs Python for the scripts in its folder. Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Site Reliability Engineering is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.3k tokens (SKILL.md is roughly 5.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 136k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Site Reliability Engineering: Inference Autopilot (rednote-machine-learning/Inference-autopilot, 144 stars), Executing Distributed System Tests (shenli/distributed-system-testing, 231 stars), Alerting Irm (grafana/skills, 282 stars) and Slo Implementation (wshobson/agents, 40k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
magnus919 (a GitHub user) maintains it in magnus919/hermes-profiles, which has 289 GitHub stars. The repository holds 31 skills in this directory. The repository was last updated on June 27, 2026.
Source: magnus919/hermes-profiles on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.