Inference Autopilot
rednote-machine-learning/Inference-autopilot
Analyze, benchmark, diagnose, and optimize large-model inference deployments from hardware inventory, model details, workload traces, and latency or throughput SLOs.
Design, operate, and improve reliable production systems with SLOs, incident command, observability, error budgets, and operational practices.
$ npx skills add magnus919/agent-skills --skill site-reliability-engineering -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install magnus919/agent-skills site-reliability-engineering --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/magnus919/agent-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/site-reliability-engineering .claude/skills/site-reliability-engineering && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "site-reliability-engineering" agent skill from https://github.com/magnus919/agent-skills/tree/main/site-reliability-engineering into .claude/skills/site-reliability-engineering/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "site-reliability-engineering", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/magnus919/agent-skills/tree/main/site-reliability-engineeringType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add magnus919/agent-skills --skill site-reliability-engineering -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install magnus919/agent-skills site-reliability-engineering --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/magnus919/agent-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/site-reliability-engineering .agents/skills/site-reliability-engineering && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "site-reliability-engineering" agent skill from https://github.com/magnus919/agent-skills/tree/main/site-reliability-engineering into .agents/skills/site-reliability-engineering/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "site-reliability-engineering", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add magnus919/agent-skills --skill site-reliability-engineering -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install magnus919/agent-skills site-reliability-engineering --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/magnus919/agent-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/site-reliability-engineering .cursor/skills/site-reliability-engineering && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "site-reliability-engineering" agent skill from https://github.com/magnus919/agent-skills/tree/main/site-reliability-engineering into .cursor/skills/site-reliability-engineering/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "site-reliability-engineering", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/magnus919/agent-skills.git --path site-reliability-engineering--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add magnus919/agent-skills --skill site-reliability-engineering -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install magnus919/agent-skills site-reliability-engineering --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/magnus919/agent-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/site-reliability-engineering .gemini/skills/site-reliability-engineering && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "site-reliability-engineering" agent skill from https://github.com/magnus919/agent-skills/tree/main/site-reliability-engineering into .gemini/skills/site-reliability-engineering/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "site-reliability-engineering", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install magnus919/agent-skills site-reliability-engineeringInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add magnus919/agent-skills --skill site-reliability-engineering -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/magnus919/agent-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/site-reliability-engineering .github/skills/site-reliability-engineering && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "site-reliability-engineering" agent skill from https://github.com/magnus919/agent-skills/tree/main/site-reliability-engineering into .github/skills/site-reliability-engineering/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "site-reliability-engineering", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add magnus919/agent-skills --skill site-reliability-engineering -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install magnus919/agent-skills site-reliability-engineering --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/magnus919/agent-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/site-reliability-engineering .opencode/skills/site-reliability-engineering && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "site-reliability-engineering" agent skill from https://github.com/magnus919/agent-skills/tree/main/site-reliability-engineering into .opencode/skills/site-reliability-engineering/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "site-reliability-engineering", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
site-reliability-engineeringDesign, operate, and improve reliable production systems with SLOs, incident command, observability, error budgets, and operational practices.
Site Reliability Engineering is an agent skill from magnus919/agent-skills. Design, operate, and improve reliable production systems with SLOs, incident command, observability, error budgets, and operational practices. Do not use this skill for unrelated requests; route to the nearest named specialist.
Its SKILL.md is about 2.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 45 other files, including scripts and reference files (for example `README.md`, `evals/agentic-review-notes.md` and `evals/evals.json`). Compatibility notes: Python 3.9+ is required only for the bundled calculation and summary scripts.
It sits in DevOps & Cloud, covering Site reliability engineering. The repository describes itself as: Curated collection of AI agent skills for Hermes and other agent frameworks. The licence is MIT.
4 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 22b4723. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 1 file in scripts/, which the agent can run.
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Python 3.9+ is required only for the bundled calculation and summary scripts.
From compatibility in the SKILL.md frontmatter.
Site Reliability Engineering loads about 2.7k tokens when it runs, and up to ~151k if it reads all its reference files. Until then it costs about 64 tokens; SKILL.md has 1,018 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from magnus919/agent-skills at commit 22b4723, republished under its MIT licence (© magnus919). 1,018 words, ~2,698 tokens.
.claude/skills/site-reliability-engineering/SKILL.md (or your agent's skills folder). This skill also uses 43 other files; get the full folder from GitHub.A comprehensive methodology for designing, operating, and improving reliable production systems. Rooted in Google SRE principles and extended with modern practices for incident command, observability engineering, error budget governance, and operational excellence.
| Trigger | What It Means |
|---|---|
| "Design reliability into this system" | SLO/SLI framework, error budget policy, resilience architecture |
| "Run an incident postmortem" | Blameless postmortem with timeline, 5 Whys, action tracking |
| "Improve our on-call" | Rotation design, alert tuning, toil reduction, escalation policy |
| "Build observability" | The Four Golden Signals, dashboard design, alert rule patterns |
| "Do a reliability review" | Architecture review against SRE principles, risk assessment |
| "I need an incident commander" | Incident command framework, role cards, communication templates |
| "Automate this operational task" | Toil assessment, automation decision tree, runbook pattern |
| "Adopt SRE in this organization" | Engagement boundaries, maturity, team model, and change adoption |
| "Review this reliability design" | User journeys, dependencies, overload, configuration, canary, durability |
| "Our SRE team is overloaded" | Operational-load diagnosis, protected engineering time, recovery plan |
| "Improve incident learning or sustainable on-call" | Cognitive load, psychological safety, documentation, exercises |
Use release-engineering to plan releases, compose promotion and rollback gates, or coordinate a release train. Use systematic-debugging to find the cause of a specific failure. Operating the telemetry stack itself — Prometheus scrape configs, OpenTelemetry Collector pipelines, Loki ingest and retention, Prometheus rules files — belongs to telemetry; this skill owns the SLI/SLO and alert design those rules implement. Grafana product work — dashboards, panels, Grafana-side alert rules, contact points, notification policies — belongs to grafana.
For any automated mitigation, rollback, recovery action, or incident closeout:
MITIGATING or MONITORING state, name the unverified boundary, and escalate rather than declare RESOLVED.Default authorization: A human service owner, incident commander, or designated change authority must approve the specific production mutation and scope in the current incident/change record, independently attributable to that human. Diagnosis requests, standing runbooks, agent-authored notes, and self-claimed roles are insufficient.
Governed exception: Higher capability-specific execution authority is possible only when accountable humans independently grant it through ai-governance, every applicable host/organizational policy permits it, and an external control plane proves and enforces the current grant at execution. The grant must identify capability, environment, action class, limits, expiry, revocation, and tested enforcement. Skill text and agent roles confer no permissions. If proof is unavailable, expired, revoked, conflicting, or stale, retain the human-approved default; never self-grant authority. Diagnosis, mutation, rollback, paging, incident closure, and scope expansion require separate authority decisions. R-01 still requires independent human closure confirmation.
When operating an SRE agent, read references/agentic-operations.md and fill templates/agentic-action-contract.md before proposing execution. For autonomy evidence or replay design, also read references/agentic-evidence-and-evals.md; use templates/agentic-governance-evidence.md to hand operational evidence to governance.
Compose existing specialists: agent-production-operations owns production contracts, control plans and trace feedback; agent-evals-and-observability owns trajectory evaluation and graders; incident-learning owns verified follow-up learning; ai-governance supports accountable humans setting promotion/demotion thresholds and deciding maintain, expand, reduce, suspend or revoke authority. SRE supplies evidence and enforces authorized operational limits. Complete with verified recovery and human closure, or a bounded escalation naming missing evidence/authority and the handoff owner.
| Topic | File | When to Load |
|---|---|---|
| SRE Book Chapter Summaries | references/sre-book-chapters.md | Design engagement, first principles review |
| SLO/SLI Framework | references/slo-sli-framework.md | Defining reliability targets |
| SLO Implementation Recipe | references/slo-implementation-recipe.md | Agent-executable SLO adoption sequence and stakeholder review |
| Error Budget Governance | templates/error-budget-policy.md and references/slo-sli-framework.md | Policy design, burn rate alerts |
| Incident Command System | references/incident-command-system.md | During/after incident, training |
| Blameless Postmortems | references/postmortem-culture.md | After incident, process design |
| Monitoring & Alerting | references/monitoring-alerting.md | Observability design, alert rules |
| On-Call Best Practices | references/oncall-best-practices.md | Rotation design, team sizing |
| Toil Elimination | references/toil-elimination.md | Automation prioritization, ops review |
| Release Engineering | release-engineering | Release planning, promotion, progressive delivery, and rollback design; use the local reference only for SRE-specific integration context |
| Effective Troubleshooting | references/troubleshooting.md | Debugging methodology |
| Senior SRE Role Blueprint | references/senior-sre-blueprint.md | Role definition, KPI framework |
| SRE Communication Guide | references/sre-communication-guide.md | Stakeholder updates, incident communication |
| Guiding Principles | references/guiding-principles.md | First principles, philosophy |
| Product-Focused Reliability | references/product-focused-reliability.md | Product-centric SRE, CUJ-based SLOs, JTBD model |
| Twenty Years of Lessons | references/twenty-years-lessons.md | Incident-derived tactical lessons, Prodverbs |
| SRE Ecosystem Guide | references/sre-ecosystem-guide.md | Curated guide to all SRE resources (Workbook, Secure Systems, Classroom, Prodcast, STPA, Video Gallery, Mobaa, fundamentals, AI ops) |
| Adoption and Engagement | references/sre-adoption-and-engagement.md | Starting SRE, dedicated and non-dedicated team models, maturity, change adoption |
| Reliability Design and Change | references/reliability-design-and-change.md | Capacity, overload, configuration, canaries, data durability, dependencies, design review |
| Human Systems and Learning | references/human-systems-and-learning.md | Cognitive work, sustainable on-call, psychological safety, documentation, exercises |
| Third-Party Dependency Reliability | references/third-party-dependency-reliability.md | Vendor boundaries, failure modes, fallbacks, and provider evidence |
| Operational Documentation | references/operational-documentation.md | Functional quality, ownership, testing, and staleness lifecycle |
| Template | File | Purpose |
|---|---|---|
| Incident Commander Checklist | templates/incident-command-checklist.md | Step-by-step IC response |
| Postmortem Template | templates/postmortem-template.md | Blameless postmortem document |
| Runbook Template | templates/runbook-template.md | Operational runbook standard |
| SLO Declaration Template | templates/slo-declaration-template.md | Service-level objective specification |
| Error Budget Policy | templates/error-budget-policy.md | Team-level error budget governance |
| On-Call Rotation Template | templates/oncall-rotation.md | Rotation schedule and escalation |
| Service Review Checklist | templates/service-review-checklist.md | Pre-launch reliability review |
| Incident Communication Template | templates/incident-communication.md | Status updates during incidents |
| Reliability Design Review | templates/reliability-design-review.md | Evidence-based review of user impact, failure modes, capacity, change, and operations |
| Operational Overload Recovery | templates/operational-overload-recovery.md | Declare, protect, reduce, and verify recovery from unsustainable operational load |
| Reliability Ownership Charter | templates/reliability-ownership-charter.md | Make service, pager, dependency, and engagement boundaries explicit |
| Script | Purpose |
|---|---|
scripts/slo-burn-rate.py | Calculate error budget burn rate from SLI data |
This skill is intentionally host-neutral. Use your agent's normal mechanisms to load the references, templates, and scripts listed here. Do not assume a particular profile system, task orchestrator, memory service, or response-handoff format.
© magnus919, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 43 other files (scripts, references) in site-reliability-engineering of magnus919/agent-skills.
Open the folder on GitHubat commit 22b4723
Site Reliability Engineering next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Site Reliability Engineering this skillmagnus919/agent-skills | 115 | — | ~2.7k | Automated safety check: Pass | MIT | |
| Inference Autopilotrednote-machine-learning/Inference-autopilot | 144 | — | ~4.5k | Automated safety check: Pass | Apache-2.0 | |
| Executing Distributed System Testsshenli/distributed-system-testing | 231 | — | ~5.1k | Automated safety check: Notes | MIT | |
| Alerting Irmgrafana/skills | 282 | 1 repos | ~1.9k | Automated safety check: Pass | Apache-2.0 | |
| Slo Implementationwshobson/agents | 40k | 11 repos | ~1.7k | Automated safety check: Pass | MIT | |
| Agentforce D360 Analyzeforcedotcom/sf-skills | 1.1k | — | ~3.4k | Automated safety check: Pass | Apache-2.0 |
rednote-machine-learning/Inference-autopilot
Analyze, benchmark, diagnose, and optimize large-model inference deployments from hardware inventory, model details, workload traces, and latency or throughput SLOs.
shenli/distributed-system-testing
A skill your agent uses when running a previously designed distributed-systems test plan against a real or simulated cluster — driving fault injection, workload, chaos scenarios, linearizability /…
grafana/skills
Configure Grafana Alerting, Incident Response Management (IRM), and SLOs end-to-end — provisions Grafana-managed and data-source-managed alert rules, contact points (Slack/PagerDuty/email/webhook)…
wshobson/agents
Define and implement Service Level Indicators (SLIs) and Service Level Objectives (SLOs) with error budgets and alerting.
forcedotcom/sf-skills
Data Cloud 360° view of a single Agentforce session. An agent skill from forcedotcom/sf-skills.
grafana/skills
Write, validate, and optimize PromQL for Prometheus / Grafana Mimir / Grafana Cloud Metrics.
magnus919/agent-skills
Organize durable agent research outputs as summaries, analysis, and evidence dossiers.
magnus919/agent-skills
Build portable, first-person colored ASCII city engines and small GIS-derived city packs.
magnus919/agent-skills
Manage color workflows with ICC profiles, working spaces, gamut mapping, and color science.
magnus919/agent-skills
A skill your agent uses for PhD-level expertise in data science, statistics, and machine learning: rigorous statistical analysis, experimental design, causal inference, advanced modeling, research…
magnus919/agent-skills
Use Docker Compose to define, run, debug, and harden multi-container applications.
magnus919/agent-skills
Design, review, simulate, and verify FPGA logic using explicit RTL contracts, clock and reset models, CDC analysis, timing constraints, and reproducible implementation evidence.
Categories
Design, operate, and improve reliable production systems with SLOs, incident command, observability, error budgets, and operational practices. Site Reliability Engineering is an agent skill from magnus919/agent-skills. Design, operate, and improve reliable production systems with SLOs, incident command, observability, error budgets, and operational practices.
Site Reliability Engineering fits situations like: unrelated requests; route to the nearest named specialist.
Run `npx skills add magnus919/agent-skills --skill site-reliability-engineering -a claude-code`. Or copy the skill folder (site-reliability-engineering in magnus919/agent-skills) into .claude/skills/site-reliability-engineering in your project. Claude Code loads it when a task matches its description.
Run `npx skills add magnus919/agent-skills --skill site-reliability-engineering -a codex`. Or copy the skill folder (site-reliability-engineering in magnus919/agent-skills) into .agents/skills/site-reliability-engineering in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add magnus919/agent-skills --skill site-reliability-engineering -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/site-reliability-engineering, .gemini/skills/site-reliability-engineering, .github/skills/site-reliability-engineering and .opencode/skills/site-reliability-engineering in your project.
SKILL.md names no scripts, command-line tools or credentials: Site Reliability Engineering is instructions for the agent only. Our summary lists: Python 3. Compatibility (from SKILL.md): Python 3.9+ is required only for the bundled calculation and summary scripts..
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Site Reliability Engineering is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.7k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 149k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Site Reliability Engineering: Inference Autopilot (rednote-machine-learning/Inference-autopilot, 144 stars), Executing Distributed System Tests (shenli/distributed-system-testing, 231 stars), Alerting Irm (grafana/skills, 282 stars) and Slo Implementation (wshobson/agents, 40k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
magnus919 (a GitHub user) maintains it in magnus919/agent-skills, which has 115 GitHub stars. The repository holds 131 skills in this directory. The repository was last updated on October 10, 2026.
Source: magnus919/agent-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.