Agent skill

System Design Resilience Ops

by HoangNguyen0403 in HoangNguyen0403/agent-skills-standard

Make a design survivable and operable: eliminate single points of failure, pick failover topology, set RPO/RTO, define readiness signals and observability, choose a deployment and rollback strategy.

MITAuto-check passedDevOps & Cloud

Install System Design Resilience Ops

skills CLI
$ npx skills add HoangNguyen0403/agent-skills-standard --skill system-design-resilience-ops -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install HoangNguyen0403/agent-skills-standard system-design-resilience-ops --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/HoangNguyen0403/agent-skills-standard.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/system-design/system-design-resilience-ops .claude/skills/system-design-resilience-ops && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
system-design-resilience-ops
GitHub stars
572
Token cost
~902 tokens
SKILL.md length
428 words
Files
3 (incl. references)
Skills in repo
211
Repo updated
First seen
Licence
MIT

At a glance

Make a design survivable and operable: eliminate single points of failure, pick failover topology, set RPO/RTO, define readiness signals and observability, choose a deployment and rollback strategy.

  • Reviewing availability
  • SKILL.md covers Priority: P1 (HIGH), SPOF Elimination, Failover and Recovery and Observability, plus 3 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • Deciding rollout mechanics

What it does

System Design Resilience Ops is an agent skill from HoangNguyen0403/agent-skills-standard. Make a design survivable and operable: eliminate single points of failure, pick failover topology, set RPO/RTO, define readiness signals and observability, choose a deployment and rollback strategy. Use when reviewing availability, planning DR, or deciding rollout mechanics.

Its SKILL.md is about 900 tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files, including reference files (for example `evals/evals.json` and `references/reliability-operations.md`).

It sits in DevOps & Cloud, covering Backup and disaster recovery, Deployment and Observability. The repository describes itself as: A collection of Agent Skills Standard and Best Practice for Programming Languages, Frameworks that help our AI Agent follow best practies on frameworks and programming laguages. The licence is MIT.

When your agent uses it

  • Reviewing availability
  • Deciding rollout mechanics

Example prompts

  • “/system-design-resilience-ops”

What it can do on your machine

Read from SKILL.md and the folder at commit b529c2d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

System Design Resilience Ops loads about 902 tokens when it runs, and up to ~1.8k if it reads all its reference files. Until then it costs about 76 tokens; SKILL.md has 428 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~76
When it runs · the whole SKILL.md, loaded when a task matches
~902
With references · SKILL.md plus every file in references/, read only if the agent opens them
~1.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from HoangNguyen0403/agent-skills-standard at commit b529c2d, republished under its MIT licence (© HoangNguyen0403). 428 words, ~902 tokens.

Download SKILL.mdSave it as .claude/skills/system-design-resilience-ops/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
system-design-resilience-ops
description
Make a design survivable and operable: eliminate single points of failure, pick failover topology, set RPO/RTO, define readiness signals and observability, choose a deployment and rollback strategy. Use when reviewing availability, planning DR, or deciding rollout mechanics.

Resilience and Operations

Priority: P1 (HIGH)

A design is not done until its failure and its rollout are designed.

SPOF Elimination

  • Walk every component and ask what happens when exactly one instance dies, then when the whole zone dies.
  • Any component with one instance, one writer, or one shared config plane is a single point of failure. Name it or remove it.
  • Redundancy only helps when failure modes are independent: shared credentials, shared config, and a shared control plane cancel the benefit.
  • Blast radius: state which users or flows are affected per component failure, and cap it with cells, bulkheads, or per-tenant quotas.

Failover and Recovery

TopologyRecovery timeCostFits
Single region, multi-AZMinutes, automaticLowMost products
Active-passive across regionsMinutes to hours, drill-dependentMediumRegulated or high-value flows
Active-active across regionsSecondsHighGlobal low-latency, conflict-tolerant data
  • Set RPO (tolerable data loss) and RTO (tolerable downtime) as numbers before choosing a topology; the numbers pick the topology, not the reverse.
  • Untested failover is a hypothesis. Schedule a drill and record the measured RTO against the target.
  • Backups need a restore test. A backup that has never been restored is not a backup.

Observability

  • Instrument the four signals per service: traffic, error rate, latency percentiles, saturation.
  • Alert on user-visible symptoms and on error-budget burn rate, not on raw CPU.
  • Propagate a trace and correlation id across every hop, including queue messages.
  • Every alert needs an owner, a runbook link, and a defined next action; an alert nobody acts on is noise.
Show full SKILL.md (175 more words)Show less

Rollout

StrategyBlast radiusRollbackCost
RollingGrows during the rollRoll forward or back, slowLow
Blue-greenFull switch at cutoverInstant switch backDouble capacity
CanarySmall cohort firstStop and drain the cohortNeeds routing plus metrics
Feature flagPer user or tenantInstant, no redeployFlag lifecycle debt
  • Schema and code deploy separately: expand, migrate, contract. Never ship a migration that only the new code can read.
  • Define the rollback trigger as a metric threshold and a time box before the deploy starts.

Anti-Patterns

  • No untested failover: no DR claim without a drill date and a measured RTO.
  • No unbounded retry: retries need budget, backoff with jitter, and a stop condition, or they amplify an outage.
  • No liveness probe on dependencies: a downstream outage must not restart the fleet.
  • No deploy without rollback: irreversible releases are outages waiting for a bad build.
  • No autoscaling without a floor and ceiling: unbounded scaling turns a bug into a bill.

References

© HoangNguyen0403, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (references) in skills/system-design/system-design-resilience-ops of HoangNguyen0403/agent-skills-standard.

  • SKILL.md
  • evals/evals.json
  • references/reliability-operations.md

Open the folder on GitHubat commit b529c2d

Compare with similar skills

System Design Resilience Ops next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

System Design Resilience Ops compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
System Design Resilience Ops this skillHoangNguyen0403/agent-skills-standard572—~902Automated safety check: PassMIT
Temps CLIgotempsh/temps833—~2kAutomated safety check: PassApache-2.0
Frontmcp Production Readinessagentfront/frontmcp146—~6.5kAutomated safety check: PassApache-2.0
Kubeshark Installerkubeshark/kubeshark12k—~3.6kAutomated safety check: NotesApache-2.0
KubeSphere ServiceMesh Managerkubesphere/kubesphere17k—~2.4kAutomated safety check: PassCustom licence
Tempsgotempsh/temps833—~2kAutomated safety check: PassApache-2.0

Similar skills

  • Temps CLI

    gotempsh/temps

    Operate Temps through the pinned @temps-sdk/cli package with bunx or npx.

    833 GitHub stars~2k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Pre-production audit, hardening, and go-live checklists for FrontMCP servers.

    146 GitHub stars~6.5k tokensUpdated yesterday
    DevOps & CloudAuto-check passed
  • Kubeshark Installer

    kubeshark/kubeshark

    Installs and configures Kubeshark on a Kubernetes cluster, choosing between the quick CLI path and a Helm install with custom values.

    12k GitHub stars~3.6k tokensUpdated yesterday
    DevOps & CloudAuto-check: notes
  • KubeSphere ServiceMesh Manager

    kubesphere/kubesphere

    Installs, checks and troubleshoots the KubeSphere ServiceMesh extension (Istio, Kiali, Jaeger), including grayscale release, sidecar injection, topology and tracing issues.

    17k GitHub stars~2.4k tokensUpdated 2 mo ago
    DevOps & CloudAuto-check passed
  • Temps

    gotempsh/temps

    Manage, deploy, operate, and instrument applications with Temps.

    833 GitHub stars~2k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Google Agents CLI Observability

    pifferologo/cloud-agents-cli

    This skill should be used when the user wants to "set up tracing", "monitor my ADK agent", "configure logging", "add observability", "debug production traffic", or needs guidance on monitoring…

    129 GitHub starsUsed in 1 repo~2.5k tokens
    DevOps & CloudAuto-check passed

More from HoangNguyen0403/agent-skills-standard

All 211 skills in this repo
  • Subagent-Driven Development

    HoangNguyen0403/agent-skills-standard

    Runs a multi-task implementation plan by sending each task to a fresh implementer subagent, reviewing it independently, then reviewing the whole branch.

    572 GitHub stars~1.3k tokensUpdated yesterday
    Auto-check passed
  • draw.io Architecture Diagramming

    HoangNguyen0403/agent-skills-standard

    Draws architecture diagrams as editable draw.io files from a JSON spec, with a fixed house style, one C4 level per diagram and evidence-tagged shapes.

    572 GitHub stars~1.3k tokensUpdated yesterday
    Auto-check passed
  • Android Navigation 3 Guide

    HoangNguyen0403/agent-skills-standard

    Implements and migrates to Jetpack Navigation 3 in Compose: NavDisplay, typed route objects, a state-list back stack, deep links, multiple back stacks and dialog scenes.

    572 GitHub stars~687 tokensUpdated yesterday
    Auto-check passed
  • Angular HttpClient Standards

    HoangNguyen0403/agent-skills-standard

    Sets rules for Angular HTTP code: functional interceptors, typed requests, services that own every call, and httpResource for reactive data loading in Angular 17+.

    572 GitHub stars~652 tokensUpdated yesterday
    Auto-check passed
  • Angular Tooling

    HoangNguyen0403/agent-skills-standard

    Angular CLI usage, code generation, build configuration, and bundle optimization.

    572 GitHub stars~743 tokensUpdated yesterday
    Auto-check passed
  • Common Code Review

    HoangNguyen0403/agent-skills-standard

    Conduct high-quality, persona-driven code reviews. An agent skill from HoangNguyen0403/agent-skills-standard.

    572 GitHub stars~772 tokensUpdated yesterday
    Auto-check passed

Categories

Questions about System Design Resilience Ops

What does System Design Resilience Ops do?

Make a design survivable and operable: eliminate single points of failure, pick failover topology, set RPO/RTO, define readiness signals and observability, choose a deployment and rollback strategy. System Design Resilience Ops is an agent skill from HoangNguyen0403/agent-skills-standard. Make a design survivable and operable: eliminate single points of failure, pick failover topology, set RPO/RTO, define readiness signals and observability, choose a deployment and rollback strategy.

When should I use System Design Resilience Ops?

System Design Resilience Ops fits situations like: reviewing availability; deciding rollout mechanics.

How do I install System Design Resilience Ops in Claude Code?

Run `npx skills add HoangNguyen0403/agent-skills-standard --skill system-design-resilience-ops -a claude-code`. Or copy the skill folder (skills/system-design/system-design-resilience-ops in HoangNguyen0403/agent-skills-standard) into .claude/skills/system-design-resilience-ops in your project. Claude Code loads it when a task matches its description.

How do I install System Design Resilience Ops in Codex?

Run `npx skills add HoangNguyen0403/agent-skills-standard --skill system-design-resilience-ops -a codex`. Or copy the skill folder (skills/system-design/system-design-resilience-ops in HoangNguyen0403/agent-skills-standard) into .agents/skills/system-design-resilience-ops in your project. Codex loads it when a task matches its description.

Can I use System Design Resilience Ops in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add HoangNguyen0403/agent-skills-standard --skill system-design-resilience-ops -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/system-design-resilience-ops, .gemini/skills/system-design-resilience-ops, .github/skills/system-design-resilience-ops and .opencode/skills/system-design-resilience-ops in your project.

What does System Design Resilience Ops need to run?

SKILL.md names no scripts, command-line tools or credentials: System Design Resilience Ops is instructions for the agent only.

Does System Design Resilience Ops access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is System Design Resilience Ops safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does System Design Resilience Ops use?

System Design Resilience Ops is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does System Design Resilience Ops use?

About 902 tokens (SKILL.md is roughly 3.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 862 tokens, read only when the agent opens those files.

What are the alternatives to System Design Resilience Ops?

Skills that share tags, products or a category with System Design Resilience Ops: Temps CLI (gotempsh/temps, 833 stars), Frontmcp Production Readiness (agentfront/frontmcp, 146 stars), Kubeshark Installer (kubeshark/kubeshark, 12k stars) and KubeSphere ServiceMesh Manager (kubesphere/kubesphere, 17k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains System Design Resilience Ops?

HoangNguyen0403 (a GitHub user) maintains it in HoangNguyen0403/agent-skills-standard, which has 572 GitHub stars. The repository holds 211 skills in this directory. The repository was last updated on October 9, 2026.

Source: HoangNguyen0403/agent-skills-standard on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.