Agent skill

Chaos Engineering

by borghei in borghei/Claude-Skills

Chaos engineering: hypothesis-driven fault injection to surface weakness before users do.

MITAuto-check passedDevOps & Cloud

Install Chaos Engineering

skills CLI
$ npx skills add borghei/Claude-Skills --skill chaos-engineering -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install borghei/Claude-Skills chaos-engineering --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/borghei/Claude-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/engineering/chaos-engineering .claude/skills/chaos-engineering && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
chaos-engineering
GitHub stars
881
Token cost
~1.6k tokens
SKILL.md length
651 words
Files
9 (incl. scripts, references)
Skills in repo
349
Repo updated
First seen
Licence
MIT

At a glance

Chaos engineering: hypothesis-driven fault injection to surface weakness before users do.

  • Designing a chaos experiment
  • SKILL.md covers Core Capabilities, When to Use, Clarify First and Tools, plus 2 more sections
  • Runs Python scripts from its folder; calls python
  • Planning a gameday

What it does

Chaos Engineering is an agent skill from borghei/Claude-Skills. Chaos engineering: hypothesis-driven fault injection to surface weakness before users do. Use when designing a chaos experiment, planning a gameday, choosing what to inject, computing blast radius, or building a chaos maturity model.

Its SKILL.md is about 1.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 10 other files, including scripts and reference files (for example `references/chaos-principles-and-maturity.md`, `references/experiment-design-loop.md` and `references/fault-injection-catalog.md`).

It sits in DevOps & Cloud, covering Chaos engineering. The repository describes itself as: 385 AI skills, 77 expert agents, and 900 stdlib Python tools for every team: engineering, PM, marketing, C-level, compliance, business ops, research, and a LinkedIn toolkit… The licence is MIT.

When your agent uses it

  • Designing a chaos experiment
  • Planning a gameday
  • Choosing what to inject
  • Computing blast radius

Example prompts

  • “Use the chaos-engineering skill to chao engineering: hypothesis-driven fault injection to surface weakness before users do”
  • “/chaos-engineering”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit 4a698e8. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 3 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Chaos Engineering loads about 1.6k tokens when it runs, and up to ~20k if it reads all its reference files. Until then it costs about 63 tokens; SKILL.md has 651 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~63
When it runs · the whole SKILL.md, loaded when a task matches
~1.6k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~20k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from borghei/Claude-Skills at commit 4a698e8, republished under its MIT licence (© borghei). 651 words, ~1,558 tokens.

Download SKILL.mdSave it as .claude/skills/chaos-engineering/SKILL.md (or your agent's skills folder). This skill also uses 8 other files; get the full folder from GitHub.
name
chaos-engineering
description
Chaos engineering: hypothesis-driven fault injection to surface weakness before users do. Use when designing a chaos experiment, planning a gameday, choosing what to inject, computing blast radius, or building a chaos maturity model.
license
MIT + Commons Clause
metadata.version
1.1.0
metadata.author
borghei
metadata.category
engineering
metadata.domain
engineering
metadata.updated
2026-06-17
metadata.tags
chaos-engineering, resilience, fault-injection, gameday, sre, reliability, blast-radius

Chaos Engineering

End-to-end chaos engineering: experiment design, fault injection catalog, gameday execution, and the maturity model that turns one-off "let's break stuff" exercises into a reliable discipline. Provider-agnostic — works whether you use Litmus, Chaos Mesh, AWS FIS, Gremlin, ChaosToolkit, or hand-rolled scripts.

This skill answers four questions: what to inject, where to inject it, how to size the blast, and how to extract durable learning from each run.

Core Capabilities

  • Principles & maturity — the five Principles of Chaos and a four-level maturity model (L0 none → L4 always-on production chaos) with level-up criteria.
  • Experiment design loop — a nine-step loop (steady state → hypothesis → variables → blast radius → abort → run → analyze → act → document) with worked good/bad examples.
  • Fault catalog — what to inject per layer: pod/host, network, dependency, resource, state, and traffic, with tool mappings.
  • Blast-radius sizing — quantify worst-case affected users and recommended caps; experiments start tiny (1 pod / 1% / 1 min) and grow only after passing.
  • Gameday execution — scheduled multi-scenario exercises with roles, agendas, scenario selection, and debrief templates.
  • Discipline — anti-patterns to avoid, the "first five experiments" for new teams, and end-to-end workflows (single experiment, gameday, kill-switch verification, post-incident verification).

When to Use

SituationSkill applies
Spinning up a chaos program from scratchYes — start with maturity model + first 5 experiments
Designing a single experiment for a known concernYes — use the experiment design loop
Planning a gameday for a team or serviceYes — use scripts/gameday_planner.py
Validating a kill switch or fallback path actually worksYes — chaos is the way to test these in prod-like conditions
Post-incident verification: "did the fix really fix it?"Yes — re-inject the original fault, confirm the new behavior
Compliance evidence (SOC 2 A1 / DORA Art. 25)Yes — chaos runs produce auditable resilience-testing evidence
Improving SLOs / error budgetsPair with engineering/observability-designer — chaos surfaces SLO violations

Clarify First

Before designing the experiment, confirm these inputs. If any is unknown or vague, ASK — do not assume:

  • Target & fault type — the service and what to inject (dependency-timeout, network, pod-kill, resource exhaustion) (drives the experiment doc via --target/--fault)
  • Steady-state hypothesis — the metric that defines "healthy" and the expected behavior under fault (the hypothesis the experiment tests)
  • Blast radius caps — user count, % targeted, duration, and abort triggers (sizes the experiment via blast_radius_calculator.py)

Stop rule: ask only the 2-3 that most change the output. If the user says "just draft it," proceed and list your assumptions at the top of the artifact.

Show full SKILL.md (256 more words)Show less

Tools

ToolPurposeCommand
chaos_experiment_designer.pyScaffold a chaos experiment doc (hypothesis stub, steady-state template, abort criteria, checklist)python scripts/chaos_experiment_designer.py --target payments-svc --fault dependency-timeout --duration 5m
blast_radius_calculator.pyCompute worst-case affected users, recommended caps, and abort triggerspython scripts/blast_radius_calculator.py --users 100000 --percent-targeted 1 --duration 60
gameday_planner.pyGenerate a tailored gameday agenda with roles, timeline, and debrief templatepython scripts/gameday_planner.py --service search-api --duration full --scenarios region-failover,dep-outage

References

Load the reference that matches the task — keep this file lean and pull detail on demand:

  • references/experiment-design-loop.md — the five principles, the four-level maturity model, the full nine-step design loop with worked steady-state/hypothesis/abort examples, and the "first five experiments." Read when designing or running an experiment.
  • references/gameday-workflows-and-antipatterns.md — the gameday agenda and roles, scenario selection, the chaos anti-patterns, all four end-to-end workflows, and the script tooling-output table. Read when planning a gameday or running a workflow.
  • references/chaos-principles-and-maturity.md — per-level maturity scorecard, the L0→L1 "first five experiments" list, and the org/SRE prerequisites for each level. Read when assessing or leveling up a program.
  • references/gameday-playbook.md — 12 scenario templates by domain (API service, database, frontend, async pipeline, multi-region, etc.), half-day/full-day agendas, and post-gameday writeup templates. Read when choosing gameday scenarios.
  • references/fault-injection-catalog.md — fault types per layer (network, host, dependency, resource, state, traffic) with tool mappings (Chaos Mesh / Litmus / AWS FIS / Gremlin / ChaosToolkit). Read when choosing what to inject.
  • engineering/observability-designer — wire metrics needed to define steady state
  • engineering/incident-commander — chaos is rehearsal for the incident response you'll need
  • engineering/feature-flags-architect — kill switches verified by chaos; chaos verified by flags
  • ra-qm-team/dora-compliance-expert — DORA Article 25 requires resilience testing; chaos runs are evidence

© borghei, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 8 other files (scripts, references) in engineering/chaos-engineering of borghei/Claude-Skills.

  • SKILL.md
  • references/chaos-principles-and-maturity.md
  • references/experiment-design-loop.md
  • references/fault-injection-catalog.md
  • references/gameday-playbook.md
  • references/gameday-workflows-and-antipatterns.md
  • scripts/blast_radius_calculator.py
  • scripts/chaos_experiment_designer.py
  • scripts/gameday_planner.py

Open the folder on GitHubat commit 4a698e8

Compare with similar skills

Chaos Engineering next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Chaos Engineering compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Chaos Engineering this skillborghei/Claude-Skills881—~1.6kAutomated safety check: PassMIT
Executing Distributed System Testsshenli/distributed-system-testing231—~5.1kAutomated safety check: NotesMIT
Audit Reviewtestflows/TestFlows-GitHub-Hetzner-Runners102—~2.1kAutomated safety check: PassCustom licence
Chaos Dr Testharness/harness-skills115—~2.6kAutomated safety check: PassApache-2.0
Chaos Experimentharness/harness-skills115—~1.6kAutomated safety check: PassApache-2.0
Chaos EngineerJeffallan/claude-skills12k—~1.8kAutomated safety check: PassMIT

Similar skills

  • Executing Distributed System Tests

    shenli/distributed-system-testing

    A skill your agent uses when running a previously designed distributed-systems test plan against a real or simulated cluster — driving fault injection, workload, chaos scenarios, linearizability /…

    231 GitHub stars~5.1k tokensUpdated 2 mo ago
    DevOps & CloudAuto-check: notes
  • Audit Review

    testflows/TestFlows-GitHub-Hetzner-Runners

    Perform deep feature audits with transition-matrix and logical fault-injection validation.

    102 GitHub stars~2.1k tokensUpdated 17 days ago
    DevOps & CloudAuto-check passed
  • Chaos Dr Test

    harness/harness-skills

    A skill your agent uses when working with Chaos Engineering steps inside a Harness pipeline.

    115 GitHub stars~2.6k tokensUpdated 2 days ago
    DevOps & CloudAuto-check passed
  • Chaos Experiment

    harness/harness-skills

    A skill your agent uses when the user asks to create, edit, update, design, or configure a Harness Chaos Experiment — including faults, probes, actions, experiment YAML, fault injection, pod-delete…

    115 GitHub stars~1.6k tokensUpdated 2 days ago
    DevOps & CloudAuto-check passed
  • Chaos Engineer

    Jeffallan/claude-skills

    Designs chaos experiments, failure injection and game days for distributed systems, with blast radius limits, rollback plans and written learnings.

    12k GitHub stars~1.8k tokensUpdated 5 days ago
    DevOps & CloudAuto-check passed
  • SRE Engineer

    Jeffallan/claude-skills

    Defines SLIs, SLOs and error budgets, and sets up golden-signal monitoring, blameless postmortems, toil automation and chaos experiments for production systems.

    12k GitHub stars~1.7k tokensUpdated 5 days ago
    DevOps & CloudAuto-check passed

More from borghei/Claude-Skills

All 349 skills in this repo
  • Agents In The Team

    borghei/Claude-Skills

    Run delivery when AI coding and ops agents take tickets. An agent skill from borghei/Claude-Skills.

    881 GitHub stars~4.2k tokensUpdated yesterday
    Auto-check passed
  • AI Content Disclosure

    borghei/Claude-Skills

    Check AI-generated marketing content and reviews for required disclosures under the EU AI Act, FTC rules and platform AI-label policies.

    881 GitHub stars~3.4k tokensUpdated yesterday
    Auto-check passed
  • AI Prototyping

    borghei/Claude-Skills

    Idea to AI-generated prototype to customer validation to engineering handoff.

    881 GitHub stars~3.6k tokensUpdated yesterday
    Auto-check passed
  • Analytics Engineer

    borghei/Claude-Skills

    Analytics engineering across data modeling, dbt, transformation, and semantic layers.

    881 GitHub stars~3.4k tokensUpdated yesterday
    Auto-check passed
  • Ansoff Matrix

    borghei/Claude-Skills

    Ansoff Matrix — 4-quadrant framework for growth options: market penetration, market/product development, and diversification.

    881 GitHub stars~2.2k tokensUpdated yesterday
    Auto-check passed
  • Brainstorm Okrs

    borghei/Claude-Skills

    OKR brainstorming and validation using the Radical Focus framework — outcome objectives, measurable key results, counter-metrics.

    881 GitHub stars~1.4k tokensUpdated yesterday
    Auto-check passed

Categories

Questions about Chaos Engineering

What does Chaos Engineering do?

Chaos engineering: hypothesis-driven fault injection to surface weakness before users do. Chaos Engineering is an agent skill from borghei/Claude-Skills. Chaos engineering: hypothesis-driven fault injection to surface weakness before users do.

When should I use Chaos Engineering?

Chaos Engineering fits situations like: designing a chaos experiment; planning a gameday; choosing what to inject; computing blast radius.

How do I install Chaos Engineering in Claude Code?

Run `npx skills add borghei/Claude-Skills --skill chaos-engineering -a claude-code`. Or copy the skill folder (engineering/chaos-engineering in borghei/Claude-Skills) into .claude/skills/chaos-engineering in your project. Claude Code loads it when a task matches its description.

How do I install Chaos Engineering in Codex?

Run `npx skills add borghei/Claude-Skills --skill chaos-engineering -a codex`. Or copy the skill folder (engineering/chaos-engineering in borghei/Claude-Skills) into .agents/skills/chaos-engineering in your project. Codex loads it when a task matches its description.

Can I use Chaos Engineering in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add borghei/Claude-Skills --skill chaos-engineering -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/chaos-engineering, .gemini/skills/chaos-engineering, .github/skills/chaos-engineering and .opencode/skills/chaos-engineering in your project.

What does Chaos Engineering need to run?

Going by SKILL.md and its folder, Chaos Engineering needs Python for the scripts in its folder and the command-line tools its instructions call (python). Our summary lists: Python 3.

Does Chaos Engineering access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Chaos Engineering safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Chaos Engineering use?

Chaos Engineering is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Chaos Engineering use?

About 1.6k tokens (SKILL.md is roughly 6.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 18k tokens, read only when the agent opens those files.

What are the alternatives to Chaos Engineering?

Skills that share tags, products or a category with Chaos Engineering: Executing Distributed System Tests (shenli/distributed-system-testing, 231 stars), Audit Review (testflows/TestFlows-GitHub-Hetzner-Runners, 102 stars), Chaos Dr Test (harness/harness-skills, 115 stars) and Chaos Experiment (harness/harness-skills, 115 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Chaos Engineering?

borghei (a GitHub user) maintains it in borghei/Claude-Skills, which has 881 GitHub stars. The repository holds 349 skills in this directory. The repository was last updated on October 7, 2026.

Source: borghei/Claude-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.