Agent skill

Chaos Engineering

by alirezarezvani in alirezarezvani/claude-skills

A skill your agent uses when planning, running, or learning from chaos engineering experiments.

MITAuto-check passedDevOps & Cloud

Install Chaos Engineering

skills CLI
$ npx skills add alirezarezvani/claude-skills --skill chaos-engineering -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install alirezarezvani/claude-skills chaos-engineering --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/alirezarezvani/claude-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/engineering/skills/chaos-engineering .claude/skills/chaos-engineering && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
chaos-engineering
GitHub stars
28k
Token cost
~2.7k tokens
SKILL.md length
850 words
Files
10 (incl. scripts, references, assets)
Skills in repo
342
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when planning, running, or learning from chaos engineering experiments.

  • Works in 4 steps: Build a hypothesis around steady-state… → Vary real-world events. Inject realistic… → Run experiments in production. Staging… → …
  • Learning from chaos engineering experiments
  • SKILL.md covers When to use, When NOT to use, Core principle: chaos without… and Quick start, plus 10 more sections
  • Runs Python scripts from its folder; calls python

What it does

Chaos Engineering is an agent skill from alirezarezvani/claude-skills. Use when planning, running, or learning from chaos engineering experiments. Triggers on "chaos experiment", "fault injection", "gameday", "resilience test", "blast radius", "steady state", "abort criteria", "Chaos Toolkit", "Chaos Mesh", "Litmus", "Gremlin", "AWS FIS", or any deliberate failure-injection question. Ships experiment designer, blast-radius calculator, and postmortem generator (all stdlib Python), 4 references on chaos principles + experiment design + attack taxonomy + tooling landscape, and a…

Its SKILL.md is about 2.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 12 other files, including scripts, reference files and assets (for example `assets/experiment_template.md`, `assets/postmortem_template.md` and `references/attack_taxonomy.md`).

It sits in DevOps & Cloud, covering Chaos engineering. It works with Python, Amazon Web Services and Kubernetes. The repository describes itself as: 380 Claude Code skills & agent skills & plugins (30+ Agents, 70+ custom commands, 380+ skills, customizable references, scripts)for Claude Code, Codex, Gemini CLI, Cursor, and 8… The licence is MIT.

When your agent uses it

  • Learning from chaos engineering experiments
  • Chaos experiment
  • Fault injection
  • Resilience test

Example prompts

  • “chaos experiment”
  • “fault injection”
  • “gameday”
  • “/chaos-engineering”

Requirements

  • Python 3

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Build a hypothesis around steady-state behavior. Not "what breaks?" but "X holds; will it still hold under fault Y?"
  2. Vary real-world events. Inject realistic failures: kill nodes, slow networks, lose cache, throttle dependencies.
  3. Run experiments in production. Staging never has the same failure modes. Start small.
  4. Automate experiments to run continuously. One-off chaos is a press release; continuous chaos is engineering.

What it can do on your machine

Read from SKILL.md and the folder at commit 19392f7. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 3 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Chaos Engineering loads about 2.7k tokens when it runs, and up to ~8.5k if it reads all its reference files. Until then it costs about 171 tokens; SKILL.md has 850 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~171
When it runs · the whole SKILL.md, loaded when a task matches
~2.7k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~8.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from alirezarezvani/claude-skills at commit 19392f7, republished under its MIT licence (© alirezarezvani). 850 words, ~2,659 tokens.

Download SKILL.mdSave it as .claude/skills/chaos-engineering/SKILL.md (or your agent's skills folder). This skill also uses 9 other files; get the full folder from GitHub.
name
chaos-engineering
description
Use when planning, running, or learning from chaos engineering experiments. Triggers on "chaos experiment", "fault injection", "gameday", "resilience test", "blast radius", "steady state", "abort criteria", "Chaos Toolkit", "Chaos Mesh", "Litmus", "Gremlin", "AWS FIS", or any deliberate failure-injection question. Ships experiment designer, blast-radius calculator, and postmortem generator (all stdlib Python), 4 references on chaos principles + experiment design + attack taxonomy + tooling landscape, and a /chaos-experiment slash command. Composes with feature-flags-architect (kill switches as abort triggers) and kubernetes-operator (common chaos targets).
context
fork
version
2.9.0
author
claude-code-skills
license
MIT
tags
chaos-engineering, resilience, fault-injection, gameday, sre, reliability, chaos-toolkit, chaos-mesh, litmus, gremlin, aws-fis
compatible_tools
claude-code, codex-cli, cursor, antigravity, opencode, gemini-cli

Chaos Engineering

Design experiments that surface real weaknesses in production systems — without becoming outages. Most "chaos engineering" attempts skip steady-state measurement, define no abort criteria, and have no blast-radius bound. This skill enforces the discipline that makes chaos experiments safe and useful.

When to use

  • Planning a chaos experiment (what to break, where, when, how to abort)
  • Calculating blast radius before running the experiment
  • Reviewing an existing experiment plan for safety
  • Choosing a chaos tool (Chaos Toolkit / Chaos Mesh / Litmus / Gremlin / AWS FIS)
  • Writing a chaos experiment postmortem
  • Running a Game Day exercise

When NOT to use

  • General incident response (use incident-response)
  • Threat hunting / red-team (use red-team, threat-detection)
  • Performance load testing (different goal — chaos is about failure modes, not capacity)
  • Production debugging (chaos discovers weaknesses preemptively, not after-the-fact)

Core principle: chaos without abort criteria is an outage

The 4 Principles of Chaos Engineering (Netflix, 2016):

  1. Build a hypothesis around steady-state behavior. Not "what breaks?" but "X holds; will it still hold under fault Y?"
  2. Vary real-world events. Inject realistic failures: kill nodes, slow networks, lose cache, throttle dependencies.
  3. Run experiments in production. Staging never has the same failure modes. Start small.
  4. Automate experiments to run continuously. One-off chaos is a press release; continuous chaos is engineering.

Add a fifth: Define abort criteria up front. A chaos experiment with no abort criteria is an outage by another name.

Quick start

bash
SKILL=engineering/chaos-engineering/skills/chaos-engineering

# 1. Design an experiment
python "$SKILL/scripts/experiment_designer.py" --target "checkout-svc" --hypothesis "p99 latency stays <500ms" --attack latency --duration-min 15

# 2. Calculate blast radius
python "$SKILL/scripts/blast_radius_calculator.py" --traffic-share 0.05 --user-pop 1000000 --duration-min 15

# 3. Generate postmortem after the experiment
python "$SKILL/scripts/experiment_postmortem.py" --plan experiment.json --result-log results.txt

The 3 Python tools

All stdlib-only. Run with --help.

experiment_designer.py

Generates a structured experiment plan from inputs. Enforces the required sections (hypothesis, steady-state metric, blast radius, abort criteria, rollback).

bash
python scripts/experiment_designer.py \
  --target "checkout-svc" \
  --hypothesis "p99 latency stays <500ms when payment-svc is slow" \
  --attack latency \
  --magnitude "+200ms" \
  --duration-min 15 \
  --blast-radius "5% of US traffic" \
  --abort-if "p99 > 1000ms OR error_rate > baseline + 1pp"

Outputs a markdown plan with: hypothesis, steady-state, attack, magnitude, duration, blast radius, abort criteria, rollback procedure, monitoring dashboards, and learning question.

blast_radius_calculator.py

Computes the blast radius of a planned experiment. Given traffic share + user population + duration, calculates expected affected users, expected error budget burn, and a risk score.

bash
python scripts/blast_radius_calculator.py \
  --traffic-share 0.05 \
  --user-pop 1000000 \
  --duration-min 15 \
  --baseline-availability 0.999 \
  --expected-impact-availability 0.95

Outputs:

  • Expected affected users
  • Error budget consumed (in minutes of error budget)
  • Risk score: GREEN / YELLOW / RED
  • Recommendation: PROCEED / REDUCE / ABORT

GREEN = <1% error budget; YELLOW = 1-10%; RED = >10%.

experiment_postmortem.py

Produces a structured postmortem from an experiment plan + results. Catches the common postmortem failure modes: no learning recorded, no follow-up actions, blame-laden language.

bash
python scripts/experiment_postmortem.py --plan experiment.json --result-log results.txt

Outputs markdown with: summary, hypothesis (was it confirmed/refuted?), what we learned, what surprised us, follow-up actions with owners, and link to next experiment.

The 7 attack types (taxonomy)

Different attacks reveal different weaknesses. See references/attack_taxonomy.md for full detail.

AttackWhat it testsTooling
LatencyTimeouts, retries, circuit breakerstc, Chaos Mesh NetworkChaos
ErrorError handling, fallback pathsChaos Mesh HTTPChaos, Toxiproxy
Resource (CPU, memory, disk)Saturation handling, autoscalingChaos Mesh StressChaos, stress-ng
Network partitionSplit-brain, consensus, failoverChaos Mesh NetworkChaos partition
Dependency failureGraceful degradation, fallbackService mesh fault injection
TimeClock skew, NTP issueslibfaketime, Chaos Mesh TimeChaos
Infrastructure (kill instance)Auto-recovery, failoverAWS FIS, Chaos Monkey

Pick the attack that matches the hypothesis. "What happens if X is slow?" → latency. "What happens if X loses network?" → partition.

Show full SKILL.md (358 more words)Show less

Tooling chooser

ToolBest forPricingStack
Chaos ToolkitLightweight, language-agnostic, JSON experimentsOSSAny
Chaos MeshKubernetes-native, rich CRDs, in-clusterOSSKubernetes
LitmusKubernetes, Argo-integrated, large libraryOSS + EnterpriseKubernetes
GremlinEnterprise SaaS, multi-cloud, auditPaidAny
AWS FISAWS-native, IAM-integrated, EC2/ECS/EKSPaid (AWS)AWS
CustomNiche needs, single-cloud, low budgetNoneAny

Decision rules:

  • k8s-only stack + OSS → Chaos Mesh or Litmus (Litmus has bigger experiment library)
  • Multi-cloud + OSS → Chaos Toolkit
  • AWS-heavy + simple needs → AWS FIS
  • Enterprise + audit/compliance → Gremlin

See references/tooling_landscape.md for trade-offs.

Workflows

Workflow 1: Design and run a single experiment
1. State a hypothesis: "When [fault], steady-state metric X stays within Y."
2. Identify the steady-state metric — must be measurable BEFORE the experiment.
3. Run blast_radius_calculator.py — confirm GREEN before proceeding.
4. Run experiment_designer.py to produce the plan.
5. Get a peer review of the plan; confirm abort criteria are concrete.
6. Notify the on-call team in #incidents (or whatever channel).
7. Run the experiment with monitoring open.
8. If abort criteria are hit, abort immediately; record what happened.
9. Run experiment_postmortem.py to capture learnings.
10. File follow-up actions; link to next experiment.
Workflow 2: Game Day exercise
1. Pick a scenario (e.g., "primary database fails over").
2. Identify all dependent services that should keep working.
3. Build a multi-experiment plan covering each layer.
4. Schedule with stakeholders; on-call coverage required.
5. Run with a facilitator who manages the scenario.
6. Capture observations in a shared doc as they happen.
7. Single combined postmortem covering all observations.
8. Track follow-up actions in a board with owners.
Workflow 3: Continuous chaos (game days → daily)
1. Start: weekly Game Day in staging.
2. Move to: weekly Game Day in production with limited blast radius.
3. Mature to: continuous chaos via scheduled experiments (Litmus chaos schedule, Gremlin scenarios).
4. Wire to deployment: every prod deploy triggers a baseline chaos sweep.
5. Track: experiments per week, weaknesses discovered, MTTR trend.

Composition with other skills

This skill explicitly composes with two others in this library:

SkillComposition
feature-flags-architectKill switches defined there are the abort triggers here
kubernetes-operatorOperators are common chaos targets (test reconcile under fault)
incident-responseChaos experiments that escalate become incidents

Anti-patterns

  • No hypothesis — "let's break things" is sabotage, not engineering
  • No steady-state metric — without a baseline, you can't tell if X broke
  • No blast radius bound — full-prod experiment without limits = outage
  • No abort criteria — see above; this is mandatory
  • No on-call coverage — chaos without monitoring is unmonitored production
  • Chaos in staging only — staging never has prod failure modes
  • Chaos in dev — useless; dev has different failure modes from prod
  • One-off chaos — single experiment is a press release; learning requires recurrence
  • Blame-laden postmortem — record causes, not blame; teams stop running chaos otherwise

References

  • references/chaos_principles.md — the 4 principles, history, when to start
  • references/experiment_design.md — hypothesis structure, steady-state metrics, abort criteria
  • references/attack_taxonomy.md — 7 attack types with examples and tooling
  • references/tooling_landscape.md — Chaos Toolkit / Mesh / Litmus / Gremlin / FIS / DIY

Slash command

/chaos-experiment — interactive experiment design wizard that runs all 3 tools.

Asset templates

  • assets/experiment_template.md — fill-in plan template
  • assets/postmortem_template.md — structured postmortem template

Verifiable success

A team using this skill should achieve:

  • 100% of chaos experiments have a written hypothesis, abort criteria, and blast-radius calculation
  • Blast radius for any single experiment never exceeds 10% of error budget
  • Mean time between chaos experiments <14 days (continuous, not one-off)
  • Each experiment produces ≥1 follow-up action that gets shipped
  • No chaos experiment escalates to a customer-impacting incident in trailing 90 days

© alirezarezvani, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 9 other files (scripts, references, assets) in engineering/skills/chaos-engineering of alirezarezvani/claude-skills.

  • SKILL.md
  • assets/experiment_template.md
  • assets/postmortem_template.md
  • references/attack_taxonomy.md
  • references/chaos_principles.md
  • references/experiment_design.md
  • references/tooling_landscape.md
  • scripts/blast_radius_calculator.py
  • scripts/experiment_designer.py
  • scripts/experiment_postmortem.py

Open the folder on GitHubat commit 19392f7

Compare with similar skills

Chaos Engineering next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Chaos Engineering compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Chaos Engineering this skillalirezarezvani/claude-skills28k—~2.7kAutomated safety check: PassMIT
AWS Cdk Developmentzxkane/aws-skills3672 repos~2.5kAutomated safety check: PassMIT
Senior DevOps Toolkitmaslennikov-ig/claude-code-orchestrator-kit2606 repos~1.1kAutomated safety check: NotesCustom licence
Vetatilladeniz/Kubeli3872 repos~1.6kAutomated safety check: PassMIT
Install Boltmcpboltmcp/boltmcp371—~2.3kAutomated safety check: PassNone
Provider Bug Reviewmondoohq/mql412—~2.9kAutomated safety check: PassCustom licence

Similar skills

  • AWS Cdk Development

    zxkane/aws-skills

    AWS Cloud Development Kit (CDK) expert for building cloud infrastructure with TypeScript/Python.

    367 GitHub starsUsed in 2 repos~2.5k tokens
    DevOps & CloudAuto-check passed
  • Senior DevOps Toolkit

    maslennikov-ig/claude-code-orchestrator-kit

    Comprehensive DevOps skill for CI/CD, infrastructure automation, containerization, and cloud platforms (AWS, GCP, Azure). Includes pipeline setup…

    260 GitHub starsUsed in 6 repos~1.1k tokens
    DevOps & CloudAuto-check: notes
  • Vet

    atilladeniz/Kubeli

    Run vet immediately after ANY logical unit of code changes. An agent skill from atilladeniz/Kubeli.

    387 GitHub starsUsed in 2 repos~1.6k tokens
    DevOps & CloudAuto-check passed
  • Install Boltmcp

    boltmcp/boltmcp

    A skill your agent uses when asked to help install or uninstall BoltMCP

    371 GitHub stars~2.3k tokensUpdated yesterday
    DevOps & CloudAuto-check passed
  • Deep static code review of an mql provider for logic errors, nil-handling bugs, pagination truncation, caching/id collisions, and other defects that silently give users wrong data.

    412 GitHub stars~2.9k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Dstack Presets

    dstackai/dstack

    Create and manage dstack presets: a toolkit that streamlines model inference optimization with agents, and a portable preset format.

    2.3k GitHub stars~403 tokensUpdated today
    DevOps & CloudAuto-check passed

More from alirezarezvani/claude-skills

All 342 skills in this repo
  • Agile Product Owner

    alirezarezvani/claude-skills

    Writes INVEST-checked user stories with acceptance criteria, splits epics, plans sprints from velocity and ranks the backlog with a weighted score.

    28k GitHub starsUsed in 3 repos~3.2k tokens
    Auto-check passed
  • Product Strategist

    alirezarezvani/claude-skills

    OKR cascade toolkit for product leaders: generates aligned company-to-team OKRs from five strategy types and scores how well they line up.

    28k GitHub starsUsed in 2 repos~1.8k tokens
    Auto-check passed
  • App Store Optimization

    alirezarezvani/claude-skills

    App Store Optimization (ASO) toolkit for researching keywords, analyzing competitor rankings, generating metadata suggestions, and improving app visibility on Apple App Store and Google Play Store.

    28k GitHub starsUsed in 1 repo~4.2k tokens
    Auto-check passed
  • AWS Solution Architect

    alirezarezvani/claude-skills

    Design AWS architectures for startups using serverless patterns and IaC templates.

    28k GitHub starsUsed in 1 repo~2.5k tokens
    Auto-check passed
  • Campaign Analytics

    alirezarezvani/claude-skills

    Calculates attribution, funnel and ROI figures for marketing campaigns with three Python scripts that need only the standard library.

    28k GitHub starsUsed in 1 repo~2.1k tokens
    Auto-check passed
  • Code to PRD

    alirezarezvani/claude-skills

    Reverse-engineers a frontend, backend or fullstack codebase into a product requirements document with per-page docs, an enum dictionary and an API inventory.

    28k GitHub starsUsed in 1 repo~4.9k tokens
    Auto-check passed

Categories

Questions about Chaos Engineering

What does Chaos Engineering do?

A skill your agent uses when planning, running, or learning from chaos engineering experiments. Chaos Engineering is an agent skill from alirezarezvani/claude-skills. Use when planning, running, or learning from chaos engineering experiments.

When should I use Chaos Engineering?

Chaos Engineering fits situations like: learning from chaos engineering experiments; chaos experiment; fault injection; resilience test.

How do I install Chaos Engineering in Claude Code?

Run `npx skills add alirezarezvani/claude-skills --skill chaos-engineering -a claude-code`. Or copy the skill folder (engineering/skills/chaos-engineering in alirezarezvani/claude-skills) into .claude/skills/chaos-engineering in your project. Claude Code loads it when a task matches its description.

How do I install Chaos Engineering in Codex?

Run `npx skills add alirezarezvani/claude-skills --skill chaos-engineering -a codex`. Or copy the skill folder (engineering/skills/chaos-engineering in alirezarezvani/claude-skills) into .agents/skills/chaos-engineering in your project. Codex loads it when a task matches its description.

Can I use Chaos Engineering in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add alirezarezvani/claude-skills --skill chaos-engineering -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/chaos-engineering, .gemini/skills/chaos-engineering, .github/skills/chaos-engineering and .opencode/skills/chaos-engineering in your project.

What does Chaos Engineering need to run?

Going by SKILL.md and its folder, Chaos Engineering needs Python for the scripts in its folder and the command-line tools its instructions call (python). Our summary lists: Python 3.

Does Chaos Engineering access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Chaos Engineering safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Chaos Engineering use?

Chaos Engineering is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Chaos Engineering use?

About 2.7k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 5.8k tokens, read only when the agent opens those files.

What are the alternatives to Chaos Engineering?

Skills that share tags, products or a category with Chaos Engineering: AWS Cdk Development (zxkane/aws-skills, 367 stars), Senior DevOps Toolkit (maslennikov-ig/claude-code-orchestrator-kit, 260 stars), Vet (atilladeniz/Kubeli, 387 stars) and Install Boltmcp (boltmcp/boltmcp, 371 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Chaos Engineering?

alirezarezvani (a GitHub user) maintains it in alirezarezvani/claude-skills, which has 27,891 GitHub stars. The repository holds 342 skills in this directory. The repository was last updated on August 30, 2026.

Source: alirezarezvani/claude-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.