Agent skill

Chaos Engineer

by Jeffallan in Jeffallan/claude-skills

Designs chaos experiments, failure injection and game days for distributed systems, with blast radius limits, rollback plans and written learnings.

MITAuto-check passedDevOps & Cloud

Install Chaos Engineer

skills CLI
$ npx skills add Jeffallan/claude-skills --skill chaos-engineer -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Jeffallan/claude-skills chaos-engineer --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Jeffallan/claude-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/chaos-engineer .claude/skills/chaos-engineer && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
chaos-engineer
GitHub stars
12k
Token cost
~1.8k tokens
SKILL.md length
368 words
Files
6 (incl. references)
Skills in repo
58
Repo updated
First seen
Licence
MIT

At a glance

Designs chaos experiments, failure injection and game days for distributed systems, with blast radius limits, rollback plans and written learnings.

  • Works in 4 steps: Define steady state and apply the… → Create and apply a Litmus ChaosEngine… → Monitor during the experiment → …
  • Designing a chaos experiment with a hypothesis, steady state and blast radius
  • SKILL.md covers When to Use This Skill, Core Workflow, Reference Guide and Safety Checklist, plus 4 more sections
  • Calls kubectl and brew

What it does

This skill helps design and run chaos experiments, build failure injection frameworks and plan game days for distributed systems. The workflow is system analysis (mapping architecture, dependencies, critical paths and failure modes), experiment design (hypothesis, steady state, blast radius and safety controls), controlled execution with monitoring and quick rollback, documenting and fixing what was learned, and automating chaos tests in CI/CD.

Reference notes cover experiment design, infrastructure failures such as server, network, zone and region loss, Kubernetes pod and node experiments with Litmus and Chaos Mesh, tools such as Chaos Monkey, Gremlin and Pumba, and game days. A safety checklist applies to every experiment: define steady-state metrics first, start with the smallest blast radius, script and test a rollback of 30 seconds or less before starting, change one variable at a time, use circuit breakers, feature flags or canary isolation in customer-facing environments, and close each experiment with a written learning summary and a tracked improvement.

When your agent uses it

  • Designing a chaos experiment with a hypothesis, steady state and blast radius
  • Implementing failure injection with Chaos Monkey or Litmus
  • Planning and running a game day exercise
  • Adding continuous chaos testing to CI/CD
  • Testing Kubernetes pod and node failures

Example prompts

  • “Design a chaos experiment for our checkout service that kills one pod and watches latency.”
  • “Plan a game day for a region failover, including rollback steps and a post-mortem template.”
  • “Add a chaos test stage to our CI/CD pipeline with a safe blast radius.”

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Define steady state and apply the experiment
  2. Create and apply a Litmus ChaosEngine manifest
  3. Monitor during the experiment
  4. Rollback / abort if steady state is violated

What it can do on your machine

Read from SKILL.md and the folder at commit 1be15d8. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • kubectl
    • brew

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • github.com
    • synergetic.solutions
    • jeffallan.github.io

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Chaos Engineer loads about 1.8k tokens when it runs, and up to ~15k if it reads all its reference files. Until then it costs about 124 tokens; SKILL.md has 368 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~124
When it runs · the whole SKILL.md, loaded when a task matches
~1.8k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~15k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Jeffallan/claude-skills at commit 1be15d8, republished under its MIT licence (© Jeffallan). 368 words, ~1,759 tokens.

Download SKILL.mdSave it as .claude/skills/chaos-engineer/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.
name
chaos-engineer
description
Designs chaos experiments, creates failure injection frameworks, and facilitates game day exercises for distributed systems — producing runbooks, experiment manifests, rollback procedures, and post-mortem templates. Use when designing chaos experiments, implementing failure injection frameworks, or conducting game day exercises. Invoke for chaos experiments, resilience testing, blast radius control, game days, antifragile systems, fault injection, Chaos Monkey, Litmus Chaos.
license
MIT
metadata.author
https://github.com/Jeffallan
metadata.company
https://synergetic.solutions
metadata.version
1.1.0
metadata.domain
devops
metadata.triggers
chaos engineering, resilience testing, failure injection, game day, blast radius, chaos experiment, fault injection, Chaos Monkey, Litmus Chaos, antifragile
metadata.role
specialist
metadata.scope
implementation
metadata.output-format
code
metadata.related-skills
sre-engineer, devops-engineer, kubernetes-specialist

Chaos Engineer

When to Use This Skill

  • Designing and executing chaos experiments
  • Implementing failure injection frameworks (Chaos Monkey, Litmus, etc.)
  • Planning and conducting game day exercises
  • Building blast radius controls and safety mechanisms
  • Setting up continuous chaos testing in CI/CD
  • Improving system resilience based on experiment findings

Core Workflow

  1. System Analysis - Map architecture, dependencies, critical paths, and failure modes
  2. Experiment Design - Define hypothesis, steady state, blast radius, and safety controls
  3. Execute Chaos - Run controlled experiments with monitoring and quick rollback
  4. Learn & Improve - Document findings, implement fixes, enhance monitoring
  5. Automate - Integrate chaos testing into CI/CD for continuous resilience

Reference Guide

Load detailed guidance based on context:

TopicReferenceLoad When
Experimentsreferences/experiment-design.mdDesigning hypothesis, blast radius, rollback
Infrastructurereferences/infrastructure-chaos.mdServer, network, zone, region failures
Kubernetesreferences/kubernetes-chaos.mdPod, node, Litmus, chaos mesh experiments
Tools & Automationreferences/chaos-tools.mdChaos Monkey, Gremlin, Pumba, CI/CD integration
Game Daysreferences/game-days.mdPlanning, executing, learning from game days

Safety Checklist

Non-obvious constraints that must be enforced on every experiment:

  • Steady state first — define and verify baseline metrics before injecting any failure
  • Blast radius cap — start with the smallest possible impact scope; expand only after validation
  • Automated rollback ≤ 30 seconds — abort path must be scripted and tested before the experiment begins
  • Single variable — change only one failure condition at a time until behaviour is well understood
  • No production without safety nets — customer-facing environments require circuit breakers, feature flags, or canary isolation
  • Close the loop — every experiment must produce a written learning summary and at least one tracked improvement
Show full SKILL.md (115 more words)Show less

Output Templates

When implementing chaos engineering, provide:

  1. Experiment design document (hypothesis, metrics, blast radius)
  2. Implementation code (failure injection scripts/manifests)
  3. Monitoring setup and alert configuration
  4. Rollback procedures and safety controls
  5. Learning summary and improvement recommendations

Concrete Example: Pod Failure Experiment (Litmus Chaos)

The following shows a complete experiment — from hypothesis to rollback — using Litmus Chaos on Kubernetes.

Step 1 — Define steady state and apply the experiment
bash
# Verify baseline: p99 latency < 200ms, error rate < 0.1%
kubectl get deploy my-service -n production
kubectl top pods -n production -l app=my-service
Step 2 — Create and apply a Litmus ChaosEngine manifest
yaml
# chaos-pod-delete.yaml
apiVersion: litmuschaos.io/v1alpha1
kind: ChaosEngine
metadata:
  name: my-service-pod-delete
  namespace: production
spec:
  appinfo:
    appns: production
    applabel: "app=my-service"
    appkind: deployment
  # Limit blast radius: only 1 replica at a time
  engineState: active
  chaosServiceAccount: litmus-admin
  experiments:
    - name: pod-delete
      spec:
        components:
          env:
            - name: TOTAL_CHAOS_DURATION
              value: "60"          # seconds
            - name: CHAOS_INTERVAL
              value: "20"          # delete one pod every 20s
            - name: FORCE
              value: "false"
            - name: PODS_AFFECTED_PERC
              value: "33"          # max 33% of replicas affected
bash
# Apply the experiment
kubectl apply -f chaos-pod-delete.yaml

# Watch experiment status
kubectl describe chaosengine my-service-pod-delete -n production
kubectl get chaosresult my-service-pod-delete-pod-delete -n production -w
Step 3 — Monitor during the experiment
bash
# Tail application logs for errors
kubectl logs -l app=my-service -n production --since=2m -f

# Check ChaosResult verdict when complete
kubectl get chaosresult my-service-pod-delete-pod-delete \
  -n production -o jsonpath='{.status.experimentStatus.verdict}'
Step 4 — Rollback / abort if steady state is violated
bash
# Immediately stop the experiment
kubectl patch chaosengine my-service-pod-delete \
  -n production --type merge -p '{"spec":{"engineState":"stop"}}'

# Confirm all pods are healthy
kubectl rollout status deployment/my-service -n production

Concrete Example: Network Latency with toxiproxy

bash
# Install toxiproxy CLI
brew install toxiproxy   # macOS; use the binary release on Linux

# Start toxiproxy server (runs alongside your service)
toxiproxy-server &

# Create a proxy for your downstream dependency
toxiproxy-cli create -l 0.0.0.0:22222 -u downstream-db:5432 db-proxy

# Inject 300ms latency with 10% jitter — blast radius: this proxy only
toxiproxy-cli toxic add db-proxy -t latency -a latency=300 -a jitter=30

# Run your load test / observe metrics here ...

# Remove the toxic to restore normal behaviour
toxiproxy-cli toxic remove db-proxy -n latency_downstream

Concrete Example: Chaos Monkey (Spinnaker / standalone)

bash
# chaos-monkey-config.yml — restrict to a single ASG
deployment:
  enabled: true
  regionIndependence: false
chaos:
  enabled: true
  meanTimeBetweenKillsInWorkDays: 2
  minTimeBetweenKillsInWorkDays: 1
  grouping: APP           # kill one instance per app, not per cluster
  exceptions:
    - account: production
      region: us-east-1
      detail: "*-canary"  # never kill canary instances

# Apply and trigger a manual kill for testing
chaos-monkey --app my-service --account staging --dry-run false

Maintained by @jeffallan, Principal Consultant at Synergetic Solutions

Documentation

© Jeffallan, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 5 other files (references) in skills/chaos-engineer of Jeffallan/claude-skills.

  • SKILL.md
  • references/chaos-tools.md
  • references/experiment-design.md
  • references/game-days.md
  • references/infrastructure-chaos.md
  • references/kubernetes-chaos.md

Open the folder on GitHubat commit 1be15d8

Compare with similar skills

Chaos Engineer next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Chaos Engineer compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Chaos Engineer this skillJeffallan/claude-skills12k—~1.8kAutomated safety check: PassMIT
Error HandlerEliasOulkadi/shokunin114—~3.6kAutomated safety check: NotesMIT
Executing Distributed System Testsshenli/distributed-system-testing231—~5.1kAutomated safety check: NotesMIT
Incident ResponderDokhacgiakhoa/Agent-Skills-4-Vibe-Coding-CLI508—~706Automated safety check: PassCustom licence
Mint Enrollfullsend-ai/fullsend149—~3.6kAutomated safety check: NotesApache-2.0
Chaos Engineeringalirezarezvani/claude-skills28k—~2.7kAutomated safety check: PassMIT

Similar skills

  • Error Handler

    EliasOulkadi/shokunin

    Design error handling, structured logging, and observability with OpenTelemetry (traces, metrics, logs), error classification, recovery patterns (retry with jitter, circuit breaker, bulkhead…

    114 GitHub stars~3.6k tokensUpdated 5 days ago
    DevOps & CloudAuto-check: notes
  • Executing Distributed System Tests

    shenli/distributed-system-testing

    A skill your agent uses when running a previously designed distributed-systems test plan against a real or simulated cluster — driving fault injection, workload, chaos scenarios, linearizability /…

    231 GitHub stars~5.1k tokensUpdated 2 mo ago
    DevOps & CloudAuto-check: notes
  • Incident Responder

    Dokhacgiakhoa/Agent-Skills-4-Vibe-Coding-CLI

    Expert SRE incident responder specializing in rapid problem resolution.

    508 GitHub stars~706 tokensUpdated 4 mo ago
    DevOps & CloudAuto-check passed
  • Mint Enroll

    fullsend-ai/fullsend

    SRE runbook for enrolling new GitHub repos into the fullsend token mint service using go run ./cmd/fullsend from this checkout.

    149 GitHub stars~3.6k tokensUpdated today
    DevOps & CloudAuto-check: notes
  • Chaos Engineering

    alirezarezvani/claude-skills

    A skill your agent uses when planning, running, or learning from chaos engineering experiments.

    28k GitHub stars~2.7k tokensUpdated 1 mo ago
    DevOps & CloudAuto-check passed
  • Knowledge Ops

    alirezarezvani/claude-skills

    A skill your agent uses when a Head of Ops, Knowledge Manager, or TPM-Internal needs to author, validate, or clean up company SOPs and internal runbooks (procurement intake, vendor offboarding…

    28k GitHub stars~4.1k tokensUpdated 1 mo ago
    DevOps & CloudAuto-check passed

More from Jeffallan/claude-skills

All 58 skills in this repo
  • API Designer

    Jeffallan/claude-skills

    Designs REST and GraphQL APIs from resource modeling to an OpenAPI 3.1 contract, with versioning, pagination and RFC 7807 error handling.

    12k GitHub starsUsed in 1 repo~2k tokens
    Auto-check passed
  • CLI Developer

    Jeffallan/claude-skills

    Walks through designing, building and polishing a command-line tool: user workflow and command hierarchy, implementation in commander, click, typer or cobra, completions and cross-platform testing.

    12k GitHub starsUsed in 1 repo~1.2k tokens
    Auto-check passed
  • Kubernetes Specialist

    Jeffallan/claude-skills

    Creates and checks Kubernetes manifests, Helm charts, RBAC and network policies, and helps debug pod problems, with kubectl checks and rollback steps.

    12k GitHub starsUsed in 1 repo~2.1k tokens
    Auto-check passed
  • Laravel Specialist

    Jeffallan/claude-skills

    Builds Laravel 10+ applications with Eloquent models, Sanctum authentication, Horizon queues, API resources and Livewire components, tested with Pest or PHPUnit.

    12k GitHub starsUsed in 1 repo~2.1k tokens
    Auto-check passed
  • Pandas Pro

    Jeffallan/claude-skills

    Handles pandas DataFrame work: cleaning, merging, groupby aggregation, pivots, time-series resampling and memory tuning, with checks on dtypes, shapes and nulls.

    12k GitHub starsUsed in 1 repo~1.5k tokens
    Auto-check passed
  • Apache Spark Engineer

    Jeffallan/claude-skills

    Guides writing and tuning Apache Spark jobs: DataFrame and RDD code, Spark SQL, partitioning, caching, shuffle tuning and structured streaming.

    12k GitHub starsUsed in 1 repo~1.7k tokens
    Auto-check passed

Works with

Categories

Questions about Chaos Engineer

What does Chaos Engineer do?

Designs chaos experiments, failure injection and game days for distributed systems, with blast radius limits, rollback plans and written learnings. This skill helps design and run chaos experiments, build failure injection frameworks and plan game days for distributed systems. The workflow is system analysis (mapping architecture, dependencies, critical paths and failure modes), experiment design (hypothesis, steady state, blast radius and safety controls), controlled execution with monitoring and quick rollback, documenting and fixing what was learned, and automating chaos tests in CI/CD.

When should I use Chaos Engineer?

Chaos Engineer fits situations like: designing a chaos experiment with a hypothesis, steady state and blast radius; implementing failure injection with Chaos Monkey or Litmus; planning and running a game day exercise; adding continuous chaos testing to CI/CD.

How do I install Chaos Engineer in Claude Code?

Run `npx skills add Jeffallan/claude-skills --skill chaos-engineer -a claude-code`. Or copy the skill folder (skills/chaos-engineer in Jeffallan/claude-skills) into .claude/skills/chaos-engineer in your project. Claude Code loads it when a task matches its description.

How do I install Chaos Engineer in Codex?

Run `npx skills add Jeffallan/claude-skills --skill chaos-engineer -a codex`. Or copy the skill folder (skills/chaos-engineer in Jeffallan/claude-skills) into .agents/skills/chaos-engineer in your project. Codex loads it when a task matches its description.

Can I use Chaos Engineer in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Jeffallan/claude-skills --skill chaos-engineer -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/chaos-engineer, .gemini/skills/chaos-engineer, .github/skills/chaos-engineer and .opencode/skills/chaos-engineer in your project.

What does Chaos Engineer need to run?

Going by SKILL.md and its folder, Chaos Engineer needs the command-line tools its instructions call (kubectl and brew).

Does Chaos Engineer access the network?

SKILL.md names 3 domains. As links in the text: github.com, synergetic.solutions and jeffallan.github.io. This is read from the text; nothing was executed.

Is Chaos Engineer safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Chaos Engineer use?

Chaos Engineer is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Chaos Engineer use?

About 1.8k tokens (SKILL.md is roughly 7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 14k tokens, read only when the agent opens those files.

What are the alternatives to Chaos Engineer?

Skills that share tags, products or a category with Chaos Engineer: Error Handler (EliasOulkadi/shokunin, 114 stars), Executing Distributed System Tests (shenli/distributed-system-testing, 231 stars), Incident Responder (Dokhacgiakhoa/Agent-Skills-4-Vibe-Coding-CLI, 508 stars) and Mint Enroll (fullsend-ai/fullsend, 149 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Chaos Engineer?

Jeffallan (a GitHub user) maintains it in Jeffallan/claude-skills, which has 11,802 GitHub stars. The repository holds 58 skills in this directory. The repository was last updated on October 3, 2026.

Source: Jeffallan/claude-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.