Chaos Engineering
alirezarezvani/claude-skills
A skill your agent uses when planning, running, or learning from chaos engineering experiments.
Validate system resilience through controlled fault injection.
$ npx skills add petrkindlmann/qa-skills --skill chaos-engineering -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install petrkindlmann/qa-skills chaos-engineering --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/petrkindlmann/qa-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/chaos-engineering .claude/skills/chaos-engineering && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "chaos-engineering" agent skill from https://github.com/petrkindlmann/qa-skills/tree/main/skills/chaos-engineering into .claude/skills/chaos-engineering/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "chaos-engineering", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/petrkindlmann/qa-skills/tree/main/skills/chaos-engineeringType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add petrkindlmann/qa-skills --skill chaos-engineering -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install petrkindlmann/qa-skills chaos-engineering --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/petrkindlmann/qa-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/chaos-engineering .agents/skills/chaos-engineering && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "chaos-engineering" agent skill from https://github.com/petrkindlmann/qa-skills/tree/main/skills/chaos-engineering into .agents/skills/chaos-engineering/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "chaos-engineering", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add petrkindlmann/qa-skills --skill chaos-engineering -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install petrkindlmann/qa-skills chaos-engineering --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/petrkindlmann/qa-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/chaos-engineering .cursor/skills/chaos-engineering && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "chaos-engineering" agent skill from https://github.com/petrkindlmann/qa-skills/tree/main/skills/chaos-engineering into .cursor/skills/chaos-engineering/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "chaos-engineering", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/petrkindlmann/qa-skills.git --path skills/chaos-engineering--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add petrkindlmann/qa-skills --skill chaos-engineering -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install petrkindlmann/qa-skills chaos-engineering --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/petrkindlmann/qa-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/chaos-engineering .gemini/skills/chaos-engineering && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "chaos-engineering" agent skill from https://github.com/petrkindlmann/qa-skills/tree/main/skills/chaos-engineering into .gemini/skills/chaos-engineering/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "chaos-engineering", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install petrkindlmann/qa-skills chaos-engineeringInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add petrkindlmann/qa-skills --skill chaos-engineering -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/petrkindlmann/qa-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/chaos-engineering .github/skills/chaos-engineering && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "chaos-engineering" agent skill from https://github.com/petrkindlmann/qa-skills/tree/main/skills/chaos-engineering into .github/skills/chaos-engineering/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "chaos-engineering", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add petrkindlmann/qa-skills --skill chaos-engineering -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install petrkindlmann/qa-skills chaos-engineering --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/petrkindlmann/qa-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/chaos-engineering .opencode/skills/chaos-engineering && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "chaos-engineering" agent skill from https://github.com/petrkindlmann/qa-skills/tree/main/skills/chaos-engineering into .opencode/skills/chaos-engineering/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "chaos-engineering", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
chaos-engineeringValidate system resilience through controlled fault injection.
Chaos Engineering is an agent skill from petrkindlmann/qa-skills. Validate system resilience through controlled fault injection. Covers hypothesis-driven chaos experiments, failure injection types (network, service, infrastructure, dependency), LitmusChaos/Chaos Mesh/AWS FIS/Gremlin/toxiproxy tooling, automated abort gating, game day planning, and progressive chaos adoption. Use when: "chaos engineering," "fault injection," "resilience test," "game day," "failure recovery," "system reliability," "blast radius." Not for: safe rollout flags/canary/dark launch during a release —…
Its SKILL.md is about 6.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including reference files (for example `references/fault-injection.md`).
It sits in DevOps & Cloud, covering Chaos engineering. It works with Amazon Web Services. The repository describes itself as: 50 QA and test-automation skills for Claude Code, Codex, Cursor, and any Agent Skills Standard runtime. The licence is MIT.
9 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit b3bb61b. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
kubectlFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use kubectl, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Chaos Engineering loads about 6.2k tokens when it runs, and up to ~7.5k if it reads all its reference files. Until then it costs about 191 tokens; SKILL.md has 2,450 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from petrkindlmann/qa-skills at commit b3bb61b, republished under its MIT licence (© petrkindlmann). 2,450 words, ~6,200 tokens.
.claude/skills/chaos-engineering/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.<objective>
Chaos engineering is the discipline of experimenting on a system to build confidence in its ability to withstand turbulent conditions. It is not random destruction -- it is hypothesis-driven, controlled experimentation that reveals weaknesses before they cause outages. A retry that "works in the demo" silently double-charges customers when the payment API times out; the only way to know is to inject the timeout and watch.
</objective>
| Situation | Go to |
|---|---|
| First experiment ever, team is new | Starting Small → First Three Experiments |
| Designing one experiment | Chaos Experiment Workflow (5 steps) |
| Picking a tool for your environment | Tools → Choosing a tool decision tree |
| Running a team session | Game Day Planning |
| Need runnable injection commands/configs | references/fault-injection.md |
| Want the abort to fire without a human | references/fault-injection.md → Automated abort |
Check .agents/qa-project-context.md first. If it exists, use it as context and skip questions already answered there.
Environment and readiness:
Architecture:
Current resilience practices:
Team and culture:
Every chaos experiment starts with a hypothesis: "We believe that if [failure X occurs], the system will [expected behavior Y]." Without a hypothesis, you are just breaking things. The hypothesis names concrete steady-state metrics (baseline metrics) — error rate, latency, throughput — and the bound each may move to.
Example hypothesis: "We believe that if the primary database becomes unavailable, the application will serve cached data for read requests and queue write requests for up to 5 minutes without user-visible errors. Blast radius: staging, one service. Steady-state baseline: error rate <0.1%, P95 latency <300ms."
The first chaos experiment should not be "shut down production." It should be "add 200ms latency to one non-critical service in staging." Increase scope gradually as confidence and tooling mature.
If you cannot detect problems in real time, you cannot safely inject failures. Chaos experiments without monitoring are just outages with extra steps. Verify dashboards, alerts, and on-call processes before running any experiment.
Running chaos experiments in automated pipelines is valuable, but game days -- scheduled sessions where the team runs experiments together and practices response -- build the human skills that matter during real incidents.
Every chaos experiment follows this five-step process.
Identify the metrics that define "normal" and predict what should happen during the experiment.
Experiment: Database failover
Steady state:
- Error rate: < 0.1%
- P95 latency: < 300ms
- Successful orders per minute: > 50
Hypothesis: When the primary database fails over to the replica,
- Error rate will spike to < 2% for < 30 seconds
- P95 latency will increase to < 1s for < 60 seconds
- No orders will be permanently lost
- The application will recover without manual interventionInject the failure in a controlled way with a clear scope and duration.
Injection:
Target: primary database (PostgreSQL)
Method: block TCP port 5432 on the primary instance
Scope: single database instance
Duration: 60 seconds
Blast radius: staging environment only (first run)
Abort conditions:
- Error rate > 10% for > 2 minutes
- Any data corruption detected
- Manual abort by experiment ownerDuring the experiment, monitor all relevant metrics in real time. Assign observers to specific dashboards.
Observation assignments:
- Engineer A: application error rate and latency dashboard
- Engineer B: database metrics (connections, replication lag, failover status)
- Engineer C: application logs (search for database connection errors)
- Engineer D: business metrics (order count, payment processing)After the experiment, analyze what happened versus what was expected.
Analysis checklist:
- Did the system behave as hypothesized? (Y/N, with details)
- How long was the impact? (Expected vs. actual duration)
- Were any errors visible to users?
- Was any data lost or corrupted?
- Did monitoring and alerting detect the problem correctly?
- How long before alerts fired?
- What was the recovery time?Document findings, fix resilience gaps, and schedule a re-run to verify the fix.
Findings document:
Experiment: Database failover (2026-03-20)
Hypothesis: Confirmed / Partially confirmed / Disproved
Recovery time: 45s (expected vs actual: expected <10s, actual 45s)
Data integrity: no rows lost; 3 writes returned 500 instead of queueing
Findings:
- Connection pool did not detect stale connections for 45 seconds (expected: <10s)
- Retry logic worked correctly for read operations
- Write operations returned 500 errors for 38 seconds (expected: queued)
Action items (every one has an owner and a due date — no item is deferred):
- [ ] Configure connection pool health checks — assigned to @maria, due 2026-03-31
- [ ] Implement write queue with 5-minute buffer — assigned to @dan, due 2026-04-02
- [ ] Re-run experiment after fixes deployed (re-run scheduled 2026-04-03);
specific metrics to check on re-run: stale-connection detection <10s,
zero write 500s, error rate <2%| Failure | Tool | Use Case |
|---|---|---|
| Latency injection | tc, toxiproxy, Gremlin | Simulate slow network, distant regions |
| Packet loss | tc netem, Chaos Mesh | Simulate unreliable network |
| DNS failure | iptables, CoreDNS manipulation | Simulate DNS outage |
| Network partition | iptables, Chaos Mesh | Simulate split-brain scenarios |
| Bandwidth restriction | tc, toxiproxy | Simulate congested network |
See references/fault-injection.md for the tc netem latency/packet-loss commands and the toxiproxy latency config.
| Failure | Method | Use Case |
|---|---|---|
| Service crash | Kill process, pod delete | Simulate unexpected crash |
| Service slowdown | CPU stress, thread pool exhaustion | Simulate overloaded service |
| Error injection | Return 500/503, throw exceptions | Simulate application errors |
| Memory pressure | stress-ng, Chaos Mesh | Simulate memory leaks |
See references/fault-injection.md for the kubectl delete pod command and the LitmusChaos pod-delete ChaosEngine manifest.
| Failure | Method | Use Case |
|---|---|---|
| Disk full | fallocate, dd | Simulate disk exhaustion |
| CPU exhaustion | stress-ng | Simulate CPU saturation |
| Memory exhaustion | stress-ng | Simulate OOM conditions |
| Clock skew | chrony manipulation, timedatectl | Simulate time drift |
See references/fault-injection.md for the fallocate disk-fill and stress-ng CPU/memory commands.
| Failure | Method | Use Case |
|---|---|---|
| API down | toxiproxy, mock server | Simulate third-party outage |
| Database unavailable | block port, kill process | Simulate database outage |
| Cache unavailable | block Redis port | Simulate cache miss storm |
| Message queue full | fill queue, block consumers | Simulate backpressure |
See references/fault-injection.md for the programmatic toxiproxy integration test that disables Redis and asserts graceful degradation.
| Tool | Type | Best For |
|---|---|---|
| LitmusChaos (3.29.x) | Kubernetes-native, CNCF | K8s environments, CI/CD integration; ChaosCenter UI; Workflows for GameDay-as-code; MCP Server (Oct 2025) drives experiments from an AI assistant |
| Chaos Mesh (2.8.x) | Kubernetes-native, CNCF | K8s with fine-grained control; eBPF chaos via bpfki runtime for kernel-precision faults |
| AWS FIS | Managed AWS service | Cloud-chaos for AWS workloads (EC2, ECS, RDS, EKS); CloudWatch-alarm stop-conditions for auto-abort — primary cloud-native option |
| Gremlin | Managed platform | Teams wanting guided experiments + compliance reporting; Health Checks halt-and-rollback on SLO breach |
| Steadybit | Managed platform | Reliability hub spanning Kubernetes + cloud + on-prem; direct alternative to Gremlin |
| kube-monkey | Open source | Lightweight K8s alternative when Litmus/Chaos Mesh feel heavy |
| Pumba | Open source | Docker-only chaos (containers, networks); pre-K8s and edge |
| toxiproxy | Network proxy, open source | Network fault injection in integration tests |
| tc (traffic control) | Linux kernel | Network latency and packet loss |
| stress-ng | Linux utility | CPU, memory, disk stress testing |
| k6 (+ xk6-disruptor) | Load testing tool | Combined load + chaos scenarios |
Avoid: Chaos Monkey (Netflix) for new projects — low activity, Spinnaker-only path (as of mid-2026). It still works and the repo is not archived, but it only injects instance termination and requires a Spinnaker deployment pipeline. Greenfield work should pick Chaos Mesh, LitmusChaos, or AWS FIS. (The older SimianArmy repo was archived in 2021; don't confuse the two.)
Decision tree:
Running on Kubernetes?
→ Cloud-managed AWS workloads: AWS FIS (cloud-native, IAM-integrated)
→ On K8s with sidecar tolerance: Chaos Mesh (eBPF, fine-grained)
→ On K8s wanting workflows + UI: LitmusChaos (ChaosCenter, Workflows)
→ On K8s lightweight: kube-monkey
Running on plain VMs / Docker?
→ Docker only: Pumba
→ Linux: tc + stress-ng (manual)
Need network fault injection in integration tests?
→ toxiproxy (lightweight, programmatic API)
Need to combine load testing with chaos?
→ k6 with xk6-disruptor extension
Need managed platform with UI and compliance?
→ Gremlin or Steadybit (both commercial)The 2026 trend is treating chaos as scheduled CI jobs rather than ad-hoc events: Litmus Workflows, Steadybit reliability hub, and Gremlin Scenarios all let you define a chaos run as YAML and trigger it from CI on a cron. Pair with the Game Day Planning section below — the human practice still matters; the automation just removes the bottleneck of "we never had time to schedule one." For a concrete cron-gated pipeline (nightly pod-delete on an off-peak window) plus automated abort/stop-condition examples, see references/fault-injection.md.
LitmusChaos shipped an MCP Server in October 2025 that connects an AI assistant such as Claude directly to ChaosCenter: you can list, run, and stop experiments in natural language ("run pod-delete on the frontend pods," "stop the network latency experiment") instead of hand-writing YAML. Relevant if your team already drives ops through an AI agent.
A game day is a scheduled session where the team runs chaos experiments together, practices incident response, and builds confidence in the system's resilience.
2 weeks before:
- [ ] Define 2-3 experiments to run (don't overload the schedule)
- [ ] Write hypotheses for each experiment
- [ ] Get approval from engineering leadership and affected teams
- [ ] Notify support team and stakeholders
- [ ] Verify monitoring and alerting are working
- [ ] Identify rollback procedures for each experiment
- [ ] Schedule 3-4 hour block (experiments + analysis + retro)
1 day before:
- [ ] Confirm all participants and their roles
- [ ] Test that fault injection tools work in the target environment
- [ ] Verify rollback procedures work (dry run)
- [ ] Prepare dashboards and observation assignments
- [ ] Brief the on-call team
- [ ] Confirm abort criteria for each experimentCommunicate before (schedule, scope, abort authority), during (live updates every 15 minutes in a dedicated channel), and after (summary within 24 hours with findings and action items).
Assign roles per experiment: experiment owner (runs it, makes abort decisions), observers (application metrics, infrastructure metrics, logs, user experience), and a scribe (records timeline and decisions).
For each experiment: was the hypothesis confirmed? What surprised us? What action items do we have? For the process: did monitoring detect problems? Did alerts fire? Were we comfortable with the blast radius? Close with action items (with owners and due dates) and schedule the next game day.
For teams new to chaos engineering, start with these three experiments in a pre-production environment.
Why first: Database latency is the most common cause of user-facing slowness, and the experiment is easy to set up and reverse.
Hypothesis: When database latency increases by 500ms, the application
will remain functional with response times under 3 seconds.
Injection: Add 500ms latency to the database connection using toxiproxy.
Duration: 5 minutes.
Environment: staging.
What to observe:
- Application response times (should increase by ~500ms, not 10x)
- Connection pool behavior (should not exhaust connections)
- Timeout handling (requests should not hang indefinitely)
- Circuit breaker activation (if implemented)
- Cache effectiveness (cached reads should be unaffected)Why second: Third-party dependencies fail regularly, and the application's handling of those failures is often untested.
Hypothesis: When the payment provider returns 500 errors, the
application will show a user-friendly error message and allow
retry without duplicate charges.
Injection: Configure mock/proxy to return 500 for payment API calls.
Duration: 10 minutes.
Environment: staging.
What to observe:
- Error message quality (user-friendly, not stack traces)
- Retry behavior (does the application retry? How many times?)
- Idempotency (retries don't create duplicate transactions)
- Fallback (is there an alternative payment path?)
- Monitoring (does the payment failure show up in alerts?)Why third: Cache failures cause "thundering herd" problems where all traffic suddenly hits the database, often causing cascading failures.
Hypothesis: When Redis becomes unavailable, the application will fall
back to direct database queries with degraded but functional performance.
Injection: Block Redis port using toxiproxy or iptables.
Duration: 5 minutes.
Environment: staging.
What to observe:
- Database query volume (should increase but not overwhelm)
- Response times (should increase but remain under 5 seconds)
- Error rate (cache miss should not cause errors)
- Connection pool (database connections should not exhaust)
- Recovery (when cache returns, does the application resume normal behavior?)Injecting failures without the ability to observe their impact is not chaos engineering -- it is sabotage. You will not know if the experiment revealed a problem until a user complains.
Fix: Before any chaos experiment, verify that you can see error rates, latency, throughput, and dependency health in real time. If you cannot, invest in monitoring first. Go one step further and wire the monitor into the experiment so it auto-aborts on breach — AWS FIS CloudWatch stop-conditions, Gremlin Health Checks, or a Litmus promProbe in mode: Continuous (see references/fault-injection.md).
The first chaos experiment should not be "kill the production database." Starting with high-impact experiments before the team has practiced with low-impact ones creates anxiety and potential real outages.
Fix: Start with staging. Start with non-critical services. Start with reversible injections (latency, not data corruption). Build confidence gradually. Graduate to production only after multiple successful staging experiments.
"The experiment is only 60 seconds, we don't need a rollback plan." Then the fault injection tool crashes and the failure persists indefinitely (exactly how a 5-minute tc latency injection becomes a 2-hour outage when the command fails midway).
Fix: Every experiment must have a documented rollback procedure that can be executed in under 30 seconds. Test the rollback before running the experiment, and have a second person ready to abort if the experiment owner is unable to. Prefer a tool-enforced stop-condition over a human finger on the kill switch (see Automated abort in references/fault-injection.md) — and pick injections with a built-in timeout (stress-ng --timeout, Litmus TOTAL_CHAOS_DURATION) so the fault self-clears even if nobody is watching.
Running chaos experiments in production without explicit approval from engineering leadership and affected teams destroys trust and careers.
Fix: Before any production run: get explicit, documented engineering leadership approval; send affected teams notification and brief the on-call team and support team; communicate scope and duration and share the abort criteria with named abort authority; and confirm the blast radius is bounded to the smallest viable target. Start with staging first. Anything skipped here is what turns a controlled experiment into an incident.
Adding controlled failures to a system that is already experiencing problems makes diagnosis harder and extends the outage.
Fix: Cancel or postpone chaos experiments if the system is not in steady state. Check for active incidents before starting. If an unrelated incident starts during an experiment, abort the experiment immediately.
The experiment revealed that the circuit breaker does not work correctly. The team says "interesting" and moves on. The finding is never fixed. The next real outage triggers the same failure.
Fix: Every chaos experiment finding gets a ticket with an owner and a due date. Re-run the experiment after the fix to verify. Track the backlog of chaos findings alongside production incident action items.
Prove the safety machinery works before trusting a single live experiment — an abort path you have never triggered is a hypothesis, not a control. Smallest check first:
tc qdisc add dev eth0 root netem delay 200ms) against one staging service, confirm the steady-state dashboard moves, then run the documented rollback (tc qdisc del dev eth0 root) and confirm metrics return to baseline. Time the rollback — it must complete inside its stated window (under 30s).promProbe in mode: Continuous. Confirm the experiment halts on its own with no human action. If it does not fire, the abort is unverified (see references/fault-injection.md).stress-ng --cpu 4 --timeout 60s, or Litmus TOTAL_CHAOS_DURATION: '60'), then walk away. Confirm the fault clears itself at the deadline even if nobody aborts.If steps 1–3 cannot be demonstrated in staging, the experiment is not safe to promote toward production.
references/)tc netem and toxiproxy network faults, kubectl/LitmusChaos service faults (ChaosEngine with engineState: active), fallocate/stress-ng infrastructure faults, the programmatic toxiproxy dependency-failure test, plus automated abort / stop-condition examples (AWS FIS, Gremlin, Litmus probes, Litmus MCP) and a cron-gated continuous-chaos CI job.© petrkindlmann, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 1 other file (references) in skills/chaos-engineering of petrkindlmann/qa-skills.
Open the folder on GitHubat commit b3bb61b
Chaos Engineering next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Chaos Engineering this skillpetrkindlmann/qa-skills | 170 | — | ~6.2k | Automated safety check: Pass | MIT | |
| Chaos Engineeringalirezarezvani/claude-skills | 28k | — | ~2.7k | Automated safety check: Pass | MIT | |
| AWS Resilience Lifecycleaws/agent-toolkit-for-aws | 2.8k | — | ~1.6k | Automated safety check: Pass | Apache-2.0 | |
| AWS Fault Injection Serviceaws/agent-toolkit-for-aws | 2.8k | — | ~2.2k | Automated safety check: Pass | Apache-2.0 | |
| Cloud Cost Optimizationwshobson/agents | 40k | 14 repos | ~1.7k | Automated safety check: Pass | MIT | |
| Review Docshashicorp/terraform-provider-aws | 11k | — | ~1.3k | Automated safety check: Pass | MPL-2.0 |
alirezarezvani/claude-skills
A skill your agent uses when planning, running, or learning from chaos engineering experiments.
aws/agent-toolkit-for-aws
Guides the end-to-end AWS resilience lifecycle integrating Resilience Hub v2, Fault Injection Service, and Application Recovery Controller.
aws/agent-toolkit-for-aws
Plans, builds, runs, and analyzes fault injection experiments with AWS Fault Injection Service (AWS FIS) to validate application resilience through chaos engineering.
wshobson/agents
Cuts cloud spend across AWS, Azure, GCP and OCI with cost tagging, rightsizing, commitment and spot pricing models, and architecture changes.
hashicorp/terraform-provider-aws
Review a Terraform AWS Provider PR's end-user documentation (website/docs//.markdown): whether docs are needed, description openings, argument/attribute style, section structure, tags wording, code…
zxkane/aws-skills
AWS Cloud Development Kit (CDK) expert for building cloud infrastructure with TypeScript/Python.
petrkindlmann/qa-skills
Test for WCAG 2.2 AA compliance with axe-core + Playwright, keyboard navigation audits, screen reader testing, ARIA pattern validation, and legal compliance mapping (ADA, EAA, Section 508).
petrkindlmann/qa-skills
Goal-driven E2E testing where a browser agent (Playwright MCP / computer-use) reads a natural-language goal and explores the app via the accessibility tree to assert outcomes — no pre-written script.
petrkindlmann/qa-skills
Use AI to write NEW test code from specs, PRDs, user stories, code diffs, bug reports, or OpenAPI specs.
petrkindlmann/qa-skills
Test REST and GraphQL APIs with Playwright APIRequestContext, Supertest, or standalone HTTP clients.
petrkindlmann/qa-skills
Design CI/CD pipelines that run test suites. An agent skill from petrkindlmann/qa-skills.
petrkindlmann/qa-skills
Test for regulatory compliance: GDPR/CMP consent verification, Google Consent Mode v2, Global Privacy Control (GPC), CCPA/US state opt-out, EU AI Act Article 50 transparency, Better Ads Standards…
Works with
Categories
Validate system resilience through controlled fault injection. Chaos Engineering is an agent skill from petrkindlmann/qa-skills. Validate system resilience through controlled fault injection.
Chaos Engineering fits situations like: : chaos engineering; fault injection; resilience test; failure recovery.
Run `npx skills add petrkindlmann/qa-skills --skill chaos-engineering -a claude-code`. Or copy the skill folder (skills/chaos-engineering in petrkindlmann/qa-skills) into .claude/skills/chaos-engineering in your project. Claude Code loads it when a task matches its description.
Run `npx skills add petrkindlmann/qa-skills --skill chaos-engineering -a codex`. Or copy the skill folder (skills/chaos-engineering in petrkindlmann/qa-skills) into .agents/skills/chaos-engineering in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add petrkindlmann/qa-skills --skill chaos-engineering -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/chaos-engineering, .gemini/skills/chaos-engineering, .github/skills/chaos-engineering and .opencode/skills/chaos-engineering in your project.
Going by SKILL.md and its folder, Chaos Engineering needs the command-line tools its instructions call (kubectl). Our summary lists: Docker.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Chaos Engineering is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 6.2k tokens (SKILL.md is roughly 25k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.3k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Chaos Engineering: Chaos Engineering (alirezarezvani/claude-skills, 28k stars), AWS Resilience Lifecycle (aws/agent-toolkit-for-aws, 2.8k stars), AWS Fault Injection Service (aws/agent-toolkit-for-aws, 2.8k stars) and Cloud Cost Optimization (wshobson/agents, 40k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
petrkindlmann (a GitHub user) maintains it in petrkindlmann/qa-skills, which has 170 GitHub stars. The repository holds 45 skills in this directory. The repository was last updated on June 10, 2026.
Source: petrkindlmann/qa-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.