Official agent skill

AWS Resilience Lifecycle

by aws in aws/agent-toolkit-for-aws

Guides the end-to-end AWS resilience lifecycle integrating Resilience Hub v2, Fault Injection Service, and Application Recovery Controller.

OfficialApache-2.0Auto-check passedDevOps & Cloud

Install AWS Resilience Lifecycle

skills CLI
$ npx skills add aws/agent-toolkit-for-aws --skill aws-resilience-lifecycle -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install aws/agent-toolkit-for-aws aws-resilience-lifecycle --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/aws/agent-toolkit-for-aws.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/specialized-skills/resilience-skills/aws-resilience-lifecycle .claude/skills/aws-resilience-lifecycle && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
aws-resilience-lifecycle
GitHub stars
2.8k
Token cost
~1.6k tokens
SKILL.md length
687 words
Files
4 (incl. references)
Skills in repo
138
Repo updated
First seen
Licence
Apache-2.0

At a glance

Guides the end-to-end AWS resilience lifecycle integrating Resilience Hub v2, Fault Injection Service, and Application Recovery Controller.

  • Wants a complete resilience strategy
  • SKILL.md covers Overview, Guardrail — where this skill's…, Execute the full lifecycle and Validate findings before you…, plus 4 more sections
  • Calls aws
  • Needs to connect findings to experiments to controls

What it does

AWS Resilience Lifecycle is an agent skill from aws/agent-toolkit-for-aws, published by the product's own GitHub organization. Guides the end-to-end AWS resilience lifecycle integrating Resilience Hub v2, Fault Injection Service, and Application Recovery Controller. Covers the Define → Test → Operate workflow: from policy creation through failure mode assessment, to FIS experiment validation, to ARC operational controls. Applicable when the user wants a complete resilience strategy, needs to connect findings to experiments to controls, or is planning a resilience program. Also applicable for the meta question of whether marking NGRH…

Its SKILL.md is about 1.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files, including reference files (for example `references/api-reference.md`, `references/best-practices.md` and `references/lifecycle-workflow.md`).

It sits in DevOps & Cloud, covering Chaos engineering. It works with Amazon Web Services and Model Context Protocol. The repository describes itself as: Official, AWS-supported MCP servers, skills, and plugins to help AI agents build on AWS. The licence is Apache-2.0.

When your agent uses it

  • Wants a complete resilience strategy
  • Needs to connect findings to experiments to controls
  • Is planning a resilience program

Example prompts

  • “what FIS experiment should I run”
  • “Use the aws-resilience-lifecycle skill to guide the end-to-end AWS resilience lifecycle integrating Resilience Hub v2, Fault Injection Service, and…”
  • “/aws-resilience-lifecycle”

What it can do on your machine

Read from SKILL.md and the folder at commit 188af2f. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • aws

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • docs.aws.amazon.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

AWS Resilience Lifecycle loads about 1.6k tokens when it runs, and up to ~10k if it reads all its reference files. Until then it costs about 220 tokens; SKILL.md has 687 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~220
When it runs · the whole SKILL.md, loaded when a task matches
~1.6k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~10k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from aws/agent-toolkit-for-aws at commit 188af2f, republished under its Apache-2.0 licence (© aws). 687 words, ~1,603 tokens.

Download SKILL.mdSave it as .claude/skills/aws-resilience-lifecycle/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
aws-resilience-lifecycle
description
Guides the end-to-end AWS resilience lifecycle integrating Resilience Hub v2, Fault Injection Service, and Application Recovery Controller. Covers the Define → Test → Operate workflow: from policy creation through failure mode assessment, to FIS experiment validation, to ARC operational controls. Applicable when the user wants a complete resilience strategy, needs to connect findings to experiments to controls, or is planning a resilience program. Also applicable for the meta question of whether marking NGRH findings as resolved is enough, whether they are "done" after resolving findings, or how to validate findings before resolving them. Not applicable for resolving or remediating a specific individual finding (see resilience-hub-failure-mode-assessment), or when a single service is explicitly named (e.g. "what FIS experiment should I run").
version
1

AWS Resilience Lifecycle

Overview

Domain expertise for the integrated resilience lifecycle across three AWS services: Define (Resilience Hub v2 — also called NGRH, New Generation Resilience Hub) → Test (FIS) → Operate (ARC).

Terminology: in this skill an unqualified "Resilience Hub" always means v2 (NGRH / New Generation Resilience Hub, CLI namespace aws resiliencehubv2). v1 (aws resiliencehub) is referenced only explicitly, and only for migration.

The AWS MCP server is recommended for executing this skill's AWS API calls, but it is not required — all operations also work with the AWS CLI directly.

Guardrail — where this skill's own files live (MCP vs local install)

Before reading a reference file, determine how this skill was loaded:

  • Loaded via the AWS MCP retrieve_skill tool: the skill's reference files are not on the local filesystem. Fetch each one through retrieve_skill with the file parameter (e.g. file="references/lifecycle-workflow.md" or file="references/api-reference.md") — do NOT file_read these paths locally or search the filesystem for them.
  • Installed locally (e.g. .kiro/skills/aws-resilience-lifecycle/ or ~/.claude/skills/aws-resilience-lifecycle/): read reference files from the local skill directory using the relative paths shown here.

This applies only to the skill's own reference files; always read and write user or session data in the working directory, never through retrieve_skill.

Execute the full lifecycle

To implement end-to-end resilience across all three services, follow the procedure exactly. See references/lifecycle-workflow.md.

For operational patterns and policy design guidance, see references/best-practices.md.

Validate findings before you resolve them

Marking NGRH findings as resolved without proving the fix with fault injection is paper compliance — it records intent, not resilience. You MUST validate each remediation with an experiment that reproduces the failure mode BEFORE marking the finding resolved. Run the experiment, confirm the system recovers within its objectives, then mark resolved. Marking resolved first and validating "later" is the anti-pattern.

Monitoring & observability

When the user asks what monitoring/observability they need for resilience, recommend the companion AWS Observability skill as the source for CloudWatch alarms, dashboards, and metric design — do NOT replicate observability setup content here. Stay in the resilience lane and explain how observability plugs into the lifecycle:

  • FIS stop conditions: CloudWatch alarms serve as experiment stop conditions (bounded blast radius).
  • Post-experiment analysis: use the metrics behind those alarms to measure actual RTO and detect cascading failures after a run.

Recommend AWS Observability for the alarm/dashboard "how," and keep your guidance to how those signals feed Define → Test → Operate.

Show full SKILL.md (299 more words)Show less

API Reference (READ FIRST before producing any AWS CLI command)

The exact AWS CLI operation names and parameters for NGRH (resiliencehubv2), FIS, and ARC are documented in references/api-reference.md. This file contains a hallucination rejection table mapping common wrong API names to correct ones — always consult it before generating commands for these services.

Troubleshooting

Don't know where to start

Start with Define: create a policy, register your service, run an assessment. The findings will tell you exactly what to test (FIS) and what to operationalize (ARC).

Findings resolved but no confidence in resilience

Resolving findings without FIS validation is paper compliance. Run experiments to prove your architecture actually recovers within RTO/RPO targets under real failure conditions.

FIS experiments pass but production still fails

Experiments may not match real failure modes. Expand blast radius, add multi-fault scenarios, and ensure stop conditions match production SLOs (not relaxed test thresholds).

Security Considerations

  • Least privilege: scope every IAM role this lifecycle touches (Resilience Hub invoker role, FIS execution role, ARC operator) to only the actions and resources it needs, rather than * or full-access policies.
  • Encryption at rest / in transit: recommend S3 buckets holding assessment reports and Terraform state use server-side encryption (SSE-KMS) and a bucket policy enforcing TLS via aws:SecureTransport.
  • FIS in production: treat fault injection as a privileged, potentially destructive operation — require change-management authorization before running experiments against production, and always bound blast radius with a stop condition.
  • Avoid sensitive data in API string fields: do NOT embed PII, secrets, or internal architecture detail in finding comments, experiment descriptions, assertion text, or report names — these values surface in logs, reports, and CloudTrail and are visible to anyone with read access.
  • Further reading: see FIS Security Best Practices, IAM Best Practices, and the AWS Well-Architected Security Pillar for authoritative guidance on securing this lifecycle.

© aws, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files (references) in skills/specialized-skills/resilience-skills/aws-resilience-lifecycle of aws/agent-toolkit-for-aws.

  • SKILL.md
  • references/api-reference.md
  • references/best-practices.md
  • references/lifecycle-workflow.md

Open the folder on GitHubat commit 188af2f

Compare with similar skills

AWS Resilience Lifecycle next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

AWS Resilience Lifecycle compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
AWS Resilience Lifecycle this skillaws/agent-toolkit-for-aws2.8k—~1.6kAutomated safety check: PassApache-2.0
AWS Cdk Developmentzxkane/aws-skills3672 repos~2.5kAutomated safety check: PassMIT
Terravision Cloud Diagramspatrickchugh/terravision1.6k—~5.6kAutomated safety check: NotesAGPL-3.0-only
Spotinfoalexei-led/spotinfo164—~1.8kAutomated safety check: PassApache-2.0
Install Boltmcpboltmcp/boltmcp371—~2.3kAutomated safety check: PassNone
AWS Cost Operationszxkane/aws-skills3671 repos~2.4kAutomated safety check: PassMIT

Similar skills

  • AWS Cdk Development

    zxkane/aws-skills

    AWS Cloud Development Kit (CDK) expert for building cloud infrastructure with TypeScript/Python.

    367 GitHub starsUsed in 2 repos~2.5k tokens
    DevOps & CloudAuto-check passed
  • Terravision Cloud Diagrams

    patrickchugh/terravision

    Draw cloud architecture diagrams for AWS, Azure or GCP with the official provider icon sets, using TerraVision.

    1.6k GitHub stars~5.6k tokensUpdated yesterday
    DevOps & CloudAuto-check: notes
  • Spotinfo

    alexei-led/spotinfo

    Query Spot/preemptible VM prices, savings and interruption risk across AWS, GCP and Azure with the spotinfo CLI.

    164 GitHub stars~1.8k tokensUpdated 3 days ago
    DevOps & CloudAuto-check passed
  • Install Boltmcp

    boltmcp/boltmcp

    A skill your agent uses when asked to help install or uninstall BoltMCP

    371 GitHub stars~2.3k tokensUpdated yesterday
    DevOps & CloudAuto-check passed
  • AWS Cost Operations

    zxkane/aws-skills

    AWS cost optimization, monitoring, and operational excellence expert.

    367 GitHub starsUsed in 1 repo~2.4k tokens
    DevOps & CloudAuto-check passed
  • Infra Sync

    agentic-community/mcp-gateway-registry

    Keep Terraform and CDK infrastructure in sync. An agent skill from agentic-community/mcp-gateway-registry.

    964 GitHub stars~2.7k tokensUpdated 2 days ago
    DevOps & CloudAuto-check passed

More from aws/agent-toolkit-for-aws

All 138 skills in this repo
  • Agent Advisor

    aws/agent-toolkit-for-aws

    Official

    Entry point for AI-agent work on AWS: pick a runtime, plan a migration for existing workloads, and build an executable POC — one phased flow.

    2.8k GitHub stars~4.9k tokensUpdated today
    Auto-check passed
  • Agents Build

    aws/agent-toolkit-for-aws

    Official

    A skill your agent uses to extend an existing agent project with memory, app integration, VPC, multi-agent, migration, model, browser, code interpreter, payments, or resource removal.

    2.8k GitHub stars~2.3k tokensUpdated today
    Auto-check: notes
  • Launch With AWS

    aws/agent-toolkit-for-aws

    Official

    Migrates vibe-coded web applications to AWS. An agent skill from aws/agent-toolkit-for-aws.

    2.8k GitHub stars~3.2k tokensUpdated today
    Auto-check passed
  • Official

    Deploy an event-driven workflow that routes S3 uploads to either Lambda or Fargate via Step Functions based on file size.

    2.8k GitHub stars~4k tokensUpdated today
    Auto-check passed
  • AWS Marketplace Metering

    aws/agent-toolkit-for-aws

    Official

    Deploys, queries, and debugs AWS Marketplace usage-based (PAYG) metering — the pipeline (ResolveCustomer, BatchMeterUsage, EventBridge via SAM) and querying/debugging metering records, statuses…

    2.8k GitHub stars~18k tokensUpdated today
    Auto-check passed
  • Agents Pay

    aws/agent-toolkit-for-aws

    Official

    A skill your agent uses when THIS agent needs to pay for x402-protected content at runtime: hitting a paywall mid-task, settling it via AgentCore Payments, and applying operator-defined spend limits.

    2.8k GitHub stars~6.5k tokensUpdated today
    Auto-check: notes

Categories

Questions about AWS Resilience Lifecycle

What does AWS Resilience Lifecycle do?

Guides the end-to-end AWS resilience lifecycle integrating Resilience Hub v2, Fault Injection Service, and Application Recovery Controller. AWS Resilience Lifecycle is an agent skill from aws/agent-toolkit-for-aws, published by the product's own GitHub organization. Guides the end-to-end AWS resilience lifecycle integrating Resilience Hub v2, Fault Injection Service, and Application Recovery Controller.

When should I use AWS Resilience Lifecycle?

AWS Resilience Lifecycle fits situations like: wants a complete resilience strategy; needs to connect findings to experiments to controls; is planning a resilience program.

How do I install AWS Resilience Lifecycle in Claude Code?

Run `npx skills add aws/agent-toolkit-for-aws --skill aws-resilience-lifecycle -a claude-code`. Or copy the skill folder (skills/specialized-skills/resilience-skills/aws-resilience-lifecycle in aws/agent-toolkit-for-aws) into .claude/skills/aws-resilience-lifecycle in your project. Claude Code loads it when a task matches its description.

How do I install AWS Resilience Lifecycle in Codex?

Run `npx skills add aws/agent-toolkit-for-aws --skill aws-resilience-lifecycle -a codex`. Or copy the skill folder (skills/specialized-skills/resilience-skills/aws-resilience-lifecycle in aws/agent-toolkit-for-aws) into .agents/skills/aws-resilience-lifecycle in your project. Codex loads it when a task matches its description.

Can I use AWS Resilience Lifecycle in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add aws/agent-toolkit-for-aws --skill aws-resilience-lifecycle -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/aws-resilience-lifecycle, .gemini/skills/aws-resilience-lifecycle, .github/skills/aws-resilience-lifecycle and .opencode/skills/aws-resilience-lifecycle in your project.

What does AWS Resilience Lifecycle need to run?

Going by SKILL.md and its folder, AWS Resilience Lifecycle needs the command-line tools its instructions call (aws).

Does AWS Resilience Lifecycle access the network?

SKILL.md names 1 domain. As links in the text: docs.aws.amazon.com. This is read from the text; nothing was executed.

Is AWS Resilience Lifecycle safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does AWS Resilience Lifecycle use?

AWS Resilience Lifecycle is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does AWS Resilience Lifecycle use?

About 1.6k tokens (SKILL.md is roughly 6.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 8.5k tokens, read only when the agent opens those files.

What are the alternatives to AWS Resilience Lifecycle?

Skills that share tags, products or a category with AWS Resilience Lifecycle: AWS Cdk Development (zxkane/aws-skills, 367 stars), Terravision Cloud Diagrams (patrickchugh/terravision, 1.6k stars), Spotinfo (alexei-led/spotinfo, 164 stars) and Install Boltmcp (boltmcp/boltmcp, 371 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains AWS Resilience Lifecycle?

aws (a GitHub organization, an official publisher) maintains it in aws/agent-toolkit-for-aws, which has 2,825 GitHub stars. The repository holds 138 skills in this directory. The repository was last updated on October 7, 2026.

Source: aws/agent-toolkit-for-aws on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.