Debugging Lambda Timeouts
aws/agent-toolkit-for-aws
Debugs AWS Lambda function timeout failures by systematically analyzing function configuration, CloudWatch logs and metrics, VPC/networking, cold starts, memory constraints, and downstream…
ALWAYS use this skill in the beginning of any incident investigation, root cause analysis, or operational troubleshooting.
$ npx skills add aws/tools-for-devops-agent --skill aws-health-events -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install aws/tools-for-devops-agent aws-health-events --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/aws/tools-for-devops-agent.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/aws-health-events .claude/skills/aws-health-events && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "aws-health-events" agent skill from https://github.com/aws/tools-for-devops-agent/tree/main/skills/aws-health-events into .claude/skills/aws-health-events/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "aws-health-events", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/aws/tools-for-devops-agent/tree/main/skills/aws-health-eventsType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add aws/tools-for-devops-agent --skill aws-health-events -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install aws/tools-for-devops-agent aws-health-events --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/aws/tools-for-devops-agent.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/aws-health-events .agents/skills/aws-health-events && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "aws-health-events" agent skill from https://github.com/aws/tools-for-devops-agent/tree/main/skills/aws-health-events into .agents/skills/aws-health-events/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "aws-health-events", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add aws/tools-for-devops-agent --skill aws-health-events -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install aws/tools-for-devops-agent aws-health-events --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/aws/tools-for-devops-agent.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/aws-health-events .cursor/skills/aws-health-events && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "aws-health-events" agent skill from https://github.com/aws/tools-for-devops-agent/tree/main/skills/aws-health-events into .cursor/skills/aws-health-events/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "aws-health-events", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/aws/tools-for-devops-agent.git --path skills/aws-health-events--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add aws/tools-for-devops-agent --skill aws-health-events -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install aws/tools-for-devops-agent aws-health-events --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/aws/tools-for-devops-agent.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/aws-health-events .gemini/skills/aws-health-events && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "aws-health-events" agent skill from https://github.com/aws/tools-for-devops-agent/tree/main/skills/aws-health-events into .gemini/skills/aws-health-events/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "aws-health-events", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install aws/tools-for-devops-agent aws-health-eventsInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add aws/tools-for-devops-agent --skill aws-health-events -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/aws/tools-for-devops-agent.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/aws-health-events .github/skills/aws-health-events && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "aws-health-events" agent skill from https://github.com/aws/tools-for-devops-agent/tree/main/skills/aws-health-events into .github/skills/aws-health-events/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "aws-health-events", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add aws/tools-for-devops-agent --skill aws-health-events -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install aws/tools-for-devops-agent aws-health-events --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/aws/tools-for-devops-agent.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/aws-health-events .opencode/skills/aws-health-events && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "aws-health-events" agent skill from https://github.com/aws/tools-for-devops-agent/tree/main/skills/aws-health-events into .opencode/skills/aws-health-events/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "aws-health-events", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
aws-health-eventsALWAYS use this skill in the beginning of any incident investigation, root cause analysis, or operational troubleshooting.
AWS Health Events is an agent skill from aws/tools-for-devops-agent, published by the product's own GitHub organization. ALWAYS use this skill in the beginning of any incident investigation, root cause analysis, or operational troubleshooting. This skill retrieves and analyzes AWS Health events (service issues, scheduled changes, and account notifications) to identify AWS-side events that may explain or correlate with observed operational issues. Activate this skill when investigating an issue and you observe service degradation, elevated error rates, latency spikes, connection failures, throttling, capacity issues…
Its SKILL.md is about 4.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 11 other files, including reference files (for example `.skilleval.yaml`, `CHANGELOG.md` and `README.md`).
It sits in Development, covering Root cause analysis. It works with Amazon Web Services. The repository describes itself as: Open-source tools for AWS DevOps Agent - extend DevOps Agent with ready-to-use skills, custom agents, and other tools, for incident response, root cause analysis, and operational…. The licence is Apache-2.0.
7 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit ddda70b. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
awsFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use aws, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
AWS Health Events loads about 4.6k tokens when it runs, and up to ~6.4k if it reads all its reference files. Until then it costs about 234 tokens; SKILL.md has 2,002 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from aws/tools-for-devops-agent at commit ddda70b, republished under its Apache-2.0 licence (© aws). 2,002 words, ~4,579 tokens.
.claude/skills/aws-health-events/SKILL.md (or your agent's skills folder). This skill also uses 9 other files; get the full folder from GitHub.Use this skill when investigating an incident and you need to check for AWS-side service events that may be causing or contributing to the observed issue. Also use this skill when a user requests a summary report of AWS Health events over a configurable time period.
Incident Investigation (automatic activation):
Chat Reporting (on-demand activation):
health:DescribeEventshealth:DescribeEventDetailshealth:DescribeAffectedEntitieshealth:DescribeEventTypesus-east-1 endpoint regardless of where the affected
resources are located.Before searching Health events, extract key details from the current incident:
Use these details as filter criteria in subsequent steps.
Use the AWS Health API DescribeEvents operation to retrieve events matching
the incident context. All calls must target the us-east-1 endpoint.
aws health describe-events \
--region us-east-1 \
--filter '{
"services": ["<SERVICE_CODE>"],
"startTimes": [{"from": "<ISO-8601-start>"}],
"regions": ["<affected-region>"],
"eventStatusCodes": ["open", "closed"],
"eventTypeCategories": ["issue", "scheduledChange", "accountNotification"]
}' \
--max-results 100| Strategy | How to Apply |
|---|---|
| By service | Use the services filter with the AWS Health service code (e.g., EC2, RDS, ELASTICLOADBALANCING), because service-specific events are most likely to correlate with the incident. |
| By time range | Use startTimes with a from value set to 7 days before the incident start, because events that started before the incident may still be active and causing impact. |
| By region | Use the regions filter to scope events to the affected region, because regional events are more likely to impact the specific resources under investigation. |
| By availability zone | Use the availabilityZones filter when the incident is isolated to a specific AZ, because AZ-scoped events have the highest correlation with AZ-specific failures. |
| By status | Include both open and closed statuses, because recently closed events may have caused residual impact that is still being observed. |
| By event scope | Include both ACCOUNT_SPECIFIC and PUBLIC events, because public service events affect all accounts in the region while account-specific events target your resources directly. |
nextToken from each response to retrieve subsequent pages.nextToken is null or a maximum of 500 events
have been collected.maxResults to 100 per page for efficient retrieval.Before retrieving full event details, filter the events returned in Step 2 to
identify only those relevant to the current investigation. This avoids
unnecessary DescribeEventDetails calls for events that are clearly unrelated.
Evaluate each event from the DescribeEvents response using these fields
(available without calling DescribeEventDetails):
| Field | Relevance Signal |
|---|---|
service | Must match one of the affected services from the incident context, or a related service from the Service Dependency Map |
eventTypeCategory | Prioritize issue events for active incidents; include scheduledChange if the incident coincides with a maintenance window |
eventTypeCode | Match against known operational event patterns (e.g., AWS_EC2_OPERATIONAL_ISSUE, AWS_RDS_MAINTENANCE) |
statusCode | Prioritize open events; include closed only if the event ended within 2 hours of the incident start |
startTime / endTime | The event's active period must overlap with the incident timeframe |
region / availabilityZone | Must match the incident's affected region or AZ |
service matches an affected service or a related
service from the Service Dependency Map.region or availabilityZone matches the
incident's affected region/AZ.accountNotification events unless the incident context
specifically suggests an account-level issue (e.g., abuse notification,
certificate expiry).closed events that ended more than 2 hours before the incident
started (unlikely to be contributing).After filtering, proceed to Step 4 only with the relevant subset of events. If all events are filtered out, report that no relevant Health events were found and suggest alternative investigation paths (see Step 7).
For each relevant event identified in Step 3, retrieve full descriptions
and timelines using DescribeEventDetails.
aws health describe-event-details \
--region us-east-1 \
--event-arns '["<arn-1>", "<arn-2>", ..., "<arn-10>"]'latestDescription text explaining the event.successfulSet and a failedSet.failedSet, report the failed ARN and error
message to the operator.successfulSet without blocking on
failures.For events with eventScopeCode of ACCOUNT_SPECIFIC, retrieve the list of
affected resources using DescribeAffectedEntities.
Important: Only call DescribeAffectedEntities for ACCOUNT_SPECIFIC events. PUBLIC events do not return entity data.
aws health describe-affected-entities \
--region us-east-1 \
--filter '{"eventArns": ["<event-arn>"]}'
--max-results 100When the incident context includes specific resource identifiers:
nextToken).entityValue against the
incident context resource identifiers.IMPAIRED, UNIMPAIRED, UNKNOWN, PENDING) and
last updated time for each entity.DescribeAffectedEntities returns an error for a specific event ARN,
report the event ARN that failed and continue processing remaining events.Score each Health event for relevance to the current incident using the following criteria:
| Classification | Criteria | Label |
|---|---|---|
| High | Matching service + overlapping timeframe + matching affected resource (or matching region/AZ if no resource IDs available) | Likely contributing factor (if event is open) |
| Medium | Matching service + overlapping timeframe (no resource match) | Likely contributing factor (if event is open) |
| Low | Matching service only (no timeframe overlap) | Background context |
entityValue matches a
resource identifier from the incident context.If the incident context does not include specific resource identifiers, score relevance using only service, timeframe, and region/AZ factors:
Present findings in a clear, structured format organized for quick comprehension and action.
If no relevant Health events are identified:
Is this a chat-based health report request?
├── YES → Search the user-specified time period (default 30 days, max 90 days)
│ Organize results by category, service, and status
│ Present as a summary report
└── NO → Continue with incident investigation flow below
Is the affected AWS service known?
├── YES → Search events for that service within the past 7 days
│ ├── Events found → Proceed to Step 3 (Filter Relevant Events)
│ └── No events found → Broaden to related services (see Service Dependency Map)
│ ├── Events found → Proceed to Step 3
│ └── No events found → Expand time window to 14 days and retry
│ ├── Events found → Proceed to Step 3
│ └── No events found → Report no events found, suggest other investigation paths
└── NO → Search all services filtered by region and availability zone (past 7 days)
├── Events found → Proceed to Step 3
└── No events found → Expand time window to 14 days
├── Events found → Proceed to Step 3
└── No events found → Report no events found, suggest other investigation paths
Does the incident involve a specific availability zone?
├── YES → Include the AZ filter in all searches above
└── NO → Filter by region onlyWhen the initial service-specific search returns no results, broaden the search to related services that share infrastructure dependencies:
| Primary Service | Related Services to Check |
|---|---|
| ELB / ALB / NLB | EC2, VPC, Route 53 |
| RDS | EC2, EBS |
| ECS / EKS | EC2, VPC, ELB |
| Lambda | VPC, CloudWatch |
| CloudFront | S3, Route 53 |
| API Gateway | Lambda, VPC |
| ElastiCache | EC2, VPC |
| DynamoDB | VPC (if VPC endpoints used) |
| S3 | CloudFront, VPC (if VPC endpoints used) |
| Kinesis | EC2, VPC |
Search up to 3 related services when broadening. Use the Health API service
codes from the references document (e.g., ELASTICLOADBALANCING for ELB,
ROUTE53 for Route 53).
| Error Condition | Agent Behavior |
|---|---|
Missing health:Describe* permissions | Report the missing permissions and specify the required IAM actions: health:DescribeEvents, health:DescribeEventDetails, health:DescribeAffectedEntities, health:DescribeEventTypes. Provide the IAM policy snippet needed. |
| Throttling (HTTP 429) | Retry with exponential backoff: wait 1s → 2s → 4s (max 3 retries). If still throttled after 3 retries, report that the Health API is currently rate-limited and recommend trying again shortly. |
| Service error (HTTP 5xx) | Report the error code and recommend the operator check the AWS Health Dashboard directly as a fallback. |
| Timeout (30 seconds) | Cancel the request and report a timeout error. Suggest the operator check the Health Dashboard directly or retry with narrower filters. |
| Zero events found | Report that no events matched the specified filters. Confirm the search parameters used. Suggest broadening the search or checking other investigation paths. |
| Invalid time range (start > end) | Report the invalid time range error. Ask the operator to provide corrected timestamps. |
| DescribeEventDetails failedSet | Report the failed event ARNs and error messages. Continue processing events from the successfulSet. |
| DescribeAffectedEntities error | Report the event ARN for which entity retrieval failed. Continue processing remaining events. |
| Unknown service name (chat report) | Inform the user the service was not recognized. List services that have events in the requested time period. |
© aws, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 9 other files (references) in skills/aws-health-events of aws/tools-for-devops-agent.
Open the folder on GitHubat commit ddda70b
AWS Health Events next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| AWS Health Events this skillaws/tools-for-devops-agent | 103 | — | ~4.6k | Automated safety check: Pass | Apache-2.0 | |
| Debugging Lambda Timeoutsaws/agent-toolkit-for-aws | 2.8k | — | ~502 | Automated safety check: Pass | Apache-2.0 | |
| Troubleshooting Application Failuresaws/agent-toolkit-for-aws | 2.8k | — | ~334 | Automated safety check: Pass | Apache-2.0 | |
| Debugging Mwaa Workflowaws/agent-toolkit-for-aws | 2.8k | — | ~2.3k | Automated safety check: Pass | Apache-2.0 | |
| AWS Cloudformationaws/agent-toolkit-for-aws | 2.8k | — | ~3.6k | Automated safety check: Pass | Apache-2.0 | |
| RStudio Node.js Version Updaterstudio/rstudio | 5.1k | — | ~3k | Automated safety check: Pass | Custom licence |
aws/agent-toolkit-for-aws
Debugs AWS Lambda function timeout failures by systematically analyzing function configuration, CloudWatch logs and metrics, VPC/networking, cold starts, memory constraints, and downstream…
aws/agent-toolkit-for-aws
Troubleshoots failing applications by discovering and analyzing CloudWatch log groups to identify error patterns, root causes, and actionable solutions.
aws/agent-toolkit-for-aws
Diagnoses and root-causes Amazon MWAA workflow failures across Provisioned (Python DAG) and Serverless (YAML workflow) environments.
aws/agent-toolkit-for-aws
Authors, validates, and troubleshoots AWS CloudFormation templates.
rstudio/rstudio
Bumps the build-time and installed Node.js versions in the RStudio repository, uploads the binaries to S3, verifies the install and opens a PR.
Burla-Cloud/burla
Sets up an isolated Burla dev cluster per git worktree so several agents can work in parallel, and explains when to use local-dev or remote-dev.
aws/tools-for-devops-agent
A skill your agent uses for GPU training or inference clusters on SageMaker HyperPod (Slurm or EKS), ParallelCluster, or self-managed EC2/EKS GPU instances.
aws/tools-for-devops-agent
AWS Database Migration Service (DMS) operational review and troubleshooting skill.
aws/tools-for-devops-agent
Performs a comprehensive Amazon ECS operations review across the 6 review pillars (Resiliency & HA, Observability, Security, Operations, Performance, Additional Analysis) using read-only AWS APIs…
aws/tools-for-devops-agent
Comprehensive Amazon RDS and Aurora operational review aligned with the AWS Well-Architected Framework and RDS/Aurora best practices.
aws/tools-for-devops-agent
Amazon SageMaker AI Operational Review. An agent skill from aws/tools-for-devops-agent.
aws/tools-for-devops-agent
Use this skill during any incident investigation, capacity planning, or operational troubleshooting when the issue may be caused by hitting AWS service limits.
Works with
Categories
ALWAYS use this skill in the beginning of any incident investigation, root cause analysis, or operational troubleshooting. AWS Health Events is an agent skill from aws/tools-for-devops-agent, published by the product's own GitHub organization. ALWAYS use this skill in the beginning of any incident investigation, root cause analysis, or operational troubleshooting.
AWS Health Events fits situations like: tasks that involve Root cause analysis.
Run `npx skills add aws/tools-for-devops-agent --skill aws-health-events -a claude-code`. Or copy the skill folder (skills/aws-health-events in aws/tools-for-devops-agent) into .claude/skills/aws-health-events in your project. Claude Code loads it when a task matches its description.
Run `npx skills add aws/tools-for-devops-agent --skill aws-health-events -a codex`. Or copy the skill folder (skills/aws-health-events in aws/tools-for-devops-agent) into .agents/skills/aws-health-events in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add aws/tools-for-devops-agent --skill aws-health-events -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/aws-health-events, .gemini/skills/aws-health-events, .github/skills/aws-health-events and .opencode/skills/aws-health-events in your project.
Going by SKILL.md and its folder, AWS Health Events needs the command-line tools its instructions call (aws).
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
AWS Health Events is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 4.6k tokens (SKILL.md is roughly 18k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.9k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with AWS Health Events: Debugging Lambda Timeouts (aws/agent-toolkit-for-aws, 2.8k stars), Troubleshooting Application Failures (aws/agent-toolkit-for-aws, 2.8k stars), Debugging Mwaa Workflow (aws/agent-toolkit-for-aws, 2.8k stars) and AWS Cloudformation (aws/agent-toolkit-for-aws, 2.8k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
aws (a GitHub organization, an official publisher) maintains it in aws/tools-for-devops-agent, which has 103 GitHub stars. The repository holds 31 skills in this directory. The repository was last updated on October 9, 2026.
Source: aws/tools-for-devops-agent on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.