AWS Observability
aws/agent-toolkit-for-aws
Builds, configures, debugs, and optimizes AWS observability - operator-symptom questions and detecting Omni vs classic CloudWatch.
Reusable investigation patterns for AWS CloudWatch: Logs Insights query templates, alarm-to-deployment correlation, blast-radius narrowing decision tree, and PromQL-style metric query patterns for…
$ npx skills add github/awesome-copilot --skill aws-cloudwatch-investigation -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install github/awesome-copilot aws-cloudwatch-investigation --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/github/awesome-copilot.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/aws-cloudwatch-investigation .claude/skills/aws-cloudwatch-investigation && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "aws-cloudwatch-investigation" agent skill from https://github.com/github/awesome-copilot/tree/main/skills/aws-cloudwatch-investigation into .claude/skills/aws-cloudwatch-investigation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "aws-cloudwatch-investigation", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/github/awesome-copilot/tree/main/skills/aws-cloudwatch-investigationType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add github/awesome-copilot --skill aws-cloudwatch-investigation -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install github/awesome-copilot aws-cloudwatch-investigation --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/github/awesome-copilot.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/aws-cloudwatch-investigation .agents/skills/aws-cloudwatch-investigation && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "aws-cloudwatch-investigation" agent skill from https://github.com/github/awesome-copilot/tree/main/skills/aws-cloudwatch-investigation into .agents/skills/aws-cloudwatch-investigation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "aws-cloudwatch-investigation", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add github/awesome-copilot --skill aws-cloudwatch-investigation -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install github/awesome-copilot aws-cloudwatch-investigation --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/github/awesome-copilot.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/aws-cloudwatch-investigation .cursor/skills/aws-cloudwatch-investigation && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "aws-cloudwatch-investigation" agent skill from https://github.com/github/awesome-copilot/tree/main/skills/aws-cloudwatch-investigation into .cursor/skills/aws-cloudwatch-investigation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "aws-cloudwatch-investigation", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/github/awesome-copilot.git --path skills/aws-cloudwatch-investigation--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add github/awesome-copilot --skill aws-cloudwatch-investigation -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install github/awesome-copilot aws-cloudwatch-investigation --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/github/awesome-copilot.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/aws-cloudwatch-investigation .gemini/skills/aws-cloudwatch-investigation && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "aws-cloudwatch-investigation" agent skill from https://github.com/github/awesome-copilot/tree/main/skills/aws-cloudwatch-investigation into .gemini/skills/aws-cloudwatch-investigation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "aws-cloudwatch-investigation", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install github/awesome-copilot aws-cloudwatch-investigationInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add github/awesome-copilot --skill aws-cloudwatch-investigation -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/github/awesome-copilot.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/aws-cloudwatch-investigation .github/skills/aws-cloudwatch-investigation && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "aws-cloudwatch-investigation" agent skill from https://github.com/github/awesome-copilot/tree/main/skills/aws-cloudwatch-investigation into .github/skills/aws-cloudwatch-investigation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "aws-cloudwatch-investigation", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add github/awesome-copilot --skill aws-cloudwatch-investigation -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install github/awesome-copilot aws-cloudwatch-investigation --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/github/awesome-copilot.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/aws-cloudwatch-investigation .opencode/skills/aws-cloudwatch-investigation && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "aws-cloudwatch-investigation" agent skill from https://github.com/github/awesome-copilot/tree/main/skills/aws-cloudwatch-investigation into .opencode/skills/aws-cloudwatch-investigation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "aws-cloudwatch-investigation", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
aws-cloudwatch-investigationReusable investigation patterns for AWS CloudWatch: Logs Insights query templates, alarm-to-deployment correlation, blast-radius narrowing decision tree, and PromQL-style metric query patterns for…
AWS Cloudwatch Investigation is an agent skill from github/awesome-copilot, published by the product's own GitHub organization. Reusable investigation patterns for AWS CloudWatch: Logs Insights query templates, alarm-to-deployment correlation, blast-radius narrowing decision tree, and PromQL-style metric query patterns for structured incident triage.
Its SKILL.md is about 2.6k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in DevOps & Cloud, covering Deployment. It works with Amazon Web Services and Prometheus. The repository describes itself as: Community-contributed instructions, agents, skills, and configurations to help you make the most of GitHub Copilot. The licence is MIT.
4 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 727ff2e. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md.
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
AWS Cloudwatch Investigation loads about 2.6k tokens when it runs. Until then it costs about 63 tokens; SKILL.md has 511 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from github/awesome-copilot at commit 727ff2e, republished under its MIT licence (© github). 511 words, ~2,633 tokens.
.claude/skills/aws-cloudwatch-investigation/SKILL.md (or your agent's skills folder).Reusable patterns for investigating production incidents using CloudWatch Logs, Metrics, and Alarms. These patterns are designed to be composed together during incident triage.
Find the top errors in a time window, grouped by error type:
fields @timestamp, @message, @logStream
| filter @message like /(?i)(error|exception|fatal|critical)/
| stats count(*) as errorCount by bin(5m), @logStream
| sort errorCount desc
| limit 20Identify which operations are driving latency spikes:
fields @timestamp, @duration, operation
| filter ispresent(@duration)
| stats avg(@duration) as avgMs,
pct(@duration, 50) as p50Ms,
pct(@duration, 95) as p95Ms,
pct(@duration, 99) as p99Ms,
count(*) as invocations
by operation
| sort p99Ms desc
| limit 15Quantify cold start impact during an incident:
fields @timestamp, @duration, @initDuration, @memorySize, @maxMemoryUsed
| filter ispresent(@initDuration)
| stats count(*) as coldStarts,
avg(@initDuration) as avgInitMs,
max(@initDuration) as maxInitMs,
avg(@duration) as avgDurationMs
by bin(5m)
| sort @timestamp descFind Lambda functions or containers killed by memory pressure:
fields @timestamp, @message, @logStream, @memorySize, @maxMemoryUsed
| filter @message like /Runtime exited|out of memory|OOMKilled|Cannot allocate memory|MemoryError/
| stats count(*) as oomEvents by @logStream, bin(10m)
| sort oomEvents desc
| limit 10For memory utilization trending before OOM:
fields @timestamp, @maxMemoryUsed, @memorySize
| filter ispresent(@maxMemoryUsed)
| stats max(@maxMemoryUsed / @memorySize * 100) as peakMemPct,
avg(@maxMemoryUsed / @memorySize * 100) as avgMemPct
by bin(5m)
| sort @timestamp descFind invocations that hit the configured timeout:
fields @timestamp, @duration, @logStream, @requestId
| filter @message like /Task timed out/ or @duration > 28000
| stats count(*) as timeouts by @logStream, bin(5m)
| sort timeouts desc# CloudTrail Lake query for deployment events
SELECT eventTime, eventName, userIdentity.arn, requestParameters
FROM <event-data-store-id>
WHERE eventTime > '<alarm_time_minus_30m>'
AND eventTime < '<alarm_time>'
AND eventName IN (
'UpdateFunctionCode', 'UpdateFunctionConfiguration',
'UpdateService', 'CreateDeployment', 'RegisterTaskDefinition',
'CreateChangeSet', 'ExecuteChangeSet',
'StartPipelineExecution', 'PutImage'
)
ORDER BY eventTime DESCCorrelation criteria — a deploy is "correlated" if:
Strengthening the correlation:
Deploy Correlation:
Event: UpdateFunctionCode
Time: 2024-03-15T14:23:07Z (12 min before alarm)
Actor: arn:aws:sts::123456789012:assumed-role/github-actions-deploy/session
Resource: arn:aws:lambda:us-east-1:123456789012:function:payment-processor
Correlation: STRONG — same resource, CI/CD actor, alarm was OK prior cycleUse this tree to systematically scope an incident from broadest to most specific:
START
|
v
[1] ACCOUNT — Which account(s) show the alarm?
| - Check: Are alarms firing in multiple accounts?
| - If yes → suspect shared service (SSO, networking, shared deployment pipeline)
| - If no → proceed to Region
v
[2] REGION — Which region(s) are affected?
| - Check: Same alarm in other regions?
| - If multi-region → suspect global service (IAM, Route53, S3 global)
| - If single-region → proceed to Service
v
[3] SERVICE — Which service namespace shows degradation?
| - Check CloudWatch namespace: AWS/Lambda, AWS/ECS, AWS/ApiGateway, etc.
| - If multiple services → suspect shared dependency (VPC, NAT, DNS, IAM)
| - If single service → proceed to Operation
v
[4] OPERATION — Which API action or function is failing?
| - For Lambda: which function name?
| - For ECS: which service/task definition?
| - For API GW: which stage/resource/method?
| - If all operations → suspect service-level issue (throttling, quota)
| - If specific operation → proceed to Resource
v
[5] RESOURCE — Which specific resource instance?
- Function ARN, Task ID, DB instance identifier
- This is your investigation target
- Proceed to log and trace analysis scoped to this resourceWhen blast radius spans multiple services, investigate in this order:
These patterns use CloudWatch metric math and GetMetricData to build composite signals. Express them as metric queries for dashboards or programmatic retrieval.
MetricDataQueries:
- Id: errors
MetricStat:
Metric:
Namespace: AWS/Lambda
MetricName: Errors
Dimensions: [{Name: FunctionName, Value: TARGET}]
Period: 60
Stat: Sum
- Id: invocations
MetricStat:
Metric:
Namespace: AWS/Lambda
MetricName: Invocations
Dimensions: [{Name: FunctionName, Value: TARGET}]
Period: 60
Stat: Sum
- Id: error_rate
Expression: "errors / invocations * 100"
Label: "Error Rate %"MetricDataQueries:
- Id: current_p99
MetricStat:
Metric:
Namespace: AWS/Lambda
MetricName: Duration
Dimensions: [{Name: FunctionName, Value: TARGET}]
Period: 300
Stat: p99
- Id: baseline_p99
MetricStat:
Metric:
Namespace: AWS/Lambda
MetricName: Duration
Dimensions: [{Name: FunctionName, Value: TARGET}]
Period: 300
Stat: p99
# Use StartTime/EndTime set to same window last week
- Id: anomaly_ratio
Expression: "current_p99 / baseline_p99"
Label: "Latency vs Baseline (ratio > 2 = anomaly)"Combine multiple throttling signals into a single pressure metric:
MetricDataQueries:
- Id: lambda_throttles
MetricStat:
Metric: {Namespace: AWS/Lambda, MetricName: Throttles}
Period: 60
Stat: Sum
- Id: api_gw_429s
MetricStat:
Metric: {Namespace: AWS/ApiGateway, MetricName: 4XXError, Dimensions: [{Name: ApiName, Value: TARGET}]}
Period: 60
Stat: Sum
- Id: dynamo_throttles
MetricStat:
Metric: {Namespace: AWS/DynamoDB, MetricName: ThrottledRequests, Dimensions: [{Name: TableName, Value: TARGET}]}
Period: 60
Stat: Sum
- Id: throttle_pressure
Expression: "lambda_throttles + api_gw_429s + dynamo_throttles"
Label: "Combined Throttle Pressure"MetricDataQueries:
- Id: concurrent
MetricStat:
Metric: {Namespace: AWS/Lambda, MetricName: ConcurrentExecutions}
Period: 60
Stat: Maximum
- Id: headroom
Expression: "1000 - concurrent"
Label: "Remaining Concurrency (account limit 1000)"Reconstruct a precise timeline by merging data from multiple sources:
| Source | Query | Yields |
|---|---|---|
| CloudWatch Alarms | Alarm history API | State transition times |
| CloudWatch Metrics | GetMetricData with 1-min period | First anomaly point |
| CloudWatch Logs | Logs Insights with earliest(@timestamp) | First error occurrence |
| CloudTrail | LookupEvents filtered by time | Deployment/change events |
| AWS Health | DescribeEvents | AWS-side incidents |
fields @timestamp, @message
| filter @message like /ERROR|WARN|timeout|refused|denied/
| stats earliest(@timestamp) as firstSeen, latest(@timestamp) as lastSeen, count(*) as occurrences
by @message
| sort firstSeen asc
| limit 20Timeline:
T-15m: CloudTrail — UpdateFunctionCode by CI/CD role
T-12m: Logs — first error "Connection refused to payments-api.internal"
T-10m: Metrics — Error count crosses 5/min threshold
T-8m: Alarm — PaymentProcessorErrors enters ALARM
T-5m: Metrics — p99 latency spikes to 28s (timeout)
T-0: Current — error rate at 45%, alarm still firingeventTime, not ingestion time.© github, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in skills/aws-cloudwatch-investigation of github/awesome-copilot.
Open the folder on GitHubat commit 727ff2e
AWS Cloudwatch Investigation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| AWS Cloudwatch Investigation this skillgithub/awesome-copilot | 40k | — | ~2.6k | Automated safety check: Pass | MIT | |
| AWS Observabilityaws/agent-toolkit-for-aws | 2.8k | — | ~7.2k | Automated safety check: Pass | Apache-2.0 | |
| AWS Cdk Developmentzxkane/aws-skills | 367 | 2 repos | ~2.5k | Automated safety check: Pass | MIT | |
| Senior DevOps Toolkitmaslennikov-ig/claude-code-orchestrator-kit | 259 | 6 repos | ~1.1k | Automated safety check: Notes | Custom licence | |
| Ecspressokayac/ecspresso | 1.1k | — | ~1.4k | Automated safety check: Pass | MIT | |
| Spa Create Configsplunk/splunk-platform-automator | 137 | — | ~3.5k | Automated safety check: Pass | Proprietary |
aws/agent-toolkit-for-aws
Builds, configures, debugs, and optimizes AWS observability - operator-symptom questions and detecting Omni vs classic CloudWatch.
zxkane/aws-skills
AWS Cloud Development Kit (CDK) expert for building cloud infrastructure with TypeScript/Python.
maslennikov-ig/claude-code-orchestrator-kit
Comprehensive DevOps skill for CI/CD, infrastructure automation, containerization, and cloud platforms (AWS, GCP, Azure). Includes pipeline setup…
kayac/ecspresso
ECS deployment tool - deploy, manage, and troubleshoot ECS services
splunk/splunk-platform-automator
A skill your agent uses when creating or updating splunkconfig.yml, designing Splunk Enterprise lab topology, multisite IDXC, SHC layout, architecture plan before config, or AWS Terraform block for…
zxkane/aws-skills
AWS Bedrock AgentCore comprehensive expert for deploying and managing AI agents at scale.
github/awesome-copilot
Maps an unfamiliar codebase into seven evidence-backed documents in docs/codebase/, using a scan script and templates, for onboarding or architecture write-ups.
github/awesome-copilot
Designs Azure infrastructure from a natural-language description, or diagrams an existing resource group, then refines the design through conversation and deploys it with Bicep.
github/awesome-copilot
Generates, edits and validates draw.io files with correct mxGraph XML, covering flowcharts, architecture, sequence, ER and UML class diagrams.
github/awesome-copilot
Cleans raw credit data and screens variables before loan modeling, dropping unstable, noisy or redundant features and writing an Excel report of every step.
github/awesome-copilot
Builds a warm, browser-based daily focus board the user updates by talking to their agent, with Eisenhower priorities, a brain-dump box and kind not-today carryover.
github/awesome-copilot
End-to-end skill for building, testing, linting, versioning, and publishing a production-grade Python library to PyPI.
Works with
Categories
Reusable investigation patterns for AWS CloudWatch: Logs Insights query templates, alarm-to-deployment correlation, blast-radius narrowing decision tree, and PromQL-style metric query patterns for…. AWS Cloudwatch Investigation is an agent skill from github/awesome-copilot, published by the product's own GitHub organization. Reusable investigation patterns for AWS CloudWatch: Logs Insights query templates, alarm-to-deployment correlation, blast-radius narrowing decision tree, and PromQL-style metric query patterns for structured incident triage.
AWS Cloudwatch Investigation fits situations like: tasks that involve Deployment.
Run `npx skills add github/awesome-copilot --skill aws-cloudwatch-investigation -a claude-code`. Or copy the skill folder (skills/aws-cloudwatch-investigation in github/awesome-copilot) into .claude/skills/aws-cloudwatch-investigation in your project. Claude Code loads it when a task matches its description.
Run `npx skills add github/awesome-copilot --skill aws-cloudwatch-investigation -a codex`. Or copy the skill folder (skills/aws-cloudwatch-investigation in github/awesome-copilot) into .agents/skills/aws-cloudwatch-investigation in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add github/awesome-copilot --skill aws-cloudwatch-investigation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/aws-cloudwatch-investigation, .gemini/skills/aws-cloudwatch-investigation, .github/skills/aws-cloudwatch-investigation and .opencode/skills/aws-cloudwatch-investigation in your project.
SKILL.md names no scripts, command-line tools or credentials: AWS Cloudwatch Investigation is instructions for the agent only.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
AWS Cloudwatch Investigation is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.6k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with AWS Cloudwatch Investigation: AWS Observability (aws/agent-toolkit-for-aws, 2.8k stars), AWS Cdk Development (zxkane/aws-skills, 367 stars), Senior DevOps Toolkit (maslennikov-ig/claude-code-orchestrator-kit, 259 stars) and Ecspresso (kayac/ecspresso, 1.1k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
github (a GitHub organization, an official publisher) maintains it in github/awesome-copilot, which has 39,748 GitHub stars. The repository holds 417 skills in this directory. The repository was last updated on October 7, 2026.
Source: github/awesome-copilot on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.