Official agent skill

AWS Cloudwatch Investigation

by github in github/awesome-copilot

Reusable investigation patterns for AWS CloudWatch: Logs Insights query templates, alarm-to-deployment correlation, blast-radius narrowing decision tree, and PromQL-style metric query patterns for…

OfficialMITAuto-check passedDevOps & Cloud

Install AWS Cloudwatch Investigation

skills CLI
$ npx skills add github/awesome-copilot --skill aws-cloudwatch-investigation -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install github/awesome-copilot aws-cloudwatch-investigation --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/github/awesome-copilot.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/aws-cloudwatch-investigation .claude/skills/aws-cloudwatch-investigation && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
aws-cloudwatch-investigation
GitHub stars
40k
Token cost
~2.6k tokens
SKILL.md length
511 words
Files
1
Skills in repo
417
Repo updated
First seen
Licence
MIT

At a glance

Reusable investigation patterns for AWS CloudWatch: Logs Insights query templates, alarm-to-deployment correlation, blast-radius narrowing decision tree, and PromQL-style metric query patterns for…

  • Works in 4 steps: Get alarm transition time — note the… → Query CloudTrail for deployment-related… → Correlation criteria — a deploy is… → …
  • Tasks that involve Deployment
  • SKILL.md covers Pattern 1: Logs Insights Query…, Pattern 2: Alarm History to…, Pattern 3: Narrow the Blast… and Pattern 4: PromQL-Style Metric…, plus 1 more section
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

AWS Cloudwatch Investigation is an agent skill from github/awesome-copilot, published by the product's own GitHub organization. Reusable investigation patterns for AWS CloudWatch: Logs Insights query templates, alarm-to-deployment correlation, blast-radius narrowing decision tree, and PromQL-style metric query patterns for structured incident triage.

Its SKILL.md is about 2.6k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in DevOps & Cloud, covering Deployment. It works with Amazon Web Services and Prometheus. The repository describes itself as: Community-contributed instructions, agents, skills, and configurations to help you make the most of GitHub Copilot. The licence is MIT.

When your agent uses it

  • Tasks that involve Deployment

Example prompts

  • “/aws-cloudwatch-investigation”

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Get alarm transition time — note the exact timestamp when the alarm entered ALARM state.
  2. Query CloudTrail for deployment-related events in a window of [alarm_time - 30min, alarm_time]
  3. Correlation criteria — a deploy is "correlated" if
  4. Strengthening the correlation

What it can do on your machine

Read from SKILL.md and the folder at commit 727ff2e. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

AWS Cloudwatch Investigation loads about 2.6k tokens when it runs. Until then it costs about 63 tokens; SKILL.md has 511 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~63
When it runs · the whole SKILL.md, loaded when a task matches
~2.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from github/awesome-copilot at commit 727ff2e, republished under its MIT licence (© github). 511 words, ~2,633 tokens.

Download SKILL.mdSave it as .claude/skills/aws-cloudwatch-investigation/SKILL.md (or your agent's skills folder).
name
aws-cloudwatch-investigation
description
Reusable investigation patterns for AWS CloudWatch: Logs Insights query templates, alarm-to-deployment correlation, blast-radius narrowing decision tree, and PromQL-style metric query patterns for structured incident triage.

AWS CloudWatch Investigation Skill

Reusable patterns for investigating production incidents using CloudWatch Logs, Metrics, and Alarms. These patterns are designed to be composed together during incident triage.


Pattern 1: Logs Insights Query Templates

Error Spike Detection

Find the top errors in a time window, grouped by error type:

fields @timestamp, @message, @logStream
| filter @message like /(?i)(error|exception|fatal|critical)/
| stats count(*) as errorCount by bin(5m), @logStream
| sort errorCount desc
| limit 20
P99 Latency Breakdown by Operation

Identify which operations are driving latency spikes:

fields @timestamp, @duration, operation
| filter ispresent(@duration)
| stats avg(@duration) as avgMs,
        pct(@duration, 50) as p50Ms,
        pct(@duration, 95) as p95Ms,
        pct(@duration, 99) as p99Ms,
        count(*) as invocations
  by operation
| sort p99Ms desc
| limit 15
Lambda Cold Start Detection

Quantify cold start impact during an incident:

fields @timestamp, @duration, @initDuration, @memorySize, @maxMemoryUsed
| filter ispresent(@initDuration)
| stats count(*) as coldStarts,
        avg(@initDuration) as avgInitMs,
        max(@initDuration) as maxInitMs,
        avg(@duration) as avgDurationMs
  by bin(5m)
| sort @timestamp desc
Out-of-Memory (OOM) Detection

Find Lambda functions or containers killed by memory pressure:

fields @timestamp, @message, @logStream, @memorySize, @maxMemoryUsed
| filter @message like /Runtime exited|out of memory|OOMKilled|Cannot allocate memory|MemoryError/
| stats count(*) as oomEvents by @logStream, bin(10m)
| sort oomEvents desc
| limit 10

For memory utilization trending before OOM:

fields @timestamp, @maxMemoryUsed, @memorySize
| filter ispresent(@maxMemoryUsed)
| stats max(@maxMemoryUsed / @memorySize * 100) as peakMemPct,
        avg(@maxMemoryUsed / @memorySize * 100) as avgMemPct
  by bin(5m)
| sort @timestamp desc
Timeout Detection

Find invocations that hit the configured timeout:

fields @timestamp, @duration, @logStream, @requestId
| filter @message like /Task timed out/ or @duration > 28000
| stats count(*) as timeouts by @logStream, bin(5m)
| sort timeouts desc

Pattern 2: Alarm History to Deploy-Event Correlation

Process
  1. Get alarm transition time — note the exact timestamp when the alarm entered ALARM state.
  2. Query CloudTrail for deployment-related events in a window of [alarm_time - 30min, alarm_time]:
# CloudTrail Lake query for deployment events
SELECT eventTime, eventName, userIdentity.arn, requestParameters
FROM <event-data-store-id>
WHERE eventTime > '<alarm_time_minus_30m>'
  AND eventTime < '<alarm_time>'
  AND eventName IN (
    'UpdateFunctionCode', 'UpdateFunctionConfiguration',
    'UpdateService', 'CreateDeployment', 'RegisterTaskDefinition',
    'CreateChangeSet', 'ExecuteChangeSet',
    'StartPipelineExecution', 'PutImage'
  )
ORDER BY eventTime DESC
  1. Correlation criteria — a deploy is "correlated" if:

    • It targets the same service/resource as the alarm
    • It completed within 15 minutes before the alarm transition
    • The deployer identity matches a CI/CD role (not a human applying a hotfix)
  2. Strengthening the correlation:

    • Check if the same alarm was healthy in the previous deployment cycle
    • Verify no other environmental changes (scaling events, config changes) in the same window
    • Look for canary/synthetic monitor failures that started at the same time
Output Format
Deploy Correlation:
  Event: UpdateFunctionCode
  Time: 2024-03-15T14:23:07Z (12 min before alarm)
  Actor: arn:aws:sts::123456789012:assumed-role/github-actions-deploy/session
  Resource: arn:aws:lambda:us-east-1:123456789012:function:payment-processor
  Correlation: STRONG — same resource, CI/CD actor, alarm was OK prior cycle

Pattern 3: Narrow the Blast Radius Decision Tree

Use this tree to systematically scope an incident from broadest to most specific:

START
  |
  v
[1] ACCOUNT — Which account(s) show the alarm?
  |  - Check: Are alarms firing in multiple accounts?
  |  - If yes → suspect shared service (SSO, networking, shared deployment pipeline)
  |  - If no → proceed to Region
  v
[2] REGION — Which region(s) are affected?
  |  - Check: Same alarm in other regions?
  |  - If multi-region → suspect global service (IAM, Route53, S3 global)
  |  - If single-region → proceed to Service
  v
[3] SERVICE — Which service namespace shows degradation?
  |  - Check CloudWatch namespace: AWS/Lambda, AWS/ECS, AWS/ApiGateway, etc.
  |  - If multiple services → suspect shared dependency (VPC, NAT, DNS, IAM)
  |  - If single service → proceed to Operation
  v
[4] OPERATION — Which API action or function is failing?
  |  - For Lambda: which function name?
  |  - For ECS: which service/task definition?
  |  - For API GW: which stage/resource/method?
  |  - If all operations → suspect service-level issue (throttling, quota)
  |  - If specific operation → proceed to Resource
  v
[5] RESOURCE — Which specific resource instance?
     - Function ARN, Task ID, DB instance identifier
     - This is your investigation target
     - Proceed to log and trace analysis scoped to this resource
Shared Dependency Investigation

When blast radius spans multiple services, investigate in this order:

  1. VPC/Networking — NAT Gateway ErrorPortAllocation, packet drops, DNS resolution failures
  2. IAM/STS — ThrottlingException on AssumeRole, token vending latency
  3. Downstream dependency — shared database, cache, or external API
  4. Deployment pipeline — simultaneous deploys across services from same pipeline run
  5. AWS service event — check AWS Health Dashboard and Service Health for the region

Show full SKILL.md (211 more words)Show less

Pattern 4: PromQL-Style Metric Query Patterns

These patterns use CloudWatch metric math and GetMetricData to build composite signals. Express them as metric queries for dashboards or programmatic retrieval.

Error Rate as Percentage
MetricDataQueries:
  - Id: errors
    MetricStat:
      Metric:
        Namespace: AWS/Lambda
        MetricName: Errors
        Dimensions: [{Name: FunctionName, Value: TARGET}]
      Period: 60
      Stat: Sum
  - Id: invocations
    MetricStat:
      Metric:
        Namespace: AWS/Lambda
        MetricName: Invocations
        Dimensions: [{Name: FunctionName, Value: TARGET}]
      Period: 60
      Stat: Sum
  - Id: error_rate
    Expression: "errors / invocations * 100"
    Label: "Error Rate %"
Latency Anomaly Detection (Compare to Baseline)
MetricDataQueries:
  - Id: current_p99
    MetricStat:
      Metric:
        Namespace: AWS/Lambda
        MetricName: Duration
        Dimensions: [{Name: FunctionName, Value: TARGET}]
      Period: 300
      Stat: p99
  - Id: baseline_p99
    MetricStat:
      Metric:
        Namespace: AWS/Lambda
        MetricName: Duration
        Dimensions: [{Name: FunctionName, Value: TARGET}]
      Period: 300
      Stat: p99
    # Use StartTime/EndTime set to same window last week
  - Id: anomaly_ratio
    Expression: "current_p99 / baseline_p99"
    Label: "Latency vs Baseline (ratio > 2 = anomaly)"
Throttling Pressure Score

Combine multiple throttling signals into a single pressure metric:

MetricDataQueries:
  - Id: lambda_throttles
    MetricStat:
      Metric: {Namespace: AWS/Lambda, MetricName: Throttles}
      Period: 60
      Stat: Sum
  - Id: api_gw_429s
    MetricStat:
      Metric: {Namespace: AWS/ApiGateway, MetricName: 4XXError, Dimensions: [{Name: ApiName, Value: TARGET}]}
      Period: 60
      Stat: Sum
  - Id: dynamo_throttles
    MetricStat:
      Metric: {Namespace: AWS/DynamoDB, MetricName: ThrottledRequests, Dimensions: [{Name: TableName, Value: TARGET}]}
      Period: 60
      Stat: Sum
  - Id: throttle_pressure
    Expression: "lambda_throttles + api_gw_429s + dynamo_throttles"
    Label: "Combined Throttle Pressure"
Concurrent Execution Headroom
MetricDataQueries:
  - Id: concurrent
    MetricStat:
      Metric: {Namespace: AWS/Lambda, MetricName: ConcurrentExecutions}
      Period: 60
      Stat: Maximum
  - Id: headroom
    Expression: "1000 - concurrent"
    Label: "Remaining Concurrency (account limit 1000)"

Pattern 5: Incident Timeline Reconstruction

Process

Reconstruct a precise timeline by merging data from multiple sources:

  1. Collect timestamps:
SourceQueryYields
CloudWatch AlarmsAlarm history APIState transition times
CloudWatch MetricsGetMetricData with 1-min periodFirst anomaly point
CloudWatch LogsLogs Insights with earliest(@timestamp)First error occurrence
CloudTrailLookupEvents filtered by timeDeployment/change events
AWS HealthDescribeEventsAWS-side incidents
  1. Build the timeline:
fields @timestamp, @message
| filter @message like /ERROR|WARN|timeout|refused|denied/
| stats earliest(@timestamp) as firstSeen, latest(@timestamp) as lastSeen, count(*) as occurrences
  by @message
| sort firstSeen asc
| limit 20
  1. Identify the sequence:
Timeline:
  T-15m: CloudTrail — UpdateFunctionCode by CI/CD role
  T-12m: Logs — first error "Connection refused to payments-api.internal"
  T-10m: Metrics — Error count crosses 5/min threshold
  T-8m:  Alarm — PaymentProcessorErrors enters ALARM
  T-5m:  Metrics — p99 latency spikes to 28s (timeout)
  T-0:   Current — error rate at 45%, alarm still firing
  1. Determine root event — the earliest change that preceded all symptoms. Walk backward from the first symptom to the most recent mutation (deploy, config change, scaling event, or external dependency shift).
Gotchas
  • CloudWatch metric timestamps are end-of-period. A 1-minute datapoint at 14:05 covers 14:04-14:05.
  • CloudTrail events can have up to 15-minute delivery delay. Use eventTime, not ingestion time.
  • Log group timestamps depend on the agent/SDK flush interval. Allow for 30-60s of clock skew.
  • Alarm state changes have a built-in evaluation delay (periods x evaluation periods). The actual anomaly started earlier.

© github, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/aws-cloudwatch-investigation of github/awesome-copilot.

Open the folder on GitHubat commit 727ff2e

Compare with similar skills

AWS Cloudwatch Investigation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

AWS Cloudwatch Investigation compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
AWS Cloudwatch Investigation this skillgithub/awesome-copilot40k—~2.6kAutomated safety check: PassMIT
AWS Observabilityaws/agent-toolkit-for-aws2.8k—~7.2kAutomated safety check: PassApache-2.0
AWS Cdk Developmentzxkane/aws-skills3672 repos~2.5kAutomated safety check: PassMIT
Senior DevOps Toolkitmaslennikov-ig/claude-code-orchestrator-kit2596 repos~1.1kAutomated safety check: NotesCustom licence
Ecspressokayac/ecspresso1.1k—~1.4kAutomated safety check: PassMIT
Spa Create Configsplunk/splunk-platform-automator137—~3.5kAutomated safety check: PassProprietary

Similar skills

  • AWS Observability

    aws/agent-toolkit-for-aws

    Official

    Builds, configures, debugs, and optimizes AWS observability - operator-symptom questions and detecting Omni vs classic CloudWatch.

    2.8k GitHub stars~7.2k tokensUpdated today
    DevOps & CloudAuto-check passed
  • AWS Cdk Development

    zxkane/aws-skills

    AWS Cloud Development Kit (CDK) expert for building cloud infrastructure with TypeScript/Python.

    367 GitHub starsUsed in 2 repos~2.5k tokens
    DevOps & CloudAuto-check passed
  • Senior DevOps Toolkit

    maslennikov-ig/claude-code-orchestrator-kit

    Comprehensive DevOps skill for CI/CD, infrastructure automation, containerization, and cloud platforms (AWS, GCP, Azure). Includes pipeline setup…

    259 GitHub starsUsed in 6 repos~1.1k tokens
    DevOps & CloudAuto-check: notes
  • Ecspresso

    kayac/ecspresso

    ECS deployment tool - deploy, manage, and troubleshoot ECS services

    1.1k GitHub stars~1.4k tokensUpdated 2 days ago
    DevOps & CloudAuto-check passed
  • Spa Create Config

    splunk/splunk-platform-automator

    A skill your agent uses when creating or updating splunkconfig.yml, designing Splunk Enterprise lab topology, multisite IDXC, SHC layout, architecture plan before config, or AWS Terraform block for…

    137 GitHub stars~3.5k tokensUpdated yesterday
    DevOps & CloudAuto-check passed
  • AWS Agentic AI

    zxkane/aws-skills

    AWS Bedrock AgentCore comprehensive expert for deploying and managing AI agents at scale.

    367 GitHub starsUsed in 1 repo~2.5k tokens
    DevOps & CloudAuto-check passed

More from github/awesome-copilot

All 417 skills in this repo
  • Acquire Codebase Knowledge

    github/awesome-copilot

    Official

    Maps an unfamiliar codebase into seven evidence-backed documents in docs/codebase/, using a scan script and templates, for onboarding or architecture write-ups.

    40k GitHub starsUsed in 1 repo~2.3k tokens
    Auto-check passed
  • Azure Architecture Autopilot

    github/awesome-copilot

    Official

    Designs Azure infrastructure from a natural-language description, or diagrams an existing resource group, then refines the design through conversation and deploys it with Bicep.

    40k GitHub starsUsed in 1 repo~1.9k tokens
    Auto-check passed
  • Draw.io Diagram Generator

    github/awesome-copilot

    Official

    Generates, edits and validates draw.io files with correct mxGraph XML, covering flowcharts, architecture, sequence, ER and UML class diagrams.

    40k GitHub starsUsed in 1 repo~4.9k tokens
    Auto-check passed
  • Credit Risk Data Cleaning

    github/awesome-copilot

    Official

    Cleans raw credit data and screens variables before loan modeling, dropping unstable, noisy or redundant features and writing an Excel report of every step.

    40k GitHub starsUsed in 1 repo~1.5k tokens
    Auto-check passed
  • Daily Focus Board

    github/awesome-copilot

    Official

    Builds a warm, browser-based daily focus board the user updates by talking to their agent, with Eisenhower priorities, a brain-dump box and kind not-today carryover.

    40k GitHub stars~3k tokensUpdated today
    Auto-check passed
  • Python Pypi Package Builder

    github/awesome-copilot

    Official

    End-to-end skill for building, testing, linting, versioning, and publishing a production-grade Python library to PyPI.

    40k GitHub starsUsed in 1 repo~4.6k tokens
    Auto-check passed

Categories

Questions about AWS Cloudwatch Investigation

What does AWS Cloudwatch Investigation do?

Reusable investigation patterns for AWS CloudWatch: Logs Insights query templates, alarm-to-deployment correlation, blast-radius narrowing decision tree, and PromQL-style metric query patterns for…. AWS Cloudwatch Investigation is an agent skill from github/awesome-copilot, published by the product's own GitHub organization. Reusable investigation patterns for AWS CloudWatch: Logs Insights query templates, alarm-to-deployment correlation, blast-radius narrowing decision tree, and PromQL-style metric query patterns for structured incident triage.

When should I use AWS Cloudwatch Investigation?

AWS Cloudwatch Investigation fits situations like: tasks that involve Deployment.

How do I install AWS Cloudwatch Investigation in Claude Code?

Run `npx skills add github/awesome-copilot --skill aws-cloudwatch-investigation -a claude-code`. Or copy the skill folder (skills/aws-cloudwatch-investigation in github/awesome-copilot) into .claude/skills/aws-cloudwatch-investigation in your project. Claude Code loads it when a task matches its description.

How do I install AWS Cloudwatch Investigation in Codex?

Run `npx skills add github/awesome-copilot --skill aws-cloudwatch-investigation -a codex`. Or copy the skill folder (skills/aws-cloudwatch-investigation in github/awesome-copilot) into .agents/skills/aws-cloudwatch-investigation in your project. Codex loads it when a task matches its description.

Can I use AWS Cloudwatch Investigation in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add github/awesome-copilot --skill aws-cloudwatch-investigation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/aws-cloudwatch-investigation, .gemini/skills/aws-cloudwatch-investigation, .github/skills/aws-cloudwatch-investigation and .opencode/skills/aws-cloudwatch-investigation in your project.

What does AWS Cloudwatch Investigation need to run?

SKILL.md names no scripts, command-line tools or credentials: AWS Cloudwatch Investigation is instructions for the agent only.

Does AWS Cloudwatch Investigation access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is AWS Cloudwatch Investigation safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does AWS Cloudwatch Investigation use?

AWS Cloudwatch Investigation is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does AWS Cloudwatch Investigation use?

About 2.6k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to AWS Cloudwatch Investigation?

Skills that share tags, products or a category with AWS Cloudwatch Investigation: AWS Observability (aws/agent-toolkit-for-aws, 2.8k stars), AWS Cdk Development (zxkane/aws-skills, 367 stars), Senior DevOps Toolkit (maslennikov-ig/claude-code-orchestrator-kit, 259 stars) and Ecspresso (kayac/ecspresso, 1.1k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains AWS Cloudwatch Investigation?

github (a GitHub organization, an official publisher) maintains it in github/awesome-copilot, which has 39,748 GitHub stars. The repository holds 417 skills in this directory. The repository was last updated on October 7, 2026.

Source: github/awesome-copilot on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.