Official agent skill

AWS Observability

by aws in aws/agent-toolkit-for-aws

Builds, configures, debugs, and optimizes AWS observability - operator-symptom questions and detecting Omni vs classic CloudWatch.

OfficialApache-2.0Auto-check passedDevOps & Cloud

Install AWS Observability

skills CLI
$ npx skills add aws/agent-toolkit-for-aws --skill aws-observability -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install aws/agent-toolkit-for-aws aws-observability --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/aws/agent-toolkit-for-aws.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/core-skills/aws-observability .claude/skills/aws-observability && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
aws-observability
GitHub stars
2.8k
Token cost
~7.2k tokens
SKILL.md length
3,277 words
Files
54 (incl. scripts, references, assets)
Skills in repo
138
Repo updated
First seen
Licence
Apache-2.0

At a glance

Builds, configures, debugs, and optimizes AWS observability - operator-symptom questions and detecting Omni vs classic CloudWatch.

  • Tasks that involve Observability
  • SKILL.md covers Overview, Scope guard — is this an…, Step 0 — CloudWatch or… and Routing — CloudWatch…, plus 2 more sections
  • Runs TypeScript scripts from its folder; calls aws, brew and pip
  • Tasks that involve Infrastructure as code

What it does

AWS Observability is an agent skill from aws/agent-toolkit-for-aws, published by the product's own GitHub organization. Builds, configures, debugs, and optimizes AWS observability - operator-symptom questions and detecting Omni vs classic CloudWatch. CloudWatch on an already-reporting service: Log Insights, metric/composite/anomaly alarms, custom metrics/EMF, dashboards, X-Ray/ADOT tracing, canaries, CloudTrail, Dynamic Instrumentation (live breakpoints/snapshots), the Application Signals service map, and fleet health views. CloudWatch Omni on an existing Space: SQL over logs and traces, PromQL over metrics, Omni dashboards, Omni…

Its SKILL.md is about 7.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 59 other files, including scripts, reference files and assets (for example `assets/cloudwatch/alarm-template.ts`, `assets/cloudwatch/otel-config.yaml` and `references/cloudwatch-omni/agent-evaluation.md`).

It sits in DevOps & Cloud, covering Observability, Infrastructure as code and Deployment. It works with Amazon Web Services, Prometheus, SQL and AWS CloudFormation. The repository describes itself as: Official, AWS-supported MCP servers, skills, and plugins to help AI agents build on AWS. The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve Observability
  • Tasks that involve Infrastructure as code
  • Tasks that involve Deployment

Example prompts

  • “/aws-observability”

Requirements

  • Python 3
  • Node.js

What it can do on your machine

Read from SKILL.md and the folder at commit bd49cc8. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (TypeScript, from the files we listed), which the agent can run.

    Shell commands in SKILL.md call:

    • aws
    • brew
    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • docs.aws.amazon.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

AWS Observability loads about 7.2k tokens when it runs, and up to ~158k if it reads all its reference files. Until then it costs about 260 tokens; SKILL.md has 3,277 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~260
When it runs · the whole SKILL.md, loaded when a task matches
~7.2k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~158k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from aws/agent-toolkit-for-aws at commit bd49cc8, republished under its Apache-2.0 licence (© aws). 3,277 words, ~7,207 tokens.

Download SKILL.mdSave it as .claude/skills/aws-observability/SKILL.md (or your agent's skills folder). This skill also uses 53 other files; get the full folder from GitHub.
name
aws-observability
description
Builds, configures, debugs, and optimizes AWS observability - operator-symptom questions and detecting Omni vs classic CloudWatch. CloudWatch on an already-reporting service: Log Insights, metric/composite/anomaly alarms, custom metrics/EMF, dashboards, X-Ray/ADOT tracing, canaries, CloudTrail, Dynamic Instrumentation (live breakpoints/snapshots), the Application Signals service map, and fleet health views. CloudWatch Omni on an existing Space: SQL over logs and traces, PromQL over metrics, Omni dashboards, Omni alerts, context graph for root cause, programmatic/IaC access (API/SDK/CLI/CloudFormation), driving Omni from a coding agent or skills, and evaluating AI agent quality from traces - on-demand and online scoring of live traffic, readback, custom evaluators. For first-time setup - creating an Omni Space, granting access, ingestion, or ADOT instrumentation for Application Signals (ServiceEvents, CI/CD metadata) or Omni - use setting-up-cloudwatch-observability. Not for app logging or threat detection.
metadata.version
7

AWS Observability

Overview

Domain expertise for AWS observability across metrics, logs, and traces, for two products that share the CloudWatch name but are separate services with separate control planes, data models, and APIs:

CloudWatchCloudWatch Omni
What it isLog groups, metric namespaces, alarms, Log Insights, X-Ray, Application SignalsApplication Observability / Agent Observability. A Space per account per Region is the access boundary over the account's CloudWatch Dataset (OpenTelemetry logs, traces, and metrics); the Dataset is a CloudWatch resource the Space reads, not something the Space contains
Control planeaws cloudwatch, aws logs, aws xray, aws application-signalsaws cloudwatchomni (endpoint prefix cloudwatch-omni, signing name cloudwatch)
QueryLog Insights query language; GetMetricDataSQL over logs.default / traces.default; PromQL over metrics; named views
NotifyAlarms (metric, composite, anomaly)Alerts (SQL/PromQL rule, contributors, OK/WARNING/CRITICAL/NODATA)
TopologyApplication Signals service mapContext graph (GetContextGraph)
AccessIAM onlyDomain → Space → access grants and access profiles (set up in setting-up-cloudwatch-observability)
Only hereDynamic Instrumentation, Synthetics canaries, CloudTrail auditing, EMFAgent-quality evaluation, views, context graph
First-time setupApplication Signals onboarding — ADOT auto-instrumentation, the amazon-cloudwatch-observability add-on, ServiceEvents, CI/CD metadataDomain → Space → grants → telemetry in → plain-ADOT instrumentation
Referencesreferences/cloudwatch/references/cloudwatch-omni/

This skill covers using both products, never setting either one up. Both first-time setup paths above live in setting-up-cloudwatch-observability (same two-folder split); route there rather than describing an onboarding procedure from memory.

Enabling Omni does not replace CloudWatch; log groups, metrics, and alarms keep working, and most customers use both.

Works best with the AWS MCP server — enables running CLI commands, querying CloudWatch, and validating configurations directly. All guidance also works with standard AWS CLI access.

Note: Reference files contain specific runtime versions, quota values, and feature matrices that may change. When precision matters (e.g. deploying to production, choosing a runtime, or checking a quota), confirm values against current AWS documentation rather than relying solely on the values in these files.

Scope guard — is this an observability request?

This skill owns the scope decision for every request routed to it, and this section is its sole home.

If a request is NOT about AWS observability with CloudWatch or CloudWatch Omni (weather, trivia, general chit-chat, non-observability coding), decline it in one sentence: state plainly that it is out of scope, never fabricate an answer, and never claim a false capability limitation (no "no internet access", no "no weather data") — the reason is simply that it is out of scope. Then redirect by naming what this skill does cover: querying CloudWatch and CloudWatch Omni logs, traces, and metrics, building dashboards, configuring alerts, investigating a service, or agent evaluation.

Do not attempt the off-topic task, and do not call a tool or run a query in pursuit of it. Keep it brief: no lecture, no long refusal.

Step 0 — CloudWatch or CloudWatch Omni?

Decide this before routing. The natural wording ("set up an alert for high latency", "build a dashboard", "query my logs") does not say which product the customer means.

  1. The customer names the product — "Omni", "Application Observability", "Agent Observability", a Space, Domain, Dataset, access grant, spaceId, cloudwatch-omni, Omni SQL, PromQL, views, context graph, evaluators, agent evaluation / scoring traces / online or continuous evaluation / gen_ai.evaluation, and OTel span vocabulary (traces.default, logs.default, spans, durationNano, status.code, resource attributes, service.name) → Omni. "Log Insights", "log group", "metric namespace", "CloudWatch alarm", "metric alarm", "composite alarm", "anomaly alarm", "X-Ray", "Application Signals", "canary", "CloudTrail", "Dynamic Instrumentation" → CloudWatch. A "log group" named inside an agent-evaluation request is the online-evaluation data source to verify, not a CloudWatch signal. A bare "alarm" (or "alert") with no other product signal is ambiguous — fall through to rule 3 and probe: a Space → Omni alert (alerts.md); no Space → CloudWatch alarm (cloudwatch/alarms.md). Exception: a "PromQL alarm" is a CloudWatch alarm on OTel metrics (cloudwatch/alarms.md) — "alarm" wins over "PromQL".

  2. Knowledge or how-to question ("what is an Omni alert", "does Omni have an API", "how do alerts differ from alarms") → answer from the reference files directly. Do not probe the account, and do not divert to the other product. Whether a capability exists is a fact about the product, not the account. concepts.md carries the full feature-equivalence matrix. 2a. Authoring / how-to alert request — "create / set up / write an Omni alert that fires when X", "how should I alert on Y" — with no Space or Region supplied and no go-ahead to actually create it, is a HOW-TO request. Deliver the authoring guidance (the alert row of "Must-state checklists" below, then alerts.md) FIRST, from the reference files. Do NOT stall on a Region/Space clarifying question and do NOT fall back to a CloudWatch alarm for a request that explicitly says "Omni alert". Probe the account only once the user supplies a Space/Region or asks you to create it. 2b. A telemetry object is already in the prompt — the user pastes, or the UI passes as context, a span, trace, log record, or query result and asks what it means, what errors it has, or how long it took. This is not a live-data request: do not probe the account, run a query, or ask the user to fetch it again. Read it in place per the "Reading a span or trace you already have" section of query/sql-logs-traces.md and state every item in its must-state list. 2c. A query request whose live result is empty or unreachable ("show me the slowest spans", "which traces failed") is still answered with the methodology — the exact query and the field rules that make it correct (see the must-state callouts in query/sql-logs-traces.md). Zero rows, a wrong-Region Space, or an unreachable endpoint is reported as a finding alongside the query, never as the whole answer.

  3. Request that must act on live data, and the wording is ambiguous (not a Step 0.5 content question, which is answered from the catalog) → probe the target Region first:

    aws cloudwatchomni list-domains
    aws cloudwatchomni list-spaces --region <target-region> --query "items[?region=='<target-region>']"  # account-global; filter to the Region
    • list-spaces is account-global (the --region flag only selects the endpoint), so filter its result to the target Region as shown rather than reading a non-empty list as proof; a Space is one per account per Region. See references/cloudwatch-omni/concepts.md for the full Region-probe rationale.
    • A Domain and a Space in that Region (a non-empty filtered list) → Omni.
    • No Space in that Region (an empty filtered list) → CloudWatch: alarms → cloudwatch/alarms.md, dashboards → cloudwatch/dashboards.md, queries → cloudwatch/log-insights.md, metrics → cloudwatch/metrics.md. If the customer explicitly asked for Omni and has no Space, that is a first-time setup — see step 4.
    • The probe itself errors ("not yet supported", unknown service, endpoint does not resolve) → the CLI/SDK model in use lacks cloudwatchomni. That is not evidence Omni is absent and must not be reported as "Omni is unavailable". The customer's installed AWS CLI/SDK most likely predates the service. If the request carried any Omni signal, give the customer the upgrade command to run (AWS CLI v2 reinstall or brew upgrade awscli; pip install -U boto3 botocore) and have them re-run aws cloudwatchomni list-domains — exact steps in programmatic-access.md. Never run a package-manager upgrade or installer on the host yourself (brew, pip install -U, .pkg/MSI); it mutates the customer's machine beyond the request and can break unrelated tooling — hand over the command and continue with the guidance. Never substitute a CloudWatch or X-Ray command for an Omni request. If the request carried no Omni signal, do not block on the upgrade. When you are already going back to the customer for missing inputs (for example an alert's threshold and period), ask which product they want in that same question: say the Space probe could not run on this CLI, describe both options, and build neither until they answer. When the request is fully specified or says not to ask, proceed on the CloudWatch path (the pre-Omni default) and mention the upgrade only in passing.
    • Still inconclusive → ask the customer.
  4. First-time setup, on EITHER product → STOP and route to the setting-up-cloudwatch-observability skill. This skill covers a service or Space that already reports and holds no onboarding procedure for either product, so there is nothing here to fall back on: not Omni setup (creating a Domain or Space, granting access, ingestion, forwarding, Slack, instrumenting an app or AI agent for a Space) and not CloudWatch setup (onboarding a service to Application Signals — ADOT auto-instrumentation, the amazon-cloudwatch-observability add-on, monitored service, ServiceEvents, CI/CD metadata, the per-platform enablement guides). Any instrumentation / ADOT / collector request is a setup request either way; the setup skill owns the product decision too, so hand the whole request over rather than probing list-spaces yourself. Application Signals once it is reporting — service map, alarms on its metrics, Dynamic Instrumentation debugging — belongs here.

Under-specified alert requests: When an alert or alarm request names what to watch (a symptom or a service) but not the inputs it needs — the threshold value and the evaluation period — ask for those rather than inventing them. Notifications are optional (per alerts.md), so ask for a notification destination only if the user wants to be notified. This holds on both paths (an Omni alert or a CloudWatch alarm).

Under-specified dashboard requests: When a dashboard request names what to show but not which metrics, panels, or layout, ground those against the data and confirm the panel set rather than inventing panels; a dashboard has no threshold, period, or notification. For ANY dashboard authoring/save request, also open dashboards.md and surface its "Facts you MUST surface when building or saving an Omni dashboard" checklist (see "Must-state checklists" below). Which signals or panels a named resource type needs is a Step 0.5 catalog question, not a dashboards-file question.

Step 0.5 — Service-health investigation (routing)

A large share of real questions are phrased as an operator symptom, not as a tool: "is my <service> throttled / slow / erroring / unhealthy," "which of my <service>s are <symptom>," "what's the health of my <service>," "what does <service> depend on and which is broken," "what signals / what should be on a dashboard or view for <service>." These are observability-data questions — answer them from the telemetry surface, not from the resource's control plane, and answer with the methodology (the correct signals, aggregation, scoping, and caveats) even when you also pull live numbers and even when no matching resource exists in the account.

Route by the symptom, then open the reference and surface every applicable item in its "facts you MUST surface" checklist — the checklist is the output contract, and it lives in the reference file, not here:

  • A metric symptom about one AWS service (throttling, latency, error rate, restarts, saturation, "what signals should it get", which signals or panels a named resource type needs on a dashboard or view) → query/promql-metrics.md. Open its "Service-health question — facts you MUST surface" section and surface every applicable item. Dashboard or view wording does not send this to either dashboards file (neither carries the signal facts) nor to rule 3's probe — it is a content question.
  • "What does <service> depend on / what's broken downstream" / blast radius / who is affected / which direction do I walk the graph / what do CALLS, ACCESSES, RUNS_ON mean → context-graph.md. Open its "Dependency / blast-radius question — facts you MUST surface" section and surface every applicable item. A slowness or error symptom phrased in terms of the graph, dependencies, or edge types routes here, not to the metric bullet above.
  • Individual slow or failing spans / traces ("slowest spans for <service>", "which traces failed") are trace SQL, not a metric aggregate → query/sql-logs-traces.md. Open its "Span Duration" and "Finding failed spans" sections and state every item in their must-state callouts — a request about a service's latency or error rate (an aggregate signal) is the PromQL bullet above instead.
Must-state checklists — the output contract

Three Omni tasks carry a checklist the answer text must carry, not just the plan. Open the reference and surface every item relevant to the request; the reference holds the per-item detail and is the source of truth, so do not restate it here.

TaskOpenSection to surface
Authoring or advising on an Omni alertalerts.md"Facts you MUST surface when authoring an alert"
Scoring traces, choosing an evaluator, reading stored scores, online/continuous evaluation, evaluation datasetsagent-evaluation.mdevery "tell the user ALL of this" callout
Authoring, saving, reading back, or debugging an Omni dashboarddashboards.md"Facts you MUST surface when building or saving an Omni dashboard"
Show full SKILL.md (1,264 more words)Show less

Routing — CloudWatch (references/cloudwatch/)

User needAction
Enabling/onboarding a service to Application Signals (auto-instrumentation, the amazon-cloudwatch-observability add-on, monitored service), propagating ServiceEvents git/deployment metadata through CI/CD, or the per-platform × per-language enablement stepsSTOP — this is first-time CloudWatch setup and no reference here covers it. Route to the setting-up-cloudwatch-observability skill
Writing Log Insights queries (pipe-delimited syntax: fields, filter, stats, sort, parse, display)Read log-insights.md
Configuring alarms (metric, composite, anomaly)Read alarms.md. For an Omni alert, see the Omni table
Publishing custom metrics or using EMFRead metrics.md
X-Ray / ADOT tracing behaviour — X-Ray-SDK-vs-ADOT choice, trace and segment structure, annotations vs metadata, sampling rules, collector pipeline config, X-Ray→OTel migration trapsRead tracing.md. Onboarding an un-instrumented service is the setup skill's job (see the first row)
Building CloudWatch dashboards (widget mechanics; which signals a given AWS service needs is Step 0.5)Read dashboards.md
Debugging observability issuesRead troubleshooting.md — starts with the 5 most common fixes
Debugging canary failuresRead synthetics.md — see Common failures table
CloudTrail operational auditingRead cloudtrail.md
Setting up Lambda monitoring with CDKUse alarm-template.ts as a starting point
Creating synthetic canariesRead synthetics.md
Configuring ADOT collectorUse otel-config.yaml as a starting point
Debugging a running service with breakpoints/snapshots — Dynamic Instrumentation (modifies live services and captures live data)Read dynamic-instrumentation.md in full before acting. Confirm with the user before any create/delete, and narrate before significant actions: observation → hypothesis → proposed action → expected result. Source inspection alone identifies hypotheses, not confirmed root causes; keep suspected causes tentative until runtime evidence confirms them.

Routing — CloudWatch Omni (references/cloudwatch-omni/)

Rows that act on live Space data assume Step 0 found a Space. Knowledge questions are answered from the file directly.

User needAction
Concepts. What Omni is, what a Domain / Space / Dataset / grant / profile / view / alert / context graph is, whether a feature is Omni or CloudWatch, where setup startsRead concepts.md
Query logs or traces — SQL (SELECT … FROM logs.default / traces.default / default), field access, schema discovery, slowest / failed spans (durationNano, status.code), TABLESAMPLERead query/sql-logs-traces.md. For slowest or failed spans, state every item in its "Span Duration" / "Finding failed spans" must-state callouts, even when the live result is empty
A span, trace, or log record supplied in the prompt — "I have this span open, what errors are in it", a pasted telemetry objectRead the "Reading a span or trace you already have" section of query/sql-logs-traces.md. Answer from the object's own fields; do not probe, query, or ask the user to fetch it
Query metrics — PromQL, which metric answers which symptom per AWS service, why a metric is missing, gauge vs counterRead query/promql-metrics.md. Metrics are PromQL, never SQL
Views — create, manage, or query named reusable SQL (FROM view.<name>)Read query/views.md
Dashboards in Omni — compose, ground panel queries, author panels[], lay out the grid, the API save semantics (an unknown root- or panel-level key, a missing type/layout, or a bad variant is REJECTED at save with a 400 ValidationException; a bad enum VALUE, x+w>60, or an unknown key inside config saves 200 and fails or is ignored at render; validate before save), fix an empty or blank panel, the *OmniDashboard APIsRead dashboards.md
Alerts in Omni — any mention of an Omni alert, CreateAlert / GetAlert / ListAlerts / UpdateAlert / DeleteAlert, a profileId, an alert ARN, or how alerts differ from alarms; create, tune, tag, list, delete; notificationsRead alerts.md. The alert API is real and first-class — do NOT redirect to CloudWatch alarms. For CloudWatch alarms when Omni is not enabled, read cloudwatch/alarms.md
Context graph — why is service X slow or failing, what depends on it, upstream/downstream, which direction to walk, edge types CALLS / ACCESSES / RUNS_ON, blast radius, walking from an insight or anomaly to a root cause, GetContextGraphRead context-graph.md and state every applicable item in its "facts you MUST surface" section
Agent evaluation — score traces on demand, choose an evaluator, read back stored gen_ai.evaluation.* scores ("which evaluators are doing worst", "which online evaluators are unhealthy / underperforming"), build datasets from traces, set up online evaluation, author a custom evaluator, audit whether an agent's traces are flowingRead agent-evaluation.md and state every applicable item in its "tell the user ALL of this" callouts
Programmatic access — "is there an API or SDK for Omni", calling Omni from code, CI, IaC, or an AI coding agentRead programmatic-access.md. Omni has a real public SigV4 API; never answer that it has none, never substitute the CloudWatch or X-Ray CLI/SDK, and answer without probing for a Space
Who has access to a Space, granting or revoking access, access profiles, creating a Space or Domain, ingestion, forwarding, Slack, Azure, instrumenting an app or AI agent — and, on the CloudWatch side, onboarding a service to Application SignalsRoute to the setting-up-cloudwatch-observability skill. If it is not installed locally, load it with the AWS MCP retrieve_skill tool (skill_name: setting-up-cloudwatch-observability; pass file for a reference it cites)
Spans multiple areasRead the most specific reference first, then consult others as needed

Files

references/cloudwatch/
FileContent
alarms.mdMetric, composite, anomaly detection alarms — configuration, constraints, recommended defaults
log-insights.mdComplete query syntax, commands, functions, known issues, reusable query library
metrics.mdCustom metrics, EMF spec, metric filters, high-resolution, retention
tracing.mdX-Ray → ADOT migration, sampling rules, annotations vs metadata, collector config
dashboards.mdWidget types, cross-account/region, dynamic labels, sharing
troubleshooting.mdError → cause → fix for all observability services
cloudtrail.mdOperational auditing, event types, S3+Athena queries
synthetics.mdCanary runtime/blueprint constraints, VPC networking, common failures
dynamic-instrumentation.mdDynamic Instrumentation debugging loop — breakpoints/probes on live code, snapshot capture + correlation analysis, create/delete gating, snapshot PII handling. Runs via scripts/cloudwatch/di_instrumentation.py + scripts/cloudwatch/di_snapshots.py; details in dynamic-instrumentation/
alarm-template.tsBest-practice CDK Lambda monitoring (alarms + dashboard)
otel-config.yamlADOT collector config for X-Ray traces + CloudWatch EMF metrics

Not here, by design: Application Signals onboarding (the enablement procedure, the ServiceEvents CI/CD chain, the 16 per-platform guides) is first-time setup and lives in setting-up-cloudwatch-observability. Do not reconstruct it here from memory; route.

references/cloudwatch-omni/
FileContent
concepts.mdWhat Omni is and is not; glossary (Domain, Space, Dataset, grant, profile, view, alert, dashboard, context graph, evaluator); Omni-vs-CloudWatch feature-equivalence matrix; how to tell which product the customer means; the setup sequence and where it lives
context-graph.mdThe service/resource topology Omni builds from traces and metrics; GetContextGraph request/response and CLI; reading upstream vs downstream and blast radius; walking from an insight or anomaly hop-by-hop to a root cause, then pivoting to queries
programmatic-access.mdThe public SigV4 API (cloudwatch-omni endpoint prefix, cloudwatch signing name), how access grants authorize a programmatic caller, CLI/SDK access (and why an unsupported-service error is a client-version issue), CloudFormation/CDK, AI coding agents, and the wrong answers to avoid
query/sql-logs-traces.mdSQL over logs and traces — table addressing, required time range, system fields, field access and quoting, schema discovery, supported operations, functions, common patterns (including durationNano span duration), constraints, TABLESAMPLE
query/promql-metrics.mdMetrics in Omni are PromQL — what is queryable (OTLP, span RED, OTel-enriched vended metrics) and what is not, label conventions, __name__ matcher, rate() on counters, per-AWS-service metric catalog with derived formulas and dimension traps
query/views.mdNamed SQL views: CreateView / UpdateView / DeleteView / ListViews, FROM view.<name>, naming and definition rules, composition patterns
dashboards.mdOmni dashboards — composition recipes, grounding panel queries, the panels[] body and panel types, visualizations, the 60-column grid, the API save semantics (unknown root/panel keys are rejected 400; unknown keys inside config save 200 and are ignored at render; validate before save), troubleshooting empty/blank panels, the Create/Get/List/Update/DeleteOmniDashboard APIs, archetype templates
alerts.mdOmni alerts — alert vs alarm, evaluation (FIELD_VALUE / COUNT_OF_RESULTS, contributors), states and no-data treatment, notification rules, step-by-step create / update / delete / tag / fetch, and the alert APIs
agent-evaluation.mdAgent-quality evaluation on OTel traces — instrumentation health audit, evaluator selection, on-demand scoring, online evaluation, custom evaluators, datasets from traces, and reading back stored gen_ai.evaluation.* scores (retrieval plan + SQL mechanics). Uses scripts/cloudwatch-omni/evaluate_traces.py and scripts/cloudwatch-omni/capture_dataset_from_traces.py

© aws, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 53 other files (scripts, references, assets) in skills/core-skills/aws-observability of aws/agent-toolkit-for-aws.

  • SKILL.md
  • assets/cloudwatch/alarm-template.ts
  • assets/cloudwatch/otel-config.yaml
  • references/cloudwatch-omni/agent-evaluation.md
  • references/cloudwatch-omni/alerts.md
  • references/cloudwatch-omni/concepts.md
  • references/cloudwatch-omni/context-graph.md
  • references/cloudwatch-omni/dashboards.md
  • references/cloudwatch-omni/programmatic-access.md
  • references/cloudwatch-omni/query/promql-metrics.md
  • references/cloudwatch-omni/query/sql-logs-traces.md
  • references/cloudwatch-omni/query/views.md
  • references/cloudwatch/alarms.md
  • references/cloudwatch/cloudtrail.md
  • references/cloudwatch/dashboards.md
  • … and 39 more

Open the folder on GitHubat commit bd49cc8

Compare with similar skills

AWS Observability next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

AWS Observability compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
AWS Observability this skillaws/agent-toolkit-for-aws2.8k—~7.2kAutomated safety check: PassApache-2.0
AWS Cdk Developmentzxkane/aws-skills3672 repos~2.5kAutomated safety check: PassMIT
Cloudformationitsmostafa/aws-agent-skills1.2k—~2.5kAutomated safety check: PassMIT
AWS Cloudformation Task Ecs Deploy Ghgiuseppe-trisciuoglio/developer-kit355—~2.8kAutomated safety check: NotesMIT
AWS Sam Bootstrapgiuseppe-trisciuoglio/developer-kit355—~954Automated safety check: NotesMIT
Senior DevOps Toolkitmaslennikov-ig/claude-code-orchestrator-kit2596 repos~1.1kAutomated safety check: NotesCustom licence

Similar skills

  • AWS Cdk Development

    zxkane/aws-skills

    AWS Cloud Development Kit (CDK) expert for building cloud infrastructure with TypeScript/Python.

    367 GitHub starsUsed in 2 repos~2.5k tokens
    DevOps & CloudAuto-check passed
  • Cloudformation

    itsmostafa/aws-agent-skills

    AWS CloudFormation infrastructure as code for stack management.

    1.2k GitHub stars~2.5k tokensUpdated 2 days ago
    DevOps & CloudAuto-check passed
  • AWS Cloudformation Task Ecs Deploy Gh

    giuseppe-trisciuoglio/developer-kit

    Provides patterns to deploy ECS tasks and services with GitHub Actions CI/CD.

    355 GitHub stars~2.8k tokensUpdated 27 days ago
    DevOps & CloudAuto-check: notes
  • AWS Sam Bootstrap

    giuseppe-trisciuoglio/developer-kit

    Provides AWS SAM bootstrap patterns: generates template.yaml and samconfig.toml for new projects via sam init, creates SAM templates for existing Lambda/CloudFormation code migration, validates…

    355 GitHub stars~954 tokensUpdated 27 days ago
    Backend & APIsAuto-check: notes
  • Senior DevOps Toolkit

    maslennikov-ig/claude-code-orchestrator-kit

    Comprehensive DevOps skill for CI/CD, infrastructure automation, containerization, and cloud platforms (AWS, GCP, Azure). Includes pipeline setup…

    259 GitHub starsUsed in 6 repos~1.1k tokens
    DevOps & CloudAuto-check: notes
  • AWS Cloud Advisor

    tech-leads-club/agent-skills

    Answers AWS architecture, security and service-selection questions by searching AWS documentation through MCP tools first, then adapting advice to your stack and team.

    7k GitHub stars~2.1k tokensUpdated 17 days ago
    DevOps & CloudAuto-check passed

More from aws/agent-toolkit-for-aws

All 138 skills in this repo
  • Agent Advisor

    aws/agent-toolkit-for-aws

    Official

    Entry point for AI-agent work on AWS: pick a runtime, plan a migration for existing workloads, and build an executable POC — one phased flow.

    2.8k GitHub stars~4.9k tokensUpdated today
    Auto-check passed
  • Agents Build

    aws/agent-toolkit-for-aws

    Official

    A skill your agent uses to extend an existing agent project with memory, app integration, VPC, multi-agent, migration, model, browser, code interpreter, payments, or resource removal.

    2.8k GitHub stars~2.3k tokensUpdated today
    Auto-check: notes
  • Launch With AWS

    aws/agent-toolkit-for-aws

    Official

    Migrates vibe-coded web applications to AWS. An agent skill from aws/agent-toolkit-for-aws.

    2.8k GitHub stars~3.2k tokensUpdated today
    Auto-check passed
  • Official

    Deploy an event-driven workflow that routes S3 uploads to either Lambda or Fargate via Step Functions based on file size.

    2.8k GitHub stars~4k tokensUpdated today
    Auto-check passed
  • AWS Marketplace Metering

    aws/agent-toolkit-for-aws

    Official

    Deploys, queries, and debugs AWS Marketplace usage-based (PAYG) metering — the pipeline (ResolveCustomer, BatchMeterUsage, EventBridge via SAM) and querying/debugging metering records, statuses…

    2.8k GitHub stars~18k tokensUpdated today
    Auto-check passed
  • Agents Pay

    aws/agent-toolkit-for-aws

    Official

    A skill your agent uses when THIS agent needs to pay for x402-protected content at runtime: hitting a paywall mid-task, settling it via AgentCore Payments, and applying operator-defined spend limits.

    2.8k GitHub stars~6.5k tokensUpdated today
    Auto-check: notes

Categories

Questions about AWS Observability

What does AWS Observability do?

Builds, configures, debugs, and optimizes AWS observability - operator-symptom questions and detecting Omni vs classic CloudWatch. AWS Observability is an agent skill from aws/agent-toolkit-for-aws, published by the product's own GitHub organization. Builds, configures, debugs, and optimizes AWS observability - operator-symptom questions and detecting Omni vs classic CloudWatch.

When should I use AWS Observability?

AWS Observability fits situations like: tasks that involve Observability; tasks that involve Infrastructure as code; tasks that involve Deployment.

How do I install AWS Observability in Claude Code?

Run `npx skills add aws/agent-toolkit-for-aws --skill aws-observability -a claude-code`. Or copy the skill folder (skills/core-skills/aws-observability in aws/agent-toolkit-for-aws) into .claude/skills/aws-observability in your project. Claude Code loads it when a task matches its description.

How do I install AWS Observability in Codex?

Run `npx skills add aws/agent-toolkit-for-aws --skill aws-observability -a codex`. Or copy the skill folder (skills/core-skills/aws-observability in aws/agent-toolkit-for-aws) into .agents/skills/aws-observability in your project. Codex loads it when a task matches its description.

Can I use AWS Observability in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add aws/agent-toolkit-for-aws --skill aws-observability -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/aws-observability, .gemini/skills/aws-observability, .github/skills/aws-observability and .opencode/skills/aws-observability in your project.

What does AWS Observability need to run?

Going by SKILL.md and its folder, AWS Observability needs TypeScript for the scripts in its folder and the command-line tools its instructions call (aws, brew and pip). Our summary lists: Python 3; Node.js.

Does AWS Observability access the network?

SKILL.md names 1 domain. As links in the text: docs.aws.amazon.com. This is read from the text; nothing was executed.

Is AWS Observability safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does AWS Observability use?

AWS Observability is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does AWS Observability use?

About 7.2k tokens (SKILL.md is roughly 29k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 151k tokens, read only when the agent opens those files.

What are the alternatives to AWS Observability?

Skills that share tags, products or a category with AWS Observability: AWS Cdk Development (zxkane/aws-skills, 367 stars), Cloudformation (itsmostafa/aws-agent-skills, 1.2k stars), AWS Cloudformation Task Ecs Deploy Gh (giuseppe-trisciuoglio/developer-kit, 355 stars) and AWS Sam Bootstrap (giuseppe-trisciuoglio/developer-kit, 355 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains AWS Observability?

aws (a GitHub organization, an official publisher) maintains it in aws/agent-toolkit-for-aws, which has 2,816 GitHub stars. The repository holds 138 skills in this directory. The repository was last updated on October 7, 2026.

Source: aws/agent-toolkit-for-aws on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.