AWS Cdk Development
zxkane/aws-skills
AWS Cloud Development Kit (CDK) expert for building cloud infrastructure with TypeScript/Python.
Builds, configures, debugs, and optimizes AWS observability - operator-symptom questions and detecting Omni vs classic CloudWatch.
$ npx skills add aws/agent-toolkit-for-aws --skill aws-observability -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install aws/agent-toolkit-for-aws aws-observability --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/aws/agent-toolkit-for-aws.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/core-skills/aws-observability .claude/skills/aws-observability && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "aws-observability" agent skill from https://github.com/aws/agent-toolkit-for-aws/tree/main/skills/core-skills/aws-observability into .claude/skills/aws-observability/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "aws-observability", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/aws/agent-toolkit-for-aws/tree/main/skills/core-skills/aws-observabilityType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add aws/agent-toolkit-for-aws --skill aws-observability -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install aws/agent-toolkit-for-aws aws-observability --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/aws/agent-toolkit-for-aws.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/core-skills/aws-observability .agents/skills/aws-observability && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "aws-observability" agent skill from https://github.com/aws/agent-toolkit-for-aws/tree/main/skills/core-skills/aws-observability into .agents/skills/aws-observability/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "aws-observability", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add aws/agent-toolkit-for-aws --skill aws-observability -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install aws/agent-toolkit-for-aws aws-observability --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/aws/agent-toolkit-for-aws.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/core-skills/aws-observability .cursor/skills/aws-observability && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "aws-observability" agent skill from https://github.com/aws/agent-toolkit-for-aws/tree/main/skills/core-skills/aws-observability into .cursor/skills/aws-observability/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "aws-observability", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/aws/agent-toolkit-for-aws.git --path skills/core-skills/aws-observability--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add aws/agent-toolkit-for-aws --skill aws-observability -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install aws/agent-toolkit-for-aws aws-observability --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/aws/agent-toolkit-for-aws.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/core-skills/aws-observability .gemini/skills/aws-observability && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "aws-observability" agent skill from https://github.com/aws/agent-toolkit-for-aws/tree/main/skills/core-skills/aws-observability into .gemini/skills/aws-observability/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "aws-observability", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install aws/agent-toolkit-for-aws aws-observabilityInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add aws/agent-toolkit-for-aws --skill aws-observability -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/aws/agent-toolkit-for-aws.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/core-skills/aws-observability .github/skills/aws-observability && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "aws-observability" agent skill from https://github.com/aws/agent-toolkit-for-aws/tree/main/skills/core-skills/aws-observability into .github/skills/aws-observability/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "aws-observability", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add aws/agent-toolkit-for-aws --skill aws-observability -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install aws/agent-toolkit-for-aws aws-observability --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/aws/agent-toolkit-for-aws.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/core-skills/aws-observability .opencode/skills/aws-observability && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "aws-observability" agent skill from https://github.com/aws/agent-toolkit-for-aws/tree/main/skills/core-skills/aws-observability into .opencode/skills/aws-observability/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "aws-observability", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
aws-observabilityBuilds, configures, debugs, and optimizes AWS observability - operator-symptom questions and detecting Omni vs classic CloudWatch.
AWS Observability is an agent skill from aws/agent-toolkit-for-aws, published by the product's own GitHub organization. Builds, configures, debugs, and optimizes AWS observability - operator-symptom questions and detecting Omni vs classic CloudWatch. CloudWatch on an already-reporting service: Log Insights, metric/composite/anomaly alarms, custom metrics/EMF, dashboards, X-Ray/ADOT tracing, canaries, CloudTrail, Dynamic Instrumentation (live breakpoints/snapshots), the Application Signals service map, and fleet health views. CloudWatch Omni on an existing Space: SQL over logs and traces, PromQL over metrics, Omni dashboards, Omni…
Its SKILL.md is about 7.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 59 other files, including scripts, reference files and assets (for example `assets/cloudwatch/alarm-template.ts`, `assets/cloudwatch/otel-config.yaml` and `references/cloudwatch-omni/agent-evaluation.md`).
It sits in DevOps & Cloud, covering Observability, Infrastructure as code and Deployment. It works with Amazon Web Services, Prometheus, SQL and AWS CloudFormation. The repository describes itself as: Official, AWS-supported MCP servers, skills, and plugins to help AI agents build on AWS. The licence is Apache-2.0.
Read from SKILL.md and the folder at commit bd49cc8. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 1 file in scripts/ (TypeScript, from the files we listed), which the agent can run.
Shell commands in SKILL.md call:
awsbrewpipFrom the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
docs.aws.amazon.comFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
AWS Observability loads about 7.2k tokens when it runs, and up to ~158k if it reads all its reference files. Until then it costs about 260 tokens; SKILL.md has 3,277 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from aws/agent-toolkit-for-aws at commit bd49cc8, republished under its Apache-2.0 licence (© aws). 3,277 words, ~7,207 tokens.
.claude/skills/aws-observability/SKILL.md (or your agent's skills folder). This skill also uses 53 other files; get the full folder from GitHub.Domain expertise for AWS observability across metrics, logs, and traces, for two products that share the CloudWatch name but are separate services with separate control planes, data models, and APIs:
| CloudWatch | CloudWatch Omni | |
|---|---|---|
| What it is | Log groups, metric namespaces, alarms, Log Insights, X-Ray, Application Signals | Application Observability / Agent Observability. A Space per account per Region is the access boundary over the account's CloudWatch Dataset (OpenTelemetry logs, traces, and metrics); the Dataset is a CloudWatch resource the Space reads, not something the Space contains |
| Control plane | aws cloudwatch, aws logs, aws xray, aws application-signals | aws cloudwatchomni (endpoint prefix cloudwatch-omni, signing name cloudwatch) |
| Query | Log Insights query language; GetMetricData | SQL over logs.default / traces.default; PromQL over metrics; named views |
| Notify | Alarms (metric, composite, anomaly) | Alerts (SQL/PromQL rule, contributors, OK/WARNING/CRITICAL/NODATA) |
| Topology | Application Signals service map | Context graph (GetContextGraph) |
| Access | IAM only | Domain → Space → access grants and access profiles (set up in setting-up-cloudwatch-observability) |
| Only here | Dynamic Instrumentation, Synthetics canaries, CloudTrail auditing, EMF | Agent-quality evaluation, views, context graph |
| First-time setup | Application Signals onboarding — ADOT auto-instrumentation, the amazon-cloudwatch-observability add-on, ServiceEvents, CI/CD metadata | Domain → Space → grants → telemetry in → plain-ADOT instrumentation |
| References | references/cloudwatch/ | references/cloudwatch-omni/ |
This skill covers using both products, never setting either one up. Both first-time
setup paths above live in setting-up-cloudwatch-observability (same two-folder split);
route there rather than describing an onboarding procedure from memory.
Enabling Omni does not replace CloudWatch; log groups, metrics, and alarms keep working, and most customers use both.
Works best with the AWS MCP server — enables running CLI commands, querying CloudWatch, and validating configurations directly. All guidance also works with standard AWS CLI access.
Note: Reference files contain specific runtime versions, quota values, and feature matrices that may change. When precision matters (e.g. deploying to production, choosing a runtime, or checking a quota), confirm values against current AWS documentation rather than relying solely on the values in these files.
This skill owns the scope decision for every request routed to it, and this section is its sole home.
If a request is NOT about AWS observability with CloudWatch or CloudWatch Omni (weather, trivia, general chit-chat, non-observability coding), decline it in one sentence: state plainly that it is out of scope, never fabricate an answer, and never claim a false capability limitation (no "no internet access", no "no weather data") — the reason is simply that it is out of scope. Then redirect by naming what this skill does cover: querying CloudWatch and CloudWatch Omni logs, traces, and metrics, building dashboards, configuring alerts, investigating a service, or agent evaluation.
Do not attempt the off-topic task, and do not call a tool or run a query in pursuit of it. Keep it brief: no lecture, no long refusal.
Decide this before routing. The natural wording ("set up an alert for high latency", "build a dashboard", "query my logs") does not say which product the customer means.
The customer names the product — "Omni", "Application Observability", "Agent
Observability", a Space, Domain, Dataset, access grant, spaceId, cloudwatch-omni,
Omni SQL, PromQL, views, context graph, evaluators, agent evaluation / scoring traces /
online or continuous evaluation / gen_ai.evaluation, and OTel span vocabulary
(traces.default, logs.default, spans, durationNano, status.code, resource
attributes, service.name) → Omni. "Log Insights", "log group", "metric
namespace", "CloudWatch alarm", "metric alarm", "composite alarm", "anomaly alarm",
"X-Ray", "Application Signals", "canary", "CloudTrail", "Dynamic Instrumentation" →
CloudWatch. A "log group" named inside an agent-evaluation request is the
online-evaluation data source to verify, not a CloudWatch signal. A bare "alarm" (or "alert") with no other product
signal is ambiguous — fall through to rule 3 and probe: a Space → Omni alert
(alerts.md); no Space → CloudWatch alarm
(cloudwatch/alarms.md). Exception: a "PromQL
alarm" is a CloudWatch alarm on OTel metrics
(cloudwatch/alarms.md) — "alarm" wins over "PromQL".
Knowledge or how-to question ("what is an Omni alert", "does Omni have an API", "how do alerts differ from alarms") → answer from the reference files directly. Do not probe the account, and do not divert to the other product. Whether a capability exists is a fact about the product, not the account. concepts.md carries the full feature-equivalence matrix. 2a. Authoring / how-to alert request — "create / set up / write an Omni alert that fires when X", "how should I alert on Y" — with no Space or Region supplied and no go-ahead to actually create it, is a HOW-TO request. Deliver the authoring guidance (the alert row of "Must-state checklists" below, then alerts.md) FIRST, from the reference files. Do NOT stall on a Region/Space clarifying question and do NOT fall back to a CloudWatch alarm for a request that explicitly says "Omni alert". Probe the account only once the user supplies a Space/Region or asks you to create it. 2b. A telemetry object is already in the prompt — the user pastes, or the UI passes as context, a span, trace, log record, or query result and asks what it means, what errors it has, or how long it took. This is not a live-data request: do not probe the account, run a query, or ask the user to fetch it again. Read it in place per the "Reading a span or trace you already have" section of query/sql-logs-traces.md and state every item in its must-state list. 2c. A query request whose live result is empty or unreachable ("show me the slowest spans", "which traces failed") is still answered with the methodology — the exact query and the field rules that make it correct (see the must-state callouts in query/sql-logs-traces.md). Zero rows, a wrong-Region Space, or an unreachable endpoint is reported as a finding alongside the query, never as the whole answer.
Request that must act on live data, and the wording is ambiguous (not a Step 0.5 content question, which is answered from the catalog) → probe the target Region first:
aws cloudwatchomni list-domains
aws cloudwatchomni list-spaces --region <target-region> --query "items[?region=='<target-region>']" # account-global; filter to the Regionlist-spaces is account-global (the --region flag only selects the endpoint), so filter its result to the target Region as shown rather than reading a non-empty list as proof; a Space is one per account per Region. See references/cloudwatch-omni/concepts.md for the full Region-probe rationale.cloudwatchomni. That is not evidence
Omni is absent and must not be reported as "Omni is unavailable". The customer's
installed AWS CLI/SDK most likely predates the service. If the request carried
any Omni signal, give the customer the upgrade command to run (AWS CLI v2
reinstall or brew upgrade awscli; pip install -U boto3 botocore) and have them
re-run aws cloudwatchomni list-domains — exact steps in
programmatic-access.md. Never
run a package-manager upgrade or installer on the host yourself (brew, pip install -U, .pkg/MSI); it mutates the customer's machine beyond the request and
can break unrelated tooling — hand over the command and continue with the guidance.
Never substitute a CloudWatch or X-Ray command for an Omni request. If the request
carried no Omni signal, do not block on the upgrade. When you are already going
back to the customer for missing inputs (for example an alert's threshold and
period), ask which product they want in that same question: say the Space probe could
not run on this CLI, describe both options, and build neither until they answer. When
the request is fully specified or says not to ask, proceed on the CloudWatch path (the
pre-Omni default) and mention the upgrade only in passing.First-time setup, on EITHER product → STOP and route to the
setting-up-cloudwatch-observability skill. This skill covers a service or Space that
already reports and holds no onboarding procedure for either product, so there is nothing
here to fall back on: not Omni setup (creating a Domain or Space, granting access,
ingestion, forwarding, Slack, instrumenting an app or AI agent for a Space) and not
CloudWatch setup (onboarding a service to Application Signals — ADOT
auto-instrumentation, the amazon-cloudwatch-observability add-on, monitored service,
ServiceEvents, CI/CD metadata, the per-platform enablement guides). Any
instrumentation / ADOT / collector request is a setup request either way; the setup skill
owns the product decision too, so hand the whole request over rather than probing
list-spaces yourself. Application Signals once it is reporting — service map,
alarms on its metrics, Dynamic Instrumentation debugging — belongs here.
Under-specified alert requests: When an alert or alarm request names what to watch (a symptom or a service) but not the inputs it needs — the threshold value and the evaluation period — ask for those rather than inventing them. Notifications are optional (per alerts.md), so ask for a notification destination only if the user wants to be notified. This holds on both paths (an Omni alert or a CloudWatch alarm).
Under-specified dashboard requests: When a dashboard request names what to show but not which metrics, panels, or layout, ground those against the data and confirm the panel set rather than inventing panels; a dashboard has no threshold, period, or notification. For ANY dashboard authoring/save request, also open dashboards.md and surface its "Facts you MUST surface when building or saving an Omni dashboard" checklist (see "Must-state checklists" below). Which signals or panels a named resource type needs is a Step 0.5 catalog question, not a dashboards-file question.
A large share of real questions are phrased as an operator symptom, not as a tool: "is my
<service> throttled / slow / erroring / unhealthy," "which of my <service>s are
<symptom>," "what's the health of my <service>," "what does <service> depend on and
which is broken," "what signals / what should be on a dashboard or view for <service>."
These are observability-data questions — answer them from the telemetry surface, not
from the resource's control plane, and answer with the methodology (the correct
signals, aggregation, scoping, and caveats) even when you also pull live numbers and even
when no matching resource exists in the account.
Route by the symptom, then open the reference and surface every applicable item in its "facts you MUST surface" checklist — the checklist is the output contract, and it lives in the reference file, not here:
<service> depend on / what's broken downstream" / blast radius / who is
affected / which direction do I walk the graph / what do CALLS, ACCESSES, RUNS_ON
mean → context-graph.md. Open its
"Dependency / blast-radius question — facts you MUST surface" section and surface
every applicable item. A slowness or error symptom phrased in terms of the graph,
dependencies, or edge types routes here, not to the metric bullet above.<service>", "which
traces failed") are trace SQL, not a metric aggregate →
query/sql-logs-traces.md. Open its
"Span Duration" and "Finding failed spans" sections and state every item in their
must-state callouts — a request about a service's latency or error rate (an aggregate
signal) is the PromQL bullet above instead.Three Omni tasks carry a checklist the answer text must carry, not just the plan. Open the reference and surface every item relevant to the request; the reference holds the per-item detail and is the source of truth, so do not restate it here.
| Task | Open | Section to surface |
|---|---|---|
| Authoring or advising on an Omni alert | alerts.md | "Facts you MUST surface when authoring an alert" |
| Scoring traces, choosing an evaluator, reading stored scores, online/continuous evaluation, evaluation datasets | agent-evaluation.md | every "tell the user ALL of this" callout |
| Authoring, saving, reading back, or debugging an Omni dashboard | dashboards.md | "Facts you MUST surface when building or saving an Omni dashboard" |
references/cloudwatch/)| User need | Action |
|---|---|
Enabling/onboarding a service to Application Signals (auto-instrumentation, the amazon-cloudwatch-observability add-on, monitored service), propagating ServiceEvents git/deployment metadata through CI/CD, or the per-platform × per-language enablement steps | STOP — this is first-time CloudWatch setup and no reference here covers it. Route to the setting-up-cloudwatch-observability skill |
| Writing Log Insights queries (pipe-delimited syntax: fields, filter, stats, sort, parse, display) | Read log-insights.md |
| Configuring alarms (metric, composite, anomaly) | Read alarms.md. For an Omni alert, see the Omni table |
| Publishing custom metrics or using EMF | Read metrics.md |
| X-Ray / ADOT tracing behaviour — X-Ray-SDK-vs-ADOT choice, trace and segment structure, annotations vs metadata, sampling rules, collector pipeline config, X-Ray→OTel migration traps | Read tracing.md. Onboarding an un-instrumented service is the setup skill's job (see the first row) |
| Building CloudWatch dashboards (widget mechanics; which signals a given AWS service needs is Step 0.5) | Read dashboards.md |
| Debugging observability issues | Read troubleshooting.md — starts with the 5 most common fixes |
| Debugging canary failures | Read synthetics.md — see Common failures table |
| CloudTrail operational auditing | Read cloudtrail.md |
| Setting up Lambda monitoring with CDK | Use alarm-template.ts as a starting point |
| Creating synthetic canaries | Read synthetics.md |
| Configuring ADOT collector | Use otel-config.yaml as a starting point |
| Debugging a running service with breakpoints/snapshots — Dynamic Instrumentation (modifies live services and captures live data) | Read dynamic-instrumentation.md in full before acting. Confirm with the user before any create/delete, and narrate before significant actions: observation → hypothesis → proposed action → expected result. Source inspection alone identifies hypotheses, not confirmed root causes; keep suspected causes tentative until runtime evidence confirms them. |
references/cloudwatch-omni/)Rows that act on live Space data assume Step 0 found a Space. Knowledge questions are answered from the file directly.
| User need | Action |
|---|---|
| Concepts. What Omni is, what a Domain / Space / Dataset / grant / profile / view / alert / context graph is, whether a feature is Omni or CloudWatch, where setup starts | Read concepts.md |
Query logs or traces — SQL (SELECT … FROM logs.default / traces.default / default), field access, schema discovery, slowest / failed spans (durationNano, status.code), TABLESAMPLE | Read query/sql-logs-traces.md. For slowest or failed spans, state every item in its "Span Duration" / "Finding failed spans" must-state callouts, even when the live result is empty |
| A span, trace, or log record supplied in the prompt — "I have this span open, what errors are in it", a pasted telemetry object | Read the "Reading a span or trace you already have" section of query/sql-logs-traces.md. Answer from the object's own fields; do not probe, query, or ask the user to fetch it |
| Query metrics — PromQL, which metric answers which symptom per AWS service, why a metric is missing, gauge vs counter | Read query/promql-metrics.md. Metrics are PromQL, never SQL |
Views — create, manage, or query named reusable SQL (FROM view.<name>) | Read query/views.md |
Dashboards in Omni — compose, ground panel queries, author panels[], lay out the grid, the API save semantics (an unknown root- or panel-level key, a missing type/layout, or a bad variant is REJECTED at save with a 400 ValidationException; a bad enum VALUE, x+w>60, or an unknown key inside config saves 200 and fails or is ignored at render; validate before save), fix an empty or blank panel, the *OmniDashboard APIs | Read dashboards.md |
Alerts in Omni — any mention of an Omni alert, CreateAlert / GetAlert / ListAlerts / UpdateAlert / DeleteAlert, a profileId, an alert ARN, or how alerts differ from alarms; create, tune, tag, list, delete; notifications | Read alerts.md. The alert API is real and first-class — do NOT redirect to CloudWatch alarms. For CloudWatch alarms when Omni is not enabled, read cloudwatch/alarms.md |
Context graph — why is service X slow or failing, what depends on it, upstream/downstream, which direction to walk, edge types CALLS / ACCESSES / RUNS_ON, blast radius, walking from an insight or anomaly to a root cause, GetContextGraph | Read context-graph.md and state every applicable item in its "facts you MUST surface" section |
Agent evaluation — score traces on demand, choose an evaluator, read back stored gen_ai.evaluation.* scores ("which evaluators are doing worst", "which online evaluators are unhealthy / underperforming"), build datasets from traces, set up online evaluation, author a custom evaluator, audit whether an agent's traces are flowing | Read agent-evaluation.md and state every applicable item in its "tell the user ALL of this" callouts |
| Programmatic access — "is there an API or SDK for Omni", calling Omni from code, CI, IaC, or an AI coding agent | Read programmatic-access.md. Omni has a real public SigV4 API; never answer that it has none, never substitute the CloudWatch or X-Ray CLI/SDK, and answer without probing for a Space |
| Who has access to a Space, granting or revoking access, access profiles, creating a Space or Domain, ingestion, forwarding, Slack, Azure, instrumenting an app or AI agent — and, on the CloudWatch side, onboarding a service to Application Signals | Route to the setting-up-cloudwatch-observability skill. If it is not installed locally, load it with the AWS MCP retrieve_skill tool (skill_name: setting-up-cloudwatch-observability; pass file for a reference it cites) |
| Spans multiple areas | Read the most specific reference first, then consult others as needed |
references/cloudwatch/| File | Content |
|---|---|
| alarms.md | Metric, composite, anomaly detection alarms — configuration, constraints, recommended defaults |
| log-insights.md | Complete query syntax, commands, functions, known issues, reusable query library |
| metrics.md | Custom metrics, EMF spec, metric filters, high-resolution, retention |
| tracing.md | X-Ray → ADOT migration, sampling rules, annotations vs metadata, collector config |
| dashboards.md | Widget types, cross-account/region, dynamic labels, sharing |
| troubleshooting.md | Error → cause → fix for all observability services |
| cloudtrail.md | Operational auditing, event types, S3+Athena queries |
| synthetics.md | Canary runtime/blueprint constraints, VPC networking, common failures |
| dynamic-instrumentation.md | Dynamic Instrumentation debugging loop — breakpoints/probes on live code, snapshot capture + correlation analysis, create/delete gating, snapshot PII handling. Runs via scripts/cloudwatch/di_instrumentation.py + scripts/cloudwatch/di_snapshots.py; details in dynamic-instrumentation/ |
| alarm-template.ts | Best-practice CDK Lambda monitoring (alarms + dashboard) |
| otel-config.yaml | ADOT collector config for X-Ray traces + CloudWatch EMF metrics |
Not here, by design: Application Signals onboarding (the enablement procedure, the
ServiceEvents CI/CD chain, the 16 per-platform guides) is first-time setup and lives in
setting-up-cloudwatch-observability. Do not reconstruct it here from memory; route.
references/cloudwatch-omni/| File | Content |
|---|---|
| concepts.md | What Omni is and is not; glossary (Domain, Space, Dataset, grant, profile, view, alert, dashboard, context graph, evaluator); Omni-vs-CloudWatch feature-equivalence matrix; how to tell which product the customer means; the setup sequence and where it lives |
| context-graph.md | The service/resource topology Omni builds from traces and metrics; GetContextGraph request/response and CLI; reading upstream vs downstream and blast radius; walking from an insight or anomaly hop-by-hop to a root cause, then pivoting to queries |
| programmatic-access.md | The public SigV4 API (cloudwatch-omni endpoint prefix, cloudwatch signing name), how access grants authorize a programmatic caller, CLI/SDK access (and why an unsupported-service error is a client-version issue), CloudFormation/CDK, AI coding agents, and the wrong answers to avoid |
| query/sql-logs-traces.md | SQL over logs and traces — table addressing, required time range, system fields, field access and quoting, schema discovery, supported operations, functions, common patterns (including durationNano span duration), constraints, TABLESAMPLE |
| query/promql-metrics.md | Metrics in Omni are PromQL — what is queryable (OTLP, span RED, OTel-enriched vended metrics) and what is not, label conventions, __name__ matcher, rate() on counters, per-AWS-service metric catalog with derived formulas and dimension traps |
| query/views.md | Named SQL views: CreateView / UpdateView / DeleteView / ListViews, FROM view.<name>, naming and definition rules, composition patterns |
| dashboards.md | Omni dashboards — composition recipes, grounding panel queries, the panels[] body and panel types, visualizations, the 60-column grid, the API save semantics (unknown root/panel keys are rejected 400; unknown keys inside config save 200 and are ignored at render; validate before save), troubleshooting empty/blank panels, the Create/Get/List/Update/DeleteOmniDashboard APIs, archetype templates |
| alerts.md | Omni alerts — alert vs alarm, evaluation (FIELD_VALUE / COUNT_OF_RESULTS, contributors), states and no-data treatment, notification rules, step-by-step create / update / delete / tag / fetch, and the alert APIs |
| agent-evaluation.md | Agent-quality evaluation on OTel traces — instrumentation health audit, evaluator selection, on-demand scoring, online evaluation, custom evaluators, datasets from traces, and reading back stored gen_ai.evaluation.* scores (retrieval plan + SQL mechanics). Uses scripts/cloudwatch-omni/evaluate_traces.py and scripts/cloudwatch-omni/capture_dataset_from_traces.py |
© aws, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 53 other files (scripts, references, assets) in skills/core-skills/aws-observability of aws/agent-toolkit-for-aws.
Open the folder on GitHubat commit bd49cc8
AWS Observability next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| AWS Observability this skillaws/agent-toolkit-for-aws | 2.8k | — | ~7.2k | Automated safety check: Pass | Apache-2.0 | |
| AWS Cdk Developmentzxkane/aws-skills | 367 | 2 repos | ~2.5k | Automated safety check: Pass | MIT | |
| Cloudformationitsmostafa/aws-agent-skills | 1.2k | — | ~2.5k | Automated safety check: Pass | MIT | |
| AWS Cloudformation Task Ecs Deploy Ghgiuseppe-trisciuoglio/developer-kit | 355 | — | ~2.8k | Automated safety check: Notes | MIT | |
| AWS Sam Bootstrapgiuseppe-trisciuoglio/developer-kit | 355 | — | ~954 | Automated safety check: Notes | MIT | |
| Senior DevOps Toolkitmaslennikov-ig/claude-code-orchestrator-kit | 259 | 6 repos | ~1.1k | Automated safety check: Notes | Custom licence |
zxkane/aws-skills
AWS Cloud Development Kit (CDK) expert for building cloud infrastructure with TypeScript/Python.
itsmostafa/aws-agent-skills
AWS CloudFormation infrastructure as code for stack management.
giuseppe-trisciuoglio/developer-kit
Provides patterns to deploy ECS tasks and services with GitHub Actions CI/CD.
giuseppe-trisciuoglio/developer-kit
Provides AWS SAM bootstrap patterns: generates template.yaml and samconfig.toml for new projects via sam init, creates SAM templates for existing Lambda/CloudFormation code migration, validates…
maslennikov-ig/claude-code-orchestrator-kit
Comprehensive DevOps skill for CI/CD, infrastructure automation, containerization, and cloud platforms (AWS, GCP, Azure). Includes pipeline setup…
tech-leads-club/agent-skills
Answers AWS architecture, security and service-selection questions by searching AWS documentation through MCP tools first, then adapting advice to your stack and team.
aws/agent-toolkit-for-aws
Entry point for AI-agent work on AWS: pick a runtime, plan a migration for existing workloads, and build an executable POC — one phased flow.
aws/agent-toolkit-for-aws
A skill your agent uses to extend an existing agent project with memory, app integration, VPC, multi-agent, migration, model, browser, code interpreter, payments, or resource removal.
aws/agent-toolkit-for-aws
Migrates vibe-coded web applications to AWS. An agent skill from aws/agent-toolkit-for-aws.
aws/agent-toolkit-for-aws
Deploy an event-driven workflow that routes S3 uploads to either Lambda or Fargate via Step Functions based on file size.
aws/agent-toolkit-for-aws
Deploys, queries, and debugs AWS Marketplace usage-based (PAYG) metering — the pipeline (ResolveCustomer, BatchMeterUsage, EventBridge via SAM) and querying/debugging metering records, statuses…
aws/agent-toolkit-for-aws
A skill your agent uses when THIS agent needs to pay for x402-protected content at runtime: hitting a paywall mid-task, settling it via AgentCore Payments, and applying operator-defined spend limits.
Categories
Builds, configures, debugs, and optimizes AWS observability - operator-symptom questions and detecting Omni vs classic CloudWatch. AWS Observability is an agent skill from aws/agent-toolkit-for-aws, published by the product's own GitHub organization. Builds, configures, debugs, and optimizes AWS observability - operator-symptom questions and detecting Omni vs classic CloudWatch.
AWS Observability fits situations like: tasks that involve Observability; tasks that involve Infrastructure as code; tasks that involve Deployment.
Run `npx skills add aws/agent-toolkit-for-aws --skill aws-observability -a claude-code`. Or copy the skill folder (skills/core-skills/aws-observability in aws/agent-toolkit-for-aws) into .claude/skills/aws-observability in your project. Claude Code loads it when a task matches its description.
Run `npx skills add aws/agent-toolkit-for-aws --skill aws-observability -a codex`. Or copy the skill folder (skills/core-skills/aws-observability in aws/agent-toolkit-for-aws) into .agents/skills/aws-observability in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add aws/agent-toolkit-for-aws --skill aws-observability -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/aws-observability, .gemini/skills/aws-observability, .github/skills/aws-observability and .opencode/skills/aws-observability in your project.
Going by SKILL.md and its folder, AWS Observability needs TypeScript for the scripts in its folder and the command-line tools its instructions call (aws, brew and pip). Our summary lists: Python 3; Node.js.
SKILL.md names 1 domain. As links in the text: docs.aws.amazon.com. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
AWS Observability is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 7.2k tokens (SKILL.md is roughly 29k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 151k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with AWS Observability: AWS Cdk Development (zxkane/aws-skills, 367 stars), Cloudformation (itsmostafa/aws-agent-skills, 1.2k stars), AWS Cloudformation Task Ecs Deploy Gh (giuseppe-trisciuoglio/developer-kit, 355 stars) and AWS Sam Bootstrap (giuseppe-trisciuoglio/developer-kit, 355 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
aws (a GitHub organization, an official publisher) maintains it in aws/agent-toolkit-for-aws, which has 2,816 GitHub stars. The repository holds 138 skills in this directory. The repository was last updated on October 7, 2026.
Source: aws/agent-toolkit-for-aws on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.