Official agent skill

Debugging Mwaa Workflow

by aws in aws/agent-toolkit-for-aws

Diagnoses and root-causes Amazon MWAA workflow failures across Provisioned (Python DAG) and Serverless (YAML workflow) environments.

OfficialApache-2.0Auto-check passedBackend & APIs

Install Debugging Mwaa Workflow

skills CLI
$ npx skills add aws/agent-toolkit-for-aws --skill debugging-mwaa-workflow -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install aws/agent-toolkit-for-aws debugging-mwaa-workflow --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/aws/agent-toolkit-for-aws.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/aws-data-analytics/skills/debugging-mwaa-workflow .claude/skills/debugging-mwaa-workflow && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
debugging-mwaa-workflow
GitHub stars
2.8k
Token cost
~2.3k tokens
SKILL.md length
988 words
Files
4 (incl. references)
Skills in repo
138
Repo updated
First seen
Licence
Apache-2.0

At a glance

Diagnoses and root-causes Amazon MWAA workflow failures across Provisioned (Python DAG) and Serverless (YAML workflow) environments.

  • Works in 5 steps: Detect Flavor and Scope the Failure → Identify the Failure → Get Error Details and Categorize → …
  • Workflow run failed
  • SKILL.md covers Guardrail — where this skill's…, Step 0: Detect Flavor and…, Step 1: Identify the Failure and Step 2: Get Error Details and…, plus 6 more sections
  • Calls aws

What it does

Debugging Mwaa Workflow is an agent skill from aws/agent-toolkit-for-aws, published by the product's own GitHub organization. Diagnoses and root-causes Amazon MWAA workflow failures across Provisioned (Python DAG) and Serverless (YAML workflow) environments. Provisioned uses aws mwaa invoke-rest-api, CloudWatch log groups, and get-environment; Serverless uses aws mwaa-serverless API (GetWorkflowRun, ListWorkflowRuns, GetTaskInstance) and CloudWatch logs. Covers failed runs and tasks, DAGs not appearing, import errors, worker OOM, IAM denials, and dependency drift. Triggers on: DAG failed, task failed, workflow run failed, MWAA error…

Its SKILL.md is about 2.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files, including reference files (for example `references/failure-catalog.md`, `references/provisioned-diagnostics.md` and `references/serverless-diagnostics.md`).

It sits in Backend & APIs, covering Serverless, Debugging and Root cause analysis. It works with Amazon Web Services, Python and Model Context Protocol. The repository describes itself as: Official, AWS-supported MCP servers, skills, and plugins to help AI agents build on AWS. The licence is Apache-2.0.

When your agent uses it

  • Workflow run failed
  • Why did my workflow fail
  • DAG not showing up
  • MWAA import error

Example prompts

  • “Use the debugging-mwaa-workflow skill to diagnose and root-causes Amazon MWAA workflow failures across Provisioned (Python DAG) and Serverless (YAML…”
  • “/debugging-mwaa-workflow”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Detect Flavor and Scope the Failure
  2. Identify the Failure
  3. Get Error Details and Categorize
  4. Check Context (Why It Happened)
  5. Provide Actionable Output

What it can do on your machine

Read from SKILL.md and the folder at commit bd49cc8. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • aws

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use aws, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Debugging Mwaa Workflow loads about 2.3k tokens when it runs, and up to ~7.5k if it reads all its reference files. Until then it costs about 213 tokens; SKILL.md has 988 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~213
When it runs · the whole SKILL.md, loaded when a task matches
~2.3k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~7.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from aws/agent-toolkit-for-aws at commit bd49cc8, republished under its Apache-2.0 licence (© aws). 988 words, ~2,316 tokens.

Download SKILL.mdSave it as .claude/skills/debugging-mwaa-workflow/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
debugging-mwaa-workflow
description
Diagnoses and root-causes Amazon MWAA workflow failures across Provisioned (Python DAG) and Serverless (YAML workflow) environments. Provisioned uses aws mwaa invoke-rest-api, CloudWatch log groups, and get-environment; Serverless uses aws mwaa-serverless API (GetWorkflowRun, ListWorkflowRuns, GetTaskInstance) and CloudWatch logs. Covers failed runs and tasks, DAGs not appearing, import errors, worker OOM, IAM denials, and dependency drift. Triggers on: DAG failed, task failed, workflow run failed, MWAA error, debug my DAG, why did my workflow fail, DAG not showing up, MWAA import error, requirements failing, worker crashed, serverless run failed. Not applicable to authoring workflows (handled by authoring-mwaa-workflow), running or smoke-testing a workflow (handled by testing-mwaa-workflow), or CI-CD deploy failures.
metadata.version
1

Debugging MWAA Workflows

AWS MCP server (optional but recommended): running the AWS CLI commands in this skill through the AWS MCP server gives sandboxed execution and audit logging. Every command here also works with the plain AWS CLI, so the skill does not require the MCP server or any MCP-only tools.

Diagnose and root-cause Amazon MWAA workflow failures, then report root cause, impact, and recommended remediation. Routes by flavor, then runs a shared 4-step diagnostic spine.

Guardrail — where this skill's own files live (MCP vs local install)

This skill can be loaded two ways, and they resolve the skill's own bundled files from different places. Determine how the skill was loaded before reading a reference:

  • Loaded through the AWS MCP retrieve_skill tool: The skill is not installed on the local filesystem. You MUST fetch each reference via retrieve_skill with the file parameter (e.g. file="references/failure-catalog.md") and read the returned content. Do NOT file_read these paths locally — they do not exist on disk.
  • Installed locally (e.g. .kiro/skills/debugging-mwaa-workflow/ or ~/.claude/skills/debugging-mwaa-workflow/): Read the files from the local skill directory using relative paths.

This distinction applies only to the skill's own packaged files. User data and session artifacts are always read from and written to the user's working directory. Never fetch or write customer data through retrieve_skill.

Step 0: Detect Flavor and Scope the Failure

Detect flavor
  1. An environment name resolvable via aws mwaa get-environment means the environment is Provisioned.
  2. A workflow/... ARN or any aws mwaa-serverless context means the environment is Serverless. A bare run identifier does NOT indicate flavor — Provisioned DAG runs also have run ids.
  3. If neither signal is present, ask: is the target MWAA Provisioned (Python DAG) or MWAA Serverless (YAML workflow)?
Route by complexity
  • Simple — a single named task or run failed with a clear exception. Jump to Step 2 for that task.
  • Standard — a run failed and the cause is unknown. Run the full Step 1 to Step 4 sweep.
  • Complex — intermittent or environment-wide (multiple DAGs, "worked yesterday", nothing appearing). Run the full sweep with emphasis on Step 3.

Step 1: Identify the Failure

Provisioned: list failed DAG runs and task instances via aws mwaa invoke-rest-api (paths /dags/{id}/dagRuns and /dags/{id}/dagRuns/{run_id}/taskInstances). If invoke-rest-api errors (RestApiClientException), fall back to the Scheduler and DAGProcessing log groups. Get version and config from aws mwaa get-environment. See references/provisioned-diagnostics.md.

Serverless: aws mwaa-serverless list-workflow-runs, then get-workflow-run. Read RunDetail.ErrorMessage — an empty TaskInstances with a parser message is a definition error; Workflow execution failed with populated TaskInstances is a task-execution failure. See references/serverless-diagnostics.md.

Step 2: Get Error Details and Categorize

Pull the real exception past boilerplate:

Provisioned: read the Task log group first, then Worker/Scheduler/ DAGProcessing as the symptom directs.

Serverless: list-task-instances then get-task-instance to get each task's LogStream, then read that stream in CloudWatch.

Then categorize in priority order — infra, then drift, then code-data — using references/failure-catalog.md. The category determines the Step 3 checks.

Step 3: Check Context (Why It Happened)

Run the context checks for the matched category from references/failure-catalog.md. Do not stop at the surface exception: a SIGKILL is an OOM story, a fresh import error on unchanged code is a drift story, a sensor timeout is an upstream-health story.

Step 4: Provide Actionable Output

Report in this exact structure:

Root Cause: <one-line diagnosis with the evidence that proves it>
Impact: <what failed, which runs, blast radius>
Immediate Fix: <the smallest change that unblocks>
Prevention: <the change that stops recurrence>
Commands: <exact read-only commands run, plus remediation commands for the user to run>

Run only read-only operations. Present state-mutating remediation (clear/rerun/backfill for Provisioned; start-workflow-run or fix-and-redeploy for Serverless) as commands for the user to run, with the impact stated. Never execute them autonomously (production safety).

For the fix-and-redeploy path, use authoring-mwaa-workflow to regenerate a compliant artifact.

Show full SKILL.md (408 more words)Show less

Gotchas

  • Serverless has no Airflow web UI, no REST API, and no CLI token. Do not attempt create-web-login-token, invoke-rest-api, or any Airflow REST path for Serverless.
  • For Provisioned, always use aws mwaa invoke-rest-api (not create-web-login-token + curl). invoke-rest-api reaches VPC-only web servers without network access.
  • GetWorkflowRun.RunDetail.ErrorMessage distinguishes a definition error (empty TaskInstances) from a task-execution failure (Workflow execution failed, populated TaskInstances). Read it before pulling task logs.
  • The Serverless log group defaults to /aws/mwaa-serverless/{workflow-id}/ but can be a custom group; confirm via get-workflow LoggingConfiguration before assuming the path.
  • A DAG not appearing has several causes — an import/parse error, the scheduler scan interval (scheduler.dag_dir_list_interval, or dag_processor.refresh_interval on Airflow 3.x) not yet elapsed, a dag_id collision, or S3-sync delay — and is rarely a broken DAG. Check GET /importErrors and GET /dags/{dag_id} via invoke-rest-api (and the DAGProcessing logs); see the failure catalog's "DAG not appearing in the UI" checklist before concluding the code is wrong.
  • A worker SIGKILL is an OOM signal. Recommend moving work to Glue/EMR/Lambda; scaling workers alone does not fix per-task memory pressure.
  • MWAA re-resolves dependencies on environment update, so an unchanged DAG can start failing on import with no code change. Treat no-code-change import failures as drift.
  • Serverless PythonOperator/BashOperator tasks run custom code from a --code package. A run that fails to extract the package or hits ModuleNotFoundError is a packaging problem (wrong-platform wheel, missing dep, bad layout), not a YAML definition error. See the failure catalog's serverless custom-code section.

Troubleshooting

ErrorCauseFix
RestApiClientException (Provisioned)Mis-scoped execution role or service errorFall back to Scheduler/DAGProcessing log groups
ResourceNotFoundException on get-workflow-runWrong workflow ARN or run idRe-list with list-workflow-runs
Task log stream empty (Serverless)Wrong log group assumedRead LoggingConfiguration from get-workflow
No task logs but run FAILEDDefinition/parse errorRead RunDetail.ErrorMessage; fix the YAML

References

Security Considerations

  • Read-only by default: diagnosis uses only read/list/describe calls. Remediation (clear/rerun/backfill, IAM or key-policy changes) is presented as commands for the user to run, never executed autonomously.
  • Least-privilege IAM: when an AccessDenied is a genuine permission gap, recommend the minimal Action/Resource from the error — never a wildcard; distinguish it from a nonexistent-resource typo (do not broaden IAM then).
  • Cross-account: KMS key-policy / assume-role changes are human-gated and coordinated with the resource owner.
  • No secret exposure: do not surface credentials or connection strings from logs or API responses in the diagnosis output.

© aws, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files (references) in plugins/aws-data-analytics/skills/debugging-mwaa-workflow of aws/agent-toolkit-for-aws.

  • SKILL.md
  • references/failure-catalog.md
  • references/provisioned-diagnostics.md
  • references/serverless-diagnostics.md

Open the folder on GitHubat commit bd49cc8

Compare with similar skills

Debugging Mwaa Workflow next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Debugging Mwaa Workflow compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Debugging Mwaa Workflow this skillaws/agent-toolkit-for-aws2.8k—~2.3kAutomated safety check: PassApache-2.0
AWS Serverless Edazxkane/aws-skills3674 repos~3.2kAutomated safety check: PassMIT
AWS Lambda Durable Functionsawslabs/agent-plugins912—~2.3kAutomated safety check: PassApache-2.0
AWS Cdk Developmentzxkane/aws-skills3672 repos~2.5kAutomated safety check: PassMIT
QA Find Bugs MCPbex-co/beancount-io294—~3kAutomated safety check: PassMIT
LangBot Plugin Developmentlangbot-app/LangBot18k—~3.9kAutomated safety check: PassApache-2.0

Similar skills

  • AWS Serverless Eda

    zxkane/aws-skills

    AWS serverless and event-driven architecture expert based on Well-Architected Framework.

    367 GitHub starsUsed in 4 repos~3.2k tokens
    Backend & APIsAuto-check passed
  • AWS Lambda Durable Functions

    awslabs/agent-plugins

    Official

    Build resilient, long-running, multi-step applications with AWS Lambda durable functions with automatic state persistence, retry logic, and orchestration for long-running executions.

    912 GitHub stars~2.3k tokensUpdated 2 days ago
    Backend & APIsAuto-check passed
  • AWS Cdk Development

    zxkane/aws-skills

    AWS Cloud Development Kit (CDK) expert for building cloud infrastructure with TypeScript/Python.

    367 GitHub starsUsed in 2 repos~2.5k tokens
    DevOps & CloudAuto-check passed
  • QA Find Bugs MCP

    bex-co/beancount-io

    Hunt bugs in the Beancount.io remote MCP server by driving the real POST /api-gateway/mcp endpoint with JSON-RPC and real MCP clients, checking transport, discovery, credential boundaries, tool and…

    294 GitHub stars~3k tokensUpdated today
    Backend & APIsAuto-check passed
  • LangBot Plugin Development

    langbot-app/LangBot

    Guides building, debugging and testing LangBot plugins: components, SDK calls, README and locale rules, SDK pitfalls and WebSocket-based testing.

    18k GitHub stars~3.9k tokensUpdated today
    DevelopmentAuto-check passed
  • Send and receive transactional emails with Cloudflare Email Service (Email Sending + Email Routing).

    127 GitHub starsUsed in 3 repos~2k tokens
    Backend & APIsAuto-check passed

More from aws/agent-toolkit-for-aws

All 138 skills in this repo
  • Agent Advisor

    aws/agent-toolkit-for-aws

    Official

    Entry point for AI-agent work on AWS: pick a runtime, plan a migration for existing workloads, and build an executable POC — one phased flow.

    2.8k GitHub stars~4.9k tokensUpdated today
    Auto-check passed
  • Agents Build

    aws/agent-toolkit-for-aws

    Official

    A skill your agent uses to extend an existing agent project with memory, app integration, VPC, multi-agent, migration, model, browser, code interpreter, payments, or resource removal.

    2.8k GitHub stars~2.3k tokensUpdated today
    Auto-check: notes
  • Launch With AWS

    aws/agent-toolkit-for-aws

    Official

    Migrates vibe-coded web applications to AWS. An agent skill from aws/agent-toolkit-for-aws.

    2.8k GitHub stars~3.2k tokensUpdated today
    Auto-check passed
  • Official

    Deploy an event-driven workflow that routes S3 uploads to either Lambda or Fargate via Step Functions based on file size.

    2.8k GitHub stars~4k tokensUpdated today
    Auto-check passed
  • AWS Marketplace Metering

    aws/agent-toolkit-for-aws

    Official

    Deploys, queries, and debugs AWS Marketplace usage-based (PAYG) metering — the pipeline (ResolveCustomer, BatchMeterUsage, EventBridge via SAM) and querying/debugging metering records, statuses…

    2.8k GitHub stars~18k tokensUpdated today
    Auto-check passed
  • Agents Pay

    aws/agent-toolkit-for-aws

    Official

    A skill your agent uses when THIS agent needs to pay for x402-protected content at runtime: hitting a paywall mid-task, settling it via AgentCore Payments, and applying operator-defined spend limits.

    2.8k GitHub stars~6.5k tokensUpdated today
    Auto-check: notes

Questions about Debugging Mwaa Workflow

What does Debugging Mwaa Workflow do?

Diagnoses and root-causes Amazon MWAA workflow failures across Provisioned (Python DAG) and Serverless (YAML workflow) environments. Debugging Mwaa Workflow is an agent skill from aws/agent-toolkit-for-aws, published by the product's own GitHub organization. Diagnoses and root-causes Amazon MWAA workflow failures across Provisioned (Python DAG) and Serverless (YAML workflow) environments.

When should I use Debugging Mwaa Workflow?

Debugging Mwaa Workflow fits situations like: workflow run failed; why did my workflow fail; DAG not showing up; MWAA import error.

How do I install Debugging Mwaa Workflow in Claude Code?

Run `npx skills add aws/agent-toolkit-for-aws --skill debugging-mwaa-workflow -a claude-code`. Or copy the skill folder (plugins/aws-data-analytics/skills/debugging-mwaa-workflow in aws/agent-toolkit-for-aws) into .claude/skills/debugging-mwaa-workflow in your project. Claude Code loads it when a task matches its description.

How do I install Debugging Mwaa Workflow in Codex?

Run `npx skills add aws/agent-toolkit-for-aws --skill debugging-mwaa-workflow -a codex`. Or copy the skill folder (plugins/aws-data-analytics/skills/debugging-mwaa-workflow in aws/agent-toolkit-for-aws) into .agents/skills/debugging-mwaa-workflow in your project. Codex loads it when a task matches its description.

Can I use Debugging Mwaa Workflow in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add aws/agent-toolkit-for-aws --skill debugging-mwaa-workflow -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/debugging-mwaa-workflow, .gemini/skills/debugging-mwaa-workflow, .github/skills/debugging-mwaa-workflow and .opencode/skills/debugging-mwaa-workflow in your project.

What does Debugging Mwaa Workflow need to run?

Going by SKILL.md and its folder, Debugging Mwaa Workflow needs the command-line tools its instructions call (aws). Our summary lists: Python 3.

Does Debugging Mwaa Workflow access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Debugging Mwaa Workflow safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Debugging Mwaa Workflow use?

Debugging Mwaa Workflow is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Debugging Mwaa Workflow use?

About 2.3k tokens (SKILL.md is roughly 9.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 5.2k tokens, read only when the agent opens those files.

What are the alternatives to Debugging Mwaa Workflow?

Skills that share tags, products or a category with Debugging Mwaa Workflow: AWS Serverless Eda (zxkane/aws-skills, 367 stars), AWS Lambda Durable Functions (awslabs/agent-plugins, 912 stars), AWS Cdk Development (zxkane/aws-skills, 367 stars) and QA Find Bugs MCP (bex-co/beancount-io, 294 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Debugging Mwaa Workflow?

aws (a GitHub organization, an official publisher) maintains it in aws/agent-toolkit-for-aws, which has 2,816 GitHub stars. The repository holds 138 skills in this directory. The repository was last updated on October 7, 2026.

Source: aws/agent-toolkit-for-aws on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.