Agent skill

Debugging Dags

by astronomer in astronomer/agents

Comprehensive DAG failure diagnosis and root-cause analysis with structured investigation and prevention recommendations.

Apache-2.0Auto-check passedDevelopment

Install Debugging Dags

skills CLI
$ npx skills add astronomer/agents --skill debugging-dags -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install astronomer/agents debugging-dags --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/astronomer/agents.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/debugging-dags .claude/skills/debugging-dags && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
debugging-dags
GitHub stars
450
Token cost
~1.9k tokens
SKILL.md length
941 words
Files
1
Skills in repo
34
Repo updated
First seen
Licence
Apache-2.0

At a glance

Comprehensive DAG failure diagnosis and root-cause analysis with structured investigation and prevention recommendations.

  • Works in 4 steps: Identify the Failure → Get the Error Details → Check Context → …
  • Deep failure investigation is needed
  • SKILL.md covers Running the CLI, Step 1: Identify the Failure, Step 2: Get the Error Details and Step 3: Check Context, plus 1 more section
  • Calls docker, pip and curl; reaches pypi.org

What it does

Debugging Dags is an agent skill from astronomer/agents. Comprehensive DAG failure diagnosis and root-cause analysis with structured investigation and prevention recommendations. Use when deep failure investigation is needed, a DAG fails to import/parse or 'airflow dags list' errors on a file; a task or run is failing and must be diagnosed and fixed; requests like 'why did X fail', 'my dag keeps failing — find and fix it', or fixing a broken DAG so it loads cleanly. For simple 'why did it fail / show logs', the airflow skill handles it directly.

Its SKILL.md is about 1.9k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Development, covering Data pipelines and ETL, Root cause analysis and Debugging. It works with Apache Airflow and Astro. The repository describes itself as: AI agent tooling for data engineering workflows. The licence is Apache-2.0.

When your agent uses it

  • Deep failure investigation is needed
  • A DAG fails to import/parse
  • Airflow dags list errors on a file
  • Run is failing and must be diagnosed and fixed

Example prompts

  • “airflow dags list”
  • “why did X fail”
  • “my dag keeps failing — find and fix it”
  • “/debugging-dags”

Requirements

  • Python 3
  • Docker

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Identify the Failure
  2. Get the Error Details
  3. Check Context
  4. Provide Actionable Output

What it can do on your machine

Read from SKILL.md and the folder at commit 1ec1a1f. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • docker
    • pip
    • curl
    • jq
    • uv

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • pypi.org

    Also links to:

    • astronomer.io

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Debugging Dags loads about 1.9k tokens when it runs. Until then it costs about 128 tokens; SKILL.md has 941 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~128
When it runs · the whole SKILL.md, loaded when a task matches
~1.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from astronomer/agents at commit 1ec1a1f, republished under its Apache-2.0 licence (© astronomer). 941 words, ~1,868 tokens.

Download SKILL.mdSave it as .claude/skills/debugging-dags/SKILL.md (or your agent's skills folder).
name
debugging-dags
description
Comprehensive DAG failure diagnosis and root-cause analysis with structured investigation and prevention recommendations. Use when deep failure investigation is needed, a DAG fails to import/parse or 'airflow dags list' errors on a file; a task or run is failing and must be diagnosed and fixed; requests like 'why did X fail', 'my dag keeps failing — find and fix it', or fixing a broken DAG so it loads cleanly. For simple 'why did it fail / show logs', the airflow skill handles it directly.

DAG Diagnosis

You are a data engineer debugging a failed Airflow DAG. Follow this systematic approach to identify the root cause and provide actionable remediation.

Running the CLI

These commands assume af is on PATH. Run via astro otto to get it automatically, or install standalone with uv tool install astro-airflow-mcp.


Step 1: Identify the Failure

If a specific DAG was mentioned:

  • Run af runs diagnose <dag_id> <dag_run_id> (if run_id is provided)
  • If no run_id specified, run af dags stats to find recent failures

If no DAG was specified:

  • Run af health to find recent failures across all DAGs
  • Check for import errors with af dags errors
  • Show DAGs with recent failures
  • Ask which DAG to investigate further

Step 2: Get the Error Details

Once you have identified a failed task:

  1. Get task logs using af tasks logs <dag_id> <dag_run_id> <task_id>
  2. Look for the actual exception - scroll past the Airflow boilerplate to find the real error
  3. Categorize the failure type:
    • Data issue: Missing data, schema change, null values, constraint violation
    • Code issue: Bug, syntax error, import failure, type error
    • Infrastructure issue: Connection timeout, resource exhaustion, permission denied
    • Dependency issue: Upstream failure, external API down, rate limiting

Step 3: Check Context

Gather additional context to understand WHY this happened:

  1. Recent changes: Was there a code deploy? Check git history if available
  2. Package version changes: Was a package upgraded — in the image, in a venv-style operator, or at the index? See Package version changes below.
  3. Data volume: Did data volume spike? Run a quick count on source tables
  4. Upstream health: Did upstream tasks succeed but produce unexpected data?
  5. Historical pattern: Is this a recurring failure? Check if same task failed before
  6. Timing: Did this fail at an unusual time? (resource contention, maintenance windows)

Use af runs get <dag_id> <dag_run_id> to compare the failed run against recent successful runs.

Package version changes

A common cause of failures with no git activity is dependency drift — the user's code didn't change, but a package they depend on did. Check in this order:

  1. Worker image diff (preferred when available). Every Astro deploy = new image tag, so the registry has a "before" and "after". Diff pip freeze between current and previous image — that's ground truth for what changed:

    docker run --rm <current_image> pip freeze > /tmp/now.txt
    docker run --rm <previous_image> pip freeze > /tmp/prev.txt
    diff /tmp/prev.txt /tmp/now.txt

    Also compare docker run --rm <image> python --version between the two — a Python minor-version bump (3.11 → 3.12, or even a patch) can break wheel compatibility even when pip freeze looks identical. af config providers lists currently installed provider versions, useful for cross-checking against modules named in the traceback.

  2. Venv-style operators bypass the worker image. @task.virtualenv, PythonVirtualenvOperator, ExternalPythonOperator, and KubernetesPodOperator build their environment per task run, so an image diff won't catch failures inside them. If the failed task is one of these, read its requirements / image / python_version / python args directly:

    • Unbounded specifier (e.g. pandas>=2.0.0 with no upper bound, or no specifier at all) → a new upstream release is the prime suspect.
    • image="foo:latest" or no tag → the image moved underneath you.
    • python_version="3.11" (on @task.virtualenv / PythonVirtualenvOperator) or a python path (on ExternalPythonOperator) resolving to a different interpreter than it used to — a Python minor-version change can break wheel compatibility for unchanged requirements. Same vector applies to the worker image itself if the base Python changed there.

    Fix is to pin: pandas>=2.0.0,<3.0.0, a lockfile, a specific image SHA, or a fully-qualified Python version (python_version="3.11.7" instead of "3.11").

  3. Index lookup when image diff isn't conclusive (no image history, or a venv-style operator). Identify the configured index first — it may not be PyPI:

    • Env vars: UV_INDEX_URL, PIP_INDEX_URL, PIP_EXTRA_INDEX_URL
    • pyproject.toml → [[tool.uv.index]]
    • ~/.pip/pip.conf, /etc/pip.conf
    • Dockerfile --index-url flags

    Then query for releases of the suspect package since the first failure started. PyPI:

    curl -s https://pypi.org/pypi/<pkg>/json | jq '.releases | to_entries | map({version: .key, uploaded: .value[0].upload_time}) | sort_by(.uploaded) | reverse | .[:5]'

    Private indexes usually expose the same /pypi/<pkg>/json shape; fall back to the Simple API (/simple/<pkg>/) or ask the user if neither works.

A release timestamp landing between the last green run and the first red run, for a package named in the traceback, is the answer.

Show full SKILL.md (278 more words)Show less
On Astro

If you're running on Astro, these additional tools can help with diagnosis:

  • Deployment activity log: Check the Astro UI for recent deploys — a failed deploy or recent code change is often the cause of sudden failures
  • Astro alerts: Configure alerts in the Astro UI for proactive failure monitoring (DAG failure, task duration, SLA miss)
  • Observability: Use the Astro observability dashboard to track DAG health trends and spot recurring issues
On OSS Airflow
  • Airflow UI: Use the DAGs page, Graph view, and task logs to inspect recent runs and failures

Step 4: Provide Actionable Output

Structure your diagnosis as:

Root Cause

What actually broke? Be specific - not "the task failed" but "the task failed because column X was null in 15% of rows when the code expected 0%".

Impact Assessment
  • What data is affected? Which tables didn't get updated?
  • What downstream processes are blocked?
  • Is this blocking production dashboards or reports?
Immediate Fix

Specific steps to resolve RIGHT NOW:

  1. If it's a data issue: SQL to fix or skip bad records
  2. If it's a code issue: The exact code change needed
  3. If it's infra: Who to contact or what to restart
Prevention

How to prevent this from happening again:

  • Add data quality checks?
  • Add better error handling?
  • Add alerting for edge cases?
  • Update documentation?
  • Pin dependencies (constraints file, lockfile, or upper-bound specifiers on venv/external/pod operators) to avoid silent upstream drift?
Quick Commands

Provide ready-to-use commands:

  • To clear and rerun the entire DAG run: af runs clear <dag_id> <run_id>
  • To clear and rerun specific failed tasks: af tasks clear <dag_id> <run_id> <task_ids> -D
  • To delete a stuck or unwanted run: af runs delete <dag_id> <run_id>

© astronomer, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/debugging-dags of astronomer/agents.

Open the folder on GitHubat commit 1ec1a1f

Compare with similar skills

Debugging Dags next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Debugging Dags compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Debugging Dags this skillastronomer/agents450—~1.9kAutomated safety check: PassApache-2.0
Upgrading Mwaa Environmentsaws/agent-toolkit-for-aws2.8k—~7.3kAutomated safety check: PassApache-2.0
Chart Testsastronomer/airflow-chart297—~2.8kAutomated safety check: PassCustom licence
Functional Testsastronomer/airflow-chart297—~2.2kAutomated safety check: PassCustom licence
Helm Chartastronomer/airflow-chart297—~6.4kAutomated safety check: PassCustom licence
Testing Mwaa Workflowaws/agent-toolkit-for-aws2.8k—~3.8kAutomated safety check: PassApache-2.0

Similar skills

  • Upgrading Mwaa Environments

    aws/agent-toolkit-for-aws

    Official

    Upgrades an MWAA environment to a newer Airflow version — within 2.x, within 3.x, or across the 2.x-to-3.x boundary.

    2.8k GitHub stars~7.3k tokensUpdated today
    Data & AnalyticsAuto-check passed
  • Chart Tests

    astronomer/airflow-chart

    A skill your agent uses when writing, editing, reviewing, or running Helm chart tests for the Astronomer airflow-chart repository.

    297 GitHub stars~2.8k tokensUpdated today
    Data & AnalyticsAuto-check passed
  • Functional Tests

    astronomer/airflow-chart

    A skill your agent uses when writing, editing, reviewing, or running functional (end-to-end) tests for the Astronomer airflow-chart repository.

    297 GitHub stars~2.2k tokensUpdated today
    Data & AnalyticsAuto-check passed
  • Helm Chart

    astronomer/airflow-chart

    A skill your agent uses for Helm chart work - creating charts, modifying existing charts, values design, testing.

    297 GitHub stars~6.4k tokensUpdated today
    Data & AnalyticsAuto-check passed
  • Testing Mwaa Workflow

    aws/agent-toolkit-for-aws

    Official

    Tests Amazon MWAA workflow execution end-to-end: trigger a run and monitor it to completion for Provisioned (Python DAG, via Airflow REST API) and Serverless (YAML workflow, via StartWorkflowRun).

    2.8k GitHub stars~3.8k tokensUpdated today
    Backend & APIsAuto-check passed
  • Stereopy Maintainer

    STOmics/Stereopy

    Stereopy project maintenance guide for code review, bug fixing, and feature development.

    293 GitHub stars~1.7k tokensUpdated 2 mo ago
    DevelopmentAuto-check passed

More from astronomer/agents

All 34 skills in this repo
  • Analyzing Data

    astronomer/agents

    Queries the data warehouse with SQL and answers business questions about data.

    450 GitHub stars~1.3k tokensUpdated yesterday
    Auto-check passed
  • Airflow

    astronomer/agents

    Queries, manages, and troubleshoots Apache Airflow using the af CLI.

    450 GitHub starsUsed in 1 repo~3.8k tokens
    Auto-check passed
  • Guide for migrating Dagster projects to Apache Airflow 3 on Astro.

    450 GitHub stars~3.8k tokensUpdated yesterday
    Auto-check passed
  • Authoring Dags

    astronomer/agents

    Workflow and best practices for writing Apache Airflow DAGs.

    450 GitHub starsUsed in 1 repo~1.8k tokens
    Auto-check passed
  • Deploying Airflow

    astronomer/agents

    Deploys Airflow DAGs and projects. An agent skill from astronomer/agents.

    450 GitHub starsUsed in 1 repo~2.8k tokens
    Auto-check passed
  • Trace downstream data lineage and impact analysis. An agent skill from astronomer/agents.

    450 GitHub starsUsed in 1 repo~1.2k tokens
    Auto-check passed

Questions about Debugging Dags

What does Debugging Dags do?

Comprehensive DAG failure diagnosis and root-cause analysis with structured investigation and prevention recommendations. Debugging Dags is an agent skill from astronomer/agents. Comprehensive DAG failure diagnosis and root-cause analysis with structured investigation and prevention recommendations.

When should I use Debugging Dags?

Debugging Dags fits situations like: deep failure investigation is needed; A DAG fails to import/parse; airflow dags list errors on a file; run is failing and must be diagnosed and fixed.

How do I install Debugging Dags in Claude Code?

Run `npx skills add astronomer/agents --skill debugging-dags -a claude-code`. Or copy the skill folder (skills/debugging-dags in astronomer/agents) into .claude/skills/debugging-dags in your project. Claude Code loads it when a task matches its description.

How do I install Debugging Dags in Codex?

Run `npx skills add astronomer/agents --skill debugging-dags -a codex`. Or copy the skill folder (skills/debugging-dags in astronomer/agents) into .agents/skills/debugging-dags in your project. Codex loads it when a task matches its description.

Can I use Debugging Dags in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add astronomer/agents --skill debugging-dags -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/debugging-dags, .gemini/skills/debugging-dags, .github/skills/debugging-dags and .opencode/skills/debugging-dags in your project.

What does Debugging Dags need to run?

Going by SKILL.md and its folder, Debugging Dags needs the command-line tools its instructions call (docker, pip, curl, jq and uv). Our summary lists: Python 3; Docker.

Does Debugging Dags access the network?

SKILL.md names 2 domains. In commands or code: pypi.org; the agent is likely to contact it when it follows the instructions. As links in the text: astronomer.io. This is read from the text; nothing was executed.

Is Debugging Dags safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Debugging Dags use?

Debugging Dags is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Debugging Dags use?

About 1.9k tokens (SKILL.md is roughly 7.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Debugging Dags?

Skills that share tags, products or a category with Debugging Dags: Upgrading Mwaa Environments (aws/agent-toolkit-for-aws, 2.8k stars), Chart Tests (astronomer/airflow-chart, 297 stars), Functional Tests (astronomer/airflow-chart, 297 stars) and Helm Chart (astronomer/airflow-chart, 297 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Debugging Dags?

astronomer (a GitHub organization) maintains it in astronomer/agents, which has 450 GitHub stars. The repository holds 34 skills in this directory. The repository was last updated on October 5, 2026.

Source: astronomer/agents on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.