Agent skill

Testing Dags

by astronomer in astronomer/agents

Complex DAG testing workflows with debugging and fixing cycles.

Apache-2.0Auto-check passedData & Analytics

Install Testing Dags

skills CLI
$ npx skills add astronomer/agents --skill testing-dags -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install astronomer/agents testing-dags --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/astronomer/agents.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/testing-dags .claude/skills/testing-dags && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
testing-dags
GitHub stars
450
Token cost
~2.5k tokens
SKILL.md length
840 words
Files
1
Skills in repo
34
Repo updated
First seen
Licence
Apache-2.0

At a glance

Complex DAG testing workflows with debugging and fixing cycles.

  • Works in 3 steps: Trigger and Wait → Debug Failures (Only If Needed) → Fix and Retest
  • Multi-step testing requests like test this dag and fix it if it fails
  • SKILL.md covers Running the CLI, Quick Validation with Astro CLI, FIRST ACTION: Just Trigger the… and Testing Workflow Overview, plus 7 more sections
  • Calls uv

What it does

Testing Dags is an agent skill from astronomer/agents. Complex DAG testing workflows with debugging and fixing cycles. Use for multi-step testing requests like "test this dag and fix it if it fails", "test and debug", "run the pipeline and troubleshoot issues". For simple test requests ("test dag", "run dag"), the airflow entrypoint skill handles it directly. This skill is for iterative test-debug-fix cycles.

Its SKILL.md is about 2.5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Data & Analytics, covering Data pipelines and ETL and Debugging. It works with Apache Airflow and Astro. The repository describes itself as: AI agent tooling for data engineering workflows. The licence is Apache-2.0.

When your agent uses it

  • Multi-step testing requests like test this dag and fix it if it fails
  • Run the pipeline and troubleshoot issues

Example prompts

  • “test this dag and fix it if it fails”
  • “test and debug”
  • “run the pipeline and troubleshoot issues”
  • “/testing-dags”

Workflow steps

3 steps, taken from the step headings in SKILL.md.

  1. Trigger and Wait
  2. Debug Failures (Only If Needed)
  3. Fix and Retest

What it can do on your machine

Read from SKILL.md and the folder at commit 1ec1a1f. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • uv

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use uv, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Testing Dags loads about 2.5k tokens when it runs. Until then it costs about 93 tokens; SKILL.md has 840 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~93
When it runs · the whole SKILL.md, loaded when a task matches
~2.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from astronomer/agents at commit 1ec1a1f, republished under its Apache-2.0 licence (© astronomer). 840 words, ~2,481 tokens.

Download SKILL.mdSave it as .claude/skills/testing-dags/SKILL.md (or your agent's skills folder).
name
testing-dags
description
Complex DAG testing workflows with debugging and fixing cycles. Use for multi-step testing requests like "test this dag and fix it if it fails", "test and debug", "run the pipeline and troubleshoot issues". For simple test requests ("test dag", "run dag"), the airflow entrypoint skill handles it directly. This skill is for iterative test-debug-fix cycles.

DAG Testing Skill

Use af commands to test, debug, and fix DAGs in iterative cycles.

Running the CLI

These commands assume af is on PATH. Run via astro otto to get it automatically, or install standalone with uv tool install astro-airflow-mcp.


Quick Validation with Astro CLI

If the user has the Astro CLI available, these commands provide fast feedback without needing a running Airflow instance:

bash
# Parse DAGs to catch import errors, syntax issues, and DAG-level problems
astro dev parse

# Run pytest against DAGs (runs tests in tests/ directory)
astro dev pytest

Use these for quick validation during development. For full end-to-end testing against a live Airflow instance, continue to the trigger-and-wait workflow below.


FIRST ACTION: Just Trigger the DAG

When the user asks to test a DAG, your FIRST AND ONLY action should be:

bash
af runs trigger-wait <dag_id>

DO NOT:

  • Call af dags list first
  • Call af dags get first
  • Call af dags errors first
  • Use grep or ls or any other bash command
  • Do any "pre-flight checks"

Just trigger the DAG. If it fails, THEN debug.


Testing Workflow Overview

┌─────────────────────────────────────┐
│ 1. TRIGGER AND WAIT                 │
│    Run DAG, wait for completion     │
└─────────────────────────────────────┘
                 ↓
        ┌───────┴───────┐
        ↓               ↓
   ┌─────────┐    ┌──────────┐
   │ SUCCESS │    │ FAILED   │
   │ Done!   │    │ Debug... │
   └─────────┘    └──────────┘
                       ↓
        ┌─────────────────────────────────────┐
        │ 2. DEBUG (only if failed)           │
        │    Get logs, identify root cause    │
        └─────────────────────────────────────┘
                       ↓
        ┌─────────────────────────────────────┐
        │ 3. FIX AND RETEST                   │
        │    Apply fix, restart from step 1   │
        └─────────────────────────────────────┘

Philosophy: Try first, debug on failure. Don't waste time on pre-flight checks — just run the DAG and diagnose if something goes wrong.


Phase 1: Trigger and Wait

Use af runs trigger-wait to test the DAG:

Primary Method: Trigger and Wait
bash
af runs trigger-wait <dag_id> --timeout 300

Example:

bash
af runs trigger-wait my_dag --timeout 300

Why this is the preferred method:

  • Single command handles trigger + monitoring
  • Returns immediately when DAG completes (success or failure)
  • Includes failed task details if run fails
  • No manual polling required
Response Interpretation

Success:

json
{
  "dag_run": {
    "dag_id": "my_dag",
    "dag_run_id": "manual__2025-01-14T...",
    "state": "success",
    "start_date": "...",
    "end_date": "..."
  },
  "timed_out": false,
  "elapsed_seconds": 45.2
}

Failure:

json
{
  "dag_run": {
    "state": "failed"
  },
  "timed_out": false,
  "elapsed_seconds": 30.1,
  "failed_tasks": [
    {
      "task_id": "extract_data",
      "state": "failed",
      "try_number": 2
    }
  ]
}

Timeout:

json
{
  "dag_id": "my_dag",
  "dag_run_id": "manual__...",
  "state": "running",
  "timed_out": true,
  "elapsed_seconds": 300.0,
  "message": "Timed out after 300 seconds. DAG run is still running."
}
Alternative: Trigger and Monitor Separately

Use this only when you need more control:

bash
# Step 1: Trigger
af runs trigger my_dag
# Returns: {"dag_run_id": "manual__...", "state": "queued"}

# Step 2: Check status
af runs get my_dag manual__2025-01-14T...
# Returns current state

Handling Results

If Success

The DAG ran successfully. Summarize for the user:

  • Total elapsed time
  • Number of tasks completed
  • Any notable outputs (if visible in logs)

You're done!

If Timed Out

The DAG is still running. Options:

  1. Check current status: af runs get <dag_id> <dag_run_id>
  2. Ask user if they want to continue waiting
  3. Increase timeout and try again
If Failed

Move to Phase 2 (Debug) to identify the root cause.


Phase 2: Debug Failures (Only If Needed)

When a DAG run fails, use these commands to diagnose:

Get Comprehensive Diagnosis
bash
af runs diagnose <dag_id> <dag_run_id>

Returns in one call:

  • Run metadata (state, timing)
  • All task instances with states
  • Summary of failed tasks
  • State counts (success, failed, skipped, etc.)
Get Task Logs
bash
af tasks logs <dag_id> <dag_run_id> <task_id>

Example:

bash
af tasks logs my_dag manual__2025-01-14T... extract_data

For specific retry attempt:

bash
af tasks logs my_dag manual__2025-01-14T... extract_data --try 2

Look for:

  • Exception messages and stack traces
  • Connection errors (database, API, S3)
  • Permission errors
  • Timeout errors
  • Missing dependencies
Check Upstream Tasks

If a task shows upstream_failed, the root cause is in an upstream task. Use af runs diagnose to find which task actually failed.

Check Import Errors (If DAG Didn't Run)

If the trigger failed because the DAG doesn't exist:

bash
af dags errors

This reveals syntax errors or missing dependencies that prevented the DAG from loading.


Phase 3: Fix and Retest

Once you identify the issue:

Common Fixes
IssueFix
Missing importAdd to DAG file
Missing packageAdd to requirements.txt
Connection errorCheck af config connections, verify credentials
Variable missingCheck af config variables, create if needed
TimeoutIncrease task timeout or optimize query
Permission errorCheck credentials in connection
After Fixing
  1. Save the file
  2. Retest: af runs trigger-wait <dag_id>

Repeat the test → debug → fix loop until the DAG succeeds.


Show full SKILL.md (327 more words)Show less

CLI Quick Reference

PhaseCommandPurpose
Testaf runs trigger-wait <dag_id>Primary test method — start here
Testaf runs trigger <dag_id>Start run (alternative)
Testaf runs get <dag_id> <run_id>Check run status
Debugaf runs diagnose <dag_id> <run_id>Comprehensive failure diagnosis
Debugaf tasks logs <dag_id> <run_id> <task_id>Get task output/errors
Debugaf dags errorsCheck for parse errors (if DAG won't load)
Debugaf dags get <dag_id>Verify DAG config
Debugaf dags explore <dag_id>Full DAG inspection
Configaf config connectionsList connections
Configaf config variablesList variables

Testing Scenarios

Scenario 1: Test a DAG (Happy Path)
bash
af runs trigger-wait my_dag
# Success! Done.
Scenario 2: Test a DAG (With Failure)
bash
# 1. Run and wait
af runs trigger-wait my_dag
# Failed...

# 2. Find failed tasks
af runs diagnose my_dag manual__2025-01-14T...

# 3. Get error details
af tasks logs my_dag manual__2025-01-14T... extract_data

# 4. [Fix the issue in DAG code]

# 5. Retest
af runs trigger-wait my_dag
Scenario 3: DAG Doesn't Exist / Won't Load
bash
# 1. Trigger fails - DAG not found
af runs trigger-wait my_dag
# Error: DAG not found

# 2. Find parse error
af dags errors

# 3. [Fix the issue in DAG code]

# 4. Retest
af runs trigger-wait my_dag
Scenario 4: Debug a Failed Scheduled Run
bash
# 1. Get failure summary
af runs diagnose my_dag scheduled__2025-01-14T...

# 2. Get error from failed task
af tasks logs my_dag scheduled__2025-01-14T... failed_task_id

# 3. [Fix the issue]

# 4. Retest
af runs trigger-wait my_dag
Scenario 5: Test with Custom Configuration
bash
af runs trigger-wait my_dag --conf '{"env": "staging", "batch_size": 100}' --timeout 600
Scenario 6: Long-Running DAG
bash
# Wait up to 1 hour
af runs trigger-wait my_dag --timeout 3600

# If timed out, check current state
af runs get my_dag manual__2025-01-14T...

Debugging Tips

Common Error Patterns

Connection Refused / Timeout:

  • Check af config connections for correct host/port
  • Verify network connectivity to external system
  • Check if connection credentials are correct

ModuleNotFoundError:

  • Package missing from requirements.txt
  • After adding, may need environment restart

PermissionError:

  • Check IAM roles, database grants, API keys
  • Verify connection has correct credentials

Task Timeout:

  • Query or operation taking too long
  • Consider adding timeout parameter to task
  • Optimize underlying query/operation
Reading Task Logs

Task logs typically show:

  1. Task start timestamp
  2. Any print/log statements from task code
  3. Return value (for @task decorated functions)
  4. Exception + full stack trace (if failed)
  5. Task end timestamp and duration

Focus on the exception at the bottom of failed task logs.

On Astro

Astro deployments support environment promotion, which helps structure your testing workflow:

  • Dev deployment: Test DAGs freely with astro deploy --dags for fast iteration
  • Staging deployment: Run integration tests against production-like data
  • Production deployment: Deploy only after validation in lower environments
  • Use separate Astro deployments for each environment and promote code through them

  • authoring-dags: For creating new DAGs (includes validation before testing)
  • debugging-dags: For general Airflow troubleshooting
  • deploying-airflow: For deploying DAGs to production after testing

© astronomer, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/testing-dags of astronomer/agents.

Open the folder on GitHubat commit 1ec1a1f

Compare with similar skills

Testing Dags next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Testing Dags compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Testing Dags this skillastronomer/agents450—~2.5kAutomated safety check: PassApache-2.0
Chart Testsastronomer/airflow-chart297—~2.8kAutomated safety check: PassCustom licence
Functional Testsastronomer/airflow-chart297—~2.2kAutomated safety check: PassCustom licence
Helm Chartastronomer/airflow-chart297—~6.4kAutomated safety check: PassCustom licence
Upgrading Mwaa Environmentsaws/agent-toolkit-for-aws2.8k—~7.3kAutomated safety check: PassApache-2.0
Testing Mwaa Workflowaws/agent-toolkit-for-aws2.8k—~3.8kAutomated safety check: PassApache-2.0

Similar skills

  • Chart Tests

    astronomer/airflow-chart

    A skill your agent uses when writing, editing, reviewing, or running Helm chart tests for the Astronomer airflow-chart repository.

    297 GitHub stars~2.8k tokensUpdated today
    Data & AnalyticsAuto-check passed
  • Functional Tests

    astronomer/airflow-chart

    A skill your agent uses when writing, editing, reviewing, or running functional (end-to-end) tests for the Astronomer airflow-chart repository.

    297 GitHub stars~2.2k tokensUpdated today
    Data & AnalyticsAuto-check passed
  • Helm Chart

    astronomer/airflow-chart

    A skill your agent uses for Helm chart work - creating charts, modifying existing charts, values design, testing.

    297 GitHub stars~6.4k tokensUpdated today
    Data & AnalyticsAuto-check passed
  • Upgrading Mwaa Environments

    aws/agent-toolkit-for-aws

    Official

    Upgrades an MWAA environment to a newer Airflow version — within 2.x, within 3.x, or across the 2.x-to-3.x boundary.

    2.8k GitHub stars~7.3k tokensUpdated today
    Data & AnalyticsAuto-check passed
  • Testing Mwaa Workflow

    aws/agent-toolkit-for-aws

    Official

    Tests Amazon MWAA workflow execution end-to-end: trigger a run and monitor it to completion for Provisioned (Python DAG, via Airflow REST API) and Serverless (YAML workflow, via StartWorkflowRun).

    2.8k GitHub stars~3.8k tokensUpdated today
    Backend & APIsAuto-check passed
  • Create Example

    godatadriven/whirl

    Create a new Whirl example project in the examples/ directory.

    205 GitHub stars~1.1k tokensUpdated 6 days ago
    Data & AnalyticsAuto-check passed

More from astronomer/agents

All 34 skills in this repo
  • Analyzing Data

    astronomer/agents

    Queries the data warehouse with SQL and answers business questions about data.

    450 GitHub stars~1.3k tokensUpdated yesterday
    Auto-check passed
  • Airflow

    astronomer/agents

    Queries, manages, and troubleshoots Apache Airflow using the af CLI.

    450 GitHub starsUsed in 1 repo~3.8k tokens
    Auto-check passed
  • Guide for migrating Dagster projects to Apache Airflow 3 on Astro.

    450 GitHub stars~3.8k tokensUpdated yesterday
    Auto-check passed
  • Authoring Dags

    astronomer/agents

    Workflow and best practices for writing Apache Airflow DAGs.

    450 GitHub starsUsed in 1 repo~1.8k tokens
    Auto-check passed
  • Deploying Airflow

    astronomer/agents

    Deploys Airflow DAGs and projects. An agent skill from astronomer/agents.

    450 GitHub starsUsed in 1 repo~2.8k tokens
    Auto-check passed
  • Trace downstream data lineage and impact analysis. An agent skill from astronomer/agents.

    450 GitHub starsUsed in 1 repo~1.2k tokens
    Auto-check passed

Questions about Testing Dags

What does Testing Dags do?

Complex DAG testing workflows with debugging and fixing cycles. Testing Dags is an agent skill from astronomer/agents. Complex DAG testing workflows with debugging and fixing cycles.

When should I use Testing Dags?

Testing Dags fits situations like: multi-step testing requests like test this dag and fix it if it fails; run the pipeline and troubleshoot issues.

How do I install Testing Dags in Claude Code?

Run `npx skills add astronomer/agents --skill testing-dags -a claude-code`. Or copy the skill folder (skills/testing-dags in astronomer/agents) into .claude/skills/testing-dags in your project. Claude Code loads it when a task matches its description.

How do I install Testing Dags in Codex?

Run `npx skills add astronomer/agents --skill testing-dags -a codex`. Or copy the skill folder (skills/testing-dags in astronomer/agents) into .agents/skills/testing-dags in your project. Codex loads it when a task matches its description.

Can I use Testing Dags in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add astronomer/agents --skill testing-dags -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/testing-dags, .gemini/skills/testing-dags, .github/skills/testing-dags and .opencode/skills/testing-dags in your project.

What does Testing Dags need to run?

Going by SKILL.md and its folder, Testing Dags needs the command-line tools its instructions call (uv).

Does Testing Dags access the network?

SKILL.md contains no URLs. Its commands use uv, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Testing Dags safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Testing Dags use?

Testing Dags is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Testing Dags use?

About 2.5k tokens (SKILL.md is roughly 9.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Testing Dags?

Skills that share tags, products or a category with Testing Dags: Chart Tests (astronomer/airflow-chart, 297 stars), Functional Tests (astronomer/airflow-chart, 297 stars), Helm Chart (astronomer/airflow-chart, 297 stars) and Upgrading Mwaa Environments (aws/agent-toolkit-for-aws, 2.8k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Testing Dags?

astronomer (a GitHub organization) maintains it in astronomer/agents, which has 450 GitHub stars. The repository holds 34 skills in this directory. The repository was last updated on October 5, 2026.

Source: astronomer/agents on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.