Agent skill

Tracing Upstream Lineage

by astronomer in astronomer/agents

Trace upstream data lineage. An agent skill from astronomer/agents.

Apache-2.0Auto-check passedData & Analytics

Install Tracing Upstream Lineage

skills CLI
$ npx skills add astronomer/agents --skill tracing-upstream-lineage -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install astronomer/agents tracing-upstream-lineage --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/astronomer/agents.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/tracing-upstream-lineage .claude/skills/tracing-upstream-lineage && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
tracing-upstream-lineage
GitHub stars
451
Token cost
~1.1k tokens
SKILL.md length
447 words
Files
1
Skills in repo
34
Repo updated
First seen
Licence
Apache-2.0

At a glance

Trace upstream data lineage. An agent skill from astronomer/agents.

  • Works in 5 steps: Identify the Target Type → Find the Producing DAG → Trace Data Sources → …
  • The user asks where data comes from
  • SKILL.md covers Lineage Investigation, Lineage for Columns and Output: Lineage Report
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Tracing Upstream Lineage is an agent skill from astronomer/agents. Trace upstream data lineage. Use when the user asks where data comes from, what feeds a table, upstream dependencies, data sources, or needs to understand data origins.

Its SKILL.md is about 1.1k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Data & Analytics, covering Data pipelines and ETL and Data governance. It works with Astro. The repository describes itself as: AI agent tooling for data engineering workflows. The licence is Apache-2.0.

When your agent uses it

  • The user asks where data comes from
  • What feeds a table
  • Upstream dependencies
  • Needs to understand data origins

Example prompts

  • “/tracing-upstream-lineage”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Identify the Target Type
  2. Find the Producing DAG
  3. Trace Data Sources
  4. Build the Lineage Chain
  5. Check Source Health

What it can do on your machine

Read from SKILL.md and the folder at commit 486ee63. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Tracing Upstream Lineage loads about 1.1k tokens when it runs. Until then it costs about 48 tokens; SKILL.md has 447 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~48
When it runs · the whole SKILL.md, loaded when a task matches
~1.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from astronomer/agents at commit 486ee63, republished under its Apache-2.0 licence (© astronomer). 447 words, ~1,125 tokens.

Download SKILL.mdSave it as .claude/skills/tracing-upstream-lineage/SKILL.md (or your agent's skills folder).
name
tracing-upstream-lineage
description
Trace upstream data lineage. Use when the user asks where data comes from, what feeds a table, upstream dependencies, data sources, or needs to understand data origins.

Upstream Lineage: Sources

Trace the origins of data - answer "Where does this data come from?"

Lineage Investigation

Step 1: Identify the Target Type

Determine what we're tracing:

  • Table: Trace what populates this table
  • Column: Trace where this specific column comes from
  • DAG: Trace what data sources this DAG reads from
Step 2: Find the Producing DAG

Tables are typically populated by Airflow DAGs. Find the connection:

  1. Search DAGs by name: Use af dags list and look for DAG names matching the table name

    • load_customers -> customers table
    • etl_daily_orders -> orders table
  2. Explore DAG source code: Use af dags source <dag_id> to read the DAG definition

    • Look for INSERT, MERGE, CREATE TABLE statements
    • Find the target table in the code
  3. Check DAG tasks: Use af tasks list <dag_id> to see what operations the DAG performs

On Astro

If you're running on Astro, the Lineage tab in the Astro UI provides visual lineage exploration across DAGs and datasets. Use it to quickly trace upstream dependencies without manually searching DAG source code.

On OSS Airflow

Use DAG source code and task logs to trace lineage (no built-in cross-DAG UI).

Step 3: Trace Data Sources

From the DAG code, identify source tables and systems:

SQL Sources (look for FROM clauses):

python
# In DAG code:
SELECT * FROM source_schema.source_table  # <- This is an upstream source

External Sources (look for connection references):

  • S3Operator -> S3 bucket source
  • PostgresOperator -> Postgres database source
  • SalesforceOperator -> Salesforce API source
  • HttpOperator -> REST API source

File Sources:

  • CSV/Parquet files in object storage
  • SFTP drops
  • Local file paths
Step 4: Build the Lineage Chain

Recursively trace each source:

TARGET: analytics.orders_daily
    ^
    +-- DAG: etl_daily_orders
            ^
            +-- SOURCE: raw.orders (table)
            |       ^
            |       +-- DAG: ingest_orders
            |               ^
            |               +-- SOURCE: Salesforce API (external)
            |
            +-- SOURCE: dim.customers (table)
                    ^
                    +-- DAG: load_customers
                            ^
                            +-- SOURCE: PostgreSQL (external DB)
Step 5: Check Source Health

For each upstream source:

  • Tables: Check freshness with the checking-freshness skill
  • DAGs: Check recent run status with af dags stats
  • External systems: Note connection info from DAG code
Show full SKILL.md (163 more words)Show less

Lineage for Columns

When tracing a specific column:

  1. Find the column in the target table schema
  2. Search DAG source code for references to that column name
  3. Trace through transformations:
    • Direct mappings: source.col AS target_col
    • Transformations: COALESCE(a.col, b.col) AS target_col
    • Aggregations: SUM(detail.amount) AS total_amount

Output: Lineage Report

Summary

One-line answer: "This table is populated by DAG X from sources Y and Z"

Lineage Diagram
[Salesforce] --> [raw.opportunities] --> [stg.opportunities] --> [fct.sales]
                        |                        |
                   DAG: ingest_sfdc         DAG: transform_sales
Source Details
SourceTypeConnectionFreshnessOwner
raw.ordersTableInternal2h agodata-team
SalesforceAPIsalesforce_connReal-timesales-ops
Transformation Chain

Describe how data flows and transforms:

  1. Raw data lands in raw.orders via Salesforce API sync
  2. DAG transform_orders cleans and dedupes into stg.orders
  3. DAG build_order_facts joins with dimensions into fct.orders
Data Quality Implications
  • Single points of failure?
  • Stale upstream sources?
  • Complex transformation chains that could break?
  • Check source freshness: checking-freshness skill
  • Debug source DAG: debugging-dags skill
  • Trace downstream impacts: tracing-downstream-lineage skill
  • Add manual lineage annotations: annotating-task-lineage skill
  • Build custom lineage extractors: creating-openlineage-extractors skill

© astronomer, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/tracing-upstream-lineage of astronomer/agents.

Open the folder on GitHubat commit 486ee63

Compare with similar skills

Tracing Upstream Lineage next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Tracing Upstream Lineage compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Tracing Upstream Lineage this skillastronomer/agents451—~1.1kAutomated safety check: PassApache-2.0
Chart Testsastronomer/airflow-chart297—~2.8kAutomated safety check: PassCustom licence
Functional Testsastronomer/airflow-chart297—~2.2kAutomated safety check: PassCustom licence
Data Quality Frameworkswshobson/agents40k11 repos~1.1kAutomated safety check: PassMIT
Helm Chartastronomer/airflow-chart297—~6.4kAutomated safety check: PassCustom licence
Monte Carlo Preventsickn33/agentic-awesome-skills47k1 repos~3.3kAutomated safety check: PassMIT

Similar skills

  • Chart Tests

    astronomer/airflow-chart

    A skill your agent uses when writing, editing, reviewing, or running Helm chart tests for the Astronomer airflow-chart repository.

    297 GitHub stars~2.8k tokensUpdated 3 days ago
    Data & AnalyticsAuto-check passed
  • Functional Tests

    astronomer/airflow-chart

    A skill your agent uses when writing, editing, reviewing, or running functional (end-to-end) tests for the Astronomer airflow-chart repository.

    297 GitHub stars~2.2k tokensUpdated 3 days ago
    Data & AnalyticsAuto-check passed
  • Sets up data quality checks with Great Expectations, dbt tests and data contracts, with checkpoints and pass-fail reports for pipelines.

    40k GitHub starsUsed in 11 repos~1.1k tokens
    Data & AnalyticsAuto-check passed
  • Helm Chart

    astronomer/airflow-chart

    A skill your agent uses for Helm chart work - creating charts, modifying existing charts, values design, testing.

    297 GitHub stars~6.4k tokensUpdated 3 days ago
    Data & AnalyticsAuto-check passed
  • Monte Carlo Prevent

    sickn33/agentic-awesome-skills

    Surfaces Monte Carlo data observability context (table health, alerts, lineage, blast radius) before SQL/dbt edits.

    47k GitHub starsUsed in 1 repo~3.3k tokens
    Data & AnalyticsAuto-check passed
  • Official

    Shared foundations for building reusable data models in PostHog, on either of two stacks: PostHog-native data-warehouse views / materialized views (HogQL, via the view- MCP tools), or an external…

    40k GitHub stars~2.1k tokensUpdated today
    Data & AnalyticsAuto-check passed

More from astronomer/agents

All 34 skills in this repo
  • Analyzing Data

    astronomer/agents

    Queries the data warehouse with SQL and answers business questions about data.

    451 GitHub stars~1.3k tokensUpdated 3 days ago
    Auto-check passed
  • Airflow

    astronomer/agents

    Queries, manages, and troubleshoots Apache Airflow using the af CLI.

    451 GitHub starsUsed in 1 repo~3.8k tokens
    Auto-check passed
  • Guide for migrating Dagster projects to Apache Airflow 3 on Astro.

    451 GitHub stars~3.8k tokensUpdated 3 days ago
    Auto-check passed
  • Authoring Dags

    astronomer/agents

    Workflow and best practices for writing Apache Airflow DAGs.

    451 GitHub starsUsed in 1 repo~1.8k tokens
    Auto-check passed
  • Deploying Airflow

    astronomer/agents

    Deploys Airflow DAGs and projects. An agent skill from astronomer/agents.

    451 GitHub starsUsed in 1 repo~2.8k tokens
    Auto-check passed
  • Airflow Hitl

    astronomer/agents

    Builds human-in-the-loop (HITL) Airflow workflows - approval gates, form input, and human-driven branching.

    451 GitHub stars~1.8k tokensUpdated 3 days ago
    Auto-check passed

Works with

Questions about Tracing Upstream Lineage

What does Tracing Upstream Lineage do?

Trace upstream data lineage. An agent skill from astronomer/agents. Tracing Upstream Lineage is an agent skill from astronomer/agents. Trace upstream data lineage.

When should I use Tracing Upstream Lineage?

Tracing Upstream Lineage fits situations like: the user asks where data comes from; what feeds a table; upstream dependencies; needs to understand data origins.

How do I install Tracing Upstream Lineage in Claude Code?

Run `npx skills add astronomer/agents --skill tracing-upstream-lineage -a claude-code`. Or copy the skill folder (skills/tracing-upstream-lineage in astronomer/agents) into .claude/skills/tracing-upstream-lineage in your project. Claude Code loads it when a task matches its description.

How do I install Tracing Upstream Lineage in Codex?

Run `npx skills add astronomer/agents --skill tracing-upstream-lineage -a codex`. Or copy the skill folder (skills/tracing-upstream-lineage in astronomer/agents) into .agents/skills/tracing-upstream-lineage in your project. Codex loads it when a task matches its description.

Can I use Tracing Upstream Lineage in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add astronomer/agents --skill tracing-upstream-lineage -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/tracing-upstream-lineage, .gemini/skills/tracing-upstream-lineage, .github/skills/tracing-upstream-lineage and .opencode/skills/tracing-upstream-lineage in your project.

What does Tracing Upstream Lineage need to run?

SKILL.md names no scripts, command-line tools or credentials: Tracing Upstream Lineage is instructions for the agent only. Our summary lists: Python 3.

Does Tracing Upstream Lineage access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Tracing Upstream Lineage safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Tracing Upstream Lineage use?

Tracing Upstream Lineage is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Tracing Upstream Lineage use?

About 1.1k tokens (SKILL.md is roughly 4.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Tracing Upstream Lineage?

Skills that share tags, products or a category with Tracing Upstream Lineage: Chart Tests (astronomer/airflow-chart, 297 stars), Functional Tests (astronomer/airflow-chart, 297 stars), Data Quality Frameworks (wshobson/agents, 40k stars) and Helm Chart (astronomer/airflow-chart, 297 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Tracing Upstream Lineage?

astronomer (a GitHub organization) maintains it in astronomer/agents, which has 451 GitHub stars. The repository holds 34 skills in this directory. The repository was last updated on October 7, 2026.

Source: astronomer/agents on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.