Agent skill

Creating Openlineage Extractors

by astronomer in astronomer/agents

Create custom OpenLineage extractors for Airflow operators. An agent skill from astronomer/agents.

Apache-2.0Auto-check passedData & Analytics

Install Creating Openlineage Extractors

skills CLI
$ npx skills add astronomer/agents --skill creating-openlineage-extractors -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install astronomer/agents creating-openlineage-extractors --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/astronomer/agents.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/creating-openlineage-extractors .claude/skills/creating-openlineage-extractors && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
creating-openlineage-extractors
GitHub stars
451
Token cost
~3.3k tokens
SKILL.md length
445 words
Files
1
Skills in repo
34
Repo updated
First seen
Licence
Apache-2.0

At a glance

Create custom OpenLineage extractors for Airflow operators. An agent skill from astronomer/agents.

  • Works in 5 steps: OpenLineage Methods (Recommended) → Custom Extractors → Circular Imports → …
  • The user needs lineage from unsupported
  • SKILL.md covers When to Use Each Approach, Two Approaches, Approach 1: OpenLineage… and Approach 2: Custom Extractors, plus 5 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Creating Openlineage Extractors is an agent skill from astronomer/agents. Create custom OpenLineage extractors for Airflow operators. Use when the user needs lineage from unsupported or third-party operators, wants column-level lineage, or needs complex extraction logic beyond what inlets/outlets provide.

Its SKILL.md is about 3.3k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Data & Analytics, covering Data pipelines and ETL. It works with Apache Airflow and Astro. The repository describes itself as: AI agent tooling for data engineering workflows. The licence is Apache-2.0.

When your agent uses it

  • The user needs lineage from unsupported
  • Third-party operators
  • Wants column-level lineage
  • Needs complex extraction logic beyond what inlets/outlets provide

Example prompts

  • “/creating-openlineage-extractors”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. OpenLineage Methods (Recommended)
  2. Custom Extractors
  3. Circular Imports
  4. Wrong Import Path
  5. Not Handling None

What it can do on your machine

Read from SKILL.md and the folder at commit 486ee63. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python, bash and ini).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • airflow.apache.org

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Creating Openlineage Extractors loads about 3.3k tokens when it runs. Until then it costs about 66 tokens; SKILL.md has 445 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~66
When it runs · the whole SKILL.md, loaded when a task matches
~3.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from astronomer/agents at commit 486ee63, republished under its Apache-2.0 licence (© astronomer). 445 words, ~3,290 tokens.

Download SKILL.mdSave it as .claude/skills/creating-openlineage-extractors/SKILL.md (or your agent's skills folder).
name
creating-openlineage-extractors
description
Create custom OpenLineage extractors for Airflow operators. Use when the user needs lineage from unsupported or third-party operators, wants column-level lineage, or needs complex extraction logic beyond what inlets/outlets provide.

Creating OpenLineage Extractors

This skill guides you through creating custom OpenLineage extractors to capture lineage from Airflow operators that don't have built-in support.

Reference: See the OpenLineage provider developer guide for the latest patterns and list of supported operators/hooks.

When to Use Each Approach

ScenarioApproach
Operator you own/maintainOpenLineage Methods (recommended, simplest)
Third-party operator you can't modifyCustom Extractor
Need column-level lineageOpenLineage Methods or Custom Extractor
Complex extraction logicOpenLineage Methods or Custom Extractor
Simple table-level lineageInlets/Outlets (simplest, but lowest priority)

Important: Always prefer OpenLineage methods over custom extractors when possible. Extractors are harder to write, easier to diverge from operator behavior after changes, and harder to debug.

On Astro

Astro includes built-in OpenLineage integration — no additional transport configuration is needed. Lineage events are automatically collected and displayed in the Astro UI's Lineage tab. Custom extractors deployed to an Astro project are automatically picked up, so you only need to register them in airflow.cfg or via environment variable and deploy.


Two Approaches

Use when you can add methods directly to your custom operator. This is the go-to solution for operators you own.

2. Custom Extractors

Use when you need lineage from third-party or provider operators that you cannot modify.


When you own the operator, add OpenLineage methods directly:

python
from airflow.models import BaseOperator


class MyCustomOperator(BaseOperator):
    """Custom operator with built-in OpenLineage support."""

    def __init__(self, source_table: str, target_table: str, **kwargs):
        super().__init__(**kwargs)
        self.source_table = source_table
        self.target_table = target_table
        self._rows_processed = 0  # Set during execution

    def execute(self, context):
        # Do the actual work
        self._rows_processed = self._process_data()
        return self._rows_processed

    def get_openlineage_facets_on_start(self):
        """Called when task starts. Return known inputs/outputs."""
        # Import locally to avoid circular imports
        from openlineage.client.event_v2 import Dataset
        from airflow.providers.openlineage.extractors import OperatorLineage

        return OperatorLineage(
            inputs=[Dataset(namespace="postgres://db", name=self.source_table)],
            outputs=[Dataset(namespace="postgres://db", name=self.target_table)],
        )

    def get_openlineage_facets_on_complete(self, task_instance):
        """Called after success. Add runtime metadata."""
        from openlineage.client.event_v2 import Dataset
        from openlineage.client.facet_v2 import output_statistics_output_dataset
        from airflow.providers.openlineage.extractors import OperatorLineage

        return OperatorLineage(
            inputs=[Dataset(namespace="postgres://db", name=self.source_table)],
            outputs=[
                Dataset(
                    namespace="postgres://db",
                    name=self.target_table,
                    facets={
                        "outputStatistics": output_statistics_output_dataset.OutputStatisticsOutputDatasetFacet(
                            rowCount=self._rows_processed
                        )
                    },
                )
            ],
        )

    def get_openlineage_facets_on_failure(self, task_instance):
        """Called after failure. Optional - for partial lineage."""
        return None
OpenLineage Methods Reference
MethodWhen CalledRequired
get_openlineage_facets_on_start()Task enters RUNNINGNo
get_openlineage_facets_on_complete(ti)Task succeedsNo
get_openlineage_facets_on_failure(ti)Task failsNo

Implement only the methods you need. Unimplemented methods fall through to Hook-Level Lineage or inlets/outlets.


Approach 2: Custom Extractors

Use this approach only when you cannot modify the operator (e.g., third-party or provider operators).

Show full SKILL.md (171 more words)Show less
Basic Structure
python
from airflow.providers.openlineage.extractors.base import BaseExtractor, OperatorLineage
from openlineage.client.event_v2 import Dataset


class MyOperatorExtractor(BaseExtractor):
    """Extract lineage from MyCustomOperator."""

    @classmethod
    def get_operator_classnames(cls) -> list[str]:
        """Return operator class names this extractor handles."""
        return ["MyCustomOperator"]

    def _execute_extraction(self) -> OperatorLineage | None:
        """Called BEFORE operator executes. Use for known inputs/outputs."""
        # Access operator properties via self.operator
        source_table = self.operator.source_table
        target_table = self.operator.target_table

        return OperatorLineage(
            inputs=[
                Dataset(
                    namespace="postgres://mydb:5432",
                    name=f"public.{source_table}",
                )
            ],
            outputs=[
                Dataset(
                    namespace="postgres://mydb:5432",
                    name=f"public.{target_table}",
                )
            ],
        )

    def extract_on_complete(self, task_instance) -> OperatorLineage | None:
        """Called AFTER operator executes. Use for runtime-determined lineage."""
        # Access properties set during execution
        # Useful for operators that determine outputs at runtime
        return None
OperatorLineage Structure
python
from airflow.providers.openlineage.extractors.base import OperatorLineage
from openlineage.client.event_v2 import Dataset
from openlineage.client.facet_v2 import sql_job

lineage = OperatorLineage(
    inputs=[Dataset(namespace="...", name="...")],      # Input datasets
    outputs=[Dataset(namespace="...", name="...")],     # Output datasets
    run_facets={"sql": sql_job.SQLJobFacet(query="SELECT...")},  # Run metadata
    job_facets={},                                      # Job metadata
)
Extraction Methods
MethodWhen CalledUse For
_execute_extraction()Before operator runsStatic/known lineage
extract_on_complete(task_instance)After successRuntime-determined lineage
extract_on_failure(task_instance)After failurePartial lineage on errors
Registering Extractors

Option 1: Configuration file (airflow.cfg)

ini
[openlineage]
extractors = mypackage.extractors.MyOperatorExtractor;mypackage.extractors.AnotherExtractor

Option 2: Environment variable

bash
AIRFLOW__OPENLINEAGE__EXTRACTORS='mypackage.extractors.MyOperatorExtractor;mypackage.extractors.AnotherExtractor'

Important: The path must be importable from the Airflow worker. Place extractors in your DAGs folder or installed package.


Common Patterns

SQL Operator Extractor
python
from airflow.providers.openlineage.extractors.base import BaseExtractor, OperatorLineage
from openlineage.client.event_v2 import Dataset
from openlineage.client.facet_v2 import sql_job


class MySqlOperatorExtractor(BaseExtractor):
    @classmethod
    def get_operator_classnames(cls) -> list[str]:
        return ["MySqlOperator"]

    def _execute_extraction(self) -> OperatorLineage | None:
        sql = self.operator.sql
        conn_id = self.operator.conn_id

        # Parse SQL to find tables (simplified example)
        # In practice, use a SQL parser like sqlglot
        inputs, outputs = self._parse_sql(sql)

        namespace = f"postgres://{conn_id}"

        return OperatorLineage(
            inputs=[Dataset(namespace=namespace, name=t) for t in inputs],
            outputs=[Dataset(namespace=namespace, name=t) for t in outputs],
            job_facets={
                "sql": sql_job.SQLJobFacet(query=sql)
            },
        )

    def _parse_sql(self, sql: str) -> tuple[list[str], list[str]]:
        """Parse SQL to extract table names. Use sqlglot for real parsing."""
        # Simplified example - use proper SQL parser in production
        inputs = []
        outputs = []
        # ... parsing logic ...
        return inputs, outputs
File Transfer Extractor
python
from airflow.providers.openlineage.extractors.base import BaseExtractor, OperatorLineage
from openlineage.client.event_v2 import Dataset


class S3ToSnowflakeExtractor(BaseExtractor):
    @classmethod
    def get_operator_classnames(cls) -> list[str]:
        return ["S3ToSnowflakeOperator"]

    def _execute_extraction(self) -> OperatorLineage | None:
        s3_bucket = self.operator.s3_bucket
        s3_key = self.operator.s3_key
        table = self.operator.table
        schema = self.operator.schema

        return OperatorLineage(
            inputs=[
                Dataset(
                    namespace=f"s3://{s3_bucket}",
                    name=s3_key,
                )
            ],
            outputs=[
                Dataset(
                    namespace="snowflake://myaccount.snowflakecomputing.com",
                    name=f"{schema}.{table}",
                )
            ],
        )
Dynamic Lineage from Execution
python
from openlineage.client.event_v2 import Dataset


class DynamicOutputExtractor(BaseExtractor):
    @classmethod
    def get_operator_classnames(cls) -> list[str]:
        return ["DynamicOutputOperator"]

    def _execute_extraction(self) -> OperatorLineage | None:
        # Only inputs known before execution
        return OperatorLineage(
            inputs=[Dataset(namespace="...", name=self.operator.source)],
        )

    def extract_on_complete(self, task_instance) -> OperatorLineage | None:
        # Outputs determined during execution
        # Access via operator properties set in execute()
        outputs = self.operator.created_tables  # Set during execute()

        return OperatorLineage(
            inputs=[Dataset(namespace="...", name=self.operator.source)],
            outputs=[Dataset(namespace="...", name=t) for t in outputs],
        )

Common Pitfalls

1. Circular Imports

Problem: Importing Airflow modules at the top level causes circular imports.

python
# ❌ BAD - can cause circular import issues
from airflow.models import TaskInstance
from openlineage.client.event_v2 import Dataset

class MyExtractor(BaseExtractor):
    ...
python
# ✅ GOOD - import inside methods
class MyExtractor(BaseExtractor):
    def _execute_extraction(self):
        from openlineage.client.event_v2 import Dataset
        # ...
2. Wrong Import Path

Problem: Extractor path doesn't match actual module location.

bash
# ❌ Wrong - path doesn't exist
AIRFLOW__OPENLINEAGE__EXTRACTORS='extractors.MyExtractor'

# ✅ Correct - full importable path
AIRFLOW__OPENLINEAGE__EXTRACTORS='dags.extractors.my_extractor.MyExtractor'
3. Not Handling None

Problem: Extraction fails when operator properties are None.

python
# ✅ Handle optional properties
def _execute_extraction(self) -> OperatorLineage | None:
    if not self.operator.source_table:
        return None  # Skip extraction

    return OperatorLineage(...)

Testing Extractors

Unit Testing
python
import pytest
from unittest.mock import MagicMock
from mypackage.extractors import MyOperatorExtractor


def test_extractor():
    # Mock the operator
    operator = MagicMock()
    operator.source_table = "input_table"
    operator.target_table = "output_table"

    # Create extractor
    extractor = MyOperatorExtractor(operator)

    # Test extraction
    lineage = extractor._execute_extraction()

    assert len(lineage.inputs) == 1
    assert lineage.inputs[0].name == "input_table"
    assert len(lineage.outputs) == 1
    assert lineage.outputs[0].name == "output_table"

Precedence Rules

OpenLineage checks for lineage in this order:

  1. Custom Extractors (highest priority)
  2. OpenLineage Methods on operator
  3. Hook-Level Lineage (from HookLineageCollector)
  4. Inlets/Outlets (lowest priority)

If a custom extractor exists, it overrides built-in extraction and inlets/outlets.


  • annotating-task-lineage: For simple table-level lineage with inlets/outlets
  • tracing-upstream-lineage: Investigate data origins
  • tracing-downstream-lineage: Investigate data dependencies

© astronomer, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/creating-openlineage-extractors of astronomer/agents.

Open the folder on GitHubat commit 486ee63

Compare with similar skills

Creating Openlineage Extractors next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Creating Openlineage Extractors compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Creating Openlineage Extractors this skillastronomer/agents451—~3.3kAutomated safety check: PassApache-2.0
Chart Testsastronomer/airflow-chart297—~2.8kAutomated safety check: PassCustom licence
Functional Testsastronomer/airflow-chart297—~2.2kAutomated safety check: PassCustom licence
Helm Chartastronomer/airflow-chart297—~6.4kAutomated safety check: PassCustom licence
Create Examplegodatadriven/whirl205—~1.1kAutomated safety check: PassApache-2.0
Senior Data Engineerbenchflow-ai/skillsbench1.8k—~5.9kAutomated safety check: PassMIT

Similar skills

  • Chart Tests

    astronomer/airflow-chart

    A skill your agent uses when writing, editing, reviewing, or running Helm chart tests for the Astronomer airflow-chart repository.

    297 GitHub stars~2.8k tokensUpdated 3 days ago
    Data & AnalyticsAuto-check passed
  • Functional Tests

    astronomer/airflow-chart

    A skill your agent uses when writing, editing, reviewing, or running functional (end-to-end) tests for the Astronomer airflow-chart repository.

    297 GitHub stars~2.2k tokensUpdated 3 days ago
    Data & AnalyticsAuto-check passed
  • Helm Chart

    astronomer/airflow-chart

    A skill your agent uses for Helm chart work - creating charts, modifying existing charts, values design, testing.

    297 GitHub stars~6.4k tokensUpdated 3 days ago
    Data & AnalyticsAuto-check passed
  • Create Example

    godatadriven/whirl

    Create a new Whirl example project in the examples/ directory.

    205 GitHub stars~1.1k tokensUpdated 9 days ago
    Data & AnalyticsAuto-check passed
  • Senior Data Engineer

    benchflow-ai/skillsbench

    World-class data engineering skill for building scalable data pipelines, ETL/ELT systems, real-time streaming, and data infrastructure.

    1.8k GitHub stars~5.9k tokensUpdated 2 mo ago
    Data & AnalyticsAuto-check passed
  • Airflow DAG Patterns

    wshobson/agents

    Patterns for writing production-ready Apache Airflow DAGs: task dependencies, custom operators and sensors, local testing, and rules for what to avoid.

    40k GitHub starsUsed in 9 repos~784 tokens
    Data & AnalyticsAuto-check passed

More from astronomer/agents

All 34 skills in this repo
  • Analyzing Data

    astronomer/agents

    Queries the data warehouse with SQL and answers business questions about data.

    451 GitHub stars~1.3k tokensUpdated 3 days ago
    Auto-check passed
  • Airflow

    astronomer/agents

    Queries, manages, and troubleshoots Apache Airflow using the af CLI.

    451 GitHub starsUsed in 1 repo~3.8k tokens
    Auto-check passed
  • Guide for migrating Dagster projects to Apache Airflow 3 on Astro.

    451 GitHub stars~3.8k tokensUpdated 3 days ago
    Auto-check passed
  • Authoring Dags

    astronomer/agents

    Workflow and best practices for writing Apache Airflow DAGs.

    451 GitHub starsUsed in 1 repo~1.8k tokens
    Auto-check passed
  • Deploying Airflow

    astronomer/agents

    Deploys Airflow DAGs and projects. An agent skill from astronomer/agents.

    451 GitHub starsUsed in 1 repo~2.8k tokens
    Auto-check passed
  • Airflow Hitl

    astronomer/agents

    Builds human-in-the-loop (HITL) Airflow workflows - approval gates, form input, and human-driven branching.

    451 GitHub stars~1.8k tokensUpdated 3 days ago
    Auto-check passed

Questions about Creating Openlineage Extractors

What does Creating Openlineage Extractors do?

Create custom OpenLineage extractors for Airflow operators. An agent skill from astronomer/agents. Creating Openlineage Extractors is an agent skill from astronomer/agents. Create custom OpenLineage extractors for Airflow operators.

When should I use Creating Openlineage Extractors?

Creating Openlineage Extractors fits situations like: the user needs lineage from unsupported; third-party operators; wants column-level lineage; needs complex extraction logic beyond what inlets/outlets provide.

How do I install Creating Openlineage Extractors in Claude Code?

Run `npx skills add astronomer/agents --skill creating-openlineage-extractors -a claude-code`. Or copy the skill folder (skills/creating-openlineage-extractors in astronomer/agents) into .claude/skills/creating-openlineage-extractors in your project. Claude Code loads it when a task matches its description.

How do I install Creating Openlineage Extractors in Codex?

Run `npx skills add astronomer/agents --skill creating-openlineage-extractors -a codex`. Or copy the skill folder (skills/creating-openlineage-extractors in astronomer/agents) into .agents/skills/creating-openlineage-extractors in your project. Codex loads it when a task matches its description.

Can I use Creating Openlineage Extractors in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add astronomer/agents --skill creating-openlineage-extractors -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/creating-openlineage-extractors, .gemini/skills/creating-openlineage-extractors, .github/skills/creating-openlineage-extractors and .opencode/skills/creating-openlineage-extractors in your project.

What does Creating Openlineage Extractors need to run?

SKILL.md names no scripts, command-line tools or credentials: Creating Openlineage Extractors is instructions for the agent only. Our summary lists: Python 3.

Does Creating Openlineage Extractors access the network?

SKILL.md names 1 domain. As links in the text: airflow.apache.org. This is read from the text; nothing was executed.

Is Creating Openlineage Extractors safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Creating Openlineage Extractors use?

Creating Openlineage Extractors is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Creating Openlineage Extractors use?

About 3.3k tokens (SKILL.md is roughly 13k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Creating Openlineage Extractors?

Skills that share tags, products or a category with Creating Openlineage Extractors: Chart Tests (astronomer/airflow-chart, 297 stars), Functional Tests (astronomer/airflow-chart, 297 stars), Helm Chart (astronomer/airflow-chart, 297 stars) and Create Example (godatadriven/whirl, 205 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Creating Openlineage Extractors?

astronomer (a GitHub organization) maintains it in astronomer/agents, which has 451 GitHub stars. The repository holds 34 skills in this directory. The repository was last updated on October 7, 2026.

Source: astronomer/agents on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.