Chart Tests
astronomer/airflow-chart
A skill your agent uses when writing, editing, reviewing, or running Helm chart tests for the Astronomer airflow-chart repository.
Create custom OpenLineage extractors for Airflow operators. An agent skill from astronomer/agents.
$ npx skills add astronomer/agents --skill creating-openlineage-extractors -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install astronomer/agents creating-openlineage-extractors --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/astronomer/agents.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/creating-openlineage-extractors .claude/skills/creating-openlineage-extractors && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "creating-openlineage-extractors" agent skill from https://github.com/astronomer/agents/tree/main/skills/creating-openlineage-extractors into .claude/skills/creating-openlineage-extractors/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "creating-openlineage-extractors", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/astronomer/agents/tree/main/skills/creating-openlineage-extractorsType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add astronomer/agents --skill creating-openlineage-extractors -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install astronomer/agents creating-openlineage-extractors --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/astronomer/agents.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/creating-openlineage-extractors .agents/skills/creating-openlineage-extractors && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "creating-openlineage-extractors" agent skill from https://github.com/astronomer/agents/tree/main/skills/creating-openlineage-extractors into .agents/skills/creating-openlineage-extractors/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "creating-openlineage-extractors", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add astronomer/agents --skill creating-openlineage-extractors -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install astronomer/agents creating-openlineage-extractors --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/astronomer/agents.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/creating-openlineage-extractors .cursor/skills/creating-openlineage-extractors && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "creating-openlineage-extractors" agent skill from https://github.com/astronomer/agents/tree/main/skills/creating-openlineage-extractors into .cursor/skills/creating-openlineage-extractors/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "creating-openlineage-extractors", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/astronomer/agents.git --path skills/creating-openlineage-extractors--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add astronomer/agents --skill creating-openlineage-extractors -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install astronomer/agents creating-openlineage-extractors --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/astronomer/agents.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/creating-openlineage-extractors .gemini/skills/creating-openlineage-extractors && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "creating-openlineage-extractors" agent skill from https://github.com/astronomer/agents/tree/main/skills/creating-openlineage-extractors into .gemini/skills/creating-openlineage-extractors/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "creating-openlineage-extractors", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install astronomer/agents creating-openlineage-extractorsInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add astronomer/agents --skill creating-openlineage-extractors -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/astronomer/agents.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/creating-openlineage-extractors .github/skills/creating-openlineage-extractors && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "creating-openlineage-extractors" agent skill from https://github.com/astronomer/agents/tree/main/skills/creating-openlineage-extractors into .github/skills/creating-openlineage-extractors/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "creating-openlineage-extractors", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add astronomer/agents --skill creating-openlineage-extractors -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install astronomer/agents creating-openlineage-extractors --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/astronomer/agents.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/creating-openlineage-extractors .opencode/skills/creating-openlineage-extractors && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "creating-openlineage-extractors" agent skill from https://github.com/astronomer/agents/tree/main/skills/creating-openlineage-extractors into .opencode/skills/creating-openlineage-extractors/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "creating-openlineage-extractors", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
creating-openlineage-extractorsCreate custom OpenLineage extractors for Airflow operators. An agent skill from astronomer/agents.
Creating Openlineage Extractors is an agent skill from astronomer/agents. Create custom OpenLineage extractors for Airflow operators. Use when the user needs lineage from unsupported or third-party operators, wants column-level lineage, or needs complex extraction logic beyond what inlets/outlets provide.
Its SKILL.md is about 3.3k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Data & Analytics, covering Data pipelines and ETL. It works with Apache Airflow and Astro. The repository describes itself as: AI agent tooling for data engineering workflows. The licence is Apache-2.0.
5 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 486ee63. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are python, bash and ini).
From the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
airflow.apache.orgFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Creating Openlineage Extractors loads about 3.3k tokens when it runs. Until then it costs about 66 tokens; SKILL.md has 445 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from astronomer/agents at commit 486ee63, republished under its Apache-2.0 licence (© astronomer). 445 words, ~3,290 tokens.
.claude/skills/creating-openlineage-extractors/SKILL.md (or your agent's skills folder).This skill guides you through creating custom OpenLineage extractors to capture lineage from Airflow operators that don't have built-in support.
Reference: See the OpenLineage provider developer guide for the latest patterns and list of supported operators/hooks.
| Scenario | Approach |
|---|---|
| Operator you own/maintain | OpenLineage Methods (recommended, simplest) |
| Third-party operator you can't modify | Custom Extractor |
| Need column-level lineage | OpenLineage Methods or Custom Extractor |
| Complex extraction logic | OpenLineage Methods or Custom Extractor |
| Simple table-level lineage | Inlets/Outlets (simplest, but lowest priority) |
Important: Always prefer OpenLineage methods over custom extractors when possible. Extractors are harder to write, easier to diverge from operator behavior after changes, and harder to debug.
Astro includes built-in OpenLineage integration — no additional transport configuration is needed. Lineage events are automatically collected and displayed in the Astro UI's Lineage tab. Custom extractors deployed to an Astro project are automatically picked up, so you only need to register them in airflow.cfg or via environment variable and deploy.
Use when you can add methods directly to your custom operator. This is the go-to solution for operators you own.
Use when you need lineage from third-party or provider operators that you cannot modify.
When you own the operator, add OpenLineage methods directly:
from airflow.models import BaseOperator
class MyCustomOperator(BaseOperator):
"""Custom operator with built-in OpenLineage support."""
def __init__(self, source_table: str, target_table: str, **kwargs):
super().__init__(**kwargs)
self.source_table = source_table
self.target_table = target_table
self._rows_processed = 0 # Set during execution
def execute(self, context):
# Do the actual work
self._rows_processed = self._process_data()
return self._rows_processed
def get_openlineage_facets_on_start(self):
"""Called when task starts. Return known inputs/outputs."""
# Import locally to avoid circular imports
from openlineage.client.event_v2 import Dataset
from airflow.providers.openlineage.extractors import OperatorLineage
return OperatorLineage(
inputs=[Dataset(namespace="postgres://db", name=self.source_table)],
outputs=[Dataset(namespace="postgres://db", name=self.target_table)],
)
def get_openlineage_facets_on_complete(self, task_instance):
"""Called after success. Add runtime metadata."""
from openlineage.client.event_v2 import Dataset
from openlineage.client.facet_v2 import output_statistics_output_dataset
from airflow.providers.openlineage.extractors import OperatorLineage
return OperatorLineage(
inputs=[Dataset(namespace="postgres://db", name=self.source_table)],
outputs=[
Dataset(
namespace="postgres://db",
name=self.target_table,
facets={
"outputStatistics": output_statistics_output_dataset.OutputStatisticsOutputDatasetFacet(
rowCount=self._rows_processed
)
},
)
],
)
def get_openlineage_facets_on_failure(self, task_instance):
"""Called after failure. Optional - for partial lineage."""
return None| Method | When Called | Required |
|---|---|---|
get_openlineage_facets_on_start() | Task enters RUNNING | No |
get_openlineage_facets_on_complete(ti) | Task succeeds | No |
get_openlineage_facets_on_failure(ti) | Task fails | No |
Implement only the methods you need. Unimplemented methods fall through to Hook-Level Lineage or inlets/outlets.
Use this approach only when you cannot modify the operator (e.g., third-party or provider operators).
from airflow.providers.openlineage.extractors.base import BaseExtractor, OperatorLineage
from openlineage.client.event_v2 import Dataset
class MyOperatorExtractor(BaseExtractor):
"""Extract lineage from MyCustomOperator."""
@classmethod
def get_operator_classnames(cls) -> list[str]:
"""Return operator class names this extractor handles."""
return ["MyCustomOperator"]
def _execute_extraction(self) -> OperatorLineage | None:
"""Called BEFORE operator executes. Use for known inputs/outputs."""
# Access operator properties via self.operator
source_table = self.operator.source_table
target_table = self.operator.target_table
return OperatorLineage(
inputs=[
Dataset(
namespace="postgres://mydb:5432",
name=f"public.{source_table}",
)
],
outputs=[
Dataset(
namespace="postgres://mydb:5432",
name=f"public.{target_table}",
)
],
)
def extract_on_complete(self, task_instance) -> OperatorLineage | None:
"""Called AFTER operator executes. Use for runtime-determined lineage."""
# Access properties set during execution
# Useful for operators that determine outputs at runtime
return Nonefrom airflow.providers.openlineage.extractors.base import OperatorLineage
from openlineage.client.event_v2 import Dataset
from openlineage.client.facet_v2 import sql_job
lineage = OperatorLineage(
inputs=[Dataset(namespace="...", name="...")], # Input datasets
outputs=[Dataset(namespace="...", name="...")], # Output datasets
run_facets={"sql": sql_job.SQLJobFacet(query="SELECT...")}, # Run metadata
job_facets={}, # Job metadata
)| Method | When Called | Use For |
|---|---|---|
_execute_extraction() | Before operator runs | Static/known lineage |
extract_on_complete(task_instance) | After success | Runtime-determined lineage |
extract_on_failure(task_instance) | After failure | Partial lineage on errors |
Option 1: Configuration file (airflow.cfg)
[openlineage]
extractors = mypackage.extractors.MyOperatorExtractor;mypackage.extractors.AnotherExtractorOption 2: Environment variable
AIRFLOW__OPENLINEAGE__EXTRACTORS='mypackage.extractors.MyOperatorExtractor;mypackage.extractors.AnotherExtractor'Important: The path must be importable from the Airflow worker. Place extractors in your DAGs folder or installed package.
from airflow.providers.openlineage.extractors.base import BaseExtractor, OperatorLineage
from openlineage.client.event_v2 import Dataset
from openlineage.client.facet_v2 import sql_job
class MySqlOperatorExtractor(BaseExtractor):
@classmethod
def get_operator_classnames(cls) -> list[str]:
return ["MySqlOperator"]
def _execute_extraction(self) -> OperatorLineage | None:
sql = self.operator.sql
conn_id = self.operator.conn_id
# Parse SQL to find tables (simplified example)
# In practice, use a SQL parser like sqlglot
inputs, outputs = self._parse_sql(sql)
namespace = f"postgres://{conn_id}"
return OperatorLineage(
inputs=[Dataset(namespace=namespace, name=t) for t in inputs],
outputs=[Dataset(namespace=namespace, name=t) for t in outputs],
job_facets={
"sql": sql_job.SQLJobFacet(query=sql)
},
)
def _parse_sql(self, sql: str) -> tuple[list[str], list[str]]:
"""Parse SQL to extract table names. Use sqlglot for real parsing."""
# Simplified example - use proper SQL parser in production
inputs = []
outputs = []
# ... parsing logic ...
return inputs, outputsfrom airflow.providers.openlineage.extractors.base import BaseExtractor, OperatorLineage
from openlineage.client.event_v2 import Dataset
class S3ToSnowflakeExtractor(BaseExtractor):
@classmethod
def get_operator_classnames(cls) -> list[str]:
return ["S3ToSnowflakeOperator"]
def _execute_extraction(self) -> OperatorLineage | None:
s3_bucket = self.operator.s3_bucket
s3_key = self.operator.s3_key
table = self.operator.table
schema = self.operator.schema
return OperatorLineage(
inputs=[
Dataset(
namespace=f"s3://{s3_bucket}",
name=s3_key,
)
],
outputs=[
Dataset(
namespace="snowflake://myaccount.snowflakecomputing.com",
name=f"{schema}.{table}",
)
],
)from openlineage.client.event_v2 import Dataset
class DynamicOutputExtractor(BaseExtractor):
@classmethod
def get_operator_classnames(cls) -> list[str]:
return ["DynamicOutputOperator"]
def _execute_extraction(self) -> OperatorLineage | None:
# Only inputs known before execution
return OperatorLineage(
inputs=[Dataset(namespace="...", name=self.operator.source)],
)
def extract_on_complete(self, task_instance) -> OperatorLineage | None:
# Outputs determined during execution
# Access via operator properties set in execute()
outputs = self.operator.created_tables # Set during execute()
return OperatorLineage(
inputs=[Dataset(namespace="...", name=self.operator.source)],
outputs=[Dataset(namespace="...", name=t) for t in outputs],
)Problem: Importing Airflow modules at the top level causes circular imports.
# ❌ BAD - can cause circular import issues
from airflow.models import TaskInstance
from openlineage.client.event_v2 import Dataset
class MyExtractor(BaseExtractor):
...# ✅ GOOD - import inside methods
class MyExtractor(BaseExtractor):
def _execute_extraction(self):
from openlineage.client.event_v2 import Dataset
# ...Problem: Extractor path doesn't match actual module location.
# ❌ Wrong - path doesn't exist
AIRFLOW__OPENLINEAGE__EXTRACTORS='extractors.MyExtractor'
# ✅ Correct - full importable path
AIRFLOW__OPENLINEAGE__EXTRACTORS='dags.extractors.my_extractor.MyExtractor'Problem: Extraction fails when operator properties are None.
# ✅ Handle optional properties
def _execute_extraction(self) -> OperatorLineage | None:
if not self.operator.source_table:
return None # Skip extraction
return OperatorLineage(...)import pytest
from unittest.mock import MagicMock
from mypackage.extractors import MyOperatorExtractor
def test_extractor():
# Mock the operator
operator = MagicMock()
operator.source_table = "input_table"
operator.target_table = "output_table"
# Create extractor
extractor = MyOperatorExtractor(operator)
# Test extraction
lineage = extractor._execute_extraction()
assert len(lineage.inputs) == 1
assert lineage.inputs[0].name == "input_table"
assert len(lineage.outputs) == 1
assert lineage.outputs[0].name == "output_table"OpenLineage checks for lineage in this order:
HookLineageCollector)If a custom extractor exists, it overrides built-in extraction and inlets/outlets.
© astronomer, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in skills/creating-openlineage-extractors of astronomer/agents.
Open the folder on GitHubat commit 486ee63
Creating Openlineage Extractors next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Creating Openlineage Extractors this skillastronomer/agents | 451 | — | ~3.3k | Automated safety check: Pass | Apache-2.0 | |
| Chart Testsastronomer/airflow-chart | 297 | — | ~2.8k | Automated safety check: Pass | Custom licence | |
| Functional Testsastronomer/airflow-chart | 297 | — | ~2.2k | Automated safety check: Pass | Custom licence | |
| Helm Chartastronomer/airflow-chart | 297 | — | ~6.4k | Automated safety check: Pass | Custom licence | |
| Create Examplegodatadriven/whirl | 205 | — | ~1.1k | Automated safety check: Pass | Apache-2.0 | |
| Senior Data Engineerbenchflow-ai/skillsbench | 1.8k | — | ~5.9k | Automated safety check: Pass | MIT |
astronomer/airflow-chart
A skill your agent uses when writing, editing, reviewing, or running Helm chart tests for the Astronomer airflow-chart repository.
astronomer/airflow-chart
A skill your agent uses when writing, editing, reviewing, or running functional (end-to-end) tests for the Astronomer airflow-chart repository.
astronomer/airflow-chart
A skill your agent uses for Helm chart work - creating charts, modifying existing charts, values design, testing.
godatadriven/whirl
Create a new Whirl example project in the examples/ directory.
benchflow-ai/skillsbench
World-class data engineering skill for building scalable data pipelines, ETL/ELT systems, real-time streaming, and data infrastructure.
wshobson/agents
Patterns for writing production-ready Apache Airflow DAGs: task dependencies, custom operators and sensors, local testing, and rules for what to avoid.
astronomer/agents
Queries the data warehouse with SQL and answers business questions about data.
astronomer/agents
Queries, manages, and troubleshoots Apache Airflow using the af CLI.
astronomer/agents
Guide for migrating Dagster projects to Apache Airflow 3 on Astro.
astronomer/agents
Workflow and best practices for writing Apache Airflow DAGs.
astronomer/agents
Deploys Airflow DAGs and projects. An agent skill from astronomer/agents.
astronomer/agents
Builds human-in-the-loop (HITL) Airflow workflows - approval gates, form input, and human-driven branching.
Works with
Categories
Create custom OpenLineage extractors for Airflow operators. An agent skill from astronomer/agents. Creating Openlineage Extractors is an agent skill from astronomer/agents. Create custom OpenLineage extractors for Airflow operators.
Creating Openlineage Extractors fits situations like: the user needs lineage from unsupported; third-party operators; wants column-level lineage; needs complex extraction logic beyond what inlets/outlets provide.
Run `npx skills add astronomer/agents --skill creating-openlineage-extractors -a claude-code`. Or copy the skill folder (skills/creating-openlineage-extractors in astronomer/agents) into .claude/skills/creating-openlineage-extractors in your project. Claude Code loads it when a task matches its description.
Run `npx skills add astronomer/agents --skill creating-openlineage-extractors -a codex`. Or copy the skill folder (skills/creating-openlineage-extractors in astronomer/agents) into .agents/skills/creating-openlineage-extractors in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add astronomer/agents --skill creating-openlineage-extractors -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/creating-openlineage-extractors, .gemini/skills/creating-openlineage-extractors, .github/skills/creating-openlineage-extractors and .opencode/skills/creating-openlineage-extractors in your project.
SKILL.md names no scripts, command-line tools or credentials: Creating Openlineage Extractors is instructions for the agent only. Our summary lists: Python 3.
SKILL.md names 1 domain. As links in the text: airflow.apache.org. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Creating Openlineage Extractors is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 3.3k tokens (SKILL.md is roughly 13k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Creating Openlineage Extractors: Chart Tests (astronomer/airflow-chart, 297 stars), Functional Tests (astronomer/airflow-chart, 297 stars), Helm Chart (astronomer/airflow-chart, 297 stars) and Create Example (godatadriven/whirl, 205 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
astronomer (a GitHub organization) maintains it in astronomer/agents, which has 451 GitHub stars. The repository holds 34 skills in this directory. The repository was last updated on October 7, 2026.
Source: astronomer/agents on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.