Agent skill

Analyzing Data

by astronomer in astronomer/agents

Queries the data warehouse with SQL and answers business questions about data.

Apache-2.0Auto-check passedDatabases

Install Analyzing Data

skills CLI
$ npx skills add astronomer/agents --skill analyzing-data -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install astronomer/agents analyzing-data --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/astronomer/agents.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/analyzing-data .claude/skills/analyzing-data && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
analyzing-data
GitHub stars
450
Token cost
~1.3k tokens
SKILL.md length
262 words
Files
27 (incl. scripts)
Skills in repo
34
Repo updated
First seen
Licence
Apache-2.0

At a glance

Queries the data warehouse with SQL and answers business questions about data.

  • Works in 6 steps: Pattern lookup — Check for a cached… → Concept lookup — Find known table mappings → Table discovery — If cache misses,… → …
  • Answering anything that needs warehouse data - counts
  • SKILL.md covers Workflow, Kernel Functions, CLI Reference and References
  • Runs Python scripts from its folder; calls uv

What it does

Analyzing Data is an agent skill from astronomer/agents. Queries the data warehouse with SQL and answers business questions about data. Use when answering anything that needs warehouse data - counts, metrics, trends, aggregations, joins across tables, data lookups, or ad-hoc SQL analysis (for example "who uses X", "how many Y", "show me Z", "find customers", "what is the count").

Its SKILL.md is about 1.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 30 other files, including scripts (for example `reference/common-patterns.md`, `reference/discovery-warehouse.md` and `scripts/cache.py`).

It sits in Databases, covering Data analysis, SQL and Data warehousing. It works with SQL, pandas and Polars. The repository describes itself as: AI agent tooling for data engineering workflows. The licence is Apache-2.0.

When your agent uses it

  • Answering anything that needs warehouse data - counts
  • Joins across tables
  • Ad-hoc SQL analysis (for example who uses X
  • What is the count)

Example prompts

  • “who uses X”
  • “how many Y”
  • “show me Z”
  • “/analyzing-data”

Requirements

  • Python 3

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Pattern lookup — Check for a cached query strategy
  2. Concept lookup — Find known table mappings
  3. Table discovery — If cache misses, search the codebase (Grep pattern="" glob="**/*.sql") or query INFORMATION_SCHEMA. See…
  4. Execute query
  5. Cache learnings — Always cache before presenting results
  6. Present findings to user.

What it can do on your machine

Read from SKILL.md and the folder at commit 1ec1a1f. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 14 files in scripts/ (Python, from the files we listed), which the agent can run.

    Shell commands in SKILL.md call:

    • uv

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use uv, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Analyzing Data loads about 1.3k tokens when it runs. Until then it costs about 85 tokens; SKILL.md has 262 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~85
When it runs · the whole SKILL.md, loaded when a task matches
~1.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from astronomer/agents at commit 1ec1a1f, republished under its Apache-2.0 licence (© astronomer). 262 words, ~1,267 tokens.

Download SKILL.mdSave it as .claude/skills/analyzing-data/SKILL.md (or your agent's skills folder). This skill also uses 26 other files; get the full folder from GitHub.
name
analyzing-data
description
Queries the data warehouse with SQL and answers business questions about data. Use when answering anything that needs warehouse data - counts, metrics, trends, aggregations, joins across tables, data lookups, or ad-hoc SQL analysis (for example "who uses X", "how many Y", "show me Z", "find customers", "what is the count").

Data Analysis

Answer business questions by querying the data warehouse. The kernel auto-starts on first exec call.

All CLI commands below are relative to this skill's directory. Before running any scripts/cli.py command, cd to the directory containing this file.

Workflow

  1. Pattern lookup — Check for a cached query strategy:

    bash
    uv run scripts/cli.py pattern lookup "<user's question>"

    If a pattern exists, follow its strategy. Record the outcome after executing:

    bash
    uv run scripts/cli.py pattern record <name> --success  # or --failure
  2. Concept lookup — Find known table mappings:

    bash
    uv run scripts/cli.py concept lookup <concept>
  3. Table discovery — If cache misses, search the codebase (Grep pattern="<concept>" glob="**/*.sql") or query INFORMATION_SCHEMA. See reference/discovery-warehouse.md.

  4. Execute query:

    bash
    uv run scripts/cli.py exec "df = run_sql('SELECT ...')"
    uv run scripts/cli.py exec "print(df)"
  5. Cache learnings — Always cache before presenting results:

    bash
    # Cache concept → table mapping
    uv run scripts/cli.py concept learn <concept> <TABLE> -k <KEY_COL>
    # Cache query strategy (if discovery was needed)
    uv run scripts/cli.py pattern learn <name> -q "question" -s "step" -t "TABLE" -g "gotcha"
  6. Present findings to user.

Kernel Functions

FunctionReturns
run_sql(query, limit=100)Polars DataFrame
run_sql_pandas(query, limit=100)Pandas DataFrame
run_sql_many(queries, limit=100)List of Polars DataFrames (one per query)

pl (Polars) and pd (Pandas) are pre-imported.

Run independent queries together with run_sql_many — they execute concurrently (Snowflake async / connection-pool fan-out) instead of one at a time:

bash
uv run scripts/cli.py exec "dfs = run_sql_many(['SELECT ...', 'SELECT ...']); print(dfs[0])"

run_sql_many is fail-fast: if any query errors, the call raises and the results of the queries that succeeded are discarded. Use separate run_sql calls if you need partial results.

Timeouts: exec waits up to 120s by default, then interrupts the query and returns a "client stopped waiting" message (the query may still finish server-side). Raise it for known long-running queries: uv run scripts/cli.py exec "..." -t 600.

Idle kernel: the kernel self-terminates after 2h idle (preserving state until then). Override with ASTRO_KERNEL_IDLE_TIMEOUT (seconds; 0 disables).

CLI Reference

Kernel
bash
uv run scripts/cli.py warehouse list      # List warehouses
uv run scripts/cli.py start [-w name]     # Start kernel (with optional warehouse)
uv run scripts/cli.py exec "..."          # Execute Python code
uv run scripts/cli.py status              # Kernel status
uv run scripts/cli.py restart             # Restart kernel
uv run scripts/cli.py stop                # Stop kernel
uv run scripts/cli.py install <pkg>       # Install package
Concept Cache
bash
uv run scripts/cli.py concept lookup <name>                     # Look up
uv run scripts/cli.py concept learn <name> <TABLE> -k <KEY_COL> # Learn
uv run scripts/cli.py concept list                               # List all
uv run scripts/cli.py concept import -p /path/to/warehouse.md   # Bulk import
Pattern Cache
bash
uv run scripts/cli.py pattern lookup "question"                                      # Look up
uv run scripts/cli.py pattern learn <name> -q "..." -s "..." -t "TABLE" -g "gotcha"  # Learn
uv run scripts/cli.py pattern record <name> --success                                # Record outcome
uv run scripts/cli.py pattern list                                                   # List all
uv run scripts/cli.py pattern delete <name>                                          # Delete
Table Schema Cache
bash
uv run scripts/cli.py table lookup <TABLE>            # Look up schema
uv run scripts/cli.py table cache <TABLE> -c '[...]'  # Cache schema
uv run scripts/cli.py table list                       # List cached
uv run scripts/cli.py table delete <TABLE>             # Delete
Cache Management
bash
uv run scripts/cli.py cache status                # Stats
uv run scripts/cli.py cache clear [--stale-only]  # Clear

References

© astronomer, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 26 other files (scripts) in skills/analyzing-data of astronomer/agents.

  • SKILL.md
  • reference/common-patterns.md
  • reference/discovery-warehouse.md
  • scripts/.gitignore
  • scripts/cache.py
  • scripts/cli.py
  • scripts/config.py
  • scripts/connectors.py
  • scripts/kernel.py
  • scripts/pyproject.toml
  • scripts/templates.py
  • scripts/tests/__init__.py
  • scripts/tests/conftest.py
  • scripts/tests/integration/__init__.py
  • scripts/tests/integration/conftest.py
  • scripts/tests/integration/test_duckdb_e2e.py
  • scripts/tests/integration/test_kernel_interrupt.py
  • … and 10 more

Open the folder on GitHubat commit 1ec1a1f

Compare with similar skills

Analyzing Data next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Analyzing Data compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Analyzing Data this skillastronomer/agents450—~1.3kAutomated safety check: PassApache-2.0
Chdb SQLvemetric/vemetric3941 repos~1.2kAutomated safety check: PassApache-2.0
Analytics Engineerborghei/Claude-Skills874—~3.4kAutomated safety check: PassMIT
Duckdb Experttheneoai/awesome-skills183—~4.2kAutomated safety check: PassMIT
Snowflake Developmentsickn33/agentic-awesome-skills47k2 repos~2.1kAutomated safety check: PassMIT
Snowflake Developmentalirezarezvani/claude-skills28k—~3.2kAutomated safety check: PassMIT

Similar skills

  • Chdb SQL

    vemetric/vemetric

    A skill your agent uses when the user wants to run SQL — especially analytical SQL — on local files (parquet/csv/json), URLs, S3 paths, or remote databases (Postgres, MySQL, MongoDB, ClickHouse…

    394 GitHub starsUsed in 1 repo~1.2k tokens
    DatabasesAuto-check passed
  • Analytics Engineer

    borghei/Claude-Skills

    Analytics engineering across data modeling, dbt, transformation, and semantic layers.

    874 GitHub stars~3.4k tokensUpdated today
    Data & AnalyticsAuto-check passed
  • Duckdb Expert

    theneoai/awesome-skills

    DuckDB expert for embedded OLAP analytics, Parquet/CSV querying, and high-performance analytical SQL on local data.

    183 GitHub stars~4.2k tokensUpdated 4 mo ago
    Data & AnalyticsAuto-check passed
  • Snowflake Development

    sickn33/agentic-awesome-skills

    Comprehensive Snowflake development assistant covering SQL best practices, data pipeline design (Dynamic Tables, Streams, Tasks, Snowpipe), Cortex AI functions, Cortex Agents, Snowpark Python, dbt…

    47k GitHub starsUsed in 2 repos~2.1k tokens
    DatabasesAuto-check passed
  • Snowflake Development

    alirezarezvani/claude-skills

    A skill your agent uses when writing Snowflake SQL, building data pipelines with Dynamic Tables or Streams/Tasks, using Cortex AI functions, creating Cortex Agents, writing Snowpark Python…

    28k GitHub stars~3.2k tokensUpdated 1 mo ago
    DatabasesAuto-check passed
  • Transforming Data

    ancoleman/ai-design-components

    Transform raw data into analytical assets using ETL/ELT patterns, SQL (dbt), Python (pandas/polars/PySpark), and orchestration (Airflow).

    526 GitHub stars~3k tokensUpdated 10 mo ago
    Data & AnalyticsAuto-check passed

More from astronomer/agents

All 34 skills in this repo
  • Airflow

    astronomer/agents

    Queries, manages, and troubleshoots Apache Airflow using the af CLI.

    450 GitHub starsUsed in 1 repo~3.8k tokens
    Auto-check passed
  • Guide for migrating Dagster projects to Apache Airflow 3 on Astro.

    450 GitHub stars~3.8k tokensUpdated 2 days ago
    Auto-check passed
  • Authoring Dags

    astronomer/agents

    Workflow and best practices for writing Apache Airflow DAGs.

    450 GitHub starsUsed in 1 repo~1.8k tokens
    Auto-check passed
  • Deploying Airflow

    astronomer/agents

    Deploys Airflow DAGs and projects. An agent skill from astronomer/agents.

    450 GitHub starsUsed in 1 repo~2.8k tokens
    Auto-check passed
  • Trace downstream data lineage and impact analysis. An agent skill from astronomer/agents.

    450 GitHub starsUsed in 1 repo~1.2k tokens
    Auto-check passed
  • Tracing Upstream Lineage

    astronomer/agents

    Trace upstream data lineage. An agent skill from astronomer/agents.

    450 GitHub starsUsed in 1 repo~1.1k tokens
    Auto-check passed

Works with

Questions about Analyzing Data

What does Analyzing Data do?

Queries the data warehouse with SQL and answers business questions about data. Analyzing Data is an agent skill from astronomer/agents. Queries the data warehouse with SQL and answers business questions about data.

When should I use Analyzing Data?

Analyzing Data fits situations like: answering anything that needs warehouse data - counts; joins across tables; ad-hoc SQL analysis (for example who uses X; what is the count).

How do I install Analyzing Data in Claude Code?

Run `npx skills add astronomer/agents --skill analyzing-data -a claude-code`. Or copy the skill folder (skills/analyzing-data in astronomer/agents) into .claude/skills/analyzing-data in your project. Claude Code loads it when a task matches its description.

How do I install Analyzing Data in Codex?

Run `npx skills add astronomer/agents --skill analyzing-data -a codex`. Or copy the skill folder (skills/analyzing-data in astronomer/agents) into .agents/skills/analyzing-data in your project. Codex loads it when a task matches its description.

Can I use Analyzing Data in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add astronomer/agents --skill analyzing-data -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/analyzing-data, .gemini/skills/analyzing-data, .github/skills/analyzing-data and .opencode/skills/analyzing-data in your project.

What does Analyzing Data need to run?

Going by SKILL.md and its folder, Analyzing Data needs Python for the scripts in its folder and the command-line tools its instructions call (uv). Our summary lists: Python 3.

Does Analyzing Data access the network?

SKILL.md contains no URLs. Its commands use uv, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Analyzing Data safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Analyzing Data use?

Analyzing Data is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Analyzing Data use?

About 1.3k tokens (SKILL.md is roughly 5.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Analyzing Data?

Skills that share tags, products or a category with Analyzing Data: Chdb SQL (vemetric/vemetric, 394 stars), Analytics Engineer (borghei/Claude-Skills, 874 stars), Duckdb Expert (theneoai/awesome-skills, 183 stars) and Snowflake Development (sickn33/agentic-awesome-skills, 47k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Analyzing Data?

astronomer (a GitHub organization) maintains it in astronomer/agents, which has 450 GitHub stars. The repository holds 34 skills in this directory. The repository was last updated on October 5, 2026.

Source: astronomer/agents on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.