Agent skill

Data Context Extractor

by w95 in w95/awesome-claude-corporate-skills

Generate or improve a company-specific data analysis skill by extracting tribal knowledge from analysts.

MITAuto-check passedData & Analytics

Install Data Context Extractor

skills CLI
$ npx skills add w95/awesome-claude-corporate-skills --skill data-context-extractor -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install w95/awesome-claude-corporate-skills data-context-extractor --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/w95/awesome-claude-corporate-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/10-data-analytics/data-context-extractor .claude/skills/data-context-extractor && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
data-context-extractor
GitHub stars
237
Used in
1 other repo
Token cost
~1.8k tokens
SKILL.md length
763 words
Files
6 (incl. scripts, references)
Skills in repo
39
Repo updated
First seen
Licence
MIT

At a glance

Generate or improve a company-specific data analysis skill by extracting tribal knowledge from analysts.

  • Works in 8 steps: Database Connection & Discovery → Core Questions (Ask These) → Generate the Skill → …
  • Data analysts want Claude to understand their companys specific data warehouse
  • SKILL.md covers How It Works, Bootstrap Mode, Iteration Mode and Reference File Standards, plus 1 more section
  • Runs Python scripts from its folder

What it does

Data Context Extractor is an agent skill from w95/awesome-claude-corporate-skills. Generate or improve a company-specific data analysis skill by extracting tribal knowledge from analysts. BOOTSTRAP MODE - Triggers: "Create a data context skill", "Set up data analysis for our warehouse", "Help me create a skill for our database", "Generate a data skill for [company]" → Discovers schemas, asks key questions, generates initial skill with reference files ITERATION MODE - Triggers: "Add context about [domain]", "The skill needs more info about [topic]", "Update the data skill with…

Its SKILL.md is about 1.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 7 other files, including scripts and reference files (for example `references/domain-template.md`, `references/example-output.md` and `references/skill-template.md`).

It sits in Data & Analytics, covering Data analysis, Data warehousing and Skill authoring. The repository describes itself as: 166 production-ready Claude AI skills organized by corporate role — executive leadership, finance, HR, marketing, sales, legal, operations, engineering, product, data, customer…. The licence is MIT.

When your agent uses it

  • Data analysts want Claude to understand their companys specific data warehouse
  • Metrics definitions
  • Common query patterns

Example prompts

  • “Create a data context skill”
  • “Set up data analysis for our warehouse”
  • “Help me create a skill for our database”
  • “/data-context-extractor”

Requirements

  • Python 3

Workflow steps

8 steps, taken from the step headings in SKILL.md.

  1. Database Connection & Discovery
  2. Core Questions (Ask These)
  3. Generate the Skill
  4. Package and Deliver
  5. Load Existing Skill
  6. Identify the Gap
  7. Targeted Discovery
  8. Update and Repackage

What it can do on your machine

Read from SKILL.md and the folder at commit 78dbc7c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Data Context Extractor loads about 1.8k tokens when it runs, and up to ~6.3k if it reads all its reference files. Until then it costs about 205 tokens; SKILL.md has 763 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~205
When it runs · the whole SKILL.md, loaded when a task matches
~1.8k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~6.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from w95/awesome-claude-corporate-skills at commit 78dbc7c, republished under its MIT licence (© w95). 763 words, ~1,794 tokens.

Download SKILL.mdSave it as .claude/skills/data-context-extractor/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.
name
data-context-extractor
description
Generate or improve a company-specific data analysis skill by extracting tribal knowledge from analysts. BOOTSTRAP MODE - Triggers: "Create a data context skill", "Set up data analysis for our warehouse", "Help me create a skill for our database", "Generate a data skill for [company]" → Discovers schemas, asks key questions, generates initial skill with reference files ITERATION MODE - Triggers: "Add context about [domain]", "The skill needs more info about [topic]", "Update the data skill with [metrics/tables/terminology]", "Improve the [domain] reference" → Loads existing skill, asks targeted questions, appends/updates reference files Use when data analysts want Claude to understand their company's specific data warehouse, terminology, metrics definitions, and common query patterns.

Data Context Extractor

A meta-skill that extracts company-specific data knowledge from analysts and generates tailored data analysis skills.

How It Works

This skill has two modes:

  1. Bootstrap Mode: Create a new data analysis skill from scratch
  2. Iteration Mode: Improve an existing skill by adding domain-specific reference files

Bootstrap Mode

Use when: User wants to create a new data context skill for their warehouse.

Phase 1: Database Connection & Discovery

Step 1: Identify the database type

Ask: "What data warehouse are you using?"

Common options:

  • BigQuery
  • Snowflake
  • PostgreSQL/Redshift
  • Databricks

Use ~~data warehouse tools (query and schema) to connect. If unclear, check available MCP tools in the current session.

Step 2: Explore the schema

Use ~~data warehouse schema tools to:

  1. List available datasets/schemas
  2. Identify the most important tables (ask user: "Which 3-5 tables do analysts query most often?")
  3. Pull schema details for those key tables

Sample exploration queries by dialect:

sql
-- BigQuery: List datasets
SELECT schema_name FROM INFORMATION_SCHEMA.SCHEMATA

-- BigQuery: List tables in a dataset
SELECT table_name FROM `project.dataset.INFORMATION_SCHEMA.TABLES`

-- Snowflake: List schemas
SHOW SCHEMAS IN DATABASE my_database

-- Snowflake: List tables
SHOW TABLES IN SCHEMA my_schema
Phase 2: Core Questions (Ask These)

After schema discovery, ask these questions conversationally (not all at once):

Entity Disambiguation (Critical)

"When people here say 'user' or 'customer', what exactly do they mean? Are there different types?"

Listen for:

  • Multiple entity types (user vs account vs organization)
  • Relationships between them (1:1, 1:many, many:many)
  • Which ID fields link them together

Primary Identifiers

"What's the main identifier for a [customer/user/account]? Are there multiple IDs for the same entity?"

Listen for:

  • Primary keys vs business keys
  • UUID vs integer IDs
  • Legacy ID systems

Key Metrics

"What are the 2-3 metrics people ask about most? How is each one calculated?"

Listen for:

  • Exact formulas (ARR = monthly_revenue × 12)
  • Which tables/columns feed each metric
  • Time period conventions (trailing 7 days, calendar month, etc.)

Data Hygiene

"What should ALWAYS be filtered out of queries? (test data, fraud, internal users, etc.)"

Listen for:

  • Standard WHERE clauses to always include
  • Flag columns that indicate exclusions (is_test, is_internal, is_fraud)
  • Specific values to exclude (status = 'deleted')

Common Gotchas

"What mistakes do new analysts typically make with this data?"

Listen for:

  • Confusing column names
  • Timezone issues
  • NULL handling quirks
  • Historical vs current state tables
Phase 3: Generate the Skill

Create a skill with this structure:

[company]-data-analyst/
├── SKILL.md
└── references/
    ├── entities.md          # Entity definitions and relationships
    ├── metrics.md           # KPI calculations
    ├── tables/              # One file per domain
    │   ├── [domain1].md
    │   └── [domain2].md
    └── dashboards.json      # Optional: existing dashboards catalog

SKILL.md Template: See references/skill-template.md

SQL Dialect Section: See references/sql-dialects.md and include the appropriate dialect notes.

Reference File Template: See references/domain-template.md

Phase 4: Package and Deliver
  1. Create all files in the skill directory
  2. Package as a zip file
  3. Present to user with summary of what was captured

Iteration Mode

Use when: User has an existing skill but needs to add more context.

Step 1: Load Existing Skill

Ask user to upload their existing skill (zip or folder), or locate it if already in the session.

Read the current SKILL.md and reference files to understand what's already documented.

Show full SKILL.md (310 more words)Show less
Step 2: Identify the Gap

Ask: "What domain or topic needs more context? What queries are failing or producing wrong results?"

Common gaps:

  • A new data domain (marketing, finance, product, etc.)
  • Missing metric definitions
  • Undocumented table relationships
  • New terminology
Step 3: Targeted Discovery

For the identified domain:

  1. Explore relevant tables: Use ~~data warehouse schema tools to find tables in that domain

  2. Ask domain-specific questions:

    • "What tables are used for [domain] analysis?"
    • "What are the key metrics for [domain]?"
    • "Any special filters or gotchas for [domain] data?"
  3. Generate new reference file: Create references/[domain].md using the domain template

Step 4: Update and Repackage
  1. Add the new reference file
  2. Update SKILL.md's "Knowledge Base Navigation" section to include the new domain
  3. Repackage the skill
  4. Present the updated skill to user

Reference File Standards

Each reference file should include:

For Table Documentation
  • Location: Full table path
  • Description: What this table contains, when to use it
  • Primary Key: How to uniquely identify rows
  • Update Frequency: How often data refreshes
  • Key Columns: Table with column name, type, description, notes
  • Relationships: How this table joins to others
  • Sample Queries: 2-3 common query patterns
For Metrics Documentation
  • Metric Name: Human-readable name
  • Definition: Plain English explanation
  • Formula: Exact calculation with column references
  • Source Table(s): Where the data comes from
  • Caveats: Edge cases, exclusions, gotchas
For Entity Documentation
  • Entity Name: What it's called
  • Definition: What it represents in the business
  • Primary Table: Where to find this entity
  • ID Field(s): How to identify it
  • Relationships: How it relates to other entities
  • Common Filters: Standard exclusions (internal, test, etc.)

Quality Checklist

Before delivering a generated skill, verify:

  • SKILL.md has complete frontmatter (name, description)
  • Entity disambiguation section is clear
  • Key terminology is defined
  • Standard filters/exclusions are documented
  • At least 2-3 sample queries per domain
  • SQL uses correct dialect syntax
  • Reference files are linked from SKILL.md navigation section

© w95, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 5 other files (scripts, references) in 10-data-analytics/data-context-extractor of w95/awesome-claude-corporate-skills.

  • SKILL.md
  • references/domain-template.md
  • references/example-output.md
  • references/skill-template.md
  • references/sql-dialects.md
  • scripts/package_data_skill.py

Open the folder on GitHubat commit 78dbc7c

Used in 1 other repository

We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in w95/awesome-claude-corporate-skills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Data Context Extractor next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Data Context Extractor compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Data Context Extractor this skillw95/awesome-claude-corporate-skills2371 repos~1.8kAutomated safety check: PassMIT
Analytics Engineerborghei/Claude-Skills881—~3.4kAutomated safety check: PassMIT
Analyzing Dataastronomer/agents451—~1.3kAutomated safety check: PassApache-2.0
Chdb Datastorevemetric/vemetric3942 repos~1.4kAutomated safety check: PassApache-2.0
Yichen Wecom Local Vaultmcncarl/yichen-skills4.3k—~1.3kAutomated safety check: PassCustom licence
Erd Studio Setupliam-machine/erd-studio165—~8.5kAutomated safety check: PassCustom licence

Similar skills

  • Analytics Engineer

    borghei/Claude-Skills

    Analytics engineering across data modeling, dbt, transformation, and semantic layers.

    881 GitHub stars~3.4k tokensUpdated today
    Data & AnalyticsAuto-check passed
  • Analyzing Data

    astronomer/agents

    Queries the data warehouse with SQL and answers business questions about data.

    451 GitHub stars~1.3k tokensUpdated today
    DatabasesAuto-check passed
  • Chdb Datastore

    vemetric/vemetric

    A skill your agent uses when the user has tabular data (pandas DataFrame, parquet, csv, Arrow, json) and wants to filter, group, aggregate, join, or speed up slow pandas.

    394 GitHub starsUsed in 2 repos~1.4k tokens
    Data & AnalyticsAuto-check passed
  • Yichen Wecom Local Vault

    mcncarl/yichen-skills

    Read, decrypt, query, search, and export local WeCom/企业微信 5.x desktop databases on macOS into a private read-only vault.

    4.3k GitHub stars~1.3k tokensUpdated 3 days ago
    Data & AnalyticsAuto-check passed
  • Erd Studio Setup

    liam-machine/erd-studio

    Friendly, step-by-step setup for ERD Studio in an existing dbt project, for people who may be new to dbt or data modelling.

    165 GitHub stars~8.5k tokensUpdated today
    Data & AnalyticsAuto-check passed
  • Sets up Dinobase, a local DuckDB database that syncs data from 100+ business sources, then answers questions across them with SQL joins and previewed write-backs.

    263 GitHub stars~1.5k tokensUpdated 3 mo ago
    Data & AnalyticsAuto-check passed

More from w95/awesome-claude-corporate-skills

All 39 skills in this repo
  • Competitive Analysis

    w95/awesome-claude-corporate-skills

    Framework for competitive landscape analysis across any industry.

    237 GitHub starsUsed in 1 repo~4k tokens
    Auto-check passed
  • Account Research

    w95/awesome-claude-corporate-skills

    Research a company using Common Room data. An agent skill from w95/awesome-claude-corporate-skills.

    237 GitHub stars~1.5k tokensUpdated 7 mo ago
    Auto-check passed
  • Call Prep

    w95/awesome-claude-corporate-skills

    Prepare for a customer or prospect call using Common Room signals.

    237 GitHub stars~1.5k tokensUpdated 7 mo ago
    Auto-check passed
  • Compose Outreach

    w95/awesome-claude-corporate-skills

    Generate personalized outreach messages using Common Room signals.

    237 GitHub stars~1.4k tokensUpdated 7 mo ago
    Auto-check passed
  • SQL Queries

    w95/awesome-claude-corporate-skills

    Write correct, performant SQL across all major data warehouse dialects (Snowflake, BigQuery, Databricks, PostgreSQL, etc.).

    237 GitHub starsUsed in 3 repos~2.8k tokens
    Auto-check passed
  • Contact Research

    w95/awesome-claude-corporate-skills

    Research a specific person using Common Room data. An agent skill from w95/awesome-claude-corporate-skills.

    237 GitHub stars~1.3k tokensUpdated 7 mo ago
    Auto-check passed

Questions about Data Context Extractor

What does Data Context Extractor do?

Generate or improve a company-specific data analysis skill by extracting tribal knowledge from analysts. Data Context Extractor is an agent skill from w95/awesome-claude-corporate-skills. Generate or improve a company-specific data analysis skill by extracting tribal knowledge from analysts.

When should I use Data Context Extractor?

Data Context Extractor fits situations like: data analysts want Claude to understand their companys specific data warehouse; metrics definitions; common query patterns.

How do I install Data Context Extractor in Claude Code?

Run `npx skills add w95/awesome-claude-corporate-skills --skill data-context-extractor -a claude-code`. Or copy the skill folder (10-data-analytics/data-context-extractor in w95/awesome-claude-corporate-skills) into .claude/skills/data-context-extractor in your project. Claude Code loads it when a task matches its description.

How do I install Data Context Extractor in Codex?

Run `npx skills add w95/awesome-claude-corporate-skills --skill data-context-extractor -a codex`. Or copy the skill folder (10-data-analytics/data-context-extractor in w95/awesome-claude-corporate-skills) into .agents/skills/data-context-extractor in your project. Codex loads it when a task matches its description.

Can I use Data Context Extractor in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add w95/awesome-claude-corporate-skills --skill data-context-extractor -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/data-context-extractor, .gemini/skills/data-context-extractor, .github/skills/data-context-extractor and .opencode/skills/data-context-extractor in your project.

What does Data Context Extractor need to run?

Going by SKILL.md and its folder, Data Context Extractor needs Python for the scripts in its folder. Our summary lists: Python 3.

Does Data Context Extractor access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Data Context Extractor safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Data Context Extractor use?

Data Context Extractor is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Data Context Extractor use?

About 1.8k tokens (SKILL.md is roughly 7.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 4.5k tokens, read only when the agent opens those files.

What are the alternatives to Data Context Extractor?

Skills that share tags, products or a category with Data Context Extractor: Analytics Engineer (borghei/Claude-Skills, 881 stars), Analyzing Data (astronomer/agents, 451 stars), Chdb Datastore (vemetric/vemetric, 394 stars) and Yichen Wecom Local Vault (mcncarl/yichen-skills, 4.3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Data Context Extractor?

w95 (a GitHub user) maintains it in w95/awesome-claude-corporate-skills, which has 237 GitHub stars. The repository holds 39 skills in this directory. The repository was last updated on February 26, 2026.

Source: w95/awesome-claude-corporate-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.