Agent skill

Data Explore

by amd in amd/gaia

Load messy tabular data into SQL scratchpad tables and answer questions with real queries instead of eyeballing.

MITAuto-check passedDocuments & Office

Install Data Explore

skills CLI
$ npx skills add amd/gaia --skill data-explore -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install amd/gaia data-explore --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/amd/gaia.git skills-src && mkdir -p .claude/skills && cp -r skills-src/hub/skills/data-explore .claude/skills/data-explore && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
data-explore
GitHub stars
1.6k
Token cost
~633 tokens
SKILL.md length
319 words
Files
1
Skills in repo
44
Repo updated
First seen
Licence
MIT

At a glance

Load messy tabular data into SQL scratchpad tables and answer questions with real queries instead of eyeballing.

  • Works in 6 steps: Look before you load. Read the first few… → Create the table with… → Insert with insert_data(table_name,… → …
  • The user has a CSV
  • SKILL.md covers Procedure, The prefix rule, Cleaning rules and Fork this
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Data Explore is an agent skill from amd/gaia. Load messy tabular data into SQL scratchpad tables and answer questions with real queries instead of eyeballing. Use when the user has a CSV, spreadsheet, export, or pasted table and asks for totals, trends, outliers, or a breakdown.

Its SKILL.md is about 630 tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Documents & Office, covering Excel spreadsheets and SQL. It works with SQL. The repository describes itself as: Build AI agents for your PC. The licence is MIT.

When your agent uses it

  • The user has a CSV
  • Pasted table and asks for totals

Example prompts

  • “/data-explore”

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Look before you load. Read the first few rows. Identify the columns, their
  2. Create the table with create_table(table_name, columns). Use explicit
  3. Insert with insert_data(table_name, data). Load everything, not a
  4. Verify the load. list_tables() to confirm the schema landed as you
  5. Answer with query_data(sql). One query per question. Show the SQL you
  6. Say what the data cannot tell you. Missing rows, nulls, and a single

What it can do on your machine

Read from SKILL.md and the folder at commit 6c3bb5c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Data Explore loads about 633 tokens when it runs. Until then it costs about 62 tokens; SKILL.md has 319 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~62
When it runs · the whole SKILL.md, loaded when a task matches
~633

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from amd/gaia at commit 6c3bb5c, republished under its MIT licence (© amd). 319 words, ~633 tokens.

Download SKILL.mdSave it as .claude/skills/data-explore/SKILL.md (or your agent's skills folder).
name
data-explore
description
Load messy tabular data into SQL scratchpad tables and answer questions with real queries instead of eyeballing. Use when the user has a CSV, spreadsheet, export, or pasted table and asks for totals, trends, outliers, or a breakdown.
license
MIT
version
1.0.0

Data Explore

An LLM reading numbers off a table gets them subtly wrong. An LLM writing SQL against those numbers does not. Always move the data into a table first.

Procedure

  1. Look before you load. Read the first few rows. Identify the columns, their types, and which one is the key. Say out loud what you think each column means and let the user correct you — a misread column poisons every later answer.
  2. Create the table with create_table(table_name, columns). Use explicit types. Store money as a number, not a string with a currency symbol; store dates as ISO YYYY-MM-DD.
  3. Insert with insert_data(table_name, data). Load everything, not a sample — the outliers are usually the point.
  4. Verify the load. list_tables() to confirm the schema landed as you intended, then query_data("SELECT COUNT(*) FROM scratch_<table>") and compare to the source row count. If they differ, find out why before answering anything.
  5. Answer with query_data(sql). One query per question. Show the SQL you ran alongside the result so the user can check your reasoning.
  6. Say what the data cannot tell you. Missing rows, nulls, and a single month of history are all limits worth naming.

The prefix rule

Every table name in a query carries the scratch_ prefix. A table created as create_table("sales", ...) is queried as SELECT ... FROM scratch_sales. Getting this wrong is the single most common failure here — the query errors instead of returning data, and no amount of rephrasing the question fixes it.

Cleaning rules

  • Trim whitespace and normalize case before comparing text keys.
  • Nulls and zeros are different. Never coerce one to the other silently.
  • If a column mixes formats (dates as both 01/02/24 and 2024-02-01), normalize on load and tell the user you did.

Fork this

Pin step 2 to your recurring export's exact schema and step 5 to the five questions you always ask of it — the skill becomes a one-command monthly report.

© amd, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in hub/skills/data-explore of amd/gaia.

Open the folder on GitHubat commit 6c3bb5c

Compare with similar skills

Data Explore next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Data Explore compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Data Explore this skillamd/gaia1.6k—~633Automated safety check: PassMIT
Forge Codegen Crudyaomindong1996/forge-admin125—~1.3kAutomated safety check: PassApache-2.0
Rust SQL Testshencangsheng/easydb_app590—~1.2kAutomated safety check: PassMIT
Data AnalysisHezaoHezao/poirot250—~997Automated safety check: PassMIT
Excel and CSV Data Analysisbytedance/deer-flow83k4 repos~2.2kAutomated safety check: PassMIT
Alumni Re Hire Trackersickn33/agentic-awesome-skills47k1 repos~3.4kAutomated safety check: PassMIT

Similar skills

  • Forge Codegen Crud

    yaomindong1996/forge-admin

    Generate or review Forge project code-generation output for CRUD modules.

    125 GitHub stars~1.3k tokensUpdated yesterday
    Documents & OfficeAuto-check passed
  • Rust SQL Test

    shencangsheng/easydb_app

    Enforces unit test requirements for the Rust data-processing modules in src-tauri/src/sql/ (generator.rs, parse.rs) and src-tauri/src/reader/ (excel.rs and other readers).

    590 GitHub stars~1.2k tokensUpdated 1 mo ago
    Documents & OfficeAuto-check passed
  • Data Analysis

    HezaoHezao/poirot

    Analyze Excel/CSV files with DuckDB SQL via bash. An agent skill from HezaoHezao/poirot.

    250 GitHub stars~997 tokensUpdated 2 mo ago
    Documents & OfficeAuto-check passed
  • Excel and CSV Data Analysis

    bytedance/deer-flow

    Analyzes uploaded Excel and CSV files with SQL through DuckDB, producing schema inspections, statistical summaries and exports to CSV, JSON or Markdown.

    83k GitHub starsUsed in 4 repos~2.2k tokens
    Data & AnalyticsAuto-check passed
  • Alumni Re Hire Tracker

    sickn33/agentic-awesome-skills

    Alumni and re-hire register: former role, last working day, re-hire eligibility, current employer and re-engagement date, as CSV, SQL, JSON Schema or Notion on request.

    47k GitHub starsUsed in 1 repo~3.4k tokens
    Documents & OfficeAuto-check passed
  • Announcement Board

    sickn33/agentic-awesome-skills

    Announcement board: author, category, department, priority, audience, publish and expiry dates, status and acknowledgements, as CSV, SQL, JSON Schema or Notion on request.

    47k GitHub starsUsed in 1 repo~3.4k tokens
    Documents & OfficeAuto-check passed

More from amd/gaia

All 44 skills in this repo
  • Adds a release eval scorecard to a GAIA hub agent by writing a harness adapter, running a real eval, and wiring the result into the agent's README and release gate.

    1.6k GitHub stars~2.6k tokensUpdated today
    Auto-check passed
  • Walks through releasing a GAIA sidecar agent as a frozen binary plus npm client through the tag-triggered Agent Hub CI pipeline, with a human gate before publishing.

    1.6k GitHub stars~3.6k tokensUpdated today
    Auto-check passed
  • Mines local Claude Code session transcripts with a deterministic Python pipeline to show what the agent is actually used for, how often it fails and what it costs.

    1.6k GitHub stars~2.3k tokensUpdated today
    Auto-check passed
  • Benchmarks AMD's GAIA agent against Claude Code and across models on quality, honesty, steps, tokens, time and real cost, using gaia eval tasks.

    1.6k GitHub stars~1.8k tokensUpdated today
    Auto-check passed
  • Guides safe code changes by finding the right file with grep or semantic search, reading before editing, reproducing bugs first, and proving a fix with a real test run.

    1.6k GitHub stars~2.1k tokensUpdated today
    Auto-check passed
  • Walks through scaffolding, writing and testing a new GAIA agent as a Python class with the SDK, from the base Agent subclass to registered tool methods.

    1.6k GitHub stars~1.5k tokensUpdated today
    Auto-check passed

Works with

Questions about Data Explore

What does Data Explore do?

Load messy tabular data into SQL scratchpad tables and answer questions with real queries instead of eyeballing. Data Explore is an agent skill from amd/gaia. Load messy tabular data into SQL scratchpad tables and answer questions with real queries instead of eyeballing.

When should I use Data Explore?

Data Explore fits situations like: the user has a CSV; pasted table and asks for totals.

How do I install Data Explore in Claude Code?

Run `npx skills add amd/gaia --skill data-explore -a claude-code`. Or copy the skill folder (hub/skills/data-explore in amd/gaia) into .claude/skills/data-explore in your project. Claude Code loads it when a task matches its description.

How do I install Data Explore in Codex?

Run `npx skills add amd/gaia --skill data-explore -a codex`. Or copy the skill folder (hub/skills/data-explore in amd/gaia) into .agents/skills/data-explore in your project. Codex loads it when a task matches its description.

Can I use Data Explore in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add amd/gaia --skill data-explore -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/data-explore, .gemini/skills/data-explore, .github/skills/data-explore and .opencode/skills/data-explore in your project.

What does Data Explore need to run?

SKILL.md names no scripts, command-line tools or credentials: Data Explore is instructions for the agent only.

Does Data Explore access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Data Explore safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Data Explore use?

Data Explore is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Data Explore use?

About 633 tokens (SKILL.md is roughly 2.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Data Explore?

Skills that share tags, products or a category with Data Explore: Forge Codegen Crud (yaomindong1996/forge-admin, 125 stars), Rust SQL Test (shencangsheng/easydb_app, 590 stars), Data Analysis (HezaoHezao/poirot, 250 stars) and Excel and CSV Data Analysis (bytedance/deer-flow, 83k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Data Explore?

amd (a GitHub organization) maintains it in amd/gaia, which has 1,580 GitHub stars. The repository holds 44 skills in this directory. The repository was last updated on October 6, 2026.

Source: amd/gaia on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.