Agent skill

Datasets

by ai-analyst-lab in ai-analyst-lab/ai-analyst

List all connected datasets with their status, table counts, and last analysis date.

MITAuto-check passedDatabases

Install Datasets

skills CLI
$ npx skills add ai-analyst-lab/ai-analyst --skill datasets -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install ai-analyst-lab/ai-analyst datasets --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/ai-analyst-lab/ai-analyst.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/datasets .claude/skills/datasets && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
datasets
GitHub stars
304
Token cost
~1k tokens
SKILL.md length
357 words
Files
1
Skills in repo
43
Repo updated
First seen
Licence
MIT

At a glance

List all connected datasets with their status, table counts, and last analysis date.

  • Works in 4 steps: Discover available datasets → Read the active pointer → Enrich with manifest data → …
  • The user invokes /datasets
  • SKILL.md covers Purpose, When to Use, Instructions and Important Notes
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Datasets is an agent skill from ai-analyst-lab/ai-analyst. List all connected datasets with their status, table counts, and last analysis date. Use this skill whenever the user invokes /datasets, or asks questions like "what datasets do I have?", "show me my data sources", "list datasets", "which datasets are connected?", "what data is available?", "show all datasets", "what datasets can I analyze?", "view my datasets", "what data sources are set up?", or any request to see what datasets exist in the system. Also trigger when users mention "switch dataset", "change…

Its SKILL.md is about 1k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Databases. The repository describes itself as: AI Product Analyst — Claude Code-powered data analysis toolkit. The licence is MIT.

When your agent uses it

  • The user invokes /datasets
  • Asks questions like what datasets do I have?
  • Show me my data sources
  • Which datasets are connected?

Example prompts

  • “what datasets do I have?”
  • “show me my data sources”
  • “list datasets”
  • “/datasets”

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Discover available datasets
  2. Read the active pointer
  3. Enrich with manifest data
  4. Display the list

What it can do on your machine

Read from SKILL.md and the folder at commit 52c0744. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Datasets loads about 1k tokens when it runs. Until then it costs about 212 tokens; SKILL.md has 357 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~212
When it runs · the whole SKILL.md, loaded when a task matches
~1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from ai-analyst-lab/ai-analyst at commit 52c0744, republished under its MIT licence (© ai-analyst-lab). 357 words, ~1,008 tokens.

Download SKILL.mdSave it as .claude/skills/datasets/SKILL.md (or your agent's skills folder).
name
datasets
description
List all connected datasets with their status, table counts, and last analysis date. Use this skill whenever the user invokes `/datasets`, or asks questions like "what datasets do I have?", "show me my data sources", "list datasets", "which datasets are connected?", "what data is available?", "show all datasets", "what datasets can I analyze?", "view my datasets", "what data sources are set up?", or any request to see what datasets exist in the system. Also trigger when users mention "switch dataset", "change dataset", "use a different dataset", or "what's my active dataset?" since seeing the list helps them choose. This is a foundational command that should be offered proactively whenever users seem unsure about what data they're working with or when they need to understand what datasets are available before starting analysis.

Skill: Datasets

Purpose

List all connected datasets with their status, table counts, and last analysis date.

When to Use

Invoke as /datasets when the user wants to see what datasets are available.

Instructions

Step 1: Discover available datasets

The system supports two discovery paths:

Path A: Registry-first (preferred)

  • Read data_sources.yaml to get the official list of registered sources
  • If the file exists and has entries, use this as your source of truth

Path B: Brain-first (fallback when registry is empty)

  • If data_sources.yaml is empty or missing, scan .knowledge/datasets/ directory
  • Each subdirectory represents a dataset (directory name = dataset ID)
  • Read each dataset's manifest.yaml to get connection details and metadata

Use whichever path yields results. Many installations have datasets in .knowledge/datasets/ but an empty data_sources.yaml registry — this is normal during initial setup or when datasets are added manually.

Step 2: Read the active pointer

Read .knowledge/active.yaml to determine which dataset is currently active.

Step 3: Enrich with manifest data

For each discovered dataset (whether from registry or directory scan), read .knowledge/datasets/{name}/manifest.yaml to get:

  • display_name — human-readable name
  • connection.type — connection type (csv, duckdb, postgres, snowflake, bigquery, databricks, redshift, mssql, mysql)
  • connection.database or other connection-specific fields
  • summary.table_count — number of tables
  • summary.date_range — temporal coverage (if available)
  • summary.row_counts — per-table row counts (if profiled)
  • summary.last_updated — when manifest was last written

If a manifest is missing or incomplete, show what you can determine from the directory structure and note that the dataset needs profiling.

Show full SKILL.md (122 more words)Show less
Step 4: Display the list
Connected Datasets:

  * your_dataset (active)
    Your Dataset Name — {table_count} tables, {date_range}
    Connection: {type} ({database})
    Analyses: 0

  - {other_dataset}
    {display_name} — {table_count} tables, {date_range}
    Connection: {type} ({details})
    Analyses: {count}

Commands:
  /switch-dataset {name}  — switch active dataset
  /connect-data           — connect a new dataset
  /data                   — inspect active dataset schema

Mark the active dataset with *. Mark others with -.

Important Notes

  1. Security: Never display connection credentials (tokens, passwords, API keys) — show only connection type and database/schema names
  2. Discovery: If data_sources.yaml is empty but .knowledge/datasets/ has content, scan the directory and show what's available — this is a normal state during development or manual dataset setup
  3. Incomplete manifests: If a dataset directory exists but has no manifest or an incomplete manifest, include it in the list with status "Not yet profiled" and show whatever metadata is available
  4. Duplicate detection: If you notice two datasets pointing to the same underlying data (same path or same database), mention this in a Notes section to help users clean up

© ai-analyst-lab, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/datasets of ai-analyst-lab/ai-analyst.

Open the folder on GitHubat commit 52c0744

Compare with similar skills

Datasets next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Datasets compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Datasets this skillai-analyst-lab/ai-analyst304—~1kAutomated safety check: PassMIT
Keeper Stress AnalysisClickHouse/ClickHouse50k—~4.7kAutomated safety check: PassApache-2.0
Perf ComparisonClickHouse/ClickHouse50k—~3.9kAutomated safety check: NotesApache-2.0
Patch Release CheckClickHouse/ClickHouse50k—~4kAutomated safety check: NotesApache-2.0
Evolving The Data ModelTriliumNext/Trilium38k—~2.1kAutomated safety check: PassAGPL-3.0
Hybrid Cloud Outboxesgetsentry/sentry46k—~4.8kAutomated safety check: PassCustom licence

Similar skills

  • Keeper Stress Analysis

    ClickHouse/ClickHouse

    Analyze ClickHouse Keeper stress-test results from play.clickhouse.com / keeperstresstests data warehouse.

    50k GitHub stars~4.7k tokensUpdated today
    DatabasesAuto-check passed
  • Perf Comparison

    ClickHouse/ClickHouse

    Evaluate ClickHouse performance test results from existing CI/dashboard data or local perf.py runs.

    50k GitHub stars~3.9k tokensUpdated today
    DatabasesAuto-check: notes
  • Patch Release Check

    ClickHouse/ClickHouse

    Check whether ClickHouse's supported versions (last 3 majors + latest LTS) have recent stable patch releases, diagnose why the scheduled AutoReleases pipeline failed, and identify which releases…

    50k GitHub stars~4k tokensUpdated today
    DatabasesAuto-check: notes
  • Evolving The Data Model

    TriliumNext/Trilium

    A skill your agent uses when adding a DB migration or a new column/field to a Becca entity in Trilium ("add a migration", "new column on notes/attributes", "ALTER TABLE", "add a field to…

    38k GitHub stars~2.1k tokensUpdated today
    DatabasesAuto-check passed
  • Hybrid Cloud Outboxes

    getsentry/sentry

    Official

    Guide for creating and maintaining outbox-based eventually consistent operations in Sentry.

    46k GitHub stars~4.8k tokensUpdated today
    DatabasesAuto-check passed
  • Adaptive Spin Wait

    microsoft/garnet

    Official

    Selects Garnet spin-wait policies based on the slow-path cost.

    12k GitHub stars~1k tokensUpdated today
    DatabasesAuto-check passed

More from ai-analyst-lab/ai-analyst

All 43 skills in this repo
  • Always Compare

    ai-analyst-lab/ai-analyst

    Never present a metric or number in isolation; anchor every number to a comparison (prior period, benchmark, or another segment) or state that none is available.

    304 GitHub stars~1.4k tokensUpdated 7 days ago
    Auto-check passed
  • Archaeology

    ai-analyst-lab/ai-analyst

    Retrieve proven SQL patterns, table cheatsheets, and join patterns from .knowledge/query-archaeology/ so past work gets reused.

    304 GitHub stars~1.3k tokensUpdated 7 days ago
    Auto-check passed
  • Archive Analysis

    ai-analyst-lab/ai-analyst

    Save completed analyses to the knowledge system's analysis archive for future reference.

    304 GitHub stars~2.7k tokensUpdated 7 days ago
    Auto-check passed
  • Auth Preflight

    ai-analyst-lab/ai-analyst

    Verify Google Workspace MCP authentication at the start of any session that needs Google APIs (Docs, Slides, Drive).

    304 GitHub stars~3.1k tokensUpdated 7 days ago
    Auto-check passed
  • Causal

    ai-analyst-lab/ai-analyst

    Causal inference toolkit for when experiments are not possible: estimate treatment effects from observational data with assumption checks and mandatory caveats.

    304 GitHub stars~1.8k tokensUpdated 7 days ago
    Auto-check passed
  • Chart To Drive

    ai-analyst-lab/ai-analyst

    Standardized workflow for uploading local chart PNGs to Google Drive and making them available for insertion into Google Docs and Slides.

    304 GitHub stars~1.4k tokensUpdated 7 days ago
    Auto-check passed

Categories

Questions about Datasets

What does Datasets do?

List all connected datasets with their status, table counts, and last analysis date. Datasets is an agent skill from ai-analyst-lab/ai-analyst. List all connected datasets with their status, table counts, and last analysis date.

When should I use Datasets?

Datasets fits situations like: the user invokes /datasets; asks questions like what datasets do I have?; show me my data sources; which datasets are connected?.

How do I install Datasets in Claude Code?

Run `npx skills add ai-analyst-lab/ai-analyst --skill datasets -a claude-code`. Or copy the skill folder (.claude/skills/datasets in ai-analyst-lab/ai-analyst) into .claude/skills/datasets in your project. Claude Code loads it when a task matches its description.

How do I install Datasets in Codex?

Run `npx skills add ai-analyst-lab/ai-analyst --skill datasets -a codex`. Or copy the skill folder (.claude/skills/datasets in ai-analyst-lab/ai-analyst) into .agents/skills/datasets in your project. Codex loads it when a task matches its description.

Can I use Datasets in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ai-analyst-lab/ai-analyst --skill datasets -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/datasets, .gemini/skills/datasets, .github/skills/datasets and .opencode/skills/datasets in your project.

What does Datasets need to run?

SKILL.md names no scripts, command-line tools or credentials: Datasets is instructions for the agent only.

Does Datasets access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Datasets safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Datasets use?

Datasets is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Datasets use?

About 1k tokens (SKILL.md is roughly 4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Datasets?

Skills that share tags, products or a category with Datasets: Keeper Stress Analysis (ClickHouse/ClickHouse, 50k stars), Perf Comparison (ClickHouse/ClickHouse, 50k stars), Patch Release Check (ClickHouse/ClickHouse, 50k stars) and Evolving The Data Model (TriliumNext/Trilium, 38k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Datasets?

ai-analyst-lab (a GitHub organization) maintains it in ai-analyst-lab/ai-analyst, which has 304 GitHub stars. The repository holds 43 skills in this directory. The repository was last updated on September 30, 2026.

Source: ai-analyst-lab/ai-analyst on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.