Agent skill

Dataset Profiler

by brevdev in brevdev/workshop-build-an-agent

Systematic procedure for profiling, summarizing, or exploring an unfamiliar CSV file or DataFrame

Apache-2.0Auto-check passedData & Analytics

Install Dataset Profiler

skills CLI
$ npx skills add brevdev/workshop-build-an-agent --skill dataset-profiler -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install brevdev/workshop-build-an-agent dataset-profiler --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/brevdev/workshop-build-an-agent.git skills-src && mkdir -p .claude/skills && cp -r skills-src/code/7-agent-harnesses/skills/.examples/dataset-profiler .claude/skills/dataset-profiler && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
dataset-profiler
GitHub stars
144
Token cost
~409 tokens
SKILL.md length
183 words
Files
1
Skills in repo
11
Repo updated
First seen
Licence
Apache-2.0

At a glance

Systematic procedure for profiling, summarizing, or exploring an unfamiliar CSV file or DataFrame

  • Works in 7 steps: Shape & size — row count, column count,… → Schema — every column with its dtype;… → Nulls — per-column null counts; call out… → …
  • Tasks that involve Performance optimization
  • SKILL.md covers Procedure, Output format and Notes
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Dataset Profiler is an agent skill from brevdev/workshop-build-an-agent. Systematic procedure for profiling, summarizing, or exploring an unfamiliar CSV file or DataFrame

Its SKILL.md is about 410 tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Data & Analytics, covering Performance optimization, DataFrames and CSV and tabular files. The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve Performance optimization
  • Tasks that involve DataFrames
  • Tasks that involve CSV and tabular files

Example prompts

  • “/dataset-profiler”

Workflow steps

7 steps, taken from the first numbered list in SKILL.md.

  1. Shape & size — row count, column count, file size on disk.
  2. Schema — every column with its dtype; flag columns whose dtype looks
  3. Nulls — per-column null counts; call out any column over 5% null.
  4. Duplicates — count of fully duplicated rows.
  5. Numeric distributions — min / max / mean / std for numeric columns;
  6. Cardinality — unique-value counts for object columns; identify likely
  7. Three surprising facts — the three most decision-relevant things a

What it can do on your machine

Read from SKILL.md and the folder at commit b5689a7. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Dataset Profiler loads about 409 tokens when it runs. Until then it costs about 29 tokens; SKILL.md has 183 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~29
When it runs · the whole SKILL.md, loaded when a task matches
~409

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from brevdev/workshop-build-an-agent at commit b5689a7, republished under its Apache-2.0 licence (© brevdev). 183 words, ~409 tokens.

Download SKILL.mdSave it as .claude/skills/dataset-profiler/SKILL.md (or your agent's skills folder).
name
dataset-profiler
description
Systematic procedure for profiling, summarizing, or exploring an unfamiliar CSV file or DataFrame

Dataset Profiler Skill

You are profiling an unfamiliar dataset. Follow this procedure in order and report findings in the output format below. Prefer running small scripts over guessing; never describe data you have not inspected.

Procedure

  1. Shape & size — row count, column count, file size on disk.
  2. Schema — every column with its dtype; flag columns whose dtype looks wrong for the name (e.g. numeric IDs parsed as float, dates as object).
  3. Nulls — per-column null counts; call out any column over 5% null.
  4. Duplicates — count of fully duplicated rows.
  5. Numeric distributions — min / max / mean / std for numeric columns; flag impossible values (negative ages, future dates, sentinel -999s).
  6. Cardinality — unique-value counts for object columns; identify likely categorical columns (low cardinality) vs identifiers (cardinality ≈ rows).
  7. Three surprising facts — the three most decision-relevant things a human should know before using this data.

Output format

## Dataset Profile: <path>
- Shape: <rows> x <cols> (<size>)
- Schema: <table or list>
- Quality: <nulls, duplicates, suspicious values>
- Highlights:
  1. ...
  2. ...
  3. ...

Notes

  • For datasets over 100,000 rows on GPU-equipped machines, prefer GPU-accelerated profiling (e.g. cudf.pandas) and keep intermediate results on the GPU.
  • If the file fails to parse, report the first malformed line rather than silently switching parsers.

© brevdev, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in code/7-agent-harnesses/skills/.examples/dataset-profiler of brevdev/workshop-build-an-agent.

Open the folder on GitHubat commit b5689a7

Compare with similar skills

Dataset Profiler next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Dataset Profiler compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Dataset Profiler this skillbrevdev/workshop-build-an-agent144—~409Automated safety check: PassApache-2.0
Antv L7antvis/L74.1k—~1.4kAutomated safety check: PassMIT
Fla Ascend Performancefla-org/flash-linear-attention5.8k—~6.3kAutomated safety check: PassMIT
CSV Data Summarizercoffeefuelbump/csv-data-summarizer-claude-skill4682 repos~1.4kAutomated safety check: PassNone
Paper FiguresEvoScientist/EvoSkills4751 repos~4.4kAutomated safety check: PassApache-2.0
Pandas ProJeffallan/claude-skills12k1 repos~1.5kAutomated safety check: PassMIT

Similar skills

  • Antv L7

    antvis/L7

    Comprehensive guide for AntV L7 geospatial visualization library.

    4.1k GitHub stars~1.4k tokensUpdated 29 days ago
    Data & AnalyticsAuto-check passed
  • Fla Ascend Performance

    fla-org/flash-linear-attention

    Guidelines for Ascend NPU kernel / Triton-Ascend backend performance work in the FLA repo.

    5.8k GitHub stars~6.3k tokensUpdated today
    SecurityAuto-check passed
  • CSV Data Summarizer

    coffeefuelbump/csv-data-summarizer-claude-skill

    Analyzes CSV files, generates summary stats, and plots quick visualizations using Python and pandas.

    468 GitHub starsUsed in 2 repos~1.4k tokens
    Data & AnalyticsAuto-check passed
  • Paper Figures

    EvoScientist/EvoSkills

    A skill your agent uses to produce standalone, publication-ready PNG graphics and reproducible matplotlib scripts from tabular data (CSVs or DataFrames).

    475 GitHub starsUsed in 1 repo~4.4k tokens
    Data & AnalyticsAuto-check passed
  • Pandas Pro

    Jeffallan/claude-skills

    Handles pandas DataFrame work: cleaning, merging, groupby aggregation, pivots, time-series resampling and memory tuning, with checks on dtypes, shapes and nulls.

    12k GitHub starsUsed in 1 repo~1.5k tokens
    Data & AnalyticsAuto-check passed
  • Verified Data Analysis with pandas

    pipeshub-ai/pipeshub-ai

    Loads, cleans, aggregates and joins tabular data with pandas under a verification rule: every number reported must be one that the code actually printed.

    3.8k GitHub stars~1.2k tokensUpdated today
    Data & AnalyticsAuto-check passed

More from brevdev/workshop-build-an-agent

All 11 skills in this repo
  • Setup Workshop

    brevdev/workshop-build-an-agent

    This skill should be used when the user wants to set up, install, deploy, bootstrap, or "spin up" the Build-an-Agent workshop (a.k.a.

    144 GitHub stars~2.3k tokensUpdated yesterday
    Auto-check: notes
  • Setup Workshop Nemoclaw

    brevdev/workshop-build-an-agent

    Set up the NVIDIA "Build an Agent" DevX workshop as a working JupyterLab environment from INSIDE a locked-down OpenShell/NemoClaw sandbox, and hand the user the token URL + access commands.

    144 GitHub stars~5.2k tokensUpdated yesterday
    Auto-check passed
  • Nvwb

    brevdev/workshop-build-an-agent

    Manage NVIDIA AI Workbench projects, contexts, builds, and environments via the nvwb CLI.

    144 GitHub stars~1.1k tokensUpdated yesterday
    Auto-check passed
  • Module 1

    brevdev/workshop-build-an-agent

    This skill should be used when a learner is working through Module 1 ("Build an Agent") of the Build-an-Agent workshop and wants help understanding the concepts, notebooks, or code — e.g.

    144 GitHub stars~2.7k tokensUpdated yesterday
    Auto-check passed
  • Module 1

    brevdev/workshop-build-an-agent

    This skill should be used when a learner is working through Module 1 ("Build an Agent") of the Build-an-Agent workshop and wants help understanding the concepts, notebooks, or code — e.g.

    144 GitHub stars~2.7k tokensUpdated yesterday
    Auto-check passed
  • Workshop

    brevdev/workshop-build-an-agent

    This skill should be used when a learner wants to navigate or understand the Build-an-Agent workshop as a whole — e.g.

    144 GitHub stars~1.4k tokensUpdated yesterday
    Auto-check passed

Questions about Dataset Profiler

What does Dataset Profiler do?

Systematic procedure for profiling, summarizing, or exploring an unfamiliar CSV file or DataFrame. Dataset Profiler is an agent skill from brevdev/workshop-build-an-agent.

When should I use Dataset Profiler?

Dataset Profiler fits situations like: tasks that involve Performance optimization; tasks that involve DataFrames; tasks that involve CSV and tabular files.

How do I install Dataset Profiler in Claude Code?

Run `npx skills add brevdev/workshop-build-an-agent --skill dataset-profiler -a claude-code`. Or copy the skill folder (code/7-agent-harnesses/skills/.examples/dataset-profiler in brevdev/workshop-build-an-agent) into .claude/skills/dataset-profiler in your project. Claude Code loads it when a task matches its description.

How do I install Dataset Profiler in Codex?

Run `npx skills add brevdev/workshop-build-an-agent --skill dataset-profiler -a codex`. Or copy the skill folder (code/7-agent-harnesses/skills/.examples/dataset-profiler in brevdev/workshop-build-an-agent) into .agents/skills/dataset-profiler in your project. Codex loads it when a task matches its description.

Can I use Dataset Profiler in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add brevdev/workshop-build-an-agent --skill dataset-profiler -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/dataset-profiler, .gemini/skills/dataset-profiler, .github/skills/dataset-profiler and .opencode/skills/dataset-profiler in your project.

What does Dataset Profiler need to run?

SKILL.md names no scripts, command-line tools or credentials: Dataset Profiler is instructions for the agent only.

Does Dataset Profiler access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Dataset Profiler safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Dataset Profiler use?

Dataset Profiler is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Dataset Profiler use?

About 409 tokens (SKILL.md is roughly 1.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Dataset Profiler?

Skills that share tags, products or a category with Dataset Profiler: Antv L7 (antvis/L7, 4.1k stars), Fla Ascend Performance (fla-org/flash-linear-attention, 5.8k stars), CSV Data Summarizer (coffeefuelbump/csv-data-summarizer-claude-skill, 468 stars) and Paper Figures (EvoScientist/EvoSkills, 475 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Dataset Profiler?

brevdev (a GitHub organization) maintains it in brevdev/workshop-build-an-agent, which has 144 GitHub stars. The repository holds 11 skills in this directory. The repository was last updated on October 7, 2026.

Source: brevdev/workshop-build-an-agent on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.