Systematic dataset profiling protocol for empirical research.

No licenceAuto-check passedData & Analytics

Install Data Profiler

skills CLI
$ npx skills add aspi6246/Claude-Code-Skills-for-Academics --skill data-profiler -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install aspi6246/Claude-Code-Skills-for-Academics data-profiler --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/aspi6246/Claude-Code-Skills-for-Academics.git skills-src && mkdir -p .claude/skills && cp -r skills-src/data-profiler .claude/skills/data-profiler && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
data-profiler
GitHub stars
157
Token cost
~2k tokens
SKILL.md length
628 words
Files
2 (incl. references)
Skills in repo
15
Repo updated
First seen
Licence
None found

At a glance

Systematic dataset profiling protocol for empirical research.

  • Works in 5 steps: Structure Discovery → Unit of Observation → Variable Profiling → …
  • The user has a new dataset and wants to understand it before analysis — including unit of observation
  • SKILL.md covers Purpose, Critical Rules, Phase 1: Structure Discovery and Phase 2: Unit of Observation, plus 3 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Data Profiler is an agent skill from aspi6246/Claude-Code-Skills-for-Academics. Systematic dataset profiling protocol for empirical research. Use this skill when the user has a new dataset and wants to understand it before analysis — including unit of observation, variable definitions, panel structure, data quality, and descriptive statistics. Trigger on phrases like "explore this data", "profile this dataset", "what's in this data", "understand this dataset", "describe this data", "what are the variables", "check the panel structure", or any request to examine a dataset before running…

Its SKILL.md is about 2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including reference files (for example `references/profile-template.md`).

It sits in Data & Analytics, covering Performance optimization, Data cleaning and Statistics. The repository describes itself as: A collection of Skills for academics.

When your agent uses it

  • The user has a new dataset and wants to understand it before analysis — including unit of observation
  • Variable definitions
  • Panel structure
  • Descriptive statistics

Example prompts

  • “explore this data”
  • “profile this dataset”
  • “s in this data”
  • “/data-profiler”

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Structure Discovery
  2. Unit of Observation
  3. Variable Profiling
  4. Relationships and Red Flags
  5. Profile Document

What it can do on your machine

Read from SKILL.md and the folder at commit 8a8f112. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are r).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Data Profiler loads about 2k tokens when it runs, and up to ~3k if it reads all its reference files. Until then it costs about 135 tokens; SKILL.md has 628 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~135
When it runs · the whole SKILL.md, loaded when a task matches
~2k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

Without a licence we can't republish the file, so here is its outline and opening line. It has 628 words (~1,955 tokens).

“This skill implements a structured protocol for profiling a dataset before any analysis begins. The goal is to produce a dataset profile document that records your understanding of the data and becomes permanent reference material for the project.”

— opening of SKILL.md by aspi6246
name
data-profiler

Read the full SKILL.md on GitHub

Files

SKILL.md and 1 other file (references) in data-profiler of aspi6246/Claude-Code-Skills-for-Academics.

  • SKILL.md
  • references/profile-template.md

Open the folder on GitHubat commit 8a8f112

Compare with similar skills

Data Profiler next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Data Profiler compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Data Profiler this skillaspi6246/Claude-Code-Skills-for-Academics157—~2kAutomated safety check: PassNone
Pandas ProJeffallan/claude-skills12k1 repos~1.5kAutomated safety check: PassMIT
Code EngineeropenJiuwen-ai/sciencediscovery148—~2.8kAutomated safety check: PassApache-2.0
Data Analysisxiaoyuge886/aigc1981 repos~794Automated safety check: PassMIT
Data Explorerliangdabiao/claude-data-analysis-ultra-main290—~2.1kAutomated safety check: PassNone
Profiling Tablesastronomer/agents450—~964Automated safety check: PassApache-2.0

Similar skills

  • Pandas Pro

    Jeffallan/claude-skills

    Handles pandas DataFrame work: cleaning, merging, groupby aggregation, pivots, time-series resampling and memory tuning, with checks on dtypes, shapes and nulls.

    12k GitHub starsUsed in 1 repo~1.5k tokens
    Data & AnalyticsAuto-check passed
  • Code Engineer

    openJiuwen-ai/sciencediscovery

    A skill your agent uses when you need to write and execute Python/R code to process, transform, and analyze data, delivering reproducible computational results with complete code-level methodology…

    148 GitHub stars~2.8k tokensUpdated 6 days ago
    Data & AnalyticsAuto-check passed
  • Data Analysis

    xiaoyuge886/aigc

    Perform data analysis tasks including data cleaning, statistical analysis, visualization, and insight generation.

    198 GitHub starsUsed in 1 repo~794 tokens
    Data & AnalyticsAuto-check passed
  • Data Explorer

    liangdabiao/claude-data-analysis-ultra-main

    Performs exploratory data analysis, statistical analysis, and pattern discovery.

    290 GitHub stars~2.1k tokensUpdated 5 mo ago
    Data & AnalyticsAuto-check passed
  • Profiling Tables

    astronomer/agents

    Deep-dive data profiling for a specific table. An agent skill from astronomer/agents.

    450 GitHub stars~964 tokensUpdated 2 days ago
    Data & AnalyticsAuto-check passed
  • Turns a shotgun profiler table (MetaPhlAn relative abundance, Bracken counts, HUMAnN function tables) into honest figures and defensible community statistics with phyloseq, vegan, microViz, and…

    1.2k GitHub starsUsed in 1 repo~3.7k tokens
    Data & AnalyticsAuto-check passed

More from aspi6246/Claude-Code-Skills-for-Academics

All 15 skills in this repo
  • Beamer Check

    aspi6246/Claude-Code-Skills-for-Academics

    Audit and fix Beamer slide decks for compilation errors, layout problems, and visual issues.

    157 GitHub stars~1.7k tokensUpdated 1 mo ago
    Auto-check passed
  • Beamer Slides Teaching

    aspi6246/Claude-Code-Skills-for-Academics

    Generate LaTeX Beamer slide decks using my defined style and conventions.

    157 GitHub stars~1.3k tokensUpdated 1 mo ago
    Auto-check passed
  • Edmans Audit

    aspi6246/Claude-Code-Skills-for-Academics

    Audit a finance paper against Alex Edmans' "Learnings From 1,000 Rejections" (2025, Financial Management) three-part framework: Contribution, Execution, Exposition.

    157 GitHub stars~1.6k tokensUpdated 1 mo ago
    Auto-check passed
  • R Empirical Finance

    aspi6246/Claude-Code-Skills-for-Academics

    Conventions for writing empirical finance R code with data.table, fixest, arrow, and ggplot2.

    157 GitHub stars~1.3k tokensUpdated 1 mo ago
    Auto-check passed
  • Top5 Framing

    aspi6246/Claude-Code-Skills-for-Academics

    Diagnose and raise the general-interest ceiling of an academic economics or finance paper so it reads like a top-5 (or top-3 finance) submission.

    157 GitHub stars~2k tokensUpdated 1 mo ago
    Auto-check passed
  • Canvas Lms API

    aspi6246/Claude-Code-Skills-for-Academics

    A skill your agent uses whenever working with the Canvas LMS REST API — including creating or updating modules, pages, and files for a course, building sync scripts, managing course structure…

    157 GitHub stars~3.7k tokensUpdated 1 mo ago
    Auto-check passed

Questions about Data Profiler

What does Data Profiler do?

Systematic dataset profiling protocol for empirical research. Data Profiler is an agent skill from aspi6246/Claude-Code-Skills-for-Academics. Systematic dataset profiling protocol for empirical research.

When should I use Data Profiler?

Data Profiler fits situations like: the user has a new dataset and wants to understand it before analysis — including unit of observation; variable definitions; panel structure; descriptive statistics.

How do I install Data Profiler in Claude Code?

Run `npx skills add aspi6246/Claude-Code-Skills-for-Academics --skill data-profiler -a claude-code`. Or copy the skill folder (data-profiler in aspi6246/Claude-Code-Skills-for-Academics) into .claude/skills/data-profiler in your project. Claude Code loads it when a task matches its description.

How do I install Data Profiler in Codex?

Run `npx skills add aspi6246/Claude-Code-Skills-for-Academics --skill data-profiler -a codex`. Or copy the skill folder (data-profiler in aspi6246/Claude-Code-Skills-for-Academics) into .agents/skills/data-profiler in your project. Codex loads it when a task matches its description.

Can I use Data Profiler in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add aspi6246/Claude-Code-Skills-for-Academics --skill data-profiler -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/data-profiler, .gemini/skills/data-profiler, .github/skills/data-profiler and .opencode/skills/data-profiler in your project.

What does Data Profiler need to run?

SKILL.md names no scripts, command-line tools or credentials: Data Profiler is instructions for the agent only.

Does Data Profiler access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Data Profiler safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Data Profiler use?

No licence was found for Data Profiler or its repository. Without one, default copyright applies: ask the author before reusing or redistributing it.

How many tokens does Data Profiler use?

About 2k tokens (SKILL.md is roughly 7.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.1k tokens, read only when the agent opens those files.

What are the alternatives to Data Profiler?

Skills that share tags, products or a category with Data Profiler: Pandas Pro (Jeffallan/claude-skills, 12k stars), Code Engineer (openJiuwen-ai/sciencediscovery, 148 stars), Data Analysis (xiaoyuge886/aigc, 198 stars) and Data Explorer (liangdabiao/claude-data-analysis-ultra-main, 290 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Data Profiler?

aspi6246 (a GitHub user) maintains it in aspi6246/Claude-Code-Skills-for-Academics, which has 157 GitHub stars. The repository holds 15 skills in this directory. The repository was last updated on August 21, 2026.

Source: aspi6246/Claude-Code-Skills-for-Academics on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.