Agent skill

Dataset Curation

by fcakyon in fcakyon/phd-skills

A skill your agent uses when the user wants to analyze dataset bias, create stratified samples, evaluate fairness, or plan dataset collection.

MITAuto-check passedResearch & Science

Install Dataset Curation

skills CLI
$ npx skills add fcakyon/phd-skills --skill dataset-curation -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install fcakyon/phd-skills dataset-curation --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/fcakyon/phd-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugin/skills/dataset-curation .claude/skills/dataset-curation && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
dataset-curation
GitHub stars
414
Token cost
~1k tokens
SKILL.md length
468 words
Files
1
Skills in repo
12
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when the user wants to analyze dataset bias, create stratified samples, evaluate fairness, or plan dataset collection.

  • Works in 6 steps: Distribution Analysis → Bias Assessment → Stratified Sampling → …
  • The user wants to analyze dataset bias
  • SKILL.md covers Step 1: Distribution Analysis, Step 2: Bias Assessment, Step 3: Stratified Sampling and Step 4: Quality Assessment, plus 3 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Dataset Curation is an agent skill from fcakyon/phd-skills. Use when the user wants to analyze dataset bias, create stratified samples, evaluate fairness, or plan dataset collection. Triggers on phrases like "dataset bias", "stratified sample", "class imbalance", "data distribution", "fairness analysis", or "ethical review".

Its SKILL.md is about 1k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Research & Science. The repository describes itself as: PhD Research Skills for Claude Code: paper reproduction, experiment design, paper review, result comparison and more. The licence is MIT.

When your agent uses it

  • The user wants to analyze dataset bias
  • Create stratified samples
  • Evaluate fairness
  • Plan dataset collection

Example prompts

  • “dataset bias”
  • “stratified sample”
  • “class imbalance”
  • “/dataset-curation”

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Distribution Analysis
  2. Bias Assessment
  3. Stratified Sampling
  4. Quality Assessment
  5. Expansion Recommendations
  6. Ethical Review Checklist

What it can do on your machine

Read from SKILL.md and the folder at commit 67acd61. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Dataset Curation loads about 1k tokens when it runs. Until then it costs about 71 tokens; SKILL.md has 468 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~71
When it runs · the whole SKILL.md, loaded when a task matches
~1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from fcakyon/phd-skills at commit 67acd61, republished under its MIT licence (© fcakyon). 468 words, ~1,032 tokens.

Download SKILL.mdSave it as .claude/skills/dataset-curation/SKILL.md (or your agent's skills folder).
name
dataset-curation
description
Use when the user wants to analyze dataset bias, create stratified samples, evaluate fairness, or plan dataset collection. Triggers on phrases like "dataset bias", "stratified sample", "class imbalance", "data distribution", "fairness analysis", or "ethical review".

Dataset Curation Methodology

You are helping a researcher curate, analyze, or expand a dataset with attention to bias, fairness, and quality.

Step 1: Distribution Analysis

Before any curation action, understand the current state:

Per-Class Distribution
  • Count instances per class/label/tag
  • Compute imbalance ratio (max_count / min_count)
  • Identify severely underrepresented classes (< 5% of max class)
  • Visualize: bar chart of class frequencies sorted by count
Co-occurrence Analysis
  • Build co-occurrence matrix: which labels appear together
  • Identify spurious correlations (e.g., "violence" always co-occurs with "male")
  • Check for label leakage between splits
Metadata Distribution
  • Source diversity: how many sources/movies/documents contribute
  • Temporal distribution: are all time periods represented?
  • Content diversity: genre, style, domain coverage

Step 2: Bias Assessment

For each identified imbalance or correlation:

  1. Is it real-world reflective? Some imbalances reflect genuine phenomena
  2. Is it harmful? Would a model trained on this data make unfair predictions?
  3. Is it fixable? Can we collect more data, resample, or reweight?
Fairness Dimensions

Check for bias along relevant protected attributes:

  • Gender representation (if applicable)
  • Racial/ethnic representation (if applicable)
  • Age distribution (if applicable)
  • Geographic/cultural diversity (if applicable)
Bias Metrics
  • Demographic parity: equal positive rates across groups
  • Equalized odds: equal TPR and FPR across groups
  • Representation ratio: group proportion in data vs population

Step 3: Stratified Sampling

When creating splits (train/val/test):

  1. Primary stratification: by label/class distribution
  2. Secondary stratification: by source (prevent source leakage across splits)
  3. Validation:
    • Chi-squared test for label distribution similarity across splits
    • No source overlap between splits
    • Rare classes have minimum representation in each split

Split ratios depend on dataset size:

  • Large (>50k): 80/10/10 or 90/5/5
  • Medium (5k-50k): 70/15/15 or 80/10/10
  • Small (<5k): k-fold cross-validation preferred
Show full SKILL.md (198 more words)Show less

Step 4: Quality Assessment

For labeled datasets, assess annotation quality:

  • Inter-annotator agreement: Cohen's kappa, Fleiss' kappa, or Krippendorff's alpha
  • Label noise estimation: sample and manually verify N labels
  • Edge cases: identify ambiguous examples that annotators might disagree on
  • Consistency checks: automated rules for label validity

Step 5: Expansion Recommendations

If the dataset needs more data:

  1. Priority classes: which classes benefit most from more data
  2. Source suggestions: where to find more data for underrepresented classes
  3. Collection strategy: active learning, targeted scraping, synthetic augmentation
  4. Cost estimation: time and resources for each approach

Step 6: Ethical Review Checklist

Before using or publishing any dataset:

  • Content sensitivity: does the data contain sensitive material?
  • Consent: was data collected with appropriate consent?
  • Privacy: are individuals identifiable? Is anonymization needed?
  • Licensing: are data sources used within their license terms?
  • Potential harms: could the dataset be misused?
  • Documentation: is the dataset documented with a datasheet/data card?

Output Format

Produce:

  1. Distribution report: per-class counts, imbalance ratios, co-occurrence matrix
  2. Bias findings: identified biases with severity and actionability
  3. Split recommendation: stratification strategy with validation results
  4. Expansion plan: prioritized suggestions for addressing gaps
  5. Ethics checklist: completed checklist with notes per item

© fcakyon, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in plugin/skills/dataset-curation of fcakyon/phd-skills.

Open the folder on GitHubat commit 67acd61

Compare with similar skills

Dataset Curation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Dataset Curation compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Dataset Curation this skillfcakyon/phd-skills414—~1kAutomated safety check: PassMIT
Hypothesis Generationspacering-net/codeg3.8k15 repos~3.6kAutomated safety check: NotesMIT
GitHub Deep Researchbytedance/deer-flow83k5 repos~1.3kAutomated safety check: PassMIT
Nature Paper CardYuan1z0825/nature-skills46k2 repos~2.1kAutomated safety check: PassApache-2.0
Read arXiv Paperkarpathy/nanochat58k2 repos~494Automated safety check: PassMIT
Content Research Writerweapp-tailwindcss/weapp-tailwindcss1.9k25 repos~3.5kAutomated safety check: PassMIT

Similar skills

  • Hypothesis Generation

    spacering-net/codeg

    Structured hypothesis formulation from observations. An agent skill from spacering-net/codeg.

    3.8k GitHub starsUsed in 15 repos~3.6k tokens
    Research & ScienceAuto-check: notes
  • GitHub Deep Research

    bytedance/deer-flow

    Researches a GitHub repository over four rounds using the GitHub API and web search, then writes a structured markdown report with timeline, metrics and Mermaid diagrams.

    83k GitHub starsUsed in 5 repos~1.3k tokens
    Research & ScienceAuto-check passed
  • Nature Paper Card

    Yuan1z0825/nature-skills

    Builds a structured deep-reading card for one scientific paper, covering methods, how experiments support claims, limitations and research ideas, with a script to prepare the source.

    46k GitHub starsUsed in 2 repos~2.1k tokens
    Research & ScienceAuto-check passed
  • Read arXiv Paper

    karpathy/nanochat

    Fetches the TeX source of an arXiv paper from its URL, reads it and writes a markdown summary tied to the nanochat project.

    58k GitHub starsUsed in 2 repos~494 tokens
    Research & ScienceAuto-check passed
  • Content Research Writer

    weapp-tailwindcss/weapp-tailwindcss

    Assists in writing high-quality content by conducting research, adding citations, improving hooks, iterating on outlines, and providing real-time feedback on each section.

    1.9k GitHub starsUsed in 25 repos~3.5k tokens
    Research & ScienceAuto-check passed
  • Peer Review

    spacering-net/codeg

    Structured manuscript/grant review with checklist-based evaluation.

    3.8k GitHub starsUsed in 18 repos~5.9k tokens
    Research & ScienceAuto-check: notes

More from fcakyon/phd-skills

All 12 skills in this repo
  • Reproduce

    fcakyon/phd-skills

    End-to-end paper reproduction from arxiv URL through smoke runs to replication experiments.

    414 GitHub stars~1.1k tokensUpdated 21 days ago
    Auto-check passed
  • Compare

    fcakyon/phd-skills

    Same-epoch comparison of training runs across wandb, neptune, tensorboard, or mlflow.

    414 GitHub stars~1.2k tokensUpdated 21 days ago
    Auto-check passed
  • Debug

    fcakyon/phd-skills

    Evidence-before-action diagnosis of failing ML experiments. An agent skill from fcakyon/phd-skills.

    414 GitHub stars~1.3k tokensUpdated 21 days ago
    Auto-check passed
  • Experiment Design

    fcakyon/phd-skills

    A skill your agent uses when the user wants to design experiments, plan ablation studies, structure baselines, or create incremental evaluation strategies.

    414 GitHub stars~987 tokensUpdated 21 days ago
    Auto-check passed
  • Latex Setup

    fcakyon/phd-skills

    A skill your agent uses when the user wants to set up or troubleshoot a LaTeX environment, choose between biber and bibtex, install packages for a specific venue template, or configure compilation.

    414 GitHub stars~1.1k tokensUpdated 21 days ago
    Auto-check: notes
  • Launch

    fcakyon/phd-skills

    Pre-flight checklist for long-running ML training jobs covering config diff, run naming, path verification, monitoring setup, and restart-cleanup.

    414 GitHub stars~1.4k tokensUpdated 21 days ago
    Auto-check passed

Questions about Dataset Curation

What does Dataset Curation do?

A skill your agent uses when the user wants to analyze dataset bias, create stratified samples, evaluate fairness, or plan dataset collection. Dataset Curation is an agent skill from fcakyon/phd-skills. Use when the user wants to analyze dataset bias, create stratified samples, evaluate fairness, or plan dataset collection.

When should I use Dataset Curation?

Dataset Curation fits situations like: the user wants to analyze dataset bias; create stratified samples; evaluate fairness; plan dataset collection.

How do I install Dataset Curation in Claude Code?

Run `npx skills add fcakyon/phd-skills --skill dataset-curation -a claude-code`. Or copy the skill folder (plugin/skills/dataset-curation in fcakyon/phd-skills) into .claude/skills/dataset-curation in your project. Claude Code loads it when a task matches its description.

How do I install Dataset Curation in Codex?

Run `npx skills add fcakyon/phd-skills --skill dataset-curation -a codex`. Or copy the skill folder (plugin/skills/dataset-curation in fcakyon/phd-skills) into .agents/skills/dataset-curation in your project. Codex loads it when a task matches its description.

Can I use Dataset Curation in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add fcakyon/phd-skills --skill dataset-curation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/dataset-curation, .gemini/skills/dataset-curation, .github/skills/dataset-curation and .opencode/skills/dataset-curation in your project.

What does Dataset Curation need to run?

SKILL.md names no scripts, command-line tools or credentials: Dataset Curation is instructions for the agent only.

Does Dataset Curation access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Dataset Curation safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Dataset Curation use?

Dataset Curation is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Dataset Curation use?

About 1k tokens (SKILL.md is roughly 4.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Dataset Curation?

Skills that share tags, products or a category with Dataset Curation: Hypothesis Generation (spacering-net/codeg, 3.8k stars), GitHub Deep Research (bytedance/deer-flow, 83k stars), Nature Paper Card (Yuan1z0825/nature-skills, 46k stars) and Read arXiv Paper (karpathy/nanochat, 58k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Dataset Curation?

fcakyon (a GitHub user) maintains it in fcakyon/phd-skills, which has 414 GitHub stars. The repository holds 12 skills in this directory. The repository was last updated on September 16, 2026.

Source: fcakyon/phd-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.