Agent skill

Dataset Datasheet

by mohitagw15856 in mohitagw15856/pm-claude-skills

Document a dataset so others know what it is, how it was made, and when not to use it.

MITAuto-check passed

Install Dataset Datasheet

skills CLI
$ npx skills add mohitagw15856/pm-claude-skills --skill dataset-datasheet -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install mohitagw15856/pm-claude-skills dataset-datasheet --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/mohitagw15856/pm-claude-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/dataset-datasheet .claude/skills/dataset-datasheet && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
dataset-datasheet
GitHub stars
1.4k
Token cost
~939 tokens
SKILL.md length
453 words
Files
1
Skills in repo
1,348
Repo updated
First seen
Licence
MIT

At a glance

Document a dataset so others know what it is, how it was made, and when not to use it.

  • Asked to write a datasheet for a dataset
  • SKILL.md covers Required Inputs, Output Format, Quality Checks and Anti-Patterns, plus 2 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • Document training/eval data

What it does

Dataset Datasheet is an agent skill from mohitagw15856/pm-claude-skills. Document a dataset so others know what it is, how it was made, and when not to use it. Use when asked to write a datasheet for a dataset, document training/eval data, or assess whether a dataset is fit for a use. Produces a datasheet — motivation, composition, collection process, preprocessing, recommended uses & limits, distribution, and maintenance.

Its SKILL.md is about 940 tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

The repository describes itself as: 1255 professional Agent Skills for Claude, ChatGPT, Gemini, Cursor & Codex — PRDs, postmortems, leases, medical bills, layoffs, go-bags, new countries. Plain markdown, MIT, in… The licence is MIT.

When your agent uses it

  • Asked to write a datasheet for a dataset
  • Document training/eval data
  • Assess whether a dataset is fit for a use

Example prompts

  • “/dataset-datasheet”

What it can do on your machine

Read from SKILL.md and the folder at commit 1cbf1f0. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Dataset Datasheet loads about 939 tokens when it runs. Until then it costs about 93 tokens; SKILL.md has 453 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~93
When it runs · the whole SKILL.md, loaded when a task matches
~939

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from mohitagw15856/pm-claude-skills at commit 1cbf1f0, republished under its MIT licence (© mohitagw15856). 453 words, ~939 tokens.

Download SKILL.mdSave it as .claude/skills/dataset-datasheet/SKILL.md (or your agent's skills folder).
name
dataset-datasheet
description
Document a dataset so others know what it is, how it was made, and when not to use it. Use when asked to write a datasheet for a dataset, document training/eval data, or assess whether a dataset is fit for a use. Produces a datasheet — motivation, composition, collection process, preprocessing, recommended uses & limits, distribution, and maintenance.

Dataset Datasheet Skill

Models inherit the flaws of their data, and most data debt is invisible because nobody wrote down where the data came from. A datasheet is that record: how the dataset was collected, what's in it, what's missing, and what it should not be used for. It's the difference between a reusable asset and a liability.

Required Inputs

Ask for these only if they aren't already provided:

  • Dataset name, version, owner and what it's used for today.
  • Motivation — why it was created and for what task.
  • Composition — what an instance is, how many, fields/labels, and time range.
  • Collection — sources, method (scraped, logged, purchased, annotated), and consent/licensing basis.
  • Known issues — gaps, imbalances, label noise, sensitive attributes, duplicates.

Output Format

Datasheet: [dataset] v[version]

Owner: [team] · Created: [date] · License: [license]

1. Motivation — why this dataset exists, the task it serves, and who funded/created it.

2. Composition

  • What a single instance represents; total count; the schema (fields, label definitions).
  • Class/label balance and key distributions (and notable skews).
  • Sensitive attributes present (directly or by proxy), and whether individuals are identifiable.
  • Known missing data, duplicates, or noise.

3. Collection process — sources, mechanism (scrape/log/survey/annotation), time window, sampling strategy, and the legal/consent basis (license, ToS, opt-in).

4. Preprocessing / labelling — cleaning, dedup, filtering, and how labels were produced (who annotated, guidelines, inter-annotator agreement).

5. Recommended uses & limits

  • Appropriate uses: tasks this data supports well.
  • Do not use for: tasks where its biases/gaps would cause harm or invalid results.

6. Distribution & access — who can use it, how it's shared, and tenancy/PII handling.

7. Maintenance — owner, update cadence, versioning, and how errors get reported and fixed.

Show full SKILL.md (186 more words)Show less

Quality Checks

  • The collection method and legal/consent basis are stated — not assumed
  • Class balance and key distribution skews are quantified, not hand-waved
  • Sensitive attributes (and proxies for them) are identified explicitly
  • "Do not use for" lists concrete tasks where the data would mislead
  • Label provenance is documented (who labelled, with what guidelines, and agreement level)
  • An owner and update/error-reporting process are named

Anti-Patterns

  • Do not describe only the happy-path contents — the gaps, skews, and noise are what cause model failures
  • Do not omit the consent/licensing basis — "we scraped it" is a legal and ethical liability if undocumented
  • Do not ignore proxy variables — removing race/gender columns doesn't remove the bias if zip code or name encodes it
  • Do not present label quality as perfect — state who labelled it and the agreement rate, or note it's unmeasured
  • Do not leave the dataset ownerless — an unmaintained dataset silently rots as the world changes

Based On

Datasheets for Datasets (Gebru et al., 2018) and data-documentation practice in responsible-AI reviews.

Example Trigger Phrases

  • "Write a datasheet for a dataset."
  • "Document training/eval data."
  • "Assess whether a dataset is fit for a use."

© mohitagw15856, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/dataset-datasheet of mohitagw15856/pm-claude-skills.

Open the folder on GitHubat commit 1cbf1f0

Compare with similar skills

Dataset Datasheet next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Dataset Datasheet compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Dataset Datasheet this skillmohitagw15856/pm-claude-skills1.4k—~939Automated safety check: PassMIT
DatasetsArize-ai/phoenix12k—~1.6kAutomated safety check: PassCustom licence
Dataset Curationwshobson/agents40k—~2kAutomated safety check: PassMIT
Arize Datasetgithub/awesome-copilot40k1 repos~3.9kAutomated safety check: NotesMIT
Create Test DatasetsClickHouse/ClickHouse50k—~1kAutomated safety check: PassApache-2.0
Hugging Face Datasetssickn33/agentic-awesome-skills47k2 repos~1.1kAutomated safety check: PassMIT

Similar skills

  • Datasets

    Arize-ai/phoenix

    Understand what a Phoenix dataset is and reason well about its examples, outputs, splits, and how it feeds evaluators and experiments.

    12k GitHub stars~1.6k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Dataset Curation

    wshobson/agents

    Prepare, format, and validate datasets for supervised fine-tuning and preference training.

    40k GitHub stars~2k tokensUpdated 5 days ago
    AI & LLM EngineeringAuto-check passed
  • Arize Dataset

    github/awesome-copilot

    Official

    Creates, manages, and queries Arize datasets and examples. An agent skill from github/awesome-copilot.

    40k GitHub starsUsed in 1 repo~3.9k tokens
    Testing & QAAuto-check: notes
  • Create Test Datasets

    ClickHouse/ClickHouse

    Create test datasets (hits, visits, tpcds, tpch) from standard scripts.

    50k GitHub stars~1k tokensUpdated today
    DatabasesAuto-check passed
  • Hugging Face Datasets

    sickn33/agentic-awesome-skills

    Create and manage datasets on Hugging Face Hub. An agent skill from sickn33/agentic-awesome-skills.

    47k GitHub starsUsed in 2 repos~1.1k tokens
    AI & LLM EngineeringAuto-check passed
  • Dummy Dataset Generator

    phuryn/pm-skills

    Generates realistic test datasets with custom columns, row counts and business constraints, output as CSV, JSON, SQL inserts or a runnable Python script.

    27k GitHub stars~983 tokensUpdated 25 days ago
    Testing & QAAuto-check passed

More from mohitagw15856/pm-claude-skills

All 1,348 skills in this repo
  • Car Tco

    mohitagw15856/pm-claude-skills

    Compare the total cost of car ownership across buy-new, buy-used, lease, and keep-your-current-car — depreciation, insurance, maintenance ramp, and fuel over a real horizon, not just the monthly…

    1.4k GitHub stars~1.1k tokensUpdated yesterday
    Auto-check passed
  • Cs Health Scorecard

    mohitagw15856/pm-claude-skills

    Build a customer health scorecard for a specific account. An agent skill from mohitagw15856/pm-claude-skills.

    1.4k GitHub stars~2.4k tokensUpdated yesterday
    Auto-check passed
  • Exit Waterfall

    mohitagw15856/pm-claude-skills

    Compute who gets what at each exit price from a cap table — liquidation preferences, conversion points, and where the founders' share collapses.

    1.4k GitHub stars~1.1k tokensUpdated yesterday
    Auto-check passed
  • Feature Prioritisation

    mohitagw15856/pm-claude-skills

    Apply prioritisation frameworks (RICE, MoSCoW, Kano, ICE, Opportunity Scoring) to rank features and backlog items.

    1.4k GitHub stars~2k tokensUpdated yesterday
    Auto-check passed
  • Fire Number

    mohitagw15856/pm-claude-skills

    Compute a financial-independence (FIRE) target and years-to-reach with every assumption labeled as an assumption — plus a sensitivity table instead of a single false-precision answer.

    1.4k GitHub stars~1.1k tokensUpdated yesterday
    Auto-check passed
  • Freelance Rate

    mohitagw15856/pm-claude-skills

    Derive a freelance day/hourly rate backwards from target income, honest billable utilization, overhead, and the self-employment tax premium — the arithmetic that proves a rate is not salary÷2000.

    1.4k GitHub stars~1.2k tokensUpdated yesterday
    Auto-check passed

Questions about Dataset Datasheet

What does Dataset Datasheet do?

Document a dataset so others know what it is, how it was made, and when not to use it. Dataset Datasheet is an agent skill from mohitagw15856/pm-claude-skills. Document a dataset so others know what it is, how it was made, and when not to use it.

When should I use Dataset Datasheet?

Dataset Datasheet fits situations like: asked to write a datasheet for a dataset; document training/eval data; assess whether a dataset is fit for a use.

How do I install Dataset Datasheet in Claude Code?

Run `npx skills add mohitagw15856/pm-claude-skills --skill dataset-datasheet -a claude-code`. Or copy the skill folder (skills/dataset-datasheet in mohitagw15856/pm-claude-skills) into .claude/skills/dataset-datasheet in your project. Claude Code loads it when a task matches its description.

How do I install Dataset Datasheet in Codex?

Run `npx skills add mohitagw15856/pm-claude-skills --skill dataset-datasheet -a codex`. Or copy the skill folder (skills/dataset-datasheet in mohitagw15856/pm-claude-skills) into .agents/skills/dataset-datasheet in your project. Codex loads it when a task matches its description.

Can I use Dataset Datasheet in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add mohitagw15856/pm-claude-skills --skill dataset-datasheet -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/dataset-datasheet, .gemini/skills/dataset-datasheet, .github/skills/dataset-datasheet and .opencode/skills/dataset-datasheet in your project.

What does Dataset Datasheet need to run?

SKILL.md names no scripts, command-line tools or credentials: Dataset Datasheet is instructions for the agent only.

Does Dataset Datasheet access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Dataset Datasheet safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Dataset Datasheet use?

Dataset Datasheet is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Dataset Datasheet use?

About 939 tokens (SKILL.md is roughly 3.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Dataset Datasheet?

Skills that share tags, products or a category with Dataset Datasheet: Datasets (Arize-ai/phoenix, 12k stars), Dataset Curation (wshobson/agents, 40k stars), Arize Dataset (github/awesome-copilot, 40k stars) and Create Test Datasets (ClickHouse/ClickHouse, 50k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Dataset Datasheet?

mohitagw15856 (a GitHub user) maintains it in mohitagw15856/pm-claude-skills, which has 1,433 GitHub stars. The repository holds 1,348 skills in this directory. The repository was last updated on October 8, 2026.

Source: mohitagw15856/pm-claude-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.