Agent skill

Dataset Curation

by wshobson in wshobson/agents

Prepare, format, and validate datasets for supervised fine-tuning and preference training.

MITAuto-check passedAI & LLM Engineering

Install Dataset Curation

skills CLI
$ npx skills add wshobson/agents --skill dataset-curation -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install wshobson/agents dataset-curation --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/wshobson/agents.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/llm-finetuning/skills/dataset-curation .claude/skills/dataset-curation && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
dataset-curation
GitHub stars
40k
Token cost
~2k tokens
SKILL.md length
927 words
Files
3 (incl. references)
Skills in repo
142
Repo updated
First seen
Licence
MIT

At a glance

Prepare, format, and validate datasets for supervised fine-tuning and preference training.

  • Works in 6 steps: Format matches the method (table above). → Template applied before concatenation. → Loss masked to assistant turns only. → …
  • Converting raw data into training format
  • SKILL.md covers Format Selection, Chat Templates and Loss Masking, Packing and Synthetic Data Rules, plus 2 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Dataset Curation is an agent skill from wshobson/agents. Prepare, format, and validate datasets for supervised fine-tuning and preference training. Use when converting raw data into training format, applying chat templates, configuring sequence packing, generating synthetic training data, or writing a dataset card before a run.

Its SKILL.md is about 2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files, including reference files (for example `references/formats-and-templates.md` and `references/synthetic-data.md`).

It sits in AI & LLM Engineering, covering Fine-tuning. The repository describes itself as: Multi-harness agentic plugin marketplace for Claude Code, Codex, Cursor, OpenCode, GitHub Copilot, Google Antigravity, and Pi. The licence is MIT.

When your agent uses it

  • Converting raw data into training format
  • Applying chat templates
  • Configuring sequence packing
  • Generating synthetic training data

Example prompts

  • “/dataset-curation”

Requirements

  • Python 3

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Format matches the method (table above).
  2. Template applied before concatenation.
  3. Loss masked to assistant turns only.
  4. 5–10 packed sequences decoded and read.
  5. ≥25% real data in the final mix.
  6. Dataset card complete — all six fields.

What it can do on your machine

Read from SKILL.md and the folder at commit 46891e7. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python and json).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Dataset Curation loads about 2k tokens when it runs, and up to ~6.8k if it reads all its reference files. Until then it costs about 72 tokens; SKILL.md has 927 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~72
When it runs · the whole SKILL.md, loaded when a task matches
~2k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~6.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from wshobson/agents at commit 46891e7, republished under its MIT licence (© wshobson). 927 words, ~1,980 tokens.

Download SKILL.mdSave it as .claude/skills/dataset-curation/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
dataset-curation
description
Prepare, format, and validate datasets for supervised fine-tuning and preference training. Use when converting raw data into training format, applying chat templates, configuring sequence packing, generating synthetic training data, or writing a dataset card before a run.

Dataset Curation

This skill assumes finetuning-method-selection already routed here — the next step is preparing data, not choosing a method. What follows: format selection by target method, the template/packing mechanics behind the most common silent training failures, rules for mixing in synthetic data without collapse, and the dataset card that closes out Phase 2 before a run starts.

Input: raw examples (demonstrations, preference judgments, or task prompts) plus a routing decision from finetuning-method-selection. Output format: a formatted, packed, validated JSONL dataset plus a completed dataset card — the Phase 2 artifact /finetune checks before launching training.

Format Selection

MethodShapeRows
SFT, single-turnInstruct (instruction/response or prompt/completion)~1,000+ floor
SFT, multi-turnConversation / ChatML messages list~1,000+ floor
DPO / ORPOPreference pair (prompt, chosen, rejected)Method-dependent, see preference-optimization
KTOUnpaired (prompt, completion, label)Method-dependent, see preference-optimization
GRPO / RLVRPrompt-only (prompt + verifier metadata)Method-dependent, see grpo-rlvr-training
  • ~1,000+ rows is the recommended floor for SFT, not a target. Below it, a handful of low-quality or duplicate examples can dominate the gradient; above it, quality over quantity — a smaller verified, deduplicated set beats a larger noisy one.

  • The ChatML shape, for orientation; the other four formats plus a ShareGPT conversion note live in references/formats-and-templates.md:

    json
    {"messages": [
      {"role": "user", "content": "..."},
      {"role": "assistant", "content": "..."}
    ]}

Chat Templates and Loss Masking

Apply the target model's chat template before any concatenation or packing, never after — packing raw text and templating the packed blob afterward corrupts turn boundaries, landing role markers in the wrong place relative to each example.

  • Train on assistant responses only. Mask the loss (-100 in the labels tensor) over system/user turns and the template's own role markers — only assistant-turn content tokens contribute to loss.

  • Template/tokenizer mismatches are a top silent failure mode. A model trained against one chat template but served or evaluated with a different one degrades without erroring. Verify the same template string used in training is applied at inference and eval time.

  • Keep the dataset in messages shape and let the trainer template and mask it (assistant_only_loss=True in current TRL) — pre-rendering to a flat text field destroys the turn boundaries masking needs. Full code sketch: references/formats-and-templates.md. Sanity-check before training — decode only unmasked positions; expect only assistant text:

    python
    keep = batch["labels"][0] != -100
    print(tokenizer.decode(batch["input_ids"][0][keep]))

Packing

Without packing, 40–70% of compute is spent on padding — variable-length examples batched at a fixed sequence length waste the gap between each example's length and the batch's max. Packing concatenates multiple examples into one sequence up to the max length, cutting most of that waste.

  • Packing changes batch semantics. A packed sequence can contain several original examples, so "steps per epoch" and any LR schedule keyed to example count shift once packing is on — recompute schedule milestones against packed-sequence count.

  • MANDATORY: decode and manually inspect 5–10 packed sequences before scaling to a full run. Confirm example boundaries land where expected, template markers are intact per sub-example, and the loss mask is still assistant-only within each packed sequence. Not optional — packing bugs are silent (the loss curve looks normal) and only surface in eval quality, hours later:

    python
    for seq in packed_dataset.select(range(10)):
        print(tokenizer.decode(seq["input_ids"]))
Show full SKILL.md (436 more words)Show less

Synthetic Data Rules

  • Keep ≥25% real data as a collapse guard. Training on a growing share of model-generated data without a real-data floor drives measurable quality collapse over successive generations — 25% real is the minimum that holds the line. General-domain replay rows count toward this floor — "real" means "not generated for this task from this student," not "human-authored." An all-synthetic-by-construction dataset can meet the ≥25% floor through replay alone (see references/synthetic-data.md's Replay-Mix Construction recipe); state which rows count as "real" in the dataset card rather than leaving the floor structurally unmeetable.
  • Magpie and rejection sampling are the workhorses. Magpie extracts prompts from the model's own template prior; rejection sampling generates several candidates per prompt and keeps only the ones a filter passes. Both beat naive single-shot generation.
  • Targeted, student-aware generation beats static generation by 1.3–2x sample efficiency — aiming at the student's actual failure modes hits a quality bar with fewer filtered examples.
  • Typical accept rates after filtering run 10–30%. Plan volume accordingly — a 10,000-row target at 15% accept needs ~65,000+ raw generations.
  • Generation-method ranking, filter funnel, replay- mix construction, and distillation pattern: references/synthetic-data.md.

The Dataset Card

Every dataset that reaches training gets a card — the required Phase 2 artifact /finetune checks before launching. The card is not free-form documentation; it MUST carry these fields:

  • Provenance — where every row came from (real source(s), synthetic method(s), or both), traceable to trace-to-training-data output.
  • Counts — total rows, and rows per split (train/eval/held-out) if split.
  • Synthetic/real ratio — the measured ratio, checked against the ≥25% real floor above.
  • Dedup method — exact-match, semantic (embedding threshold), or both; see the filter funnel in references/synthetic-data.md.
  • Template used — the exact chat template string/identifier, kept consistent through inference and eval — this is what ties an eval-harness-first run back to the checkpoint.
  • Packing config — whether packing was used, max sequence length, and confirmation the 5–10-sequence manual inspection above was done.

A dataset missing any of these six fields isn't ready for /finetune — the card is a gate, not a summary written after the fact.

Phase 2 Exit Checklist

Before handing off to /finetune, confirm:

  1. Format matches the method (table above).
  2. Template applied before concatenation.
  3. Loss masked to assistant turns only.
  4. 5–10 packed sequences decoded and read.
  5. ≥25% real data in the final mix.
  6. Dataset card complete — all six fields.

References

  • references/formats-and-templates.md — JSONL examples per format, current-TRL masking code, and the ShareGPT conversion note.
  • references/synthetic-data.md — generation-method ranking, filter funnel, replay-mix construction, and teacher→student distillation pattern.

Related skills: finetuning-method-selection routes here; lora-qlora-recipes, vision-sft, and preference-optimization consume the datasets this skill produces; trace-to-training-data is the provenance source for graded-trajectory datasets; eval-harness-first grades the resulting checkpoint.

© wshobson, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (references) in plugins/llm-finetuning/skills/dataset-curation of wshobson/agents.

  • SKILL.md
  • references/formats-and-templates.md
  • references/synthetic-data.md

Open the folder on GitHubat commit 46891e7

Compare with similar skills

Dataset Curation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Dataset Curation compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Dataset Curation this skillwshobson/agents40k—~2kAutomated safety check: PassMIT
Peft Fine TuningOrchestra-Research/AI-Research-SKILLs13k9 repos~3.1kAutomated safety check: PassMIT
Hugging Face LLM Trainerhuggingface/skills11k3 repos~7.2kAutomated safety check: PassApache-2.0
Sentence-Transformers Training Routerhuggingface/skills11k1 repos~2.6kAutomated safety check: PassApache-2.0
Dataset Evaluationawslabs/agent-plugins9152 repos~1.3kAutomated safety check: PassApache-2.0
Train RlOpenPipe/ART11k—~2.4kAutomated safety check: PassApache-2.0

Similar skills

  • Peft Fine Tuning

    Orchestra-Research/AI-Research-SKILLs

    Parameter-efficient fine-tuning for LLMs using LoRA, QLoRA, and 25+ methods.

    13k GitHub starsUsed in 9 repos~3.1k tokens
    AI & LLM EngineeringAuto-check passed
  • Hugging Face LLM Trainer

    huggingface/skills

    Official

    Trains or fine-tunes language and vision models with TRL or Unsloth on Hugging Face Jobs cloud GPUs, then converts the results to GGUF.

    11k GitHub starsUsed in 3 repos~7.2k tokens
    AI & LLM EngineeringAuto-check passed
  • Official

    Routes a sentence-transformers training task to the right model type and required reference docs and example scripts, covering bi-encoders, rerankers, sparse and multi-vector models.

    11k GitHub starsUsed in 1 repo~2.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Dataset Evaluation

    awslabs/agent-plugins

    Official

    Validates dataset formatting and quality for SageMaker model fine-tuning (SFT, DPO, or RLVR).

    915 GitHub starsUsed in 2 repos~1.3k tokens
    AI & LLM EngineeringAuto-check passed
  • Train Rl

    OpenPipe/ART

    RL training reference for the ART framework. An agent skill from OpenPipe/ART.

    11k GitHub stars~2.4k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Qwopus27b Rl Training

    R6410418/Jackrong-llm-finetuning-guide

    Prepare, validate, launch-plan, monitor, resume, and stop configurable Qwopus 27B reinforcement-learning workflows for GRPO or GSPO.

    1.7k GitHub stars~830 tokensUpdated 3 mo ago
    AI & LLM EngineeringAuto-check passed

More from wshobson/agents

All 142 skills in this repo
  • Billing Automation

    wshobson/agents

    Covers building subscription billing: billing cycles, subscription states, invoice generation, proration, tax handling and dunning for failed payments.

    40k GitHub starsUsed in 14 repos~473 tokens
    Auto-check passed
  • Cuts cloud spend across AWS, Azure, GCP and OCI with cost tagging, rightsizing, commitment and spot pricing models, and architecture changes.

    40k GitHub starsUsed in 14 repos~1.7k tokens
    Auto-check passed
  • Profiles slow Python code with cProfile and memory profilers, then applies targeted fixes for CPU, memory, I/O and query bottlenecks.

    40k GitHub starsUsed in 13 repos~814 tokens
    Auto-check passed
  • Portfolio Risk Metrics

    wshobson/agents

    Covers portfolio risk measurement with VaR, CVaR, Sharpe, Sortino and drawdown, plus guidance on limits, stress tests and tail risk.

    40k GitHub starsUsed in 13 repos~502 tokens
    Auto-check passed
  • Writes unit tests for shell scripts with Bats: error-condition tests, fixtures and mocks, cross-shell checks, parallel runs, helper files and CI integration.

    40k GitHub starsUsed in 12 repos~1.3k tokens
    Auto-check passed
  • Plans memory headroom, works through out-of-memory failures and watches temperature and power during long ML training jobs on NVIDIA DGX Spark.

    40k GitHub starsUsed in 1 repo~2k tokens
    Auto-check passed

Questions about Dataset Curation

What does Dataset Curation do?

Prepare, format, and validate datasets for supervised fine-tuning and preference training. Dataset Curation is an agent skill from wshobson/agents. Prepare, format, and validate datasets for supervised fine-tuning and preference training.

When should I use Dataset Curation?

Dataset Curation fits situations like: converting raw data into training format; applying chat templates; configuring sequence packing; generating synthetic training data.

How do I install Dataset Curation in Claude Code?

Run `npx skills add wshobson/agents --skill dataset-curation -a claude-code`. Or copy the skill folder (plugins/llm-finetuning/skills/dataset-curation in wshobson/agents) into .claude/skills/dataset-curation in your project. Claude Code loads it when a task matches its description.

How do I install Dataset Curation in Codex?

Run `npx skills add wshobson/agents --skill dataset-curation -a codex`. Or copy the skill folder (plugins/llm-finetuning/skills/dataset-curation in wshobson/agents) into .agents/skills/dataset-curation in your project. Codex loads it when a task matches its description.

Can I use Dataset Curation in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add wshobson/agents --skill dataset-curation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/dataset-curation, .gemini/skills/dataset-curation, .github/skills/dataset-curation and .opencode/skills/dataset-curation in your project.

What does Dataset Curation need to run?

SKILL.md names no scripts, command-line tools or credentials: Dataset Curation is instructions for the agent only. Our summary lists: Python 3.

Does Dataset Curation access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Dataset Curation safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Dataset Curation use?

Dataset Curation is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Dataset Curation use?

About 2k tokens (SKILL.md is roughly 7.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 4.8k tokens, read only when the agent opens those files.

What are the alternatives to Dataset Curation?

Skills that share tags, products or a category with Dataset Curation: Peft Fine Tuning (Orchestra-Research/AI-Research-SKILLs, 13k stars), Hugging Face LLM Trainer (huggingface/skills, 11k stars), Sentence-Transformers Training Router (huggingface/skills, 11k stars) and Dataset Evaluation (awslabs/agent-plugins, 915 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Dataset Curation?

wshobson (a GitHub user) maintains it in wshobson/agents, which has 40,287 GitHub stars. The repository holds 142 skills in this directory. The repository was last updated on October 5, 2026.

Source: wshobson/agents on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.