Agent skill

Training Data Curation

by sundial-org in sundial-org/skills

Guidelines for creating high-quality datasets for LLM post-training (SFT/DPO/RLHF).

No licenceAuto-check passedAI & LLM Engineering

Install Training Data Curation

skills CLI
$ npx skills add sundial-org/skills --skill training-data-curation -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install sundial-org/skills training-data-curation --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/sundial-org/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/training-data-curation .claude/skills/training-data-curation && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
training-data-curation
GitHub stars
152
Token cost
~1.4k tokens
SKILL.md length
525 words
Files
1
Skills in repo
13
Repo updated
First seen
Licence
None found

At a glance

Guidelines for creating high-quality datasets for LLM post-training (SFT/DPO/RLHF).

  • Works in 8 steps: Llama 2 Paper — Touvron et al. (2023).… → TRL Library — HuggingFace trainer… → FineWeb Paper — Penedo et al. (2024).… → …
  • Preparing data for fine-tuning
  • SKILL.md covers Data Quality Principles, Format Requirements, Quality Checklist and Common Quality Issues, plus 4 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Training Data Curation is an agent skill from sundial-org/skills. Guidelines for creating high-quality datasets for LLM post-training (SFT/DPO/RLHF). Use when preparing data for fine-tuning, evaluating data quality, or designing data collection strategies.

Its SKILL.md is about 1.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering Fine-tuning and Data cleaning. The repository describes itself as: Claude Code Skills by Sundial.

When your agent uses it

  • Preparing data for fine-tuning
  • Evaluating data quality
  • Designing data collection strategies

Example prompts

  • “Use the training-data-curation skill to guideline for creating high-quality datasets for LLM post-training (SFT/DPO/RLHF)”
  • “/training-data-curation”

Workflow steps

8 steps, taken from the first numbered list in SKILL.md.

  1. Llama 2 Paper — Touvron et al. (2023). SFT/RLHF data quality practices, 27K SFT examples, >70% annotator agreement threshold
  2. TRL Library — HuggingFace trainer implementations for SFT, DPO, KTO, ORPO
  3. FineWeb Paper — Penedo et al. (2024). Large-scale filtering: MinHash dedup, language detection, quality classifiers
  4. Data-Juicer — Alibaba's quality filtering toolkit with repetition filters, n-gram analysis
  5. Tinker API — Training API using messages format for SFT, DPO/RLHF support
  6. Data Provenance Initiative — Longpre et al. (2023). Dataset licensing and attribution audit
  7. KTO Paper — Ethayarajh et al. (2024). Binary preference learning without pairs
  8. C4/T5 Paper — Raffel et al. (2020). Foundational filtering: terminal punctuation, min sentences, alpha ratio, boilerplate removal

What it can do on your machine

Read from SKILL.md and the folder at commit 5a3bf9f. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • arxiv.org
    • huggingface.co
    • github.com
    • tinker-docs.thinkingmachines.ai

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Training Data Curation loads about 1.4k tokens when it runs. Until then it costs about 53 tokens; SKILL.md has 525 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~53
When it runs · the whole SKILL.md, loaded when a task matches
~1.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

Without a licence we can't republish the file, so here is its outline and opening line. It has 525 words (~1,366 tokens).

“Best practices for gathering and preparing training data for LLM fine-tuning.”

— opening of SKILL.md by sundial-org
name
training-data-curation

Read the full SKILL.md on GitHub

Files

Just SKILL.md in skills/training-data-curation of sundial-org/skills.

Open the folder on GitHubat commit 5a3bf9f

Compare with similar skills

Training Data Curation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Training Data Curation compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Training Data Curation this skillsundial-org/skills152—~1.4kAutomated safety check: PassNone
Audit Sft Data Qualitytokenbender/agent-guides367—~2.7kAutomated safety check: PassApache-2.0
Tao Finetune Cosmos EmbedNVIDIA/skills3.5k—~3.5kAutomated safety check: NotesApache-2.0
Peft Fine TuningOrchestra-Research/AI-Research-SKILLs13k9 repos~3.1kAutomated safety check: PassMIT
Hugging Face LLM Trainerhuggingface/skills11k3 repos~7.2kAutomated safety check: PassApache-2.0
Sentence-Transformers Training Routerhuggingface/skills11k1 repos~2.6kAutomated safety check: PassApache-2.0

Similar skills

  • Audit Sft Data Quality

    tokenbender/agent-guides

    Audit supervised fine-tuning datasets against the behavior and task they are meant to teach.

    367 GitHub stars~2.7k tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check passed
  • Official

    Cosmos-Embed1 video-text embedding for text-to-video retrieval, video-to-video search, semantic deduplication, and fine-tuning.

    3.5k GitHub stars~3.5k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check: notes
  • Peft Fine Tuning

    Orchestra-Research/AI-Research-SKILLs

    Parameter-efficient fine-tuning for LLMs using LoRA, QLoRA, and 25+ methods.

    13k GitHub starsUsed in 9 repos~3.1k tokens
    AI & LLM EngineeringAuto-check passed
  • Hugging Face LLM Trainer

    huggingface/skills

    Official

    Trains or fine-tunes language and vision models with TRL or Unsloth on Hugging Face Jobs cloud GPUs, then converts the results to GGUF.

    11k GitHub starsUsed in 3 repos~7.2k tokens
    AI & LLM EngineeringAuto-check passed
  • Official

    Routes a sentence-transformers training task to the right model type and required reference docs and example scripts, covering bi-encoders, rerankers, sparse and multi-vector models.

    11k GitHub starsUsed in 1 repo~2.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Dataset Evaluation

    awslabs/agent-plugins

    Official

    Validates dataset formatting and quality for SageMaker model fine-tuning (SFT, DPO, or RLVR).

    915 GitHub starsUsed in 2 repos~1.3k tokens
    AI & LLM EngineeringAuto-check passed

More from sundial-org/skills

All 13 skills in this repo
  • AI Co Scientist

    sundial-org/skills

    Transform Claude Code into an AI Scientist that orchestrates research workflows using tree-based hypothesis exploration.

    152 GitHub stars~2.5k tokensUpdated 2 mo ago
    Auto-check passed
  • Skill

    sundial-org/skills

    Find, install, create, improve, and publish AI agent skills through the Sundial ecosystem.

    152 GitHub stars~1.2k tokensUpdated 2 mo ago
    Auto-check passed
  • Skill To Card

    sundial-org/skills

    End-to-end workflow that creates a skill from a description and attached files, publishes it to Sundial as a private skill, generates a trading card (front + back with QR code), and sends it to a…

    152 GitHub stars~1.5k tokensUpdated 2 mo ago
    Auto-check passed
  • Commit Splitter

    sundial-org/skills

    Split large sets of uncommitted changes into logical, well-organized commits.

    152 GitHub stars~860 tokensUpdated 2 mo ago
    Auto-check passed
  • Tinker Training Cost

    sundial-org/skills

    Calculate training costs for Tinker fine-tuning jobs. An agent skill from sundial-org/skills.

    152 GitHub stars~1.2k tokensUpdated 2 mo ago
    Auto-check passed
  • Codex

    sundial-org/skills

    Run OpenAI's Codex CLI agent in non-interactive mode using codex exec.

    152 GitHub stars~1.5k tokensUpdated 2 mo ago
    Auto-check passed

Questions about Training Data Curation

What does Training Data Curation do?

Guidelines for creating high-quality datasets for LLM post-training (SFT/DPO/RLHF). Training Data Curation is an agent skill from sundial-org/skills. Guidelines for creating high-quality datasets for LLM post-training (SFT/DPO/RLHF).

When should I use Training Data Curation?

Training Data Curation fits situations like: preparing data for fine-tuning; evaluating data quality; designing data collection strategies.

How do I install Training Data Curation in Claude Code?

Run `npx skills add sundial-org/skills --skill training-data-curation -a claude-code`. Or copy the skill folder (skills/training-data-curation in sundial-org/skills) into .claude/skills/training-data-curation in your project. Claude Code loads it when a task matches its description.

How do I install Training Data Curation in Codex?

Run `npx skills add sundial-org/skills --skill training-data-curation -a codex`. Or copy the skill folder (skills/training-data-curation in sundial-org/skills) into .agents/skills/training-data-curation in your project. Codex loads it when a task matches its description.

Can I use Training Data Curation in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add sundial-org/skills --skill training-data-curation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/training-data-curation, .gemini/skills/training-data-curation, .github/skills/training-data-curation and .opencode/skills/training-data-curation in your project.

What does Training Data Curation need to run?

SKILL.md names no scripts, command-line tools or credentials: Training Data Curation is instructions for the agent only.

Does Training Data Curation access the network?

SKILL.md names 4 domains. As links in the text: arxiv.org, huggingface.co, github.com and tinker-docs.thinkingmachines.ai. This is read from the text; nothing was executed.

Is Training Data Curation safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Training Data Curation use?

No licence was found for Training Data Curation or its repository. Without one, default copyright applies: ask the author before reusing or redistributing it.

How many tokens does Training Data Curation use?

About 1.4k tokens (SKILL.md is roughly 5.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Training Data Curation?

Skills that share tags, products or a category with Training Data Curation: Audit Sft Data Quality (tokenbender/agent-guides, 367 stars), Tao Finetune Cosmos Embed (NVIDIA/skills, 3.5k stars), Peft Fine Tuning (Orchestra-Research/AI-Research-SKILLs, 13k stars) and Hugging Face LLM Trainer (huggingface/skills, 11k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Training Data Curation?

sundial-org (a GitHub organization) maintains it in sundial-org/skills, which has 152 GitHub stars. The repository holds 13 skills in this directory. The repository was last updated on July 15, 2026.

Source: sundial-org/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.