Agent skill

Training Data

by ericrisco in ericrisco/rsc-harness

A skill your agent uses when assembling or curating the corpus a fine-tune trains on — turning raw examples into the JSONL shape a trainer expects (instruction, conversational, preference-pair or…

MITAuto-check passedAI & LLM Engineering

Install Training Data

skills CLI
$ npx skills add ericrisco/rsc-harness --skill training-data -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install ericrisco/rsc-harness training-data --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/ericrisco/rsc-harness.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/training-data .claude/skills/training-data && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
training-data
GitHub stars
167
Token cost
~3.7k tokens
SKILL.md length
1,427 words
Files
5 (incl. references)
Skills in repo
227
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when assembling or curating the corpus a fine-tune trains on — turning raw examples into the JSONL shape a trainer expects (instruction, conversational, preference-pair or…

  • Works in 7 steps: The format the trainer expects (pick by… → Chat templates — the silent run-wrecker → Synthetic data & distillation → …
  • Curating the corpus a fine-tune trains on — turning raw examples into the JSONL shape a trainer expects (instruction
  • SKILL.md covers 1. The format the trainer…, 2. Chat templates — the silent…, 3. Synthetic data & distillation and 4. Dedup + decontamination…, plus 7 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Training Data is an agent skill from ericrisco/rsc-harness. Use when assembling or curating the corpus a fine-tune trains on — turning raw examples into the JSONL shape a trainer expects (instruction, conversational, preference-pair or binary-label), matching the data to the model's chat template, generating synthetic or distilled examples, deduplicating and decontaminating against eval sets, and quality-filtering. NOT cleaning tabular rows, nulls and dtypes (that is data-cleaning), NOT building a retrieval corpus of chunks and embeddings (that is embeddings-search), NOT…

Its SKILL.md is about 3.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files, including reference files (for example `evals/README.md`, `evals/cases.yaml` and `references/formats.md`).

It sits in AI & LLM Engineering, covering Data cleaning and Embeddings. The repository describes itself as: Your agent invents things because it has no memory, and can't touch your database because it has no arms. rsc is the meta-harness that gives it both, plus the trade to know the… The licence is MIT.

When your agent uses it

  • Curating the corpus a fine-tune trains on — turning raw examples into the JSONL shape a trainer expects (instruction
  • Preference-pair
  • Matching the data to the models chat template
  • Generating synthetic

Example prompts

  • “/training-data”

Requirements

  • Python 3

Workflow steps

7 steps, taken from the step headings in SKILL.md.

  1. The format the trainer expects (pick by trainer, not by taste)
  2. Chat templates — the silent run-wrecker
  3. Synthetic data & distillation
  4. Dedup + decontamination (skip these and your numbers lie)
  5. Quality filtering — a few thousand clean beats a noisy dump
  6. Licensing — two separate questions
  7. Tooling

What it can do on your machine

Read from SKILL.md and the folder at commit e3d5b33. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python and jsonl).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Training Data loads about 3.7k tokens when it runs, and up to ~6.6k if it reads all its reference files. Until then it costs about 152 tokens; SKILL.md has 1,427 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~152
When it runs · the whole SKILL.md, loaded when a task matches
~3.7k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~6.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from ericrisco/rsc-harness at commit e3d5b33, republished under its MIT licence (© ericrisco). 1,427 words, ~3,660 tokens.

Download SKILL.mdSave it as .claude/skills/training-data/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
training-data
description
Use when assembling or curating the corpus a fine-tune trains on — turning raw examples into the JSONL shape a trainer expects (instruction, conversational, preference-pair or binary-label), matching the data to the model's chat template, generating synthetic or distilled examples, deduplicating and decontaminating against eval sets, and quality-filtering. NOT cleaning tabular rows, nulls and dtypes (that is `data-cleaning`), NOT building a retrieval corpus of chunks and embeddings (that is `embeddings-search`), NOT running the trainer or picking hyperparameters (that is `finetuning`).
tags
training-data, fine-tuning-dataset, jsonl, chat-template, preference-data, dpo, kto, synthetic-data, decontamination, dataset-curation
recommends
finetuning, unsloth, huggingface, data-cleaning
origin
risco

training-data — the corpus a fine-tune actually eats

You own the training corpus: the JSONL of chat turns, instruction triples, or preference pairs that a trainer reads. The deliverable is a validated, deduplicated, decontaminated, license-clean file in the exact shape the trainer expects, rendered through the target model's chat template. You stop the moment that file loads cleanly and round-trips through apply_chat_template. You do not choose LoRA rank or launch the run — that is finetuning / unsloth.

Loud boundary. This is LLM training corpora — messages, instruction triples, preference pairs. It is not:

  • Tabular row cleaning — nulls, dtypes, dedupe of CSV rows, category normalization → data-cleaning.
  • A retrieval corpus — chunking documents and embedding them for search → embeddings-search.
  • Actually training or serving — hyperparameters, the run, export → finetuning, unsloth, huggingface.

Version reality (verified July 2026 — re-verify, these move monthly). TRL is on the v1.x line (its dataset-formats doc was tagged v1.8.0 at author time); transformers is in the 4.57+ era (mixed text+vision data needs ≥4.57); datasets is 4.x (the Json() feature type needs ≥4.7). Pin whatever you install — do not trust these numbers as current.

1. The format the trainer expects (pick by trainer, not by taste)

The trainer dictates the columns. Get this wrong and TRL either errors or, worse, trains on a mangled string. Two axes: format (standard = plain strings vs conversational = messages lists) and type (the task). One JSON object per line = JSONL.

TrainerDataset typeRequired keys
SFTTrainerlanguage-modeling or prompt-completionmessages / text, or prompt+completion
DPOTrainer, ORPOTrainer, CPOTrainerpreference (explicit prompt recommended)prompt, chosen, rejected
KTOTrainer, BCOTrainerunpaired preference (binary label)prompt, completion, label
RewardTrainerpreference (implicit prompt)chosen, rejected
GRPOTrainer, RLOOTrainer, PPOTrainerprompt-onlyprompt

Tiny JSONL of each (conversational values are lists of {role, content}; label is a JSON boolean):

jsonl
# Alpaca instruction (standard) — classic; NOT a native TRL type, see below
{"instruction": "Classify the sentiment.", "input": "The battery dies in an hour.", "output": "negative"}

# Conversational messages (SFT) — the default for chat fine-tunes
{"messages": [{"role": "system", "content": "You are a terse support agent."}, {"role": "user", "content": "My order never arrived."}, {"role": "assistant", "content": "Sorry about that — what is your order number?"}]}

# Preference pair (DPO) — chosen beats rejected for the same prompt
{"prompt": [{"role": "user", "content": "Define a hash map in one sentence."}], "chosen": [{"role": "assistant", "content": "A hash map stores key-value pairs and finds a value by hashing its key to a bucket, giving average O(1) lookup."}], "rejected": [{"role": "assistant", "content": "It's a fast dictionary thing."}]}

# KTO / unpaired preference — one completion + a good/bad boolean label
{"prompt": [{"role": "user", "content": "Define a hash map in one sentence."}], "completion": [{"role": "assistant", "content": "It's a fast dictionary thing."}], "label": false}

Alpaca is not a native TRL type. {instruction, input, output} is the Stanford-Alpaca convention, still common in Unsloth notebooks, but TRL trains on text/messages/prompt+ completion. You must either (a) map it into messages (instruction+input → user, output → assistant), or (b) render it into a single text string via a prompt template — and append the EOS token yourself, or the model never learns to stop (the #1 Unsloth-Alpaca bug). Prefer (a) messages for chat models. Full field matrix, tool-calling (tools column) and vision (images) extras, and every type→type conversion live in references/formats.md.

2. Chat templates — the silent run-wrecker

A chat template is a Jinja string stored in the tokenizer (in tokenizer_config.json under chat_template, or a standalone chat_template.jinja in newer tokenizers). It maps a messages list to the exact token string the model was trained on — special tokens (<|im_start|>, [INST], <|start_header_id|>, …) and all. You render it, you never hand-type it:

python
from transformers import AutoTokenizer
tok = AutoTokenizer.from_pretrained("<target-model>")   # the model you will fine-tune

# TRAINING: no trailing generation prompt — the assistant turn is already in the data
text = tok.apply_chat_template(msgs, tokenize=False, add_generation_prompt=False)

# INFERENCE: add_generation_prompt=True appends the assistant turn-start so the model continues
prompt = tok.apply_chat_template(msgs, tokenize=False, add_generation_prompt=True)

Three ways this silently destroys a run — no error, just a worse model:

  • Hand-formatting the tokens. Writing <|im_start|>user\n… strings yourself and getting one token, one newline, or the BOS wrong. Train-time string ≠ inference-time string → the model learns a distribution it is never served. Always render via apply_chat_template.
  • Using the wrong model's template. The template must be the one of the model you are fine-tuning. Copy Llama's template onto a Qwen fine-tune and every example is subtly malformed.
  • A base model with no template at all. Base (non-instruct) checkpoints often ship chat_template = None. apply_chat_template then raises — you must choose and attach a template (e.g. ChatML) and use that same one at inference forever after.

Also decide loss masking: for chat SFT you usually train only on the assistant tokens (completion_only_loss / assistant-only masking in SFTTrainer, or a completion-only collator), so the model is not penalized for "predicting" the user's words. TRL applies the template for you when the dataset is conversational — let it, rather than pre-flattening to text.

3. Synthetic data & distillation

Not enough real examples? Generate them. Two workhorses: Self-Instruct (seed a few hand-written examples, prompt a strong model to produce more, filter) and Evol-Instruct (iteratively mutate prompts to be harder/deeper). Wrap them in a pipeline framework rather than ad-hoc loops (see §7).

Licensing trap — read before you distill. Generating your training data from another model's outputs ("distillation") is a terms-of-service question, not just a quality one. Some providers' terms restrict using their outputs to train competing models; some open-weight licenses carry naming/derivative obligations (e.g. Llama-derived data/models may inherit naming requirements). Never assert a model's license from memory — check the specific model card and provider ToS at author time (licenses change). If in doubt, distill from an openly-licensed-for-this-use model, and record the provenance per example.

4. Dedup + decontamination (skip these and your numbers lie)

  • Dedup. Exact dedup is trivial (hash the text). Real corpora need near-dup removal: MinHash + LSH (Jaccard similarity over shingles) catches templated/boilerplate repeats that inflate a few patterns. Dupes waste compute and bias the model toward whatever is over-represented.
  • Decontamination — the one people forget. Remove any training example that overlaps your eval / benchmark test sets (n-gram overlap, e.g. long-n-gram match against MMLU, GSM8K, your own held-out set). If test items leak into training, your eval score is inflated and meaningless — you measured memorization, not capability. Decontaminate against every metric you will report, including your private eval. Code for both in references/synthesis-dedup-quality.md.
Show full SKILL.md (588 more words)Show less

5. Quality filtering — a few thousand clean beats a noisy dump

LIMA (Less Is More for Alignment, arXiv 2305.11206) is the anchor: ~1,000 carefully curated examples produced a strong instruction-follower — for alignment/style SFT, quality and diversity dominate raw volume. (This is about teaching behavior/format, not injecting a lot of new knowledge — a broad knowledge shift still wants scale.) Cheap, high-leverage filters, applied before you spend GPU hours:

  • Length/format: drop empty or truncated turns, runaway-length outliers, malformed JSON, wrong-role sequences (two assistant turns in a row, missing final assistant turn for SFT).
  • Dedup + decontam from §4.
  • Diversity: cluster/embed and prune near-identical intents so the set is not 80% one task.
  • Model/heuristic scoring: rate helpfulness/correctness (a reward model or an LLM judge) and keep the top slice — but audit the judge, LLM-as-judge has its own biases.

6. Licensing — two separate questions

  1. The dataset's own license — what you release the JSONL under, and whether you can release it (aggregating others' data does not launder their licenses).
  2. Source-usage restrictions — the terms on where each example came from: scraped-site ToS, the license of any base dataset you built on, and the model-output ToS from §3. These bind even if you never publish. Keep a provenance column so an audit can trace every row.

State the license class and point at the source; never freeze a license as bare fact.

7. Tooling

  • distilabel (Argilla, now under Hugging Face) — the go-to synthetic-data / AI-feedback pipeline framework: composable Step/Task graphs (TextGeneration, UltraFeedback, EvolInstruct), serializable to YAML/JSON, outputs a Distiset you push to the Hub. v1.x.
  • Argilla — human-in-the-loop annotation/review UI to label and vet examples.
  • HF datasets — load/map/filter/push_to_hub; the substrate everything else speaks.
  • Lilac — dataset exploration/clustering for quality triage. [verify — the open-source repo was archived (read-only) around July 2025 after the Databricks acquisition]; treat as unmaintained OSS and confirm before depending on it.

Worked lifecycle (build → validate → dedup → decontaminate → format → push)

python
from datasets import load_dataset
from transformers import AutoTokenizer

ds  = load_dataset("json", data_files="raw.jsonl", split="train")
tok = AutoTokenizer.from_pretrained("<target-model>")

# 1. VALIDATE shape + render every row through the template (catches template errors NOW,
#    not after 3 GPU-hours). A base model with chat_template=None raises here — attach one.
def render(ex):
    return {"text": tok.apply_chat_template(ex["messages"], tokenize=False,
                                            add_generation_prompt=False)}
ds = ds.filter(lambda ex: isinstance(ex.get("messages"), list) and ex["messages"]
               and ex["messages"][-1]["role"] == "assistant")   # SFT: must end on assistant
ds = ds.map(render)

# 2. DEDUP (near-dup) and 3. DECONTAMINATE against your eval set — see references for MinHash
#    + n-gram code; both are one filter pass each.

# 4. PUSH with a data card recording license + provenance.
ds.push_to_hub("me/support-sft", private=True)

Deep code — MinHash/LSH dedup, n-gram decontamination, a distilabel Self-Instruct pipeline, and the full conversion matrix — is in references/.

Guardrails / gotchas

  • Wrong format for the trainer = hard error or silent garbage. Match the table in §1 to your trainer before generating a single row.
  • Hand-typed chat tokens train a distribution you never serve. Render via apply_chat_template.
  • No EOS in Alpaca-text formatting → the model never stops. Append it.
  • Skipped decontamination → inflated eval; you measured leakage. Non-negotiable.
  • Distilling model outputs can violate ToS. Provenance + license check first.
  • Quantity worship — a noisy 500k dump loses to a curated few-thousand for alignment (LIMA).
  • label in KTO/unpaired data is a JSON boolean (true/false), not the strings "true"/"1".
  • finetuning — consumes this corpus: chooses SFT vs DPO vs KTO, LoRA/QLoRA, hyperparameters, runs trl/peft. You hand it the file; it trains.
  • unsloth — one fast single-GPU training backend + GGUF export; its notebooks expect exactly the Alpaca/messages shapes you produce here.
  • huggingface — the Hub you push_to_hub the dataset to, model cards, and hosted/routed inference of the result.
  • data-cleaning — upstream when your raw source is dirty tabular rows; it hands you clean rows, you turn rows into training examples.

Checklist

  • Format matches the target trainer (§1 table); one JSON object per line.
  • Every row renders through the target model's apply_chat_template without error.
  • SFT rows end on an assistant turn; preference rows have distinct chosen/rejected; KTO label is a boolean.
  • Loss masking / EOS handling decided (assistant-only loss; EOS appended if flattening to text).
  • Near-duplicates removed (MinHash/LSH); exact dupes gone.
  • Decontaminated against every eval/benchmark you will report.
  • Quality-filtered (length/format/diversity/score) — curated over bulk.
  • Dataset license set and source-usage/model-output ToS checked; provenance recorded.
  • Data card written; pushed (private first).

© ericrisco, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files (references) in skills/training-data of ericrisco/rsc-harness.

  • SKILL.md
  • evals/README.md
  • evals/cases.yaml
  • references/formats.md
  • references/synthesis-dedup-quality.md

Open the folder on GitHubat commit e3d5b33

Compare with similar skills

Training Data next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Training Data compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Training Data this skillericrisco/rsc-harness167—~3.7kAutomated safety check: PassMIT
Cocoindexdavila7/claude-code-templates32k2 repos~6.4kAutomated safety check: NotesMIT
Tao Finetune Cosmos EmbedNVIDIA/skills3.5k—~3.5kAutomated safety check: NotesApache-2.0
KtxKaelio/ktx1.6k1 repos~3.2kAutomated safety check: PassApache-2.0
Dingo VerifyMigoXLab/dingo757—~833Automated safety check: PassApache-2.0
Unimoljinzhezenggroup/computational-chemistry-agent-skills1481 repos~1.5kAutomated safety check: PassLGPL-3.0-or-later

Similar skills

  • Cocoindex

    davila7/claude-code-templates

    Comprehensive toolkit for developing with the CocoIndex library.

    32k GitHub starsUsed in 2 repos~6.4k tokens
    AI & LLM EngineeringAuto-check: notes
  • Official

    Cosmos-Embed1 video-text embedding for text-to-video retrieval, video-to-video search, semantic deduplication, and fine-tuning.

    3.5k GitHub stars~3.5k tokensUpdated today
    AI & LLM EngineeringAuto-check: notes
  • Ktx

    Kaelio/ktx

    Installs and configures ktx, the open-source context layer for data agents — runs ktx setup non-interactively with hidden CLI flags, configures database connections and embeddings, installs agent…

    1.6k GitHub starsUsed in 1 repo~3.2k tokens
    AI & LLM EngineeringAuto-check passed
  • Dingo Verify

    MigoXLab/dingo

    A skill your agent uses when the user wants to fact-check an article or verify factual claims in a document.

    757 GitHub stars~833 tokensUpdated 10 days ago
    AI & LLM EngineeringAuto-check passed
  • Unimol

    jinzhezenggroup/computational-chemistry-agent-skills

    A standardized CLI wrapper for Uni-Mol molecular ML workflows that handles representation extraction (embeddings), model training (regression/classification), and property prediction with built-in…

    148 GitHub starsUsed in 1 repo~1.5k tokens
    AI & LLM EngineeringAuto-check passed
  • Audit Sft Data Quality

    tokenbender/agent-guides

    Audit supervised fine-tuning datasets against the behavior and task they are meant to teach.

    367 GitHub stars~2.7k tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check passed

More from ericrisco/rsc-harness

All 227 skills in this repo
  • Ab Testing

    ericrisco/rsc-harness

    A skill your agent uses when designing or analyzing a controlled experiment — falsifiable hypothesis, sample size from an MDE, reading significance/CI/power, CUPED, or rescuing tests that won't go…

    167 GitHub stars~2.4k tokensUpdated today
    Auto-check passed
  • Accessibility

    ericrisco/rsc-harness

    A skill your agent uses when making a web UI conform to WCAG 2.2 Level AA — axe-core or Lighthouse a11y violations, keyboard operability, focus management, ARIA roles/names/live regions, contrast…

    167 GitHub stars~3.4k tokensUpdated today
    Auto-check passed
  • Ads

    ericrisco/rsc-harness

    A skill your agent uses when running or fixing paid acquisition on Google or Meta — campaign structure (Performance Max, Demand Gen, Search, Advantage+), platform-fit creative, budget/scaling rules…

    167 GitHub stars~2.2k tokensUpdated today
    Auto-check passed
  • Agent Eval

    ericrisco/rsc-harness

    A skill your agent uses when measuring whether an LLM or agent system actually got better and gating merges on it: golden sets, fixing an inflated LLM-as-judge, scoring RAG (faithfulness, contextual…

    167 GitHub stars~3.2k tokensUpdated today
    Auto-check passed
  • AI Media

    ericrisco/rsc-harness

    A skill your agent uses when a creative goal must become a finished media file: pick and order generative-media models per modality — AI voiceover, image-to-video clips, score — then glue them with…

    167 GitHub stars~3.3k tokensUpdated today
    Auto-check passed
  • Analytics

    ericrisco/rsc-harness

    A skill your agent uses when instrumenting product or web analytics — GA4/PostHog SDK wiring, event taxonomy, funnels, double-counted events, consent gating, PII scrubbing.

    167 GitHub stars~2.8k tokensUpdated today
    Auto-check passed

Questions about Training Data

What does Training Data do?

A skill your agent uses when assembling or curating the corpus a fine-tune trains on — turning raw examples into the JSONL shape a trainer expects (instruction, conversational, preference-pair or…. Training Data is an agent skill from ericrisco/rsc-harness. Use when assembling or curating the corpus a fine-tune trains on — turning raw examples into the JSONL shape a trainer expects (instruction, conversational, preference-pair or binary-label), matching the data to the model's chat template, generating synthetic or distilled examples, deduplicating and decontaminating against eval sets, and quality-filtering.

When should I use Training Data?

Training Data fits situations like: curating the corpus a fine-tune trains on — turning raw examples into the JSONL shape a trainer expects (instruction; preference-pair; matching the data to the models chat template; generating synthetic.

How do I install Training Data in Claude Code?

Run `npx skills add ericrisco/rsc-harness --skill training-data -a claude-code`. Or copy the skill folder (skills/training-data in ericrisco/rsc-harness) into .claude/skills/training-data in your project. Claude Code loads it when a task matches its description.

How do I install Training Data in Codex?

Run `npx skills add ericrisco/rsc-harness --skill training-data -a codex`. Or copy the skill folder (skills/training-data in ericrisco/rsc-harness) into .agents/skills/training-data in your project. Codex loads it when a task matches its description.

Can I use Training Data in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ericrisco/rsc-harness --skill training-data -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/training-data, .gemini/skills/training-data, .github/skills/training-data and .opencode/skills/training-data in your project.

What does Training Data need to run?

SKILL.md names no scripts, command-line tools or credentials: Training Data is instructions for the agent only. Our summary lists: Python 3.

Does Training Data access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Training Data safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Training Data use?

Training Data is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Training Data use?

About 3.7k tokens (SKILL.md is roughly 15k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3k tokens, read only when the agent opens those files.

What are the alternatives to Training Data?

Skills that share tags, products or a category with Training Data: Cocoindex (davila7/claude-code-templates, 32k stars), Tao Finetune Cosmos Embed (NVIDIA/skills, 3.5k stars), Ktx (Kaelio/ktx, 1.6k stars) and Dingo Verify (MigoXLab/dingo, 757 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Training Data?

ericrisco (a GitHub user) maintains it in ericrisco/rsc-harness, which has 167 GitHub stars. The repository holds 227 skills in this directory. The repository was last updated on October 7, 2026.

Source: ericrisco/rsc-harness on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.