Cocoindex
davila7/claude-code-templates
Comprehensive toolkit for developing with the CocoIndex library.
A skill your agent uses when assembling or curating the corpus a fine-tune trains on — turning raw examples into the JSONL shape a trainer expects (instruction, conversational, preference-pair or…
$ npx skills add ericrisco/rsc-harness --skill training-data -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install ericrisco/rsc-harness training-data --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/ericrisco/rsc-harness.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/training-data .claude/skills/training-data && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "training-data" agent skill from https://github.com/ericrisco/rsc-harness/tree/main/skills/training-data into .claude/skills/training-data/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "training-data", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/ericrisco/rsc-harness/tree/main/skills/training-dataType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add ericrisco/rsc-harness --skill training-data -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install ericrisco/rsc-harness training-data --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ericrisco/rsc-harness.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/training-data .agents/skills/training-data && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "training-data" agent skill from https://github.com/ericrisco/rsc-harness/tree/main/skills/training-data into .agents/skills/training-data/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "training-data", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add ericrisco/rsc-harness --skill training-data -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install ericrisco/rsc-harness training-data --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ericrisco/rsc-harness.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/training-data .cursor/skills/training-data && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "training-data" agent skill from https://github.com/ericrisco/rsc-harness/tree/main/skills/training-data into .cursor/skills/training-data/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "training-data", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/ericrisco/rsc-harness.git --path skills/training-data--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add ericrisco/rsc-harness --skill training-data -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install ericrisco/rsc-harness training-data --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ericrisco/rsc-harness.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/training-data .gemini/skills/training-data && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "training-data" agent skill from https://github.com/ericrisco/rsc-harness/tree/main/skills/training-data into .gemini/skills/training-data/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "training-data", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install ericrisco/rsc-harness training-dataInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add ericrisco/rsc-harness --skill training-data -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/ericrisco/rsc-harness.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/training-data .github/skills/training-data && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "training-data" agent skill from https://github.com/ericrisco/rsc-harness/tree/main/skills/training-data into .github/skills/training-data/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "training-data", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add ericrisco/rsc-harness --skill training-data -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install ericrisco/rsc-harness training-data --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ericrisco/rsc-harness.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/training-data .opencode/skills/training-data && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "training-data" agent skill from https://github.com/ericrisco/rsc-harness/tree/main/skills/training-data into .opencode/skills/training-data/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "training-data", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
training-dataA skill your agent uses when assembling or curating the corpus a fine-tune trains on — turning raw examples into the JSONL shape a trainer expects (instruction, conversational, preference-pair or…
Training Data is an agent skill from ericrisco/rsc-harness. Use when assembling or curating the corpus a fine-tune trains on — turning raw examples into the JSONL shape a trainer expects (instruction, conversational, preference-pair or binary-label), matching the data to the model's chat template, generating synthetic or distilled examples, deduplicating and decontaminating against eval sets, and quality-filtering. NOT cleaning tabular rows, nulls and dtypes (that is data-cleaning), NOT building a retrieval corpus of chunks and embeddings (that is embeddings-search), NOT…
Its SKILL.md is about 3.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files, including reference files (for example `evals/README.md`, `evals/cases.yaml` and `references/formats.md`).
It sits in AI & LLM Engineering, covering Data cleaning and Embeddings. The repository describes itself as: Your agent invents things because it has no memory, and can't touch your database because it has no arms. rsc is the meta-harness that gives it both, plus the trade to know the… The licence is MIT.
7 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit e3d5b33. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are python and jsonl).
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Training Data loads about 3.7k tokens when it runs, and up to ~6.6k if it reads all its reference files. Until then it costs about 152 tokens; SKILL.md has 1,427 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from ericrisco/rsc-harness at commit e3d5b33, republished under its MIT licence (© ericrisco). 1,427 words, ~3,660 tokens.
.claude/skills/training-data/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.You own the training corpus: the JSONL of chat turns, instruction triples, or preference
pairs that a trainer reads. The deliverable is a validated, deduplicated, decontaminated,
license-clean file in the exact shape the trainer expects, rendered through the target
model's chat template. You stop the moment that file loads cleanly and round-trips through
apply_chat_template. You do not choose LoRA rank or launch the run — that is
finetuning / unsloth.
Loud boundary. This is LLM training corpora — messages, instruction triples, preference pairs. It is not:
data-cleaning.embeddings-search.finetuning, unsloth, huggingface.Version reality (verified July 2026 — re-verify, these move monthly). TRL is on the v1.x
line (its dataset-formats doc was tagged v1.8.0 at author time); transformers is in the
4.57+ era (mixed text+vision data needs ≥4.57); datasets is 4.x (the Json() feature
type needs ≥4.7). Pin whatever you install — do not trust these numbers as current.
The trainer dictates the columns. Get this wrong and TRL either errors or, worse, trains on a
mangled string. Two axes: format (standard = plain strings vs conversational =
messages lists) and type (the task). One JSON object per line = JSONL.
| Trainer | Dataset type | Required keys |
|---|---|---|
SFTTrainer | language-modeling or prompt-completion | messages / text, or prompt+completion |
DPOTrainer, ORPOTrainer, CPOTrainer | preference (explicit prompt recommended) | prompt, chosen, rejected |
KTOTrainer, BCOTrainer | unpaired preference (binary label) | prompt, completion, label |
RewardTrainer | preference (implicit prompt) | chosen, rejected |
GRPOTrainer, RLOOTrainer, PPOTrainer | prompt-only | prompt |
Tiny JSONL of each (conversational values are lists of {role, content}; label is a JSON
boolean):
# Alpaca instruction (standard) — classic; NOT a native TRL type, see below
{"instruction": "Classify the sentiment.", "input": "The battery dies in an hour.", "output": "negative"}
# Conversational messages (SFT) — the default for chat fine-tunes
{"messages": [{"role": "system", "content": "You are a terse support agent."}, {"role": "user", "content": "My order never arrived."}, {"role": "assistant", "content": "Sorry about that — what is your order number?"}]}
# Preference pair (DPO) — chosen beats rejected for the same prompt
{"prompt": [{"role": "user", "content": "Define a hash map in one sentence."}], "chosen": [{"role": "assistant", "content": "A hash map stores key-value pairs and finds a value by hashing its key to a bucket, giving average O(1) lookup."}], "rejected": [{"role": "assistant", "content": "It's a fast dictionary thing."}]}
# KTO / unpaired preference — one completion + a good/bad boolean label
{"prompt": [{"role": "user", "content": "Define a hash map in one sentence."}], "completion": [{"role": "assistant", "content": "It's a fast dictionary thing."}], "label": false}Alpaca is not a native TRL type. {instruction, input, output} is the Stanford-Alpaca
convention, still common in Unsloth notebooks, but TRL trains on text/messages/prompt+
completion. You must either (a) map it into messages (instruction+input → user, output →
assistant), or (b) render it into a single text string via a prompt template — and append
the EOS token yourself, or the model never learns to stop (the #1 Unsloth-Alpaca bug). Prefer
(a) messages for chat models. Full field matrix, tool-calling (tools column) and vision
(images) extras, and every type→type conversion live in references/formats.md.
A chat template is a Jinja string stored in the tokenizer (in tokenizer_config.json under
chat_template, or a standalone chat_template.jinja in newer tokenizers). It maps a messages
list to the exact token string the model was trained on — special tokens (<|im_start|>,
[INST], <|start_header_id|>, …) and all. You render it, you never hand-type it:
from transformers import AutoTokenizer
tok = AutoTokenizer.from_pretrained("<target-model>") # the model you will fine-tune
# TRAINING: no trailing generation prompt — the assistant turn is already in the data
text = tok.apply_chat_template(msgs, tokenize=False, add_generation_prompt=False)
# INFERENCE: add_generation_prompt=True appends the assistant turn-start so the model continues
prompt = tok.apply_chat_template(msgs, tokenize=False, add_generation_prompt=True)Three ways this silently destroys a run — no error, just a worse model:
<|im_start|>user\n… strings yourself and getting one
token, one newline, or the BOS wrong. Train-time string ≠ inference-time string → the model
learns a distribution it is never served. Always render via apply_chat_template.chat_template = None. apply_chat_template then raises — you must choose and attach a
template (e.g. ChatML) and use that same one at inference forever after.Also decide loss masking: for chat SFT you usually train only on the assistant tokens
(completion_only_loss / assistant-only masking in SFTTrainer, or a completion-only collator),
so the model is not penalized for "predicting" the user's words. TRL applies the template for
you when the dataset is conversational — let it, rather than pre-flattening to text.
Not enough real examples? Generate them. Two workhorses: Self-Instruct (seed a few hand-written examples, prompt a strong model to produce more, filter) and Evol-Instruct (iteratively mutate prompts to be harder/deeper). Wrap them in a pipeline framework rather than ad-hoc loops (see §7).
Licensing trap — read before you distill. Generating your training data from another model's outputs ("distillation") is a terms-of-service question, not just a quality one. Some providers' terms restrict using their outputs to train competing models; some open-weight licenses carry naming/derivative obligations (e.g. Llama-derived data/models may inherit naming requirements). Never assert a model's license from memory — check the specific model card and provider ToS at author time (licenses change). If in doubt, distill from an openly-licensed-for-this-use model, and record the provenance per example.
references/synthesis-dedup-quality.md.LIMA (Less Is More for Alignment, arXiv 2305.11206) is the anchor: ~1,000 carefully curated examples produced a strong instruction-follower — for alignment/style SFT, quality and diversity dominate raw volume. (This is about teaching behavior/format, not injecting a lot of new knowledge — a broad knowledge shift still wants scale.) Cheap, high-leverage filters, applied before you spend GPU hours:
assistant turns in a row, missing final assistant turn for SFT).State the license class and point at the source; never freeze a license as bare fact.
Step/Task graphs (TextGeneration, UltraFeedback,
EvolInstruct), serializable to YAML/JSON, outputs a Distiset you push to the Hub. v1.x.datasets — load/map/filter/push_to_hub; the substrate everything else speaks.from datasets import load_dataset
from transformers import AutoTokenizer
ds = load_dataset("json", data_files="raw.jsonl", split="train")
tok = AutoTokenizer.from_pretrained("<target-model>")
# 1. VALIDATE shape + render every row through the template (catches template errors NOW,
# not after 3 GPU-hours). A base model with chat_template=None raises here — attach one.
def render(ex):
return {"text": tok.apply_chat_template(ex["messages"], tokenize=False,
add_generation_prompt=False)}
ds = ds.filter(lambda ex: isinstance(ex.get("messages"), list) and ex["messages"]
and ex["messages"][-1]["role"] == "assistant") # SFT: must end on assistant
ds = ds.map(render)
# 2. DEDUP (near-dup) and 3. DECONTAMINATE against your eval set — see references for MinHash
# + n-gram code; both are one filter pass each.
# 4. PUSH with a data card recording license + provenance.
ds.push_to_hub("me/support-sft", private=True)Deep code — MinHash/LSH dedup, n-gram decontamination, a distilabel Self-Instruct pipeline, and
the full conversion matrix — is in references/.
apply_chat_template.text formatting → the model never stops. Append it.label in KTO/unpaired data is a JSON boolean (true/false), not the strings "true"/"1".finetuning — consumes this corpus: chooses SFT vs DPO vs KTO,
LoRA/QLoRA, hyperparameters, runs trl/peft. You hand it the file; it trains.unsloth — one fast single-GPU training backend + GGUF export;
its notebooks expect exactly the Alpaca/messages shapes you produce here.huggingface — the Hub you push_to_hub the dataset to, model
cards, and hosted/routed inference of the result.data-cleaning — upstream when your raw source is dirty tabular
rows; it hands you clean rows, you turn rows into training examples.apply_chat_template without error.assistant turn; preference rows have distinct chosen/rejected; KTO label is a boolean.text).© ericrisco, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 4 other files (references) in skills/training-data of ericrisco/rsc-harness.
Open the folder on GitHubat commit e3d5b33
Training Data next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Training Data this skillericrisco/rsc-harness | 167 | — | ~3.7k | Automated safety check: Pass | MIT | |
| Cocoindexdavila7/claude-code-templates | 32k | 2 repos | ~6.4k | Automated safety check: Notes | MIT | |
| Tao Finetune Cosmos EmbedNVIDIA/skills | 3.5k | — | ~3.5k | Automated safety check: Notes | Apache-2.0 | |
| KtxKaelio/ktx | 1.6k | 1 repos | ~3.2k | Automated safety check: Pass | Apache-2.0 | |
| Dingo VerifyMigoXLab/dingo | 757 | — | ~833 | Automated safety check: Pass | Apache-2.0 | |
| Unimoljinzhezenggroup/computational-chemistry-agent-skills | 148 | 1 repos | ~1.5k | Automated safety check: Pass | LGPL-3.0-or-later |
davila7/claude-code-templates
Comprehensive toolkit for developing with the CocoIndex library.
NVIDIA/skills
Cosmos-Embed1 video-text embedding for text-to-video retrieval, video-to-video search, semantic deduplication, and fine-tuning.
Kaelio/ktx
Installs and configures ktx, the open-source context layer for data agents — runs ktx setup non-interactively with hidden CLI flags, configures database connections and embeddings, installs agent…
MigoXLab/dingo
A skill your agent uses when the user wants to fact-check an article or verify factual claims in a document.
jinzhezenggroup/computational-chemistry-agent-skills
A standardized CLI wrapper for Uni-Mol molecular ML workflows that handles representation extraction (embeddings), model training (regression/classification), and property prediction with built-in…
tokenbender/agent-guides
Audit supervised fine-tuning datasets against the behavior and task they are meant to teach.
ericrisco/rsc-harness
A skill your agent uses when designing or analyzing a controlled experiment — falsifiable hypothesis, sample size from an MDE, reading significance/CI/power, CUPED, or rescuing tests that won't go…
ericrisco/rsc-harness
A skill your agent uses when making a web UI conform to WCAG 2.2 Level AA — axe-core or Lighthouse a11y violations, keyboard operability, focus management, ARIA roles/names/live regions, contrast…
ericrisco/rsc-harness
A skill your agent uses when running or fixing paid acquisition on Google or Meta — campaign structure (Performance Max, Demand Gen, Search, Advantage+), platform-fit creative, budget/scaling rules…
ericrisco/rsc-harness
A skill your agent uses when measuring whether an LLM or agent system actually got better and gating merges on it: golden sets, fixing an inflated LLM-as-judge, scoring RAG (faithfulness, contextual…
ericrisco/rsc-harness
A skill your agent uses when a creative goal must become a finished media file: pick and order generative-media models per modality — AI voiceover, image-to-video clips, score — then glue them with…
ericrisco/rsc-harness
A skill your agent uses when instrumenting product or web analytics — GA4/PostHog SDK wiring, event taxonomy, funnels, double-counted events, consent gating, PII scrubbing.
Categories
A skill your agent uses when assembling or curating the corpus a fine-tune trains on — turning raw examples into the JSONL shape a trainer expects (instruction, conversational, preference-pair or…. Training Data is an agent skill from ericrisco/rsc-harness. Use when assembling or curating the corpus a fine-tune trains on — turning raw examples into the JSONL shape a trainer expects (instruction, conversational, preference-pair or binary-label), matching the data to the model's chat template, generating synthetic or distilled examples, deduplicating and decontaminating against eval sets, and quality-filtering.
Training Data fits situations like: curating the corpus a fine-tune trains on — turning raw examples into the JSONL shape a trainer expects (instruction; preference-pair; matching the data to the models chat template; generating synthetic.
Run `npx skills add ericrisco/rsc-harness --skill training-data -a claude-code`. Or copy the skill folder (skills/training-data in ericrisco/rsc-harness) into .claude/skills/training-data in your project. Claude Code loads it when a task matches its description.
Run `npx skills add ericrisco/rsc-harness --skill training-data -a codex`. Or copy the skill folder (skills/training-data in ericrisco/rsc-harness) into .agents/skills/training-data in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ericrisco/rsc-harness --skill training-data -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/training-data, .gemini/skills/training-data, .github/skills/training-data and .opencode/skills/training-data in your project.
SKILL.md names no scripts, command-line tools or credentials: Training Data is instructions for the agent only. Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Training Data is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 3.7k tokens (SKILL.md is roughly 15k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Training Data: Cocoindex (davila7/claude-code-templates, 32k stars), Tao Finetune Cosmos Embed (NVIDIA/skills, 3.5k stars), Ktx (Kaelio/ktx, 1.6k stars) and Dingo Verify (MigoXLab/dingo, 757 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
ericrisco (a GitHub user) maintains it in ericrisco/rsc-harness, which has 167 GitHub stars. The repository holds 227 skills in this directory. The repository was last updated on October 7, 2026.
Source: ericrisco/rsc-harness on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.