Agent skill

Golden Dataset

by yonatangross in yonatangross/orchestkit

Golden dataset lifecycle patterns for curation, versioning, quality validation, and CI integration.

MITAuto-check passedDevOps & Cloud

Install Golden Dataset

skills CLI
$ npx skills add yonatangross/orchestkit --skill golden-dataset -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install yonatangross/orchestkit golden-dataset --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/yonatangross/orchestkit.git skills-src && mkdir -p .claude/skills && cp -r skills-src/src/skills/golden-dataset .claude/skills/golden-dataset && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
golden-dataset
GitHub stars
292
Token cost
~2k tokens
SKILL.md length
676 words
Files
17 (incl. scripts, references)
Skills in repo
108
Repo updated
First seen
Licence
MIT

At a glance

Golden dataset lifecycle patterns for curation, versioning, quality validation, and CI integration.

  • Works in 8 steps: Using placeholder URLs instead of… → Skipping embedding regeneration after… → Not validating referential integrity… → …
  • Building evaluation datasets
  • SKILL.md covers Quick Reference, Curation, Management and Validation, plus 9 more sections
  • Runs Python scripts from its folder

What it does

Golden Dataset is an agent skill from yonatangross/orchestkit. Golden dataset lifecycle patterns for curation, versioning, quality validation, and CI integration. Use when building evaluation datasets, managing dataset versions, validating quality scores, or integrating golden tests into pipelines.

Its SKILL.md is about 2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 19 other files, including scripts and reference files (for example `metadata.json`, `references/ork-delta.md` and `references/quality-metrics.md`). Compatibility notes: Claude Code 2.1.277+.

It sits in DevOps & Cloud. The repository describes itself as: The Complete AI Development Toolkit for Claude Code. 106 skills, 36 agents, 171 hooks. Install ork for stable (v9.x), or ork-alpha for the v10 line, which ships daily. The licence is MIT.

When your agent uses it

  • Building evaluation datasets
  • Managing dataset versions
  • Validating quality scores
  • Integrating golden tests into pipelines

Example prompts

  • “/golden-dataset”

Requirements

  • Python 3
  • Compatibility (from SKILL.md): Claude Code 2.1.277+.
  • Pre-approved tools (allowed-tools): Read, Glob, Grep, WebFetch, WebSearch

Workflow steps

8 steps, taken from the first numbered list in SKILL.md.

  1. Using placeholder URLs instead of canonical source URLs
  2. Skipping embedding regeneration after restore
  3. Not validating referential integrity between documents and queries
  4. Over-indexing on articles (neglecting tutorials, research papers)
  5. Missing difficulty distribution balance in test queries
  6. Not running verification after backup/restore operations
  7. Testing restore procedures in production instead of staging
  8. Committing SQL dumps instead of JSON (not version-control friendly)

What it can do on your machine

Read from SKILL.md and the folder at commit e4ff8d9. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Glob
    • Grep
    • WebFetch
    • WebSearch

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 2 files in scripts/ (Python), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • json-schema.org
    • zod.dev
    • github.com
    • docs.github.com
    • langfuse.com
    • postgresql.org

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Claude Code 2.1.277+.

    From compatibility in the SKILL.md frontmatter.

Context cost

Golden Dataset loads about 2k tokens when it runs, and up to ~8.1k if it reads all its reference files. Until then it costs about 63 tokens; SKILL.md has 676 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~63
When it runs · the whole SKILL.md, loaded when a task matches
~2k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~8.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from yonatangross/orchestkit at commit e4ff8d9, republished under its MIT licence (© yonatangross). 676 words, ~2,042 tokens.

Download SKILL.mdSave it as .claude/skills/golden-dataset/SKILL.md (or your agent's skills folder). This skill also uses 16 other files; get the full folder from GitHub.
name
golden-dataset
description
Golden dataset lifecycle patterns for curation, versioning, quality validation, and CI integration. Use when building evaluation datasets, managing dataset versions, validating quality scores, or integrating golden tests into pipelines.
allowed-tools
Read, Glob, Grep, WebFetch, WebSearch
compatibility
Claude Code 2.1.277+.
license
MIT
user-invocable
false
disable-model-invocation
true
metadata.owner-agent
data-pipeline-engineer
metadata.category
document-asset-creation
metadata.version
2.0.0
metadata.author
OrchestKit
metadata.complexity
medium
metadata.tags
golden-dataset, evaluation, dataset-curation, dataset-validation, quality, llm-testing

Golden Dataset

Comprehensive patterns for building, managing, and validating golden datasets for AI/ML evaluation. Each category has individual rule files in rules/ loaded on-demand.

Quick Reference

CategoryRulesImpactWhen to Use
Curation2HIGHContent collection, annotation pipelines
Management2HIGHVersioning, backup/restore
Validation1CRITICALRegression testing
Add Workflow1HIGH9-phase curation, quality scoring, bias detection, silver-to-gold

Total: 6 rules across 4 categories. House thresholds and scars: references/ork-delta.md.

Curation

Content collection, multi-agent annotation, and diversity analysis for golden datasets.

RuleFileKey Pattern
Collectionrules/curation-collection.mdContent type classification, quality thresholds, duplicate prevention
Annotationrules/curation-annotation.mdMulti-agent pipeline, consensus aggregation, Langfuse tracing

Difficulty ladder, coverage floors, and duplicate thresholds: references/ork-delta.md.

Management

Versioning, storage, and CI/CD automation for golden datasets.

RuleFileKey Pattern
Versioningrules/management-versioning.mdJSON backup format, embedding regeneration, disaster recovery
Storagerules/management-storage.mdBackup strategies, URL contract, data integrity checks

CI automation for backups is upstream's job; see "Upstream coverage" below.

Validation

Quality scoring, drift detection, and regression testing for golden datasets.

RuleFileKey Pattern
Regressionrules/validation-regression.mdDifficulty distribution, pre-commit hooks, full dataset validation

Schema validation and duplicate detection are upstream's job (see "Upstream coverage" below); the house thresholds they must enforce live in references/ork-delta.md.

Add Workflow

Structured workflow for adding new documents to the golden dataset.

RuleFileKey Pattern
Add Documentrules/curation-add-workflow.md9-phase curation, parallel quality analysis, bias detection

Quick Start Example

python
async def validate_before_add(document: dict, source_url_map: dict) -> dict:
    """Pre-addition validation for golden dataset entries."""
    errors = []

    # 1. URL contract check
    if "placeholder" in document.get("source_url", ""):
        errors.append("URL must be canonical, not a placeholder")

    # 2. Content quality
    if len(document.get("title", "")) < 10:
        errors.append("Title too short (min 10 chars)")

    # 3. Tag requirements
    if len(document.get("tags", [])) < 2:
        errors.append("At least 2 domain tags required")

    return {"valid": len(errors) == 0, "errors": errors}

Key Decisions

DecisionRecommendation
Backup formatJSON (version controlled, portable)
Embedding storageExclude from backup (regenerate on restore)
Quality threshold>= 0.70 quality score for inclusion
Confidence threshold>= 0.65 for auto-include
Duplicate threshold>= 0.90 similarity blocks, >= 0.85 warns
Min tags per entry2 domain tags
Min test queries3 per document
Difficulty balanceTrivial 3, Easy 3, Medium 5, Hard 3 minimum
CI frequencyWeekly automated backup (Sunday 2am UTC)

Common Mistakes

  1. Using placeholder URLs instead of canonical source URLs
  2. Skipping embedding regeneration after restore
  3. Not validating referential integrity between documents and queries
  4. Over-indexing on articles (neglecting tutorials, research papers)
  5. Missing difficulty distribution balance in test queries
  6. Not running verification after backup/restore operations
  7. Testing restore procedures in production instead of staging
  8. Committing SQL dumps instead of JSON (not version-control friendly)
Show full SKILL.md (312 more words)Show less

Running a dataset as an experiment

Curating a dataset is half the job; the other half is running something against it and scoring the result. Both Langfuse SDKs ship a runner, and their shapes differ.

Python (SDK 4.x): see monitoring-observability/references/experiments-api.md.

JS/TS (SDK 5.x): @langfuse/client exposes the runner directly on a fetched dataset.

typescript
import { LangfuseClient } from "@langfuse/client";

const langfuse = new LangfuseClient();
const dataset = await langfuse.dataset.get("my-evaluation-dataset");

const result = await dataset.runExperiment({
  name: "Retrieval quality",
  task: myTask,               // (params) => Promise<any>
  evaluators: [myEvaluator],  // per-item: (params) => Promise<Evaluation | Evaluation[]>
});
TypeScoresUse for
Evaluatorone itemPer-example quality (faithfulness, relevance)
RunEvaluatorthe whole runAggregate assertions — pass rate, mean score, regression checks
Evaluation—{ name, value, comment?, metadata?, dataType?, configId? }

A per-item Evaluator cannot see the other items, so anything comparative belongs in a RunEvaluator. createEvaluatorFromAutoevals wraps an autoevals scorer instead of hand-writing one, and RegressionError is thrown when a run regresses against a configured baseline — catch it to fail CI on a quality drop rather than only on an exception.

Full JS surface: monitoring-observability/references/langfuse-js-v5.md.

Evaluations

See test-cases.json for 9 test cases across all categories.

Upstream coverage (do not restate)

TopicFirst-party source
Dataset schema validation (JSON Schema, field constraints)https://json-schema.org and https://zod.dev
Duplicate detection via embeddings, cosine similarityhttps://github.com/pgvector/pgvector
Scheduled backup automation (cron workflows, commit bots)https://docs.github.com/actions/using-workflows/events-that-trigger-workflows#schedule
Dataset runs, experiment scoring, annotation queueshttps://langfuse.com/docs/datasets
Backup and restore mechanics for postgres datasetshttps://www.postgresql.org/docs/current/backup.html

House thresholds these must enforce: references/ork-delta.md.

  • ork:rag-retrieval - Retrieval evaluation using golden dataset
  • ork:monitoring-observability - Langfuse tracing patterns for curation workflows
  • ork:testing-llm - Evaluation harnesses that consume golden datasets
  • ork:testing-unit - Unit testing patterns and strategies

Capability Details

curation

Keywords: golden dataset, curation, content collection, annotation, quality criteria

Solves:

  • Classify document content types for golden dataset
  • Run multi-agent quality analysis pipelines
  • Generate test queries for new documents
management

Keywords: golden dataset, backup, restore, versioning, disaster recovery

Solves:

  • Backup and restore golden datasets with JSON
  • Regenerate embeddings after restore
  • Automate backups with CI/CD
validation

Keywords: golden dataset, validation, schema, duplicate detection, quality metrics

Solves:

  • Validate entries against document schema
  • Detect duplicate or near-duplicate entries
  • Analyze dataset coverage and distribution gaps

© yonatangross, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 16 other files (scripts, references) in src/skills/golden-dataset of yonatangross/orchestkit.

  • SKILL.md
  • metadata.json
  • references/ork-delta.md
  • references/quality-metrics.md
  • references/storage-patterns.md
  • references/versioning.md
  • rules/_sections.md
  • rules/_template.md
  • rules/curation-add-workflow.md
  • rules/curation-annotation.md
  • rules/curation-collection.md
  • rules/management-storage.md
  • rules/management-versioning.md
  • rules/validation-regression.md
  • scripts/backup-golden-dataset.md
  • scripts/backup-script.py
  • test-cases.json

Open the folder on GitHubat commit e4ff8d9

Compare with similar skills

Golden Dataset next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Golden Dataset compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Golden Dataset this skillyonatangross/orchestkit292—~2kAutomated safety check: PassMIT
Monitor CInrwl/nx29k6 repos~4.7kAutomated safety check: PassMIT
Terraform and OpenTofu Guideagentscope-ai/QwenPaw36k6 repos~4.2kAutomated safety check: PassApache-2.0
Vercel Optimize Auditvercel-labs/agent-skills32k8 repos~4.3kAutomated safety check: PassNone
Analyze GitHub Action Logswithastro/astro63k1 repos~1.3kAutomated safety check: PassCustom licence
Openclaw Live Updateropenclaw/openclaw392k—~3.7kAutomated safety check: PassMIT

Similar skills

  • Monitor CI

    nrwl/nx

    Monitor Nx Cloud CI pipeline and handle self-healing fixes. An agent skill from nrwl/nx.

    29k GitHub starsUsed in 6 repos~4.7k tokens
    DevOps & CloudAuto-check passed
  • Terraform and OpenTofu Guide

    agentscope-ai/QwenPaw

    Guidance for writing and testing Terraform and OpenTofu code: module structure, naming, test approaches, CI/CD workflows, state handling and security scanning.

    36k GitHub starsUsed in 6 repos~4.2k tokens
    DevOps & CloudAuto-check passed
  • Vercel Optimize Audit

    vercel-labs/agent-skills

    Official

    Runs a metrics-first audit of a deployed Vercel project, gating investigations on real signals to produce ranked, citation-backed cost and performance recommendations.

    32k GitHub starsUsed in 8 repos~4.3k tokens
    DevOps & CloudAuto-check passed
  • Official

    Analyze recent GitHub Actions workflow runs to identify patterns, mistakes, and improvements.

    63k GitHub starsUsed in 1 repo~1.3k tokens
    DevOps & CloudAuto-check passed
  • Openclaw Live Updater

    openclaw/openclaw

    Maintain the canonical live OpenClaw main checkout, macOS LaunchAgent-managed Gateway, local macOS app, exact-head main CI, and recurring full release validation.

    392k GitHub stars~3.7k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Docs Learn PR Preview

    netdata/netdata

    Use only when the user explicitly asks to build, run, preview, inspect, or validate learn.netdata.cloud locally using the contents of a PR or documentation branch before merge.

    81k GitHub stars~2k tokensUpdated today
    DevOps & CloudAuto-check passed

More from yonatangross/orchestkit

All 107 skills in this repo
  • API Design

    yonatangross/orchestkit

    API contract design for REST and GraphQL, covering resource shape, URL and header versioning with deprecation windows, RFC 9457 Problem Details error handling, and OpenAPI specs.

    292 GitHub stars~2.9k tokensUpdated today
    Auto-check passed
  • Architecture Decision Record

    yonatangross/orchestkit

    ADR templates in the Nygard format with context, decision, consequences, and alternatives.

    292 GitHub stars~2k tokensUpdated today
    Auto-check passed
  • Audit Full

    yonatangross/orchestkit

    Single-pass codebase analysis leveraging a 1M-token context window for comprehensive security scanning, architecture review, and dependency auditing.

    292 GitHub stars~3.5k tokensUpdated today
    Auto-check: notes
  • Code Review Playbook

    yonatangross/orchestkit

    Structured review processes, conventional comments, language-specific checklists, and feedback templates.

    292 GitHub stars~2.2k tokensUpdated today
    Auto-check passed
  • Create PR

    yonatangross/orchestkit

    Creates GitHub pull requests with pre-flight validation, conventional title formatting, and structured summary generation.

    292 GitHub stars~4.5k tokensUpdated today
    Auto-check: notes
  • Explore

    yonatangross/orchestkit

    Multi-angle codebase exploration spawning 3-5 parallel agents for code structure, data flow, architecture patterns, and health assessment.

    292 GitHub stars~3.9k tokensUpdated today
    Auto-check: notes

Categories

Questions about Golden Dataset

What does Golden Dataset do?

Golden dataset lifecycle patterns for curation, versioning, quality validation, and CI integration. Golden Dataset is an agent skill from yonatangross/orchestkit. Golden dataset lifecycle patterns for curation, versioning, quality validation, and CI integration.

When should I use Golden Dataset?

Golden Dataset fits situations like: building evaluation datasets; managing dataset versions; validating quality scores; integrating golden tests into pipelines.

How do I install Golden Dataset in Claude Code?

Run `npx skills add yonatangross/orchestkit --skill golden-dataset -a claude-code`. Or copy the skill folder (src/skills/golden-dataset in yonatangross/orchestkit) into .claude/skills/golden-dataset in your project. Claude Code loads it when a task matches its description.

How do I install Golden Dataset in Codex?

Run `npx skills add yonatangross/orchestkit --skill golden-dataset -a codex`. Or copy the skill folder (src/skills/golden-dataset in yonatangross/orchestkit) into .agents/skills/golden-dataset in your project. Codex loads it when a task matches its description.

Can I use Golden Dataset in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add yonatangross/orchestkit --skill golden-dataset -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/golden-dataset, .gemini/skills/golden-dataset, .github/skills/golden-dataset and .opencode/skills/golden-dataset in your project.

What does Golden Dataset need to run?

Going by SKILL.md and its folder, Golden Dataset needs Python for the scripts in its folder. Our summary lists: Python 3. Its frontmatter pre-approves these tools: Read, Glob, Grep, WebFetch, WebSearch. Compatibility (from SKILL.md): Claude Code 2.1.277+..

Does Golden Dataset access the network?

SKILL.md names 6 domains. As links in the text: json-schema.org, zod.dev, github.com, docs.github.com, langfuse.com and postgresql.org. This is read from the text; nothing was executed.

Is Golden Dataset safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Golden Dataset use?

Golden Dataset is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Golden Dataset use?

About 2k tokens (SKILL.md is roughly 8.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 6k tokens, read only when the agent opens those files.

What are the alternatives to Golden Dataset?

Skills that share tags, products or a category with Golden Dataset: Monitor CI (nrwl/nx, 29k stars), Terraform and OpenTofu Guide (agentscope-ai/QwenPaw, 36k stars), Vercel Optimize Audit (vercel-labs/agent-skills, 32k stars) and Analyze GitHub Action Logs (withastro/astro, 63k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Golden Dataset?

yonatangross (a GitHub user) maintains it in yonatangross/orchestkit, which has 292 GitHub stars. The repository holds 108 skills in this directory. The repository was last updated on October 10, 2026.

Source: yonatangross/orchestkit on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.