Monitor CI
nrwl/nx
Monitor Nx Cloud CI pipeline and handle self-healing fixes. An agent skill from nrwl/nx.
Golden dataset lifecycle patterns for curation, versioning, quality validation, and CI integration.
$ npx skills add yonatangross/orchestkit --skill golden-dataset -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install yonatangross/orchestkit golden-dataset --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/yonatangross/orchestkit.git skills-src && mkdir -p .claude/skills && cp -r skills-src/src/skills/golden-dataset .claude/skills/golden-dataset && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "golden-dataset" agent skill from https://github.com/yonatangross/orchestkit/tree/main/src/skills/golden-dataset into .claude/skills/golden-dataset/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "golden-dataset", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/yonatangross/orchestkit/tree/main/src/skills/golden-datasetType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add yonatangross/orchestkit --skill golden-dataset -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install yonatangross/orchestkit golden-dataset --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/yonatangross/orchestkit.git skills-src && mkdir -p .agents/skills && cp -r skills-src/src/skills/golden-dataset .agents/skills/golden-dataset && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "golden-dataset" agent skill from https://github.com/yonatangross/orchestkit/tree/main/src/skills/golden-dataset into .agents/skills/golden-dataset/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "golden-dataset", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add yonatangross/orchestkit --skill golden-dataset -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install yonatangross/orchestkit golden-dataset --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/yonatangross/orchestkit.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/src/skills/golden-dataset .cursor/skills/golden-dataset && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "golden-dataset" agent skill from https://github.com/yonatangross/orchestkit/tree/main/src/skills/golden-dataset into .cursor/skills/golden-dataset/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "golden-dataset", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/yonatangross/orchestkit.git --path src/skills/golden-dataset--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add yonatangross/orchestkit --skill golden-dataset -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install yonatangross/orchestkit golden-dataset --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/yonatangross/orchestkit.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/src/skills/golden-dataset .gemini/skills/golden-dataset && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "golden-dataset" agent skill from https://github.com/yonatangross/orchestkit/tree/main/src/skills/golden-dataset into .gemini/skills/golden-dataset/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "golden-dataset", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install yonatangross/orchestkit golden-datasetInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add yonatangross/orchestkit --skill golden-dataset -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/yonatangross/orchestkit.git skills-src && mkdir -p .github/skills && cp -r skills-src/src/skills/golden-dataset .github/skills/golden-dataset && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "golden-dataset" agent skill from https://github.com/yonatangross/orchestkit/tree/main/src/skills/golden-dataset into .github/skills/golden-dataset/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "golden-dataset", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add yonatangross/orchestkit --skill golden-dataset -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install yonatangross/orchestkit golden-dataset --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/yonatangross/orchestkit.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/src/skills/golden-dataset .opencode/skills/golden-dataset && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "golden-dataset" agent skill from https://github.com/yonatangross/orchestkit/tree/main/src/skills/golden-dataset into .opencode/skills/golden-dataset/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "golden-dataset", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
golden-datasetGolden dataset lifecycle patterns for curation, versioning, quality validation, and CI integration.
Golden Dataset is an agent skill from yonatangross/orchestkit. Golden dataset lifecycle patterns for curation, versioning, quality validation, and CI integration. Use when building evaluation datasets, managing dataset versions, validating quality scores, or integrating golden tests into pipelines.
Its SKILL.md is about 2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 19 other files, including scripts and reference files (for example `metadata.json`, `references/ork-delta.md` and `references/quality-metrics.md`). Compatibility notes: Claude Code 2.1.277+.
It sits in DevOps & Cloud. The repository describes itself as: The Complete AI Development Toolkit for Claude Code. 106 skills, 36 agents, 171 hooks. Install ork for stable (v9.x), or ork-alpha for the v10 line, which ships daily. The licence is MIT.
8 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit e4ff8d9. It shows what the files ask for, not the result of running them.
Pre-approves these tools, so the agent can use them without asking each time:
ReadGlobGrepWebFetchWebSearchFrom allowed-tools in the SKILL.md frontmatter.
Ships 2 files in scripts/ (Python), which the agent can run.
From the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
json-schema.orgzod.devgithub.comdocs.github.comlangfuse.compostgresql.orgFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Claude Code 2.1.277+.
From compatibility in the SKILL.md frontmatter.
Golden Dataset loads about 2k tokens when it runs, and up to ~8.1k if it reads all its reference files. Until then it costs about 63 tokens; SKILL.md has 676 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from yonatangross/orchestkit at commit e4ff8d9, republished under its MIT licence (© yonatangross). 676 words, ~2,042 tokens.
.claude/skills/golden-dataset/SKILL.md (or your agent's skills folder). This skill also uses 16 other files; get the full folder from GitHub.Comprehensive patterns for building, managing, and validating golden datasets for AI/ML evaluation. Each category has individual rule files in rules/ loaded on-demand.
| Category | Rules | Impact | When to Use |
|---|---|---|---|
| Curation | 2 | HIGH | Content collection, annotation pipelines |
| Management | 2 | HIGH | Versioning, backup/restore |
| Validation | 1 | CRITICAL | Regression testing |
| Add Workflow | 1 | HIGH | 9-phase curation, quality scoring, bias detection, silver-to-gold |
Total: 6 rules across 4 categories. House thresholds and scars: references/ork-delta.md.
Content collection, multi-agent annotation, and diversity analysis for golden datasets.
| Rule | File | Key Pattern |
|---|---|---|
| Collection | rules/curation-collection.md | Content type classification, quality thresholds, duplicate prevention |
| Annotation | rules/curation-annotation.md | Multi-agent pipeline, consensus aggregation, Langfuse tracing |
Difficulty ladder, coverage floors, and duplicate thresholds: references/ork-delta.md.
Versioning, storage, and CI/CD automation for golden datasets.
| Rule | File | Key Pattern |
|---|---|---|
| Versioning | rules/management-versioning.md | JSON backup format, embedding regeneration, disaster recovery |
| Storage | rules/management-storage.md | Backup strategies, URL contract, data integrity checks |
CI automation for backups is upstream's job; see "Upstream coverage" below.
Quality scoring, drift detection, and regression testing for golden datasets.
| Rule | File | Key Pattern |
|---|---|---|
| Regression | rules/validation-regression.md | Difficulty distribution, pre-commit hooks, full dataset validation |
Schema validation and duplicate detection are upstream's job (see "Upstream coverage"
below); the house thresholds they must enforce live in references/ork-delta.md.
Structured workflow for adding new documents to the golden dataset.
| Rule | File | Key Pattern |
|---|---|---|
| Add Document | rules/curation-add-workflow.md | 9-phase curation, parallel quality analysis, bias detection |
async def validate_before_add(document: dict, source_url_map: dict) -> dict:
"""Pre-addition validation for golden dataset entries."""
errors = []
# 1. URL contract check
if "placeholder" in document.get("source_url", ""):
errors.append("URL must be canonical, not a placeholder")
# 2. Content quality
if len(document.get("title", "")) < 10:
errors.append("Title too short (min 10 chars)")
# 3. Tag requirements
if len(document.get("tags", [])) < 2:
errors.append("At least 2 domain tags required")
return {"valid": len(errors) == 0, "errors": errors}| Decision | Recommendation |
|---|---|
| Backup format | JSON (version controlled, portable) |
| Embedding storage | Exclude from backup (regenerate on restore) |
| Quality threshold | >= 0.70 quality score for inclusion |
| Confidence threshold | >= 0.65 for auto-include |
| Duplicate threshold | >= 0.90 similarity blocks, >= 0.85 warns |
| Min tags per entry | 2 domain tags |
| Min test queries | 3 per document |
| Difficulty balance | Trivial 3, Easy 3, Medium 5, Hard 3 minimum |
| CI frequency | Weekly automated backup (Sunday 2am UTC) |
Curating a dataset is half the job; the other half is running something against it and scoring the result. Both Langfuse SDKs ship a runner, and their shapes differ.
Python (SDK 4.x): see monitoring-observability/references/experiments-api.md.
JS/TS (SDK 5.x): @langfuse/client exposes the runner directly on a fetched dataset.
import { LangfuseClient } from "@langfuse/client";
const langfuse = new LangfuseClient();
const dataset = await langfuse.dataset.get("my-evaluation-dataset");
const result = await dataset.runExperiment({
name: "Retrieval quality",
task: myTask, // (params) => Promise<any>
evaluators: [myEvaluator], // per-item: (params) => Promise<Evaluation | Evaluation[]>
});| Type | Scores | Use for |
|---|---|---|
Evaluator | one item | Per-example quality (faithfulness, relevance) |
RunEvaluator | the whole run | Aggregate assertions — pass rate, mean score, regression checks |
Evaluation | — | { name, value, comment?, metadata?, dataType?, configId? } |
A per-item Evaluator cannot see the other items, so anything comparative belongs in a
RunEvaluator. createEvaluatorFromAutoevals wraps an autoevals scorer instead of hand-writing
one, and RegressionError is thrown when a run regresses against a configured baseline — catch it
to fail CI on a quality drop rather than only on an exception.
Full JS surface: monitoring-observability/references/langfuse-js-v5.md.
See test-cases.json for 9 test cases across all categories.
| Topic | First-party source |
|---|---|
| Dataset schema validation (JSON Schema, field constraints) | https://json-schema.org and https://zod.dev |
| Duplicate detection via embeddings, cosine similarity | https://github.com/pgvector/pgvector |
| Scheduled backup automation (cron workflows, commit bots) | https://docs.github.com/actions/using-workflows/events-that-trigger-workflows#schedule |
| Dataset runs, experiment scoring, annotation queues | https://langfuse.com/docs/datasets |
| Backup and restore mechanics for postgres datasets | https://www.postgresql.org/docs/current/backup.html |
House thresholds these must enforce: references/ork-delta.md.
ork:rag-retrieval - Retrieval evaluation using golden datasetork:monitoring-observability - Langfuse tracing patterns for curation workflowsork:testing-llm - Evaluation harnesses that consume golden datasetsork:testing-unit - Unit testing patterns and strategiesKeywords: golden dataset, curation, content collection, annotation, quality criteria
Solves:
Keywords: golden dataset, backup, restore, versioning, disaster recovery
Solves:
Keywords: golden dataset, validation, schema, duplicate detection, quality metrics
Solves:
© yonatangross, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 16 other files (scripts, references) in src/skills/golden-dataset of yonatangross/orchestkit.
Open the folder on GitHubat commit e4ff8d9
Golden Dataset next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Golden Dataset this skillyonatangross/orchestkit | 292 | — | ~2k | Automated safety check: Pass | MIT | |
| Monitor CInrwl/nx | 29k | 6 repos | ~4.7k | Automated safety check: Pass | MIT | |
| Terraform and OpenTofu Guideagentscope-ai/QwenPaw | 36k | 6 repos | ~4.2k | Automated safety check: Pass | Apache-2.0 | |
| Vercel Optimize Auditvercel-labs/agent-skills | 32k | 8 repos | ~4.3k | Automated safety check: Pass | None | |
| Analyze GitHub Action Logswithastro/astro | 63k | 1 repos | ~1.3k | Automated safety check: Pass | Custom licence | |
| Openclaw Live Updateropenclaw/openclaw | 392k | — | ~3.7k | Automated safety check: Pass | MIT |
nrwl/nx
Monitor Nx Cloud CI pipeline and handle self-healing fixes. An agent skill from nrwl/nx.
agentscope-ai/QwenPaw
Guidance for writing and testing Terraform and OpenTofu code: module structure, naming, test approaches, CI/CD workflows, state handling and security scanning.
vercel-labs/agent-skills
Runs a metrics-first audit of a deployed Vercel project, gating investigations on real signals to produce ranked, citation-backed cost and performance recommendations.
withastro/astro
Analyze recent GitHub Actions workflow runs to identify patterns, mistakes, and improvements.
openclaw/openclaw
Maintain the canonical live OpenClaw main checkout, macOS LaunchAgent-managed Gateway, local macOS app, exact-head main CI, and recurring full release validation.
netdata/netdata
Use only when the user explicitly asks to build, run, preview, inspect, or validate learn.netdata.cloud locally using the contents of a PR or documentation branch before merge.
yonatangross/orchestkit
API contract design for REST and GraphQL, covering resource shape, URL and header versioning with deprecation windows, RFC 9457 Problem Details error handling, and OpenAPI specs.
yonatangross/orchestkit
ADR templates in the Nygard format with context, decision, consequences, and alternatives.
yonatangross/orchestkit
Single-pass codebase analysis leveraging a 1M-token context window for comprehensive security scanning, architecture review, and dependency auditing.
yonatangross/orchestkit
Structured review processes, conventional comments, language-specific checklists, and feedback templates.
yonatangross/orchestkit
Creates GitHub pull requests with pre-flight validation, conventional title formatting, and structured summary generation.
yonatangross/orchestkit
Multi-angle codebase exploration spawning 3-5 parallel agents for code structure, data flow, architecture patterns, and health assessment.
Categories
Golden dataset lifecycle patterns for curation, versioning, quality validation, and CI integration. Golden Dataset is an agent skill from yonatangross/orchestkit. Golden dataset lifecycle patterns for curation, versioning, quality validation, and CI integration.
Golden Dataset fits situations like: building evaluation datasets; managing dataset versions; validating quality scores; integrating golden tests into pipelines.
Run `npx skills add yonatangross/orchestkit --skill golden-dataset -a claude-code`. Or copy the skill folder (src/skills/golden-dataset in yonatangross/orchestkit) into .claude/skills/golden-dataset in your project. Claude Code loads it when a task matches its description.
Run `npx skills add yonatangross/orchestkit --skill golden-dataset -a codex`. Or copy the skill folder (src/skills/golden-dataset in yonatangross/orchestkit) into .agents/skills/golden-dataset in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add yonatangross/orchestkit --skill golden-dataset -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/golden-dataset, .gemini/skills/golden-dataset, .github/skills/golden-dataset and .opencode/skills/golden-dataset in your project.
Going by SKILL.md and its folder, Golden Dataset needs Python for the scripts in its folder. Our summary lists: Python 3. Its frontmatter pre-approves these tools: Read, Glob, Grep, WebFetch, WebSearch. Compatibility (from SKILL.md): Claude Code 2.1.277+..
SKILL.md names 6 domains. As links in the text: json-schema.org, zod.dev, github.com, docs.github.com, langfuse.com and postgresql.org. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Golden Dataset is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 2k tokens (SKILL.md is roughly 8.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 6k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Golden Dataset: Monitor CI (nrwl/nx, 29k stars), Terraform and OpenTofu Guide (agentscope-ai/QwenPaw, 36k stars), Vercel Optimize Audit (vercel-labs/agent-skills, 32k stars) and Analyze GitHub Action Logs (withastro/astro, 63k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
yonatangross (a GitHub user) maintains it in yonatangross/orchestkit, which has 292 GitHub stars. The repository holds 108 skills in this directory. The repository was last updated on October 10, 2026.
Source: yonatangross/orchestkit on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.