Agent skill

Version Dataset

by Aperivue in Aperivue/medsci-skills

A skill your agent uses when you must prove an analysis ran on the intended data or lock a dataset version.

MITAuto-check passedResearch & Science

Install Version Dataset

skills CLI
$ npx skills add Aperivue/medsci-skills --skill version-dataset -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Aperivue/medsci-skills version-dataset --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Aperivue/medsci-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/version-dataset .claude/skills/version-dataset && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
version-dataset
GitHub stars
329
Token cost
~1.1k tokens
SKILL.md length
467 words
Files
5 (incl. scripts, references)
Skills in repo
54
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when you must prove an analysis ran on the intended data or lock a dataset version.

  • Works in 3 steps: Lock the version (gate) → Verify before each run (gate) → Diff across versions
  • You must prove an analysis ran on the intended data
  • SKILL.md covers Deterministic Script, Workflow, Non-Deterministic Artifacts and Scope
  • Runs Python and Shell scripts from its folder; calls python

What it does

Version Dataset is an agent skill from Aperivue/medsci-skills. Use when you must prove an analysis ran on the intended data or lock a dataset version. Builds a deterministic content-hash manifest (file SHA-256, schema, per-column value hashes), verifies later copies against it for drift, and diffs two manifests.

Its SKILL.md is about 1.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 7 other files, including scripts and reference files (for example `references/manifest_schema.md`, `scripts/version_dataset.py` and `skill.yml`).

It sits in Research & Science. The repository describes itself as: Agent Skills for medical research — literature search, reporting-guideline & citation checks, statistics, publication figures, submission. Works with Claude Code, Codex, Cursor &… The licence is MIT.

When your agent uses it

  • You must prove an analysis ran on the intended data
  • Lock a dataset version

Example prompts

  • “/version-dataset”

Requirements

  • Python 3
  • A Bash shell

Workflow steps

3 steps, taken from the step headings in SKILL.md.

  1. Lock the version (gate)
  2. Verify before each run (gate)
  3. Diff across versions

What it can do on your machine

Read from SKILL.md and the folder at commit 3b14ae2. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python and Shell), which the agent can run.

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Version Dataset loads about 1.1k tokens when it runs, and up to ~2.4k if it reads all its reference files. Until then it costs about 67 tokens; SKILL.md has 467 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~67
When it runs · the whole SKILL.md, loaded when a task matches
~1.1k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~2.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from Aperivue/medsci-skills at commit 3b14ae2, republished under its MIT licence (© Aperivue). 467 words, ~1,055 tokens.

Download SKILL.mdSave it as .claude/skills/version-dataset/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
version-dataset
description
Use when you must prove an analysis ran on the intended data or lock a dataset version. Builds a deterministic content-hash manifest (file SHA-256, schema, per-column value hashes), verifies later copies against it for drift, and diffs two manifests.
metadata.triggers
version dataset, dataset version, data manifest, data hash, dataset drift, reproducibility lock, verify dataset, data provenance, did my data change…

Version Dataset Skill

An analysis must run on the data it claims to, with a fixed seed. This skill makes drift between runs loud instead of silent: it records a deterministic fingerprint (file SHA-256 plus, for tabular files, schema and per-column value hashes; no timestamp unless explicitly passed) so a later run can prove the inputs are unchanged. It never alters data, and manifests hold hashes, not the data itself. Provenance notes are in English.

Deterministic Script

bash
# Build a manifest (record the analysis seed + provenance)
python "${CLAUDE_SKILL_DIR}/scripts/version_dataset.py" manifest data.csv \
  --out manifest.json --seed 42 --provenance "KNHANES 2018 extract v1"

# Verify a later copy against it (CI / pre-analysis gate)
python "${CLAUDE_SKILL_DIR}/scripts/version_dataset.py" verify --manifest manifest.json --strict

# Compare two manifests (what changed between versions)
python "${CLAUDE_SKILL_DIR}/scripts/version_dataset.py" diff --old v1.json --new v2.json

File hashing is stdlib-only; tabular schema/column hashing uses pandas when present. --ignore-cols excludes volatile columns; --base makes manifest keys relative.

Workflow

Step 1: Lock the version (gate)

Build the manifest at the moment the dataset is frozen for analysis. Gate: confirm with the user the seed and provenance note are correct before locking — the manifest is the record they will cite as "this is the data the results came from." The provenance note is the user's text; never invent it.

Step 2: Verify before each run (gate)

Before re-running an analysis (or in CI), verify --strict. Never say a dataset is unchanged without running verify. Gate: if drift is reported (MANIFEST_DRIFT, exit 1), stop and show the user the drift report exactly as computed — never downplay a changed column hash. Do not proceed on changed data without their explicit acknowledgement and a re-lock. Silent re-run on drifted data is the failure this skill exists to prevent.

Read ${CLAUDE_SKILL_DIR}/references/manifest_schema.md before interpreting drift — it defines each drift category. A CSV/TSV file is compared on its logical content (schema + per-column value hashes), not raw bytes, so re-quoting, reordering columns, or an --ignore-cols volatile column does not trip a false drift. Binary tabular files (Parquet/Stata/SAS/Excel) are also checked at the byte level, since labels, metadata and other sheets are not in the column hashes, unless an --ignore-cols column is present in the file (its changes alter the bytes).

Show full SKILL.md (156 more words)Show less
Step 3: Diff across versions

When a dataset is intentionally updated, diff the old and new manifests and present the change set (added/removed/changed columns, row-count delta) so the user can record what changed and re-lock. Gate: the user approves the new version before it replaces the locked one.

Non-Deterministic Artifacts

Do not put PPTX/DOCX or figure binaries under strict byte verification — they embed timestamps or render metadata and change on every build. Manifest only the deterministic inputs and tabular outputs (data files, result CSVs), or use --ignore-cols for volatile columns; the policy is in references/manifest_schema.md. Each bundled demo/*/ carries a manifest.lock.json (input data

  • deterministic result tables) that verify --strict checks; the reference gives the demo command.

Scope

  • Run /deidentify before a manifest is shared: example values are not stored, but provenance notes may carry context.
  • Cleaning, profiling or de-identifying the data → /clean-data, /generate-codebook, /deidentify. /generate-codebook documents what is in the data; this skill locks which version.

© Aperivue, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files (scripts, references) in skills/version-dataset of Aperivue/medsci-skills.

  • SKILL.md
  • references/manifest_schema.md
  • scripts/version_dataset.py
  • skill.yml
  • tests/test_version_dataset.sh

Open the folder on GitHubat commit 3b14ae2

Compare with similar skills

Version Dataset next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Version Dataset compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Version Dataset this skillAperivue/medsci-skills329—~1.1kAutomated safety check: PassMIT
Hypothesis Generationspacering-net/codeg3.8k15 repos~3.6kAutomated safety check: NotesMIT
GitHub Deep Researchbytedance/deer-flow83k5 repos~1.3kAutomated safety check: PassMIT
Nature Paper CardYuan1z0825/nature-skills46k2 repos~2.1kAutomated safety check: PassApache-2.0
Read arXiv Paperkarpathy/nanochat58k2 repos~494Automated safety check: PassMIT
Content Research Writerweapp-tailwindcss/weapp-tailwindcss1.9k25 repos~3.5kAutomated safety check: PassMIT

Similar skills

  • Hypothesis Generation

    spacering-net/codeg

    Structured hypothesis formulation from observations. An agent skill from spacering-net/codeg.

    3.8k GitHub starsUsed in 15 repos~3.6k tokens
    Research & ScienceAuto-check: notes
  • GitHub Deep Research

    bytedance/deer-flow

    Researches a GitHub repository over four rounds using the GitHub API and web search, then writes a structured markdown report with timeline, metrics and Mermaid diagrams.

    83k GitHub starsUsed in 5 repos~1.3k tokens
    Research & ScienceAuto-check passed
  • Nature Paper Card

    Yuan1z0825/nature-skills

    Builds a structured deep-reading card for one scientific paper, covering methods, how experiments support claims, limitations and research ideas, with a script to prepare the source.

    46k GitHub starsUsed in 2 repos~2.1k tokens
    Research & ScienceAuto-check passed
  • Read arXiv Paper

    karpathy/nanochat

    Fetches the TeX source of an arXiv paper from its URL, reads it and writes a markdown summary tied to the nanochat project.

    58k GitHub starsUsed in 2 repos~494 tokens
    Research & ScienceAuto-check passed
  • Content Research Writer

    weapp-tailwindcss/weapp-tailwindcss

    Assists in writing high-quality content by conducting research, adding citations, improving hooks, iterating on outlines, and providing real-time feedback on each section.

    1.9k GitHub starsUsed in 25 repos~3.5k tokens
    Research & ScienceAuto-check passed
  • Peer Review

    spacering-net/codeg

    Structured manuscript/grant review with checklist-based evaluation.

    3.8k GitHub starsUsed in 18 repos~5.9k tokens
    Research & ScienceAuto-check: notes

More from Aperivue/medsci-skills

All 54 skills in this repo
  • Model Assessment

    Aperivue/medsci-skills

    A skill your agent uses when validating or evaluating a trained medical-imaging model.

    329 GitHub starsUsed in 1 repo~4.5k tokens
    Auto-check passed
  • Obsidian Paper Vault

    Aperivue/medsci-skills

    A skill your agent uses when turning a folder of research PDFs into Obsidian notes, even if Obsidian is not named.

    329 GitHub stars~1.6k tokensUpdated 3 days ago
    Auto-check passed
  • Radiomics ML

    Aperivue/medsci-skills

    A skill your agent uses when building or auditing a radiomics or tabular clinical-ML prediction model with a classical learner (LASSO, SVM, random forest, XGBoost and similar).

    329 GitHub starsUsed in 1 repo~2.7k tokens
    Auto-check passed
  • Verify Refs

    Aperivue/medsci-skills

    A skill your agent uses when checking whether a manuscript's references are real.

    329 GitHub starsUsed in 1 repo~3.1k tokens
    Auto-check passed
  • Clean Data

    Aperivue/medsci-skills

    A skill your agent uses when a clinical CSV/Excel dataset needs profiling and cleaning before analysis (missing values, outliers, duplicates, type mismatches).

    329 GitHub stars~2k tokensUpdated 3 days ago
    Auto-check passed
  • Design Study

    Aperivue/medsci-skills

    A skill your agent uses when checking a radiology or medical AI study design before drafting or submission.

    329 GitHub stars~3.9k tokensUpdated 3 days ago
    Auto-check passed

Questions about Version Dataset

What does Version Dataset do?

A skill your agent uses when you must prove an analysis ran on the intended data or lock a dataset version. Version Dataset is an agent skill from Aperivue/medsci-skills. Use when you must prove an analysis ran on the intended data or lock a dataset version.

When should I use Version Dataset?

Version Dataset fits situations like: you must prove an analysis ran on the intended data; lock a dataset version.

How do I install Version Dataset in Claude Code?

Run `npx skills add Aperivue/medsci-skills --skill version-dataset -a claude-code`. Or copy the skill folder (skills/version-dataset in Aperivue/medsci-skills) into .claude/skills/version-dataset in your project. Claude Code loads it when a task matches its description.

How do I install Version Dataset in Codex?

Run `npx skills add Aperivue/medsci-skills --skill version-dataset -a codex`. Or copy the skill folder (skills/version-dataset in Aperivue/medsci-skills) into .agents/skills/version-dataset in your project. Codex loads it when a task matches its description.

Can I use Version Dataset in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Aperivue/medsci-skills --skill version-dataset -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/version-dataset, .gemini/skills/version-dataset, .github/skills/version-dataset and .opencode/skills/version-dataset in your project.

What does Version Dataset need to run?

Going by SKILL.md and its folder, Version Dataset needs Python and a shell for the scripts in its folder and the command-line tools its instructions call (python). Our summary lists: Python 3; A Bash shell.

Does Version Dataset access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Version Dataset safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Version Dataset use?

Version Dataset is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Version Dataset use?

About 1.1k tokens (SKILL.md is roughly 4.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.4k tokens, read only when the agent opens those files.

What are the alternatives to Version Dataset?

Skills that share tags, products or a category with Version Dataset: Hypothesis Generation (spacering-net/codeg, 3.8k stars), GitHub Deep Research (bytedance/deer-flow, 83k stars), Nature Paper Card (Yuan1z0825/nature-skills, 46k stars) and Read arXiv Paper (karpathy/nanochat, 58k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Version Dataset?

Aperivue (a GitHub organization) maintains it in Aperivue/medsci-skills, which has 329 GitHub stars. The repository holds 54 skills in this directory. The repository was last updated on October 5, 2026.

Source: Aperivue/medsci-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.