Hypothesis Generation
spacering-net/codeg
Structured hypothesis formulation from observations. An agent skill from spacering-net/codeg.
A skill your agent uses when you must prove an analysis ran on the intended data or lock a dataset version.
$ npx skills add Aperivue/medsci-skills --skill version-dataset -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install Aperivue/medsci-skills version-dataset --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/Aperivue/medsci-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/version-dataset .claude/skills/version-dataset && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "version-dataset" agent skill from https://github.com/Aperivue/medsci-skills/tree/main/skills/version-dataset into .claude/skills/version-dataset/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "version-dataset", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/Aperivue/medsci-skills/tree/main/skills/version-datasetType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add Aperivue/medsci-skills --skill version-dataset -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install Aperivue/medsci-skills version-dataset --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Aperivue/medsci-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/version-dataset .agents/skills/version-dataset && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "version-dataset" agent skill from https://github.com/Aperivue/medsci-skills/tree/main/skills/version-dataset into .agents/skills/version-dataset/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "version-dataset", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Aperivue/medsci-skills --skill version-dataset -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install Aperivue/medsci-skills version-dataset --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Aperivue/medsci-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/version-dataset .cursor/skills/version-dataset && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "version-dataset" agent skill from https://github.com/Aperivue/medsci-skills/tree/main/skills/version-dataset into .cursor/skills/version-dataset/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "version-dataset", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/Aperivue/medsci-skills.git --path skills/version-dataset--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add Aperivue/medsci-skills --skill version-dataset -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install Aperivue/medsci-skills version-dataset --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Aperivue/medsci-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/version-dataset .gemini/skills/version-dataset && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "version-dataset" agent skill from https://github.com/Aperivue/medsci-skills/tree/main/skills/version-dataset into .gemini/skills/version-dataset/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "version-dataset", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install Aperivue/medsci-skills version-datasetInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add Aperivue/medsci-skills --skill version-dataset -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/Aperivue/medsci-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/version-dataset .github/skills/version-dataset && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "version-dataset" agent skill from https://github.com/Aperivue/medsci-skills/tree/main/skills/version-dataset into .github/skills/version-dataset/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "version-dataset", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Aperivue/medsci-skills --skill version-dataset -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install Aperivue/medsci-skills version-dataset --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Aperivue/medsci-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/version-dataset .opencode/skills/version-dataset && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "version-dataset" agent skill from https://github.com/Aperivue/medsci-skills/tree/main/skills/version-dataset into .opencode/skills/version-dataset/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "version-dataset", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
version-datasetA skill your agent uses when you must prove an analysis ran on the intended data or lock a dataset version.
Version Dataset is an agent skill from Aperivue/medsci-skills. Use when you must prove an analysis ran on the intended data or lock a dataset version. Builds a deterministic content-hash manifest (file SHA-256, schema, per-column value hashes), verifies later copies against it for drift, and diffs two manifests.
Its SKILL.md is about 1.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 7 other files, including scripts and reference files (for example `references/manifest_schema.md`, `scripts/version_dataset.py` and `skill.yml`).
It sits in Research & Science. The repository describes itself as: Agent Skills for medical research — literature search, reporting-guideline & citation checks, statistics, publication figures, submission. Works with Claude Code, Codex, Cursor &… The licence is MIT.
3 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 3b14ae2. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 1 file in scripts/ (Python and Shell), which the agent can run.
Shell commands in SKILL.md call:
pythonFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Version Dataset loads about 1.1k tokens when it runs, and up to ~2.4k if it reads all its reference files. Until then it costs about 67 tokens; SKILL.md has 467 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from Aperivue/medsci-skills at commit 3b14ae2, republished under its MIT licence (© Aperivue). 467 words, ~1,055 tokens.
.claude/skills/version-dataset/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.An analysis must run on the data it claims to, with a fixed seed. This skill makes drift between runs loud instead of silent: it records a deterministic fingerprint (file SHA-256 plus, for tabular files, schema and per-column value hashes; no timestamp unless explicitly passed) so a later run can prove the inputs are unchanged. It never alters data, and manifests hold hashes, not the data itself. Provenance notes are in English.
# Build a manifest (record the analysis seed + provenance)
python "${CLAUDE_SKILL_DIR}/scripts/version_dataset.py" manifest data.csv \
--out manifest.json --seed 42 --provenance "KNHANES 2018 extract v1"
# Verify a later copy against it (CI / pre-analysis gate)
python "${CLAUDE_SKILL_DIR}/scripts/version_dataset.py" verify --manifest manifest.json --strict
# Compare two manifests (what changed between versions)
python "${CLAUDE_SKILL_DIR}/scripts/version_dataset.py" diff --old v1.json --new v2.jsonFile hashing is stdlib-only; tabular schema/column hashing uses pandas when present.
--ignore-cols excludes volatile columns; --base makes manifest keys relative.
Build the manifest at the moment the dataset is frozen for analysis. Gate: confirm with the user the seed and provenance note are correct before locking — the manifest is the record they will cite as "this is the data the results came from." The provenance note is the user's text; never invent it.
Before re-running an analysis (or in CI), verify --strict. Never say a dataset is unchanged
without running verify. Gate: if drift is reported (MANIFEST_DRIFT, exit 1), stop and
show the user the drift report exactly as computed — never downplay a changed column hash. Do not
proceed on changed data without their explicit acknowledgement and a re-lock. Silent re-run on
drifted data is the failure this skill exists to prevent.
Read ${CLAUDE_SKILL_DIR}/references/manifest_schema.md before interpreting drift — it defines
each drift category. A CSV/TSV file is compared on its logical content (schema + per-column
value hashes), not raw bytes, so re-quoting, reordering columns, or an --ignore-cols volatile
column does not trip a false drift. Binary tabular files (Parquet/Stata/SAS/Excel) are also
checked at the byte level, since labels, metadata and other sheets are not in the column hashes,
unless an --ignore-cols column is present in the file (its changes alter the bytes).
When a dataset is intentionally updated, diff the old and new manifests and
present the change set (added/removed/changed columns, row-count delta) so the
user can record what changed and re-lock. Gate: the user approves the new
version before it replaces the locked one.
Do not put PPTX/DOCX or figure binaries under strict byte verification — they embed timestamps
or render metadata and change on every build. Manifest only the deterministic inputs and tabular
outputs (data files, result CSVs), or use --ignore-cols for volatile columns; the policy is in
references/manifest_schema.md. Each bundled demo/*/ carries a manifest.lock.json (input data
verify --strict checks; the reference gives the demo
command./deidentify before a manifest is shared: example values are not stored, but provenance
notes may carry context./clean-data, /generate-codebook,
/deidentify. /generate-codebook documents what is in the data; this skill locks which
version.© Aperivue, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 4 other files (scripts, references) in skills/version-dataset of Aperivue/medsci-skills.
Open the folder on GitHubat commit 3b14ae2
Version Dataset next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Version Dataset this skillAperivue/medsci-skills | 329 | — | ~1.1k | Automated safety check: Pass | MIT | |
| Hypothesis Generationspacering-net/codeg | 3.8k | 15 repos | ~3.6k | Automated safety check: Notes | MIT | |
| GitHub Deep Researchbytedance/deer-flow | 83k | 5 repos | ~1.3k | Automated safety check: Pass | MIT | |
| Nature Paper CardYuan1z0825/nature-skills | 46k | 2 repos | ~2.1k | Automated safety check: Pass | Apache-2.0 | |
| Read arXiv Paperkarpathy/nanochat | 58k | 2 repos | ~494 | Automated safety check: Pass | MIT | |
| Content Research Writerweapp-tailwindcss/weapp-tailwindcss | 1.9k | 25 repos | ~3.5k | Automated safety check: Pass | MIT |
spacering-net/codeg
Structured hypothesis formulation from observations. An agent skill from spacering-net/codeg.
bytedance/deer-flow
Researches a GitHub repository over four rounds using the GitHub API and web search, then writes a structured markdown report with timeline, metrics and Mermaid diagrams.
Yuan1z0825/nature-skills
Builds a structured deep-reading card for one scientific paper, covering methods, how experiments support claims, limitations and research ideas, with a script to prepare the source.
karpathy/nanochat
Fetches the TeX source of an arXiv paper from its URL, reads it and writes a markdown summary tied to the nanochat project.
weapp-tailwindcss/weapp-tailwindcss
Assists in writing high-quality content by conducting research, adding citations, improving hooks, iterating on outlines, and providing real-time feedback on each section.
spacering-net/codeg
Structured manuscript/grant review with checklist-based evaluation.
Aperivue/medsci-skills
A skill your agent uses when validating or evaluating a trained medical-imaging model.
Aperivue/medsci-skills
A skill your agent uses when turning a folder of research PDFs into Obsidian notes, even if Obsidian is not named.
Aperivue/medsci-skills
A skill your agent uses when building or auditing a radiomics or tabular clinical-ML prediction model with a classical learner (LASSO, SVM, random forest, XGBoost and similar).
Aperivue/medsci-skills
A skill your agent uses when checking whether a manuscript's references are real.
Aperivue/medsci-skills
A skill your agent uses when a clinical CSV/Excel dataset needs profiling and cleaning before analysis (missing values, outliers, duplicates, type mismatches).
Aperivue/medsci-skills
A skill your agent uses when checking a radiology or medical AI study design before drafting or submission.
Categories
A skill your agent uses when you must prove an analysis ran on the intended data or lock a dataset version. Version Dataset is an agent skill from Aperivue/medsci-skills. Use when you must prove an analysis ran on the intended data or lock a dataset version.
Version Dataset fits situations like: you must prove an analysis ran on the intended data; lock a dataset version.
Run `npx skills add Aperivue/medsci-skills --skill version-dataset -a claude-code`. Or copy the skill folder (skills/version-dataset in Aperivue/medsci-skills) into .claude/skills/version-dataset in your project. Claude Code loads it when a task matches its description.
Run `npx skills add Aperivue/medsci-skills --skill version-dataset -a codex`. Or copy the skill folder (skills/version-dataset in Aperivue/medsci-skills) into .agents/skills/version-dataset in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Aperivue/medsci-skills --skill version-dataset -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/version-dataset, .gemini/skills/version-dataset, .github/skills/version-dataset and .opencode/skills/version-dataset in your project.
Going by SKILL.md and its folder, Version Dataset needs Python and a shell for the scripts in its folder and the command-line tools its instructions call (python). Our summary lists: Python 3; A Bash shell.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Version Dataset is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.1k tokens (SKILL.md is roughly 4.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.4k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Version Dataset: Hypothesis Generation (spacering-net/codeg, 3.8k stars), GitHub Deep Research (bytedance/deer-flow, 83k stars), Nature Paper Card (Yuan1z0825/nature-skills, 46k stars) and Read arXiv Paper (karpathy/nanochat, 58k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
Aperivue (a GitHub organization) maintains it in Aperivue/medsci-skills, which has 329 GitHub stars. The repository holds 54 skills in this directory. The repository was last updated on October 5, 2026.
Source: Aperivue/medsci-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.