Agent skill

Generate Codebook

by Aperivue in Aperivue/medsci-skills

A skill your agent uses when a tabular dataset (CSV, Excel, Parquet, Stata, SAS) needs a data dictionary.

MITAuto-check passedDocuments & Office

Install Generate Codebook

skills CLI
$ npx skills add Aperivue/medsci-skills --skill generate-codebook -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Aperivue/medsci-skills generate-codebook --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Aperivue/medsci-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/generate-codebook .claude/skills/generate-codebook && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
generate-codebook
GitHub stars
329
Token cost
~1.1k tokens
SKILL.md length
527 words
Files
5 (incl. scripts, references)
Skills in repo
54
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when a tabular dataset (CSV, Excel, Parquet, Stata, SAS) needs a data dictionary.

  • Works in 4 steps: Profile (deterministic) → Review with the researcher (gate) → Resolve [NEEDS DICTIONARY] items (gate) → …
  • A tabular dataset (CSV
  • SKILL.md covers Deterministic Script, Workflow, Known limits and Output Format
  • Runs Python and Shell scripts from its folder; calls python

What it does

Generate Codebook is an agent skill from Aperivue/medsci-skills. Use when a tabular dataset (CSV, Excel, Parquet, Stata, SAS) needs a data dictionary. Profiles every variable (type, levels, range, missingness) into codebook.md and codebook.json and flags coded values of unknown meaning as [NEEDS DICTIONARY] instead of guessing.

Its SKILL.md is about 1.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 7 other files, including scripts and reference files (for example `references/codebook_schema.md`, `scripts/generate_codebook.py` and `skill.yml`).

It sits in Documents & Office, covering DataFrames, Excel spreadsheets and Econometrics and empirical research. It works with Microsoft Excel. The repository describes itself as: Agent Skills for medical research — literature search, reporting-guideline & citation checks, statistics, publication figures, submission. Works with Claude Code, Codex, Cursor &… The licence is MIT.

When your agent uses it

  • A tabular dataset (CSV
  • SAS) needs a data dictionary

Example prompts

  • “/generate-codebook”

Requirements

  • Python 3
  • A Bash shell

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Profile (deterministic)
  2. Review with the researcher (gate)
  3. Resolve [NEEDS DICTIONARY] items (gate)
  4. Hand off

What it can do on your machine

Read from SKILL.md and the folder at commit 3b14ae2. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python and Shell), which the agent can run.

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Generate Codebook loads about 1.1k tokens when it runs, and up to ~2.1k if it reads all its reference files. Until then it costs about 71 tokens; SKILL.md has 527 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~71
When it runs · the whole SKILL.md, loaded when a task matches
~1.1k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~2.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from Aperivue/medsci-skills at commit 3b14ae2, republished under its MIT licence (© Aperivue). 527 words, ~1,113 tokens.

Download SKILL.mdSave it as .claude/skills/generate-codebook/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
generate-codebook
description
Use when a tabular dataset (CSV, Excel, Parquet, Stata, SAS) needs a data dictionary. Profiles every variable (type, levels, range, missingness) into codebook.md and codebook.json and flags coded values of unknown meaning as [NEEDS DICTIONARY] instead of guessing.
metadata.triggers
generate codebook, data dictionary, codebook, profile variables, variable dictionary, describe dataset, what variables, column dictionary, build codebook

Generate Codebook Skill

Turn a raw tabular dataset into a structured, citable data dictionary (codebook). This is the generator side of the dictionary-first workflow: it produces the artifact that /define-variables and dictionary-first QC later consume. Distributions, types, and missingness are observable and the bundled script profiles them; the meaning of a coded value (fatty_liver_grade = 0) is not observable from the data and lives only in the authoritative data dictionary. You generate code and review output — you do not invent the meaning of coded values.

Deterministic Script

Run the bundled profiler rather than describing columns from memory:

bash
python "${CLAUDE_SKILL_DIR}/scripts/generate_codebook.py" data.csv --out-dir .

Supports .csv/.tsv/.xlsx/.parquet/.dta/.sas7bdat. Flags: --max-levels N (categorical cutoff, default 20), --json-only, --md-only. The script is pandas-only, runs locally, and never sends data anywhere.

Workflow

Step 1: Profile (deterministic)

Run generate_codebook.py on the dataset. It writes codebook.json (machine- readable) and codebook.md (review table), reporting per variable: role (id / continuous / categorical / binary / date / text), dtype, missingness, unique count, level frequencies or quantile summary, and a needs_dictionary flag.

Step 2: Review with the researcher (gate)

Read ${CLAUDE_SKILL_DIR}/references/codebook_schema.md (the codebook.json schema, the role-inference heuristics, the needs_dictionary rule) before interpreting the output. Present codebook.md and walk the user through it. Gate: role inference is a heuristic, so the user confirms the inferred roles (e.g., an integer-coded scale mis-read as continuous, or an id column). Do not proceed to definition work until the user approves the role assignments.

Step 3: Resolve [NEEDS DICTIONARY] items (gate)

For every variable flagged needs_dictionary: true, the level codes are uninterpretable without the authoritative source. Gate: ask the user to supply the meaning of each code from the real data dictionary (file/sheet/row), or to confirm none exists. Fill label, units, and per-level meanings into the codebook only from that source, citing file > sheet > row — never from inference. If the user cannot supply it, leave the [NEEDS DICTIONARY] marker in place; do not erase it.

Example: in a cohort file, sex (levels 1/2) and fatty_liver_grade (0..4) are flagged because their levels are bare codes; smoking_status (never/former/current) is not. Never write sex: 1 = male because "that is the usual coding" — if the dictionary is unavailable, the flag stays.

Show full SKILL.md (177 more words)Show less
Step 4: Hand off

The completed codebook.json becomes the input dictionary for /define-variables (operationalization) and the citation source for dictionary-first QC. Gate: confirm with the user that no needs_dictionary flags remain unresolved before the codebook is treated as authoritative for downstream analysis.

Run /deidentify on the raw data before a codebook is shared externally. Cleaning or transforming data is /clean-data.

Known limits

  • Codes longer than 3 characters that are not numbers (C50.9, S001) are not recognised as codes, so their column is not flagged needs_dictionary.
  • An integer-coded column with more distinct values than --max-levels is classed continuous and is not flagged. The Step 2 role review is where such columns are caught.
  • The same holds for a text column holding number strings: any number value (45, 88) is read as a measurement, so integer codes exported as strings with more than --max-levels values are not flagged.

Output Format

codebook.json (schema in references) and codebook.md (review table with a "Columns requiring dictionary lookup" section). Summarize the counts (rows, columns, needs_dictionary_count) in chat; do not paste the full JSON.

© Aperivue, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files (scripts, references) in skills/generate-codebook of Aperivue/medsci-skills.

  • SKILL.md
  • references/codebook_schema.md
  • scripts/generate_codebook.py
  • skill.yml
  • tests/test_generate_codebook.sh

Open the folder on GitHubat commit 3b14ae2

Compare with similar skills

Generate Codebook next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Generate Codebook compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Generate Codebook this skillAperivue/medsci-skills329—~1.1kAutomated safety check: PassMIT
Convert Fileduckdb/duckdb-skills5991 repos~720Automated safety check: NotesMIT
Excel ParserHarryoung/efka104—~2.3kAutomated safety check: PassApache-2.0
Sn Da Image CaptionMichaelYang-lyx/AIDABench1111 repos~2kAutomated safety check: PassNone
Tabular Cleanupgaasher/Agent-Loop-Skills174—~4kAutomated safety check: PassMIT
Matlab Import Export Datamatlab/matlab-agentic-toolkit1.1k—~3.6kAutomated safety check: PassCustom licence

Similar skills

  • Convert File

    duckdb/duckdb-skills

    Official

    Convert any data file to another format: CSV, Parquet, JSON, Excel, GeoJSON, and more.

    599 GitHub starsUsed in 1 repo~720 tokens
    Documents & OfficeAuto-check: notes
  • Excel Parser

    Harryoung/efka

    Smart Excel/CSV file parsing with intelligent routing based on file complexity analysis.

    104 GitHub stars~2.3k tokensUpdated 6 mo ago
    Documents & OfficeAuto-check passed
  • Sn Da Image Caption

    MichaelYang-lyx/AIDABench

    图片理解与数据提取 skill。当图片文件(.png/.jpg/.jpeg/.gif/.webp/.bmp)是主要输入且用户需要理解、提取数据或分析图片内容时使用。提供预配置的 caption 脚本(scripts/caption.py),通过 vision 模型将图片转为文本描述,无需额外配置 API Key。覆盖:(1) 通过 scripts/caption.py…

    111 GitHub starsUsed in 1 repo~2k tokens
    Documents & OfficeAuto-check passed
  • Tabular Cleanup

    gaasher/Agent-Loop-Skills

    A skill your agent uses when the user has a messy tabular data dump (CSV/TSV/parquet/Excel/JSON) and wants it iteratively cleaned to an inferred data contract — a checklist of deterministic…

    174 GitHub stars~4k tokensUpdated 3 mo ago
    Documents & OfficeAuto-check passed
  • Matlab Import Export Data

    matlab/matlab-agentic-toolkit

    Read or write data files in MATLAB. An agent skill from matlab/matlab-agentic-toolkit.

    1.1k GitHub stars~3.6k tokensUpdated 7 days ago
    Documents & OfficeAuto-check passed
  • Codebook

    brycewang-stanford/Auto-Empirical-Research-Skills

    Auto-generates a Markdown codebook from a dataset (CSV, DTA, Excel, Parquet) with types and summary statistics.

    4.5k GitHub stars~527 tokensUpdated 2 days ago
    Documents & OfficeAuto-check: notes

More from Aperivue/medsci-skills

All 54 skills in this repo
  • Model Assessment

    Aperivue/medsci-skills

    A skill your agent uses when validating or evaluating a trained medical-imaging model.

    329 GitHub starsUsed in 1 repo~4.5k tokens
    Auto-check passed
  • Obsidian Paper Vault

    Aperivue/medsci-skills

    A skill your agent uses when turning a folder of research PDFs into Obsidian notes, even if Obsidian is not named.

    329 GitHub stars~1.6k tokensUpdated 2 days ago
    Auto-check passed
  • Radiomics ML

    Aperivue/medsci-skills

    A skill your agent uses when building or auditing a radiomics or tabular clinical-ML prediction model with a classical learner (LASSO, SVM, random forest, XGBoost and similar).

    329 GitHub starsUsed in 1 repo~2.7k tokens
    Auto-check passed
  • Clean Data

    Aperivue/medsci-skills

    A skill your agent uses when a clinical CSV/Excel dataset needs profiling and cleaning before analysis (missing values, outliers, duplicates, type mismatches).

    329 GitHub stars~2k tokensUpdated 2 days ago
    Auto-check passed
  • Design Study

    Aperivue/medsci-skills

    A skill your agent uses when checking a radiology or medical AI study design before drafting or submission.

    329 GitHub stars~3.9k tokensUpdated 2 days ago
    Auto-check passed
  • Fill Icmje Coi

    Aperivue/medsci-skills

    A skill your agent uses when each author needs an ICMJE Conflict of Interest disclosure form (coidisclosure.docx) for submission.

    329 GitHub stars~1.5k tokensUpdated 2 days ago
    Auto-check passed

Works with

Questions about Generate Codebook

What does Generate Codebook do?

A skill your agent uses when a tabular dataset (CSV, Excel, Parquet, Stata, SAS) needs a data dictionary. Generate Codebook is an agent skill from Aperivue/medsci-skills. Use when a tabular dataset (CSV, Excel, Parquet, Stata, SAS) needs a data dictionary.

When should I use Generate Codebook?

Generate Codebook fits situations like: A tabular dataset (CSV; SAS) needs a data dictionary.

How do I install Generate Codebook in Claude Code?

Run `npx skills add Aperivue/medsci-skills --skill generate-codebook -a claude-code`. Or copy the skill folder (skills/generate-codebook in Aperivue/medsci-skills) into .claude/skills/generate-codebook in your project. Claude Code loads it when a task matches its description.

How do I install Generate Codebook in Codex?

Run `npx skills add Aperivue/medsci-skills --skill generate-codebook -a codex`. Or copy the skill folder (skills/generate-codebook in Aperivue/medsci-skills) into .agents/skills/generate-codebook in your project. Codex loads it when a task matches its description.

Can I use Generate Codebook in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Aperivue/medsci-skills --skill generate-codebook -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/generate-codebook, .gemini/skills/generate-codebook, .github/skills/generate-codebook and .opencode/skills/generate-codebook in your project.

What does Generate Codebook need to run?

Going by SKILL.md and its folder, Generate Codebook needs Python and a shell for the scripts in its folder and the command-line tools its instructions call (python). Our summary lists: Python 3; A Bash shell.

Does Generate Codebook access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Generate Codebook safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Generate Codebook use?

Generate Codebook is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Generate Codebook use?

About 1.1k tokens (SKILL.md is roughly 4.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 995 tokens, read only when the agent opens those files.

What are the alternatives to Generate Codebook?

Skills that share tags, products or a category with Generate Codebook: Convert File (duckdb/duckdb-skills, 599 stars), Excel Parser (Harryoung/efka, 104 stars), Sn Da Image Caption (MichaelYang-lyx/AIDABench, 111 stars) and Tabular Cleanup (gaasher/Agent-Loop-Skills, 174 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Generate Codebook?

Aperivue (a GitHub organization) maintains it in Aperivue/medsci-skills, which has 329 GitHub stars. The repository holds 54 skills in this directory. The repository was last updated on October 5, 2026.

Source: Aperivue/medsci-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.