---
name: generate-codebook
description: Use when a tabular dataset (CSV, Excel, Parquet, Stata, SAS) needs a data dictionary. Profiles every variable (type, levels, range, missingness) into codebook.md and codebook.json and flags coded values of unknown meaning as [NEEDS DICTIONARY] instead of guessing.
metadata:
  triggers: "generate codebook, data dictionary, codebook, profile variables, variable dictionary, describe dataset, what variables, column dictionary, build codebook"
---

# Generate Codebook Skill

Turn a raw tabular dataset into a structured, **citable** data dictionary (codebook). This is the
*generator* side of the dictionary-first workflow: it produces the artifact that `/define-variables`
and dictionary-first QC later consume. Distributions, types, and missingness are observable and the
bundled script profiles them; the **meaning** of a coded value (`fatty_liver_grade = 0`) is not
observable from the data and lives only in the authoritative data dictionary. You generate code and
review output — you do **not** invent the meaning of coded values.

## Deterministic Script

Run the bundled profiler rather than describing columns from memory:

```bash
python "${CLAUDE_SKILL_DIR}/scripts/generate_codebook.py" data.csv --out-dir .
```

Supports `.csv/.tsv/.xlsx/.parquet/.dta/.sas7bdat`. Flags: `--max-levels N`
(categorical cutoff, default 20), `--json-only`, `--md-only`. The script is
pandas-only, runs locally, and never sends data anywhere.

## Workflow

### Step 1: Profile (deterministic)

Run `generate_codebook.py` on the dataset. It writes `codebook.json` (machine-
readable) and `codebook.md` (review table), reporting per variable: role
(id / continuous / categorical / binary / date / text), dtype, missingness,
unique count, level frequencies or quantile summary, and a `needs_dictionary` flag.

### Step 2: Review with the researcher (gate)

Read `${CLAUDE_SKILL_DIR}/references/codebook_schema.md` (the codebook.json schema, the
role-inference heuristics, the `needs_dictionary` rule) before interpreting the output. Present
`codebook.md` and walk the user through it. **Gate:** role inference is a heuristic, so the user
confirms the inferred roles (e.g., an integer-coded scale mis-read as continuous, or an id column).
Do not proceed to definition work until the user approves the role assignments.

### Step 3: Resolve [NEEDS DICTIONARY] items (gate)

For every variable flagged `needs_dictionary: true`, the level codes are
uninterpretable without the authoritative source. **Gate:** ask the user to
supply the meaning of each code from the real data dictionary (file/sheet/row),
or to confirm none exists. Fill `label`, `units`, and per-level meanings into the
codebook **only** from that source, citing file > sheet > row — never from inference. If the user
cannot supply it, leave the `[NEEDS DICTIONARY]` marker in place; do not erase it.

Example: in a cohort file, `sex` (levels `1/2`) and `fatty_liver_grade` (`0..4`) are flagged because
their levels are bare codes; `smoking_status` (`never/former/current`) is not. Never write
`sex: 1 = male` because "that is the usual coding" — if the dictionary is unavailable, the flag stays.

### Step 4: Hand off

The completed `codebook.json` becomes the input dictionary for `/define-variables`
(operationalization) and the citation source for dictionary-first QC. **Gate:**
confirm with the user that no `needs_dictionary` flags remain unresolved before
the codebook is treated as authoritative for downstream analysis.

Run `/deidentify` on the raw data before a codebook is shared externally. Cleaning or transforming
data is `/clean-data`.

## Known limits

- Codes longer than 3 characters that are not numbers (`C50.9`, `S001`) are not recognised as
  codes, so their column is not flagged `needs_dictionary`.
- An integer-coded column with more distinct values than `--max-levels` is classed `continuous`
  and is not flagged. The Step 2 role review is where such columns are caught.
- The same holds for a `text` column holding number strings: any number value
  (`45`, `88`) is read as a measurement, so integer codes exported as strings
  with more than `--max-levels` values are not flagged.

## Output Format

`codebook.json` (schema in references) and `codebook.md` (review table with a
"Columns requiring dictionary lookup" section). Summarize the counts
(rows, columns, `needs_dictionary_count`) in chat; do not paste the full JSON.
