---
name: imaging-data
description: Use when preparing a medical-imaging dataset (DICOM/NIfTI) for modelling. Profiles spacing, orientation, intensity, label integrity, foreground fraction and target volume, gates them against the plan, then plans and audits preprocessing and augmentation for leakage.
metadata:
  triggers: "profile dataset, dataset profile, EDA, exploratory data analysis, explore the data, what does the data look like, imaging dataset, NIfTI, voxel spacing, slice thickness, orientation, intensity distribution, Hounsfield, class imbalance, foreground fraction, label sanity, empty label, label QC, dataset QC, data audit, before training, target volume, organ volume, is my test set labelled, research direction, where do I start, profile imaging, preprocess imaging, preprocessing, data pipeline, DICOM, resample, spacing, intensity normalization, intensity normalisation, windowing, HU window, z-score, histogram matching, augmentation, augmentation plan, TorchIO, MONAI transforms, data leakage, normalization leakage, preprocessing manifest, fit on train, per-image normalization, patient-level split, slice-level leakage, imaging data prep"
---

# Imaging-Data Skill

The dataset decides more of a study than the architecture does, and it decides it first. Phases 1–3
establish what the data is and what it will not support, while that is still cheap; Phases 4–7 design
and audit the preparation pipeline so it is leakage-safe before `/model-scaffold` builds the repo.
Describe-and-audit only: never modify, resample, reorient, split or write image data, never run
preprocessing on real patient data, and wire MONAI / TorchIO transforms by reference rather than
writing a new normalisation or resampling implementation.

Elsewhere: tabular/clinical variables → `/generate-codebook`, `/clean-data`; auditing the split
table, held-out metrics, calibration, subgroup results → `/model-assessment`; choosing an
architecture → `/model-selection`; building the repo → `/model-scaffold`.

## Workflow

### Phase 1 — Profile every case
```bash
python3 ${CLAUDE_SKILL_DIR}/scripts/profile_imaging_dataset.py \
    --split train:imagesTr:labelsTr \
    --split test:imagesTs \
    --dataset "MSD Task09 Spleen" \
    --declared-labels 0=background,1=spleen \
    --target-label 1 \
    --plan resample=true,reorient=false,loss=dice_ce,metrics=dice+hd95 \
    --out eda/profile.json
```
One record per case: grid, spacing, orientation, intensity percentiles, the label values actually
present, foreground fraction, and target volume in mL. A `--split` given no label directory is
recorded as **unlabelled** — itself a finding. Requires `nibabel` + `numpy`; the gate does not.
Every profile figure comes from opening the files — never from a dataset's README, a similar dataset,
or memory. A README can be wrong about its own label indices; the labels cannot.

**`--target-label` on a multi-structure atlas.** Foreground defaults to every non-zero index — the
whole annotated anatomy. Measured on the AMOS22 CT cases, that pools to 3.2 % instead of the spleen's 0.20 %, so the
pooled figure sits above the 1 % imbalance threshold while the target sits far below it and the
imbalance verdicts go quiet exactly where the risk is. Naming the target also makes `LABEL_EMPTY` mean
*this case has no spleen*. Pass `--target-label all` for a genuinely multi-class study; leave it out
on a multi-structure atlas and the gate raises `TARGET_LABEL_UNDECLARED`.

### Phase 2 — Gate the profile against the declared plan
```bash
python3 ${CLAUDE_SKILL_DIR}/scripts/check_dataset_profile.py --profile eda/profile.json \
    --out qc/dataset_profile.json --strict
```
Stdlib-only, so the audit re-runs anywhere the JSON travels. Never report a profile "pass" without
running it.

| Verdict | Severity | Fires when |
|---|---|---|
| `LABEL_SHAPE_MISMATCH` | Major | label grid ≠ image grid |
| `LABEL_EMPTY` | Major | a labelled case has zero foreground |
| `LABEL_VALUE_UNEXPECTED` | Major | label values outside the declared set |
| `TEST_SET_UNLABELLED` | Major | a split named test/held-out/external/eval carries no labels |
| `ACCURACY_UNDER_IMBALANCE` | Major | accuracy is planned while the target is a sliver of the volume |
| `LABEL_MISSING` | Minor | a case in a labelled split has no label file |
| `SPACING_HETEROGENEOUS` | Minor | spacing spans ≥ ratio on an axis and no resampling is declared |
| `ORIENTATION_MIXED` | Minor | >1 orientation code and no reorientation declared |
| `INTENSITY_SCALE_INCONSISTENT` | Minor | some cases on the HU scale, others not |
| `EXTREME_IMBALANCE` | Minor | median foreground below the threshold with no Dice-family loss |
| `TARGET_LABEL_UNDECLARED` | Minor | >1 structure declared, no target named, so foreground pools them all |

The gate flags an **undeclared decision, not variability**: 5× spacing spread and two orientation
codes pass once resampling and reorientation are declared (the clean challenge fixture proves this).
`--spacing-ratio` (default 2.0) and `--imbalance-frac` (default 0.01) are **screening defaults, not
published cut-points** — never present them as such; the values applied are printed in the output
and belong in the Methods. A split the profile shows unlabelled is never a held-out test set, however
the directory is named.

### Phase 3 — Turn the profile into research decisions
Write these decision notes into the study record, so `/design-study`, the preparation phases below,
and `/write-paper` inherit them instead of re-deriving them:
1. **Resampling target** — from the spacing distribution, not a tutorial default (carried into Phase 5).
2. **Loss and metric family** — from the foreground fraction. Segmentation reports Dice **and** a
   boundary metric per structure (`/model-assessment`); accuracy is not on the list.
3. **Pre-specified subgroups** — from the clinical spread the profile shows (target volume, slice
   thickness, modality). Pre-specifying them here is what separates a subgroup finding from a post-hoc one.
4. **Where the held-out set comes from** — especially when the shipped "test" directory is unlabelled.
5. **What the cohort cannot support** — n, single-source acquisition, absent subgroups: the seed of
   the Limitations paragraph, written before results can bias it.

### Phase 4 — Inventory the preparation steps and fix fit scope
Collect the modality, the data manifest (one row per image/slice with a `patient_id`), the resample
spacing, the intensity transform (fixed HU window vs a fitted z-score / min-max / histogram match),
and the augmentation plan. Read `${CLAUDE_SKILL_DIR}/references/preprocessing_guide.md` for the
modality-aware normalisation, physiology-preserving vs -breaking augmentation, and MONAI / TorchIO wiring.
- Fit dataset-level normalisation on the **training split only** — never all/full/test.
- Run any data-fitted transform **after** the split; before it there is no train/test distinction.
- Prefer per-image (per-sample) normalisation where clinically appropriate — leakage-free even before the split.
- Keep augmentation **train-only**; augmenting val/test folds undisclosed test-time augmentation into the metric.
- Split at the **patient** level, then map slices to their patient's split.

### Phase 5 — Emit the preprocessing manifest
Write `preprocessing_manifest.json`, which `/model-scaffold` consumes and the gate checks. Every value
comes from the real data manifest and the declared pipeline — never invented patient IDs or split
assignments.

```json
{
  "split_seed": 42,
  "transforms": [
    {"name": "hu_window", "type": "clip", "fit_scope": "none", "stage": "before_split"},
    {"name": "train_zscore", "type": "standardize", "fit_scope": "train", "stage": "after_split"},
    {"name": "flip_rotate", "type": "augmentation", "stage": "after_split", "applies_to": ["train"]}
  ],
  "split_assignment": [
    {"patient_id": "P001", "unit_id": "P001_s1", "split": "train"}
  ]
}
```

`fit_scope`: `train` (OK) · `all`/`full`/`dataset`/`test` (leak) · `sample`/`per_image`/`none`/`fixed`
(not data-fitted). `stage`: `before_split` / `after_split`. The fields must describe what the code
actually does — never tag a dataset-fitted transform per-sample to clear the gate; that hides the leak.
A declared dataset-level `fit_scope` is judged whatever the `type` is, so a library class name
(`HistogramStandardization`, `NormalizeIntensityd`) fit on `all` is a leak like `standardize` would be.

**Declare the fit scope of resampling too.** A target spacing chosen in advance is `fit_scope: fixed`
and never leaks. A target derived from the cohort does: nnU-Net sets its target spacing from a
percentile of the dataset fingerprint, so a resample fitted over every case carries held-out geometry
into the training grid exactly as an intensity statistic would. The fingerprint's scope decides which
you have, not the word "resample".

### Phase 6 — Gate the manifest
```bash
python3 ${CLAUDE_SKILL_DIR}/scripts/check_preprocessing_leakage.py --manifest preprocessing_manifest.json \
    --out qc/preprocessing_leakage.json --strict
```
Verdicts: `PREPROCESS_BEFORE_SPLIT`, `NORMALIZATION_LEAKAGE`, `PATIENT_CROSS_SPLIT` (Major);
`AUGMENTATION_ON_EVAL`, `UNSPECIFIED_FIT_SCOPE`, `MISSING_SEED` (Minor), reproduced by set arithmetic
and rule on the manifest. A green gate is the precondition for handing the manifest to
`/model-scaffold`; its `split_assignment` is the same patient-level split `/model-assessment` later
re-verifies. Never report a pass without running it.

### Phase 7 — Before inference on a new cohort: check the normaliser's domain
Phase 6 asks whether a transform was fit on the right **scope**. Before running a trained model on a
cohort it was not trained on, ask whether that cohort sits in the intensity **domain** the trained
normaliser assumes:

```bash
python3 ${CLAUDE_SKILL_DIR}/scripts/check_normalizer_domain.py \
    --profile eda/<cohort>_profile.json \
    --contract work/nnUNet_results/.../plans.json \
    --splits external_mri --out qc/normalizer_domain.json --strict
```
Its challenge card holds a cohort in the contract's own domain that must come back clean, an
arbitrary-unit cohort that must raise a Major, and an unreadable contract that must refuse rather than pass.

## Outputs and hand-off

- `eda/profile.json` (Phase 1), `qc/dataset_profile.json` (Phase 2), and the decision notes (Phase 3).
- `preprocessing_manifest.json` with the augmentation-appropriateness and normalisation fit-scope
  notes (Phases 4–5), `qc/preprocessing_leakage.json` (Phase 6), `qc/normalizer_domain.json` (Phase 7).

The manifest feeds `/model-scaffold`, which also **reads the `qc/` reports**: keep them in `qc/` beside
the manifest (or `../qc/`), or pass them with `--imaging-qc`. An unresolved Major there refuses the
scaffold until it is fixed and re-gated or acknowledged with a stated reason (`--ack-qc`); Minor and
Flag claims are carried into the repo's `IMAGING_QC.md`, and a missing report is recorded as not
assessed. Re-run a gate after fixing its finding — a stale report still blocks. The manifest documents the CLAIM 2024 / TRIPOD+AI data-preprocessing items
for `/check-reporting`; `/self-review`'s `model_development` probe looks for exactly this pipeline in a
finished manuscript. Regression: `bash ${CLAUDE_SKILL_DIR}/scripts/check_dataset_profile_challenge/verify.sh`,
`bash ${CLAUDE_SKILL_DIR}/scripts/check_preprocessing_leakage_challenge/verify.sh`,
`bash ${CLAUDE_SKILL_DIR}/scripts/check_normalizer_domain_challenge/verify.sh`,
`bash ${CLAUDE_SKILL_DIR}/tests/test_dataset_profile.sh`,
`bash ${CLAUDE_SKILL_DIR}/tests/test_preprocessing_leakage.sh`.
