Agent skill

Imaging Data

by Aperivue in Aperivue/medsci-skills

A skill your agent uses when preparing a medical-imaging dataset (DICOM/NIfTI) for modelling.

MITAuto-check passedResearch & Science

Install Imaging Data

skills CLI
$ npx skills add Aperivue/medsci-skills --skill imaging-data -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Aperivue/medsci-skills imaging-data --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Aperivue/medsci-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/imaging-data .claude/skills/imaging-data && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
imaging-data
GitHub stars
331
Token cost
~2.9k tokens
SKILL.md length
1,190 words
Files
32 (incl. scripts, references)
Skills in repo
54
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when preparing a medical-imaging dataset (DICOM/NIfTI) for modelling.

  • Works in 7 steps: Profile every case → Gate the profile against the declared plan → Turn the profile into research decisions → …
  • Preparing a medical-imaging dataset (DICOM/NIfTI) for modelling
  • SKILL.md covers Workflow and Outputs and hand-off
  • Runs Python and Shell scripts from its folder; calls bash and python3

What it does

Imaging Data is an agent skill from Aperivue/medsci-skills. Use when preparing a medical-imaging dataset (DICOM/NIfTI) for modelling. Profiles spacing, orientation, intensity, label integrity, foreground fraction and target volume, gates them against the plan, then plans and audits preprocessing and augmentation for leakage.

Its SKILL.md is about 2.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 39 other files, including scripts and reference files (for example `references/preprocessing_guide.md`, `scripts/check_dataset_profile.py` and `scripts/check_dataset_profile_challenge/fixture/profile_clean.json`).

It sits in Research & Science, covering Clinical and healthcare research. The repository describes itself as: Agent Skills for medical research — literature search, reporting-guideline & citation checks, statistics, publication figures, submission. Works with Claude Code, Codex, Cursor &… The licence is MIT.

When your agent uses it

  • Preparing a medical-imaging dataset (DICOM/NIfTI) for modelling
  • Tasks that involve Clinical and healthcare research

Example prompts

  • “/imaging-data”

Requirements

  • Python 3
  • A Bash shell

Workflow steps

7 steps, taken from the step headings in SKILL.md.

  1. Profile every case
  2. Gate the profile against the declared plan
  3. Turn the profile into research decisions
  4. Inventory the preparation steps and fix fit scope
  5. Emit the preprocessing manifest
  6. Gate the manifest
  7. Before inference on a new cohort: check the normaliser's domain

What it can do on your machine

Read from SKILL.md and the folder at commit 3b14ae2. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 11 files in scripts/ (Python and Shell, from the files we listed), which the agent can run.

    Shell commands in SKILL.md call:

    • bash
    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Imaging Data loads about 2.9k tokens when it runs, and up to ~4.2k if it reads all its reference files. Until then it costs about 70 tokens; SKILL.md has 1,190 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~70
When it runs · the whole SKILL.md, loaded when a task matches
~2.9k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~4.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from Aperivue/medsci-skills at commit 3b14ae2, republished under its MIT licence (© Aperivue). 1,190 words, ~2,915 tokens.

Download SKILL.mdSave it as .claude/skills/imaging-data/SKILL.md (or your agent's skills folder). This skill also uses 31 other files; get the full folder from GitHub.
name
imaging-data
description
Use when preparing a medical-imaging dataset (DICOM/NIfTI) for modelling. Profiles spacing, orientation, intensity, label integrity, foreground fraction and target volume, gates them against the plan, then plans and audits preprocessing and augmentation for leakage.
metadata.triggers
profile dataset, dataset profile, EDA, exploratory data analysis, explore the data, what does the data look like, imaging dataset, NIfTI, voxel spacing, slice…

Imaging-Data Skill

The dataset decides more of a study than the architecture does, and it decides it first. Phases 1–3 establish what the data is and what it will not support, while that is still cheap; Phases 4–7 design and audit the preparation pipeline so it is leakage-safe before /model-scaffold builds the repo. Describe-and-audit only: never modify, resample, reorient, split or write image data, never run preprocessing on real patient data, and wire MONAI / TorchIO transforms by reference rather than writing a new normalisation or resampling implementation.

Elsewhere: tabular/clinical variables → /generate-codebook, /clean-data; auditing the split table, held-out metrics, calibration, subgroup results → /model-assessment; choosing an architecture → /model-selection; building the repo → /model-scaffold.

Workflow

Phase 1 — Profile every case
bash
python3 ${CLAUDE_SKILL_DIR}/scripts/profile_imaging_dataset.py \
    --split train:imagesTr:labelsTr \
    --split test:imagesTs \
    --dataset "MSD Task09 Spleen" \
    --declared-labels 0=background,1=spleen \
    --target-label 1 \
    --plan resample=true,reorient=false,loss=dice_ce,metrics=dice+hd95 \
    --out eda/profile.json

One record per case: grid, spacing, orientation, intensity percentiles, the label values actually present, foreground fraction, and target volume in mL. A --split given no label directory is recorded as unlabelled — itself a finding. Requires nibabel + numpy; the gate does not. Every profile figure comes from opening the files — never from a dataset's README, a similar dataset, or memory. A README can be wrong about its own label indices; the labels cannot.

--target-label on a multi-structure atlas. Foreground defaults to every non-zero index — the whole annotated anatomy. Measured on the AMOS22 CT cases, that pools to 3.2 % instead of the spleen's 0.20 %, so the pooled figure sits above the 1 % imbalance threshold while the target sits far below it and the imbalance verdicts go quiet exactly where the risk is. Naming the target also makes LABEL_EMPTY mean this case has no spleen. Pass --target-label all for a genuinely multi-class study; leave it out on a multi-structure atlas and the gate raises TARGET_LABEL_UNDECLARED.

Phase 2 — Gate the profile against the declared plan
bash
python3 ${CLAUDE_SKILL_DIR}/scripts/check_dataset_profile.py --profile eda/profile.json \
    --out qc/dataset_profile.json --strict

Stdlib-only, so the audit re-runs anywhere the JSON travels. Never report a profile "pass" without running it.

VerdictSeverityFires when
LABEL_SHAPE_MISMATCHMajorlabel grid ≠ image grid
LABEL_EMPTYMajora labelled case has zero foreground
LABEL_VALUE_UNEXPECTEDMajorlabel values outside the declared set
TEST_SET_UNLABELLEDMajora split named test/held-out/external/eval carries no labels
ACCURACY_UNDER_IMBALANCEMajoraccuracy is planned while the target is a sliver of the volume
LABEL_MISSINGMinora case in a labelled split has no label file
SPACING_HETEROGENEOUSMinorspacing spans ≥ ratio on an axis and no resampling is declared
ORIENTATION_MIXEDMinor>1 orientation code and no reorientation declared
INTENSITY_SCALE_INCONSISTENTMinorsome cases on the HU scale, others not
EXTREME_IMBALANCEMinormedian foreground below the threshold with no Dice-family loss
TARGET_LABEL_UNDECLAREDMinor>1 structure declared, no target named, so foreground pools them all

The gate flags an undeclared decision, not variability: 5× spacing spread and two orientation codes pass once resampling and reorientation are declared (the clean challenge fixture proves this). --spacing-ratio (default 2.0) and --imbalance-frac (default 0.01) are screening defaults, not published cut-points — never present them as such; the values applied are printed in the output and belong in the Methods. A split the profile shows unlabelled is never a held-out test set, however the directory is named.

Phase 3 — Turn the profile into research decisions

Write these decision notes into the study record, so /design-study, the preparation phases below, and /write-paper inherit them instead of re-deriving them:

  1. Resampling target — from the spacing distribution, not a tutorial default (carried into Phase 5).
  2. Loss and metric family — from the foreground fraction. Segmentation reports Dice and a boundary metric per structure (/model-assessment); accuracy is not on the list.
  3. Pre-specified subgroups — from the clinical spread the profile shows (target volume, slice thickness, modality). Pre-specifying them here is what separates a subgroup finding from a post-hoc one.
  4. Where the held-out set comes from — especially when the shipped "test" directory is unlabelled.
  5. What the cohort cannot support — n, single-source acquisition, absent subgroups: the seed of the Limitations paragraph, written before results can bias it.
Phase 4 — Inventory the preparation steps and fix fit scope

Collect the modality, the data manifest (one row per image/slice with a patient_id), the resample spacing, the intensity transform (fixed HU window vs a fitted z-score / min-max / histogram match), and the augmentation plan. Read ${CLAUDE_SKILL_DIR}/references/preprocessing_guide.md for the modality-aware normalisation, physiology-preserving vs -breaking augmentation, and MONAI / TorchIO wiring.

  • Fit dataset-level normalisation on the training split only — never all/full/test.
  • Run any data-fitted transform after the split; before it there is no train/test distinction.
  • Prefer per-image (per-sample) normalisation where clinically appropriate — leakage-free even before the split.
  • Keep augmentation train-only; augmenting val/test folds undisclosed test-time augmentation into the metric.
  • Split at the patient level, then map slices to their patient's split.
Show full SKILL.md (440 more words)Show less
Phase 5 — Emit the preprocessing manifest

Write preprocessing_manifest.json, which /model-scaffold consumes and the gate checks. Every value comes from the real data manifest and the declared pipeline — never invented patient IDs or split assignments.

json
{
  "split_seed": 42,
  "transforms": [
    {"name": "hu_window", "type": "clip", "fit_scope": "none", "stage": "before_split"},
    {"name": "train_zscore", "type": "standardize", "fit_scope": "train", "stage": "after_split"},
    {"name": "flip_rotate", "type": "augmentation", "stage": "after_split", "applies_to": ["train"]}
  ],
  "split_assignment": [
    {"patient_id": "P001", "unit_id": "P001_s1", "split": "train"}
  ]
}

fit_scope: train (OK) · all/full/dataset/test (leak) · sample/per_image/none/fixed (not data-fitted). stage: before_split / after_split. The fields must describe what the code actually does — never tag a dataset-fitted transform per-sample to clear the gate; that hides the leak. A declared dataset-level fit_scope is judged whatever the type is, so a library class name (HistogramStandardization, NormalizeIntensityd) fit on all is a leak like standardize would be.

Declare the fit scope of resampling too. A target spacing chosen in advance is fit_scope: fixed and never leaks. A target derived from the cohort does: nnU-Net sets its target spacing from a percentile of the dataset fingerprint, so a resample fitted over every case carries held-out geometry into the training grid exactly as an intensity statistic would. The fingerprint's scope decides which you have, not the word "resample".

Phase 6 — Gate the manifest
bash
python3 ${CLAUDE_SKILL_DIR}/scripts/check_preprocessing_leakage.py --manifest preprocessing_manifest.json \
    --out qc/preprocessing_leakage.json --strict

Verdicts: PREPROCESS_BEFORE_SPLIT, NORMALIZATION_LEAKAGE, PATIENT_CROSS_SPLIT (Major); AUGMENTATION_ON_EVAL, UNSPECIFIED_FIT_SCOPE, MISSING_SEED (Minor), reproduced by set arithmetic and rule on the manifest. A green gate is the precondition for handing the manifest to /model-scaffold; its split_assignment is the same patient-level split /model-assessment later re-verifies. Never report a pass without running it.

Phase 7 — Before inference on a new cohort: check the normaliser's domain

Phase 6 asks whether a transform was fit on the right scope. Before running a trained model on a cohort it was not trained on, ask whether that cohort sits in the intensity domain the trained normaliser assumes:

bash
python3 ${CLAUDE_SKILL_DIR}/scripts/check_normalizer_domain.py \
    --profile eda/<cohort>_profile.json \
    --contract work/nnUNet_results/.../plans.json \
    --splits external_mri --out qc/normalizer_domain.json --strict

Its challenge card holds a cohort in the contract's own domain that must come back clean, an arbitrary-unit cohort that must raise a Major, and an unreadable contract that must refuse rather than pass.

Outputs and hand-off

  • eda/profile.json (Phase 1), qc/dataset_profile.json (Phase 2), and the decision notes (Phase 3).
  • preprocessing_manifest.json with the augmentation-appropriateness and normalisation fit-scope notes (Phases 4–5), qc/preprocessing_leakage.json (Phase 6), qc/normalizer_domain.json (Phase 7).

The manifest feeds /model-scaffold, which also reads the qc/ reports: keep them in qc/ beside the manifest (or ../qc/), or pass them with --imaging-qc. An unresolved Major there refuses the scaffold until it is fixed and re-gated or acknowledged with a stated reason (--ack-qc); Minor and Flag claims are carried into the repo's IMAGING_QC.md, and a missing report is recorded as not assessed. Re-run a gate after fixing its finding — a stale report still blocks. The manifest documents the CLAIM 2024 / TRIPOD+AI data-preprocessing items for /check-reporting; /self-review's model_development probe looks for exactly this pipeline in a finished manuscript. Regression: bash ${CLAUDE_SKILL_DIR}/scripts/check_dataset_profile_challenge/verify.sh, bash ${CLAUDE_SKILL_DIR}/scripts/check_preprocessing_leakage_challenge/verify.sh, bash ${CLAUDE_SKILL_DIR}/scripts/check_normalizer_domain_challenge/verify.sh, bash ${CLAUDE_SKILL_DIR}/tests/test_dataset_profile.sh, bash ${CLAUDE_SKILL_DIR}/tests/test_preprocessing_leakage.sh.

© Aperivue, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 31 other files (scripts, references) in skills/imaging-data of Aperivue/medsci-skills.

  • SKILL.md
  • references/preprocessing_guide.md
  • scripts/check_dataset_profile.py
  • scripts/check_dataset_profile_challenge/expected/clean.txt
  • scripts/check_dataset_profile_challenge/expected/defect.txt
  • scripts/check_dataset_profile_challenge/fixture/profile_clean.json
  • scripts/check_dataset_profile_challenge/fixture/profile_defect.json
  • scripts/check_dataset_profile_challenge/problem.md
  • scripts/check_dataset_profile_challenge/verify.sh
  • scripts/check_normalizer_domain.py
  • scripts/check_normalizer_domain_challenge/expected/arbitrary_mismatch.txt
  • scripts/check_normalizer_domain_challenge/expected/hu_clean.txt
  • scripts/check_normalizer_domain_challenge/fixture/plan_ct.json
  • … and 19 more

Open the folder on GitHubat commit 3b14ae2

Compare with similar skills

Imaging Data next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Imaging Data compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Imaging Data this skillAperivue/medsci-skills331—~2.9kAutomated safety check: PassMIT
Clinical Trials Databasegoogle-deepmind/science-skills3.2k2 repos~3.2kAutomated safety check: PassApache-2.0
CHARLS Paper Reproduction Guidexjtulyc/MedgeClaw6171 repos~1.8kAutomated safety check: PassNone
Biomedical Analysis Dispatchxjtulyc/MedgeClaw6171 repos~2kAutomated safety check: PassNone
Research Proposalluwill/research-skills858—~4.4kAutomated safety check: NotesNone
Medical Imaging ReviewLeonChaoX/qinyan-academic-skills9433 repos~1.1kAutomated safety check: NotesMIT

Similar skills

  • Clinical Trials Database

    google-deepmind/science-skills

    Query ClinicalTrials.gov via APIv2. An agent skill from google-deepmind/science-skills.

    3.2k GitHub starsUsed in 2 repos~3.2k tokens
    Research & ScienceAuto-check passed
  • Guides an agent through reproducing papers built on the CHARLS health and retirement survey, from variable mapping to cognition, depression and isolation scores.

    617 GitHub starsUsed in 1 repo~1.8k tokens
    Research & ScienceAuto-check passed
  • Routes bioinformatics, drug discovery, clinical and multi-omics tasks from a chat interface to Claude Code sessions running K-Dense scientific skills, with a live dashboard per task.

    617 GitHub starsUsed in 1 repo~2k tokens
    Research & ScienceAuto-check passed
  • Research Proposal

    luwill/research-skills

    A skill your agent uses when the user asks to write or draft a PhD / doctoral research proposal, research plan, 研究计划书, or 开题报告 — a forward-looking plan of background, gap, research questions…

    858 GitHub stars~4.4k tokensUpdated 9 days ago
    Research & ScienceAuto-check: notes
  • Medical Imaging Review

    LeonChaoX/qinyan-academic-skills

    Write comprehensive literature reviews for medical imaging AI research.

    943 GitHub starsUsed in 3 repos~1.1k tokens
    Research & ScienceAuto-check: notes
  • Drug Research

    lamm-mit/scienceclaw

    Generates comprehensive drug research reports with compound disambiguation, evidence grading, and mandatory completeness sections.

    244 GitHub starsUsed in 3 repos~1.7k tokens
    Research & ScienceAuto-check passed

More from Aperivue/medsci-skills

All 54 skills in this repo
  • Obsidian Paper Vault

    Aperivue/medsci-skills

    A skill your agent uses when turning a folder of research PDFs into Obsidian notes, even if Obsidian is not named.

    331 GitHub stars~1.6k tokensUpdated 4 days ago
    Auto-check passed
  • Clean Data

    Aperivue/medsci-skills

    A skill your agent uses when a clinical CSV/Excel dataset needs profiling and cleaning before analysis (missing values, outliers, duplicates, type mismatches).

    331 GitHub stars~2k tokensUpdated 4 days ago
    Auto-check passed
  • Design Study

    Aperivue/medsci-skills

    A skill your agent uses when checking a radiology or medical AI study design before drafting or submission.

    331 GitHub stars~3.9k tokensUpdated 4 days ago
    Auto-check passed
  • Fill Icmje Coi

    Aperivue/medsci-skills

    A skill your agent uses when each author needs an ICMJE Conflict of Interest disclosure form (coidisclosure.docx) for submission.

    331 GitHub stars~1.5k tokensUpdated 4 days ago
    Auto-check passed
  • Fill Protocol

    Aperivue/medsci-skills

    A skill your agent uses when an institutional Word form (.doc/.docx IRB protocol, ethics application, grant template) must be filled without breaking its styles, tables, fonts or page layout.

    331 GitHub stars~1.7k tokensUpdated 4 days ago
    Auto-check passed
  • Find Cohort Gap

    Aperivue/medsci-skills

    A skill your agent uses when looking for research topics a longitudinal cohort database can answer (NHIS, UK Biobank, an institutional EMR or registry).

    331 GitHub stars~2.9k tokensUpdated 4 days ago
    Auto-check passed

Questions about Imaging Data

What does Imaging Data do?

A skill your agent uses when preparing a medical-imaging dataset (DICOM/NIfTI) for modelling. Imaging Data is an agent skill from Aperivue/medsci-skills. Use when preparing a medical-imaging dataset (DICOM/NIfTI) for modelling.

When should I use Imaging Data?

Imaging Data fits situations like: preparing a medical-imaging dataset (DICOM/NIfTI) for modelling; tasks that involve Clinical and healthcare research.

How do I install Imaging Data in Claude Code?

Run `npx skills add Aperivue/medsci-skills --skill imaging-data -a claude-code`. Or copy the skill folder (skills/imaging-data in Aperivue/medsci-skills) into .claude/skills/imaging-data in your project. Claude Code loads it when a task matches its description.

How do I install Imaging Data in Codex?

Run `npx skills add Aperivue/medsci-skills --skill imaging-data -a codex`. Or copy the skill folder (skills/imaging-data in Aperivue/medsci-skills) into .agents/skills/imaging-data in your project. Codex loads it when a task matches its description.

Can I use Imaging Data in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Aperivue/medsci-skills --skill imaging-data -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/imaging-data, .gemini/skills/imaging-data, .github/skills/imaging-data and .opencode/skills/imaging-data in your project.

What does Imaging Data need to run?

Going by SKILL.md and its folder, Imaging Data needs Python and a shell for the scripts in its folder and the command-line tools its instructions call (bash and python3). Our summary lists: Python 3; A Bash shell.

Does Imaging Data access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Imaging Data safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Imaging Data use?

Imaging Data is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Imaging Data use?

About 2.9k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.3k tokens, read only when the agent opens those files.

What are the alternatives to Imaging Data?

Skills that share tags, products or a category with Imaging Data: Clinical Trials Database (google-deepmind/science-skills, 3.2k stars), CHARLS Paper Reproduction Guide (xjtulyc/MedgeClaw, 617 stars), Biomedical Analysis Dispatch (xjtulyc/MedgeClaw, 617 stars) and Research Proposal (luwill/research-skills, 858 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Imaging Data?

Aperivue (a GitHub organization) maintains it in Aperivue/medsci-skills, which has 331 GitHub stars. The repository holds 54 skills in this directory. The repository was last updated on October 5, 2026.

Source: Aperivue/medsci-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.