---
name: design-study
description: Use when checking a radiology or medical AI study design before drafting or submission. Reviews the analysis unit, cohort logic, leakage risks, comparator, validation strategy and reporting-guideline fit. AI-vs-expert benchmarks are /design-ai-benchmarking.
metadata:
  triggers: "study design, leakage check, cohort design, analysis plan, validation strategy, comparator design, bias check"
---

# Design-Study Skill

## Core Review Questions

Always inspect:

1. The exact research question.
2. The analysis unit: patient, lesion, exam, study, phase, report.
3. The index date or decision point.
4. How inclusion and exclusion criteria are applied.
5. Any information leakage.
6. The reference standard or endpoint definition.
7. A clinically meaningful comparator.
8. The validation strategy.
9. The uncertainty reporting required.
10. The best-fitting reporting guideline.
11. Whether exposure/outcome/covariate **definitions are literature-grounded** or invented ad hoc from the data dictionary. If ad hoc, defer to `/define-variables` before drafting Methods.

---

## Standard Output

```text
## Study Design Review
Question: ...
Study type: ...
Analysis unit: ...
Index date / prediction timepoint: ...

### Strengths
- ...

### Major validity risks
1. ...
2. ...

### Minimal fixes
- ...

### Reporting fit
- Recommended guideline: ...

### Decision
- Ready for analysis / Needs redesign / Drafting can proceed with limitations
```

Cite a reference only with a `/search-lit`-confirmed DOI or PMID; mark any other `[UNVERIFIED - NEEDS MANUAL CHECK]`. Never invent a clinical definition, diagnostic criterion, or guideline recommendation: flag an unconfirmed one `[VERIFY]` and ask the user.

---

## Workflow

### Phase 1: Reconstruct the study

Extract from the protocol, draft, slides, tables, or notes: clinical problem, intended use case,
population, inputs, outputs, outcome definition, and timing of variable availability.

**Gate:** Present the reconstructed study summary (question, analysis unit, intended use) to the
user and confirm it before proceeding — a wrong reconstruction misdirects the entire validity review.

### Phase 2: Check structural validity

Read a reference only when its condition holds:

| File | Read it when |
|---|---|
| `references/dag_adjustment.md` | confounding control needs an explicit adjustment set |
| `references/target_trial_emulation.md` | the design emulates a target trial |
| `references/venue_accept_recipe.md` | a clinical DL / AI-validation study must decide **which venue tier the achievable design can be accepted at, and the one design move that reaches the tier above** (the bridge into `/find-journal`); skip when there is no publication-tier decision |
| `references/combine_models_ablation_design.md` | the model is built by **combining / adapting / fine-tuning existing models** (nnU-Net, TotalSegmentator, SAM/MedSAM, a pretrained backbone) and the comparator must be designed as an ablation; skip for a model trained de novo |
| `references/multi_model_comparison_design.md` | the contribution is **comparing several models / architectures head-to-head** and the comparison must be fair; skip for a single-model study (one model's ablation → the row above; AI-vs-human → `/design-ai-benchmarking`) |
| `references/segmentation_failure_characterization_design.md` | the claim is that a segmentation model is **clinically usable**, not that it scores well; skip when the endpoint is benchmark accuracy (metric choice → `/model-assessment`; abstention / risk–coverage → `/model-assessment`) |

#### A. Analysis unit

Look for mismatches: a patient-level claim from lesion-level analysis; an exam-level split with
patient overlap; phase-level samples treated as independent.

#### B. Leakage

Look for:
- postoperative features used for preoperative prediction
- normalization or thresholding performed before the data split
- repeated exams across train/test
- reader annotations derived from outcome information
- **input-text contamination for NLP/LLM extraction tasks**: if the input includes clinical history,
  indication, impression, prior diagnosis, or referral text, confirm it does not name or strongly imply
  the target label. If it does, the task is information retrieval under label leakage, not phenotype
  inference: redesign the input mask, report a sensitivity analysis excluding those fields, or reframe.
- **construct dependence** (a predictor that is a definitional component of the outcome): (i) an input
  that computes the outcome (fasting insulin and glucose when the outcome is HOMA-IR), or (ii) a
  near-tautological composite of the outcome's defining components (inflated, near-circular
  association). Test: "could this predictor be derived, in whole or part, from the outcome's definition
  or the same measurement?" If yes, exclude it, or keep it only as a labeled calibration probe.

#### F. Time origin & survivorship (incident / transition models)

For any time-to-event or incident/transition design, check before drafting:
- **Time origin per model.** Each incident model starts its at-risk clock at the correct origin.
  Watch for **immortal-time bias** (a span in which the event cannot occur, misattributed to one
  group) and **left-truncation / delayed entry**.
- **Mediator-ascertainment-window survivorship.** A "progressor" / transition label conditional on
  *surviving to* a later ascertainment (a second scan, a follow-up visit) is survivorship-biased; plan
  a landmark time or an explicit intermediate-state (multistate / illness-death) model.
- **Primary-analysis-set selection.** If the primary is not the full cohort, pre-specify why, from
  the missingness mechanism *relative to the outcome*: complete-case regression is unbiased when
  missingness is independent of the outcome given the covariates (even with a covariate MNAR) and
  generally biased when it depends on the outcome, including under MAR, where multiple imputation
  is the usual primary (Hughes et al., *Int J Epidemiol* 2019; logistic-regression odds ratios are
  the exception when missingness depends on the outcome alone). Never pick it for significance.
- A design that cannot yet answer these should say so — but at review time a Methods/Limitations
  admission that it was *"not formally assessed"* is escalated to MAJOR by the survival probe (S1).

#### C. Reference standard

Check who established ground truth, when, whether blinding was possible, and whether only a subset
had gold-standard verification. Then:
- **Construct ↔ nominal-definition match.** Restate each construct's nominal definition and confirm
  every included case satisfies it. An "incidentaloma" defined as an *indeterminate* finding must not
  include frank malignancy reads; a label that overshoots its definition inflates the cohort and
  breaks the κ.
- **Per-flag reference-standard concordance.** Report concordance *per flag category*, not only overall;
  a construct where most flags (e.g., ~86%) miss the reference standard measures something else.
- **Manuscript definition ↔ `variable_operationalization.md`.** Methods definitions must match the
  `/define-variables` operationalization table verbatim, and a blinded re-classification form must
  quote the analytic protocol's definition verbatim. A paraphrase or "common-sense extension" in the
  form is the documented cause of a low κ that is a *definition mismatch*, not real disagreement.

#### D. Validation

Classify: apparent only, internal split, cross-validation, temporal, external, or multi-center external.

#### E. Reader / expert-elicitation studies

When the study elicits expert ratings — a reader study, an annotation panel, an AI-output evaluation —
read `references/reader_elicitation_design.md` before data collection. The acceptance ceiling of a
perceptual / reader AI study is fixed at design time: no quality of execution lifts a ceiling baked
into the comparator, the estimand, or the reader cohort. For an AI-system-versus-human-expert
benchmark, route to `/design-ai-benchmarking` (arm definition, LLM-as-judge versus human-as-judge
adjudication, structured export schema).

### Phase 3: Clinical framing

Ask whether the comparator and endpoint support the stated claim:
- Is the model better than current practice, or just another model?
- Is the endpoint clinically meaningful?
- Does performance translate to action?
- **incremental value**: if the model/marker is framed as adding value *beyond* / *on top of* /
  *incremental to* an existing tool (a clinical score, a routine test, a baseline model), pre-specify
  the baseline comparator built from the in-routine-use predictors **and** how added value will be
  shown: test it with the likelihood-ratio (or Wald) test of the new term in the development data;
  report ΔC-index / ΔAUC with a paired CI as the size of the gain, not as a second test (a DeLong
  test of nested models fitted and evaluated on the same data is invalid; DeLong is for two models
  frozen before a held-out evaluation); prefer the change in decision-curve net benefit at a
  pre-specified threshold; NRI only as a pre-specified categorical NRI reported separately for
  events and non-events, never the continuous NRI (analyze-stats `incremental_value.md`; Kerr et
  al., *Epidemiology* 2014). A standalone discrimination number does not support a "beyond X"
  claim, and the nested-model comparison cannot be added post hoc without the baseline model.
- **fine-tuning contribution baseline**: if an NLP/LLM study claims that fine-tuning, LoRA, prompt
  engineering, or a multi-agent wrapper improves extraction/classification, pre-specify a
  same-backbone zero-shot or few-shot comparator on the identical input, output schema, and test
  split; a weaker or unrelated baseline cannot show the adaptation adds value. For imaging:
  - a model built by combining / adapting / fine-tuning existing models → the full ablation ladder
    (un-adapted base, best single component, direct-train vs transfer), per
    `references/combine_models_ablation_design.md`;
  - a **head-to-head comparison of several models** → comparison *fairness*: one frozen
    split/preprocessing through every model, a strong fairly-tuned baseline, a matched (or disclosed)
    compute budget, and a paired delta test, per `references/multi_model_comparison_design.md`;
  - a claim that a segmentation model is **clinically usable** → a pre-specified failure taxonomy, an
    acceptability endpoint with a named judge and adjudication rule, the tail beside the mean, and edit
    effort paired against manual-from-scratch, per
    `references/segmentation_failure_characterization_design.md`. A mean DSC cannot be converted into
    a usability claim after the fact.
- **endpoint↔conclusion scope**: decide up front what *kind* of conclusion the design can support. A
  cross-sectional / single-visit / prevalence design cannot support a prognostic or surveillance claim
  (rescreen interval, disease progression) — that needs longitudinal follow-up. A binary surrogate
  endpoint (present/absent, >0, dichotomized) is risk stratification, not a patient-care directive
  (defer/withhold/initiate therapy). At review time `/self-review` §D + `check_scope_coherence.py` flag
  `CROSS_SECTIONAL_PROGNOSTIC` / `SURROGATE_CARE_DIRECTIVE` against the conclusion.

### Phase 4: Reporting fit

Recommend one primary guideline — `TRIPOD+AI`, `CLAIM`, `STARD`, `STROBE` (`TARGET` for a
target-trial emulation), `PRISMA`, `CARE`, or `ARRIVE` — plus journal-specific additions if needed.

---

## Frequent Failure Modes

### Diagnostic AI
No clinically relevant comparator; exam-level instead of patient-level split; unclear reference
standard; AUROC-only reporting without threshold metrics.

### Prognostic modeling
Unclear time zero; immortal time bias; feature timing mismatch; no calibration.

### Retrospective cohort / screening database

- **Time zero misalignment**: cohort entry ≠ follow-up start → immortal time bias.
- Interval-censored outcomes treated as exact → underestimated event times.
- Unacknowledged healthy-volunteer bias → inflated external-validity claims.
- Surveillance bias from unequal follow-up frequency between groups.
- Map each threat explicitly to one of the three bias classes (Hernán/Robins): selection bias (who
  enters), information bias (how measured), confounding (what else differs).
- **Comparative / causal question → emulate a target trial.** For a treatment-vs-treatment,
  screening-vs-no-screening, or drug-A-vs-drug-B question on routinely-collected data, specify the
  seven target-trial components (eligibility, strategies, assignment, **time zero**, outcome, causal
  contrast, analysis plan) before extraction. This prevents immortal-time, prevalent-user, and
  confounding-by-indication bias and turns an association into a defensible causal contrast.
  New-user + active-comparator design, grace-period clone-censor-weight, and negative controls are in
  `references/target_trial_emulation.md`.
- **Confounding completeness**: pre-specify the adjustment set from a DAG (not a Table-1 p < 0.05
  rule), and plan an extended-adjustment sensitivity model to show whether any measured covariate that
  turns out imbalanced by exposure but lies outside the adjustment set leaves the primary estimate
  robust. Pre-screen the proposed covariates with `scripts/adjustment_set_helper.py` (flags mediator /
  collider / descendant adjustment and omitted confounders, and proposes a candidate backdoor set),
  then derive the **minimal** sufficient set with dagitty — see `references/dag_adjustment.md`. At
  review time `/self-review` Phase 2.5e and the O-probes in `observational_confounding.md` (O1–O18)
  check this against Table 1 — including O7 over-adjustment, O10 overlapping-subset-gradient
  discipline, O11 design-based weighting for complex-survey data, O12 data-driven-threshold mining,
  O13 (a cross-sectional mediation claim cannot order X→M→Y), and O14 (a synergy/joint-effect claim
  needs the additive interaction scale — RERI/AP/S — not a multiplicative-only test).

### Multimodal LLM / report generation
No clear rubric for clinical correctness; benchmark labels derived from noisy reports without
adjudication; unsupported claims about safety or workflow benefit; input text containing the target
label or diagnosis being predicted; no same-backbone zero-shot/few-shot baseline for a fine-tuning or
prompt-engineering claim.

### Imaging meta-analysis
Overlapping cohorts; paired modalities analyzed as independent; heterogeneity metrics missing;
zero-cell handling unspecified.

---

## Minimal-Fix Principle

Recommend the smallest feasible repair first: clarify the claim, narrow the target population, add a
limitation statement, add a clinically relevant baseline, re-run one key sensitivity analysis, or
redefine the endpoint more explicitly. Escalate to redesign only when the central claim is not
defensible otherwise.

---

## Handoff Rules

- route to `analyze-stats` when the design is basically sound but analysis details need refinement
- route to `check-reporting` after the design is locked
- route to `self-review` when the user wants a pre-submission quality check on their own manuscript
- route back to `write-paper` only after the main validity risks are documented

This skill does not compute statistics, draft full manuscript prose, resolve raw data-engineering
issues, or replace a full peer review when journal-facing tone is required.
