---
name: self-review
description: Use when checking your own manuscript before submission from a reviewer's perspective. Returns anticipated Major/Minor comments with fixes, including numerical, citation and leakage checks, with an optional multi-reviewer panel. Someone else's paper is /peer-review.
metadata:
  triggers: "self-review, pre-submission check, check my paper, reviewer perspective, manuscript self-check"
---

# Self-Review Skill

Check the user's own manuscript before submission and produce an actionable list of anticipated
reviewer comments, each with a specific fix — not a written review.

## Optional Flags

- `--fix`: after the report, apply fixes for every issue with `fixable_by_ai: true`, editing the manuscript in place, then report a diff summary. Never fix `fixable_by_ai: false` issues (missing data, design flaws). Maximum 2 fix-and-re-review iterations; if the score is still below threshold after the second, stop and report what remains as structural (inside `/write-paper` this routes to Phase 7.4a Audit Recovery).
- `--json`: also emit the structured JSON block (Phase 3c). Default when called from `/write-paper` Phase 7.
- `--panel`: run the multi-agent panel review (Phase 2.6) — domain-expert reviewers in parallel plus an editor synthesis — instead of the single-pass review. Opt-in and **off by default**, because it spawns N reviewer agents + 1 editor and costs several times more tokens; reserve it for a high-stakes final pass on a top-tier target. Do **not** combine with `--fix`: a panel diagnoses and prioritizes; run `--fix` as a separate pass once the author has triaged the panel's findings.

## Severity Framing

- **Fatal**: a design flaw existing data cannot fix (data leakage that invalidates all results, no reference standard, label–feature circularity). Submission would likely be rejected. Reserve Fatal for true design-level problems.
- **Fixable**: significant but addressable with existing data (missing calibration, unclear exclusion criteria, absent CIs, incomplete reporting). Most issues are Fixable.

## Two Objectives: the Floor and the Ceiling

The **floor** (categories A–K and the gates of Phases 2.5–2.5f) minimizes rejection-for-cause —
fabricated citations, numbers that do not reconcile, overclaims, missing checklist items, leakage —
and much of it works by **adding** a hedge, caveat, disclosure or audit trail. Iterated, that
over-hardens a manuscript: every finding is correct, yet the whole reads as a defensive audit
(over-hedged, audit trail in the body, Abstract buried under caveats, the strongest robustness
result hidden in Limitations). The **ceiling** pass (category L / Phase 2.5g) runs **after** the
floor gates, reads the accurate manuscript as a whole, and recommends only SUBTRACTION — REMOVE,
MOVE, or TIGHTEN. It is advisory, never blocks, and cannot relax a floor gate. Report its findings
as their own block (Phase 3), not folded into the "add this" comments. Phase 2.5i then reads the
floor + ceiling state to declare the loop done, including a zero-edit PASS.

## Workflow

### Phase 1: Intake

1. Get the manuscript — PDF, Word doc, or pasted text.
2. Ask the user: target journal (sets reporting standards and scope); manuscript type (original research / review / perspective / technical note / letter / meta-analysis / case report); anything they are already worried about. On an interactive run, offer the `--panel` review (Phase 2.6) **once**, in one line, then proceed single-pass unless the user opts in. Never offer or apply the panel under `--json` or when called from `/write-paper`.
3. Read the full manuscript.
4. **SSOT gate — confirm there is one manuscript, not several.** Self-review reads a single file,
   so drift between a legacy working copy and the live submission copy is invisible to it. Before
   a `--panel` run or any pre-submission pass, check for multiple copies:

   ```bash
   find . \( -path '*manuscript*' -o -path '*main_document*' \) -name '*.md' | grep -v node_modules
   ```

   If more than one manuscript-like file exists, confirm which is the SSOT and run
   `/sync-submission`'s divergence gate — a `STALE_COPY` (an SSOT numeric claim or heading that did
   not propagate to the other copy) is a P0 that must clear first:

   ```bash
   python3 "${CLAUDE_SKILL_DIR}/../sync-submission/scripts/detect_copy_divergence.py" \
     --ssot <ssot>.md --copy <other-copy>.md
   ```

   Review the SSOT copy, never a stale one. **Under `--panel` this is blocking:** if `find` returns
   more than one file and the SSOT is not pinned (no `SSOT.yaml` with `truth.manuscript_md`, no
   explicit `--ssot <path>`), STOP before spawning any reviewer and have the user name the SSOT
   (and clear any `STALE_COPY`), because a panel on a stale copy wastes the whole pass. Do not
   auto-pick the longest/newest file. A single-pass review may proceed on the one file it was given.

### Phase 2: Systematic Check

Work every category the Research-Type Adaptation table (below) marks as applicable, and for each
item decide whether a reviewer would raise it as a Major or Minor comment. The per-item check
tables are in `references/phases/phase2_systematic_check.md` — read it once you have the
manuscript and know its type.

| | Category | What it asks |
|---|---|---|
| **A** | Study Design & Data Integrity | patient-level splits, leakage, input-text contamination, analysis unit |
| **B** | Reference Standard & Ground Truth | definition specificity, timing, annotator independence |
| **C** | Validation & Statistical Reporting | CIs, **calibration**, comparator, effect size, power-aware nulls, equivalence margins, interaction anchoring |
| **D** | Clinical Framing & Importance | intended use, overclaiming, novelty, **endpoint↔conclusion scope** |
| **E** | Reproducibility | preprocessing, model detail, hardware/software, data & code availability |
| **F** | Reporting Completeness | abstract↔body consistency, flow diagram, ethics, missing data, word cap |
| **G** | Reporting Guideline Compliance | match the type to its checklist; `/check-reporting` does the item-level audit |
| **H** | Circularity | label–feature overlap, tautological prediction, circular validation |
| **I** | Protocol Heterogeneity | multi-site acquisition, harmonization, temporal protocol drift |
| **J** | Method Transparency | model provenance, fine-tuning, classical-style body conventions |
| **K** | Reviewer-team consistency | *SR/MA only* — dual-vs-single conjunction, LLM-as-reviewer (both fabrication-grade) |
| **L** | Editorial impression & defensiveness | *advisory, never blocking* — the ceiling category: REMOVE / MOVE / TIGHTEN |

**Run the deterministic gates at Phase 2 entry, on every path:**

```bash
# D. endpoint↔conclusion scope
python3 "${CLAUDE_SKILL_DIR}/scripts/check_scope_coherence.py" \
  --manuscript manuscript.md --out qc/scope_coherence.json --strict

# J. classical-style body conventions
python3 "${CLAUDE_SKILL_DIR}/scripts/check_classical_style.py" \
  --manuscript manuscript.md --out qc/classical_style.json --strict

# K. reviewer-team consistency (SR/MA only; pass the extraction JSON file or directory)
python "${CLAUDE_SKILL_DIR}/scripts/check_reviewer_team_consistency.py" \
    --manuscript manuscript.md --prospero prospero/record.md \
    --extraction-json extraction/ --out _audit_self/reviewer_team_consistency.md

# J/D. Perspective structure (genre-gated: silent unless article_type is a Perspective).
# Pass the known type via --type; it also self-detects from the front-matter article_type.
python3 "${CLAUDE_SKILL_DIR}/scripts/check_perspective_structure.py" \
  --manuscript manuscript.md --type "${TYPE:-}" --out qc/perspective_structure.json

# Was every reported analysis ever defined (outcome, reference standard)?
python3 "${CLAUDE_SKILL_DIR}/scripts/check_analysis_definitions.py" \
  --manuscript manuscript.md --json --strict > qc/analysis_definitions.json
```

Verdict mapping: `CROSS_SECTIONAL_PROGNOSTIC`, `SURROGATE_CARE_DIRECTIVE`, `SECTION_SYMBOL`,
`INBODY_AI_DISCLOSURE`, any reviewer-team hit (exit 1), `MODEL_OUTCOME_UNDEFINED` (a Cox /
Fine–Gray / logistic model with no outcome named), `MODEL_NOT_IN_METHODS`, and
`REFERENCE_STANDARD_UNDEFINED` (discrimination or calibration with nothing to score against) are
Anticipated **Major** Comments. `CROSS_SECTIONAL_YIELD_LANGUAGE`, `ELIGIBILITY_PROSE`,
`DECIMAL_INCONSISTENCY`, `EM_DASH_OVERUSE`, `PERSPECTIVE_HEADING_NOT_ASSERTION`,
`PERSPECTIVE_ABSTRACT_NO_AUTHORIAL_MOVE`, and `TIER_LABEL_UNDEFINED` are **Minor**. The
per-verdict rationale and resolution paths are in the Phase 2 reference file.

`ANALYSIS_LOAD` is **informational, never a verdict**: many analyses are the cause of omitted
definitions, not the defect. Do not cut analyses to satisfy it; restore the definitions they
crowded out, and if load is genuinely high, move the defensive analyses to the supplement.

### Research-Type Adaptation

| Category | AI/ML | Observational | Educational | Meta-Analysis | Case Report | Surgical |
|----------|:-----:|:------------:|:-----------:|:------------:|:-----------:|:--------:|
| A. Study Design | Full | Full | Partial | N/A | N/A | Full |
| B. Reference Standard | Full | Full | N/A | Per-study | Partial | Full |
| C. Validation & Stats | Full | Full | Full | Special* | Partial | Full |
| D. Clinical Framing | Full | Full | Full | Full | Full | Full |
| E. Reproducibility | Full | Partial | Partial | Partial | N/A | Full |
| F. Reporting | Full | Full | Full | Full | Full | Full |
| G. Guideline Compliance | Full | Full | Full | Full | Full | Full |
| H. Circularity | Full | Partial | N/A | N/A | N/A | Partial |
| I. Protocol Heterogeneity | Full | Full | N/A | Per-study | N/A | Full |
| J. Method Transparency | Full | Partial | Partial | N/A | N/A | Partial |
| K. Reviewer-team consistency | N/A | N/A | N/A | Full | N/A | N/A |
| L. Editorial impression | Full | Full | Full | Full | Full | Full |

*Meta-analysis: replace C with heterogeneity (I², prediction intervals), publication bias (funnel
plot, Egger), and sensitivity/subgroup analyses.

**Type-specific additional checks:**

- **Observational**: confounding (DAG or adjustment strategy), selection bias, exposure measurement validity. Run **Phase 2.5e (Confounding Completeness)**, then apply the O-probes in `references/domain-probes/observational_confounding.md` — O1 (covariate imbalanced by exposure in Table 1 yet absent from the adjustment set) and O8 (records > subjects with the analysis unit undisclosed; `check_cohort_arithmetic.py --id-col`) are the deterministic ones, and O7 (adjusting for a mediator/consequence of the outcome) is their opposite-direction twin. For a **clinical prediction model** (TRIPOD / TRIPOD+AI, nested predictor-set comparison), also apply the CP-probes in `references/domain-probes/clinical_prediction_model.md`.
- **Educational**: learning-outcome measurement validity, Kirkpatrick level, control-group adequacy, curriculum fidelity.
- **Meta-analyses**: search comprehensiveness (2+ databases), screening reproducibility (2 reviewers), per-study RoB, GRADE certainty.
- **Case reports**: diagnostic-reasoning transparency, timeline completeness, informed consent, generalizability disclaimer.
- **Surgical**: learning curve, surgeon volume/experience, complication grading (Clavien-Dindo), operative detail.

**Domain probe modules** — the same probes `/peer-review` uses, vendored here. Load every module
whose row matches:

| Manuscript type / signal | Probe module |
|---|---|
| Systematic Review / Meta-Analysis | `references/domain-probes/sr_ma.md` (P0–P19) |
| Time-to-event / survival / prognostic model (Cox, Fine-Gray, DeepSurv, nomogram, risk-stratification cutoff) | `references/domain-probes/survival_prognostic.md` (S1–S9) |
| Radiomic feature reproducibility / acquisition-parameter sweep / reliability-based feature filtering | `references/domain-probes/radiomics.md` (R1–R4) |
| Cross-modality image synthesis (MRI→PET / MRI→CT / non-contrast→contrast / low-dose→full-dose) claiming functional/molecular information or target-modality substitution | `references/domain-probes/image_synthesis.md` (IS1–IS4) |
| Narrative / review article / primer / state-of-the-art | `references/domain-probes/narrative_review.md` (RV1–RV9) |
| Perspective / opinion / viewpoint (npj DM long-essay, Lancet Comment, NEJM AI / RYAI short-structured) | `references/domain-probes/narrative_review.md` (RV1–RV9) + the `check_perspective_structure.py` gate above |
| AI/ML primary study with a clinical claim (generalizable / outperforms clinicians / deployment-ready / can replace a reader) | `references/domain-probes/ai_overclaiming.md` (AO0–AO7) |
| Engineer-built medical-imaging model (segmentation / classification / detection) being validated — partition/leakage, seed & run variance, metric selection, reproducibility, reference-standard quality; plus saliency faithfulness, uncertainty/OOD/abstention and deployment feasibility when a clinical-use claim is made | `references/domain-probes/model_development.md` (MD0–MD11) |
| LLM / MLLM evaluated on a clinical task (report generation, VQA, clinical text extraction/classification; closed API or open weights) | `references/domain-probes/mllm_evaluation.md` (ME0–ME8) |
| Randomised controlled trial (parallel / crossover / cluster / stepped-wedge) | `references/domain-probes/rct_trial.md` (RC0–RC7) |
| Diagnostic test accuracy (DTA) primary study / multi-reader multi-case (MRMC) reader study (AI-vs-reader, AI-assisted reading, modality comparison) | `references/domain-probes/diagnostic_accuracy.md` (D1–D12) |
| Case report / case series (incl. adverse-event/pharmacovigilance and imaging-led radiology/nuclear-medicine/IR reports) | `references/domain-probes/case_report.md` (CR1–CR9) |
| AI/ML, prediction, or diagnostic study claiming cross-population performance, or presenting subgroup analyses as a fairness/equity argument | `references/domain-probes/equity_fairness.md` (EQ0–EQ6) |
| Mendelian randomization (two-sample, one-sample, multivariable, drug-target / cis-MR, non-linear MR) | `references/domain-probes/mendelian_randomization.md` (MR1–MR8) |
| Polygenic risk score (PRS / PGS) developed, validated, or applied as a predictor or risk-stratifier | `references/domain-probes/polygenic_risk_score.md` (PG1–PG8) |
| Network meta-analysis (≥3 interventions via direct + indirect evidence, treatment ranking, incl. component NMA) | `references/domain-probes/network_meta_analysis.md` (NM1–NM8) |
| Health economic evaluation (cost-effectiveness / cost-utility / cost-benefit / budget-impact; trial- or model-based — decision tree, Markov, DES) | `references/domain-probes/health_economic_evaluation.md` (HE1–HE8) |
| Observational study using routinely-collected health data (claims / EHR / registry / health-checkup DB, linked or not) | `references/domain-probes/record_routinely_collected_data.md` (RD1–RD8) |
| Self-report survey / questionnaire study (KAP, physician/patient survey, web/e-survey) | `references/domain-probes/survey_research.md` (SV1–SV8) |
| Scoping review (maps breadth of evidence; PCC framing, charting — not a focused effectiveness/accuracy question) | `references/domain-probes/scoping_review.md` (SC1–SC8) |
| Qualitative study (interviews, focus groups, ethnography, grounded theory, phenomenology, document analysis) | `references/domain-probes/qualitative_research.md` (QL1–QL8) |
| **Self-improving / self-evaluating system** (an agent that critiques and rewrites its own output; training on model-generated data; an LLM judge scoring the training signal; "self-evolving" clinical agents) | `references/domain-probes/self_improving_system.md` (SI1–SI7) + `${CLAUDE_SKILL_DIR}/../peer-review/scripts/check_self_improvement_claims.py` |

Apply each probe as an additional source of comments, complementing (not replacing) categories
A–K: a conclusion-threatening or design-level finding becomes a **Fatal** Anticipated Major
Comment, a reporting-level finding a **Fixable** Anticipated Minor Comment, each tagged with the
closest category letter (A–K).

For a **classifier / NLP / tabular ML** manuscript, also run the feature-selection-leakage gate — a
data-driven selection (feature selection, univariate filtering, vocabulary, a threshold) fit on the
FULL dataset before cross-validation inflates the CV metric:

```bash
python3 "${CLAUDE_SKILL_DIR}/scripts/check_cv_leakage.py" \
  --manuscript manuscript.md --out qc/cv_leakage.json
```

`CV_SELECTION_LEAKAGE` (Major) fires when a selection token co-occurs with cross-validation and no
fold-nesting is disclosed ("within each fold" / "nested CV" suppresses it). Patient-vs-image split
leakage is a different check (`model-assessment/check_split_leakage.py`).

### Phase 2.5: Numerical Cross-Verification (Internal)

Before the report, verify internal consistency:

1. **Abstract vs Body**: every Abstract number matches Results and Tables.
2. **Table vs Text**: sample sizes, primary outcomes and p-values agree between tables and narrative.
3. **Figure vs Text**: figure legends match the data described in Results.
4. **Percentage arithmetic**: n/N percentages are correct (23/150 = 15.3%, not 15.0%). The gate
   holds each cell to its printed precision (half a unit in the last printed place).
5. **CI plausibility**: confidence intervals are reasonable for the sample sizes.
6. **Rate back-calculation**: every rate inverts to its own numerator/denominator (incidence rate ≈ events / person-years × scale, ±rounding). A rate that does not recompute, or implies more events than the cohort can supply, is a Major.
7. **Exclusion-cascade and complete-case arithmetic** (cohort/observational): start N − Σ(exclusions) == final analytic N, and total − missing == complete. A footnote N that does not equal the subtraction is a Major.

For cohort/observational manuscripts, run the gate instead of eyeballing it (it parses prose
equations + GFM tables, and recomputes from a committed CSV when given one):

```bash
python3 "${CLAUDE_SKILL_DIR}/scripts/check_cohort_arithmetic.py" \
  --manuscript manuscript.md --data analysis/cohort.csv --id-col mockid \
  --out qc/cohort_arithmetic.json --strict
```

`RATE_BACKCALC` / `CASCADE_SUM` / `PARTITION_OVERLAP` are Anticipated Major Comments (category A).
On health-screening / EMR / registry data pass `--id-col` (or let it auto-detect a subject-ID
column) so it also checks the analysis unit: `records > unique subjects` with neither the unit nor a
one-record-per-subject sensitivity stated emits `ANALYSIS_UNIT_UNDISCLOSED` (Major — non-independent
observations give anti-conservative CIs; probe O8). Flag any remaining internal-consistency
discrepancies as Anticipated Minor Comments (category F).

Then recompute what a reviewer recomputes by hand:

```bash
# Every "n (%)" in a table, recomputed against its own denominator.
python3 "${CLAUDE_SKILL_DIR}/scripts/check_table_percentages.py" \
  --manuscript manuscript.md --out qc/table_percentages.json --strict

# Every reported P beside a 2×2 (or r×c) count, recomputed from the counts themselves.
python3 "${CLAUDE_SKILL_DIR}/scripts/check_reported_p_from_counts.py" \
  --manuscript manuscript.md --json --strict > qc/reported_p.json

# Every t(df) / F(df1, df2) / χ2(df) / z with a P in its sentence, P recomputed from the statistic.
python3 "${CLAUDE_SKILL_DIR}/scripts/check_test_statistic_p.py" \
  --manuscript manuscript.md --json --strict > qc/test_statistic_p.json   # add --grim grim.json

# Diagnostic-accuracy only: sensitivity/specificity against the reference-standard denominators.
python3 "${CLAUDE_SKILL_DIR}/scripts/check_dta_denominators.py" \
  --manuscript manuscript.md --json --strict > qc/dta_denominators.json
```

`PERCENT_MISMATCH`, `P_NOT_REPRODUCIBLE`, and `DTA_DENOMINATOR_MISMATCH` / `STAGE_ROWSUM` are
**P0 Major** — not a rounding disagreement: one of the two numbers is wrong. Run the first two on
**every** manuscript with a table; the third only on diagnostic-accuracy work.

`check_test_statistic_p.py` reads the statistic and the P at their printed precision: a P no
statistic in the rounding interval can give is `P_STAT_INCONSISTENT` (Minor), and Major
`P_STAT_DECISION_ERROR` when the two also fall on opposite sides of alpha (`--alpha`). Adjusted or
non-standard P values (corrected, exact, Welch, …) and results that match only the unstated
one-sided P are never Major. `--grim` takes declared means
(`[{"label", "mean", "n", "items"?, "decimals"?}]`): `GRIM_INCONSISTENT` (Major) when no integer
sum gives the mean; `GRIM_NOT_ASSESSED` when n·items >= 10^decimals.

To check each P against its own test, declare the tests as `p_tests.json` (copy
`${CLAUDE_SKILL_DIR}/templates/p_tests.json`; schema in `references/p_tests_schema.md`) and add
`--tests p_tests.json`. The row's P is recomputed with the declared test (`fisher`, `chi2`,
`chi2_yates`). A reported P whose whole printed-precision interval lies on the other side of alpha
from the recomputed P is `P_ALPHA_CROSSING` (**Major**). An `adjusted`, `paired` or `other:` test,
a row whose printed percentage is not count/n (another denominator), or a table of 3+ groups (after dropping
a Total column) is `P_NOT_ASSESSED` (Minor).

Known limits: without `--tests`, `check_reported_p_from_counts.py` flags a row only when its reported
P differs from the crude 2x2 P by more than one order of magnitude under every test family, and only
in tables with at least two count rows. It does not flag a reported `<` bound (`<0.001`) that the
crude P exceeds, or a P on the other side of alpha, by less than that, nor a lone count row (open
finding SR-01 in prose-only mode): the table cannot show whether the P is adjusted, paired or on
another denominator. Check such P values against the stated test by hand, or declare them.
`check_table_percentages.py` reads one denominator per column. A cell that misses only at its printed
precision (within 0.5 pp) on a footnoted row (`Current smoker^a | 23 (15.4%)` under n = 150 with 149
known) is a Minor `PERCENT_PRECISION_NOTE`; an unfootnoted row, or a larger miss, is a
`PERCENT_MISMATCH`. Confirm footnoted rows against the footnote.

### Phase 2.5a: Numerical Source-Fidelity Audit (External)

Numbers can be self-consistent everywhere and still wrong at the source; only a traversal back to
the primary source catches that. First run the **displayed-arithmetic** gate — a stated difference
must equal the subtraction of its two displayed components at the same precision:

```bash
python3 "${CLAUDE_SKILL_DIR}/scripts/check_rounded_delta.py" \
  --manuscript manuscript.md --out qc/rounded_delta.json
```

`ROUNDED_DELTA_MISMATCH` (Minor) fires when AUCs shown as `0.79` and `0.82` are reported with a
difference of `0.02`. A higher-precision pair (`0.794` vs `0.816`) with a 2-dp delta is legitimate
and not flagged.

**When to run the external audit:** MA revisions, submissions, or when the user says "check against
the source", "verify extraction", or "random sample". Skip otherwise.

**The audit:** draw a stratified sample of 5 numerical claims — always including one
comparative-arm value and one revision-introduced number — and trace each through three layers
(manuscript → extraction CSV → primary-source page; plus analysis script → CSV where a script
produced it). **Any mismatch is a Major Comment**; one that reverses a direction or crosses a
significance boundary is a P0 blocker. Every `[VERIFY-CSV]` tag is a mandatory audit item regardless
of sample size.

Read `references/phases/phase2_5a_source_fidelity.md` when running the external audit — it has the
traversal procedure, recording table, sampling strata, and four prose-judgement rules (hand-entered
analysis-script inputs; prose↔table statistic-type mismatch, e.g. a median in the text against a
mean in Table 1; stale derived CSVs after a model/adjustment-set change; direction reversals
internal consistency cannot see).

### Phase 2.5a-2: Design & Power Statistic Provenance

Applies only when the manuscript states a sample-size calculation, a power figure, or a
detectable-effect claim. These are computed, not copied, so Phase 2.5a cannot check them: re-derive
each from the manuscript's own inputs. A value not reproduced by committed code, or reproducible
only by a method the committed script does not implement, is a Major Comment (P0 if a headline
claim). Read `references/phases/phase2_5a2_design_power.md` when the manuscript reports a
sample-size / power / MDE calculation.

### Phase 2.5b: Screening-Count Reconciliation from ID Sets (SR/MA + observational tier/stratum)

**When to run:** any SR/MA manuscript revision, at any stage (before Phase 3); or any observational
manuscript presenting an ordinal tier / mutually-exclusive stratum split. Skip otherwise. A wrong
prose total survives every other pass because Abstract, Methods, Results, captions and supplement
all cite it back to each other; only a recount from the **ID sets** catches it.

**A. SR/MA — recount from the ID sets.** Derive every study count from the screening TSV and the
consensus sheet, never from prose, and **list the narrative-only IDs explicitly** (turning "10
narrative-only studies" into "2 (IDs 120, 474)"). A derived total that disagrees with the Abstract,
Methods, Results, Figure 1 caption or Limitations is a **P0 Major, blocking submission**; an `N → M`
transition claim not backed by an enumerable ID addition/subtraction set is a **Major**.

**B. Observational tier/stratum.** A disjoint partition must satisfy `Σ(stratum N) == unique total`
and `Σ(stratum events) == total events`. Denominators summing *above* the unique cohort double-count
subjects; every stratum n equal to the grand total is a mis-entry. Confirm the reference (baseline)
row of any stratified hazard/odds table is present and labelled.

```bash
python3 "${CLAUDE_SKILL_DIR}/scripts/check_cohort_arithmetic.py" \
  --manuscript manuscript.md --data analysis/strata.csv --strict
```

**C. Cross-script cut-point consistency.** When one cohort is re-stratified in more than one
script, the derived categorical must use one cut definition (same breaks, same `right=` closure,
same labels) — otherwise per-stratum Ns drift while the grand total still reconciles. The same gate
covers a derived 0/1 composite rebuilt in a second script with a clause dropped.

```bash
python3 "${CLAUDE_SKILL_DIR}/scripts/check_binning_consistency.py" \
  --root analysis --root scripts --strict
```

`PARTITION_OVERLAP`, `BINNING_DRIFT`, and `DERIVED_DEF_DRIFT` are all **P0 Major**. Read
`references/phases/phase2_5b_screening_counts.md` when doing the SR/MA ID-set recount or a
stratified-cohort recount (set definitions, derivation formulas, reconciliation-block template).

### Phase 2.5c: Reference Scans (hallucination + adequacy)

**2.5c** catches a citation that does not exist or has an invented first author; **2.5c-2** catches a
claim with no citation. Both need a bibliography — skip them if there is no `refs.bib` and no
reference list. Run `/verify-refs --strict` first (these scans read its audit), then the adequacy
checker:

```bash
python3 "${CLAUDE_SKILL_DIR}/scripts/check_reference_adequacy.py" \
  --manuscript manuscript/manuscript.md --bib "$BIB" \
  --article-type "$TYPE" ${CAP:+--journal-cap "$CAP"} \
  --out qc/reference_adequacy.json --strict
```

A `FABRICATED` record or any `duplicate_findings[]` entry in `qc/reference_audit.json` is a P0 Major
Comment that blocks submission. Read `references/phases/phase2_5c_reference_scans.md` when the
manuscript has a bibliography and you are auditing citations.

### Phase 2.5d: Cross-Reference QC (Manuscript ↔ rendered DOCX)

In-text Table/Figure citations can resolve to a *different* caption in the rendered DOCX when the
build script carries its own legacy SSOT; no other phase sees this.

**Markdown stage (always).** Every captioned `Figure N.` / `Table N.` must be cited elsewhere in the
body:

```bash
python3 "${CLAUDE_SKILL_DIR}/scripts/check_figure_citation.py" \
  --manuscript manuscript.md --out qc/figure_citation.json
```

`FIGURE_ORPHAN` / `TABLE_ORPHAN` (Minor): a float with a legend but no in-text citation.

**DOCX stage (only when a rendered DOCX exists** — circulation drafts, post-build checks):

```bash
python3 "${CLAUDE_SKILL_DIR}/../manage-refs/scripts/check_xref.py" \
  --md manuscript/manuscript.md --docx manuscript/manuscript_final.docx \
  --out qc/xref_audit.json [--allow-separate-attachments]
```

`MISMATCH` is always **Major (P0)**. `MISSING_DOCX` and `MISSING_BODY` are **Major (P0)** by default.
For journals that take figures/tables as separate attachments (European Radiology, Radiology, AJR),
pass `--allow-separate-attachments`: it downgrades `MISSING_DOCX` to Minor, and `MISSING_BODY` to
Minor only when no `--docx` was supplied — nothing was checked, so treat those rows as unverified
(`summary.downgraded_unchecked`) and re-run with `--docx` before submission. `MISSING_BODY` for a
float that IS in the rendered DOCX stays P0 (SSOT drift). `UNCITED` is Minor.

**Do NOT auto-fix cross-reference defects in `--fix` mode**, because rewriting a body caption
without re-running the DOCX build only moves the mismatch. Emit each P0 row as its own `M`-numbered
Major Comment with `category: "F"` and `fixable_by_ai: false`, and route the user to `/write-paper`
Step 7.6a. Read `references/phases/phase2_5d_xref_qc.md` when the xref gate fired and you are writing
up the reconciliation.

### Phase 2.5e: Confounding Completeness (observational only)

**When to run:** observational manuscripts (cohort, case-control, cross-sectional,
health-screening registry) whose central claim is an adjusted exposure–outcome association. **Skip
for RCTs, diagnostic-accuracy, SR/MA, and descriptive studies.**

A covariate that is measured, imbalanced across exposure groups in Table 1, and absent from the
adjustment set (probe O1) is invisible to a prose pass. Run the gate and treat each
`UNADJUSTED_IMBALANCED` covariate as an Anticipated Major Comment (category A):

```bash
python3 "${CLAUDE_SKILL_DIR}/scripts/check_confounding_completeness.py" \
  --table1 table1_by_<exposure>.csv \
  --adjusted-list "age, sex, BMI, hypertension, diabetes" \
  --exposure-defining-list "body mass index, waist, fasting glucose, triglycerides, HDL cholesterol" \
  --out qc/confounding_completeness.json --strict
```

For observational manuscripts, **read `references/phases/confounding_completeness.md`** for the full
procedure: the `--exposure-defining-list` over-adjustment exemption for guideline-defined exposures
(MASLD / metabolic syndrome / CKM / sarcopenia / frailty), the SMD-from-`mean ± SD` fallback, the
extended-adjustment sensitivity model (refit the unadjusted estimate on the reduced complete-case
frame, not the full frame), and probes O2–O10 of `references/domain-probes/observational_confounding.md`.

### Phase 2.5f: Claim-vs-Artifact Cross-Check

This phase checks claims against the artifacts they should trace to — the pre-registration, the
protocol, the analysis outputs — where the prose is internally consistent yet disagrees with them.
Run the gates (pass the supplement so the corpus is complete):

```bash
# 1. claims ↔ pre-registration/protocol: estimand provenance + E-value arithmetic
python3 "${CLAUDE_SKILL_DIR}/scripts/check_claim_artifact.py" \
  --manuscript manuscript.md --prereg prereg.md \
  --out qc/claim_artifact.json --strict   # + --evalues evalues.json (templates/) to recompute declared E-values

# 2. Methods ↔ Results ↔ disk coverage (both directions: promised-absent AND run-but-unreported)
python3 "${CLAUDE_SKILL_DIR}/scripts/check_artifact_coverage.py" \
  --manuscript manuscript.md --supplement supplement.md --analysis-dir output/analysis \
  --out qc/artifact_coverage.json --strict

# 3. reader-facing residue in EVERY rendered artifact, not just the body
python3 "${CLAUDE_SKILL_DIR}/scripts/check_supplement_hygiene.py" \
  --supplement supplement.md --supplement tables.md --supplement captions.md \
  --manuscript manuscript.md --out qc/supplement_hygiene.json --strict

# 4. float AND in-text reference-number ([N]) citation order — a desk-reject item the hygiene gate does not cover
python3 "${CLAUDE_SKILL_DIR}/scripts/check_citation_order.py" \
  --manuscript manuscript.md --out qc/citation_order.json --strict

# 5. a headline null is uninterpretable without a precision statement
python3 "${CLAUDE_SKILL_DIR}/scripts/check_null_calibration.py" \
  --manuscript manuscript.md --out qc/null_calibration.json --strict

# 5b. a headline OR/HR/RR whose 95% CI spans an order of magnitude (a direction, not a magnitude), or events/covariates < 10 (EPV)
python3 "${CLAUDE_SKILL_DIR}/scripts/check_effect_stability.py" \
  --manuscript manuscript.md --out qc/effect_stability.json --strict

# 5c. incorporation bias — a trajectory-defined reference standard with a trajectory predictor (growth) reported as associated with the outcome
python3 "${CLAUDE_SKILL_DIR}/scripts/check_incorporation_bias.py" \
  --manuscript manuscript.md --out qc/incorporation_bias.json --strict

# 6. reader/observer study only — prove the (call × confidence) → score encoding is strictly
#    monotonic; a folded score silently mis-estimates the AUC and no prose review can see it
python3 "${CLAUDE_SKILL_DIR}/../analyze-stats/scripts/rating_monotonicity.py" \
  --encoding score_def.json
```

| Verdict | Severity |
|---|---|
| `PRIMARY_REASSIGNED` | **Major** — the primary was re-designated after results were known |
| `EVALUE_ARITHMETIC` | **Major** — recompute for the *declared primary* estimate; `EVALUE_NON_PRIMARY` is an advisory flag (check which estimate the E-value bounds) |
| `PROMISED_ABSENT`, `DISK_UNREPORTED`, `PROMISED_STAT_NO_VALUE` | **Major** |
| `SUPP_INTERNAL_LABEL`, `SUPP_PLACEHOLDER`, `SUPP_BUILD_MARKER`, `SUPP_RESPONSE_FRAMING`, `SUPP_PLANNING_RESIDUE`, `SUPP_XREF_UNRESOLVED` | **Major** — a slip in a supplement is as fatal at a technical check as one in the body |
| `CITATION_ORDER` | **Major**; `CITATION_GAP` **Minor** |
| `CONFIRM_NULL_NO_MDE` | **Major** |
| `ESTIMAND_DRIFT`, `PRIMARY_DISCLOSURE_NOTE` | **Advisory Minor — never a blocker.** The provenance match is fuzzy (token overlap); confirm against the actual registration first. `PRIMARY_DISCLOSURE_NOTE` flags an honest disclosure the guidance recommends — do not penalise it. |

With `--evalues` (schema `references/evalues_schema.md`), each declared E-value is recomputed from
its declared RR and CI (VanderWeele–Ding; the CI E-value from the limit nearest 1) over the printed
precision of every number: `EVALUE_DECLARED_MISMATCH` is **Major**, `EVALUE_DECLARED_NOT_IN_TEXT`
Minor; an OR/HR must be converted to an RR or declared `other:` (Minor, not recomputed).
Known limits: in prose-only mode the E-value check splits sentences at every '.', so the decimal in
"HR 1.52" can cut the estimate out of its window (`EVALUE_UNVERIFIABLE`); "E-value = 3.10",
"(E-value 3.10)" and the plural are not read, and a CI-limit E-value is not recognised (open
finding SR-02 without `--evalues`). Check by hand, or declare them.

**Checks no script makes** (prose judgement):

1. **Primary-change guard** — two models for one contrast, one significant and one null, the
   significant one foregrounded: confirm which was pre-specified.
2. **Headline vs own-sensitivity direction** — a headline claim pointing the opposite way from the
   authors' own sensitivity estimate means the paper contradicts its own robustness check (Major).
3. **Figure-embedded numbers are grep-blind** — every numeric audit above misses numbers *inside* a
   rasterised figure. Read each figure page visually before submission.

Re-run `/sync-submission`'s `check_cross_artifact_stale.py` **after** any reframe, not just once at
the start. For time-to-event manuscripts, apply probe **S8 (estimand provenance)** of
`references/domain-probes/survival_prognostic.md`. Read `references/phases/phase2_5f_claim_artifact.md`
when a gate above fired and you need the rationale and resolution path, or there is a
pre-registration to reconcile.

### Phase 2.5g: Editorial-Impression / Defensiveness Scan (the ceiling pass)

Run this **after** the floor gates (Phases 2.5–2.5f): it reads the accurate manuscript and recommends
what to take back out. It is advisory and **non-blocking** — it never produces a Major and never
gates submission.

```bash
python3 "${CLAUDE_SKILL_DIR}/scripts/check_editorial_impression.py" \
  --manuscript manuscript.md --out qc/editorial_impression.json
```

It exits 0 even under `--strict` and emits up to six verdicts, each tagged with a SUBTRACTION
`action`:

| Verdict | Reads as | Action |
|---|---|---|
| `HEDGE_DENSITY` | defensive-caveat tokens per 1,000 narrative words over threshold | TIGHTEN |
| `HEDGE_REPEAT` | one caveat motif repeated across body + Abstract | TIGHTEN |
| `AUDIT_IN_BODY` | SHA / commit / unit-test / post-lock / manifest / seed in the narrative | MOVE (→ Methods/supplement) |
| `LIMITATIONS_VOLUME` | a long enumerated Limitations list | TIGHTEN (consolidate) |
| `ABSTRACT_CAVEAT_LOAD` | several caveat clauses in the Abstract | TIGHTEN |
| `BURIED_DEFENSE` | strong numeric robustness result only in Limitations/supplement | MOVE (→ Results) |

Each finding becomes a Minor `issues[]` entry with `category: "L"`, `category_name: "Editorial
impression"`, `issue_type: "editorial_impression"`, `subtype: <verdict>`, and `action: "REMOVE" |
"MOVE" | "TIGHTEN"`, reported in the Phase 3 "Editorial-Impression Risks" block. Mark them
`fixable_by_ai: false` — tightening a hedge or moving a result is the author's voice-and-judgment
edit — except a clearly redundant `HEDGE_REPEAT`, which `--fix` may collapse to one statement.

When an earlier phase recommends adding a caveat or disclosure, weigh it against L: an
integrity-critical disclosure is a must, stated once and crisply; a defensive over-disclosure is a
cut or move. Place it once and point to the supplement rather than repeating it at every claim site.

### Phase 2.5h: Baseline Drift (anchor to the last human-approved version)

Run after the ceiling pass and **before** the loop controller (Phase 2.5i). Each refine pass takes
the previous AI output as its baseline, so framing bias compounds unseen; this gate compares the
manuscript against the **last human-approved version** (the frozen `v_N` circulated to senior
authors/co-authors — **not** the last AI output). With no baseline (a first draft), skip it.

```bash
python3 "${CLAUDE_SKILL_DIR}/scripts/check_baseline_drift.py" \
  --manuscript manuscript.md --baseline "$BASELINE_MD" \
  --out qc/baseline_drift.json
```

| Verdict | Signal (baseline → current) | Fold into report as |
|---|---|---|
| `STRENGTH_INFLATION` | certainty markers up while hedges fall | Minor — tone back to the approved strength |
| `SIGNIFICANCE_INFLATION_DRIFT` | novel/pivotal/unprecedented tokens added | Minor — remove the inflation |
| `SCOPE_INFLATION_DRIFT` | new generalization phrases ("in clinical practice") | Minor — the estimand did not widen; re-scope |
| `HEDGE_ACCRETION` | hedge/caveat density up | Minor — cumulative over-hardening; TIGHTEN |

Every finding is **Minor and advisory**; the gate never blocks. Treat drift as a prompt to review
against the approved anchor, not an instruction to revert — new analysis can justify a stronger
claim, but the author should confirm it. `qc/baseline_drift.json` feeds the loop controller, so a
drifted draft does not read as a zero-edit PASS.

### Phase 2.5i: Refinement Terminal-State (the loop controller)

Run this **last**, after the floor gates and the ceiling pass: it reads their `qc/*.json` artifacts
and classifies whether the review → revise → review loop is done, so an accurate manuscript is not
over-hardened by a pass it does not need. It is advisory and **never blocks**; do not treat it as a
second gate on the floor detectors, which already fail under `--strict` on their own Majors.

```bash
python3 "${CLAUDE_SKILL_DIR}/scripts/refinement_stop.py" \
  --qc-dir qc --out qc/refinement_stop.json
```

| Verdict | Meaning | What the harness must do |
|---|---|---|
| `CONTINUE` | a floor gate still reports a Major | genuine work remains — keep going |
| `STOP_OVERHARDENING` | floor clean, ceiling flags accumulation | STOP adding; only optional SUBTRACTION (REMOVE/MOVE/TIGHTEN) remains — do **not** run another additive pass |
| `STOP_MINOR_OPTIONAL` | floor clean, only optional Minor polish left | stop the required-work loop; present the Minor items as an optional menu, do not loop for them |
| `STOP_ZERO_EDIT` | floor at fixed point, ceiling clean | the manuscript is submission-ready as-is — **NO EDITS REQUIRED. Do not manufacture changes.** Report the zero-edit PASS as a first-class outcome |
| `INDETERMINATE` | no gate artifacts yet, no floor gate parsed (e.g. only the ceiling ran), or an empty / invalid `qc/*.json` (a gate that crashed under `--json > file`) | run, or re-run, the floor + ceiling gates first |

Known limits: a detector-keyed artifact whose findings are not under `claims` / `findings`
(for example `/verify-refs`' `qc/reference_audit.json`) is listed as *Unparsed* with a WARNING but
does not by itself block a `STOP_*`; read its own verdict before acting on the stop signal.

Once the verdict is any `STOP_*`, stop the additive cycle: surface the terminal state in the Phase 3
report and do not re-run self-review to find "one more thing". A zero-edit or minor-optional result
is a legitimate outcome, not a failure to try harder.

### Phase 2.5j: Refinement Regression (fixed vs broke, across runs)

Run each round, after the loop controller. A revision that resolves finding X can introduce finding
Y, and the pass-rate hides it; this step reads a run-history ledger (one line per run, the
`verdict@where` fingerprints of its findings) and reports what the revision *fixed* vs what it
*broke*.

```bash
python3 "${CLAUDE_SKILL_DIR}/scripts/refinement_regression.py" \
  --qc-dir qc --ledger qc/refinement_ledger.jsonl --append \
  --out qc/refinement_regression.json
```

Use `--append` on a real run so the current findings become the next entry; omit it to classify
without recording.

| Verdict | Meaning | What the harness must do |
|---|---|---|
| `PROGRESSING` | findings resolved, none new | continue |
| `REGRESSION` | the revision introduced new finding(s) | review the new findings before accepting the fix — the pass-rate went up but something broke |
| `CHURNING` | a resolved finding reappeared (Mirror Loop) | **stop revising and re-anchor** — more passes re-derive, they do not converge |
| `CONVERGED` | nothing new, nothing carried | the loop is done |
| `INDETERMINATE` | first run, no prior entry | re-run after a revision |

It is advisory and **never blocks**. Report both axes in Phase 3: a revision is an improvement only
if it resolved findings **and** the `new`/`churn` columns are empty.

### Phase 2.6: Multi-Agent Panel Review (--panel, opt-in)

Run this phase **only when `--panel` is passed**, after the numerical audits (Phases 2.5–2.5d) so the
reviewers see source-verified numbers, and before the Phase 3 report, which it feeds. Two things bind
before you spawn anything: the **SSOT must be singular** (the Phase 1 step 4 gate — halt and ask if
more than one manuscript-like `.md` is unpinned), and the roster must not be a **substrate
monoculture** (a panel sharing the drafter's model inherits its blind spots; route at least one lens
to Codex or a human co-author). `check_panel_diversity.py --strict` enforces the second, and fires
`PANEL_UNDERRETURN` when fewer reviewers returned than were spawned — a panel with <2 returned reviews
is a failed run, not a thin one.

Read `${CLAUDE_SKILL_DIR}/references/phases/phase2_6_panel.md` when `--panel` is passed — it has the
reviewer-set table, roster manifest, editor synthesis and lens-diversity gate.

### Phase 3: Report

Before writing the comments, skim `references/exemplar_findings/` for the finding at hand
(cohort-arithmetic mismatch, unadjusted confounder, cross-sectional scope overreach, post-hoc
primary / estimand drift). Each models the full shape — gate fired, the comment in a reviewer's
words, Fatal/Fixable, category letter, fix, `fixable_by_ai`, R0-ready line. Match the structure, not
the wording; they are synthetic.

Write the report to `qc/self_review.md` with this structure:

```markdown
# Self-Review Report: {manuscript title}

**Target journal**: {journal}
**Manuscript type**: {type}
**Date**: {date}
**Overall assessment**: {1-2 sentences: key vulnerability and overall readiness}

## Anticipated Major Comments (fix before submission)

M1. **{Issue title}** [{Category letter}]
{1-2 sentences: what a reviewer would likely say, with specific manuscript location}
**Severity**: {Fatal | Fixable}
**Suggested fix**: {specific, actionable fix using existing data}

M2. ...

## Anticipated Minor Comments (address proactively)

m1. **{Issue}** [{Category}]: {1 sentence with location + fix}
m2. ...

## Editorial-Impression Risks (REMOVE / MOVE / TIGHTEN)

*The subtraction axis — what to take out, move, or tighten so the accurate manuscript reads
confidently. Advisory and non-blocking; from Phase 2.5g / category L. Omit this block only if the
scan returned nothing.*

L1. **{Issue}** [{REMOVE | MOVE | TIGHTEN}]: {1 sentence — what reads as over-defensive and where, with the subtraction to make}
L2. ...

## Strengths (emphasize in cover letter)

- {Specific strength 1}
- {Specific strength 2}
- ...
```

Keep the ADD / FIX axis (Major / Minor Comments) and the SUBTRACTION axis (Editorial-Impression
Risks) visually separate; never fold L items into the Minor Comments.

In suggested fixes and in any text drafted in Phase 4:
- **References** only from `/search-lit` with a confirmed DOI or PMID; mark any other reference `[UNVERIFIED - NEEDS MANUAL CHECK]` (Phase 2.5c blocks a `FABRICATED` one).
- **Clinical definitions, diagnostic criteria and guideline recommendations** you cannot verify: flag with `[VERIFY]` and ask the user; never invent them.

**Conciseness**: each comment says what a reviewer would raise, where, and the fix: a Major in a
few lines, a Minor in a sentence or two, an Editorial-Impression Risk in one sentence. List every
Major the gates and the review produced (each P0 gate row is its own Major); never merge or drop
one to fit a count. Strengths: the few worth repeating in the cover letter — never a sentence that
pre-empts an objection (anticipated objections go to the response-letter bank; see §L).

### Phase 3b: R0 Numbering (Optional)

If the user plans to use `/revise` once real reviews arrive, offer to append R0-numbered output so
they can later tell anticipated (R0) from novel (R1-only) comments:

```markdown
## R0 Pre-Submission Findings (for /revise cross-reference)

R0-1 [MAJ] {mapped from M1}: {issue title}
R0-2 [MAJ] {mapped from M2}: {issue title}
R0-3 [MIN] {mapped from m1}: {issue title}
...
```

### Phase 3c: Structured JSON Output (--json)

Emit machine-readable JSON **only when `--json` is passed** (or another skill consumes this run):
append the block to the report and also write it to `qc/self_review.json`. Read
`${CLAUDE_SKILL_DIR}/references/phases/phase3c_json_output.md` for the schema, field semantics and
worked example when --json was passed, or a downstream skill consumes this run.
The JSON carries a `coverage` ledger (status per category A–L and per loaded probe module), and
**PASS is not allowed while any applicable entry is `not_assessed`**. Finalise it with
`scripts/check_review_coverage.py --review qc/self_review.json --out qc/review_coverage.json --strict`
(exit 1 = PASS over an unworked category; fix the JSON before any consumer reads it).

### Phase 4: Fix Support (on request)

The review ends at Phase 3. Read `${CLAUDE_SKILL_DIR}/references/phases/phase4_fix_support.md` when
`--fix` was passed or the user asks you to apply or draft fixes for the findings.
