Agent skill

Check Reporting

by Aperivue in Aperivue/medsci-skills

A skill your agent uses when auditing a manuscript item by item against a reporting guideline or risk-of-bias tool.

MITAuto-check passedDatabases

Install Check Reporting

skills CLI
$ npx skills add Aperivue/medsci-skills --skill check-reporting -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Aperivue/medsci-skills check-reporting --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Aperivue/medsci-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/check-reporting .claude/skills/check-reporting && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
check-reporting
GitHub stars
329
Used in
1 other repo
Token cost
~7.9k tokens
SKILL.md length
3,741 words
Files
86 (incl. scripts, references)
Skills in repo
54
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when auditing a manuscript item by item against a reporting guideline or risk-of-bias tool.

  • Works in 6 steps: Existing-checklist staleness pre-check → Select Guideline → Load Checklist → …
  • Auditing a manuscript item by item against a reporting guideline
  • SKILL.md covers Reference Files, Workflow, Skill Interactions and Gates
  • Calls python3 and python

What it does

Check Reporting is an agent skill from Aperivue/medsci-skills. Use when auditing a manuscript item by item against a reporting guideline or risk-of-bias tool. Covers 49 reporting guidelines and risk-of-bias tools (STROBE, CONSORT, STARD, TRIPOD+AI, PRISMA, QUADAS and more), marking each item PRESENT, PARTIAL or MISSING. Not a reviewer critique (/self-review).

Its SKILL.md is about 7.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 88 other files, including scripts and reference files (for example `references/LICENSES.md`, `references/appraisal_tools/METRICS.md` and `references/appraisal_tools/METRICS_RELOADED.md`).

It sits in Databases, covering ORMs and data access. It works with Prisma. The repository describes itself as: Agent Skills for medical research — literature search, reporting-guideline & citation checks, statistics, publication figures, submission. Works with Claude Code, Codex, Cursor &… The licence is MIT.

When your agent uses it

  • Auditing a manuscript item by item against a reporting guideline
  • Risk-of-bias tool

Example prompts

  • “/check-reporting”

Requirements

  • Python 3

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Existing-checklist staleness pre-check
  2. Select Guideline
  3. Load Checklist
  4. Scan Manuscript
  5. Assess Each Item
  6. Generate Report

What it can do on your machine

Read from SKILL.md and the folder at commit 3b14ae2. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/, which the agent can run.

    Shell commands in SKILL.md call:

    • python3
    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Check Reporting loads about 7.9k tokens when it runs, and up to ~117k if it reads all its reference files. Until then it costs about 79 tokens; SKILL.md has 3,741 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~79
When it runs · the whole SKILL.md, loaded when a task matches
~7.9k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~117k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from Aperivue/medsci-skills at commit 3b14ae2, republished under its MIT licence (© Aperivue). 3,741 words, ~7,860 tokens.

Download SKILL.mdSave it as .claude/skills/check-reporting/SKILL.md (or your agent's skills folder). This skill also uses 85 other files; get the full folder from GitHub.
name
check-reporting
description
Use when auditing a manuscript item by item against a reporting guideline or risk-of-bias tool. Covers 49 reporting guidelines and risk-of-bias tools (STROBE, CONSORT, STARD, TRIPOD+AI, PRISMA, QUADAS and more), marking each item PRESENT, PARTIAL or MISSING. Not a reviewer critique (/self-review).
metadata.triggers
checklist, QUADAS-3, abstract checklist, structured abstract, reporting guideline, STROBE, STROBE-MR, Mendelian randomization, CONSORT, CONSORT-AI, STARD…

Check-Reporting Skill

Reference Files

Checklists are vendored under ${CLAUDE_SKILL_DIR}/references/checklists/, one file per instrument. Each file's header gives its version, source citation and licence, and says when its item text is an own-words summary rather than the published wording.

  • STROBE.md -- STROBE
  • STROBE_MR.md -- STROBE-MR 2021
  • RECORD.md -- RECORD 2015 (RECORD-PE for drug studies)
  • REMARK.md -- REMARK
  • TARGET.md -- TARGET 2025
  • GATHER.md -- GATHER 2016
  • CHEERS_2022.md -- CHEERS 2022
  • CROSS.md -- CROSS 2021 + CHERRIES (internet surveys)
  • SRQR.md -- SRQR 2014 (all qualitative approaches)
  • COREQ.md -- COREQ 2007 (interviews and focus groups)
  • STARD.md -- STARD 2015
  • STARD_AI.md -- STARD-AI 2025
  • TRIPOD.md -- TRIPOD 2015
  • TRIPOD_AI.md -- TRIPOD+AI 2024
  • TRIPOD_LLM.md -- TRIPOD-LLM 2025
  • PGS_RS.md -- PGS-RS / PRS-RS 2021
  • CONSORT.md -- CONSORT 2025
  • CONSORT_AI.md -- CONSORT-AI 2020
  • SPIRIT.md -- SPIRIT 2025
  • SPIRIT_AI.md -- SPIRIT-AI 2020
  • CLAIM_2024.md -- CLAIM 2024
  • DECIDE_AI.md -- DECIDE-AI 2022
  • MI_CLEAR_LLM.md -- MI-CLEAR-LLM
  • CLEAR.md -- CLEAR (radiomics)
  • ARRIVE_2.md -- ARRIVE 2.0
  • CARE.md -- CARE 2013
  • SQUIRE_2.md -- SQUIRE 2.0
  • GRRAS.md -- GRRAS
  • PRISMA_2020.md -- PRISMA 2020
  • PRISMA_2020_Abstracts.md -- PRISMA 2020 for Abstracts, 12 items. A separate instrument, not a subset of the 27-item checklist (main item 2 defers to it): score it with its own denominator.
  • PRISMA_DTA.md -- PRISMA-DTA
  • PRISMA_P.md -- PRISMA-P
  • PRISMA_ScR.md -- PRISMA-ScR
  • MOOSE.md -- MOOSE
  • SWiM.md -- SWiM
  • AMSTAR2.md -- AMSTAR 2
  • QUADAS3.md -- QUADAS-3
  • QUADAS2.md -- QUADAS-2
  • QUADAS_C.md -- QUADAS-C
  • RoB2.md -- RoB 2
  • ROBINS_I.md -- ROBINS-I
  • ROBINS_E.md -- ROBINS-E
  • ROBIS.md -- ROBIS
  • ROB_ME.md -- ROB-ME
  • RoB_NMA.md -- RoB NMA
  • PROBAST.md -- PROBAST
  • PROBAST_AI.md -- PROBAST+AI
  • NOS.md -- NOS
  • COSMIN_RoB.md -- COSMIN RoB

Workflow

Step 0: Existing-checklist staleness pre-check

If a checklist already exists for this project (qc/reporting_checklist.json or a prior .md report), verify it targets the current manuscript before reusing it — one generated against an older version carries stale section/line references and a stale version label:

bash
python3 "${CLAUDE_SKILL_DIR}/scripts/check_checklist_version.py" \
  --checklist qc/reporting_checklist.json --manuscript manuscript_v8.md

A non-zero exit means the checklist is stale (older target_version, changed source_sha256, different target_manuscript) or pre-dates the version contract: regenerate it against the current manuscript (Steps 1–5) rather than reusing it. Every report you generate carries the target_manuscript / target_version / source_sha256 fields (Part A header + Part D JSON) so this check works next round.

Step 1: Select Guideline

Use the guideline the user names; otherwise auto-detect it from the study type. If the study type is ambiguous, ask the user to confirm before selecting a guideline.

Auto-detection mapping:

Study TypePrimary GuidelineAI Extension
Observational studySTROBE--
Mendelian randomization studySTROBE-MR (base STROBE + MR extension)--
Health economic evaluation (cost-effectiveness / cost-utility / cost-benefit / budget-impact)CHEERS 2022--
Observational study using routinely-collected data (claims / EHR / registry / health-checkup DB)RECORD (base STROBE + RECORD extension; RECORD-PE for drug studies)--
Survey / questionnaire study (KAP, physician/patient, cross-sectional, e-survey)CROSS (+ CHERRIES for internet surveys)--
Scoping review (maps breadth/nature of evidence, clarifies concepts, identifies gaps — not a focused effectiveness/accuracy question)PRISMA-ScR (base PRISMA + scoping-review extension)--
Qualitative study (interviews, focus groups, ethnography, grounded theory, phenomenology, document analysis)SRQR (all qualitative approaches); COREQ (interviews/focus groups specifically)--
Randomized controlled trialCONSORT 2025CONSORT-AI
Diagnostic accuracy studySTARD 2015STARD-AI
Prediction model (development/validation)TRIPODTRIPOD+AI
Polygenic (risk) score prediction studyPGS-RS (with TRIPOD / TRIPOD+AI)--
Prognostic tumor-marker / biomarker study (single or multiple markers; e.g., ctDNA / molecular residual disease)REMARK (pair with STROBE for the observational-design items; TRIPOD / TRIPOD+AI if a prognostic model is developed)--
Causal / comparative-effectiveness question emulated on observational data (treatment vs treatment, screening vs none, drug A vs B on registry / EHR / claims data)TARGET (pair with the /design-study target-trial-emulation module for design; RECORD / STROBE for the routinely-collected-data items)--
Health-estimate / burden-of-disease modeling study (GBD or GBD-satellite, comparative-risk / population-attributable-fraction, cause-of-death or prevalence/incidence estimation, with or without forecasts)GATHER (pair with /analyze-stats burden-decomposition-forecasting guide for the analytic layer)--
Systematic review / meta-analysisPRISMA 2020PRISMA 2020 for Abstracts (run on the abstract, scored separately)
DTA systematic review / meta-analysisPRISMA-DTAPRISMA 2020 for Abstracts (run on the abstract, scored separately)
Meta-analysis of observational studiesMOOSEPRISMA 2020 (use both)
Risk of bias (DTA studies)QUADAS-3 (current recommended version)QUADAS-2 only when appraising or reproducing a review that used it
Risk of bias (RCTs)RoB 2--
Risk of bias (non-randomised intervention studies)ROBINS-I--
Risk of bias (non-randomised exposure studies)ROBINS-E--
Risk of bias (comparative DTA studies)QUADAS-CQUADAS-3 (use both; apply the E&E's adaptation — see Using QUADAS-C with QUADAS-3 in QUADAS3.md)
Risk of bias (prediction models)PROBASTPROBAST+AI
Risk of bias (systematic reviews)ROBISAMSTAR 2
Risk of bias (missing evidence in MA)ROB-ME--
Risk of bias (network meta-analysis)RoB NMA--
Risk of bias (measurement properties)COSMIN RoB--
Quality assessment (observational)NOS--
Case reportCARE--
Study protocolSPIRIT 2025SPIRIT-AI
Animal studyARRIVE 2.0--
AI/ML study in clinical imagingCLAIM 2024--
Study using a large language model (develop/fine-tune/prompt/evaluate an LLM)TRIPOD-LLMMI-CLEAR-LLM (use alongside when LLM accuracy is an outcome)
Early-stage / live clinical evaluation of an AI decision-support system (human factors, workflow, safety)DECIDE-AI--
LLM accuracy evaluation in healthcareMI-CLEAR-LLMSTARD-AI or CLAIM 2024 (use alongside)
Reliability / agreement studyGRRAS--
SR protocolPRISMA-P--
Synthesis without meta-analysisSWiMPRISMA 2020 (use both)
Quality of systematic reviewsAMSTAR 2ROBIS
Radiomics studyCLEARCLAIM 2024 (if deep learning component)
Educational / QI studySQUIRE 2.0--
Generative AI images ARE the study object (realism / real-vs-synthetic reader study / model-vs-model quality)(no single guideline -- assemble)see decision aid below

QUADAS-3 has two protocol-stage phases, and this skill usually runs too late for them. Phase 1 (state the synthesis question) and phase 2 (define the ideal test accuracy trial each judgement is made against) are review-level and belong in the protocol, with the review-specific guidance for answering each signalling question. Reaching them for the first time during manuscript QC means writing the comparator after seeing the results. If they are missing, say so as a limitation rather than reconstructing them, and route the protocol work to /meta-analysis Phase 1. Phases 3–6 are what a QC pass can genuinely run.

Rules:

  • A guideline the user names is the one you score. The rules below choose the guideline only when the user names none. If they call for a different or additional instrument (STARD-AI for an AI index test, TRIPOD+AI for an ML prediction model, an AI extension), score the named guideline and put one line at the top of the report saying which instrument fits better and the command to rerun, e.g. /check-reporting STARD-AI. Do not switch or add instruments on your own: the user maps the report to the checklist form the journal asked for.
  • If the study involves AI/ML, apply the AI version of the guideline. CONSORT-AI and SPIRIT-AI are scored alongside their base guideline; TRIPOD+AI, STARD-AI and TRIPOD-LLM replace theirs (below).
    • Exception — TRIPOD: TRIPOD+AI 2024 (Collins et al., BMJ 2024) is a complete rewrite, not an addendum to TRIPOD 2015 (Moons et al., Ann Intern Med 2015). For non-AI prediction models, use TRIPOD 2015 only. For AI/ML prediction models, use TRIPOD+AI 2024 only. Do NOT apply both simultaneously.
  • STARD-AI (Sounderajah et al., Nat Med 2025) extends STARD 2015 with 14 new and 4 modified items (40 total) and incorporates all STARD 2015 items. For AI diagnostic accuracy studies use STARD-AI only — do NOT apply STARD 2015 and STARD-AI simultaneously.
  • TRIPOD-LLM (Gallifant et al., Nat Med 2025) is the reporting guideline for studies that develop, fine-tune, prompt, or evaluate a large language model for a clinical/biomedical task. It extends the TRIPOD family (TRIPOD 2015 → TRIPOD+AI 2024 → TRIPOD-LLM 2025); name the base instrument and the extension and cite each. It is modular — task-specific items (Annotation, Prompting, Summarization, Instruction-tuning) are N/A when that component is absent. Use TRIPOD-LLM for LLM studies in place of TRIPOD+AI; pair with MI-CLEAR-LLM when LLM accuracy is an evaluated outcome. The vendored checklist is an educational summary (own-words paraphrase of item intent); complete the official instrument for a submission checklist.
  • MI-CLEAR-LLM is a supplementary checklist (8 item categories in the 2025 update; the 2024 original had 6), not a standalone reporting guideline. Always pair it with the study's primary guideline (e.g., STARD-AI for AI diagnostic accuracy, CLAIM for imaging AI). Apply it whenever the study evaluates LLM accuracy as an outcome — do NOT apply it merely because the manuscript was written with LLM assistance. Its scope is LLM accuracy studies (including VLMs interpreting images); it does not apply at study level when a generative model produces the images under study (next bullet).
  • Generative-AI images as the study object (a generative model synthesizes images and the study evaluates their realism, controllability, real-vs-synthetic distinguishability, or model-vs-model quality) has no single dominant checklist. Assemble: CLAIM 2024 (imaging-AI umbrella; model-development items N/A when commercial models are used as-is) + FUTURE-AI traceability + MI-CLEAR-LLM transparency items only (prompt/model/version/params/runs — for generation provenance, not study-level compliance) on the generator side; STARD-AI (for real-vs-synthetic detection) + GRRAS (reader reliability) + MRMC reporting on the evaluation side. Map applicable items and cite base + extension; never claim wholesale compliance. Full decision aid: ${CLAUDE_SKILL_DIR}/references/genai_image_study_object_decision_aid.md.
  • If multiple guidelines apply (e.g., a diagnostic accuracy study that is also an AI study), check against all relevant guidelines and merge into one report.
Step 2: Load Checklist
  1. Run the fail-fast guard first for every guideline you intend to apply:

    bash
    python "${CLAUDE_SKILL_DIR}/scripts/check_checklist_exists.py" --guideline "STARD-AI"
    • Exit 0 → the vendored checklist exists; read it from ${CLAUDE_SKILL_DIR}/references/checklists/ and proceed.
    • Exit 1 (MISSING_CHECKLIST_CONTRACT_VIOLATION) → the guideline is routed but no checklist file is vendored. Do not construct items from memory. Halt, report the violation to the user, and stop unless they explicitly opt in (next bullet).
    • Exit 2 (UNKNOWN_GUIDELINE) → the name is not recognised; confirm the correct guideline with the user.
  2. No silent fallback. A from-memory checklist is permitted only when the user explicitly accepts it — re-run the guard with --allow-from-memory (exit 0 + a NON-AUTHORITATIVE warning). The report MUST then carry a prominent banner that the assessment was constructed from model knowledge and is not backed by a vendored checklist, and submission_safe must not be asserted on its basis.

Step 3: Scan Manuscript

Read the whole manuscript before assessing any item — including tables, figures and their captions, supplementary material, and the reference list (registration numbers and protocol references often sit there).

Step 4: Assess Each Item

Items most often missing in medical manuscripts — look for these first, whichever guideline applies: registration number and registration/amendment date consistency (run Step 4c), sample-size justification, missing-data handling, blinding, funding and conflicts of interest, ethics approval with committee name and approval number, and a data availability statement; for AI studies, the training/validation/test split, model architecture and hyperparameters, failure-mode analysis, fairness/bias assessment, and commercial interests with data/code availability.

For every checklist item, determine:

StatusCriteria
PRESENTThe item is fully addressed with sufficient detail.
PARTIALThe item is mentioned or partially addressed but lacks required detail.
MISSINGThe item is not found anywhere in the manuscript.
N/AThe item does not apply to this particular study (justify why).

For each item, record:

  • Status: PRESENT / PARTIAL / MISSING / N/A
  • Location: Section name and paragraph or approximate position (e.g., "Methods, paragraph 3")
  • Notes: What was found (if PRESENT/PARTIAL) or what should be added (if MISSING)

Be strict. PARTIAL means the item is mentioned but lacks specificity; a vague reference does not count as PRESENT — the detail level must match what the guideline expects. "We used appropriate statistical tests" = PARTIAL (which tests?); "We used the Mann-Whitney U test for continuous variables and Fisher's exact test for categorical variables" = PRESENT. If an item is genuinely unclear in its applicability, mark it N/A with justification.

Two gaps that are easy to pass:

  • Power-aware framing of a null result (STROBE 16a / 18 / 20) — for an observational study whose headline is a non-significant association, a flat "X was not associated with Y" overreads the data when the analysis is not powered to exclude a clinically meaningful effect. Mark item 18/20 PARTIAL unless the manuscript states the precision as an exclusion (e.g., "the 95% CI excluded an eGFR difference larger than ~1.7") or reports a minimum detectable effect — "no effect" vs "could not exclude an effect of size X" are different claims, and a negative conclusion needs the latter.
  • Confounder-selection rationale, not "adjust for everything that differs" (STROBE 16a explicitly asks which confounders were adjusted for and why) — flag a kitchen-sink adjustment set chosen because variables differ in Table 1. The Methods must give a causal rationale (DAG / prior literature) and must not adjust for a mediator or consequence of the outcome (over-adjustment, e.g. serum uric acid in an eGFR model); both an unjustified inclusion and an unjustified omission are item-16a gaps.

What is appraised is the source paper's reporting — never your convenience in using it. This holds for every instrument here, reporting checklists and risk-of-bias / quality tools alike, and is easiest to lose in a systematic review, where you read each paper in order to extract from it. An item asking "are the results clearly reported?" is not asking "were they reported in the unit my pool needs". If a downgrade's stated reason turns on a denominator, an analysis unit, a subgroup you needed and they did not report separately, or a format you could not parse, it is an extraction note, not a scoring reason: record it in a separate extraction-note column and restore the score. An extraction limitation often belongs in your limitations paragraph; folded into the score it makes the appraisal unreproducible, because another assessor with a different pool would score the same paper differently.

Step 4b: Section Boundary Check

In addition to checklist items, verify that:

  • Results section contains only factual findings: no interpretation, no "why" explanations, no prior literature comparisons, no evaluative adjectives without numbers.
  • Discussion section does not introduce new data not presented in Results.
  • Flag any boundary violation as a separate finding in Part C Action Items with the label [BOUNDARY].
Show full SKILL.md (1,551 more words)Show less
Step 4c: Registration / Protocol Timing Consistency Check

Applies to: systematic reviews, meta-analyses, and intervention studies with prospective registration (PRISMA 2020, PRISMA-DTA, PRISMA-P, MOOSE, CONSORT, SPIRIT). The registration identifier is a single checklist item and can pass Step 4 while the manuscript is inconsistent about when registration and amendments happened relative to the analysis.

Read ${CLAUDE_SKILL_DIR}/references/step4c_registration_timing.md (item-by-item procedure, JSON schema, flagging edge cases) and run its five checks: (1) registration identifier present in Methods, Abstract, and cover letter; (2) initial registration date precedes — or is explicitly disclosed as post-dating — the extraction milestone; (3) amendment dates appear in Methods, the described change is visible in Methods, analysis was re-run if the amendment post-dates the lock, and no amendment post-dates submission; (4) Methods agree with the registry record (PROSPERO PDF, ClinicalTrials.gov export) — a silent discrepancy is a finding; (5) a retrospective-registration disclosure paragraph when evidence suggests post-extraction filing.

Flagging: any failure is logged in Part C Action Items with label [REGISTRATION-TIMING]. fixable_by_ai: false when reconciliation requires an external amendment filing; true only when the fix is a Methods-text insertion of a date already disclosed elsewhere. Part D JSON includes a registration_timing object (registry, id, initial_registration_date, amendments[], timing_consistency, findings[]).

Registration-ID format gate: a PROSPERO ID is CRD42 + 9 digits = 14 characters (^CRD42\d{9}$, e.g. CRD42024500001). Run grep -oE 'CRD42[0-9]+' manuscript.md and assert each match is 14 characters long; a 15-character ID (a stray inserted digit) is a transcription error logged as [REGISTRATION-TIMING] (fixable_by_ai: false — verify against the live PROSPERO record, do not guess the correct digit).

Step 4d: PRISMA Figure 1 Arithmetic & Cross-Reference Audit

Applies to: systematic reviews and meta-analyses using PRISMA 2020 / PRISMA-DTA / PRISMA-P. Triggers when Item 16a (flow diagram) is PRESENT — the diagram can pass Step 4 while its numbers do not add up or disagree with the text.

  1. Choose the Figure 1 source, in this order: (a) analysis/figures/Figure1_PRISMA.md markdown manifest, (b) caption text in manuscript.md, (c) PPTX text run if a .pptx exists, (d) manual entry from PNG/SVG.

  2. Run the audit. It checks the five subtractions (screened = identified − duplicates; sought-for-retrieval = screened − excluded at screening; retrieved = sought − not retrieved; assessed for eligibility = sought − not retrieved; included = assessed for eligibility − excluded with reasons) and that the body-text PRISMA numbers match the Figure 1 boxes 1:1, and writes qc/prisma_figure_audit.json:

    bash
    python3 ${CLAUDE_SKILL_DIR}/scripts/check_prisma_figure.py \
      --md <manuscript.md> --figure <Figure 1 source: .md manifest / caption / text export> \
      --out qc/prisma_figure_audit.json

    Exit 1 = an arithmetic or cross-reference MISMATCH; exit 2 = missing/unparsable input. When the Figure 1 numbers exist only in a PNG/SVG, transcribe them by hand and run the same checks manually as ${CLAUDE_SKILL_DIR}/references/step4d_prisma_figure_audit.md specifies (regex set, JSON schema, and edge cases: duplicates across databases, the citation-searching strand, dual-reviewer screening).

  3. By hand (the script does not do this): the reasons for exclusion in Methods and the Figure legend must agree on counts and category names, and when identification is split across sources (databases, registers, other methods) the per-source counts must sum to the identified total — the script reads only the first identified count.

  4. If analysis/figures/_figure_manifest.md (from /make-figures) exists, verify that the row whose Type = prisma (or Type = prisma-dta) points at the same file used as the audit source, and that its Critic field is yes or partial (not no). A missing row, mismatched path, or Critic = no logs [MANIFEST-XREF] (advisory); the arithmetic check still runs.

Flagging: any MISMATCH or arithmetic failure logs a Part C Action Item with label [PRISMA-FIGURE], fixable_by_ai: false (the author must reconcile the numbers).

PRISMA Cascade Arithmetic Auto-Verify

When PRISMA 2020 or PRISMA-DTA is selected and round-by-round screening TSV artifacts are available, recompute the screening cascade from the raw decisions. Off-by-one errors in the prose cascade are a high-frequency reviewer red flag (e.g., 151 + 108 + 39 + 1 + 1 + 4 = 304 followed by a prose summary "305" four lines later).

bash
python "${CLAUDE_SKILL_DIR}/scripts/prisma_cascade_check.py" \
    --round1 2_Screening/round1.tsv \
    --round2 2_Screening/round2.tsv \
    --round3 2_Screening/round3_adjudication.tsv \
    --manuscript manuscript.md \
    --out qc/prisma_cascade.json --strict

The script counts INCLUDE / EXCLUDE / MAYBE decisions per round, computes the cascade, and reports per-stage drift where the manuscript's stage counts disagree. Treat any manuscript_drift entry as a P0 blocker — fix the prose to match the computed cascade and re-run. A --manuscript path that is not a file exits 2. When the manuscript is read but none of the script's stage phrases is found in it, manuscript_check is "unverifiable", submission_safe is false, an UNVERIFIED line is printed, and --strict exits 1: the drift check did not run, so compare the prose stage counts to stage_counts by hand. stages_compared lists the stages that were actually checked.

Step 4e: Reporting-Framework Naming Audit

Applies to: any manuscript that invokes an AI/extension reporting framework (PROBAST+AI, STARD-AI, TRIPOD+AI, TRIPOD-LLM, CONSORT-AI, SPIRIT-AI, PRISMA-DTA, QUADAS-C). A base reporting tool and its extension are distinct instruments with separate citations, and Step 1 does not police how the framework is named in prose. The recurring failures: invoking an extension without ever naming or citing the base instrument it extends; mixing +AI and -AI hyphenation for one family within a single document; coining item labels like "12-AI"; and waving at "recent guidance" instead of naming the framework.

bash
python3 "${CLAUDE_SKILL_DIR}/scripts/check_framework_naming.py" \
  --manuscript manuscript.md --out qc/framework_naming.json --strict

Verdicts: BASE_MISSING (extension used, base instrument never named standalone) is a Major and logs [FRAMEWORK-NAMING] in Part C with fixable_by_ai: true (insert the base name + its citation). HYPHEN_MIX, CITE_MISSING, SELF_COINED_LABEL, and VAGUE_GUIDANCE are Minor (fixable_by_ai: true). Part D JSON includes a framework_naming object mirroring the script's claims[].

Step 4f: Critical-item floor cross-check

Applies to: every guideline assessment for which the floor defines a row (load and check only those; do not invent a floor for an unlisted guideline). After the item-by-item table, read ${CLAUDE_SKILL_DIR}/references/critical_item_floor.md and check the small set of non-waivable items for this study type — and, for AI/ML and radiomics manuscripts, the methodological-quality / risk-of-bias instrument (PROBAST+AI, METRICS/RQS, APPRAISE-AI) and concerns it lists. A MISSING critical item is surfaced as a Critical gap and becomes the report's headline regardless of the overall percentage — a high percentage with a missing critical item (undefined reference standard, no leakage-controlled partition, calibration absent for a prediction model, an unreconciled flow diagram) is not "broadly acceptable."

Step 5: Generate Report

This report is an internal working audit — it carries auto-fix annotations, a machine-readable JSON block (compliance_pct, fixable_by_ai, …), and Action Items. It is NOT the official reporting checklist a journal expects (that is the blank guideline form with Item | Recommendation | Reported in page/section, which the authors fill in). Never submit this report as the submission checklist. So that the file is self-identifying and cannot be reused by filename into a later submission package, the report MUST begin with this banner as its very first line:

<!-- INTERNAL AUDIT — NOT FOR SUBMISSION. This is the /check-reporting working
report, not the official journal checklist. Do not upload to a submission portal. -->

Write the checklist content and report in English (matching the guideline originals), whatever language you use with the user. Save it as qc/reporting_checklist.md and the Part D JSON as qc/reporting_checklist.json.

Read ${CLAUDE_SKILL_DIR}/references/report_templates.md when you write the report — it holds the literal templates for the four parts:

  • Part A — Summary. Header (manuscript file, version token, guideline, date), the PRESENT/PARTIAL/MISSING/N-A count table, and overall compliance. The headline is the critical items (Step 4f), not the percentage: report {present}/{total} and name every missing critical item with the section it belongs in.
  • Part B — Item-by-item checklist. One row per item: # | Section | Item | Status | Location | Notes.
  • Part C — Action items (MISSING and PARTIAL only), ordered by: items most journals enforce strictly (ethics approval, registration, sample size) → items in Methods (easiest to fix) → everything else.
  • Part D — Machine-readable JSON, appended as a fenced block. MUST be present under --json or when called from /write-paper Phase 7, which parses it.

JSON field contract (the part other skills depend on — get these right):

  • compliance_pct — present / (total_items - na) * 100, one decimal; null when every item is N/A (total_items == na).
  • action_items — MISSING and PARTIAL only; PRESENT and N/A are excluded.
  • fixable_by_ai — true when the fix inserts or expands text using information already in the manuscript or inferable from it; false when it needs external facts the author alone holds (registration number, IRB approval number, protocol details, sample-size rationale and target).
  • suggested_fix — concrete draft text, insertable as written. A fix that still contains a bracketed placeholder ([N], [rationale]) is not insertable and is fixable_by_ai: false.
  • source_sha256 — first 12 hex chars of the SHA-256 of the manuscript bytes, so a stale report cannot be silently attributed to a newer manuscript.

In notes and suggested fixes, cite a reference only with a /search-lit-confirmed DOI or PMID and mark any other [UNVERIFIED - NEEDS MANUAL CHECK]; never invent clinical definitions, diagnostic criteria, or guideline recommendations — flag anything uncertain [VERIFY] and ask the user.

When the user asks for the filled checklist a journal requires, build it from the Submission checklist export template in ${CLAUDE_SKILL_DIR}/references/report_templates.md.


Skill Interactions

WhenCallPurpose
During manuscript writing/write-paper Phase 7Final compliance check
Need to add Methods text/write-paper Phase 3Draft missing Methods content
Need statistical details/analyze-statsGenerate missing statistical reporting
Need flow diagram/make-figuresGenerate CONSORT/STARD/PRISMA diagram

Gates

GateSeverityTriggerAction on fail
Mandatory items presentENFORCED at submission< 100% of guideline-mandatory items marked PRESENTAuto-fix MISSING items where text exists; otherwise route to /write-paper Phase 7 for re-draft
Step 4d PRISMA Figure 1 arithmetic & cross-reference audit (PRISMA / PRISMA-DTA only)ENFORCED for SR/MAflow numbers don't sum (e.g., screened ≠ included + excluded), or in-text counts mismatch flow diagramHALT; reconcile against extraction artifacts
Optional items (e.g., supplementary AI declarations)ADVISORY< 80% of optional items presentwarn; user accepts
Cross-reporting-guideline routing (study type → guideline)ENFORCEDstudy type undeclared or guideline missingAsk user; do not silently default

© Aperivue, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 85 other files (scripts, references) in skills/check-reporting of Aperivue/medsci-skills.

  • SKILL.md
  • references/LICENSES.md
  • references/appraisal_tools/METRICS.md
  • references/appraisal_tools/METRICS_RELOADED.md
  • references/checklists/AMSTAR2.md
  • references/checklists/ARRIVE_2.md
  • references/checklists/CARE.md
  • references/checklists/CHEERS_2022.md
  • references/checklists/CLAIM_2024.md
  • references/checklists/CLEAR.md
  • references/checklists/CONSORT.md
  • references/checklists/CONSORT_AI.md
  • references/checklists/COREQ.md
  • references/checklists/COSMIN_RoB.md
  • references/checklists/CROSS.md
  • references/checklists/DECIDE_AI.md
  • references/checklists/GATHER.md
  • references/checklists/GRRAS.md
  • … and 68 more

Open the folder on GitHubat commit 3b14ae2

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in Aperivue/medsci-skills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Check Reporting next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Check Reporting compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Check Reporting this skillAperivue/medsci-skills3291 repos~7.9kAutomated safety check: PassMIT
Content Create Hero Imageprisma/web1.1k—~6.9kAutomated safety check: PassNone
Prisma Client APIcurvenote/curvenote1692 repos~1.6kAutomated safety check: PassMIT
Docs Writerprisma/web1.1k—~5.2kAutomated safety check: NotesNone
Prismablencorp/claude-code-kit106—~2.7kAutomated safety check: PassMIT
Prisma Expertdavila7/claude-code-templates32k7 repos~2.6kAutomated safety check: PassMIT

Similar skills

  • Official

    A skill your agent uses when the operator wants a hero or meta image for a Prisma blog post; asks to create or generate a blog hero, cover, social card, Open Graph, or YouTube image; mentions cover…

    1.1k GitHub stars~6.9k tokensUpdated today
    DatabasesAuto-check passed
  • Prisma Client API

    curvenote/curvenote

    Prisma Client API reference covering model queries, filters, operators, and client methods.

    169 GitHub starsUsed in 2 repos~1.6k tokens
    DatabasesAuto-check passed
  • Docs Writer

    prisma/web

    Official

    A skill your agent uses when writing, rewriting, or improving technical docs (quickstarts, how-tos, tutorials, concept pages, or API references).

    1.1k GitHub stars~5.2k tokensUpdated today
    DatabasesAuto-check: notes
  • Prisma

    blencorp/claude-code-kit

    Prisma ORM patterns including Prisma Client usage, queries, mutations, relations, transactions, and schema management.

    106 GitHub stars~2.7k tokensUpdated 10 mo ago
    DatabasesAuto-check passed
  • Prisma Expert

    davila7/claude-code-templates

    Prisma ORM expert for schema design, migrations, query optimization, relations modeling, and database operations.

    32k GitHub starsUsed in 7 repos~2.6k tokens
    DatabasesAuto-check passed
  • Survey Methodology

    Muuuun/luxas

    Plan-time methodology contract for survey/review/report projects.

    1.2k GitHub stars~8.1k tokensUpdated 1 mo ago
    DatabasesAuto-check passed

More from Aperivue/medsci-skills

All 54 skills in this repo
  • Model Assessment

    Aperivue/medsci-skills

    A skill your agent uses when validating or evaluating a trained medical-imaging model.

    329 GitHub starsUsed in 1 repo~4.5k tokens
    Auto-check passed
  • Obsidian Paper Vault

    Aperivue/medsci-skills

    A skill your agent uses when turning a folder of research PDFs into Obsidian notes, even if Obsidian is not named.

    329 GitHub stars~1.6k tokensUpdated 3 days ago
    Auto-check passed
  • Radiomics ML

    Aperivue/medsci-skills

    A skill your agent uses when building or auditing a radiomics or tabular clinical-ML prediction model with a classical learner (LASSO, SVM, random forest, XGBoost and similar).

    329 GitHub starsUsed in 1 repo~2.7k tokens
    Auto-check passed
  • Verify Refs

    Aperivue/medsci-skills

    A skill your agent uses when checking whether a manuscript's references are real.

    329 GitHub starsUsed in 1 repo~3.1k tokens
    Auto-check passed
  • Clean Data

    Aperivue/medsci-skills

    A skill your agent uses when a clinical CSV/Excel dataset needs profiling and cleaning before analysis (missing values, outliers, duplicates, type mismatches).

    329 GitHub stars~2k tokensUpdated 3 days ago
    Auto-check passed
  • Design Study

    Aperivue/medsci-skills

    A skill your agent uses when checking a radiology or medical AI study design before drafting or submission.

    329 GitHub stars~3.9k tokensUpdated 3 days ago
    Auto-check passed

Works with

Categories

Questions about Check Reporting

What does Check Reporting do?

A skill your agent uses when auditing a manuscript item by item against a reporting guideline or risk-of-bias tool. Check Reporting is an agent skill from Aperivue/medsci-skills. Use when auditing a manuscript item by item against a reporting guideline or risk-of-bias tool.

When should I use Check Reporting?

Check Reporting fits situations like: auditing a manuscript item by item against a reporting guideline; risk-of-bias tool.

How do I install Check Reporting in Claude Code?

Run `npx skills add Aperivue/medsci-skills --skill check-reporting -a claude-code`. Or copy the skill folder (skills/check-reporting in Aperivue/medsci-skills) into .claude/skills/check-reporting in your project. Claude Code loads it when a task matches its description.

How do I install Check Reporting in Codex?

Run `npx skills add Aperivue/medsci-skills --skill check-reporting -a codex`. Or copy the skill folder (skills/check-reporting in Aperivue/medsci-skills) into .agents/skills/check-reporting in your project. Codex loads it when a task matches its description.

Can I use Check Reporting in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Aperivue/medsci-skills --skill check-reporting -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/check-reporting, .gemini/skills/check-reporting, .github/skills/check-reporting and .opencode/skills/check-reporting in your project.

What does Check Reporting need to run?

Going by SKILL.md and its folder, Check Reporting needs the command-line tools its instructions call (python3 and python). Our summary lists: Python 3.

Does Check Reporting access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Check Reporting safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Check Reporting use?

Check Reporting is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Check Reporting use?

About 7.9k tokens (SKILL.md is roughly 31k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 109k tokens, read only when the agent opens those files.

What are the alternatives to Check Reporting?

Skills that share tags, products or a category with Check Reporting: Content Create Hero Image (prisma/web, 1.1k stars), Prisma Client API (curvenote/curvenote, 169 stars), Docs Writer (prisma/web, 1.1k stars) and Prisma (blencorp/claude-code-kit, 106 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Check Reporting?

Aperivue (a GitHub organization) maintains it in Aperivue/medsci-skills, which has 329 GitHub stars. The repository holds 54 skills in this directory. The repository was last updated on October 5, 2026.

Source: Aperivue/medsci-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.