Agent skill

Bio Clinical Biostatistics Multiplicity Graphical

by GPTomics in GPTomics/bioSkills

Implements multiplicity control for confirmatory clinical trials using graphical procedures (Bretz-Maurer-Hommel), gatekeeping (parallel, serial, mixed), Hochberg/Hommel/Holm with PRDS, and the…

MITAuto-check passedResearch & Science

Install Bio Clinical Biostatistics Multiplicity Graphical

skills CLI
$ npx skills add GPTomics/bioSkills --skill bio-clinical-biostatistics-multiplicity-graphical -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install GPTomics/bioSkills bio-clinical-biostatistics-multiplicity-graphical --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/clinical-biostatistics/multiplicity-graphical .claude/skills/bio-clinical-biostatistics-multiplicity-graphical && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
bio-clinical-biostatistics-multiplicity-graphical
GitHub stars
1.2k
Used in
2 other repos
Token cost
~6.2k tokens
SKILL.md length
2,711 words
Files
3
Skills in repo
559
Repo updated
First seen
Licence
MIT

At a glance

Implements multiplicity control for confirmatory clinical trials using graphical procedures (Bretz-Maurer-Hommel), gatekeeping (parallel, serial, mixed), Hochberg/Hommel/Holm with PRDS, and the…

  • Designing the multiplicity strategy for confirmatory trials with multiple primary
  • SKILL.md covers Version Compatibility, The Foundational Theorem --…, Algorithmic Taxonomy and Decision Tree by Scenario, plus 9 more sections
  • Runs R scripts from its folder; calls pip
  • Key secondary endpoints

What it does

Bio Clinical Biostatistics Multiplicity Graphical is an agent skill from GPTomics/bioSkills. Implements multiplicity control for confirmatory clinical trials using graphical procedures (Bretz-Maurer-Hommel), gatekeeping (parallel, serial, mixed), Hochberg/Hommel/Holm with PRDS, and the closed-testing principle (Marcus-Peritz-Gabriel; Goeman 2021 admissibility). Covers FDA Multiple Endpoints Final Guidance (October 2022), graphical procedures via R gMCP, primary + key-secondary + subgroup hierarchies, and FWER vs FDR distinction. Use when designing the multiplicity strategy for confirmatory trials with…

Its SKILL.md is about 6.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files (for example `usage-guide.md`).

It sits in Research & Science, covering Clinical and healthcare research and PRD writing. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.

When your agent uses it

  • Designing the multiplicity strategy for confirmatory trials with multiple primary
  • Key secondary endpoints

Example prompts

  • “Use the bio-clinical-biostatistics-multiplicity-graphical skill to implement multiplicity control for confirmatory clinical trials using graphical…”
  • “/bio-clinical-biostatistics-multiplicity-graphical”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (R), which the agent can run.

    Shell commands in SKILL.md call:

    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Bio Clinical Biostatistics Multiplicity Graphical loads about 6.2k tokens when it runs. Until then it costs about 153 tokens; SKILL.md has 2,711 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~153
When it runs · the whole SKILL.md, loaded when a task matches
~6.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 2,711 words, ~6,242 tokens.

Download SKILL.mdSave it as .claude/skills/bio-clinical-biostatistics-multiplicity-graphical/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
bio-clinical-biostatistics-multiplicity-graphical
description
Implements multiplicity control for confirmatory clinical trials using graphical procedures (Bretz-Maurer-Hommel), gatekeeping (parallel, serial, mixed), Hochberg/Hommel/Holm with PRDS, and the closed-testing principle (Marcus-Peritz-Gabriel; Goeman 2021 admissibility). Covers FDA Multiple Endpoints Final Guidance (October 2022), graphical procedures via R gMCP, primary + key-secondary + subgroup hierarchies, and FWER vs FDR distinction. Use when designing the multiplicity strategy for confirmatory trials with multiple primary or key secondary endpoints.
tool_type
r
primary_tool
gMCP
goal_approach_exempt
true

Version Compatibility

Reference examples tested with: R gMCP 0.8.16+, graphicalMCP 0.2+, gatekeeping, multcomp, multxpert; Python statsmodels 0.14+ for basic FDR/FWER methods.

Before using code patterns, verify installed versions match. If versions differ:

  • R: packageVersion('<pkg>') then ?function_name
  • Python: pip show <package> then help(module.function)

If code throws ImportError, AttributeError, or TypeError, introspect the installed package and adapt the example to match the actual API rather than retrying.

Multiplicity Control for Confirmatory Trials

"Design the multiplicity strategy for my trial" -> Specify a closed-testing procedure (graphical, gatekeeping, hierarchical, or step-down Bonferroni-Holm) that controls family-wise error rate at the trial-wide level across primary endpoints, key secondary endpoints, and subgroup analyses, with provable strong FWER control.

The Foundational Theorem -- Closed Testing Is Necessary

Marcus, Peritz & Gabriel 1976 Biometrika 63:655: a hypothesis H_I (I ⊆ {1,...,m}) is rejected iff every intersection hypothesis ∩_{J⊇I} H_J is rejected by a valid α-level local test. Strong FWER control holds for ANY choice of local tests.

Goeman, Hemerik & Solari 2021 Ann Stat 49:1218 tightens this: closed testing is not merely sufficient — it is necessary for admissibility under FDP/FWER/k-FWER. Every admissible multiplicity procedure is equivalent to some closed test. Graphical procedures, gatekeepers, Hommel, fixed-sequence, fallback — all are closed tests in disguise.

FWER vs FDR philosophical divide:

  • FWER: P(any false positive among m tests) — regulatory standard for confirmatory inference (agency wants to bound per-trial false-positive rate)
  • FDR: Expected proportion of false discoveries among rejections — exploratory standard (genomics, fMRI, biomarker screens) where many true positives expected

Confirmatory clinical trials use FWER essentially universally.

Algorithmic Taxonomy

ProcedureTypeFWER controlPower profileUse case
BonferroniSingle-stepYes, any dependenceConservative; loses 30-50% power vs Hommel under positive dependenceVery small m; worst-case dependence
Holm 1979Step-downYes, any dependenceBetter than Bonferroni; uniformly dominatesDefault for any dependence pattern
Hochberg 1988Step-upYes under PRDS (Sarkar 1998)Better than Holm under PRDSPositive correlation; verify PRDS
Hommel 1988Step-up via closed testsYes under PRDSUniformly dominates Hochberg by 1-3%Whenever Hochberg is valid
Fixed-sequence (hierarchical)SequentialYes, any dependenceFull alpha for first; subsequent zero if any failWhen clear priority ordering; "key secondary" labelling
Parallel gatekeeping (Dmitrienko 2003)Multi-familyYesFamily-by-family; secondary tested if any primary rejectsPrimary family + secondary family
Serial gatekeepingSequential familiesYesStrict: family k tested only if ALL of family k-1 rejectCo-primary + secondary tiers
Mixed gatekeeping (Dmitrienko-Tamhane 2008)CombinationYesCombines closed-testing local procedures across familiesComplex hierarchies
Graphical procedures (Bretz-Maurer 2009)Closed-test as directed graphYes by constructionFlexible; allocate alpha to hypotheses via graph weightsModern standard for confirmatory SAPs
Graphical + Simes/parametric (Bretz et al 2011)Closed-test with non-Bonferroni local testsYes when Simes validGains power under correlationComplex co-primary + key secondary + subgroup hierarchies
Maurer-Bretz 2013 entangled graphsMemory-augmented graphsYes by constructionAlpha propagation depends on originParent-descendant constraints
Benjamini-Hochberg 1995FDRFDR controlled at level qHigher power than FWERExploratory only; NOT for confirmatory regulatory

Postdoc reading list:

  • Marcus R, Peritz E, Gabriel KR 1976 Biometrika 63:655 (closed testing — the foundation)
  • Goeman JJ, Hemerik J, Solari A 2021 Ann Stat 49:1218 (closed testing necessary for admissibility)
  • Holm S 1979 Scand J Stat 6:65 (step-down Bonferroni)
  • Hochberg Y 1988 Biometrika 75:800 (step-up Simes)
  • Hommel G 1988 Biometrika 75:383 (closed Simes; dominates Hochberg)
  • Sarkar SK 1998/2008 Ann Stat (PRDS for Hochberg validity)
  • Bretz F, Maurer W, Brannath W, Posch M 2009 Stat Med 28:586 (graphical procedures — foundational paper)
  • Bretz F, Posch M, Glimm E, Klinglmueller F, Maurer W, Rohmeyer K 2011 Biom J 53:894 (Simes/parametric extensions)
  • Maurer W, Bretz F 2013 Stat Med 32:1739 (entangled graphs / memory)
  • Dmitrienko A, Offen WW, Westfall PH 2003 Stat Med 22:2387 (parallel gatekeeping)
  • Dmitrienko A, Tamhane AC, Wiens B 2008 Biom J (mixed/multistage gatekeeping)
  • FDA 2022 Multiple Endpoints in Clinical Trials Final Guidance (October 2022)
  • Pocock SJ, Ariti CA, Collier TJ, Wang D 2012 Eur Heart J (win-ratio)

Decision Tree by Scenario

ScenarioRecommended procedureWhy
2 co-primary endpoints (both must succeed)No alpha split needed; per-endpoint alpha-level test; cite FDA 2022Co-primary doesn't split alpha; inflates n via joint power
2 multiple primary endpoints (any-wins)Graphical procedure or Holm with weightsAlpha must be allocated; graphical is flexible
1 primary + 2 key secondary endpointsHierarchical (serial gatekeeping) OR graphical with alpha propagationModern SAPs favour graphical
1 primary + 3 secondary + 4 subgroup analysesGraphical procedure via gMCP with pre-specified weightsComplex hierarchies benefit from graph visualisation
Primary endpoint + tipping-point sensitivityNo multiplicity adjustment needed for sensitivitySensitivity is "what if" not "another claim"
Many exploratory biomarker subgroupsBenjamini-Hochberg FDRExploratory; not for label claims
Win-ratio composite (cardiology)Single test; no multiplicityComposite captures multiple events in single hierarchy
Subgroup analysis (pre-specified)Graphical alpha allocation; small budget (≤20% by convention); see Dane 2019 for subgroup disciplineConfirmatory subgroup discovery requires explicit allocation
Adaptive trial with treatment arm droppingCombination tests (Bauer-Köhne 1994) + closed testingSee clinical-biostatistics/adaptive-designs
Group-sequential with multiple endpointsgsDesign or rpact with multivariate alpha spendingHierarchical alpha across both time and endpoints

Bretz-Maurer Graphical Procedures -- The Modern Standard

The Bretz-Maurer-Brannath-Posch 2009 Stat Med 28:586 framework recast weighted Bonferroni-Holm closed tests as directed weighted graphs:

  • Vertices = elementary null hypotheses with local weights summing to 1
  • Directed edges = alpha-propagation rule (when a hypothesis is rejected, its weight redistributes to descendants per edge weights)
  • The graph IS the procedure: a single visual fully specifies a closed-test procedure across primary, key secondary, and subgroup hierarchies
gMCP R package
r
library(gMCP)

# Construct a graph for primary + 2 key secondary endpoints
# Primary endpoint at full alpha; if rejected, alpha propagates equally to secondaries
hypotheses <- c('Primary', 'Sec1', 'Sec2')
weights <- c(1, 0, 0)  # initial alpha all on primary
# Transition matrix: rows = source, columns = target
# When Primary rejects, weight 0.5 goes to each secondary; when Sec1/Sec2 rejects, alpha returns
transitions <- matrix(c(
    0,    0.5,  0.5,
    0,    0,    1,
    0,    1,    0
), nrow = 3, byrow = TRUE, dimnames = list(hypotheses, hypotheses))

graph <- graphMCP(m = transitions, weights = weights, hnames = hypotheses)
# Note: in current gMCP, the graph constructor is `graphMCP(m=, weights=, hnames=)`;
# `matrix2graph()` appeared in older tutorials and is not the canonical exported API
# -- verify with `?graphMCP` / `?gMCP` in the installed gMCP release before scripting.
# Set p-values from the trial
p_vals <- c(Primary = 0.018, Sec1 = 0.042, Sec2 = 0.038)

# Run the graphical procedure at alpha = 0.025
result <- gMCP(graph, pvalues = p_vals, alpha = 0.025)
print(result)
# Hierarchical rejection: Primary rejects -> alpha propagates to secondaries -> ...
Standard SAP graph patterns
PatternGraph topologyUse
Pure hierarchical (fixed sequence)H1 -> H2 -> H3 with weight 1 on each transitionStrict ordering
Holm graph (equal weights)Each Hi -> Hj with weight 1/(m-1)No priority ordering
Primary + secondariesPrimary -> Sec1 (0.5), Sec2 (0.5); Sec1 ↔ Sec2 (1)Pivotal labeling claims
Co-primary chainH1 -> H2 with full weight if BOTH H1a, H1b rejectCo-primary + secondary
Subgroup branchPrimary -> Subgroup_OS (0.2), Sec1 (0.4), Sec2 (0.4)Discovery subgroup with budget
Bretz et al 2011 -- Simes and parametric extensions

When endpoints are positively correlated, replace the Bonferroni-based intersection test with Simes (for positive dependence) or parametric (using known correlation):

r
library(gMCP)
# Use Simes-based local tests at each intersection
result_simes <- gMCP(graph, pvalues = p_vals, alpha = 0.025, test = 'Simes')
# Or parametric with estimated correlation matrix
result_param <- gMCP(graph, pvalues = p_vals, alpha = 0.025, corr = correlation_matrix)
Maurer-Bretz 2013 entangled graphs

Entangled graphs add memory: the alpha propagation can depend on the origin of the alpha. This allows parent-descendant constraints that a single non-entangled graph cannot express. Example: secondary endpoint Sec1 receives alpha only from Primary, never from Sec2.

Postdoc argument: purists argue memory makes the procedure non-coherent in Gabriel's sense; Glimm/Maurer/Bretz argue it matches real-world inferential intent.

Gabriel coherence in plain terms: a coherent procedure rejects a hypothesis H consistently regardless of which superset of H is being tested. Non-entangled graphs are coherent: if H1 is rejected via path A, it would also be rejected via path B. Entangled (memory-bearing) graphs sacrifice coherence: the same H may be rejected when alpha arrives from one parent but not from another, because the propagation history changes the available alpha. The trade-off is operational power -- entangled graphs can encode "secondary X is meaningful only if primary Y rejects, not if primary Z rejects" inferential intent that flat coherent procedures cannot express. Choose based on whether the SAP needs path-dependent priority.

Gatekeeping Procedures

Serial gatekeeping (hierarchical)

Test H1 at full alpha; only if it rejects, test H2 at full alpha; etc. Maximises power for H1 but H_k becomes inferentially worthless once any H_j (j<k) fails.

r
# Hierarchical / serial: just a chain graph in gMCP
hyp <- c('H1', 'H2', 'H3', 'H4')
weights <- c(1, 0, 0, 0)
trans <- matrix(c(0, 1, 0, 0, 0, 0, 1, 0, 0, 0, 0, 1, 0, 0, 0, 0),
                 nrow=4, dimnames=list(hyp, hyp))
graph <- graphMCP(m = trans, weights = weights, hnames = hyp)

Pre-specification of order is critical — based on clinical importance, NOT expected effect size. Ordering by expected effect is data-driven and inflates Type-I.

Parallel gatekeeping (Dmitrienko 2003)

Secondary family is tested only if at least one primary rejects. Bonferroni-based parallel gatekeeper has stepwise representation (Guilbaud 2007 Biom J 49:917).

Mixed / multistage (Dmitrienko-Tamhane 2008)

Permits using any closed-testing local procedure (e.g., Holm in family 1, Hommel in family 2) and combining via closure principle. R gMCP::generalMixGatekeeping or Mediana/MultXpert packages.

Postdoc tradeoff: parallel gatekeeping power loss vs collapsing endpoints into a composite (which avoids multiplicity but dilutes effect if components move in opposite directions); whether tree gatekeeping (Dmitrienko et al 2008 Stat Med 27:3446) over-engineers vs equivalent graphical procedure.

Hochberg vs Hommel vs Holm

Holm 1979 — step-down rejective Bonferroni; FWER controlled under any joint dependence. Conservative but robust.

Hochberg 1988 — step-up using ordered Simes critical values; needs Simes inequality which requires PRDS (Sarkar 1998, 2008). Under PRDS, Hochberg uniformly dominates Holm.

Hommel 1988 — also Simes-based but uses closed-testing tableau directly (not step-up shortcut). Uniformly more powerful than Hochberg (typically 1-3% gain).

When Hochberg fails (anti-conservative)

Hochberg becomes Type-I-inflated under negative dependence — relevant when comparing endpoints mathematically constrained to move in opposite directions (LDL-C and HDL-C; complementary efficacy and safety endpoints).

Sarkar critique: when PRDS cannot be proven, fall back to Holm. The lost power is the price of robustness.

python
# Python: statsmodels supports Holm, Hochberg, Hommel
from statsmodels.stats.multitest import multipletests

p_vals = [0.018, 0.042, 0.038, 0.015]
for method in ['holm', 'hochberg', 'hommel', 'bonferroni']:
    reject, adj_p, _, _ = multipletests(p_vals, alpha=0.05, method=method)
    print(f'{method}: reject={reject}, adjusted={adj_p}')

FDA Multiple Endpoints Final Guidance (October 2022)

Federal Register 2022-22882 finalises 2017 draft. Key changes vs draft:

  • Explicit recognition of newer methods including win-ratio (Pocock 2012 Eur Heart J) and weighted composites
  • Clearer language that "key secondary" endpoints are those for which sponsor wishes to make label claims and which must be in a Type-I-error-controlled hierarchy
  • Appendix with worked graphical-procedure examples
Categories
CategoryApproachNote
CompositeSingle test; no multiplicityWin-ratio, DOOR/RADAR, time-to-first-event
Co-primary (all-win)Each at full alpha; n inflated for joint powerPower = product of marginals
Multiple primary (any-wins)Alpha must be split (Bonferroni or graphical)More n required than co-primary if effects similar
Primary + key secondaryHierarchical or graphicalModern preference: graphical for flexibility

Winner's bias warning: when post-hoc-selected endpoints are emphasised, bias-corrected effect estimates are recommended (same selection-bias issue as adaptive design).

The "Almighty Primary Endpoint" Critique

Dmitrienko-D'Agostino 2017 Stat Med 36:4423 editorial surveys progress in trial-multiplicity methodology. A recurring theme motivating that work: insisting on a single primary endpoint can lose power when a therapy has broad multi-domain benefit (heart failure drugs with effects on mortality, hospitalisation, symptoms, biomarkers) -- motivating composite endpoints, the win ratio, or multiple primary endpoints with explicit alpha allocation.

Win-ratio (Pocock-Ariti-Collier-Wang 2012) and hierarchical composite (DOOR/RADAR, Evans 2015) are responses — they preserve a single inferential test while letting multiple endpoints contribute.

FDA counter-position (Hung, O'Neill, Wang): without a designated primary, sponsors and regulators negotiate over secondary endpoints post hoc, destroying inferential meaning. Hence the FDA 2022 guidance reaffirms key-secondary hierarchies.

Show full SKILL.md (1,029 more words)Show less

Per-Method Failure Modes

Hochberg under negative dependence
  • Trigger: Endpoints constrained to move in opposite directions (LDL vs HDL; efficacy vs harm).
  • Mechanism: Hochberg's PRDS assumption fails; Simes inequality doesn't hold; Type-I inflated.
  • Symptom: Replication with Holm finds non-significant where Hochberg rejected.
  • Fix: Switch to Holm (no PRDS assumption); cite Sarkar 1998.
Fixed-sequence with wrong ordering
  • Trigger: Ordering by expected effect size rather than clinical priority.
  • Mechanism: Data-driven ordering inflates Type-I.
  • Symptom: Reviewer asks for pre-specified ordering rationale.
  • Fix: Pre-specify order by clinical priority in SAP; document rationale.
Graphical procedure without pre-specified weights
  • Trigger: Weights chosen at analysis time to favour observed results.
  • Mechanism: Equivalent to post-hoc multiplicity tuning; inflates Type-I.
  • Symptom: Multiple "what if" graph variants in CSR.
  • Fix: Pre-specify graph and weights in SAP; document at protocol design.
Bonferroni when graph would gain power
  • Trigger: Default conservative choice when no thought put into structure.
  • Mechanism: Loses 30-50% power vs Hommel/graphical when m ~ 10 correlated tests.
  • Symptom: Underpowered trial reaches non-significance where graphical procedure would.
  • Fix: Design a proper graph in gMCP; cite Bretz-Maurer 2009.
FDR used for confirmatory primary
  • Trigger: SAP specifies BH-FDR for primary multiplicity.
  • Mechanism: FDR controls expected proportion of false discoveries, not P(any false positive).
  • Symptom: Regulatory reviewer rejects as non-confirmatory.
  • Fix: FWER (graphical, Holm, Hochberg, Hommel) for confirmatory; FDR for exploratory only.
Subgroup analyses claimed without multiplicity
  • Trigger: Trial reports 10 subgroups with one significant at α=0.05.
  • Mechanism: ~40% probability of at least one false positive under global null.
  • Symptom: Cherry-picked subgroup claim in submission.
  • Fix: Pre-specified graphical alpha allocation OR explicit hypothesis-generating label; cite EMA 2019 subgroup guideline.
Win-ratio reported without hierarchical priority
  • Trigger: Win-ratio composite with unspecified component priority.
  • Mechanism: Component prioritisation drives the result; arbitrary choice = data-dependent answer.
  • Symptom: Two analysts get different results from same data depending on hierarchy.
  • Fix: Pre-specify hierarchy in SAP; sensitivity over alternative hierarchies.

Quantitative Thresholds

ThresholdSourceRationale
FWER for confirmatory; FDR for exploratoryICH E9; FDA 2022 Multiple EndpointsRegulatory standard universally
Bonferroni: ~10 tests -> 30-50% power lossSarkar 1998 PRDSConservative under positive dependence
PRDS required for Hochberg validitySarkar 2008 Ann StatOtherwise Type-I inflated; fall back to Holm
Subgroup α budget <=20% of total (convention)Dane 2019 EFSPI white paper (subgroup discipline)Discipline against subgroup fishing
Key secondary requires hierarchy in SAPFDA 2022 FinalLabeling claims need Type-I-controlled test
Composite avoids multiplicity but dilutes effectPocock 2012 Eur Heart JWin-ratio captures heterogeneity in single test

Common Errors

Error / symptomCauseSolution
Hochberg applied to negatively-dependent endpointsPRDS not checkedSwitch to Holm (cite Sarkar 1998)
Fixed-sequence ordering data-drivenPost-hoc selectionPre-specify clinical priority in SAP
Bonferroni at 10 correlated endpointsDefault conservatismGraphical procedure (gMCP); 30-50% power gain
FDR for confirmatory primaryMisunderstanding error ratesFWER mandatory for confirmatory; FDR exploratory only
Graphical procedure run with multiple weight schemesPost-hoc graph tuningPre-specify single graph in SAP
Subgroups significant without multiplicityCherry-pickingPre-specified allocation OR explicit hypothesis-generating label
Co-primary treated as multiple primaryConfused alpha allocationCo-primary: no alpha split; inflate n. Multiple primary: split alpha
Win-ratio component priority unspecifiedData-driven choicePre-specify hierarchy with rationale; sensitivity over alternatives
multipletests default method='hs' (Holm-Sidak)Common Python mistakeAlways specify method='holm', 'hommel', etc., explicitly
Sensitivity analysis listed as a "key secondary" requiring alphaConfusion about roleSensitivity is "what if" not "another claim"; no alpha needed

Anticipated Reviewer Pushback

PushbackResponse
"Why this multiplicity procedure?"Closed testing per Marcus-Peritz-Gabriel; specific implementation is graphical (Bretz-Maurer 2009) with pre-specified weights in SAP
"Why Hommel not Holm?"PRDS holds (positive correlation among endpoints); Hommel dominates Holm by 1-3% with no Type-I cost
"Why graph weights X, Y, Z?"Clinical priority: primary > key secondary > exploratory; weights reflect labelling claim hierarchy
"Are these endpoints positively correlated?"Sensitivity analyses provided: Bonferroni, Holm, Hochberg, Hommel results all in CSR appendix; concordant
"Where is alpha for the subgroup analysis?"Pre-specified 20% of primary alpha allocated (a common convention); cite Dane 2019 for subgroup discipline
"Why not just composite endpoint?"Composite would dilute differential effect on mortality vs hospitalisation; key-secondary hierarchy preserves component-level claims
"PRDS check for Hochberg?"Endpoints positively correlated via simulation under null; PRDS holds; Hochberg/Hommel valid
"Sensitivity in the hierarchy?"No — sensitivity is "what if" and does not require alpha. Listed as supportive not key secondary.

References

  • Bretz F, Maurer W, Brannath W, Posch M. 2009. A graphical approach to sequentially rejective multiple test procedures. Stat Med 28:586-604.
  • Bretz F, Posch M, Glimm E, Klinglmueller F, Maurer W, Rohmeyer K. 2011. Graphical approaches for multiple comparison procedures using weighted Bonferroni, Simes, or parametric tests. Biom J 53:894-913.
  • Burman CF, Sonesson C, Guilbaud O. 2009. A recycling framework for the construction of Bonferroni-based multiple tests. Stat Med 28:739-761.
  • Dmitrienko A, Offen WW, Westfall PH. 2003. Gatekeeping strategies for clinical trials that do not require all primary effects to be significant. Stat Med 22:2387-2400.
  • Dmitrienko A, Tamhane AC, Wiens BL. 2008. General multistage gatekeeping procedures. Biom J 50:667-677.
  • FDA. 2022. Multiple Endpoints in Clinical Trials. Final Guidance.
  • Goeman JJ, Hemerik J, Solari A. 2021. Only closed testing procedures are admissible for controlling false discovery proportions. Ann Stat 49:1218-1238.
  • Guilbaud O. 2007. Bonferroni parallel gatekeeping -- transparent generalizations, adjusted p-values, and short proofs. Biom J 49:917-927.
  • Hochberg Y. 1988. A sharper Bonferroni procedure for multiple tests of significance. Biometrika 75:800-802.
  • Holm S. 1979. A simple sequentially rejective multiple test procedure. Scand J Stat 6:65-70.
  • Hommel G. 1988. A stagewise rejective multiple test procedure based on a modified Bonferroni test. Biometrika 75:383-386.
  • Marcus R, Peritz E, Gabriel KR. 1976. On closed testing procedures with special reference to ordered analysis of variance. Biometrika 63:655-660.
  • Maurer W, Bretz F. 2013. Memory and other properties of multiple test procedures generated by entangled graphs. Stat Med 32:1739-1753.
  • Pocock SJ, Ariti CA, Collier TJ, Wang D. 2012. The win ratio: a new approach to the analysis of composite endpoints in clinical trials. Eur Heart J 33:176-182.
  • Sarkar SK. 2008. Generalizing Simes' test and Hochberg's stepup procedure. Ann Stat 36:337-363.
  • clinical-biostatistics/trial-reporting - Multiplicity strategy reporting per CONSORT 2025
  • clinical-biostatistics/subgroup-analysis - Subgroup multiplicity allocation
  • clinical-biostatistics/power-and-sample-size - Power adjustment for co-primary endpoints
  • clinical-biostatistics/adaptive-designs - Combination tests for adaptive multiplicity
  • clinical-biostatistics/effect-measures - Reporting multiple effect measures post-multiplicity adjustment
  • experimental-design/multiple-testing - General methods (FDR, FWER, q-values)

© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files in clinical-biostatistics/multiplicity-graphical of GPTomics/bioSkills.

  • SKILL.md
  • examples/multiplicity_graphical.R
  • usage-guide.md

Open the folder on GitHubat commit d91ed3d

Used in 2 other repositories

We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 2 other GitHub owners. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Bio Clinical Biostatistics Multiplicity Graphical next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Bio Clinical Biostatistics Multiplicity Graphical compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Bio Clinical Biostatistics Multiplicity Graphical this skillGPTomics/bioSkills1.2k2 repos~6.2kAutomated safety check: PassMIT
Spec Peer Reviewpedronauck/skills634—~794Automated safety check: PassNone
Pride FetchClawBio/ClawBio1.2k—~4.2kAutomated safety check: PassMIT
Clinical Trials Databasegoogle-deepmind/science-skills3.2k2 repos~3.2kAutomated safety check: PassApache-2.0
CHARLS Paper Reproduction Guidexjtulyc/MedgeClaw6171 repos~1.8kAutomated safety check: PassNone
Biomedical Analysis Dispatchxjtulyc/MedgeClaw6171 repos~2kAutomated safety check: PassNone

Similar skills

  • Spec Peer Review

    pedronauck/skills

    Run one requested external review of an approved spec, design doc, RFC, or detailed PRD; produce findings for user-selected incorporation.

    634 GitHub stars~794 tokensUpdated 26 days ago
    Research & ScienceAuto-check passed
  • Pride Fetch

    ClawBio/ClawBio

    Query metadata and download data from the PRIDE Archive, EMBL-EBI's proteomics identifications database, via the PRIDE Archive REST API v3.

    1.2k GitHub stars~4.2k tokensUpdated yesterday
    Research & ScienceAuto-check passed
  • Clinical Trials Database

    google-deepmind/science-skills

    Query ClinicalTrials.gov via APIv2. An agent skill from google-deepmind/science-skills.

    3.2k GitHub starsUsed in 2 repos~3.2k tokens
    Research & ScienceAuto-check passed
  • Guides an agent through reproducing papers built on the CHARLS health and retirement survey, from variable mapping to cognition, depression and isolation scores.

    617 GitHub starsUsed in 1 repo~1.8k tokens
    Research & ScienceAuto-check passed
  • Routes bioinformatics, drug discovery, clinical and multi-omics tasks from a chat interface to Claude Code sessions running K-Dense scientific skills, with a live dashboard per task.

    617 GitHub starsUsed in 1 repo~2k tokens
    Research & ScienceAuto-check passed
  • Research Paper

    luwill/research-skills

    A skill your agent uses when the user asks to write or draft an ORIGINAL RESEARCH ARTICLE — IMRaD paper, conference paper, short/workshop paper, 研究论文/期刊论文/会议论文 — reporting their own completed…

    862 GitHub stars~1.9k tokensUpdated today
    Research & ScienceAuto-check passed

More from GPTomics/bioSkills

All 559 skills in this repo
  • Bio Alignment Io

    GPTomics/bioSkills

    Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.

    1.2k GitHub starsUsed in 3 repos~4.9k tokens
    Auto-check passed
  • bioSkills Installer

    GPTomics/bioSkills

    Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.

    1.2k GitHub starsUsed in 1 repo~789 tokens
    Auto-check passed
  • Bio Write Sequences

    GPTomics/bioSkills

    Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.

    1.2k GitHub starsUsed in 3 repos~2.1k tokens
    Auto-check passed
  • Amplicon Primer Clipping

    GPTomics/bioSkills

    Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.

    1.2k GitHub starsUsed in 2 repos~2.2k tokens
    Auto-check passed
  • Filters BAM alignments by FLAG bits, mapping quality and regions with samtools view or pysam, with recipes for common keep and drop cases.

    1.2k GitHub starsUsed in 2 repos~3.6k tokens
    Auto-check passed
  • Bio Alignment Indexing

    GPTomics/bioSkills

    Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.

    1.2k GitHub starsUsed in 2 repos~2.4k tokens
    Auto-check passed

Questions about Bio Clinical Biostatistics Multiplicity Graphical

What does Bio Clinical Biostatistics Multiplicity Graphical do?

Implements multiplicity control for confirmatory clinical trials using graphical procedures (Bretz-Maurer-Hommel), gatekeeping (parallel, serial, mixed), Hochberg/Hommel/Holm with PRDS, and the…. Bio Clinical Biostatistics Multiplicity Graphical is an agent skill from GPTomics/bioSkills. Implements multiplicity control for confirmatory clinical trials using graphical procedures (Bretz-Maurer-Hommel), gatekeeping (parallel, serial, mixed), Hochberg/Hommel/Holm with PRDS, and the closed-testing principle (Marcus-Peritz-Gabriel; Goeman 2021 admissibility).

When should I use Bio Clinical Biostatistics Multiplicity Graphical?

Bio Clinical Biostatistics Multiplicity Graphical fits situations like: designing the multiplicity strategy for confirmatory trials with multiple primary; key secondary endpoints.

How do I install Bio Clinical Biostatistics Multiplicity Graphical in Claude Code?

Run `npx skills add GPTomics/bioSkills --skill bio-clinical-biostatistics-multiplicity-graphical -a claude-code`. Or copy the skill folder (clinical-biostatistics/multiplicity-graphical in GPTomics/bioSkills) into .claude/skills/bio-clinical-biostatistics-multiplicity-graphical in your project. Claude Code loads it when a task matches its description.

How do I install Bio Clinical Biostatistics Multiplicity Graphical in Codex?

Run `npx skills add GPTomics/bioSkills --skill bio-clinical-biostatistics-multiplicity-graphical -a codex`. Or copy the skill folder (clinical-biostatistics/multiplicity-graphical in GPTomics/bioSkills) into .agents/skills/bio-clinical-biostatistics-multiplicity-graphical in your project. Codex loads it when a task matches its description.

Can I use Bio Clinical Biostatistics Multiplicity Graphical in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-clinical-biostatistics-multiplicity-graphical -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-clinical-biostatistics-multiplicity-graphical, .gemini/skills/bio-clinical-biostatistics-multiplicity-graphical, .github/skills/bio-clinical-biostatistics-multiplicity-graphical and .opencode/skills/bio-clinical-biostatistics-multiplicity-graphical in your project.

What does Bio Clinical Biostatistics Multiplicity Graphical need to run?

Going by SKILL.md and its folder, Bio Clinical Biostatistics Multiplicity Graphical needs R for the scripts in its folder and the command-line tools its instructions call (pip). Our summary lists: Python 3.

Does Bio Clinical Biostatistics Multiplicity Graphical access the network?

SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Bio Clinical Biostatistics Multiplicity Graphical safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Bio Clinical Biostatistics Multiplicity Graphical use?

Bio Clinical Biostatistics Multiplicity Graphical is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Bio Clinical Biostatistics Multiplicity Graphical use?

About 6.2k tokens (SKILL.md is roughly 25k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Bio Clinical Biostatistics Multiplicity Graphical?

Skills that share tags, products or a category with Bio Clinical Biostatistics Multiplicity Graphical: Spec Peer Review (pedronauck/skills, 634 stars), Pride Fetch (ClawBio/ClawBio, 1.2k stars), Clinical Trials Database (google-deepmind/science-skills, 3.2k stars) and CHARLS Paper Reproduction Guide (xjtulyc/MedgeClaw, 617 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Bio Clinical Biostatistics Multiplicity Graphical?

GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,218 GitHub stars. The repository holds 559 skills in this directory. The repository was last updated on August 15, 2026.

Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.