Then: "Which endpoint triggers each analysis?"
- Give examples based on number chosen.
Immediately after establishing triggers, ask: "Which endpoints are tested at which analysis?" An endpoint that triggers an analysis is always tested there, but a non-triggering endpoint may or may not be tested at every analysis. For example, if PFS triggers the IA and OS triggers the FA, PFS might be tested only at the IA (single look) or at both IA and FA.
Present a table for confirmation showing which endpoint × analysis combinations are active. Example:
"Is PFS also tested at the FA, or only at the IA?" This determines whether PFS is a single-look (k=1) or multi-look endpoint, which affects boundary computation.
(If multi-population) Expand the table to show all hypotheses (H1–H4) instead of just endpoints.
8b. (Separate schedules) Ask for each endpoint individually.
(Single endpoint — skip Q8/8a/8b) "How many times do you want to peek at the data before the final analysis?" (same options as 8a)
(Skip if no interims) "When should the interim peek(s) happen? This is how far along the study should be — measured by what fraction of the total expected events have occurred."
- Single-look triggering endpoint (k=1): If the endpoint that triggers an analysis is only tested once at that analysis (e.g., PFS tested only at the IA), there is NO information fraction to ask — the IA timing is fully determined by the number of events needed for that endpoint's test (driven by power, alpha, and HR). The non-triggering endpoint's IF at that analysis is calculated from the shared timeline. Skip Q9 entirely for this case.
- Multi-look triggering endpoint: Only ask the information fraction for the triggering endpoint at each analysis. The non-triggering endpoint's IF is calculated, not asked.
- If co-primary separate schedules: ask fractions separately per endpoint
- Common choices: 50%, 60%, 75% (1 IA); 33%/67%, 50%/75% (2 IAs)
- Key concept: Information fraction only applies to endpoints with multiple looks. For a single-look endpoint, 100% of its events are used at its only analysis — there is no fraction to choose.
(If multi-population) After computing the design, present the derived event table for confirmation:
"Do these event counts look reasonable, or would you like to adjust?"
(Skip if no interims) "How do you want to spend your alpha (false-positive budget) across the interim and final analyses?"
- A) Conservative early, save most for the final look — Lan-DeMets O'Brien-Fleming (sfLDOF)
- B) Moderately conservative — Hwang-Shih-DeCani gamma=-4 (sfHSD, gamma=-4)
- C) Moderate — Hwang-Shih-DeCani gamma=-2 (sfHSD, gamma=-2)
- D) Aggressive early, spread evenly — Lan-DeMets Pocock (sfLDPocock)
- E) Other
(If multi-population) "Should the same spending function be used for all hypotheses?"
- A) Yes — same for all
- B) No — specify per endpoint or per population
"How sure do you need to be that the trial will detect a real benefit if it exists? Higher power means a bigger study."
- A) 80% — standard, smaller study
- B) 85% — moderate confidence
- C) 90% — higher confidence, larger study
- D) Other — specify a percentage
(If co-primary endpoints) Whether to ask power for each endpoint depends on what triggers its last look:
- If an endpoint's last look is triggered by itself → its power is a free parameter — ask the user.
- If an endpoint's last look is triggered by another endpoint → its power is derived — calculate it, don't ask.
Example 1: IA1→PFS, IA2→OS, FA→OS. Only ask OS power.
Example 2: IA→PFS, FA→OS. Ask power for both.
(If multi-population) Ask power for the lead hypothesis that drives the study size (typically OS in the subpopulation for step-down, or OS in the broadest population for alpha-split). Gated hypotheses have their power calculated.
(Skip if no interims) "Should the study also be able to stop early for futility — i.e., if the drug clearly isn't working?"
- A) Yes, with a non-binding rule (advisory, not forced)
- B) Yes, with a binding rule (must stop if crossed)
- C) No — efficacy stopping only
(If co-primary or multi-population) Also ask which hypothesis/hypotheses should have futility stopping.
(If futility requested) "How should the futility boundary be determined?"
- A) Beta spending — boundary controls probability of stopping under H1.
- B) Under the null — boundary set for high probability of stopping when HR=1.
Event inflation — "non-binding" is asymmetric in gsDesign:
- Binding (test.type=3): Beta spending inflates events (both alpha and power assume binding).
- Non-binding (test.type=4): Also inflates events, but less than binding. "Non-binding" only applies to alpha (efficacy boundaries ignore futility). Power is still computed as if futility IS binding — trials that cross the futility bound under H1 are treated as lost, so more events are needed to maintain target power. The inflation is modest (~0.5% for gamma=-20, ~3% for gamma=-4, ~5% for gamma=-2). See
reference.md → "Beta spending futility and sample size" for the full explanation.
- Under the null: No event inflation.
(If beta spending) "Which spending function for futility?"
- A) HSD gamma=-2 — moderate (aggressive boundary)
- B) HSD gamma=-4 — conservative
- C) HSD gamma=-6 — very conservative
- D) HSD gamma=-20 — minimal (boundary rarely crossed under H1)
- E) Other
"How fast do you expect to enroll patients?"
- A simple steady rate (e.g., "15 patients/month")
- A ramp-up schedule (e.g., "5/month for 3 months, then 10/month, then 15/month")
13b. "What range of total sample size do you consider feasible for this trial? This helps us check whether the computed design falls within practical limits."
- Example: "600–900 patients" or "no more than 800"
- If the user provides a range, store it. After the design is computed, compare the resulting total N against this range. If N falls outside the range, flag it and discuss options (adjust alpha split, relax power target, change enrollment, etc.).
- Infeasibility check: If the minimum N required to meet the power target exceeds the user's upper bound, this is an infeasible constraint. Do NOT silently exceed the limit — explicitly flag it: "The minimum sample size for these design parameters is N_min = XXX, which exceeds your constraint of < YYY. Options: (a) relax the N constraint, (b) reduce power target, (c) increase HR assumption, (d) adjust alpha allocation." Present trade-offs and let the user decide before proceeding.
- N sensitivity exploration must center around the user's stated range. The user's range reflects what is feasible and desired — anchor exploration there. If the user says "less than 600", explore e.g., 500–600 in steps of 20. If the user says "600–900", explore that range. Do NOT start from the computed minimum N and work up — that wastes the user's time on infeasible or uninteresting values.
- N values must be consistent with the enrollment ramp. Do NOT use arbitrary round numbers (e.g., 520, 560, 600). Instead, compute the actual N that results from varying the last enrollment period duration. With a ramp of 5/mo×2mo + 20/mo×3mo + 30/mo×Kmo, the only free parameter is K (months of steady-state enrollment). Iterate over K values and report the resulting N = 5×2 + 20×3 + 30×K = 70 + 30K. For example: K=15→520, K=16→550, K=17→580, K=18→610. This ensures every N in the sensitivity table is achievable with the stated enrollment rates.
- N sensitivity must include a "what-if" row up to 5% above the constraint. When the design is near the N constraint boundary, always include at least one row in the sensitivity table that exceeds the constraint by up to 5% (e.g., if the constraint is N < 450, include the next achievable N up to ~472). Label it clearly as "exceeds constraint" and show the improvement in study duration, IA-FA gap, or power. This gives the user concrete data on the cost of relaxing the constraint by a small amount — often a modest N increase (e.g., 420 → 450) produces a disproportionate improvement in study timeline.
- Proactively recommend the above-constraint N when the operational benefit is significant. If the "what-if" row shows a disproportionate improvement (e.g., 6+ months shorter study duration, or IA-FA gap dropping from borderline to comfortable), explicitly recommend it: "N = 450 exceeds the constraint by 7%, but it shortens the study by 8 months and provides a comfortable 9-month IA-FA gap vs the tight 4-month gap at N = 420. Consider relaxing the N constraint." The sensitivity table makes the trade-off visible; your recommendation makes the choice actionable.
- If the user says "no constraint" or declines to specify, skip and proceed.
"What percentage of patients do you expect to drop out of the study per year?"
- A) ~2% per year
- B) ~5% per year
- C) ~10% per year
- D) Other
(If co-primary endpoints) Ask dropout separately for each endpoint.
(Skip if no interims) "What is the minimum follow-up you want before the first interim analysis? This is the time between when the last patient enrolls and when the first IA occurs."
- A) 3 months
- B) 6 months
- C) Other — specify
(Skip if no interims) "What is the minimum gap you want between any two consecutive analyses (including IA→FA)? This accounts for data cleaning, database lock, and review time."
- A) 6 months
- B) 9 months
- C) 12 months
- D) Other — specify
Q15 and Q16 are minimum constraints, not fixed inputs. The actual IA and FA timing is driven by event targets (information fractions or power requirements). The min follow-up and min gap only apply as lower bounds — if the event-driven timing already satisfies them, use the event-driven timing. Only push the analysis later if the event-driven timing would violate a constraint. This prevents over-powering the design by forcing analyses later than needed.
After all inputs are collected, summarize everything in a clean table and ask the user to confirm before running any computation.
Confirmation Table Template (Single Population)
Confirmation Table Template (Multi-Population)
Pattern Selection
Based on the collected inputs, identify the closest design pattern from examples.md. Use the table below as a starting point — not all designs will match a pattern exactly.
Single-look (k=1) endpoints: When a co-primary endpoint has only one analysis (e.g., PFS tested only at the IA), gsDesign(k=1) and gsSurv(k=1) will fail. Use compute_single_look_boundary() from examples.md and nSurv() for the baseline. See reference.md → "gsDesign(k=1) failure" for details.
Multi-population designs: Use Pattern 5 for any design with 2+ populations. Boundaries are computed via compute_gsd_boundaries() per hypothesis, and multiplicity is controlled via Maurer-Bretz graphical method. See reference.md → "Multi-Population Design" for the alpha recycling priority rules and step-down vs alpha-split strategies.
Why Schoenfeld (not gsSurv/nSurv) for subgroup events: gsSurv() and nSurv() treat the subgroup as an independent trial — they size enrollment as if the subgroup were the entire study. But in a multi-population design, the subgroup is a fraction of the total enrollment. Using gsSurv() for a 70%-prevalence subgroup effectively sizes the trial as if 100% of patients are in the subgroup, inflating the total N by up to 1/prevalence (e.g., ~43% for 70% prevalence). Instead, derive events analytically:
- Compute required events per hypothesis:
events = 4 × (z_α + z_β)² / log(HR)² (Schoenfeld)
- For multi-look hypotheses, apply the GSD inflation factor:
events_FA = events_schoenfeld × gsDesign(k, alpha, beta, sfu)$n.I[k] / gsDesign(k=1, alpha, beta)$n.I[1]
- Compute per-patient event probability using
compute_event_prob() (see examples.md), which integrates over the enrollment distribution
- Derive N from the bottleneck hypothesis:
N_sub = events_FA / event_prob, then N_total = N_sub / prevalence
- Take the max N across all hypotheses
This approach correctly accounts for the fact that the subgroup sees only prevalence × N patients, and avoids the inflation that gsSurv() introduces.
Step-down gated hypotheses: For hypotheses with initial alpha = 0 (gated behind other hypotheses), compute boundaries at full alpha (0.025), not at the alpha that might flow through the gate. The gated hypothesis has no initial allocation — its effective alpha depends entirely on the cascade. Full alpha is the appropriate basis for design properties (boundaries, events, power).
If no pattern matches exactly, compose the design from building blocks across multiple patterns. The patterns are not mutually exclusive — real designs often combine elements. For example:
- A co-primary design with non-proportional hazards → design under PH (Pattern 6 or 7), then apply NPH Evaluation Workflow
- A single endpoint with multi-population subgroups → combine Patterns 1 + 5
- A multi-population co-primary with single-look PFS → combine Patterns 5 + 7 (see "Pattern 5+7 Combo" below)
- A 3-endpoint design with mixed fixed-sequence and alpha splitting → adapt Pattern 5
When composing, read reference.md → "Analysis Framework" for the 5-perspective approach (hypotheses, multiplicity, IA plan, boundaries, power) to structure any custom design systematically.
Pattern 5+7 Combo: Multi-Population with Single-Look PFS
This is a common pattern for aggressive diseases (e.g., 2L SCLC) with short PFS median. Uses the N-first algorithm (see reference.md → "N-First Design Algorithm"):
Phase A — Determine starting N:
- Compute required events per hypothesis via Schoenfeld formula
- Estimate N_min using
estimate_min_N() (see examples.md)
- Pick starting N from user's feasibility range (Q13b), close to N_min
- Derive R from enrollment ramp (iterate K for last period)
Phase B — Design at fixed N:
- Find IA time when PFS-ES reaches its Schoenfeld event target (PFS-triggered IA)
- Compute OS events at IA → derive OS IF
- Design OS boundaries with
gsSurv(gamma_es, R_fixed, minfup=NULL, T=NULL) — only for boundaries, not enrollment
- Find FA time when OS-ES reaches its required events
- Recompute all cross-endpoint events at final IA and FA times
- Compute single-look PFS boundaries via
compute_single_look_boundary()
- Compute OS GSD boundaries via
compute_gsd_boundaries()
- Gated hypotheses: compute boundaries at full alpha (0.025)
Phase C — Evaluate and adjust N:
If power < target, timing too late, or OS IF too high → present N adjustment alongside other levers (alpha reallocation, relaxed power). Re-run Phase B.
N is a top-level design parameter. All results depend on N. Do NOT let gsSurv() determine enrollment — always fix N first, then derive everything from the fixed enrollment.
Key differences from standard Pattern 7: enrollment rates are scaled by prevalence for subgroup hypotheses, events are derived from prevalence (not nSurv()), and the multiplicity graph uses Maurer-Bretz with step-down or alpha-split gating.
Then read reference.md and the relevant sections of examples.md to proceed.
IA Timing Checks
After computing the design (step 6), read post_design.md → "IA Timing Checks" for the full checklist, warning messages, and user options.
Quick summary — all must pass before proceeding:
NPH Evaluation Workflow
When the user specifies non-proportional hazards (Q7b = B), follow the "Design under PH, evaluate under NPH" approach. Do NOT size the trial directly under NPH — it breaks the IA plan.
Read reference.md → "NPH Evaluation: Design Under PH, Evaluate Under NPH" for the full rationale, step-by-step algorithm, tool selection rules, and gotchas.
Read examples.md → "NPH Evaluation with lrstat" for the complete R code pattern.
Summary of the workflow:
- Complete the PH design (steps 1–7) — this gives events, boundaries, enrollment
expected_time() → AHR and timing at each PH event target under NPH
gs_power_npe() → analytical power under NPH
lrsim() → verify analyticals (timing ±1 mo, power ±2 pp, type I error ±0.5 pp)
- Present comparison table, assessment, and options to the user
Co-Primary Shared-Timing Design Workflow
When co-primary endpoints share analysis calendar times but different endpoints trigger different analyses, the design requires an iterative workflow because the endpoints' event counts are interdependent through the shared timeline. You cannot design both endpoints independently.
Read reference.md → "Co-Primary Shared-Timing: Iterative Design Workflow" for the step-by-step algorithm.
Read examples.md → "Co-Primary Shared-Timing Iterative Workflow" for the complete R code pattern.
Key principle: design the FA-triggering endpoint first (it drives the study size), then derive everything else from the resulting timeline.
Verification
Every new design MUST be verified by simulation before delivery. Read post_design.md → "Verification" for the full procedure: what to verify, pass criteria, how to run lrsim(), and the verification log template.
Quick pass criteria:
- Power (H1): within ±2 pp of calculated
- Type I error (H0): within ±0.5 pp of alpha
- Events: within ±5% of calculated
- Timing: within ±1 month of calculated
Non-binding futility: use futilityBounds = rep(-6, k-1) in BOTH H0 and H1 lrsim() calls. See examples.md → "Verification with lrsim()" for the code.