Agent skill

Bio Clinical Biostatistics Logistic Regression

by GPTomics in GPTomics/bioSkills

Performs logistic regression for clinical trial outcomes (binary, ordinal, multinomial) with marginal-vs-conditional estimand reporting per FDA 2023 covariate adjustment guidance…

MITAuto-check passedResearch & Science

Install Bio Clinical Biostatistics Logistic Regression

skills CLI
$ npx skills add GPTomics/bioSkills --skill bio-clinical-biostatistics-logistic-regression -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install GPTomics/bioSkills bio-clinical-biostatistics-logistic-regression --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/clinical-biostatistics/logistic-regression .claude/skills/bio-clinical-biostatistics-logistic-regression && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
bio-clinical-biostatistics-logistic-regression
GitHub stars
1.2k
Used in
2 other repos
Token cost
~7.3k tokens
SKILL.md length
3,043 words
Files
3
Skills in repo
559
Repo updated
First seen
Licence
MIT

At a glance

Performs logistic regression for clinical trial outcomes (binary, ordinal, multinomial) with marginal-vs-conditional estimand reporting per FDA 2023 covariate adjustment guidance…

  • Modeling binary
  • SKILL.md covers Version Compatibility, Algorithmic Taxonomy, Decision Tree by Scenario and Standard Workflow, plus 13 more sections
  • Runs Python scripts from its folder; calls pip
  • Ordinal endpoints in confirmatory

What it does

Bio Clinical Biostatistics Logistic Regression is an agent skill from GPTomics/bioSkills. Performs logistic regression for clinical trial outcomes (binary, ordinal, multinomial) with marginal-vs-conditional estimand reporting per FDA 2023 covariate adjustment guidance, g-computation/standardisation for marginal effects, modified Poisson for RR, Brant test for proportional odds, Firth penalty for separation, and Hauck-Donner detection. Use when modeling binary or ordinal endpoints in confirmatory or exploratory clinical trials.

Its SKILL.md is about 7.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files (for example `examples/logistic_regression_clinical.py` and `usage-guide.md`).

It sits in Research & Science, covering Clinical and healthcare research. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.

When your agent uses it

  • Modeling binary
  • Ordinal endpoints in confirmatory
  • Exploratory clinical trials

Example prompts

  • “Use the bio-clinical-biostatistics-logistic-regression skill to perform logistic regression for clinical trial outcomes (binary, ordinal…”
  • “/bio-clinical-biostatistics-logistic-regression”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Bio Clinical Biostatistics Logistic Regression loads about 7.3k tokens when it runs. Until then it costs about 122 tokens; SKILL.md has 3,043 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~122
When it runs · the whole SKILL.md, loaded when a task matches
~7.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 3,043 words, ~7,308 tokens.

Download SKILL.mdSave it as .claude/skills/bio-clinical-biostatistics-logistic-regression/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
bio-clinical-biostatistics-logistic-regression
description
Performs logistic regression for clinical trial outcomes (binary, ordinal, multinomial) with marginal-vs-conditional estimand reporting per FDA 2023 covariate adjustment guidance, g-computation/standardisation for marginal effects, modified Poisson for RR, Brant test for proportional odds, Firth penalty for separation, and Hauck-Donner detection. Use when modeling binary or ordinal endpoints in confirmatory or exploratory clinical trials.
tool_type
python
primary_tool
statsmodels

Version Compatibility

Reference examples tested with: statsmodels 0.14+, scipy 1.12+, numpy 1.26+, pandas 2.1+, firthmodels 0.3+, marginaleffects 0.0.13+ (Python) / 0.20+ (R). R packages cited: RobinCar, marginaleffects, brant, MASS, VGAM.

Before using code patterns, verify installed versions match. If versions differ:

  • Python: pip show <package> then help(module.function) to check signatures
  • R: packageVersion('<pkg>') then ?function_name

If code throws ImportError, AttributeError, or TypeError, introspect the installed package and adapt the example to match the actual API rather than retrying.

Logistic Regression for Clinical Outcomes

"Model clinical outcomes with logistic regression" -> Estimate the marginal or conditional treatment effect on a binary or ordinal endpoint using a model that respects randomisation stratification, declares its estimand, and survives covariate misspecification.

Conditional vs marginal (the non-collapsibility subtlety): the OR is non-collapsible. The conditional OR from logistic regression is a different parameter than the marginal OR, even when there is NO confounding and randomisation is perfect. This is mathematical, not statistical bias. FDA 2023 favours marginal RD (via g-computation) for primary reporting to avoid parameter ambiguity. See clinical-biostatistics/effect-measures and Permutt 2020.

Algorithmic Taxonomy

ApproachEstimandInferenceStrengthFails when
Unadjusted logistic / chi-squareMarginal ORWald or LRSimple; transparentLoses efficiency vs adjusted (Senn 2013); inflates SE under stratified randomisation (Kahan-Morris 2012)
Logistic with covariates (ML, Wald CI)Conditional log-ORWaldStandard; widely availableConditional OR != marginal OR due to non-collapsibility (Permutt 2020); not the FDA 2023 primary estimand
Logistic + g-computation / standardisationMarginal RD/RR/ORInfluence function or bootstrap SEFDA 2023 recommended primary estimand for binaryRequires correct outcome model AND post-fit standardisation; needs robust SE machinery
Targeted Maximum Likelihood (TMLE)Marginal RD/RR/ORInfluence functionProvably efficient; doubly robust in observationalImplementation heavier; mostly R (tmle, tmle3); rare in confirmatory submissions
Modified Poisson with sandwich SEMarginal RRHC1/HC3 sandwichDirect RR estimation when prevalence >10%Slightly less efficient than log-binomial when log-binomial converges
Log-binomial regressionMarginal RRWaldDirect RR estimationFrequent convergence failure when predicted risk near 1
Firth penalised logisticConditional OR (penalised)Penalised LR test preferredHandles separation, rare events (<5% prevalence)Wald CI/p liberal; must use PLR test (Heinze-Schemper 2002)
Ordinal logistic (proportional odds)Common conditional OR across cut-pointsWald or LRPreserves ordering informationProportional odds assumption violation (Brant test)
Partial proportional oddsPO holds for some covariates, not othersHybridSalvages ordinal model when PO fails for one predictorIncreased complexity; harder interpretation
Multinomial logisticPer-category log-ORWaldNo PO assumption neededLoses efficiency; harder communication

Postdoc reading list: Permutt 2020 Stat Biopharm Res 12:45 (conditional vs marginal estimand); FDA May 2023 Final Guidance "Adjusting for Covariates in RCTs" (the regulatory rulebook); Tsiatis et al 2008 Stat Med 27:4658 (robust ANCOVA framework); Senn 2013 Stat Med 32:1439 (precision from adjustment even under perfect balance); Kahan-Morris 2012 Stat Med 31:328 (must adjust for stratification factors).

Decision Tree by Scenario

ScenarioRecommended approachWhy
RCT, binary primary endpoint, no covariate adjustment in SAPUnadjusted logistic with marginal OR; chi-square supportiveSimple regulatory case; consider whether covariate adjustment would gain power (Senn 2013)
RCT, binary primary, covariates pre-specified, FDA 2023 compliantLogistic adjusted + g-computation for marginal RD; conditional OR supportiveFDA 2023 final: marginal estimand for primary; mention HC3 sandwich SE
RCT, stratified randomisation (sex/region/severity)Include strata as covariates in logistic; the analytic decision is non-optionalKahan-Morris 2012: ignoring strata is over-conservative -- SE biased upward, CIs too wide, power loss
Common outcome (prevalence >10%) where RR is the policy quantityModified Poisson + HC1/HC3 OR log-binomial regressionAvoid OR misinterpretation; Zou 2004; cite EMA 2015 on covariate adjustment
Rare event (<5% prevalence) or separation observedFirth penalty + penalised LR test for p-valueHeinze-Schemper 2002; Wald p from Firth is liberal
Ordinal outcome, PO assumption supportableProportional odds; verify with Brant test (R brant::brant)Most efficient when PO holds; common in toxicity grading and PROs
Ordinal outcome, PO fails on one or two predictorsPartial PO model (R VGAM::vglm(..., cumulative(parallel=FALSE~X)))Salvages most efficiency; only the offending coefficient gets per-cut-point estimate
Ordinal outcome, PO fails widelyMultinomial logisticWorst case; loses ordering info but valid
Single arm observational with strong confoundingLogistic adjusted + g-computation + propensity weightingDoubly robust; cite Hernan-Robins; consider TMLE
Binary endpoint, longitudinal repeated measuresGEE or generalised linear mixed modelSandwich SE for GEE; mixed-model OR is subject-specific not marginal

Standard Workflow

Goal: Fit a covariate-adjusted logistic regression for a binary clinical endpoint with explicit reference category and pre-specified covariates.

Approach: Use the formula API for automatic categorical handling and explicit reference; report both the conditional log-OR and the marginal RD via g-computation.

python
import statsmodels.formula.api as smf
import pandas as pd
import numpy as np

# CRITICAL: set reference category explicitly. statsmodels default is alphabetical,
# so 'Active' sorts before 'Placebo' and the OR direction silently flips.
model = smf.logit(
    'outcome ~ C(ARM, Treatment(reference="Placebo")) + age + C(sex) + baseline_score',
    data=df
).fit()

# Conditional ORs (Wald)
or_table = pd.DataFrame({
    'OR': np.exp(model.params),
    'Lower_CI': np.exp(model.conf_int()[0]),
    'Upper_CI': np.exp(model.conf_int()[1]),
    'p_value': model.pvalues
})

# Marginal RD via g-computation (the FDA 2023-recommended primary estimand):
df_active = df.assign(ARM='Active')
df_placebo = df.assign(ARM='Placebo')
risk_active = model.predict(df_active).mean()
risk_placebo = model.predict(df_placebo).mean()
marginal_rd = risk_active - risk_placebo
# For SE, use influence-function or bootstrap; see RobinCar / marginaleffects in R

The reference-category trap is the single most common silent bug. smf.logit('y ~ C(ARM)') uses alphabetical ordering, so ARM='Active' becomes 1 and ARM='Placebo' becomes 0 -- but with ARM=['Active','Placebo','Control'], 'Active' is the reference and 'Placebo' is the comparison. Always pass Treatment(reference="Placebo") or equivalent. R glm(y ~ relevel(ARM, ref='Placebo')).

Marginal vs Conditional Estimand -- The FDA 2023 Pivot

The single most important methodological shift in clinical biostatistics 2020-2025. Permutt 2020 established and FDA 2023 codified: the maximum-likelihood coefficient on Z in glm(Y ~ Z + X, family=binomial) is the conditional log-OR -- the OR comparing Z=1 to Z=0 holding X fixed. Due to OR non-collapsibility, this is a different parameter than the marginal log-OR (the OR comparing the entire treated population to the entire control population), even under perfect randomisation.

The FDA 2023 final guidance ("Adjusting for Covariates in Randomized Clinical Trials," May 2023) endorses covariate-adjusted nonlinear-model analysis provided the analyst targets a marginal estimand. Reporting the conditional OR from a multivariable logistic regression as "the treatment effect" silently changes the estimand from what was pre-specified.

G-computation / Standardisation
python
# Marginal effects on three scales via g-computation:
df_z1 = df.assign(ARM='Active')
df_z0 = df.assign(ARM='Placebo')
p_z1 = model.predict(df_z1)
p_z0 = model.predict(df_z0)

marg_p1 = p_z1.mean()
marg_p0 = p_z0.mean()
marg_rd = marg_p1 - marg_p0
marg_rr = marg_p1 / marg_p0
marg_or = (marg_p1 / (1 - marg_p1)) / (marg_p0 / (1 - marg_p0))

For valid SE: the delta-method/influence-function variance and HC3 sandwich are required. Python marginaleffects v0.0.13+ implements this via marginaleffects.avg_comparisons(model, variables='ARM', vcov='HC3'). R is more mature: marginaleffects::avg_comparisons (Arel-Bundock-Greifer-Heiss 2024 JSS 111:9), RobinCar (purpose-built for FDA 2023), riskCommunicator.

Tsiatis et al 2008 robustness guarantee: under randomisation Z ⊥ X, the g-computation marginal estimator is consistent for the marginal ATE even if the outcome model is misspecified. Efficiency depends on model quality; consistency does not. This is why FDA 2023 accepts the marginal RD via g-computation without requiring proof of correct logistic mean structure.

Reporting template (post-FDA-2023):

"Primary estimand: marginal risk difference of -8.5 percentage points (95% CI -12.7 to -4.2, HC3 SE), computed by g-computation/standardisation from a logistic regression adjusted for age, sex, and baseline severity. Supportive: conditional OR 0.48 (95% CI 0.34-0.67, Wald) from the same model."

Modified Poisson for Common Outcomes -- Direct RR

When prevalence > 10% and the policy quantity is RR (not OR), modified Poisson with sandwich SE is the de facto modern standard (Zou 2004 AJE 159:702):

python
import statsmodels.api as sm
import numpy as np

X = sm.add_constant(df[['treatment', 'age', 'baseline_score']])
poisson_model = sm.GLM(df['outcome'], X, family=sm.families.Poisson()).fit(cov_type='HC1')
rr = np.exp(poisson_model.params)
rr_ci = np.exp(poisson_model.conf_int())

cov_type='HC1' corrects the over-dispersion that Poisson inherently assumes. Without it the variance is wrong because binary data are NOT Poisson. For n < 250, HC3 is preferred (Long-Ervin 2000). What postdocs argue about: modified Poisson vs log-binomial -- log-binomial directly models RR but frequently fails to converge when predicted risk approaches 1; modified Poisson always converges but is marginally less efficient when log-binomial works.

Proportional Odds and Brant Test

Proportional odds (PO) assumption: the effect of each predictor is constant across all cut-points of the ordinal outcome. Must be tested or the model parameters are invalid.

python
from statsmodels.miscmodels.ordinal_model import OrderedModel
from statsmodels.api import MNLogit
import pandas as pd

df['severity'] = pd.Categorical(df['severity'], categories=['mild', 'moderate', 'severe'], ordered=True)

po_model = OrderedModel.from_formula('severity ~ treatment + age', data=df, distr='logit').fit(method='bfgs', disp=0)
mn_model = MNLogit.from_formula('severity ~ treatment + age', data=df).fit(disp=0)

# LR test for PO vs MN (omnibus)
lr_stat = 2 * (mn_model.llf - po_model.llf)
lr_df = mn_model.df_model - po_model.df_model
from scipy.stats import chi2
lr_p = 1 - chi2.cdf(lr_stat, lr_df)

# Per-coefficient Brant test in R: brant::brant(po_model_object)
# The Brant test localises which predictor(s) violate PO -- essential before deciding on partial PO

The omnibus LR test is necessary but not sufficient. Brant test (Brant 1990 Biometrics 46:1171; R brant::brant) provides per-coefficient PO tests, identifying which predictor violates PO. If only one or two predictors violate PO, fit a partial proportional odds model (R VGAM::vglm(..., cumulative(parallel=FALSE~X1+X2))) rather than abandoning the model entirely.

OrderedModel intercept gotcha: do NOT add an intercept term. Threshold parameters (cut-points between ordinal levels) replace the intercept. An explicit constant causes non-identifiability and optimiser failure.

Separation and Firth Penalty

Detection:

python
from firthmodels import FirthLogisticRegression, detect_separation
import numpy as np

sep_result = detect_separation(X, y)
if sep_result.separation:
    print(sep_result.summary())
# Manual signs: coefficient > 10, SE > 100, convergence warnings

Firth penalty (Firth 1993 Biometrika 80:27):

python
firth = FirthLogisticRegression()
firth.fit(X, y)
or_firth = np.exp(firth.coef_)

# IMPORTANT: Wald p-values from Firth are LIBERAL (anti-conservative).
# Prefer the penalised likelihood-ratio test (PLRT).
# Some Python `firthlogist` releases expose a PLRT attribute (e.g. `pvalues_lrt_`)
# but its presence and name vary by release -- check `dir(firth)` against the
# installed version, otherwise compute the PLRT manually from the penalised
# log-likelihoods of nested models.

Heinze-Schemper 2002 Stat Med 21:2409 showed Wald inference from Firth is liberal (anti-conservative); the penalised likelihood-ratio test (PLRT) is the recommended inference. PLRT attributes (e.g. pvalues_lrt_) appear in some Python Firth packages but the exact attribute name varies by release -- introspect the installed package; compute the PLRT manually if no attribute is exposed:

python
def penalised_lrt(firth_full, firth_reduced):
    # Compute 2 * (penalised log-lik full - penalised log-lik reduced) ~ chi-square_df
    pass  # implementation depends on package version; see Heinze-Schemper 2002

Firth's method was originally designed for finite-sample bias reduction, not separation per se -- it adds the Jeffreys prior penalty to the likelihood, keeping coefficients finite under separation as a side effect. Also recommended for rare events (<5% prevalence) where ML bias is non-negligible.

Hauck-Donner Effect Detection

The Hauck-Donner effect (1977 JASA 72:851; revived by Yee 2022 JASA 117:1763): the Wald test statistic is non-monotonic in the parameter estimate near the boundary. A large log-OR can produce a small Wald chi-square -- so the Wald test fails to reject when LR/profile-likelihood would. Common in small samples with strong predictors.

python
# In R, detect with VGAM::hdeff() and replace Wald with profile likelihood:
# MASS::confint.glm(model) returns profile-likelihood CIs as default
# Python equivalent: bootstrap or manual profile likelihood

When to suspect Hauck-Donner: large coefficient magnitude with non-significant Wald p; large SE with finite estimate; switching from Wald to LR test changes significance. The fix is always: switch to profile likelihood or LR inference.

Reconciliation: When Methods Disagree

PatternLikely causeAction
Conditional OR from logistic vs marginal RD from g-computation give different conclusionsNon-collapsibility (OR is non-collapsible; RD is collapsible)Report marginal RD as primary per FDA 2023; conditional OR as supportive with explicit parameter label; cite Permutt 2020
ML logistic diverges (large coefficient, huge SE); Firth penalised convergesComplete or quasi-complete separation in covariate-outcome relationshipFirth penalty with penalised LR test (NOT Wald p-values); cite Heinze-Schemper 2002
Modified Poisson and log-binomial give different RR estimatesLog-binomial convergence fragility when predicted risk approaches 1Modified Poisson with sandwich SE preferred for robust convergence; cite Zou 2004
Proportional-odds (cumulative logit) and multinomial give different inferencesPO assumption violated for one or more predictors (Brant test rejects)Localise via Brant test; if 1-2 predictors violate, partial-PO model in VGAM; if widely, multinomial
Adjusted vs unadjusted treatment effect differ substantiallyConfounding (in observational) OR non-collapsibility (in RCT)In RCT, non-collapsibility expected (Permutt 2020); in observational, investigate confounding via DAG
Large OR with non-significant Wald p-valueHauck-Donner effect: Wald non-monotone near boundary (Yee 2022)Use profile-likelihood CI (R MASS::confint.glm); detect via VGAM::hdeff()
Stratified analysis significant; unstratified notAchieved SE smaller in stratified analysisInclude stratification factors in primary model (Kahan-Morris 2012); ignoring biases the SE upward and loses power (over-conservative)
Significant treatment effect on conditional model vanishes when mediator includedAdjusting for post-treatment variable on causal pathwayNever adjust for mediators in primary; use causal DAG to distinguish confounder vs mediator

Per-Method Failure Modes

Reference-category silent reversal
  • Trigger: Treatment variable coded as character with Active and Placebo; reference not explicitly set.
  • Mechanism: statsmodels defaults to alphabetical ordering; 'Active' becomes the reference.
  • Symptom: OR for "ARM[T.Placebo]" appears in output; direction is the opposite of intended.
  • Fix: Always pass C(ARM, Treatment(reference="Placebo")) or numeric encoding with explicit comment.
Adjusting for a mediator
  • Trigger: Including a post-randomisation variable that lies on the causal pathway.
  • Mechanism: Adjustment blocks the causal effect, attenuating treatment effect estimate toward null.
  • Symptom: Significant unadjusted effect becomes non-significant after adjustment for "mechanism" variable (e.g., inflammation marker).
  • Fix: Use causal DAG to distinguish confounder from mediator; never adjust for post-treatment variables in the primary analysis. See ICH E9(R1).
Show full SKILL.md (1,233 more words)Show less
Hauck-Donner non-monotonicity
  • Trigger: Small samples, strong predictor, cell counts approaching zero.
  • Mechanism: Wald test statistic is non-monotone in the parameter near the boundary.
  • Symptom: Large OR magnitude with non-significant Wald p; switching to LR test changes significance.
  • Fix: Use profile likelihood (R MASS::confint.glm) or LR inference; detect with VGAM::hdeff().
Complete or quasi-complete separation
  • Trigger: A predictor perfectly predicts outcome in a subset of the data.
  • Mechanism: Likelihood is monotone in that coefficient; ML estimate diverges to infinity.
  • Symptom: Coefficient >10, SE >100, convergence warning; convergence message says "fitted probabilities numerically 0 or 1."
  • Fix: Firth penalty (firthmodels); inference via penalised LR test, not Wald.
Proportional odds violation undetected
  • Trigger: Ordinal outcome fit with cumulative-logit (PO) model without testing assumption.
  • Mechanism: PO assumes a constant effect across cut-points; violation invalidates model parameters.
  • Symptom: Brant test rejects PO for one or more predictors; per-cut-point effects in a saturated model diverge.
  • Fix: Brant test first; if PO fails on one predictor, partial PO model; if widely fails, multinomial logistic.
Stratified randomisation not in analysis
  • Trigger: Stratification by site/region/severity at randomisation; analysis ignores strata.
  • Mechanism: Achieved SE smaller than calculated SE because randomisation removed between-stratum variability.
  • Symptom: Over-conservative inference -- SE biased upward, CIs too wide, Type-I error below nominal, and power loss (Kahan-Morris 2012).
  • Fix: Include stratification factors as covariates in the logistic; this is non-optional per ICH E9, FDA 2023, EMA 2015.

Model Diagnostics

DiagnosticMethodThresholdCaveat
DiscriminationROC-AUC (sklearn.metrics.roc_auc_score)>0.7 acceptable, >0.8 goodTrial-level, not patient-individual; depends on outcome prevalence
Calibration -- primaryCalibration plot (observed vs predicted in deciles)Curve along the diagonalVisual; preferred over H-L
Calibration -- secondaryHosmer-Lemeshow chi-squarep > 0.05Low power n<200; oversensitive n>2000
Pseudo R-squaredmodel.prsquared (McFadden)>0.2 excellentNOT comparable to OLS R-squared
Events per variablen_events / n_covariates>=10 EPV (Peduzzi 1996)Below this: bias, overfitting; consider Firth
Hosmer-Lemeshow
python
from scipy.stats import chi2
import pandas as pd

def hosmer_lemeshow(y_true, y_pred, n_groups=10):
    df_hl = pd.DataFrame({'y': y_true, 'prob': y_pred})
    df_hl['group'] = pd.qcut(df_hl['prob'], n_groups, duplicates='drop')
    grouped = df_hl.groupby('group').agg(obs=('y', 'sum'), n=('y', 'count'), pred=('prob', 'mean'))
    grouped['expected'] = grouped['n'] * grouped['pred']
    hl_stat = (((grouped['obs'] - grouped['expected']) ** 2) /
               (grouped['n'] * grouped['pred'] * (1 - grouped['pred']))).sum()
    actual_groups = len(grouped)
    return hl_stat, 1 - chi2.cdf(hl_stat, actual_groups - 2)

H-L is supplementary, not primary: Hosmer-Lemeshow has low power n<200 (rarely rejects even for poor calibration) and oversensitive n>2000 (rejects for trivial miscalibration). The decile choice is also arbitrary -- pd.qcut(..., duplicates='drop') can reduce the actual number of groups when probabilities tie, changing df. Use calibration plots (smoothed observed vs predicted) as primary; Hosmer-Lemeshow as supplementary p-value.

Quantitative Thresholds

ThresholdSourceRationale
10 events per variable (EPV)Peduzzi et al 1996 J Clin Epidemiol 49:1373Below this, coefficient bias and CI mis-coverage
Prevalence > 10% -> prefer marginal RR via modified PoissonZou 2004 AJE 159:702OR overstates RR; modified Poisson directly estimates RR
Marginal estimand for primary regulatory analysisFDA 2023 Final GuidanceConditional OR is a different parameter than marginal (Permutt 2020)
Brant test for PO assumptionBrant 1990 Biometrics 46:1171Omnibus LR test misses per-coefficient violations
Firth penalty with penalised LR test, not WaldHeinze-Schemper 2002 Stat Med 21:2409Wald is liberal under Firth penalty
HC3 sandwich SE for n <=250Long-Ervin 2000 Am Stat 54:217HC1 (Stata default) anti-conservative in small samples
Stratification factors in analysis when used in randomisationKahan-Morris 2012 Stat Med 31:328Ignoring biases SE upward -- CIs too wide, power loss (over-conservative)

Common Errors

Error / symptomCauseSolution
OR direction opposite of expectedAlphabetical reference; 'Active' sorts before 'Placebo'C(ARM, Treatment(reference="Placebo")) always
Conditional and marginal OR differ substantiallyNon-collapsibility, not confoundingCite Permutt 2020; report both with explicit labels
Treatment effect vanishes after adjustmentAdjusted for a mediator (post-treatment variable)Causal DAG check; never adjust for post-treatment in primary
Huge coefficient, huge SE, non-significant WaldSeparation OR Hauck-DonnerDetect separation: firthmodels.detect_separation; switch to Firth + PLRT
H-L p > 0.5 with obvious miscalibrationLow power at n<200Use calibration plot as primary; H-L supplementary
H-L p < 0.001 with great-looking plotOversensitive at n>2000Use calibration plot; cite Steyerberg 2019 for cautions
OrderedModel fails to converge with interceptThreshold parameters replace interceptDrop the explicit constant
firth.pvalues_ looks too lowWald p from Firth is liberalInspect the installed Firth package for a PLRT attribute (e.g. pvalues_lrt_); compute PLRT manually if none is exposed
GLM(family=Binomial) and Logit give same point but different outputsm.Logit provides pred_table(), prsquared; sm.GLM does notUse Logit unless changing link functions
Robust SE not reported for modified PoissonMissing cov_type='HC1' or 'HC3'Always specify sandwich SE for modified Poisson

Anticipated Reviewer Pushback

PushbackResponse
"Why is the marginal RD different from the conditional OR's implied RD?"Non-collapsibility of OR; cite Permutt 2020 and FDA 2023. Marginal RD via g-computation is the primary per FDA 2023.
"Adjustment for stratification factors?"Pre-specified strata included in the model. Cite Kahan-Morris 2012 and ICH E9.
"How were covariates chosen?"Pre-specified in the SAP based on prior literature/clinical knowledge. No data-driven selection (which would inflate Type-I).
"Why not use stepwise selection?"Stepwise leaks signal from the outcome and inflates Type-I. Pre-specification is the regulatory norm; cite FDA 2023.
"PH check for the longitudinal binary?"Binary doesn't have PH; if longitudinal, used GEE with sandwich SE or GLMM; see ICH E9 estimand for treatment policy.
"Why Firth not standard logistic?"Separation detected (or rare events <5%); standard ML diverges; Firth penalty + PLRT recommended (Heinze-Schemper 2002).
"Brant test result?"PO holds for predictors X1, X2, X3; PO fails for X4 -> partial PO model fit per VGAM.
"Calibration?"Calibration plot in supplement; H-L p reported as supplementary, not primary.

References

  • Arel-Bundock V, Greifer N, Heiss A. 2024. How to interpret statistical models using marginaleffects in R and Python. J Stat Softw 111:9.
  • Brant R. 1990. Assessing proportionality in the proportional odds model for ordinal logistic regression. Biometrics 46:1171-1178.
  • FDA. 2023. Adjusting for Covariates in Randomized Clinical Trials for Drugs and Biological Products. Final Guidance, May 2023.
  • Firth D. 1993. Bias reduction of maximum likelihood estimates. Biometrika 80:27-38.
  • Hauck WW, Donner A. 1977. Wald's test as applied to hypotheses in logit analysis. JASA 72:851-853.
  • Heinze G, Schemper M. 2002. A solution to the problem of separation in logistic regression. Stat Med 21:2409-2419.
  • Kahan BC, Morris TP. 2012. Improper analysis of trials randomised using stratified blocks or minimisation. Stat Med 31:328-340.
  • Lin DY, Wei LJ. 1989. The robust inference for the Cox proportional hazards model. JASA 84:1074-1078.
  • Long JS, Ervin LH. 2000. Using heteroscedasticity consistent standard errors in the linear regression model. Am Stat 54:217-224.
  • Moore KL, van der Laan MJ. 2009. Covariate adjustment in randomized trials with binary outcomes: TMLE. Stat Med 28:39-64.
  • Peduzzi P, Concato J, Kemper E, Holford TR, Feinstein AR. 1996. A simulation study of the number of events per variable in logistic regression analysis. J Clin Epidemiol 49:1373-1379.
  • Permutt T. 2020. Do covariates change the estimand? Stat Biopharm Res 12:45-53.
  • Senn S. 2013. Seven myths of randomisation in clinical trials. Stat Med 32:1439-1450.
  • Steyerberg EW. 2019. Clinical Prediction Models (2nd ed). Springer.
  • Tsiatis AA, Davidian M, Zhang M, Lu X. 2008. Covariate adjustment for two-sample treatment comparisons in randomized clinical trials. Stat Med 27:4658-4677.
  • Wang B, Susukida R, Mojtabai R, Amin-Esmaeili M, Rosenblum M. 2023. Model-robust inference for clinical trials that improve precision by stratified randomization and covariate adjustment. JASA 118:1152-1163.
  • Yee TW. 2022. On the Hauck-Donner effect in Wald tests. JASA 117:1763-1774.
  • Zou G. 2004. A modified Poisson regression approach to prospective studies with binary data. AJE 159:702-706.
  • clinical-biostatistics/cdisc-data-handling - Prepare analysis datasets from CDISC SDTM/ADaM domains
  • clinical-biostatistics/effect-measures - Modern CI methods; marginal vs conditional in depth
  • clinical-biostatistics/categorical-tests - Chi-square and Fisher alternatives for unadjusted tests
  • clinical-biostatistics/subgroup-analysis - Interaction terms and HTE detection
  • clinical-biostatistics/survival-analysis - Cox regression for time-to-event analogues
  • clinical-biostatistics/multiplicity-graphical - Bretz-Maurer graphs for multiple endpoints from one logistic model
  • clinical-biostatistics/trial-reporting - CONSORT 2025 and ICH E9(R1) reporting of logistic analyses

© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files in clinical-biostatistics/logistic-regression of GPTomics/bioSkills.

  • SKILL.md
  • examples/logistic_regression_clinical.py
  • usage-guide.md

Open the folder on GitHubat commit d91ed3d

Used in 2 other repositories

We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 2 other GitHub owners. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Bio Clinical Biostatistics Logistic Regression next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Bio Clinical Biostatistics Logistic Regression compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Bio Clinical Biostatistics Logistic Regression this skillGPTomics/bioSkills1.2k2 repos~7.3kAutomated safety check: PassMIT
CHARLS Paper Reproduction Guidexjtulyc/MedgeClaw6171 repos~1.8kAutomated safety check: PassNone
Histolab Whole Slide Image Tilingdavila7/claude-code-templates33k11 repos~5.1kAutomated safety check: PassMIT
NeuroKit2 Biosignal Processingdavila7/claude-code-templates33k11 repos~3kAutomated safety check: PassMIT
PyHealth Clinical ML Toolkitdavila7/claude-code-templates33k11 repos~4.4kAutomated safety check: PassMIT
PyhealthK-Dense-AI/scientific-agent-skills48k1 repos~2.1kAutomated safety check: PassMIT

Similar skills

  • Guides an agent through reproducing papers built on the CHARLS health and retirement survey, from variable mapping to cognition, depression and isolation scores.

    617 GitHub starsUsed in 1 repo~1.8k tokens
    Research & ScienceAuto-check passed
  • Histolab Whole Slide Image Tiling

    davila7/claude-code-templates

    Processes digital pathology whole slide images with histolab: tissue detection, mask creation, tile extraction and dataset preparation for deep learning.

    33k GitHub starsUsed in 11 repos~5.1k tokens
    Research & ScienceAuto-check passed
  • NeuroKit2 Biosignal Processing

    davila7/claude-code-templates

    Processes physiological signals with NeuroKit2 in Python: ECG, PPG, EEG, EDA, respiration, EMG and EOG, including HRV, events and complexity measures.

    33k GitHub starsUsed in 11 repos~3k tokens
    Research & ScienceAuto-check passed
  • PyHealth Clinical ML Toolkit

    davila7/claude-code-templates

    Builds machine learning pipelines on clinical data with PyHealth: EHR datasets, prediction tasks, medical code mapping, healthcare models and evaluation.

    33k GitHub starsUsed in 11 repos~4.4k tokens
    Research & ScienceAuto-check passed
  • Pyhealth

    K-Dense-AI/scientific-agent-skills

    Builds and validates PyHealth clinical machine-learning pipelines for EHR, signals, imaging, and medical codes.

    48k GitHub starsUsed in 1 repo~2.1k tokens
    Research & ScienceAuto-check passed
  • Rct Bias Assessment Rob

    aipoch/medical-research-skills

    Automates Risk of Bias 2 (ROB2) assessment for RCT papers by analyzing text against specific domains and synthesizing a report.

    1.9k GitHub stars~1.5k tokensUpdated 24 days ago
    Research & ScienceAuto-check passed

More from GPTomics/bioSkills

All 559 skills in this repo
  • Bio Alignment Io

    GPTomics/bioSkills

    Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.

    1.2k GitHub starsUsed in 3 repos~4.9k tokens
    Auto-check passed
  • bioSkills Installer

    GPTomics/bioSkills

    Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.

    1.2k GitHub starsUsed in 1 repo~789 tokens
    Auto-check passed
  • Bio Write Sequences

    GPTomics/bioSkills

    Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.

    1.2k GitHub starsUsed in 3 repos~2.1k tokens
    Auto-check passed
  • Amplicon Primer Clipping

    GPTomics/bioSkills

    Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.

    1.2k GitHub starsUsed in 2 repos~2.2k tokens
    Auto-check passed
  • Filters BAM alignments by FLAG bits, mapping quality and regions with samtools view or pysam, with recipes for common keep and drop cases.

    1.2k GitHub starsUsed in 2 repos~3.6k tokens
    Auto-check passed
  • Bio Alignment Indexing

    GPTomics/bioSkills

    Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.

    1.2k GitHub starsUsed in 2 repos~2.4k tokens
    Auto-check passed

Questions about Bio Clinical Biostatistics Logistic Regression

What does Bio Clinical Biostatistics Logistic Regression do?

Performs logistic regression for clinical trial outcomes (binary, ordinal, multinomial) with marginal-vs-conditional estimand reporting per FDA 2023 covariate adjustment guidance…. Bio Clinical Biostatistics Logistic Regression is an agent skill from GPTomics/bioSkills. Performs logistic regression for clinical trial outcomes (binary, ordinal, multinomial) with marginal-vs-conditional estimand reporting per FDA 2023 covariate adjustment guidance, g-computation/standardisation for marginal effects, modified Poisson for RR, Brant test for proportional odds, Firth penalty for separation, and Hauck-Donner detection.

When should I use Bio Clinical Biostatistics Logistic Regression?

Bio Clinical Biostatistics Logistic Regression fits situations like: modeling binary; ordinal endpoints in confirmatory; exploratory clinical trials.

How do I install Bio Clinical Biostatistics Logistic Regression in Claude Code?

Run `npx skills add GPTomics/bioSkills --skill bio-clinical-biostatistics-logistic-regression -a claude-code`. Or copy the skill folder (clinical-biostatistics/logistic-regression in GPTomics/bioSkills) into .claude/skills/bio-clinical-biostatistics-logistic-regression in your project. Claude Code loads it when a task matches its description.

How do I install Bio Clinical Biostatistics Logistic Regression in Codex?

Run `npx skills add GPTomics/bioSkills --skill bio-clinical-biostatistics-logistic-regression -a codex`. Or copy the skill folder (clinical-biostatistics/logistic-regression in GPTomics/bioSkills) into .agents/skills/bio-clinical-biostatistics-logistic-regression in your project. Codex loads it when a task matches its description.

Can I use Bio Clinical Biostatistics Logistic Regression in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-clinical-biostatistics-logistic-regression -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-clinical-biostatistics-logistic-regression, .gemini/skills/bio-clinical-biostatistics-logistic-regression, .github/skills/bio-clinical-biostatistics-logistic-regression and .opencode/skills/bio-clinical-biostatistics-logistic-regression in your project.

What does Bio Clinical Biostatistics Logistic Regression need to run?

Going by SKILL.md and its folder, Bio Clinical Biostatistics Logistic Regression needs Python for the scripts in its folder and the command-line tools its instructions call (pip). Our summary lists: Python 3.

Does Bio Clinical Biostatistics Logistic Regression access the network?

SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Bio Clinical Biostatistics Logistic Regression safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Bio Clinical Biostatistics Logistic Regression use?

Bio Clinical Biostatistics Logistic Regression is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Bio Clinical Biostatistics Logistic Regression use?

About 7.3k tokens (SKILL.md is roughly 29k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Bio Clinical Biostatistics Logistic Regression?

Skills that share tags, products or a category with Bio Clinical Biostatistics Logistic Regression: CHARLS Paper Reproduction Guide (xjtulyc/MedgeClaw, 617 stars), Histolab Whole Slide Image Tiling (davila7/claude-code-templates, 33k stars), NeuroKit2 Biosignal Processing (davila7/claude-code-templates, 33k stars) and PyHealth Clinical ML Toolkit (davila7/claude-code-templates, 33k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Bio Clinical Biostatistics Logistic Regression?

GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,218 GitHub stars. The repository holds 559 skills in this directory. The repository was last updated on August 15, 2026.

Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.