Agent skill

Bio Clinical Biostatistics Cdisc Data

by GPTomics in GPTomics/bioSkills

Reads, validates, and prepares CDISC SDTM and ADaM clinical trial data for analysis.

MITAuto-check passedResearch & Science

Install Bio Clinical Biostatistics Cdisc Data

skills CLI
$ npx skills add GPTomics/bioSkills --skill bio-clinical-biostatistics-cdisc-data -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install GPTomics/bioSkills bio-clinical-biostatistics-cdisc-data --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/clinical-biostatistics/cdisc-data-handling .claude/skills/bio-clinical-biostatistics-cdisc-data && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
bio-clinical-biostatistics-cdisc-data
GitHub stars
1.2k
Used in
2 other repos
Token cost
~7.3k tokens
SKILL.md length
2,962 words
Files
3
Skills in repo
559
Repo updated
First seen
Licence
MIT

At a glance

Reads, validates, and prepares CDISC SDTM and ADaM clinical trial data for analysis.

  • Works in 5 steps: Analysis-ready (one procedure call ->… → Traceability (every value links back to… → Clear/unambiguous communication via… → …
  • Working with clinical trial datasets in CDISC SDTM/ADaM format
  • SKILL.md covers Version Compatibility, Aggregation Strategy Taxonomy…, Decision Tree by Scenario and SDTM vs ADaM -- The Regulatory…, plus 17 more sections
  • Runs Python scripts from its folder; calls pip

What it does

Bio Clinical Biostatistics Cdisc Data is an agent skill from GPTomics/bioSkills. Reads, validates, and prepares CDISC SDTM and ADaM clinical trial data for analysis. Covers SDTM domain joins (DM, AE, EX, VS, LB, DS), ADaM architecture (ADSL, BDS, OCCDS, ADTTE) with traceability, treatment-emergent AE conventions, baseline derivation, SUPPQUAL/NSV handling, Define-XML 2.1, and Pinnacle 21 / CORE validation. Use when working with clinical trial datasets in CDISC SDTM/ADaM format, preparing analysis-ready data, or validating for regulatory submission.

Its SKILL.md is about 7.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files (for example `examples/cdisc_data_preparation.py` and `usage-guide.md`).

It sits in Research & Science, covering Clinical and healthcare research. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.

When your agent uses it

  • Working with clinical trial datasets in CDISC SDTM/ADaM format
  • Preparing analysis-ready data
  • Validating for regulatory submission

Example prompts

  • “/bio-clinical-biostatistics-cdisc-data”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Analysis-ready (one procedure call -> the analysis result)
  2. Traceability (every value links back to SDTM via metadata)
  3. Clear/unambiguous communication via Define-XML
  4. Naming conventions (PARAM, PARAMCD, AVAL, AVALC, BASE, CHG, PCHG, ABLFL, ANL01FL, ...)
  5. Structural rules (ADSL one row per subject; BDS one row per subject/parameter/timepoint/analysis flag)

What it can do on your machine

Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Bio Clinical Biostatistics Cdisc Data loads about 7.3k tokens when it runs. Until then it costs about 128 tokens; SKILL.md has 2,962 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~128
When it runs · the whole SKILL.md, loaded when a task matches
~7.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 2,962 words, ~7,320 tokens.

Download SKILL.mdSave it as .claude/skills/bio-clinical-biostatistics-cdisc-data/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
bio-clinical-biostatistics-cdisc-data
description
Reads, validates, and prepares CDISC SDTM and ADaM clinical trial data for analysis. Covers SDTM domain joins (DM, AE, EX, VS, LB, DS), ADaM architecture (ADSL, BDS, OCCDS, ADTTE) with traceability, treatment-emergent AE conventions, baseline derivation, SUPPQUAL/NSV handling, Define-XML 2.1, and Pinnacle 21 / CORE validation. Use when working with clinical trial datasets in CDISC SDTM/ADaM format, preparing analysis-ready data, or validating for regulatory submission.
tool_type
python
primary_tool
pyreadstat
goal_approach_exempt
true

Version Compatibility

Reference examples tested with: pyreadstat 1.2+, pandas 2.1+, numpy 1.26+. CDISC standards referenced: SDTM 2.0 / SDTMIG 3.4 (SDTM 3.0 / SDTMIG 4.0 in public review through April 2026); ADaMIG v1.3 (2021); OCCDS v1.1 (Nov 2021); BDS-for-TTE v1.0; Define-XML 2.1 (FDA-recommended for studies starting on/after March 15, 2023); Dataset-JSON v1.1 (Dec 2024; FDA Federal Register notice April 2025); Pinnacle 21 Community 4.0+; CORE (CDISC Open Rules Engine, 2021). Define-XML 2.1 FDA support began March 15, 2021 and is required for studies starting on/after March 15, 2023.

Before using code patterns, verify installed versions match. If versions differ:

  • Python: pip show <package> then help(module.function) to check signatures
  • R packages cited (essential for ADaM derivation): admiral (Roche/openpharma), metacore, metatools, xportr

If code throws ImportError, AttributeError, or TypeError, introspect the installed package and adapt the example to match the actual API rather than retrying.

CDISC SDTM and ADaM Data Handling

"Load clinical trial data" -> Parse CDISC SDTM domain files; build or consume ADaM analysis-ready datasets; preserve subject-level and event-level structure; respect traceability and validation expectations for regulatory submission.

  • Python: pyreadstat.read_xport(), pd.read_sas(), pd.merge()
  • R: haven::read_xpt(), admiral for ADaM derivation, Pinnacle21 or CORE for validation

Aggregation Strategy Taxonomy -- Choose the Right Question

StrategyScientific question answeredExample endpointFails when
Any event (binary)Does treatment change probability of experiencing the event at all?Had any serious AE: Yes/NoTreatment changes event burden but not anyone-event probability
Event countDoes treatment change burden of events per patient?Total AE count per subjectSubjects with 1 vs 10 events treated equivalently
Maximum severityDoes treatment shift patients toward more severe manifestations?Worst AESEV per subjectConfounded with event count (more events -> higher chance of severe)
First event + timeDoes treatment delay onset of the event?Time to first serious AE (TTE)Multiple events per subject ignored
Rate (events per person-time)What is the per-time-unit rate?AEs per subject-yearRequires exposure-time tracking; differential dropout biases rates
Composite (per ICH E9 R1)Event becomes part of endpoint definitionDeath = treatment failureDirection of components conflict; needs hierarchy

These are NOT interchangeable. A drug might not change the proportion with AEs (binary: no effect) but increase events per patient (count: harmful). The choice must be pre-specified in the SAP based on the scientific question, not analytic convenience.

Decision Tree by Scenario

ScenarioRecommended aggregationWhy
Primary safety endpoint, single SAE eventAny event (binary); analyse with logisticStandard regulatory; cite FDA Safety Reporting Guidance
Total adverse-event burden across studyEvent count per subject; analyse with Poisson or negative binomialCaptures all events; sandwich SE recommended
Toxicity grade comparison across armsMax severity per subject; ordinal logistic with PO checkPreserves grade ordering; cite Brant test for PO
Time-to-first AE (Kaplan-Meier visualisation)First event + time; censor non-eventsSee clinical-biostatistics/survival-analysis
Rate of exacerbations per patient-yearRate via negative binomial with offset for exposure timeStandard in COPD/asthma trials
Composite endpoint (e.g., MACE)Component-level definition with hierarchyPre-specify per ICH E9(R1) composite strategy
Stratification factor extractionUse STRATA1, STRATA2 from RANDB or DM SUPPMust appear in analysis (Kahan-Morris 2012)
Baseline value derivationVSBLFL='Y' / LBBLFL='Y'; derive latest pre-dose only if flag missingTrust SDTM flag when present

SDTM vs ADaM -- The Regulatory Layer Cake

LayerStandardPurposeGranularityExamples
Source CRFEDC systemRaw data captureForm/pageRave, Medidata, Veeva
SDTMSDTM 2.0 / SDTMIG 3.4Tabulation; "what happened"One row per observationDM, AE, EX, VS, LB, DS
ADaMADaMIG v1.3 (2021); v3.0 in developmentAnalysis-ready; "one-PROC-away from CSR table"Subject (ADSL), parameter-timepoint (BDS), occurrence (OCCDS)ADSL, ADAE, ADLB, ADTTE, ADRS
TLFSponsor SAS / R / PythonTables, listings, figures for CSROutputStatistical methods section, demographic table, primary efficacy

The ADaM Fundamental Principles (the "ROT" document):

  1. Analysis-ready (one procedure call -> the analysis result)
  2. Traceability (every value links back to SDTM via metadata)
  3. Clear/unambiguous communication via Define-XML
  4. Naming conventions (PARAM, PARAMCD, AVAL, AVALC, BASE, CHG, PCHG, ABLFL, ANL01FL, ...)
  5. Structural rules (ADSL one row per subject; BDS one row per subject/parameter/timepoint/analysis flag)

Postdoc reading: the ADaM IG v1.3 PDF (cdisc.org), ADaM ROT, FDA Study Data Technical Conformance Guide (current 2024 version), Pinnacle 21 validation rule catalog, PHUSE Connect 2023-2025 conference proceedings.

SDTM Domain Overview

DomainLevelDescriptionKey Variables
DMSubjectDemographics (one row per subject)USUBJID, ARM, ARMCD, ACTARM, ACTARMCD, AGE, SEX, RACE, RFSTDTC, RFXSTDTC, RFENDTC
AEEventAdverse events (multiple per subject)USUBJID, AETERM, AEDECOD, AEBODSYS, AESEV, AESER, AESTDTC, AEENDTC
EXEventDrug exposure/dosingUSUBJID, EXTRT, EXDOSE, EXSTDTC, EXENDTC
VSEventVital signsUSUBJID, VSTESTCD, VSSTRESN, VSBLFL, VISIT
LBEventLab resultsUSUBJID, LBTESTCD, LBSTRESN, LBSTRESC, LBORRES, LBBLFL, LBSPEC
DSEventDispositionUSUBJID, DSDECOD, DSSTDTC
SEEventSubject elements (treatment epochs)USUBJID, ETCD, SESTDTC, SEENDTC
MHEventMedical historyUSUBJID, MHDECOD, MHCAT
CMEventConcomitant medicationsUSUBJID, CMDECOD, CMSTDTC, CMENDTC

USUBJID = STUDYID-SITEID-SUBJID is the universal merge key. Subject-level domains (DM) have one row per USUBJID; event-level domains have multiple.

ARM vs ACTARM: ARM is planned treatment from randomisation; ACTARM is actual treatment received. In crossover designs, ARM differs from ACTARM by definition; in parallel-arm trials, they diverge when subjects are randomised to one arm but receive another (per-protocol violations). Primary analyses use ARM (ITT); safety uses ACTARM.

RFSTDTC vs RFXSTDTC: RFSTDTC is "first study activity date" (typically screening start); RFXSTDTC is "first treatment date." For treatment-emergent adverse event (TEAE) calculations, ALWAYS use RFXSTDTC (per ICH E2A) — RFSTDTC includes screening AEs which are not treatment-emergent.

Reading .xpt Files

python
import pyreadstat
import pandas as pd

# pyreadstat (recommended -- handles SAS metadata)
dm, meta = pyreadstat.read_xport('dm.xpt')
# meta.column_names, meta.column_labels, meta.variable_value_labels

# pandas built-in (SAS XPORT v5)
dm = pd.read_sas('dm.xpt', format='xport', encoding='utf-8')

# CSV fallback (common in academic datasets)
dm = pd.read_csv('DM.csv')

When pyreadstat is available, the metadata object provides column labels, value labels, and format information lost with other readers. Critical for analysis-dataset derivation: the metadata carries the controlled-terminology codelist, essential for handling values like AESEV ('MILD'/'MODERATE'/'SEVERE') with semantic ordering.

SAS XPT v5 vs Dataset-JSON -- The 2025-2026 Transition

SAS XPT v5 is the current FDA submission format but dates to 1995, with constraints:

  • 8-character variable names (so LBSTRESN is a max-length name)
  • 200-character text values
  • No UTF-8 (ASCII only) -> problematic for multilingual trials
  • Single dataset per file

Dataset-JSON v1.1 (CDISC, December 2024; FDA Federal Register notice April 2025) is the modern replacement. PHUSE-CDISC-FDA pilot has demonstrated drop-in feasibility. FDA adoption timeline pending as of mid-2026; EMA and PMDA exploring in parallel.

Pragmatic position: for the next ~2 years, SAS XPT v5 will remain the de facto submission format; sponsors should architect for Dataset-JSON migration but maintain XPT compliance.

Joining Domains -- The Right Way

python
import pandas as pd

dm = pd.read_csv('DM.csv')
ae = pd.read_csv('AE.csv')

# WRONG: merging event-level directly onto subject-level inflates rows
# RIGHT: aggregate first, then merge
any_serious = ae.groupby('USUBJID')['AESER'].apply(lambda x: (x == 'Y').any()).reset_index()
any_serious.columns = ['USUBJID', 'HAD_SERIOUS_AE']

analysis = dm.merge(any_serious, on='USUBJID', how='left')
analysis['HAD_SERIOUS_AE'] = analysis['HAD_SERIOUS_AE'].fillna(False)

Always use how='left' when merging onto DM to preserve all randomised subjects, even those with no events. Fill missing event indicators with 0 or False.

Aggregation strategy must follow the scientific question
StrategyScientific questionExample
Any event (binary)Does treatment increase probability of experiencing the event at all?Had any serious AE: yes/no
Event countDoes treatment increase event burden per patient?Total AE count per subject
Maximum severityDoes treatment shift toward more severe manifestations?Worst AESEV per subject
First event + timeDoes treatment delay onset?Time to first serious AE
Rate (events per person-time)What is the per-time-unit rate?AEs per subject-year

These are NOT interchangeable. A drug might not change the proportion with AEs (binary: no effect) but increase events per patient (count: harmful). The choice must follow the SAP, not analytic convenience.

python
# Count events per subject
ae_counts = ae.groupby('USUBJID').size().reset_index(name='AE_COUNT')

# Maximum severity per subject (map to numeric first -- string max is unreliable)
severity_map = {'MILD': 1, 'MODERATE': 2, 'SEVERE': 3}
ae['AESEV_NUM'] = ae['AESEV'].map(severity_map)
max_severity = ae.groupby('USUBJID')['AESEV_NUM'].max().reset_index()

# Specific event: COVID-19 adverse event
covid_ae = ae[ae['AEDECOD'] == 'COVID-19']
covid_ae['AESEV_NUM'] = covid_ae['AESEV'].map(severity_map)
had_covid = covid_ae.groupby('USUBJID')['AESEV_NUM'].max().reset_index()
had_covid.columns = ['USUBJID', 'COVID_SEVERITY']

analysis = dm.merge(had_covid, on='USUBJID', how='left')
analysis['HAD_COVID'] = analysis['COVID_SEVERITY'].notna().astype(int)

ADaM Architecture -- The Postdoc Deep Dive

ADSL (Subject-Level) -- The Spine

Exactly one row per subject. Every other ADaM dataset must merge to ADSL on USUBJID. Standard variables:

  • USUBJID -- universal subject ID
  • TRT01A / TRT01P / ACTARMCD / ARMCD -- planned and actual treatment, period 1
  • TRTSDT / TRTEDT -- treatment start/end dates (derived from EX, not SDTM)
  • AGE, SEX, RACE, ETHNIC -- demographics from DM
  • RANDDT -- randomisation date
  • DCSREAS / DCSREASP / DCSREASCD -- discontinuation reason (coded + verbatim)
  • Population flags: ITTFL, FASFL, SAFFL, PPROTFL, EFFFL (Y/N flags for analysis populations)
  • Stratification factors: STRATA1, STRATA2 (from randomisation)
  • Baseline covariates that will be used as model covariates downstream
BDS (Basic Data Structure) -- Long Format Analysis Data

One row per subject per parameter per analysis timepoint per analysis flag. Used for ADVS, ADLB, ADEFF, ADQS, ADTTE.

Required variables:

  • USUBJID, STUDYID -- merge keys
  • PARAM, PARAMCD, PARAMN -- parameter name, code, number
  • AVISIT, AVISITN -- analysis visit name, number
  • ADT, ADY -- analysis date, analysis day (relative to TRTSDT)
  • AVAL, AVALC -- analysis value (numeric, character)
  • BASE -- baseline value (replicated per subject/parameter)
  • CHG, PCHG -- change from baseline, percent change
  • ABLFL -- 'Y' for the baseline record
  • ANL01FL, ANL02FL -- analysis flags for primary/secondary analyses
  • DTYPE -- derivation type ('LOCF', 'WOCF', 'AVERAGE', 'BOCF', or null for original)
  • BASETYPE -- when multiple baselines per subject/parameter (crossover)
  • EPOCH -- study period (SCREENING, TREATMENT, FOLLOW-UP)
python
# Example: derive ADLB BDS structure from LB SDTM
import pandas as pd

lb = pd.read_csv('LB.csv')
adsl = pd.read_csv('ADSL.csv')

# Filter to active tests
adlb = lb[lb['LBTESTCD'].isin(['ALT', 'AST', 'CREAT', 'HGB'])].copy()
adlb['AVAL'] = adlb['LBSTRESN']
adlb['PARAM'] = adlb['LBTEST']
adlb['PARAMCD'] = adlb['LBTESTCD']

# Merge subject-level treatment from ADSL
adlb = adlb.merge(adsl[['USUBJID', 'TRT01A', 'TRTSDT']], on='USUBJID')

# Compute analysis day
adlb['ADT'] = pd.to_datetime(adlb['LBDTC'], errors='coerce')
adlb['TRTSDT'] = pd.to_datetime(adlb['TRTSDT'], errors='coerce')
adlb['ADY'] = (adlb['ADT'] - adlb['TRTSDT']).dt.days + 1  # Day 1 = first dose

# Set ABLFL from LBBLFL
adlb['ABLFL'] = (adlb['LBBLFL'] == 'Y').map({True: 'Y', False: None})

# Derive BASE per subject/parameter
baselines = adlb[adlb['ABLFL'] == 'Y'][['USUBJID', 'PARAMCD', 'AVAL']]
baselines.columns = ['USUBJID', 'PARAMCD', 'BASE']
adlb = adlb.merge(baselines, on=['USUBJID', 'PARAMCD'], how='left')

# Compute CHG and PCHG
adlb['CHG'] = adlb['AVAL'] - adlb['BASE']
adlb['PCHG'] = 100 * adlb['CHG'] / adlb['BASE']
OCCDS (Occurrence Data Structure) -- One Row per Event

OCCDS v1.1 (Nov 2021) handles AE, CM, MH. Variables: AEDECOD, AEBODSYS, AESEV, AESER, ASTDT (analysis start date), AENDT, TRTEMFL.

OCCDS v1.1 added TRTEM01FL through TRTEM##FL for multi-period treatment-emergent flags — essential for crossover and multi-phase studies where a single TRTEMFL is ambiguous.

ADTTE (Time-to-Event) -- The CNSR Convention Trap

The ADaM BDS for TTE v1.0 uses BDS structure with extra variables for survival analysis:

  • STARTDT -- time origin (typically TRTSDT for OS; RANDDT for PFS; response date for DOR)
  • ADT -- analysis date (event date if event, censoring date if censored)
  • AVAL = ADT - STARTDT (+1 if "first day = day 1" convention)
  • AVALU = 'DAYS' (or 'MONTHS' for some endpoints)
  • CNSR -- censoring indicator. CONVENTION: CNSR = 0 for events; positive integers for censoring, integer encodes censoring reason
  • EVNTDESC -- text description ('Death due to disease', 'Last alive contact')
  • CNSDTDSC, SRCDOM, SRCVAR, SRCSEQ -- traceability back to SDTM source

The CNSR convention is OPPOSITE to most statistical packages, which use 1 = event. R survival::Surv(time, event) expects event=1; SAS PROC LIFETEST takes CENSORED= statement that's opposite to CNSR convention. This is a perpetual bug source. When passing ADTTE to analysis:

python
# Convert CDISC ADTTE CNSR to R/Python statistical convention
adtte['event'] = (adtte['CNSR'] == 0).astype(int)  # 1 = event for survival packages
Define-XML 2.1

Every ADaM dataset requires variable-level metadata in Define-XML 2.1 (FDA-required for studies starting on/after March 15, 2023; support began March 15, 2021). Fields per variable:

  • Origin -- CRF, derived, predecessor SDTM variable
  • Derivation rule -- free text or controlled algorithm
  • Codelist -- linked controlled terminology
  • Length, datatype, label

The FDA reviewer's Analysis Data Reviewer's Guide (ADRG) is now expected in every NDA/BLA — walks reviewer through how each analysis dataset was built.

Two-level traceability expectation: SDTM raw -> ADaM analysis-ready, with no orphan derivations. FDA reviewers explicitly trace AE counts in CSR table -> ADAE rows -> AE SDTM rows. Any break is a flag.

Treatment-Emergent AE -- The Convention Variation

ICH E2A (1995) defines an AE generically. TEAE is sponsor-defined:

TRTEMFL = 'Y' if AE.ASTDT >= TRTSDT AND AE.ASTDT <= TRTEDT + X days

X = post-treatment follow-up window. Common values:

  • Small molecules: 28 or 30 days
  • Biologics with extended half-life: longer (e.g., 60-90 days for mAbs)
  • Cell/gene therapy: indefinite (lifelong monitoring expected)

Sponsor variation:

  • Day-of-first-dose AE included as TEAE (FDA preference) vs excluded (some EMA reviewers)
  • Partial-date imputation: impute day 15 if only month/year known, vs censor as missing
  • Worsening of pre-existing AE: flagged via SEV change vs requires new PT (preferred term)

MedDRA SOC/PT hierarchy: AEs coded to MedDRA Preferred Terms (PT), grouped by System Organ Class (SOC). Clinically related PTs (e.g., 'Diarrhea' / 'Frequent bowel movements' / 'Loose stools') often combined via Standardized MedDRA Queries (SMQs) or sponsor-defined groupings. ADAE typically carries both AEDECOD (PT) and SMQ/group flags.

Show full SKILL.md (1,151 more words)Show less

Baseline Derivation

ABLFL = 'Y' marks the record whose AVAL becomes BASE for all other records of the same subject/parameter.

Standard rule: last non-missing assessment on or before first dose (TRTSDT). If protocol mandates a specific baseline visit ('Day 1 pre-dose'), that visit's record is flagged.

python
# Derive ABLFL when SDTM baseline flag is missing/inconsistent
import pandas as pd

vs['VSDTC_dt'] = pd.to_datetime(vs['VSDTC'], errors='coerce')
vs = vs.merge(adsl[['USUBJID', 'TRTSDT']], on='USUBJID')
vs['is_pre_treatment'] = vs['VSDTC_dt'] <= pd.to_datetime(vs['TRTSDT'])

# Latest pre-treatment value per subject/parameter
baseline_records = (vs[vs['is_pre_treatment'] & vs['VSSTRESN'].notna()]
                    .sort_values('VSDTC_dt')
                    .groupby(['USUBJID', 'VSTESTCD'])
                    .tail(1))
baseline_records['derived_ABLFL'] = 'Y'

Critical detail: filter on VSBLFL='Y' (or LBBLFL='Y') as the primary source. Only fall back to derivation when the flag is missing. Trust the SDTM flag when present; CRF-level baseline designation embeds clinical judgement the analyst cannot reconstruct.

BASETYPE required when more than one baseline exists per subject/parameter (crossover studies, multi-period trials). Distinguishes "Period 1 Baseline" vs "Period 2 Baseline."

DTYPE values per CDISC controlled terminology: 'LOCF' (last observation carried forward), 'WOCF' (worst), 'AVERAGE', 'BOCF' (baseline observation carried forward), null for original. Pinnacle 21 flags any DTYPE value not in CT.

SUPPQUAL and the NSV Transition

SUPPQUAL (supplemental qualifiers) is the legacy mechanism for sponsor-defined variables that don't fit standard SDTM domain columns. Long-format QNAM/QVAL pairs:

python
supp = pd.read_sas('suppae.xpt', format='xport', encoding='utf-8')
supp_pivot = supp.pivot_table(
    index='USUBJID', columns='QNAM', values='QVAL', aggfunc='first'
).reset_index()
ae_enriched = ae.merge(supp_pivot, on='USUBJID', how='left')

For record-level SUPPQUAL (where IDVAR and IDVARVAL identify specific rows):

python
supp_record = supp[supp['IDVAR'] == 'AESEQ'].copy()
supp_record['AESEQ'] = supp_record['IDVARVAL'].astype(float)
supp_pivot_record = supp_record.pivot_table(
    index=['USUBJID', 'AESEQ'], columns='QNAM', values='QVAL', aggfunc='first'
).reset_index()
ae_enriched = ae.merge(supp_pivot_record, on=['USUBJID', 'AESEQ'], how='left')

The 2024-2026 SUPP transition: Therapeutic Area User Guides (TAUGs) increasingly use NS-- domain extensions or Non-Standard Variables (NSV) Registry-listed variables directly in the parent domain, instead of QNAM/QVAL pairs in SUPP--. Not a hard deprecation but the direction is clear. The Non-Standard Variables Registry at cdisc.org is the new canonical place to look up sponsor-extension variables.

Date Handling -- The Partial-Date Reality

python
dm['RFSTDT'] = pd.to_datetime(dm['RFSTDTC'], errors='coerce')
ae['AESTDT'] = pd.to_datetime(ae['AESTDTC'], errors='coerce')
ae['AEENDT'] = pd.to_datetime(ae['AEENDTC'], errors='coerce')

# Days from randomization to AE onset
ae_with_ref = ae.merge(dm[['USUBJID', 'RFSTDT']], on='USUBJID')
ae_with_ref['AE_ONSET_DAY'] = (ae_with_ref['AESTDT'] - ae_with_ref['RFSTDT']).dt.days

SDTM dates are ISO 8601 strings. Partial dates (e.g., '2023-03' without day) are common. errors='coerce' converts these to NaT rather than raising errors. For analysis requiring complete dates, CDISC conventions impute missing day as the 1st for start dates and the last day of the month for end dates, but imputation rules should match the SAP.

SDTM records include EPOCH (SCREENING, TREATMENT, FOLLOW-UP). For TEAEs, filter AEs to onset during or after the treatment epoch. Including pre-treatment AEs confounds the treatment effect estimate.

Validation -- Pinnacle 21 and CORE

Pinnacle 21 (Certara, formerly OpenCDISC) is the de facto FDA submission validation standard. Validates against FDA Validation Rules + CDISC IG conformance + Define-XML schema. Severity tiers:

  • Reject -- submission will not be accepted
  • Error -- must justify
  • Warning -- should investigate

FDA Validation Rules are published quarterly by FDA Office of Translational Sciences; Pinnacle 21 wraps these into its rule engine.

CORE (CDISC Open Rules Engine, 2021) is a newer open-source alternative using YAML-defined rules from the CDISC Rules Catalog. Gaining traction but not yet at Pinnacle-21 parity for confirmatory submissions.

bash
# Pinnacle 21 Community (free; appropriate for non-pivotal trials)
p21-community validate --rules sdtmig-3.4 --output-dir validation_output study_data/

Population Flags

FlagSourcePurpose
ITTFLDM all randomisedPrimary efficacy population (ICH E9 default)
FASFLITT minus eligibility failures + no post-baselinePractical primary (FAS = Full Analysis Set)
SAFFLEX (received at least one dose)Safety analysis (AE reporting)
PPROTFLSE + DS + protocol-violation listPer-protocol; sensitivity only
EFFFLSponsor-definedModified ITT variants

FAS vs ITT subtlety: FAS may exclude post-randomisation subjects (ineligibility, no post-baseline efficacy); ITT cannot. Many SAPs equate them; FDA may insist on stricter ITT at submission. Pre-specify both with explicit FAS exclusion criteria in the protocol.

Missing Data Considerations -- The Clinical Reasoning Layer

Before any imputation/complete-case decision, examine the DS (Disposition) domain to tabulate reasons for discontinuation by treatment arm. If discontinuation rates or reasons differ between arms, missing data is likely informative (MNAR) and standard MMRM-MAR is questionable.

python
ds = pd.read_csv('DS.csv')
dropouts = ds[ds['DSDECOD'] != 'COMPLETED']
dropouts_by_arm = dropouts.merge(adsl[['USUBJID', 'ARM']], on='USUBJID')
discontinuation_reasons = dropouts_by_arm.groupby(['ARM', 'DSDECOD']).size().unstack(fill_value=0)

This is the data-quality precursor to choosing the estimand strategy in trial-reporting (see ICH E9(R1)) — missing patterns informed by DS drive the choice between treatment-policy, hypothetical, or composite ICE strategies.

Common Pitfalls

PitfallSymptomSolution
Event-level merged onto subject-level without aggregationRow count inflates after mergeAggregate first, then merge
First chronological record used as baselineMisclassified baselineFilter on VSBLFL='Y' / LBBLFL='Y'; derive only if missing
Character (xxORRES) used for analysisInconsistent numeric coercionUse xxSTRESN (numeric standardised); missing xxSTRESN with present xxORRES means 'NOT DONE' or '<LLOQ'
ARM used in safety analysisCrossover or actual-treatment differsUse ACTARM for safety; ARM for ITT efficacy
RFSTDTC used as TEAE referenceIncludes screening AEsUse RFXSTDTC (first treatment); cite ICH E2A
ADTTE CNSR confused with stat-pkg conventionWrong event/censoring assignmentCDISC: CNSR=0 means event; convert: event = (CNSR == 0).astype(int)
Partial date parsing errorNaT in date columnpd.to_datetime(..., errors='coerce')
SUPPQUAL granularity confusionWrong rows mergedCheck IDVAR before choosing subject vs record-level merge
Non-standard column names treated as standard SDTMMissing variablesInspect actual columns; map to semantic roles
Pinnacle 21 not run before submissionFDA rejectAlways validate against current SDTMIG and FDA Validation Rules before submission
Common non-standard column mappings
Standard SDTMCommon alternativesRole
ARM / ARMCDTRTGRP, TRT01P, treatment, groupTreatment assignment
AEDECODAEPT, ae_term, preferred_termAE preferred term
AESEV (text)AESEV (numeric 1-4), severity, AETOXGRSeverity / toxicity grade
USUBJIDSUBJID, subject_id, patient_idSubject identifier
RFXSTDTCtrt_start, first_dose_dateFirst treatment date
LBSTRESNlab_value_num, result_numericLab numeric result

Quantitative Thresholds and Conventions

Threshold/ConventionSourceRationale
TEAE window: 28-30 days post-treatment for small moleculesICH E2A; sponsor conventionMode of action, half-life inform window
RFXSTDTC for TEAE reference (not RFSTDTC)ICH E2ARFSTDTC includes screening; TEAE is post-treatment
ABLFL='Y' for last non-missing pre-doseCDISC ADaM IG v1.3Standard baseline definition; trust SDTM flag
ARM (planned) for ITT efficacy; ACTARM for safetyICH E9Crossover/PP-violation handling
Pinnacle 21 validation before submissionFDA Study Data Technical Conformance GuideStandard quality gate; reject errors block acceptance
Define-XML 2.1 for studies starting >=March 15, 2023FDA Study Data Standards CatalogOlder 2.0 still accepted for prior studies
Dataset-JSON v1.1 (Dec 2024; FDA notice April 2025)CDISC + FDA Federal RegisterModern replacement for XPT v5; timeline pending
CNSR=0 for events, positive integers for censoringADaM BDS-for-TTE v1.0OPPOSITE of R survival and most stat packages

Anticipated Reviewer Pushback

PushbackResponse
"RFXSTDTC or RFSTDTC for TEAE?"RFXSTDTC per ICH E2A; RFSTDTC would include screening AEs
"Baseline from VSBLFL or derived?"VSBLFL when present; documented derivation rule when missing
"Pinnacle 21 errors?"All errors resolved or justified; warnings reviewed and documented
"Define-XML 2.1 ADRG provided?"Yes — analysis dataset traceability documented to variable level
"ITT vs FAS reconciliation?"Pre-specified in protocol with explicit FAS exclusion criteria
"OCCDS v1.1 multi-period flags?"TRTEM01FL...TRTEM##FL pre-specified for crossover periods
"ADTTE CNSR convention conversion documented?"Explicit: CDISC CNSR=0 means event; convert to event=1 for downstream R/Python

References

  • CDISC. 2021. Analysis Data Model Implementation Guide (ADaMIG) v1.3.
  • CDISC. 2021. Occurrence Data Structure (OCCDS) v1.1.
  • CDISC. 2012. ADaM Basic Data Structure for Time-to-Event Analyses v1.0.
  • CDISC. 2024. Dataset-JSON v1.1.
  • FDA. 2024. Study Data Technical Conformance Guide.
  • FDA Federal Register Notice. April 2025. Dataset-JSON Pilot Comment Request.
  • ICH. 1995. E2A: Clinical Safety Data Management -- Definitions and Standards for Expedited Reporting.
  • ICH. 1998. E9: Statistical Principles for Clinical Trials.
  • ICH. 2019. E9(R1) Addendum on Estimands and Sensitivity Analysis.
  • Pinnacle 21 (Certara). 2024. Community Edition Validation Rules.
  • PHUSE/CDISC. 2024. Dataset-JSON Pilot Reports.
  • clinical-biostatistics/logistic-regression - Model binary outcomes from prepared ADaM/SDTM data
  • clinical-biostatistics/trial-reporting - Use prepared analysis datasets for ICH E9(R1) estimands and CONSORT 2025 reporting
  • clinical-biostatistics/missing-data-sensitivity - DS-domain reasoning informs estimand choice
  • clinical-biostatistics/survival-analysis - ADTTE CNSR convention; time-to-event preparation
  • expression-matrix/metadata-joins - General metadata joining patterns

© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files in clinical-biostatistics/cdisc-data-handling of GPTomics/bioSkills.

  • SKILL.md
  • examples/cdisc_data_preparation.py
  • usage-guide.md

Open the folder on GitHubat commit d91ed3d

Used in 2 other repositories

We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 2 other GitHub owners. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Bio Clinical Biostatistics Cdisc Data next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Bio Clinical Biostatistics Cdisc Data compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Bio Clinical Biostatistics Cdisc Data this skillGPTomics/bioSkills1.2k2 repos~7.3kAutomated safety check: PassMIT
Clinical Trials Databasegoogle-deepmind/science-skills3.2k2 repos~3.2kAutomated safety check: PassApache-2.0
CHARLS Paper Reproduction Guidexjtulyc/MedgeClaw6171 repos~1.8kAutomated safety check: PassNone
Biomedical Analysis Dispatchxjtulyc/MedgeClaw6171 repos~2kAutomated safety check: PassNone
Research Paperluwill/research-skills862—~1.9kAutomated safety check: PassNone
Research Proposalluwill/research-skills862—~4.5kAutomated safety check: NotesNone

Similar skills

  • Clinical Trials Database

    google-deepmind/science-skills

    Query ClinicalTrials.gov via APIv2. An agent skill from google-deepmind/science-skills.

    3.2k GitHub starsUsed in 2 repos~3.2k tokens
    Research & ScienceAuto-check passed
  • Guides an agent through reproducing papers built on the CHARLS health and retirement survey, from variable mapping to cognition, depression and isolation scores.

    617 GitHub starsUsed in 1 repo~1.8k tokens
    Research & ScienceAuto-check passed
  • Routes bioinformatics, drug discovery, clinical and multi-omics tasks from a chat interface to Claude Code sessions running K-Dense scientific skills, with a live dashboard per task.

    617 GitHub starsUsed in 1 repo~2k tokens
    Research & ScienceAuto-check passed
  • Research Paper

    luwill/research-skills

    A skill your agent uses when the user asks to write or draft an ORIGINAL RESEARCH ARTICLE — IMRaD paper, conference paper, short/workshop paper, 研究论文/期刊论文/会议论文 — reporting their own completed…

    862 GitHub stars~1.9k tokensUpdated yesterday
    Research & ScienceAuto-check passed
  • Research Proposal

    luwill/research-skills

    A skill your agent uses when the user asks to write or draft a PhD / doctoral research proposal, research plan, 研究计划书, or 开题报告 — a forward-looking plan of background, gap, research questions…

    862 GitHub stars~4.5k tokensUpdated yesterday
    Research & ScienceAuto-check: notes
  • Medical Imaging Review

    LeonChaoX/qinyan-academic-skills

    Write comprehensive literature reviews for medical imaging AI research.

    944 GitHub starsUsed in 3 repos~1.1k tokens
    Research & ScienceAuto-check: notes

More from GPTomics/bioSkills

All 559 skills in this repo
  • Bio Alignment Io

    GPTomics/bioSkills

    Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.

    1.2k GitHub starsUsed in 3 repos~4.9k tokens
    Auto-check passed
  • bioSkills Installer

    GPTomics/bioSkills

    Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.

    1.2k GitHub starsUsed in 1 repo~789 tokens
    Auto-check passed
  • Bio Write Sequences

    GPTomics/bioSkills

    Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.

    1.2k GitHub starsUsed in 3 repos~2.1k tokens
    Auto-check passed
  • Amplicon Primer Clipping

    GPTomics/bioSkills

    Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.

    1.2k GitHub starsUsed in 2 repos~2.2k tokens
    Auto-check passed
  • Filters BAM alignments by FLAG bits, mapping quality and regions with samtools view or pysam, with recipes for common keep and drop cases.

    1.2k GitHub starsUsed in 2 repos~3.6k tokens
    Auto-check passed
  • Bio Alignment Indexing

    GPTomics/bioSkills

    Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.

    1.2k GitHub starsUsed in 2 repos~2.4k tokens
    Auto-check passed

Questions about Bio Clinical Biostatistics Cdisc Data

What does Bio Clinical Biostatistics Cdisc Data do?

Reads, validates, and prepares CDISC SDTM and ADaM clinical trial data for analysis. Bio Clinical Biostatistics Cdisc Data is an agent skill from GPTomics/bioSkills. Reads, validates, and prepares CDISC SDTM and ADaM clinical trial data for analysis.

When should I use Bio Clinical Biostatistics Cdisc Data?

Bio Clinical Biostatistics Cdisc Data fits situations like: working with clinical trial datasets in CDISC SDTM/ADaM format; preparing analysis-ready data; validating for regulatory submission.

How do I install Bio Clinical Biostatistics Cdisc Data in Claude Code?

Run `npx skills add GPTomics/bioSkills --skill bio-clinical-biostatistics-cdisc-data -a claude-code`. Or copy the skill folder (clinical-biostatistics/cdisc-data-handling in GPTomics/bioSkills) into .claude/skills/bio-clinical-biostatistics-cdisc-data in your project. Claude Code loads it when a task matches its description.

How do I install Bio Clinical Biostatistics Cdisc Data in Codex?

Run `npx skills add GPTomics/bioSkills --skill bio-clinical-biostatistics-cdisc-data -a codex`. Or copy the skill folder (clinical-biostatistics/cdisc-data-handling in GPTomics/bioSkills) into .agents/skills/bio-clinical-biostatistics-cdisc-data in your project. Codex loads it when a task matches its description.

Can I use Bio Clinical Biostatistics Cdisc Data in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-clinical-biostatistics-cdisc-data -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-clinical-biostatistics-cdisc-data, .gemini/skills/bio-clinical-biostatistics-cdisc-data, .github/skills/bio-clinical-biostatistics-cdisc-data and .opencode/skills/bio-clinical-biostatistics-cdisc-data in your project.

What does Bio Clinical Biostatistics Cdisc Data need to run?

Going by SKILL.md and its folder, Bio Clinical Biostatistics Cdisc Data needs Python for the scripts in its folder and the command-line tools its instructions call (pip). Our summary lists: Python 3.

Does Bio Clinical Biostatistics Cdisc Data access the network?

SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Bio Clinical Biostatistics Cdisc Data safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Bio Clinical Biostatistics Cdisc Data use?

Bio Clinical Biostatistics Cdisc Data is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Bio Clinical Biostatistics Cdisc Data use?

About 7.3k tokens (SKILL.md is roughly 29k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Bio Clinical Biostatistics Cdisc Data?

Skills that share tags, products or a category with Bio Clinical Biostatistics Cdisc Data: Clinical Trials Database (google-deepmind/science-skills, 3.2k stars), CHARLS Paper Reproduction Guide (xjtulyc/MedgeClaw, 617 stars), Biomedical Analysis Dispatch (xjtulyc/MedgeClaw, 617 stars) and Research Paper (luwill/research-skills, 862 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Bio Clinical Biostatistics Cdisc Data?

GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,218 GitHub stars. The repository holds 559 skills in this directory. The repository was last updated on August 15, 2026.

Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.