Agent skill

Clinical Data Cleaner

by aipoch in aipoch/medical-research-skills

A skill your agent uses when cleaning clinical trial data, preparing data for FDA/EMA submission, standardizing SDTM datasets, handling missing values in clinical studies, detecting outliers in lab…

MITAuto-check passedData & Analytics

Install Clinical Data Cleaner

skills CLI
$ npx skills add aipoch/medical-research-skills --skill clinical-data-cleaner -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install aipoch/medical-research-skills clinical-data-cleaner --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/aipoch/medical-research-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/'scientific-skills/Data Analysis/clinical-data-cleaner' .claude/skills/clinical-data-cleaner && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
clinical-data-cleaner
GitHub stars
2k
Token cost
~2.4k tokens
SKILL.md length
967 words
Files
9 (incl. scripts, references)
Skills in repo
578
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when cleaning clinical trial data, preparing data for FDA/EMA submission, standardizing SDTM datasets, handling missing values in clinical studies, detecting outliers in lab…

  • Works in 5 steps: SDTM Domain Validation → Missing Value Handling → Outlier Detection → …
  • Cleaning clinical trial data
  • SKILL.md covers Quick Check, Audit-Ready Commands, When to Use and Workflow, plus 17 more sections
  • Runs Python scripts from its folder; calls python

What it does

Clinical Data Cleaner is an agent skill from aipoch/medical-research-skills. Use when cleaning clinical trial data, preparing data for FDA/EMA submission, standardizing SDTM datasets, handling missing values in clinical studies, detecting outliers in lab results, or converting raw CRF data to CDISC format. Cleans and standardizes clinical trial data fo...

Its SKILL.md is about 2.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 10 other files, including scripts and reference files (for example `POLISH_CHANGELOG.md`, `eval_report_clinical-data-cleaner_result.json` and `references/common-patterns.md`).

It sits in Data & Analytics, covering Clinical and healthcare research and Data cleaning. The repository describes itself as: Hundreds of agent skills for medical research, including protocol design, data analysis, evidence insights, and academic writing. The licence is MIT.

When your agent uses it

  • Cleaning clinical trial data
  • Preparing data for FDA/EMA submission
  • Standardizing SDTM datasets
  • Handling missing values in clinical studies

Example prompts

  • “/clinical-data-cleaner”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. SDTM Domain Validation
  2. Missing Value Handling
  3. Outlier Detection
  4. Date Standardization
  5. Complete Pipeline

What it can do on your machine

Read from SKILL.md and the folder at commit 686e09d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Clinical Data Cleaner loads about 2.4k tokens when it runs, and up to ~7.1k if it reads all its reference files. Until then it costs about 76 tokens; SKILL.md has 967 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~76
When it runs · the whole SKILL.md, loaded when a task matches
~2.4k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~7.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from aipoch/medical-research-skills at commit 686e09d, republished under its MIT licence (© aipoch). 967 words, ~2,433 tokens.

Download SKILL.mdSave it as .claude/skills/clinical-data-cleaner/SKILL.md (or your agent's skills folder). This skill also uses 8 other files; get the full folder from GitHub.
name
clinical-data-cleaner
description
Use when cleaning clinical trial data, preparing data for FDA/EMA submission, standardizing SDTM datasets, handling missing values in clinical studies, detecting outliers in lab results, or converting raw CRF data to CDISC format. Cleans and standardizes clinical trial data fo...
license
MIT
author
AIPOCH

Source: https://github.com/aipoch/medical-research-skills

Clinical Data Cleaner

Clean, validate, and standardize clinical trial data to meet CDISC SDTM standards for regulatory submissions to FDA or EMA.

Quick Check

Use this command to verify that the packaged script entry point can be parsed before deeper execution.

bash
python -m py_compile scripts/main.py

Audit-Ready Commands

Use these concrete commands for validation. They are intentionally self-contained and avoid placeholder paths.

bash
python -m py_compile scripts/main.py
python scripts/main.py --help
python scripts/main.py --input "Audit validation sample with explicit symptoms, history, assessment, and next-step plan."

When to Use

  • Use this skill when the task needs Use when cleaning clinical trial data, preparing data for FDA/EMA submission, standardizing SDTM datasets, handling missing values in clinical studies, detecting outliers in lab results, or converting raw CRF data to CDISC format. Cleans and standardizes clinical trial data for regulatory compliance with audit trails.
  • Use this skill for data analysis tasks that require explicit assumptions, bounded scope, and a reproducible output format.
  • Use this skill when you need a documented fallback path for missing inputs, execution errors, or partial evidence.

Workflow

  1. Confirm the user objective, required inputs, and non-negotiable constraints before doing detailed work.
  2. Validate that the request matches the documented scope and stop early if the task would require unsupported assumptions.
  3. Use the packaged script path or the documented reasoning path with only the inputs that are actually available.
  4. Return a structured result that separates assumptions, deliverables, risks, and unresolved items.
  5. If execution fails or inputs are incomplete, switch to the fallback path and state exactly what blocked full completion.

Quick Start

python
from scripts.main import ClinicalDataCleaner

# Initialize for Demographics domain
cleaner = ClinicalDataCleaner(domain='DM')

# Clean data with default settings
cleaned = cleaner.clean(raw_data)

# Save with audit trail
cleaner.save_report('output.csv')

Core Capabilities

1. SDTM Domain Validation
python
cleaner = ClinicalDataCleaner(domain='DM')  # or 'LB', 'VS'
is_valid, missing = cleaner.validate_domain(data)

Required Fields:

  • DM: STUDYID, USUBJID, SUBJID, RFSTDTC, RFENDTC, SITEID, AGE, SEX, RACE
  • LB: STUDYID, USUBJID, LBTESTCD, LBCAT, LBORRES, LBORRESU, LBSTRESC, LBDTC
  • VS: STUDYID, USUBJID, VSTESTCD, VSORRES, VSORRESU, VSSTRESC, VSDTC
2. Missing Value Handling
python
cleaner = ClinicalDataCleaner(
    domain='DM',
    missing_strategy='median'  # mean, median, mode, forward, drop
)
cleaned = cleaner.handle_missing_values(data)
3. Outlier Detection
python
cleaner = ClinicalDataCleaner(
    domain='LB',
    outlier_method='domain',  # iqr, zscore, domain
    outlier_action='flag'     # flag, remove, cap
)
flagged = cleaner.detect_outliers(data)

Clinical Thresholds:

ParameterRangeUnit
Glucose50-500mg/dL
Hemoglobin5-20g/dL
Systolic BP70-220mmHg
4. Date Standardization
python
standardized = cleaner.standardize_dates(data)
# Converts to ISO 8601: 2023-01-15T09:30:00
5. Complete Pipeline
python
cleaner = ClinicalDataCleaner(
    domain='DM',
    missing_strategy='median',
    outlier_method='iqr',
    outlier_action='flag'
)
cleaned_data = cleaner.clean(data)
cleaner.save_report('output.csv')

Output Files:

  • output.csv - Cleaned SDTM data
  • output.report.json - Audit trail for regulatory submission

CLI Usage

text
# Clean demographics
python scripts/main.py \
  --input dm_raw.csv \
  --domain DM \
  --output dm_clean.csv \
  --missing-strategy median \
  --outlier-method iqr \
  --outlier-action flag

# Clean lab data with clinical thresholds
python scripts/main.py \
  --input lb_raw.csv \
  --domain LB \
  --output lb_clean.csv \
  --outlier-method domain

Common Patterns

See references/common-patterns.md for detailed examples:

  • Regulatory Submission Preparation
  • Interim Analysis Data Preparation
  • Database Migration Cleanup
  • External Lab Data Integration

Troubleshooting

See references/troubleshooting.md for solutions to:

  • Validation failures
  • Date parsing errors
  • Memory errors with large datasets
  • Outlier detection issues

Quality Checklist

Pre-Cleaning:

  • IACUC approval obtained (animal studies)
  • Sample size adequately powered
  • Randomization method documented

Post-Cleaning:

  • Validate against CDISC SDTM IG
  • Review all cleaning actions in audit trail
  • Test import to analysis software

References

  • references/sdtm_ig_guide.md - CDISC SDTM Implementation Guide
  • references/domain_specs.json - Domain-specific field requirements
  • references/outlier_thresholds.json - Clinical outlier thresholds
  • references/common-patterns.md - Detailed usage patterns
  • references/troubleshooting.md - Problem-solving guide

Skill ID: 189 | Version: 2.0 | License: MIT

Output Requirements

Every final response should make these items explicit when they are relevant:

  • Objective or requested deliverable
  • Inputs used and assumptions introduced
  • Workflow or decision path
  • Core result, recommendation, or artifact
  • Constraints, risks, caveats, or validation needs
  • Unresolved items and next-step checks

Error Handling

  • If required inputs are missing, state exactly which fields are missing and request only the minimum additional information.
  • If the task goes outside the documented scope, stop instead of guessing or silently widening the assignment.
  • If scripts/main.py fails, report the failure point, summarize what still can be completed safely, and provide a manual fallback.
  • Do not fabricate files, citations, data, search results, or execution outcomes.

Input Validation

This skill accepts requests that match the documented purpose of clinical-data-cleaner and include enough context to complete the workflow safely.

Do not continue the workflow when the request is out of scope, missing a critical input, or would require unsupported assumptions. Instead respond:

clinical-data-cleaner only handles its documented workflow. Please provide the missing required inputs or switch to a more suitable skill.

Show full SKILL.md (377 more words)Show less

Response Template

Use the following fixed structure for non-trivial requests:

  1. Objective
  2. Inputs Received
  3. Assumptions
  4. Workflow
  5. Deliverable
  6. Risks and Limits
  7. Next Checks

If the request is simple, you may compress the structure, but still keep assumptions and limits explicit when they affect correctness.

When Not to Use

  • Do not proceed when required input files, identifiers, parameters, or context are missing — ask the user to provide them first.
  • Do not assume capabilities beyond this skill's declared scope when the user requests external operations or inferences.
  • Do not proceed without user confirmation when overwriting existing results, executing high-cost batch operations, or expanding task scope.

Required Inputs

FieldRequiredFormat/SourceExampleIf Missing
User task descriptionYesTextResearch question, writing goal, analysis objectiveStop and ask user to provide
Primary input materialDepends on taskText, file path, ID, table, or literaturePMID, PDF, CSV, DOCX, keywords, etc.Specify which material type is missing
Output preferenceNoTextLanguage, format, target journal, templateUse skill default format

Output Contract

  • Primary output: Structured result or target file aligned with this skill's objective.
  • Optional output: Intermediate check notes, issue list, supplementary suggestions, or generated file paths.
  • Format requirement: Unless the user specifies otherwise, prefer stable, reviewable Markdown or JSON; if the skill's bundled script requires a fixed format, use that format.
  • If partially complete: Must explicitly mark as PARTIAL and state which steps are completed and which remain.

Failure Handling

  • Missing critical input: Explicitly state which fields, files, or identifiers are missing and pause.
  • Script, template, or resource execution failure: Report the failing step, likely cause, and recovery suggestions — do not silently degrade.
  • Partial completion only: Return the verified portion first, then list remaining blockers and suggested next steps.

User Checkpoints

  • Before executing batch processing, overwriting files, long-running searches, or multi-stage generation, confirm scope and output format with the user.
  • Before proceeding when a key judgment is ambiguous, evidence is insufficient, or the workflow is entering the next stage, confirm with the user.

Quick Validation

  • Check that key scripts, templates, or reference file paths this skill depends on exist.
  • Check that the final output contains the core fields, sections, or files specified for this task.
  • Check that results clearly mark assumptions, limitations, and incomplete items.

© aipoch, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 8 other files (scripts, references) in scientific-skills/Data Analysis/clinical-data-cleaner of aipoch/medical-research-skills.

  • SKILL.md
  • POLISH_CHANGELOG.md
  • eval_report_clinical-data-cleaner_result.json
  • references/common-patterns.md
  • references/domain_specs.json
  • references/outlier_thresholds.json
  • references/sdtm_ig_guide.md
  • references/troubleshooting.md
  • scripts/main.py

Open the folder on GitHubat commit 686e09d

Compare with similar skills

Clinical Data Cleaner next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Clinical Data Cleaner compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Clinical Data Cleaner this skillaipoch/medical-research-skills2k—~2.4kAutomated safety check: PassMIT
Statistical ReviewerRConsortium/pharma-skills119—~4.8kAutomated safety check: PassNone
Model CardAperivue/medsci-skills331—~1.5kAutomated safety check: PassMIT
Lab Unit Harmonizationbenchflow-ai/skillsbench1.8k—~2.7kAutomated safety check: PassApache-2.0
CHARLS Paper Reproduction Guidexjtulyc/MedgeClaw6171 repos~1.8kAutomated safety check: PassNone
Dingo VerifyMigoXLab/dingo757—~741Automated safety check: NotesApache-2.0

Similar skills

  • Statistical Reviewer

    RConsortium/pharma-skills

    Simulates an independent statistical reviewer auditing a clinical trial submission package (SDTM, ADaM, TLG/TLF, SAP, CSR).

    119 GitHub stars~4.8k tokensUpdated 5 days ago
    Data & AnalyticsAuto-check passed
  • Model Card

    Aperivue/medsci-skills

    A skill your agent uses when a trained medical-imaging model needs its documentation.

    331 GitHub stars~1.5k tokensUpdated 4 days ago
    Data & AnalyticsAuto-check passed
  • Lab Unit Harmonization

    benchflow-ai/skillsbench

    Comprehensive clinical laboratory data harmonization for multi-source healthcare analytics.

    1.8k GitHub stars~2.7k tokensUpdated 2 mo ago
    Data & AnalyticsAuto-check passed
  • Guides an agent through reproducing papers built on the CHARLS health and retirement survey, from variable mapping to cognition, depression and isolation scores.

    617 GitHub starsUsed in 1 repo~1.8k tokens
    Research & ScienceAuto-check passed
  • Dingo Verify

    MigoXLab/dingo

    A skill your agent uses when the user wants to fact-check an article or verify factual claims in a document.

    757 GitHub stars~741 tokensUpdated 11 days ago
    Data & AnalyticsAuto-check: notes
  • OpenMed ETL to OMOP CDM

    maziyarpanahi/openmed

    Maps OpenMed-extracted, terminology-coded conditions, drugs and measurements into OMOP CDM v5.4 tables for OHDSI and ATLAS analytics.

    5.5k GitHub stars~1.9k tokensUpdated today
    Data & AnalyticsAuto-check passed

More from aipoch/medical-research-skills

All 578 skills in this repo
  • Academic Poster Generator

    aipoch/medical-research-skills

    Complete workflow for generating academic research posters from PDF literature; use when you need to extract paper content from PDFs and produce a LaTeX-based poster…

    2k GitHub stars~2.2k tokensUpdated 22 days ago
    Auto-check passed
  • Diagnostic Study Quality Assessment Quadas

    aipoch/medical-research-skills

    Analyzes clinical diagnostic accuracy studies for bias using the QUADAS-2 tool.

    2k GitHub stars~1.4k tokensUpdated 22 days ago
    Auto-check passed
  • Exploratory Data Analysis

    aipoch/medical-research-skills

    Perform comprehensive exploratory data analysis on scientific data files across 200+ file formats.

    2k GitHub stars~3.7k tokensUpdated 22 days ago
    Auto-check passed
  • Iso Certification

    aipoch/medical-research-skills

    A toolkit for preparing ISO 13485:2016 certification documentation for medical device QMS.

    2k GitHub stars~1.8k tokensUpdated 22 days ago
    Auto-check passed
  • Journal Skills

    aipoch/medical-research-skills

    Recommends target journals for manuscript submission by analyzing the paper topic/abstract and the journal distribution of similar PubMed literature; use when users ask for journal…

    2k GitHub stars~1.7k tokensUpdated 22 days ago
    Auto-check passed
  • Latex Posters

    aipoch/medical-research-skills

    Creates academic-poster writing packages for LaTeX using beamerposter, tikzposter, or baposter.

    2k GitHub stars~1.3k tokensUpdated 22 days ago
    Auto-check passed

Questions about Clinical Data Cleaner

What does Clinical Data Cleaner do?

A skill your agent uses when cleaning clinical trial data, preparing data for FDA/EMA submission, standardizing SDTM datasets, handling missing values in clinical studies, detecting outliers in lab…. Clinical Data Cleaner is an agent skill from aipoch/medical-research-skills. Use when cleaning clinical trial data, preparing data for FDA/EMA submission, standardizing SDTM datasets, handling missing values in clinical studies, detecting outliers in lab results, or converting raw CRF data to CDISC format.

When should I use Clinical Data Cleaner?

Clinical Data Cleaner fits situations like: cleaning clinical trial data; preparing data for FDA/EMA submission; standardizing SDTM datasets; handling missing values in clinical studies.

How do I install Clinical Data Cleaner in Claude Code?

Run `npx skills add aipoch/medical-research-skills --skill clinical-data-cleaner -a claude-code`. Or copy the skill folder (scientific-skills/Data Analysis/clinical-data-cleaner in aipoch/medical-research-skills) into .claude/skills/clinical-data-cleaner in your project. Claude Code loads it when a task matches its description.

How do I install Clinical Data Cleaner in Codex?

Run `npx skills add aipoch/medical-research-skills --skill clinical-data-cleaner -a codex`. Or copy the skill folder (scientific-skills/Data Analysis/clinical-data-cleaner in aipoch/medical-research-skills) into .agents/skills/clinical-data-cleaner in your project. Codex loads it when a task matches its description.

Can I use Clinical Data Cleaner in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add aipoch/medical-research-skills --skill clinical-data-cleaner -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/clinical-data-cleaner, .gemini/skills/clinical-data-cleaner, .github/skills/clinical-data-cleaner and .opencode/skills/clinical-data-cleaner in your project.

What does Clinical Data Cleaner need to run?

Going by SKILL.md and its folder, Clinical Data Cleaner needs Python for the scripts in its folder and the command-line tools its instructions call (python). Our summary lists: Python 3.

Does Clinical Data Cleaner access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Clinical Data Cleaner safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Clinical Data Cleaner use?

Clinical Data Cleaner is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Clinical Data Cleaner use?

About 2.4k tokens (SKILL.md is roughly 9.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 4.6k tokens, read only when the agent opens those files.

What are the alternatives to Clinical Data Cleaner?

Skills that share tags, products or a category with Clinical Data Cleaner: Statistical Reviewer (RConsortium/pharma-skills, 119 stars), Model Card (Aperivue/medsci-skills, 331 stars), Lab Unit Harmonization (benchflow-ai/skillsbench, 1.8k stars) and CHARLS Paper Reproduction Guide (xjtulyc/MedgeClaw, 617 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Clinical Data Cleaner?

aipoch (a GitHub organization) maintains it in aipoch/medical-research-skills, which has 1,978 GitHub stars. The repository holds 578 skills in this directory. The repository was last updated on September 17, 2026.

Source: aipoch/medical-research-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.