Agent skill

Clean Data

by Aperivue in Aperivue/medsci-skills

A skill your agent uses when a clinical CSV/Excel dataset needs profiling and cleaning before analysis (missing values, outliers, duplicates, type mismatches).

MITAuto-check passedData & Analytics

Install Clean Data

skills CLI
$ npx skills add Aperivue/medsci-skills --skill clean-data -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Aperivue/medsci-skills clean-data --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Aperivue/medsci-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/clean-data .claude/skills/clean-data && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
clean-data
GitHub stars
329
Token cost
~2k tokens
SKILL.md length
883 words
Files
15 (incl. scripts, references)
Skills in repo
54
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when a clinical CSV/Excel dataset needs profiling and cleaning before analysis (missing values, outliers, duplicates, type mismatches).

  • Works in 3 steps: Profiling → Flagging → Code Generation
  • A clinical CSV/Excel dataset needs profiling and cleaning before analysis (missing values
  • SKILL.md covers Three-Stage Workflow and Output Format
  • Runs Python and Shell scripts from its folder

What it does

Clean Data is an agent skill from Aperivue/medsci-skills. Use when a clinical CSV/Excel dataset needs profiling and cleaning before analysis (missing values, outliers, duplicates, type mismatches). Profiles, flags and generates cleaning code in three stages, each gated on the researcher's approval. Never auto-cleans.

Its SKILL.md is about 2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 18 other files, including scripts and reference files (for example `references/cleaning_patterns.md`, `references/implausible_value_rules.md` and `references/profiling_template.py`).

It sits in Data & Analytics, covering Data cleaning and Excel spreadsheets. It works with Microsoft Excel. The repository describes itself as: Agent Skills for medical research — literature search, reporting-guideline & citation checks, statistics, publication figures, submission. Works with Claude Code, Codex, Cursor &… The licence is MIT.

When your agent uses it

  • A clinical CSV/Excel dataset needs profiling and cleaning before analysis (missing values
  • Type mismatches)

Example prompts

  • “/clean-data”

Requirements

  • Python 3
  • A Bash shell

Workflow steps

3 steps, taken from the step headings in SKILL.md.

  1. Profiling
  2. Flagging
  3. Code Generation

What it can do on your machine

Read from SKILL.md and the folder at commit 3b14ae2. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 2 files in scripts/ (Python and Shell), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Clean Data loads about 2k tokens when it runs, and up to ~9.6k if it reads all its reference files. Until then it costs about 68 tokens; SKILL.md has 883 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~68
When it runs · the whole SKILL.md, loaded when a task matches
~2k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~9.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from Aperivue/medsci-skills at commit 3b14ae2, republished under its MIT licence (© Aperivue). 883 words, ~2,012 tokens.

Download SKILL.mdSave it as .claude/skills/clean-data/SKILL.md (or your agent's skills folder). This skill also uses 14 other files; get the full folder from GitHub.
name
clean-data
description
Use when a clinical CSV/Excel dataset needs profiling and cleaning before analysis (missing values, outliers, duplicates, type mismatches). Profiles, flags and generates cleaning code in three stages, each gated on the researcher's approval. Never auto-cleans.
metadata.triggers
clean data, data cleaning, data preprocessing, data profiling, missing values, outliers, check my data, data quality

Data Profiling and Cleaning Skill

Profile, flag, and generate cleaning code for a clinical dataset in three stages, each ending at a user-approval gate. You generate code and reports; you do NOT auto-clean data. Every cleaning decision needs the researcher's explicit confirmation, because clinical cleaning calls need domain knowledge.

PHI first. If the dataset contains PHI or PII, run /deidentify before proceeding, and use *_deidentified.* files when they exist in the working directory. Otherwise work from the data dictionary / codebook alone, or in a local-only environment with no network access. The skill writes code that runs on the data; it does not need to see the raw rows to write it.

Three-Stage Workflow

Stage 1: Profiling

Input: CSV/Excel file path OR data dictionary/codebook

  1. Adapt ${CLAUDE_SKILL_DIR}/references/profiling_template.py (pandas) to the dataset. It reports: variable and row counts, data types; missing count and percentage per variable; unique counts for categorical variables; min/max/mean/median/SD for numeric variables; histograms and bar charts.
  2. If the user provides a codebook, cross-reference variable names, expected types, and expected ranges. Never guess a column name or coding: when a mapping is uncertain, write [VERIFY: variable_name] and ask the user to confirm it against the data dictionary.
  3. Present the summary (see Output Format). Every count and statistic comes from the executed script's output.

Gate: The user reviews the profile. Ask whether to proceed to Stage 2 (Flagging) and whether any variables should be excluded or focused on.

Stage 2: Flagging

Read ${CLAUDE_SKILL_DIR}/references/cleaning_patterns.md for missing-data mechanisms, outlier decision rules, duplicate detection, date handling, and clinical pitfalls (inequality-prefixed lab values, mixed units, sentinel values). Flag issues in these categories:

  1. Missing values: variables with >5% missing; pattern analysis (MCAR/MAR/MNAR heuristic).
  2. Statistical outliers: IQR method (Q1 - 1.5IQR, Q3 + 1.5IQR) and Z-score (|z| > 3).
  3. Duplicates: exact row duplicates AND near-duplicates (same patient ID, different dates).
  4. Type mismatches: numeric stored as string, dates in inconsistent formats.
  5. Implausible values: flag against the codebook's valid range or, when the codebook is silent, the domain-default hard physiologic bounds in references/implausible_value_rules.md §1. An implausible value is a likely data-entry/unit/sentinel error (correct or set missing); a statistical outlier (#2) is biologically possible (keep + sensitivity analysis). Check units before calling a bound violation an error. Never auto-fix. 5b. Cross-field inconsistencies: logical contradictions per references/implausible_value_rules.md §2 — temporal ordering (birth ≤ event ≤ death, admission ≤ discharge), derived-vs-source (recomputed BMI/age; subset ≤ superset; total = sum of parts), sex-/state-specific fields, and min ≤ max / diastolic < systolic pairs. Name the rule that fired; a hard contradiction is High severity.
  6. Category inconsistencies: typos and variants in categorical values ("Male", "male", "M", "MALE").
  7. Categorical-implied zeros: when a category defines a natural zero for a dose/duration variable (smoking_status == 'never' implies pack_years == 0; alcohol_use == 'never' implies grams_per_week == 0), flag records that store the implied zero as NULL. This is a contradiction, not a missing-data pattern: complete-case models silently drop those never-smokers and MICE imputes them a non-zero dose, corrupting the exposure contrast. Suggested action: "Set dose = 0 where category == reference level; impute only the residual missingness among the exposed." Detected by scripts/check_structural_zero.py given the category↔dose mapping; pairs with /analyze-stats "Covariate Pitfalls: Structural Zeros & Dose/Duration Variables".
  8. Reverse-coded scale items: in a multi-item Likert scale that mixes positively and negatively worded items, every reverse item must be recoded (min+max) - x before the scale total or Cronbach's alpha is computed. An un-recoded reverse item correlates negatively with the rest and collapses alpha, often to a negative value. A negative alpha is a reverse-coding bug, not "multidimensional structure". Suggested action: "Recode reverse-worded items, then recompute reliability." Detected by scripts/check_reverse_coding.py (negative item-rest correlation and negative raw alpha, given the scale item columns); the recode itself is applied downstream by /analyze-stats likert_summary.py --reverse-items.
Show full SKILL.md (271 more words)Show less

Present the flag report as a table:

VariableIssue TypeCountSeveritySuggested Action
ageOutlier (IQR)3MediumReview: values 150, 200, -5
pack_yearsCategorical-implied zero12421HighSet 0 where smoking_status=='never' (structural zero, not missing)

Severity levels:

  • High: likely data errors that will affect analysis (type mismatches, impossible values)
  • Medium: potential issues that need expert review (statistical outliers, moderate missingness)
  • Low: minor inconsistencies that are easy to fix (category labels, trailing whitespace)

Gate: The user marks each row (A) Approve the suggested action, (R) Reject / keep as-is, or (M) Modify the action. Only approved actions generate cleaning code.

Stage 3: Code Generation

For ONLY user-approved actions, generate Python (or R if requested) code:

  • Missing value handling: listwise deletion, mean/median imputation, or MICE setup (code only, the user runs it and picks the variables and method)
  • Outlier handling: winsorization, removal, or keep-and-flag
  • Duplicate removal: exact dedup with logging
  • Type conversion: standardize dates, numeric parsing
  • Category harmonization: mapping table for inconsistent labels

All generated code MUST include:

  • Before/after row counts printed to console
  • Logging of every modification to a cleaning log DataFrame
  • Reproducibility: np.random.seed(42) and random.seed(42) where applicable
  • Output: cleaned CSV (a new file; the input file is never overwritten) + cleaning_log.csv
  • Clear comments explaining each cleaning step

End the generated script with this notice:

"This code implements ONLY the cleaning rules you approved. Review the cleaning_log.csv output to verify all changes before proceeding to analysis."

Out of scope: free-text extraction from clinical notes, and image data or DICOM metadata. After cleaning, hand off to /analyze-stats. Take any citation from /search-lit, never from memory.

Output Format

Structure all reports using this template:

## Data Profiling Report

### Dataset Overview
- Rows: [N]
- Columns: [N]
- File size: [size]
- Date range: [if applicable]

### Variable Summary
| Variable | Type | Missing N (%) | Unique | Min | Max | Mean | SD |
|----------|------|---------------|--------|-----|-----|------|-----|
| ...      | ...  | ...           | ...    | ... | ... | ...  | ... |

### Flags
| Variable | Issue | Count | Severity | Suggested Action |
|----------|-------|-------|----------|-----------------|
| ...      | ...   | ...   | ...      | ...             |

### Cleaning Code
[Python/R script -- only for approved actions]

### Cleaning Log
[What was changed, how many rows affected, before/after counts]

© Aperivue, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 14 other files (scripts, references) in skills/clean-data of Aperivue/medsci-skills.

  • SKILL.md
  • references/cleaning_patterns.md
  • references/implausible_value_rules.md
  • references/profiling_template.py
  • scripts/check_reverse_coding.py
  • scripts/check_structural_zero.py
  • skill.yml
  • tests/fixtures/profile_na_tokens.csv
  • tests/fixtures/scale_reverse.csv
  • tests/fixtures/smoking.csv
  • tests/fixtures/smoking_clean.csv
  • tests/fixtures/smoking_sentinel.csv
  • tests/test_profiling_template.sh
  • tests/test_reverse_coding.sh
  • tests/test_structural_zero.sh

Open the folder on GitHubat commit 3b14ae2

Compare with similar skills

Clean Data next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Clean Data compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Clean Data this skillAperivue/medsci-skills329—~2kAutomated safety check: PassMIT
Dataset Quality Auditzebbern/claude-code-guide4.6k—~996Automated safety check: PassMIT
Clean Dataexplorium-ai/gtm-skills160—~2kAutomated safety check: PassMIT
Outlier Detection And Quality AssessmentMichaelYang-lyx/AIDABench1111 repos~1kAutomated safety check: PassNone
Visual Skillsnpc-live/clawfirm156—~7.4kAutomated safety check: PassNone
Invalid Data CleaningMichaelYang-lyx/AIDABench1111 repos~410Automated safety check: PassNone

Similar skills

  • Dataset Quality Audit

    zebbern/claude-code-guide

    Run comprehensive quality checks on tabular data (CSV/Excel/TSV/JSON), detecting missing values, duplicates, outliers, format issues, and type inconsistencies to produce an overall score, grade, and…

    4.6k GitHub stars~996 tokensUpdated today
    Data & AnalyticsAuto-check passed
  • Clean Data

    explorium-ai/gtm-skills

    Data cleaning, entity matching, and deduplication skill for Claude Code and Codex: triage, standardize, and validate a CSV, Excel, or JSON list of B2B companies or contacts before enrichment.

    160 GitHub stars~2k tokensUpdated 13 days ago
    Data & AnalyticsAuto-check passed
  • 执行全面的异常值检测与数据质量评估,利用 IQR 方法识别异常值并结合偏度、峰度分析数据分布特征,适用于非正态分布数据的预处理阶段。

    111 GitHub starsUsed in 1 repo~1k tokens
    Data & AnalyticsAuto-check passed
  • Visual Skills

    npc-live/clawfirm

    A skill your agent uses whenever the user provides data (CSV, JSON, table, pasted numbers, or any structured dataset) and expects a visual output — even if they don't say 'chart' or 'visualize'.

    156 GitHub stars~7.4k tokensUpdated 3 mo ago
    Data & AnalyticsAuto-check passed
  • Invalid Data Cleaning

    MichaelYang-lyx/AIDABench

    用于大规模Excel数据的预处理,通过统计总行数判断是否转换为Parquet格式以提升读写效率,并使用正则表达式清洗指定文本列(如仅保留中文字符),最后导出清洗后的文件并提供下载链接。

    111 GitHub starsUsed in 1 repo~410 tokens
    Data & AnalyticsAuto-check passed
  • Data Cleaning

    ericrisco/rsc-harness

    A skill your agent uses when a raw table is too dirty to trust — nulls, sentinels, duplicate rows, category sprawl, mixed types, bad dates — and you need a re-runnable clean() plus a schema gate…

    156 GitHub stars~3.6k tokensUpdated yesterday
    Data & AnalyticsAuto-check passed

More from Aperivue/medsci-skills

All 54 skills in this repo
  • Model Assessment

    Aperivue/medsci-skills

    A skill your agent uses when validating or evaluating a trained medical-imaging model.

    329 GitHub starsUsed in 1 repo~4.5k tokens
    Auto-check passed
  • Obsidian Paper Vault

    Aperivue/medsci-skills

    A skill your agent uses when turning a folder of research PDFs into Obsidian notes, even if Obsidian is not named.

    329 GitHub stars~1.6k tokensUpdated 2 days ago
    Auto-check passed
  • Radiomics ML

    Aperivue/medsci-skills

    A skill your agent uses when building or auditing a radiomics or tabular clinical-ML prediction model with a classical learner (LASSO, SVM, random forest, XGBoost and similar).

    329 GitHub starsUsed in 1 repo~2.7k tokens
    Auto-check passed
  • Design Study

    Aperivue/medsci-skills

    A skill your agent uses when checking a radiology or medical AI study design before drafting or submission.

    329 GitHub stars~3.9k tokensUpdated 2 days ago
    Auto-check passed
  • Fill Icmje Coi

    Aperivue/medsci-skills

    A skill your agent uses when each author needs an ICMJE Conflict of Interest disclosure form (coidisclosure.docx) for submission.

    329 GitHub stars~1.5k tokensUpdated 2 days ago
    Auto-check passed
  • Fill Protocol

    Aperivue/medsci-skills

    A skill your agent uses when an institutional Word form (.doc/.docx IRB protocol, ethics application, grant template) must be filled without breaking its styles, tables, fonts or page layout.

    329 GitHub stars~1.7k tokensUpdated 2 days ago
    Auto-check passed

Works with

Questions about Clean Data

What does Clean Data do?

A skill your agent uses when a clinical CSV/Excel dataset needs profiling and cleaning before analysis (missing values, outliers, duplicates, type mismatches). Clean Data is an agent skill from Aperivue/medsci-skills. Use when a clinical CSV/Excel dataset needs profiling and cleaning before analysis (missing values, outliers, duplicates, type mismatches).

When should I use Clean Data?

Clean Data fits situations like: A clinical CSV/Excel dataset needs profiling and cleaning before analysis (missing values; type mismatches).

How do I install Clean Data in Claude Code?

Run `npx skills add Aperivue/medsci-skills --skill clean-data -a claude-code`. Or copy the skill folder (skills/clean-data in Aperivue/medsci-skills) into .claude/skills/clean-data in your project. Claude Code loads it when a task matches its description.

How do I install Clean Data in Codex?

Run `npx skills add Aperivue/medsci-skills --skill clean-data -a codex`. Or copy the skill folder (skills/clean-data in Aperivue/medsci-skills) into .agents/skills/clean-data in your project. Codex loads it when a task matches its description.

Can I use Clean Data in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Aperivue/medsci-skills --skill clean-data -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/clean-data, .gemini/skills/clean-data, .github/skills/clean-data and .opencode/skills/clean-data in your project.

What does Clean Data need to run?

Going by SKILL.md and its folder, Clean Data needs Python and a shell for the scripts in its folder. Our summary lists: Python 3; A Bash shell.

Does Clean Data access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Clean Data safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Clean Data use?

Clean Data is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Clean Data use?

About 2k tokens (SKILL.md is roughly 8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 7.6k tokens, read only when the agent opens those files.

What are the alternatives to Clean Data?

Skills that share tags, products or a category with Clean Data: Dataset Quality Audit (zebbern/claude-code-guide, 4.6k stars), Clean Data (explorium-ai/gtm-skills, 160 stars), Outlier Detection And Quality Assessment (MichaelYang-lyx/AIDABench, 111 stars) and Visual Skills (npc-live/clawfirm, 156 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Clean Data?

Aperivue (a GitHub organization) maintains it in Aperivue/medsci-skills, which has 329 GitHub stars. The repository holds 54 skills in this directory. The repository was last updated on October 5, 2026.

Source: Aperivue/medsci-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.