Agent skill

Stata Data Cleaning

by wentorai in wentorai/research-plugins

Clean, transform, and validate messy research data using Stata

MITAuto-check passedResearch & Science

Install Stata Data Cleaning

skills CLI
$ npx skills add wentorai/research-plugins --skill stata-data-cleaning -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install wentorai/research-plugins stata-data-cleaning --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/wentorai/research-plugins.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/analysis/wrangling/stata-data-cleaning .claude/skills/stata-data-cleaning && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
stata-data-cleaning
GitHub stars
298
Used in
1 other repo
Token cost
~2.1k tokens
SKILL.md length
359 words
Files
1
Skills in repo
405
Repo updated
First seen
Licence
MIT

At a glance

Clean, transform, and validate messy research data using Stata

  • Works in 6 steps: Never modify raw data files: Always read… → Log everything: Use log using to capture… → Use assert statements: Validate… → …
  • Tasks that involve Econometrics and empirical research
  • SKILL.md covers Overview, Initial Data Assessment, String Cleaning and Missing Data Handling, plus 5 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Stata Data Cleaning is an agent skill from wentorai/research-plugins. Clean, transform, and validate messy research data using Stata

Its SKILL.md is about 2.1k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Research & Science, covering Econometrics and empirical research and Data cleaning. The repository describes itself as: 350+ academic research skills, MCP configs, and plugins for Research-Claw and AI agents. The licence is MIT.

When your agent uses it

  • Tasks that involve Econometrics and empirical research
  • Tasks that involve Data cleaning

Example prompts

  • “/stata-data-cleaning”

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Never modify raw data files: Always read raw data and write to a separate processed file.
  2. Log everything: Use log using to capture all output for audit trails.
  3. Use assert statements: Validate assumptions about the data at each stage.
  4. Document decisions: Comment every recode, drop, or imputation with the rationale.
  5. Version your cleaning scripts: Use git to track changes to .do files.
  6. Produce a data dictionary: Label every variable and value label in the final dataset.

What it can do on your machine

Read from SKILL.md and the folder at commit bf44b3c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are stata).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • stata.com
    • dimewiki.worldbank.org
    • povertyactionlab.org

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Stata Data Cleaning loads about 2.1k tokens when it runs. Until then it costs about 21 tokens; SKILL.md has 359 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~21
When it runs · the whole SKILL.md, loaded when a task matches
~2.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from wentorai/research-plugins at commit bf44b3c, republished under its MIT licence (© wentorai). 359 words, ~2,150 tokens.

Download SKILL.mdSave it as .claude/skills/stata-data-cleaning/SKILL.md (or your agent's skills folder).
name
stata-data-cleaning
description
Clean, transform, and validate messy research data using Stata

Stata Data Cleaning

Clean, transform, and validate messy research datasets in Stata. This skill covers the complete data preparation pipeline from raw survey or administrative data to analysis-ready datasets, with emphasis on documentation, reproducibility, and handling the common data quality issues encountered in social science, economics, and health research.

Overview

Data cleaning typically consumes 60-80% of research time in empirical studies, yet it is often under-documented and poorly reproducible. Stata provides a powerful set of commands for data manipulation, but knowing which commands to use and in what order requires experience with common data quality issues: inconsistent coding, duplicate observations, string formatting problems, implausible values, and complex missing data patterns.

This skill provides a systematic, step-by-step data cleaning workflow in Stata. Each step produces a log of changes made, enabling full reproducibility and audit trails. The workflow is organized around the principle that raw data should never be modified in place -- instead, cleaning scripts transform raw data into processed datasets while preserving the original.

The approach follows best practices from the World Bank's DIME Analytics team and the J-PAL research transparency guidelines, making it suitable for projects that require rigorous data documentation for peer review, replication packages, or regulatory compliance.

Initial Data Assessment

Loading and Inspecting Data
stata
* ============================================
* Data Cleaning Script: [Project Name]
* Author: [Name]
* Date: [Date]
* Input: raw/survey_data_raw.dta
* Output: processed/survey_data_clean.dta
* ============================================

clear all
set more off
log using "logs/cleaning_log.smcl", replace

* Load raw data
use "raw/survey_data_raw.dta", clear

* Basic inspection
describe
summarize
codebook, compact

* Check dimensions
display "Observations: " _N
display "Variables: " c(k)

* Check for duplicates on ID variable
duplicates report respondent_id
duplicates list respondent_id if duplicates(respondent_id) > 0
Data Quality Report
stata
* Generate a data quality summary
foreach var of varlist _all {
    quietly {
        count if missing(`var')
        local nmiss = r(N)
        local pctmiss = (`nmiss' / _N) * 100
    }
    if `pctmiss' > 0 {
        display "`var': `nmiss' missing (`pctmiss'%)"
    }
}

* Check value ranges for numeric variables
foreach var of varlist age income years_education {
    summarize `var', detail
    * Flag implausible values
    count if `var' < 0 & !missing(`var')
    count if `var' > 150 & !missing(`var')
}

String Cleaning

Standardizing Text Variables
stata
* Trim whitespace
replace name = strtrim(name)
replace name = stritrim(name)  // Remove internal multiple spaces

* Standardize case
replace city = proper(city)        // Title case
replace country = upper(country)   // Upper case
replace email = lower(email)       // Lower case

* Remove special characters
replace phone = ustrregexra(phone, "[^0-9]", "")

* Fix encoding issues
replace name = ustrfix(name)

* Standardize common variations
replace department = "Computer Science" if ///
    inlist(department, "CS", "Comp Sci", "Comp. Sci.", "CompSci")

replace gender = "Female" if inlist(gender, "F", "f", "female", "FEMALE")
replace gender = "Male" if inlist(gender, "M", "m", "male", "MALE")
Show full SKILL.md (144 more words)Show less
Parsing Complex Strings
stata
* Split full name into first and last
gen first_name = word(full_name, 1)
gen last_name = word(full_name, -1)

* Extract year from date string "March 15, 2024"
gen year = real(word(date_string, -1))

* Parse numeric values from strings like "$1,234.56"
gen income_clean = real(subinstr(subinstr(income_str, "$", "", .), ",", "", .))

Missing Data Handling

Identifying Missing Data Patterns
stata
* Install missing data analysis tools
ssc install mdesc
ssc install misstable

* Summary of missing data
mdesc

* Missing data patterns
misstable summarize
misstable patterns

* Create missing indicator variables
foreach var of varlist income education occupation {
    gen mi_`var' = missing(`var')
}

* Test whether missing is random (Little's MCAR test approximation)
* Compare means of observed variables by missing status
foreach var of varlist income education {
    ttest age, by(mi_`var')
    ttest gender_numeric, by(mi_`var')
}
Recoding Missing Values
stata
* Common survey codes for missing
* -99 = refused, -88 = don't know, -77 = not applicable
foreach var of varlist income satisfaction trust_score {
    replace `var' = .r if `var' == -99  // .r = refused
    replace `var' = .d if `var' == -88  // .d = don't know
    replace `var' = .n if `var' == -77  // .n = not applicable
}

* Extended missing values preserve the reason for missingness
* while still being treated as missing in analyses

Variable Construction

Recoding and Categorization
stata
* Create age groups
recode age (18/29 = 1 "18-29") (30/44 = 2 "30-44") ///
           (45/59 = 3 "45-59") (60/max = 4 "60+"), gen(age_group)

* Create binary indicator
gen high_income = (income > 75000) if !missing(income)

* Create composite scale (e.g., Likert items)
alpha item1 item2 item3 item4 item5, gen(scale_score) item
* Cronbach's alpha is reported; scale_score is the mean

* Standardize continuous variables
foreach var of varlist income education_years age {
    egen z_`var' = std(`var')
}

* Winsorize extreme values
winsor2 income, cuts(1 99) replace
Date Variables
stata
* Parse date strings
gen interview_date = date(date_string, "MDY")
format interview_date %td

* Extract components
gen interview_year = year(interview_date)
gen interview_month = month(interview_date)
gen interview_dow = dow(interview_date)  // 0=Sunday

* Calculate durations
gen days_since_treatment = interview_date - treatment_date
gen months_since = (interview_date - treatment_date) / 30.44

Data Validation

Assertion-Based Validation
stata
* These assertions halt execution if violated
assert _N == 5000  // Expected sample size
assert !missing(respondent_id)  // No missing IDs
assert age >= 18 & age <= 120 if !missing(age)  // Plausible age range
assert inlist(gender, "Male", "Female", "Other", "") | missing(gender)

* Cross-variable consistency checks
assert education_years >= 0 if !missing(education_years)
assert income >= 0 if !missing(income)
assert end_date >= start_date if !missing(end_date) & !missing(start_date)
Duplicate Detection and Resolution
stata
* Identify duplicates
duplicates tag respondent_id, gen(dup_flag)
list respondent_id survey_date if dup_flag > 0, sepby(respondent_id)

* Keep most recent observation per respondent
bysort respondent_id (survey_date): keep if _n == _N

* Or keep first observation
bysort respondent_id (survey_date): keep if _n == 1

Saving and Documentation

stata
* Label all variables
label variable age "Age at time of interview (years)"
label variable income "Annual household income (USD)"
label variable education_years "Total years of formal education"

* Save cleaned dataset
compress  // Reduce file size
save "processed/survey_data_clean.dta", replace

* Export codebook
codebook, compact
describe, short

* Close log
log close

Best Practices

  1. Never modify raw data files: Always read raw data and write to a separate processed file.
  2. Log everything: Use log using to capture all output for audit trails.
  3. Use assert statements: Validate assumptions about the data at each stage.
  4. Document decisions: Comment every recode, drop, or imputation with the rationale.
  5. Version your cleaning scripts: Use git to track changes to .do files.
  6. Produce a data dictionary: Label every variable and value label in the final dataset.

References

© wentorai, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/analysis/wrangling/stata-data-cleaning of wentorai/research-plugins.

Open the folder on GitHubat commit bf44b3c

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in wentorai/research-plugins, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Stata Data Cleaning next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Stata Data Cleaning compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Stata Data Cleaning this skillwentorai/research-plugins2981 repos~2.1kAutomated safety check: PassMIT
Stata Data Cleaningmeleantonio/awesome-econ-ai-stuff6462 repos~1.8kAutomated safety check: PassCustom licence
Data Cleaningbrycewang-stanford/Auto-Empirical-Research-Skills4.6k—~2.9kAutomated safety check: PassCustom licence
CHARLS Paper Reproduction Guidexjtulyc/MedgeClaw6171 repos~1.8kAutomated safety check: PassNone
Example Datasetspymc-labs/CausalPy1.2k—~587Automated safety check: PassApache-2.0
Daily PapersXiangyue-Zhang/auto-deep-researcher-24x71.3k—~309Automated safety check: PassApache-2.0

Similar skills

  • Stata Data Cleaning

    meleantonio/awesome-econ-ai-stuff

    Clean and transform messy data in Stata with reproducible workflows

    646 GitHub starsUsed in 2 repos~1.8k tokens
    Research & ScienceAuto-check passed
  • Data Cleaning

    brycewang-stanford/Auto-Empirical-Research-Skills

    Clean and transform messy data for analysis in Python, R, or Stata

    4.6k GitHub stars~2.9k tokensUpdated 5 days ago
    Data & AnalyticsAuto-check passed
  • Guides an agent through reproducing papers built on the CHARLS health and retirement survey, from variable mapping to cognition, depression and isolation scores.

    617 GitHub starsUsed in 1 repo~1.8k tokens
    Research & ScienceAuto-check passed
  • Example Datasets

    pymc-labs/CausalPy

    Load built-in CausalPy example datasets for demos, tutorials, tests, and quick causal-analysis prototypes.

    1.2k GitHub stars~587 tokensUpdated today
    Research & ScienceAuto-check passed
  • Daily Papers

    Xiangyue-Zhang/auto-deep-researcher-24x7

    Daily arXiv paper recommendations with automatic deduplication

    1.3k GitHub stars~309 tokensUpdated 4 mo ago
    Research & ScienceAuto-check passed
  • Data Profiler

    aspi6246/Claude-Code-Skills-for-Academics

    Systematic dataset profiling protocol for empirical research.

    159 GitHub stars~2k tokensUpdated 1 mo ago
    Data & AnalyticsAuto-check passed

More from wentorai/research-plugins

All 405 skills in this repo
  • Abstract Writing Guide

    wentorai/research-plugins

    Craft structured research abstracts that maximize clarity and journal acceptance

    298 GitHub starsUsed in 1 repo~1.7k tokens
    Auto-check passed
  • Academic Citation Manager

    wentorai/research-plugins

    Manage academic citations across BibTeX, APA, MLA, and Chicago formats

    298 GitHub starsUsed in 1 repo~2.7k tokens
    Auto-check passed
  • Academic Paper Summarizer

    wentorai/research-plugins

    Summarize academic papers with structured extraction of key elements

    298 GitHub starsUsed in 1 repo~1.4k tokens
    Auto-check passed
  • Academic Study Methods

    wentorai/research-plugins

    Evidence-based study techniques for academic learning and retention

    298 GitHub starsUsed in 1 repo~1.8k tokens
    Auto-check passed
  • Academic Tone Guide

    wentorai/research-plugins

    Adjust writing tone and register for academic audiences and venues

    298 GitHub starsUsed in 1 repo~1.9k tokens
    Auto-check passed
  • Academic Translation Guide

    wentorai/research-plugins

    Academic translation, post-editing, and Chinglish correction guide

    298 GitHub starsUsed in 1 repo~1.6k tokens
    Auto-check passed

Questions about Stata Data Cleaning

What does Stata Data Cleaning do?

Clean, transform, and validate messy research data using Stata. Stata Data Cleaning is an agent skill from wentorai/research-plugins.

When should I use Stata Data Cleaning?

Stata Data Cleaning fits situations like: tasks that involve Econometrics and empirical research; tasks that involve Data cleaning.

How do I install Stata Data Cleaning in Claude Code?

Run `npx skills add wentorai/research-plugins --skill stata-data-cleaning -a claude-code`. Or copy the skill folder (skills/analysis/wrangling/stata-data-cleaning in wentorai/research-plugins) into .claude/skills/stata-data-cleaning in your project. Claude Code loads it when a task matches its description.

How do I install Stata Data Cleaning in Codex?

Run `npx skills add wentorai/research-plugins --skill stata-data-cleaning -a codex`. Or copy the skill folder (skills/analysis/wrangling/stata-data-cleaning in wentorai/research-plugins) into .agents/skills/stata-data-cleaning in your project. Codex loads it when a task matches its description.

Can I use Stata Data Cleaning in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add wentorai/research-plugins --skill stata-data-cleaning -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/stata-data-cleaning, .gemini/skills/stata-data-cleaning, .github/skills/stata-data-cleaning and .opencode/skills/stata-data-cleaning in your project.

What does Stata Data Cleaning need to run?

SKILL.md names no scripts, command-line tools or credentials: Stata Data Cleaning is instructions for the agent only.

Does Stata Data Cleaning access the network?

SKILL.md names 3 domains. As links in the text: stata.com, dimewiki.worldbank.org and povertyactionlab.org. This is read from the text; nothing was executed.

Is Stata Data Cleaning safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Stata Data Cleaning use?

Stata Data Cleaning is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Stata Data Cleaning use?

About 2.1k tokens (SKILL.md is roughly 8.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Stata Data Cleaning?

Skills that share tags, products or a category with Stata Data Cleaning: Stata Data Cleaning (meleantonio/awesome-econ-ai-stuff, 646 stars), Data Cleaning (brycewang-stanford/Auto-Empirical-Research-Skills, 4.6k stars), CHARLS Paper Reproduction Guide (xjtulyc/MedgeClaw, 617 stars) and Example Datasets (pymc-labs/CausalPy, 1.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Stata Data Cleaning?

wentorai (a GitHub user) maintains it in wentorai/research-plugins, which has 298 GitHub stars. The repository holds 405 skills in this directory. The repository was last updated on June 19, 2026.

Source: wentorai/research-plugins on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.