Question2report
refraction-ray/xalpha
Turn a natural-language financial question into a polished, self-contained HTML report.
Systematic data cleaning workflows for research datasets. An agent skill from wentorai/research-plugins.
$ npx skills add wentorai/research-plugins --skill data-cleaning-pipeline -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install wentorai/research-plugins data-cleaning-pipeline --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/wentorai/research-plugins.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/analysis/wrangling/data-cleaning-pipeline .claude/skills/data-cleaning-pipeline && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "data-cleaning-pipeline" agent skill from https://github.com/wentorai/research-plugins/tree/main/skills/analysis/wrangling/data-cleaning-pipeline into .claude/skills/data-cleaning-pipeline/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "data-cleaning-pipeline", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/wentorai/research-plugins/tree/main/skills/analysis/wrangling/data-cleaning-pipelineType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add wentorai/research-plugins --skill data-cleaning-pipeline -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install wentorai/research-plugins data-cleaning-pipeline --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/wentorai/research-plugins.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/analysis/wrangling/data-cleaning-pipeline .agents/skills/data-cleaning-pipeline && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "data-cleaning-pipeline" agent skill from https://github.com/wentorai/research-plugins/tree/main/skills/analysis/wrangling/data-cleaning-pipeline into .agents/skills/data-cleaning-pipeline/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "data-cleaning-pipeline", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add wentorai/research-plugins --skill data-cleaning-pipeline -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install wentorai/research-plugins data-cleaning-pipeline --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/wentorai/research-plugins.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/analysis/wrangling/data-cleaning-pipeline .cursor/skills/data-cleaning-pipeline && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "data-cleaning-pipeline" agent skill from https://github.com/wentorai/research-plugins/tree/main/skills/analysis/wrangling/data-cleaning-pipeline into .cursor/skills/data-cleaning-pipeline/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "data-cleaning-pipeline", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/wentorai/research-plugins.git --path skills/analysis/wrangling/data-cleaning-pipeline--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add wentorai/research-plugins --skill data-cleaning-pipeline -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install wentorai/research-plugins data-cleaning-pipeline --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/wentorai/research-plugins.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/analysis/wrangling/data-cleaning-pipeline .gemini/skills/data-cleaning-pipeline && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "data-cleaning-pipeline" agent skill from https://github.com/wentorai/research-plugins/tree/main/skills/analysis/wrangling/data-cleaning-pipeline into .gemini/skills/data-cleaning-pipeline/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "data-cleaning-pipeline", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install wentorai/research-plugins data-cleaning-pipelineInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add wentorai/research-plugins --skill data-cleaning-pipeline -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/wentorai/research-plugins.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/analysis/wrangling/data-cleaning-pipeline .github/skills/data-cleaning-pipeline && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "data-cleaning-pipeline" agent skill from https://github.com/wentorai/research-plugins/tree/main/skills/analysis/wrangling/data-cleaning-pipeline into .github/skills/data-cleaning-pipeline/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "data-cleaning-pipeline", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add wentorai/research-plugins --skill data-cleaning-pipeline -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install wentorai/research-plugins data-cleaning-pipeline --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/wentorai/research-plugins.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/analysis/wrangling/data-cleaning-pipeline .opencode/skills/data-cleaning-pipeline && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "data-cleaning-pipeline" agent skill from https://github.com/wentorai/research-plugins/tree/main/skills/analysis/wrangling/data-cleaning-pipeline into .opencode/skills/data-cleaning-pipeline/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "data-cleaning-pipeline", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
data-cleaning-pipelineSystematic data cleaning workflows for research datasets. An agent skill from wentorai/research-plugins.
Data Cleaning Pipeline is an agent skill from wentorai/research-plugins. Systematic data cleaning workflows for research datasets
Its SKILL.md is about 2.1k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Data & Analytics, covering Data cleaning. The repository describes itself as: 350+ academic research skills, MCP configs, and plugins for Research-Claw and AI agents. The licence is MIT.
Read from SKILL.md and the folder at commit bf44b3c. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are python).
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Data Cleaning Pipeline loads about 2.1k tokens when it runs. Until then it costs about 20 tokens; SKILL.md has 145 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from wentorai/research-plugins at commit bf44b3c, republished under its MIT licence (© wentorai). 145 words, ~2,148 tokens.
.claude/skills/data-cleaning-pipeline/SKILL.md (or your agent's skills folder).A skill for building systematic, reproducible data cleaning pipelines for research datasets. Covers common data quality issues, step-by-step cleaning workflows, handling missing values, detecting and treating outliers, validating data integrity, and documenting cleaning decisions for reproducibility.
Data cleaning should follow a consistent, documented order. Each step builds on the previous one, and the entire pipeline should be scripted for reproducibility.
Data Cleaning Pipeline (recommended order):
1. Initial Assessment
- Load data, check dimensions, inspect dtypes
- Generate summary statistics and missing value report
- Identify structural issues (merged cells, inconsistent delimiters)
2. Structural Fixes
- Standardize column names (snake_case, no spaces)
- Fix data types (strings to numbers, dates, categories)
- Split or merge columns as needed
- Remove completely empty rows/columns
3. Deduplication
- Identify exact duplicates
- Identify near-duplicates (fuzzy matching)
- Decide keep-first, keep-last, or merge strategy
4. Missing Value Treatment
- Classify missingness mechanism (MCAR, MAR, MNAR)
- Apply appropriate imputation or exclusion strategy
- Document and justify missing data decisions
5. Outlier Detection and Treatment
- Statistical methods (IQR, z-score, Mahalanobis)
- Domain-based validation (impossible values)
- Decide: correct, cap, remove, or keep with flag
6. Consistency Checks
- Cross-field validation (age vs birth date)
- Range validation (0-100 for percentages)
- Referential integrity (foreign keys exist)
7. Documentation and Export
- Log all changes with before/after counts
- Export cleaned dataset with version number
- Save cleaning script for reproducibilityimport pandas as pd
import numpy as np
def generate_quality_report(df):
"""
Generate a comprehensive data quality report.
Run this BEFORE any cleaning to establish a baseline.
"""
report = {
"dimensions": f"{df.shape[0]} rows x {df.shape[1]} columns",
"memory_usage": f"{df.memory_usage(deep=True).sum() / 1e6:.1f} MB",
"duplicate_rows": df.duplicated().sum(),
}
col_report = []
for col in df.columns:
info = {
"column": col,
"dtype": str(df[col].dtype),
"missing_count": df[col].isna().sum(),
"missing_pct": f"{df[col].isna().mean() * 100:.1f}%",
"unique_values": df[col].nunique(),
"sample_values": str(df[col].dropna().head(3).tolist()),
}
if pd.api.types.is_numeric_dtype(df[col]):
info["min"] = df[col].min()
info["max"] = df[col].max()
info["mean"] = df[col].mean()
info["std"] = df[col].std()
col_report.append(info)
report["columns"] = col_report
return reportMissing data mechanisms (Rubin's classification):
MCAR (Missing Completely At Random):
- Missingness is unrelated to any variable
- Example: Lab samples randomly lost during transport
- Test: Little's MCAR test, compare distributions
- Safe to: Listwise delete if < 5% missing
MAR (Missing At Random):
- Missingness depends on observed variables but not the missing value
- Example: Younger participants skip income questions more often
- Test: Compare missingness patterns across groups
- Best approach: Multiple imputation, regression imputation
MNAR (Missing Not At Random):
- Missingness depends on the unobserved value itself
- Example: High-income people refuse to report income
- Cannot be tested directly from the data
- Requires: Sensitivity analysis, selection models, domain expertisefrom sklearn.impute import SimpleImputer, KNNImputer
def impute_missing_values(df, numeric_strategy="median",
categorical_strategy="mode"):
"""
Apply appropriate imputation strategies by column type.
For research data, prefer:
- Median for skewed numeric data
- Mean for normally distributed numeric data
- Mode for categorical data
- KNN for multivariate patterns
- Multiple imputation for inference (use statsmodels or mice)
"""
numeric_cols = df.select_dtypes(include=[np.number]).columns
categorical_cols = df.select_dtypes(include=["object", "category"]).columns
# Numeric imputation
if len(numeric_cols) > 0:
if numeric_strategy == "knn":
imputer = KNNImputer(n_neighbors=5)
df[numeric_cols] = imputer.fit_transform(df[numeric_cols])
else:
imputer = SimpleImputer(strategy=numeric_strategy)
df[numeric_cols] = imputer.fit_transform(df[numeric_cols])
# Categorical imputation
if len(categorical_cols) > 0:
imputer = SimpleImputer(strategy="most_frequent")
df[categorical_cols] = imputer.fit_transform(df[categorical_cols])
return dfdef detect_outliers_iqr(series, multiplier=1.5):
"""
Detect outliers using the IQR method.
Standard multiplier is 1.5 (outlier) or 3.0 (extreme outlier).
"""
q1 = series.quantile(0.25)
q3 = series.quantile(0.75)
iqr = q3 - q1
lower = q1 - multiplier * iqr
upper = q3 + multiplier * iqr
outliers = (series < lower) | (series > upper)
return outliers, lower, upper
def detect_outliers_zscore(series, threshold=3.0):
"""
Detect outliers using z-score method.
Threshold of 3.0 corresponds to 99.7% of normal distribution.
Use modified z-score (MAD-based) for skewed distributions.
"""
from scipy import stats
z_scores = np.abs(stats.zscore(series.dropna()))
outliers = z_scores > threshold
return outliersCommon domain validations:
Age: 0-120 (flag > 100)
Height (cm): 50-250
Weight (kg): 1-300
Blood pressure systolic: 60-250
Blood pressure diastolic: 30-150
Temperature (C): 30-45 for body temperature
Likert scale (1-5): only integer values 1-5
Percentage: 0-100
Latitude: -90 to 90
Longitude: -180 to 180
Year of birth: 1900-current_year
Email: matches standard regex patternclass CleaningLog:
"""
Log all cleaning operations for reproducibility.
Every step should be documented with before/after counts.
"""
def __init__(self):
self.entries = []
self.version = 0
def log_step(self, step_name, description,
rows_before, rows_after, cols_affected):
self.version += 1
self.entries.append({
"version": self.version,
"step": step_name,
"description": description,
"rows_before": rows_before,
"rows_after": rows_after,
"rows_removed": rows_before - rows_after,
"columns_affected": cols_affected,
})
def save_report(self, path):
report_df = pd.DataFrame(self.entries)
report_df.to_csv(path, index=False)Reproducibility rules:
1. Never modify the raw data file -- always save cleaned versions
2. Use version numbers (data_v1_raw, data_v2_cleaned, data_v3_final)
3. Script every step -- no manual edits in Excel
4. Document every decision (why delete, why impute, why cap)
5. Include the cleaning script in supplementary materials
6. Record software versions (pandas, numpy, R packages)
7. Set random seeds for any stochastic imputation
8. Save intermediate datasets at major checkpointsA well-documented data cleaning pipeline not only improves the quality of research findings but also strengthens the credibility of the work during peer review. Reviewers increasingly expect transparent data handling practices, and journals like PLOS ONE and Nature require data availability statements that implicitly demand reproducible preprocessing.
© wentorai, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in skills/analysis/wrangling/data-cleaning-pipeline of wentorai/research-plugins.
Open the folder on GitHubat commit bf44b3c
We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in wentorai/research-plugins, which our catalogue first saw on October 7, 2026.
Data Cleaning Pipeline next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Data Cleaning Pipeline this skillwentorai/research-plugins | 298 | 1 repos | ~2.1k | Automated safety check: Pass | MIT | |
| Question2reportrefraction-ray/xalpha | 2.7k | — | ~3.2k | Automated safety check: Pass | MIT | |
| Dingo VerifyMigoXLab/dingo | 757 | — | ~741 | Automated safety check: Notes | Apache-2.0 | |
| Data Validationplatonai/Browser4 | 1.2k | — | ~896 | Automated safety check: Pass | Apache-2.0 | |
| Pandas ProJeffallan/claude-skills | 12k | 1 repos | ~1.5k | Automated safety check: Pass | MIT | |
| Issues DeduplicationJetBrains/ideavim | 10k | — | ~1.3k | Automated safety check: Pass | MIT |
refraction-ray/xalpha
Turn a natural-language financial question into a polished, self-contained HTML report.
MigoXLab/dingo
A skill your agent uses when the user wants to fact-check an article or verify factual claims in a document.
platonai/Browser4
Validates data against common and custom rules (required fields, formats, ranges).
Jeffallan/claude-skills
Handles pandas DataFrame work: cleaning, merging, groupby aggregation, pivots, time-series resampling and memory tuning, with checks on dtypes, shapes and nulls.
JetBrains/ideavim
Handles deduplication of YouTrack issues. An agent skill from JetBrains/ideavim.
monarchjuno/vibe-investing
Fetch financial, market, economic, fundamental, news, options, crypto, ETF, index, and macro data through the OpenBB Python interface instead of the OpenBB MCP server.
wentorai/research-plugins
Craft structured research abstracts that maximize clarity and journal acceptance
wentorai/research-plugins
Manage academic citations across BibTeX, APA, MLA, and Chicago formats
wentorai/research-plugins
Summarize academic papers with structured extraction of key elements
wentorai/research-plugins
Evidence-based study techniques for academic learning and retention
wentorai/research-plugins
Adjust writing tone and register for academic audiences and venues
wentorai/research-plugins
Academic translation, post-editing, and Chinglish correction guide
Categories
Systematic data cleaning workflows for research datasets. An agent skill from wentorai/research-plugins. Data Cleaning Pipeline is an agent skill from wentorai/research-plugins.
Data Cleaning Pipeline fits situations like: tasks that involve Data cleaning.
Run `npx skills add wentorai/research-plugins --skill data-cleaning-pipeline -a claude-code`. Or copy the skill folder (skills/analysis/wrangling/data-cleaning-pipeline in wentorai/research-plugins) into .claude/skills/data-cleaning-pipeline in your project. Claude Code loads it when a task matches its description.
Run `npx skills add wentorai/research-plugins --skill data-cleaning-pipeline -a codex`. Or copy the skill folder (skills/analysis/wrangling/data-cleaning-pipeline in wentorai/research-plugins) into .agents/skills/data-cleaning-pipeline in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add wentorai/research-plugins --skill data-cleaning-pipeline -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/data-cleaning-pipeline, .gemini/skills/data-cleaning-pipeline, .github/skills/data-cleaning-pipeline and .opencode/skills/data-cleaning-pipeline in your project.
SKILL.md names no scripts, command-line tools or credentials: Data Cleaning Pipeline is instructions for the agent only. Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Data Cleaning Pipeline is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.1k tokens (SKILL.md is roughly 8.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Data Cleaning Pipeline: Question2report (refraction-ray/xalpha, 2.7k stars), Dingo Verify (MigoXLab/dingo, 757 stars), Data Validation (platonai/Browser4, 1.2k stars) and Pandas Pro (Jeffallan/claude-skills, 12k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
wentorai (a GitHub user) maintains it in wentorai/research-plugins, which has 298 GitHub stars. The repository holds 405 skills in this directory. The repository was last updated on June 19, 2026.
Source: wentorai/research-plugins on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.