Question2report
refraction-ray/xalpha
Turn a natural-language financial question into a polished, self-contained HTML report.
Clean and preprocess datasets by handling missing values, removing duplicates, correcting types, resolving outliers, and enforcing validation schemas.
$ npx skills add seb1n/awesome-ai-agent-skills --skill data-cleaning -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install seb1n/awesome-ai-agent-skills data-cleaning --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/seb1n/awesome-ai-agent-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/data-and-analytics/data-cleaning .claude/skills/data-cleaning && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "data-cleaning" agent skill from https://github.com/seb1n/awesome-ai-agent-skills/tree/main/data-and-analytics/data-cleaning into .claude/skills/data-cleaning/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "data-cleaning", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/seb1n/awesome-ai-agent-skills/tree/main/data-and-analytics/data-cleaningType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add seb1n/awesome-ai-agent-skills --skill data-cleaning -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install seb1n/awesome-ai-agent-skills data-cleaning --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/seb1n/awesome-ai-agent-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/data-and-analytics/data-cleaning .agents/skills/data-cleaning && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "data-cleaning" agent skill from https://github.com/seb1n/awesome-ai-agent-skills/tree/main/data-and-analytics/data-cleaning into .agents/skills/data-cleaning/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "data-cleaning", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add seb1n/awesome-ai-agent-skills --skill data-cleaning -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install seb1n/awesome-ai-agent-skills data-cleaning --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/seb1n/awesome-ai-agent-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/data-and-analytics/data-cleaning .cursor/skills/data-cleaning && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "data-cleaning" agent skill from https://github.com/seb1n/awesome-ai-agent-skills/tree/main/data-and-analytics/data-cleaning into .cursor/skills/data-cleaning/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "data-cleaning", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/seb1n/awesome-ai-agent-skills.git --path data-and-analytics/data-cleaning--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add seb1n/awesome-ai-agent-skills --skill data-cleaning -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install seb1n/awesome-ai-agent-skills data-cleaning --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/seb1n/awesome-ai-agent-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/data-and-analytics/data-cleaning .gemini/skills/data-cleaning && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "data-cleaning" agent skill from https://github.com/seb1n/awesome-ai-agent-skills/tree/main/data-and-analytics/data-cleaning into .gemini/skills/data-cleaning/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "data-cleaning", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install seb1n/awesome-ai-agent-skills data-cleaningInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add seb1n/awesome-ai-agent-skills --skill data-cleaning -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/seb1n/awesome-ai-agent-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/data-and-analytics/data-cleaning .github/skills/data-cleaning && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "data-cleaning" agent skill from https://github.com/seb1n/awesome-ai-agent-skills/tree/main/data-and-analytics/data-cleaning into .github/skills/data-cleaning/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "data-cleaning", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add seb1n/awesome-ai-agent-skills --skill data-cleaning -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install seb1n/awesome-ai-agent-skills data-cleaning --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/seb1n/awesome-ai-agent-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/data-and-analytics/data-cleaning .opencode/skills/data-cleaning && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "data-cleaning" agent skill from https://github.com/seb1n/awesome-ai-agent-skills/tree/main/data-and-analytics/data-cleaning into .opencode/skills/data-cleaning/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "data-cleaning", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
data-cleaningClean and preprocess datasets by handling missing values, removing duplicates, correcting types, resolving outliers, and enforcing validation schemas.
Data Cleaning is an agent skill from seb1n/awesome-ai-agent-skills. Clean and preprocess datasets by handling missing values, removing duplicates, correcting types, resolving outliers, and enforcing validation schemas. Use when the user requests data cleaning or provides relevant inputs for this workflow.
Its SKILL.md is about 2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Data & Analytics, covering Data cleaning. The repository describes itself as: 103 ready-to-use AI agent skills for Claude Code, OpenAI Codex, Gemini CLI, Cursor, GitHub Copilot, Windsurf, and other Agent Skills-compatible tools. Complete SKILL.md… The licence is MIT.
6 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 75865a5. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are python).
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Data Cleaning loads about 2k tokens when it runs. Until then it costs about 63 tokens; SKILL.md has 626 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from seb1n/awesome-ai-agent-skills at commit 75865a5, republished under its MIT licence (© seb1n). 626 words, ~1,962 tokens.
.claude/skills/data-cleaning/SKILL.md (or your agent's skills folder).This skill enables an AI agent to systematically clean and preprocess raw datasets into analysis-ready form. The agent handles missing values, duplicate records, data type mismatches, inconsistent formats, outlier treatment, and normalization. It can also enforce validation schemas to ensure ongoing data quality. The primary toolchain is pandas with support from pyjanitor and great_expectations for advanced validation.
Ingest and profile the raw data. Load the dataset and immediately generate a quality report: count nulls per column, identify duplicate rows, check data types against expected schema, and flag columns with mixed types. This profile drives every subsequent cleaning decision.
Handle missing values. Apply strategy per column based on data type and missingness pattern. For numeric columns with less than 5% missing, use median imputation. For categorical columns, use mode or a dedicated "Unknown" category. For columns missing more than 40%, flag them for potential removal and consult the user before dropping.
Remove duplicates and resolve conflicts. Identify exact duplicates and near-duplicates (e.g., rows differing only in whitespace or casing). For exact duplicates, keep the first occurrence. For near-duplicates, apply fuzzy matching with a configurable similarity threshold and merge conflicting values by recency or completeness.
Correct data types and standardize formats. Coerce columns to their intended types — parse date strings into datetime objects, convert numeric strings to floats, and normalize categorical values to a canonical form. Standardize formats such as phone numbers, postal codes, and currency representations.
Detect and treat outliers. Use the IQR method (1.5x) for symmetric distributions and z-scores for normally distributed data. Offer three treatment options: cap at boundary values (winsorization), replace with null for later imputation, or flag-only mode that annotates but preserves original values.
Validate the cleaned output. Run the cleaned dataset through validation rules — non-null constraints, range checks, uniqueness constraints, and referential integrity. Report any remaining violations and save the clean dataset alongside a cleaning log that documents every transformation applied.
Provide the agent with the file path to the raw dataset and optionally a schema definition specifying expected column types, valid ranges, and uniqueness constraints. The agent will produce a cleaned file and a transformation log.
import pandas as pd
import numpy as np
# Load raw data
df = pd.read_csv("messy_orders.csv")
print(f"Raw shape: {df.shape}") # (2340, 8)
print(df.isnull().sum())
# order_id 0
# customer_name 12
# email 45
# order_date 18
# amount 23
# status 0
# region 67
# discount 0
# 1. Fix data types — order_date has mixed formats
df["order_date"] = pd.to_datetime(df["order_date"], format="mixed", dayfirst=False)
df["amount"] = pd.to_numeric(df["amount"], errors="coerce")
# 2. Handle missing values
df["customer_name"] = df["customer_name"].fillna("Unknown")
df["email"] = df["email"].fillna("missing@placeholder.com")
df["amount"] = df["amount"].fillna(df["amount"].median())
df["region"] = df["region"].fillna(df["region"].mode()[0])
df["order_date"] = df["order_date"].fillna(method="ffill")
# 3. Remove duplicates
before = len(df)
df = df.drop_duplicates(subset=["order_id"], keep="first")
print(f"Removed {before - len(df)} duplicate orders") # Removed 34 duplicate orders
# 4. Standardize categorical values
df["status"] = df["status"].str.strip().str.lower().replace({
"shipped": "shipped", "ship": "shipped",
"cancelled": "cancelled", "canceled": "cancelled",
"pending": "pending", "pend": "pending"
})
df["region"] = df["region"].str.strip().str.title()
# 5. Outlier treatment — cap amounts at IQR bounds
Q1 = df["amount"].quantile(0.25)
Q3 = df["amount"].quantile(0.75)
IQR = Q3 - Q1
lower, upper = Q1 - 1.5 * IQR, Q3 + 1.5 * IQR
df["amount"] = df["amount"].clip(lower=lower, upper=upper)
print(f"Clean shape: {df.shape}") # (2306, 8)
df.to_csv("clean_orders.csv", index=False)import great_expectations as gx
context = gx.get_context()
# Define a validation suite
suite = context.add_expectation_suite("orders_validation")
# Add expectations
suite.add_expectation(
gx.expectations.ExpectColumnValuesToNotBeNull(column="order_id")
)
suite.add_expectation(
gx.expectations.ExpectColumnValuesToBeBetween(
column="amount", min_value=0.01, max_value=50000.00
)
)
suite.add_expectation(
gx.expectations.ExpectColumnValuesToBeInSet(
column="status", value_set=["pending", "shipped", "delivered", "cancelled"]
)
)
suite.add_expectation(
gx.expectations.ExpectColumnValuesToBeUnique(column="order_id")
)
suite.add_expectation(
gx.expectations.ExpectColumnValuesToMatchRegex(
column="email", regex=r"^[^@]+@[^@]+\.[^@]+$"
)
)
# Run validation against cleaned data
results = context.run_validation(suite, batch=gx.read_csv("clean_orders.csv"))
print(f"Success: {results.success}")
print(f"Passed: {results.statistics['successful_expectations']}/{results.statistics['evaluated_expectations']}")
# Success: True
# Passed: 5/5errors="coerce" with pd.to_numeric and pd.to_datetime to surface conversion failures as NaNs rather than crashing._1, _2) before any operations.read_csv raises a UnicodeDecodeError, retry with encoding="latin-1" then encoding="cp1252" and log which encoding succeeded.pd.to_datetime(col, format="mixed") and verify the parsed results with spot checks.$, €, commas, and whitespace before type coercion: df["price"].str.replace(r"[$€,\s]", "", regex=True).astype(float).© seb1n, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in data-and-analytics/data-cleaning of seb1n/awesome-ai-agent-skills.
Open the folder on GitHubat commit 75865a5
Data Cleaning next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Data Cleaning this skillseb1n/awesome-ai-agent-skills | 206 | — | ~2k | Automated safety check: Pass | MIT | |
| Question2reportrefraction-ray/xalpha | 2.7k | — | ~3.2k | Automated safety check: Pass | MIT | |
| Dingo VerifyMigoXLab/dingo | 757 | — | ~741 | Automated safety check: Notes | Apache-2.0 | |
| Data Validationplatonai/Browser4 | 1.2k | — | ~896 | Automated safety check: Pass | Apache-2.0 | |
| Issues DeduplicationJetBrains/ideavim | 10k | — | ~1.3k | Automated safety check: Pass | MIT | |
| Pandas ProJeffallan/claude-skills | 12k | 1 repos | ~1.5k | Automated safety check: Pass | MIT |
refraction-ray/xalpha
Turn a natural-language financial question into a polished, self-contained HTML report.
MigoXLab/dingo
A skill your agent uses when the user wants to fact-check an article or verify factual claims in a document.
platonai/Browser4
Validates data against common and custom rules (required fields, formats, ranges).
JetBrains/ideavim
Handles deduplication of YouTrack issues. An agent skill from JetBrains/ideavim.
Jeffallan/claude-skills
Handles pandas DataFrame work: cleaning, merging, groupby aggregation, pivots, time-series resampling and memory tuning, with checks on dtypes, shapes and nulls.
monarchjuno/vibe-investing
Fetch financial, market, economic, fundamental, news, options, crypto, ETF, index, and macro data through the OpenBB Python interface instead of the OpenBB MCP server.
seb1n/awesome-ai-agent-skills
Plan, execute, document, and retest authorized security assessments of AI agents and multi-agent workflows using safe adversarial cases, synthetic identities, canaries, and evidence-based findings.
seb1n/awesome-ai-agent-skills
Build a preliminary, evidence-based EU AI Act readiness assessment across AI-system inventory, territorial scope, operator roles, prohibited-practice screening, risk classification, transparency…
seb1n/awesome-ai-agent-skills
Design and verify auditable human oversight, approval gates, escalation paths, and safe state transitions for AI agent workflows.
seb1n/awesome-ai-agent-skills
Design, implement, harden, and verify Model Context Protocol (MCP) servers with precise tool contracts, least-privilege authorization, safe transports, structured errors, and interoperability tests.
seb1n/awesome-ai-agent-skills
Audit agent skills, plugins, prompts, manifests, scripts, dependencies, and bundled assets for provenance, prompt-injection, permission, execution, exfiltration, persistence, and update risk.
seb1n/awesome-ai-agent-skills
Inspect, profile, clean, reconcile, analyze, visualize, and verify spreadsheet data while preserving formulas, formatting, types, and source files.
Categories
Clean and preprocess datasets by handling missing values, removing duplicates, correcting types, resolving outliers, and enforcing validation schemas. Data Cleaning is an agent skill from seb1n/awesome-ai-agent-skills. Clean and preprocess datasets by handling missing values, removing duplicates, correcting types, resolving outliers, and enforcing validation schemas.
Data Cleaning fits situations like: the user requests data cleaning; provides relevant inputs for this workflow.
Run `npx skills add seb1n/awesome-ai-agent-skills --skill data-cleaning -a claude-code`. Or copy the skill folder (data-and-analytics/data-cleaning in seb1n/awesome-ai-agent-skills) into .claude/skills/data-cleaning in your project. Claude Code loads it when a task matches its description.
Run `npx skills add seb1n/awesome-ai-agent-skills --skill data-cleaning -a codex`. Or copy the skill folder (data-and-analytics/data-cleaning in seb1n/awesome-ai-agent-skills) into .agents/skills/data-cleaning in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add seb1n/awesome-ai-agent-skills --skill data-cleaning -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/data-cleaning, .gemini/skills/data-cleaning, .github/skills/data-cleaning and .opencode/skills/data-cleaning in your project.
SKILL.md names no scripts, command-line tools or credentials: Data Cleaning is instructions for the agent only. Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Data Cleaning is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 2k tokens (SKILL.md is roughly 7.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Data Cleaning: Question2report (refraction-ray/xalpha, 2.7k stars), Dingo Verify (MigoXLab/dingo, 757 stars), Data Validation (platonai/Browser4, 1.2k stars) and Issues Deduplication (JetBrains/ideavim, 10k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
seb1n (a GitHub user) maintains it in seb1n/awesome-ai-agent-skills, which has 206 GitHub stars. The repository holds 91 skills in this directory. The repository was last updated on August 9, 2026.
Source: seb1n/awesome-ai-agent-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.