Agent skill

Data Cleaning

by seb1n in seb1n/awesome-ai-agent-skills

Clean and preprocess datasets by handling missing values, removing duplicates, correcting types, resolving outliers, and enforcing validation schemas.

MITAuto-check passedData & Analytics

Install Data Cleaning

skills CLI
$ npx skills add seb1n/awesome-ai-agent-skills --skill data-cleaning -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install seb1n/awesome-ai-agent-skills data-cleaning --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/seb1n/awesome-ai-agent-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/data-and-analytics/data-cleaning .claude/skills/data-cleaning && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
data-cleaning
GitHub stars
206
Token cost
~2k tokens
SKILL.md length
626 words
Files
1
Skills in repo
91
Repo updated
First seen
Licence
MIT

At a glance

Clean and preprocess datasets by handling missing values, removing duplicates, correcting types, resolving outliers, and enforcing validation schemas.

  • Works in 6 steps: Ingest and profile the raw data. Load… → Handle missing values. Apply strategy… → Remove duplicates and resolve conflicts.… → …
  • The user requests data cleaning
  • SKILL.md covers Workflow, Supported Technologies, Usage and Examples, plus 2 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Data Cleaning is an agent skill from seb1n/awesome-ai-agent-skills. Clean and preprocess datasets by handling missing values, removing duplicates, correcting types, resolving outliers, and enforcing validation schemas. Use when the user requests data cleaning or provides relevant inputs for this workflow.

Its SKILL.md is about 2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Data & Analytics, covering Data cleaning. The repository describes itself as: 103 ready-to-use AI agent skills for Claude Code, OpenAI Codex, Gemini CLI, Cursor, GitHub Copilot, Windsurf, and other Agent Skills-compatible tools. Complete SKILL.md… The licence is MIT.

When your agent uses it

  • The user requests data cleaning
  • Provides relevant inputs for this workflow

Example prompts

  • “/data-cleaning”

Requirements

  • Python 3

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Ingest and profile the raw data. Load the dataset and immediately generate a quality report: count nulls per column, identify duplicate…
  2. Handle missing values. Apply strategy per column based on data type and missingness pattern. For numeric columns with less than 5%…
  3. Remove duplicates and resolve conflicts. Identify exact duplicates and near-duplicates (e.g., rows differing only in whitespace or…
  4. Correct data types and standardize formats. Coerce columns to their intended types — parse date strings into datetime objects, convert…
  5. Detect and treat outliers. Use the IQR method (1.5x) for symmetric distributions and z-scores for normally distributed data. Offer three…
  6. Validate the cleaned output. Run the cleaned dataset through validation rules — non-null constraints, range checks, uniqueness…

What it can do on your machine

Read from SKILL.md and the folder at commit 75865a5. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Data Cleaning loads about 2k tokens when it runs. Until then it costs about 63 tokens; SKILL.md has 626 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~63
When it runs · the whole SKILL.md, loaded when a task matches
~2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from seb1n/awesome-ai-agent-skills at commit 75865a5, republished under its MIT licence (© seb1n). 626 words, ~1,962 tokens.

Download SKILL.mdSave it as .claude/skills/data-cleaning/SKILL.md (or your agent's skills folder).
name
data-cleaning
description
Clean and preprocess datasets by handling missing values, removing duplicates, correcting types, resolving outliers, and enforcing validation schemas. Use when the user requests data cleaning or provides relevant inputs for this workflow.
license
MIT
metadata.author
awesome-ai-agent-skills
metadata.version
1.0.0

Data Cleaning

This skill enables an AI agent to systematically clean and preprocess raw datasets into analysis-ready form. The agent handles missing values, duplicate records, data type mismatches, inconsistent formats, outlier treatment, and normalization. It can also enforce validation schemas to ensure ongoing data quality. The primary toolchain is pandas with support from pyjanitor and great_expectations for advanced validation.

Workflow

  1. Ingest and profile the raw data. Load the dataset and immediately generate a quality report: count nulls per column, identify duplicate rows, check data types against expected schema, and flag columns with mixed types. This profile drives every subsequent cleaning decision.

  2. Handle missing values. Apply strategy per column based on data type and missingness pattern. For numeric columns with less than 5% missing, use median imputation. For categorical columns, use mode or a dedicated "Unknown" category. For columns missing more than 40%, flag them for potential removal and consult the user before dropping.

  3. Remove duplicates and resolve conflicts. Identify exact duplicates and near-duplicates (e.g., rows differing only in whitespace or casing). For exact duplicates, keep the first occurrence. For near-duplicates, apply fuzzy matching with a configurable similarity threshold and merge conflicting values by recency or completeness.

  4. Correct data types and standardize formats. Coerce columns to their intended types — parse date strings into datetime objects, convert numeric strings to floats, and normalize categorical values to a canonical form. Standardize formats such as phone numbers, postal codes, and currency representations.

  5. Detect and treat outliers. Use the IQR method (1.5x) for symmetric distributions and z-scores for normally distributed data. Offer three treatment options: cap at boundary values (winsorization), replace with null for later imputation, or flag-only mode that annotates but preserves original values.

  6. Validate the cleaned output. Run the cleaned dataset through validation rules — non-null constraints, range checks, uniqueness constraints, and referential integrity. Report any remaining violations and save the clean dataset alongside a cleaning log that documents every transformation applied.

Supported Technologies

  • pandas — core data manipulation and type coercion
  • pyjanitor — method-chaining convenience for cleaning operations
  • great_expectations — schema validation and data quality checks
  • fuzzywuzzy — fuzzy string matching for near-duplicate detection
  • numpy — numerical operations for outlier detection
Show full SKILL.md (267 more words)Show less

Usage

Provide the agent with the file path to the raw dataset and optionally a schema definition specifying expected column types, valid ranges, and uniqueness constraints. The agent will produce a cleaned file and a transformation log.

Examples

Example 1: Cleaning a messy CSV with pandas
python
import pandas as pd
import numpy as np

# Load raw data
df = pd.read_csv("messy_orders.csv")
print(f"Raw shape: {df.shape}")  # (2340, 8)
print(df.isnull().sum())
# order_id         0
# customer_name   12
# email           45
# order_date      18
# amount          23
# status           0
# region          67
# discount         0

# 1. Fix data types — order_date has mixed formats
df["order_date"] = pd.to_datetime(df["order_date"], format="mixed", dayfirst=False)
df["amount"] = pd.to_numeric(df["amount"], errors="coerce")

# 2. Handle missing values
df["customer_name"] = df["customer_name"].fillna("Unknown")
df["email"] = df["email"].fillna("missing@placeholder.com")
df["amount"] = df["amount"].fillna(df["amount"].median())
df["region"] = df["region"].fillna(df["region"].mode()[0])
df["order_date"] = df["order_date"].fillna(method="ffill")

# 3. Remove duplicates
before = len(df)
df = df.drop_duplicates(subset=["order_id"], keep="first")
print(f"Removed {before - len(df)} duplicate orders")  # Removed 34 duplicate orders

# 4. Standardize categorical values
df["status"] = df["status"].str.strip().str.lower().replace({
    "shipped": "shipped", "ship": "shipped",
    "cancelled": "cancelled", "canceled": "cancelled",
    "pending": "pending", "pend": "pending"
})
df["region"] = df["region"].str.strip().str.title()

# 5. Outlier treatment — cap amounts at IQR bounds
Q1 = df["amount"].quantile(0.25)
Q3 = df["amount"].quantile(0.75)
IQR = Q3 - Q1
lower, upper = Q1 - 1.5 * IQR, Q3 + 1.5 * IQR
df["amount"] = df["amount"].clip(lower=lower, upper=upper)

print(f"Clean shape: {df.shape}")  # (2306, 8)
df.to_csv("clean_orders.csv", index=False)
Example 2: Data validation pipeline with schema enforcement
python
import great_expectations as gx

context = gx.get_context()

# Define a validation suite
suite = context.add_expectation_suite("orders_validation")

# Add expectations
suite.add_expectation(
    gx.expectations.ExpectColumnValuesToNotBeNull(column="order_id")
)
suite.add_expectation(
    gx.expectations.ExpectColumnValuesToBeBetween(
        column="amount", min_value=0.01, max_value=50000.00
    )
)
suite.add_expectation(
    gx.expectations.ExpectColumnValuesToBeInSet(
        column="status", value_set=["pending", "shipped", "delivered", "cancelled"]
    )
)
suite.add_expectation(
    gx.expectations.ExpectColumnValuesToBeUnique(column="order_id")
)
suite.add_expectation(
    gx.expectations.ExpectColumnValuesToMatchRegex(
        column="email", regex=r"^[^@]+@[^@]+\.[^@]+$"
    )
)

# Run validation against cleaned data
results = context.run_validation(suite, batch=gx.read_csv("clean_orders.csv"))

print(f"Success: {results.success}")
print(f"Passed: {results.statistics['successful_expectations']}/{results.statistics['evaluated_expectations']}")
# Success: True
# Passed: 5/5

Best Practices

  • Always save a copy of the raw data before any transformations — cleaning should be reproducible, not destructive.
  • Log every transformation with counts (e.g., "Filled 23 nulls in amount with median 312.45") to create an auditable cleaning trail.
  • Prefer domain-informed imputation over mechanical defaults; consult the user when a column's missingness pattern is non-random (MNAR).
  • Validate early and often — run schema checks after each cleaning phase, not just at the end.
  • Treat cleaning as iterative: the first pass catches the obvious issues, but downstream analysis frequently surfaces new ones.
  • Use errors="coerce" with pd.to_numeric and pd.to_datetime to surface conversion failures as NaNs rather than crashing.

Edge Cases

  • Entirely empty columns. If a column is 100% null, drop it automatically and log a warning rather than attempting imputation on zero information.
  • Duplicate column names. Pandas silently allows duplicate column names. Detect them on load and rename with suffixes (_1, _2) before any operations.
  • Encoding issues. If read_csv raises a UnicodeDecodeError, retry with encoding="latin-1" then encoding="cp1252" and log which encoding succeeded.
  • Date columns with multiple formats. When a single column contains "2024-01-15", "01/15/2024", and "Jan 15, 2024", use pd.to_datetime(col, format="mixed") and verify the parsed results with spot checks.
  • Numeric columns stored as strings with currency symbols. Strip $, €, commas, and whitespace before type coercion: df["price"].str.replace(r"[$€,\s]", "", regex=True).astype(float).

© seb1n, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in data-and-analytics/data-cleaning of seb1n/awesome-ai-agent-skills.

Open the folder on GitHubat commit 75865a5

Compare with similar skills

Data Cleaning next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Data Cleaning compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Data Cleaning this skillseb1n/awesome-ai-agent-skills206—~2kAutomated safety check: PassMIT
Question2reportrefraction-ray/xalpha2.7k—~3.2kAutomated safety check: PassMIT
Dingo VerifyMigoXLab/dingo757—~741Automated safety check: NotesApache-2.0
Data Validationplatonai/Browser41.2k—~896Automated safety check: PassApache-2.0
Issues DeduplicationJetBrains/ideavim10k—~1.3kAutomated safety check: PassMIT
Pandas ProJeffallan/claude-skills12k1 repos~1.5kAutomated safety check: PassMIT

Similar skills

  • Question2report

    refraction-ray/xalpha

    Turn a natural-language financial question into a polished, self-contained HTML report.

    2.7k GitHub stars~3.2k tokensUpdated 2 mo ago
    Data & AnalyticsAuto-check passed
  • Dingo Verify

    MigoXLab/dingo

    A skill your agent uses when the user wants to fact-check an article or verify factual claims in a document.

    757 GitHub stars~741 tokensUpdated 10 days ago
    Data & AnalyticsAuto-check: notes
  • Data Validation

    platonai/Browser4

    Validates data against common and custom rules (required fields, formats, ranges).

    1.2k GitHub stars~896 tokensUpdated today
    Data & AnalyticsAuto-check passed
  • Issues Deduplication

    JetBrains/ideavim

    Official

    Handles deduplication of YouTrack issues. An agent skill from JetBrains/ideavim.

    10k GitHub stars~1.3k tokensUpdated yesterday
    Data & AnalyticsAuto-check passed
  • Pandas Pro

    Jeffallan/claude-skills

    Handles pandas DataFrame work: cleaning, merging, groupby aggregation, pivots, time-series resampling and memory tuning, with checks on dtypes, shapes and nulls.

    12k GitHub starsUsed in 1 repo~1.5k tokens
    Data & AnalyticsAuto-check passed
  • Openbb Data Fetcher

    monarchjuno/vibe-investing

    Fetch financial, market, economic, fundamental, news, options, crypto, ETF, index, and macro data through the OpenBB Python interface instead of the OpenBB MCP server.

    299 GitHub stars~2.9k tokensUpdated 5 mo ago
    Data & AnalyticsAuto-check: notes

More from seb1n/awesome-ai-agent-skills

All 91 skills in this repo
  • Agent Red Teaming

    seb1n/awesome-ai-agent-skills

    Plan, execute, document, and retest authorized security assessments of AI agents and multi-agent workflows using safe adversarial cases, synthetic identities, canaries, and evidence-based findings.

    206 GitHub stars~2.8k tokensUpdated 1 mo ago
    Auto-check passed
  • Eu AI Act Readiness

    seb1n/awesome-ai-agent-skills

    Build a preliminary, evidence-based EU AI Act readiness assessment across AI-system inventory, territorial scope, operator roles, prohibited-practice screening, risk classification, transparency…

    206 GitHub stars~3.3k tokensUpdated 1 mo ago
    Auto-check passed
  • Human In The Loop

    seb1n/awesome-ai-agent-skills

    Design and verify auditable human oversight, approval gates, escalation paths, and safe state transitions for AI agent workflows.

    206 GitHub stars~2.5k tokensUpdated 1 mo ago
    Auto-check passed
  • MCP Server Building

    seb1n/awesome-ai-agent-skills

    Design, implement, harden, and verify Model Context Protocol (MCP) servers with precise tool contracts, least-privilege authorization, safe transports, structured errors, and interoperability tests.

    206 GitHub stars~2.5k tokensUpdated 1 mo ago
    Auto-check passed
  • Skill Supply Chain Audit

    seb1n/awesome-ai-agent-skills

    Audit agent skills, plugins, prompts, manifests, scripts, dependencies, and bundled assets for provenance, prompt-injection, permission, execution, exfiltration, persistence, and update risk.

    206 GitHub stars~2.4k tokensUpdated 1 mo ago
    Auto-check passed
  • Spreadsheet Analysis

    seb1n/awesome-ai-agent-skills

    Inspect, profile, clean, reconcile, analyze, visualize, and verify spreadsheet data while preserving formulas, formatting, types, and source files.

    206 GitHub stars~2.5k tokensUpdated 1 mo ago
    Auto-check passed

Questions about Data Cleaning

What does Data Cleaning do?

Clean and preprocess datasets by handling missing values, removing duplicates, correcting types, resolving outliers, and enforcing validation schemas. Data Cleaning is an agent skill from seb1n/awesome-ai-agent-skills. Clean and preprocess datasets by handling missing values, removing duplicates, correcting types, resolving outliers, and enforcing validation schemas.

When should I use Data Cleaning?

Data Cleaning fits situations like: the user requests data cleaning; provides relevant inputs for this workflow.

How do I install Data Cleaning in Claude Code?

Run `npx skills add seb1n/awesome-ai-agent-skills --skill data-cleaning -a claude-code`. Or copy the skill folder (data-and-analytics/data-cleaning in seb1n/awesome-ai-agent-skills) into .claude/skills/data-cleaning in your project. Claude Code loads it when a task matches its description.

How do I install Data Cleaning in Codex?

Run `npx skills add seb1n/awesome-ai-agent-skills --skill data-cleaning -a codex`. Or copy the skill folder (data-and-analytics/data-cleaning in seb1n/awesome-ai-agent-skills) into .agents/skills/data-cleaning in your project. Codex loads it when a task matches its description.

Can I use Data Cleaning in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add seb1n/awesome-ai-agent-skills --skill data-cleaning -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/data-cleaning, .gemini/skills/data-cleaning, .github/skills/data-cleaning and .opencode/skills/data-cleaning in your project.

What does Data Cleaning need to run?

SKILL.md names no scripts, command-line tools or credentials: Data Cleaning is instructions for the agent only. Our summary lists: Python 3.

Does Data Cleaning access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Data Cleaning safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Data Cleaning use?

Data Cleaning is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Data Cleaning use?

About 2k tokens (SKILL.md is roughly 7.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Data Cleaning?

Skills that share tags, products or a category with Data Cleaning: Question2report (refraction-ray/xalpha, 2.7k stars), Dingo Verify (MigoXLab/dingo, 757 stars), Data Validation (platonai/Browser4, 1.2k stars) and Issues Deduplication (JetBrains/ideavim, 10k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Data Cleaning?

seb1n (a GitHub user) maintains it in seb1n/awesome-ai-agent-skills, which has 206 GitHub stars. The repository holds 91 skills in this directory. The repository was last updated on August 9, 2026.

Source: seb1n/awesome-ai-agent-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.