Agent skill

Data Cleaning

by ericrisco in ericrisco/rsc-harness

A skill your agent uses when a raw table is too dirty to trust — nulls, sentinels, duplicate rows, category sprawl, mixed types, bad dates — and you need a re-runnable clean() plus a schema gate…

MITAuto-check passedData & Analytics

Install Data Cleaning

skills CLI
$ npx skills add ericrisco/rsc-harness --skill data-cleaning -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install ericrisco/rsc-harness data-cleaning --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/ericrisco/rsc-harness.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/data-cleaning .claude/skills/data-cleaning && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
data-cleaning
GitHub stars
156
Token cost
~3.6k tokens
SKILL.md length
1,324 words
Files
6 (incl. scripts, references)
Skills in repo
229
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when a raw table is too dirty to trust — nulls, sentinels, duplicate rows, category sprawl, mixed types, bad dates — and you need a re-runnable clean() plus a schema gate…

  • A raw table is too dirty to trust — nulls
  • SKILL.md covers The pipeline shape, Read it right, Profile before you fix and Normalize, plus 7 more sections
  • Runs Shell scripts from its folder; calls pip
  • Category sprawl

What it does

Data Cleaning is an agent skill from ericrisco/rsc-harness. Use when a raw table is too dirty to trust — nulls, sentinels, duplicate rows, category sprawl, mixed types, bad dates — and you need a re-runnable clean() plus a schema gate that fails loud. NOT emitting .xlsx (that is spreadsheet-ops), NOT acquiring rows (that is data-scraper), NOT parsing PDF/HTML into rows (that is structured-extraction).

Its SKILL.md is about 3.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 8 other files, including scripts and reference files (for example `evals/README.md`, `evals/cases.yaml` and `references/normalization-recipes.md`).

It sits in Data & Analytics, covering Excel spreadsheets, Data cleaning and Web scraping. It works with Microsoft Excel, pandas and DuckDB. The repository describes itself as: Your agent invents things because it has no memory, and can't touch your database because it has no arms. rsc is the meta-harness that gives it both, plus the trade to know the… The licence is MIT.

When your agent uses it

  • A raw table is too dirty to trust — nulls
  • Category sprawl
  • Bad dates — and you need a re-runnable clean() plus a schema gate that fails loud

Example prompts

  • “/data-cleaning”

Requirements

  • Python 3
  • A Bash shell

What it can do on your machine

Read from SKILL.md and the folder at commit 92fde8f. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Shell), which the agent can run.

    Shell commands in SKILL.md call:

    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Data Cleaning loads about 3.6k tokens when it runs, and up to ~6.6k if it reads all its reference files. Until then it costs about 90 tokens; SKILL.md has 1,324 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~90
When it runs · the whole SKILL.md, loaded when a task matches
~3.6k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~6.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from ericrisco/rsc-harness at commit 92fde8f, republished under its MIT licence (© ericrisco). 1,324 words, ~3,632 tokens.

Download SKILL.mdSave it as .claude/skills/data-cleaning/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.
name
data-cleaning
description
Use when a raw table is too dirty to trust — nulls, sentinels, duplicate rows, category sprawl, mixed types, bad dates — and you need a re-runnable clean() plus a schema gate that fails loud. NOT emitting .xlsx (that is spreadsheet-ops), NOT acquiring rows (that is data-scraper), NOT parsing PDF/HTML into rows (that is structured-extraction).
tags
data-cleaning, data-quality, pandas, pandera, validation, deduplication, reproducibility, polars
recommends
duckdb, spreadsheet-ops, structured-extraction, data-scraper, analytics, business-intelligence, forecasting, python
origin
risco

Data cleaning — make dirty data trustworthy, and make the cleaning auditable

A clean table is typed + deduped + normalized + validated + reproducible. The deliverable here is never "I opened a notebook and fixed some rows by hand." It is a re-runnable function clean(raw) -> df plus a schema gate that fails loud when next month's file violates the contract. Reproducible means the same input always yields the same output: versions pinned, sorts deterministic, nothing random without a seed. If you can't re-run it tomorrow and get the identical result, you haven't cleaned the data — you've edited a snapshot.

Cleaning starts once you hold tabular rows and ends at a validated table/DataFrame/Parquet. Before that boundary the job is acquisition (data-scraper, structured-extraction); after it, consumption (spreadsheet-ops, analytics, business-intelligence, forecasting). Multi-GB analytical SQL is an engine choice, not a cleaning one — duckdb.

Current stack (verified 2026-06-02): pandas 3.0.x (3.0.0 shipped 2026-01-21) and pandera 0.31.1 (supports pandas ≥ 3) for in-pipeline schema validation; Polars and DuckDB when pandas runs out of RAM. Pin them: pandas==3.0.3, pandera==0.31.1.

The pipeline shape

One canonical order. Each step is positioned for a reason, not by habit.

python
import pandas as pd

def clean(raw_path: str) -> pd.DataFrame:
    df = read_typed(raw_path)     # 1. read with explicit dtypes — never let pandas guess
    df = normalize(df)            # 2. strings/categories/numbers/dates — collapse invisible variance
    df = dedupe(df)               # 3. AFTER normalize+type, so "1"/1 and "US "/"US" actually collapse
    df = handle_missing(df)       # 4. decide per column: drop / impute+flag / leave NA / quarantine
    df = Schema.validate(df, lazy=True)  # 5. the GATE — fail loud, surface every violation at once
    return df
  • Type before dedupe — otherwise "1" (string) and 1 (int) survive as two distinct keys.
  • Normalize before dedupe — "US " and "US" are the same customer; dedupe can't see that until whitespace/case are collapsed.
  • Validate last — it is the gate, not a cleaning step. It asserts the contract holds after all fixes.
  • Write to a NEW artifact — the raw file is read-only; you never overwrite your only source.

Read it right

The single most common reproducibility footgun: pandas' legacy numpy path silently casts an integer column containing one NaN to float64, so your id becomes 1001.0. Control the dtype on read.

python
# BAD — pandas guesses: ids become floats, "N/A" stays a string, "" is sometimes NaN sometimes ""
df = pd.read_csv("raw.csv")

# GOOD — explicit, deterministic, real nullable types
df = pd.read_csv(
    "raw.csv",
    dtype_backend="pyarrow",          # real nullable ints/strings; no silent float-cast
    na_values=["", "N/A", "NA", "null", "-1", "999"],  # YOUR sentinels become real NA
    keep_default_na=True,             # keep pandas' default NA tokens too
    encoding="utf-8",                 # state it; don't let locale decide
)

Two pandas 3.0 facts that read depends on. dtype_backend="pyarrow" only works if pyarrow is actually installed — PDEP-14 deliberately kept a NumPy-object fallback so PyArrow stays recommended, not required — so pip install pyarrow for the faster backed path, or pass dtype_backend="numpy_nullable" when it is absent. And the default str dtype (PyArrow-backed when pyarrow is present, NumPy-object-backed otherwise) uses NaN missing-value semantics like every other default dtype: test for null with pd.isna(), never by comparing against whichever null token happened to appear.

Profile before you fix

Let the numbers drive the plan, not a glance at df.head(). Run this first, every time.

python
def profile(df: pd.DataFrame) -> pd.DataFrame:
    return pd.DataFrame({
        "dtype":     df.dtypes.astype(str),
        "null_pct":  (df.isna().mean() * 100).round(1),
        "n_unique":  df.nunique(dropna=True),       # cardinality — catches category sprawl
        "sample":    df.apply(lambda s: s.dropna().unique()[:3].tolist()),
    })

print(profile(df))
print("rows:", len(df), "exact dupes:", df.duplicated().sum())

A column at 90% null is a drop candidate; one with 400 distinct "countries" needs a mapping table; an "age" with min -1/max 999 has sentinels to map. The profile is your TODO list.

Normalize

Each fix below: Bad → Good, with a one-line why.

Strings — invisible variance (trailing space, mixed case, lookalike unicode) silently breaks joins and dedupe.

python
# BAD: "US ", "us", "us" all look different to a join
# GOOD:
s = df["country"].str.strip().str.casefold().str.normalize("NFKC")

Categories — use a mapping table, never a tower of regex. A dict is auditable and an unmapped value gets quarantined instead of silently passing through.

python
COUNTRY = {"usa": "US", "u.s.": "US", "united states": "US", "u.s.a.": "US", "es": "ES", "españa": "ES"}
key = df["country"].str.strip().str.casefold()
df["country"] = key.map(COUNTRY)            # unmapped -> NA, which the gate below will catch (no silent pass)

Numbers — turn sentinels into NA, then choose a range policy explicitly: clip (cap to bound) when out-of-range is plausibly a recording cap, reject (→ NA / quarantine) when it is impossible.

python
df["age"] = df["age"].mask(df["age"].isin([-1, 999]))   # sentinels -> NA
df["age"] = df["age"].clip(lower=0, upper=120)          # clip policy; or .mask(~df["age"].between(0,120)) to reject

Dates — state the format, coerce, then count the casualties. Never trust dayfirst inference; 03/04/2026 is ambiguous and pandas will pick silently.

python
parsed = pd.to_datetime(df["signup"], format="%Y-%m-%d", utc=True, errors="coerce")
bad = parsed.isna() & df["signup"].notna()
assert bad.sum() == 0, f"{bad.sum()} dates failed the expected format — inspect before proceeding"
df["signup"] = parsed

Copy-paste versions of all of these — category mapping with unmapped→quarantine, a robust date parser, unicode/encoding repair, a sentinel→NA table, numeric clip-vs-reject, plus Polars equivalents — are in references/normalization-recipes.md.

Dedupe

drop_duplicates(keep="first") is meaningless without a defined key and a stable sort — "first" of what order? Define both.

python
key = ["customer_id"]                                   # the BUSINESS key, stated explicitly
df = (df.sort_values(["customer_id", "updated_at"], ascending=[True, False], kind="stable")
        .drop_duplicates(subset=key, keep="first"))     # keep most-recent per customer, deterministically

Near-duplicates ("Acme Inc" vs "Acme, Inc.") are a normalization problem — collapse them in the normalize step first; only then does exact dedupe catch them. Fuzzy matching is a separate, riskier decision — make it visible, never automatic.

Missing values — decide per column

No silent fillna(0): a zero is a value, and treating "unknown" as zero poisons every mean, sum, and model downstream. Pick deliberately.

SituationActionWhy
Column is mostly null (e.g. >70%) and not load-bearingDrop the columnImputing it invents signal that isn't there
A few rows missing a required key (id, date)Drop the row (and log/quarantine)Can't dedupe or join without the key
Numeric gap you must fill for a modelImpute and add a _was_missing flagThe model can learn "was missing"; you keep the audit trail
Genuinely optional fieldLeave NANA is information; don't fabricate a value
Value is present but invalid (unmapped category, bad date)Quarantine the rowDon't drop silently and don't let it pass the gate
python
df["income_was_missing"] = df["income"].isna()
df["income"] = df["income"].fillna(df["income"].median())   # impute + flag, never bare fillna(0)
Show full SKILL.md (598 more words)Show less

Validate — the gate

This is where cleaning becomes trustworthy. Declare the contract as a pandera DataFrameModel, validate output (and input expectations where they exist), and split valid rows from failures instead of crashing — the failures become your quarantine.

python
import pandera.pandas as pa
from pandera.typing import Series

class CustomerSchema(pa.DataFrameModel):
    customer_id: Series[int]   = pa.Field(unique=True, ge=1)
    country:     Series[str]   = pa.Field(isin=["US", "ES", "FR"])     # only mapped categories survive
    age:         Series[float] = pa.Field(ge=0, le=120, nullable=True)
    signup:      Series[pa.DateTime] = pa.Field(nullable=False)

    class Config:
        strict = True       # reject unexpected columns
        coerce = True       # coerce to declared dtype, fail loud if impossible

# lazy=True collects EVERY violation at once instead of dying on the first
try:
    valid = CustomerSchema.validate(df, lazy=True)
except pa.errors.SchemaErrors as e:
    failures = e.failure_cases          # dataframe of exactly which rows/checks failed
    failures.to_parquet("quarantine.parquet")   # keep, don't drop — someone investigates these
    valid = df.drop(index=e.failure_cases["index"].dropna().unique())  # proceed with the clean subset

coerce=True fixes types the contract expects; nullable states which columns may hold NA; field Checks (ge, le, isin, unique) are the allowed-value rules. strict catches columns that shouldn't be there. Together they are the data contract in code. Log the row-count diff on every run — in, out, coerced, quarantined — so what the pipeline changed is an auditable record, not an assumption.

When to escalate beyond pandera: reach for GX Core 1.0 (Great Expectations' rebranded OSS — Data Context → Data Source → Expectation Suite → Validation Definition → Checkpoint) when you need a shared data-quality platform across many datasets and teams with a results store and docs. Use dbt model contracts (enforced at build) plus dbt tests (post-materialization) when the cleaning lives in a SQL warehouse, not Python. The full DataFrameModel (custom @pa.check, lazy SchemaErrors report, valid/quarantine split helper), the GX checkpoint sketch, the dbt model-contract + data_tests YAML, and the "which validator" chooser are in references/validation-patterns.md.

Scale — when pandas hurts

Heuristic: pandas is fine while the data fits comfortably in RAM (roughly ≤ 1–2 GB working set). Beyond that, or when a groupby/join dominates the runtime, switch the mechanics (not the principles):

  • Polars for clean-at-scale: pl.scan_csv(...) (lazy, parallel, Rust), then .unique(), .drop_nulls(), .fill_null(...), .str.* — the same profile→normalize→dedupe→validate shape, faster. pandera validates Polars frames too, and the recipes reference has the Polars equivalent of every fix above.
  • DuckDB when the bottleneck is analytical SQL over multi-GB files — point heavy joins/aggregations there: duckdb. It is an engine choice; correctness/normalization is still this skill's job.

Anti-patterns

Anti-patternWhy it breaks
"fillna(0) to get rid of the nulls"Zero is a value; it distorts every mean/sum/model. Impute deliberately and add a _was_missing flag.
"drop_duplicates() — done"No subset, no sort → which row survives is nondeterministic. Define the key, sort_values(kind="stable"), set keep.
"pd.read_csv(path) and start cleaning"pandas guesses: ids become floats, dates become strings. Pass dtype_backend + na_values.
"I fixed the rows in a notebook cell"Not reproducible — next month's file gets nothing. Wrap it in clean(raw) -> df.
"Drop the rows that look wrong"Silent data loss with no audit trail. Quarantine to a file; someone investigates.
"A few regexes will normalize the countries"Unmaintainable and silent on new values. Use a mapping dict; unmapped → NA → caught by the gate.
"pd.to_datetime figures out the format"Ambiguous dates parse silently wrong. State format=, errors="coerce", then assert the NaT count.
"Validation passed, so we're good"A gate that never fails is a no-op. Feed it a known-bad row and confirm it rejects.
"It's slow, rewrite everything in Polars"Switch the engine, not the discipline — profile→normalize→dedupe→validate still applies.

Verify

scripts/verify.sh runs from anywhere, no network. It does static structure checks on this skill (frontmatter keys, references present) always, and — when pandas + pandera are installed — extracts the documented pattern, feeds it one clearly-good row and one clearly-bad row, and asserts the good row PASSES validation while the bad row is FLAGGED/quarantined, proving the gate is not a no-op. Without pandas/pandera it prints SKIP for the runtime check and still passes the static checks.

Project grounding (02-DOCS + CLAUDE.md)

In a project with a 02-DOCS/ layer (the harness wiki), record this dataset's cleaning decisions — the schema/contract, the category mapping tables, the dedupe key, the quarantine location, version pins — in 02-DOCS/wiki/data/<dataset>.md, link it from the root CLAUDE.md ## Knowledge map, and read it first on every re-run so the contract stays consistent. No 02-DOCS/? Skip silently. Conventions are recorded, never gated.

© ericrisco, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 5 other files (scripts, references) in skills/data-cleaning of ericrisco/rsc-harness.

  • SKILL.md
  • evals/README.md
  • evals/cases.yaml
  • references/normalization-recipes.md
  • references/validation-patterns.md
  • scripts/verify.sh

Open the folder on GitHubat commit 92fde8f

Compare with similar skills

Data Cleaning next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Data Cleaning compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Data Cleaning this skillericrisco/rsc-harness156—~3.6kAutomated safety check: PassMIT
CSV and Excel MergerOneWave-AI/claude-skills322—~1.6kAutomated safety check: PassMIT
Sn Da Excel WorkflowMichaelYang-lyx/AIDABench1111 repos~2.5kAutomated safety check: PassNone
Invalid Data CleaningMichaelYang-lyx/AIDABench1111 repos~410Automated safety check: PassNone
XLSX Spreadsheet ToolkitXiaomiMiMo/MiMo-Code14k—~2.9kAutomated safety check: PassApache-2.0
Verified Data Analysis with pandaspipeshub-ai/pipeshub-ai3.8k—~1.2kAutomated safety check: PassApache-2.0

Similar skills

  • CSV and Excel Merger

    OneWave-AI/claude-skills

    Combines CSV, TSV and Excel files into one verified table with pandas, by stacking or joining, mapping columns, normalizing keys and removing duplicates.

    322 GitHub stars~1.6k tokensUpdated 5 days ago
    Documents & OfficeAuto-check passed
  • Sn Da Excel Workflow

    MichaelYang-lyx/AIDABench

    Excel 数据分析多步编排器。覆盖:(1) 读取多 Sheet Excel 文件并统计行数,(2) 大文件检测(≥10k 行自动 Parquet 优化),(3) 数据清洗(缺失值、文本标准化、无效字符),(4) 条件筛选与分类提取,(5) 跨 Sheet 统计聚合,(6) 导出 Excel/CSV 并提供下载链接。覆盖从数据读取到报告生成全流程,按步骤编排 capability 子…

    111 GitHub starsUsed in 1 repo~2.5k tokens
    Documents & OfficeAuto-check passed
  • Invalid Data Cleaning

    MichaelYang-lyx/AIDABench

    用于大规模Excel数据的预处理,通过统计总行数判断是否转换为Parquet格式以提升读写效率,并使用正则表达式清洗指定文本列(如仅保留中文字符),最后导出清洗后的文件并提供下载链接。

    111 GitHub starsUsed in 1 repo~410 tokens
    Data & AnalyticsAuto-check passed
  • XLSX Spreadsheet Toolkit

    XiaomiMiMo/MiMo-Code

    Builds, edits, cleans, recalculates and reads Excel workbooks and CSV files with openpyxl and pandas, plus LibreOffice for recalculation and PDF export.

    14k GitHub stars~2.9k tokensUpdated 4 days ago
    Documents & OfficeAuto-check passed
  • Verified Data Analysis with pandas

    pipeshub-ai/pipeshub-ai

    Loads, cleans, aggregates and joins tabular data with pandas under a verification rule: every number reported must be one that the code actually printed.

    3.8k GitHub stars~1.2k tokensUpdated today
    Data & AnalyticsAuto-check passed
  • Exploring Data

    oaustegard/claude-skills

    Exploratory data analysis. An agent skill from oaustegard/claude-skills.

    150 GitHub stars~1.7k tokensUpdated 5 days ago
    Data & AnalyticsAuto-check passed

More from ericrisco/rsc-harness

All 229 skills in this repo
  • Ab Testing

    ericrisco/rsc-harness

    A skill your agent uses when designing or analyzing a controlled experiment — falsifiable hypothesis, sample size from an MDE, reading significance/CI/power, CUPED, or rescuing tests that won't go…

    156 GitHub stars~2.4k tokensUpdated yesterday
    Auto-check passed
  • Accessibility

    ericrisco/rsc-harness

    A skill your agent uses when making a web UI conform to WCAG 2.2 Level AA — axe-core or Lighthouse a11y violations, keyboard operability, focus management, ARIA roles/names/live regions, contrast…

    156 GitHub stars~3.4k tokensUpdated yesterday
    Auto-check passed
  • Ads

    ericrisco/rsc-harness

    A skill your agent uses when running or fixing paid acquisition on Google or Meta — campaign structure (Performance Max, Demand Gen, Search, Advantage+), platform-fit creative, budget/scaling rules…

    156 GitHub stars~2.2k tokensUpdated yesterday
    Auto-check passed
  • Agent Eval

    ericrisco/rsc-harness

    A skill your agent uses when measuring whether an LLM or agent system actually got better and gating merges on it: golden sets, fixing an inflated LLM-as-judge, scoring RAG (faithfulness, contextual…

    156 GitHub stars~3.2k tokensUpdated yesterday
    Auto-check passed
  • AI Media

    ericrisco/rsc-harness

    A skill your agent uses when a creative goal must become a finished media file: pick and order generative-media models per modality — AI voiceover, image-to-video clips, score — then glue them with…

    156 GitHub stars~3.3k tokensUpdated yesterday
    Auto-check passed
  • Analytics

    ericrisco/rsc-harness

    A skill your agent uses when instrumenting product or web analytics — GA4/PostHog SDK wiring, event taxonomy, funnels, double-counted events, consent gating, PII scrubbing.

    156 GitHub stars~2.8k tokensUpdated yesterday
    Auto-check passed

Questions about Data Cleaning

What does Data Cleaning do?

A skill your agent uses when a raw table is too dirty to trust — nulls, sentinels, duplicate rows, category sprawl, mixed types, bad dates — and you need a re-runnable clean() plus a schema gate…. Data Cleaning is an agent skill from ericrisco/rsc-harness. Use when a raw table is too dirty to trust — nulls, sentinels, duplicate rows, category sprawl, mixed types, bad dates — and you need a re-runnable clean() plus a schema gate that fails loud.

When should I use Data Cleaning?

Data Cleaning fits situations like: A raw table is too dirty to trust — nulls; category sprawl; bad dates — and you need a re-runnable clean() plus a schema gate that fails loud.

How do I install Data Cleaning in Claude Code?

Run `npx skills add ericrisco/rsc-harness --skill data-cleaning -a claude-code`. Or copy the skill folder (skills/data-cleaning in ericrisco/rsc-harness) into .claude/skills/data-cleaning in your project. Claude Code loads it when a task matches its description.

How do I install Data Cleaning in Codex?

Run `npx skills add ericrisco/rsc-harness --skill data-cleaning -a codex`. Or copy the skill folder (skills/data-cleaning in ericrisco/rsc-harness) into .agents/skills/data-cleaning in your project. Codex loads it when a task matches its description.

Can I use Data Cleaning in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ericrisco/rsc-harness --skill data-cleaning -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/data-cleaning, .gemini/skills/data-cleaning, .github/skills/data-cleaning and .opencode/skills/data-cleaning in your project.

What does Data Cleaning need to run?

Going by SKILL.md and its folder, Data Cleaning needs a shell for the scripts in its folder and the command-line tools its instructions call (pip). Our summary lists: Python 3; A Bash shell.

Does Data Cleaning access the network?

SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Data Cleaning safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Data Cleaning use?

Data Cleaning is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Data Cleaning use?

About 3.6k tokens (SKILL.md is roughly 15k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3k tokens, read only when the agent opens those files.

What are the alternatives to Data Cleaning?

Skills that share tags, products or a category with Data Cleaning: CSV and Excel Merger (OneWave-AI/claude-skills, 322 stars), Sn Da Excel Workflow (MichaelYang-lyx/AIDABench, 111 stars), Invalid Data Cleaning (MichaelYang-lyx/AIDABench, 111 stars) and XLSX Spreadsheet Toolkit (XiaomiMiMo/MiMo-Code, 14k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Data Cleaning?

ericrisco (a GitHub user) maintains it in ericrisco/rsc-harness, which has 156 GitHub stars. The repository holds 229 skills in this directory. The repository was last updated on October 6, 2026.

Source: ericrisco/rsc-harness on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.