CSV and Excel Merger
OneWave-AI/claude-skills
Combines CSV, TSV and Excel files into one verified table with pandas, by stacking or joining, mapping columns, normalizing keys and removing duplicates.
A skill your agent uses when a raw table is too dirty to trust — nulls, sentinels, duplicate rows, category sprawl, mixed types, bad dates — and you need a re-runnable clean() plus a schema gate…
$ npx skills add ericrisco/rsc-harness --skill data-cleaning -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install ericrisco/rsc-harness data-cleaning --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/ericrisco/rsc-harness.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/data-cleaning .claude/skills/data-cleaning && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "data-cleaning" agent skill from https://github.com/ericrisco/rsc-harness/tree/main/skills/data-cleaning into .claude/skills/data-cleaning/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "data-cleaning", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/ericrisco/rsc-harness/tree/main/skills/data-cleaningType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add ericrisco/rsc-harness --skill data-cleaning -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install ericrisco/rsc-harness data-cleaning --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ericrisco/rsc-harness.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/data-cleaning .agents/skills/data-cleaning && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "data-cleaning" agent skill from https://github.com/ericrisco/rsc-harness/tree/main/skills/data-cleaning into .agents/skills/data-cleaning/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "data-cleaning", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add ericrisco/rsc-harness --skill data-cleaning -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install ericrisco/rsc-harness data-cleaning --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ericrisco/rsc-harness.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/data-cleaning .cursor/skills/data-cleaning && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "data-cleaning" agent skill from https://github.com/ericrisco/rsc-harness/tree/main/skills/data-cleaning into .cursor/skills/data-cleaning/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "data-cleaning", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/ericrisco/rsc-harness.git --path skills/data-cleaning--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add ericrisco/rsc-harness --skill data-cleaning -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install ericrisco/rsc-harness data-cleaning --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ericrisco/rsc-harness.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/data-cleaning .gemini/skills/data-cleaning && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "data-cleaning" agent skill from https://github.com/ericrisco/rsc-harness/tree/main/skills/data-cleaning into .gemini/skills/data-cleaning/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "data-cleaning", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install ericrisco/rsc-harness data-cleaningInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add ericrisco/rsc-harness --skill data-cleaning -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/ericrisco/rsc-harness.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/data-cleaning .github/skills/data-cleaning && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "data-cleaning" agent skill from https://github.com/ericrisco/rsc-harness/tree/main/skills/data-cleaning into .github/skills/data-cleaning/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "data-cleaning", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add ericrisco/rsc-harness --skill data-cleaning -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install ericrisco/rsc-harness data-cleaning --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ericrisco/rsc-harness.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/data-cleaning .opencode/skills/data-cleaning && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "data-cleaning" agent skill from https://github.com/ericrisco/rsc-harness/tree/main/skills/data-cleaning into .opencode/skills/data-cleaning/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "data-cleaning", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
data-cleaningA skill your agent uses when a raw table is too dirty to trust — nulls, sentinels, duplicate rows, category sprawl, mixed types, bad dates — and you need a re-runnable clean() plus a schema gate…
Data Cleaning is an agent skill from ericrisco/rsc-harness. Use when a raw table is too dirty to trust — nulls, sentinels, duplicate rows, category sprawl, mixed types, bad dates — and you need a re-runnable clean() plus a schema gate that fails loud. NOT emitting .xlsx (that is spreadsheet-ops), NOT acquiring rows (that is data-scraper), NOT parsing PDF/HTML into rows (that is structured-extraction).
Its SKILL.md is about 3.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 8 other files, including scripts and reference files (for example `evals/README.md`, `evals/cases.yaml` and `references/normalization-recipes.md`).
It sits in Data & Analytics, covering Excel spreadsheets, Data cleaning and Web scraping. It works with Microsoft Excel, pandas and DuckDB. The repository describes itself as: Your agent invents things because it has no memory, and can't touch your database because it has no arms. rsc is the meta-harness that gives it both, plus the trade to know the… The licence is MIT.
Read from SKILL.md and the folder at commit 92fde8f. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 1 file in scripts/ (Shell), which the agent can run.
Shell commands in SKILL.md call:
pipFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Data Cleaning loads about 3.6k tokens when it runs, and up to ~6.6k if it reads all its reference files. Until then it costs about 90 tokens; SKILL.md has 1,324 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from ericrisco/rsc-harness at commit 92fde8f, republished under its MIT licence (© ericrisco). 1,324 words, ~3,632 tokens.
.claude/skills/data-cleaning/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.A clean table is typed + deduped + normalized + validated + reproducible. The deliverable here is
never "I opened a notebook and fixed some rows by hand." It is a re-runnable function clean(raw) -> df
plus a schema gate that fails loud when next month's file violates the contract. Reproducible means
the same input always yields the same output: versions pinned, sorts deterministic, nothing random without
a seed. If you can't re-run it tomorrow and get the identical result, you haven't cleaned the data — you've
edited a snapshot.
Cleaning starts once you hold tabular rows and ends at a validated table/DataFrame/Parquet. Before
that boundary the job is acquisition (data-scraper,
structured-extraction); after it, consumption
(spreadsheet-ops, analytics,
business-intelligence,
forecasting). Multi-GB analytical SQL is an engine choice, not a cleaning one
— duckdb.
Current stack (verified 2026-06-02): pandas 3.0.x (3.0.0 shipped 2026-01-21) and pandera 0.31.1
(supports pandas ≥ 3) for in-pipeline schema validation; Polars and DuckDB when pandas runs out of
RAM. Pin them: pandas==3.0.3, pandera==0.31.1.
One canonical order. Each step is positioned for a reason, not by habit.
import pandas as pd
def clean(raw_path: str) -> pd.DataFrame:
df = read_typed(raw_path) # 1. read with explicit dtypes — never let pandas guess
df = normalize(df) # 2. strings/categories/numbers/dates — collapse invisible variance
df = dedupe(df) # 3. AFTER normalize+type, so "1"/1 and "US "/"US" actually collapse
df = handle_missing(df) # 4. decide per column: drop / impute+flag / leave NA / quarantine
df = Schema.validate(df, lazy=True) # 5. the GATE — fail loud, surface every violation at once
return df"1" (string) and 1 (int) survive as two distinct keys."US " and "US" are the same customer; dedupe can't see that until
whitespace/case are collapsed.The single most common reproducibility footgun: pandas' legacy numpy path silently casts an integer column
containing one NaN to float64, so your id becomes 1001.0. Control the dtype on read.
# BAD — pandas guesses: ids become floats, "N/A" stays a string, "" is sometimes NaN sometimes ""
df = pd.read_csv("raw.csv")
# GOOD — explicit, deterministic, real nullable types
df = pd.read_csv(
"raw.csv",
dtype_backend="pyarrow", # real nullable ints/strings; no silent float-cast
na_values=["", "N/A", "NA", "null", "-1", "999"], # YOUR sentinels become real NA
keep_default_na=True, # keep pandas' default NA tokens too
encoding="utf-8", # state it; don't let locale decide
)Two pandas 3.0 facts that read depends on. dtype_backend="pyarrow" only works if pyarrow is actually
installed — PDEP-14 deliberately kept a NumPy-object fallback so PyArrow stays recommended, not required
— so pip install pyarrow for the faster backed path, or pass dtype_backend="numpy_nullable" when it is
absent. And the default str dtype (PyArrow-backed when pyarrow is present, NumPy-object-backed otherwise)
uses NaN missing-value semantics like every other default dtype: test for null with pd.isna(), never
by comparing against whichever null token happened to appear.
Let the numbers drive the plan, not a glance at df.head(). Run this first, every time.
def profile(df: pd.DataFrame) -> pd.DataFrame:
return pd.DataFrame({
"dtype": df.dtypes.astype(str),
"null_pct": (df.isna().mean() * 100).round(1),
"n_unique": df.nunique(dropna=True), # cardinality — catches category sprawl
"sample": df.apply(lambda s: s.dropna().unique()[:3].tolist()),
})
print(profile(df))
print("rows:", len(df), "exact dupes:", df.duplicated().sum())A column at 90% null is a drop candidate; one with 400 distinct "countries" needs a mapping table; an
"age" with min -1/max 999 has sentinels to map. The profile is your TODO list.
Each fix below: Bad → Good, with a one-line why.
Strings — invisible variance (trailing space, mixed case, lookalike unicode) silently breaks joins and dedupe.
# BAD: "US ", "us", "us" all look different to a join
# GOOD:
s = df["country"].str.strip().str.casefold().str.normalize("NFKC")Categories — use a mapping table, never a tower of regex. A dict is auditable and an unmapped value gets quarantined instead of silently passing through.
COUNTRY = {"usa": "US", "u.s.": "US", "united states": "US", "u.s.a.": "US", "es": "ES", "españa": "ES"}
key = df["country"].str.strip().str.casefold()
df["country"] = key.map(COUNTRY) # unmapped -> NA, which the gate below will catch (no silent pass)Numbers — turn sentinels into NA, then choose a range policy explicitly: clip (cap to bound) when out-of-range is plausibly a recording cap, reject (→ NA / quarantine) when it is impossible.
df["age"] = df["age"].mask(df["age"].isin([-1, 999])) # sentinels -> NA
df["age"] = df["age"].clip(lower=0, upper=120) # clip policy; or .mask(~df["age"].between(0,120)) to rejectDates — state the format, coerce, then count the casualties. Never trust dayfirst inference;
03/04/2026 is ambiguous and pandas will pick silently.
parsed = pd.to_datetime(df["signup"], format="%Y-%m-%d", utc=True, errors="coerce")
bad = parsed.isna() & df["signup"].notna()
assert bad.sum() == 0, f"{bad.sum()} dates failed the expected format — inspect before proceeding"
df["signup"] = parsedCopy-paste versions of all of these — category mapping with unmapped→quarantine, a robust date parser, unicode/encoding repair, a sentinel→NA table, numeric clip-vs-reject, plus Polars equivalents — are in references/normalization-recipes.md.
drop_duplicates(keep="first") is meaningless without a defined key and a stable sort — "first" of what
order? Define both.
key = ["customer_id"] # the BUSINESS key, stated explicitly
df = (df.sort_values(["customer_id", "updated_at"], ascending=[True, False], kind="stable")
.drop_duplicates(subset=key, keep="first")) # keep most-recent per customer, deterministicallyNear-duplicates ("Acme Inc" vs "Acme, Inc.") are a normalization problem — collapse them in the
normalize step first; only then does exact dedupe catch them. Fuzzy matching is a separate, riskier
decision — make it visible, never automatic.
No silent fillna(0): a zero is a value, and treating "unknown" as zero poisons every mean, sum, and model
downstream. Pick deliberately.
| Situation | Action | Why |
|---|---|---|
| Column is mostly null (e.g. >70%) and not load-bearing | Drop the column | Imputing it invents signal that isn't there |
| A few rows missing a required key (id, date) | Drop the row (and log/quarantine) | Can't dedupe or join without the key |
| Numeric gap you must fill for a model | Impute and add a _was_missing flag | The model can learn "was missing"; you keep the audit trail |
| Genuinely optional field | Leave NA | NA is information; don't fabricate a value |
| Value is present but invalid (unmapped category, bad date) | Quarantine the row | Don't drop silently and don't let it pass the gate |
df["income_was_missing"] = df["income"].isna()
df["income"] = df["income"].fillna(df["income"].median()) # impute + flag, never bare fillna(0)This is where cleaning becomes trustworthy. Declare the contract as a pandera DataFrameModel, validate
output (and input expectations where they exist), and split valid rows from failures instead of
crashing — the failures become your quarantine.
import pandera.pandas as pa
from pandera.typing import Series
class CustomerSchema(pa.DataFrameModel):
customer_id: Series[int] = pa.Field(unique=True, ge=1)
country: Series[str] = pa.Field(isin=["US", "ES", "FR"]) # only mapped categories survive
age: Series[float] = pa.Field(ge=0, le=120, nullable=True)
signup: Series[pa.DateTime] = pa.Field(nullable=False)
class Config:
strict = True # reject unexpected columns
coerce = True # coerce to declared dtype, fail loud if impossible
# lazy=True collects EVERY violation at once instead of dying on the first
try:
valid = CustomerSchema.validate(df, lazy=True)
except pa.errors.SchemaErrors as e:
failures = e.failure_cases # dataframe of exactly which rows/checks failed
failures.to_parquet("quarantine.parquet") # keep, don't drop — someone investigates these
valid = df.drop(index=e.failure_cases["index"].dropna().unique()) # proceed with the clean subsetcoerce=True fixes types the contract expects; nullable states which columns may hold NA; field
Checks (ge, le, isin, unique) are the allowed-value rules. strict catches columns that
shouldn't be there. Together they are the data contract in code. Log the row-count diff on every run —
in, out, coerced, quarantined — so what the pipeline changed is an auditable record, not an assumption.
When to escalate beyond pandera: reach for GX Core 1.0 (Great Expectations' rebranded OSS — Data
Context → Data Source → Expectation Suite → Validation Definition → Checkpoint) when you need a shared
data-quality platform across many datasets and teams with a results store and docs. Use dbt model
contracts (enforced at build) plus dbt tests (post-materialization) when the cleaning lives in a
SQL warehouse, not Python. The full DataFrameModel (custom @pa.check, lazy SchemaErrors report,
valid/quarantine split helper), the GX checkpoint sketch, the dbt model-contract + data_tests YAML, and
the "which validator" chooser are in references/validation-patterns.md.
Heuristic: pandas is fine while the data fits comfortably in RAM (roughly ≤ 1–2 GB working set). Beyond that, or when a groupby/join dominates the runtime, switch the mechanics (not the principles):
pl.scan_csv(...) (lazy, parallel, Rust), then .unique(),
.drop_nulls(), .fill_null(...), .str.* — the same profile→normalize→dedupe→validate shape, faster.
pandera validates Polars frames too, and the
recipes reference has the Polars equivalent of every fix above.duckdb. It is an engine choice; correctness/normalization is still this
skill's job.| Anti-pattern | Why it breaks |
|---|---|
"fillna(0) to get rid of the nulls" | Zero is a value; it distorts every mean/sum/model. Impute deliberately and add a _was_missing flag. |
"drop_duplicates() — done" | No subset, no sort → which row survives is nondeterministic. Define the key, sort_values(kind="stable"), set keep. |
"pd.read_csv(path) and start cleaning" | pandas guesses: ids become floats, dates become strings. Pass dtype_backend + na_values. |
| "I fixed the rows in a notebook cell" | Not reproducible — next month's file gets nothing. Wrap it in clean(raw) -> df. |
| "Drop the rows that look wrong" | Silent data loss with no audit trail. Quarantine to a file; someone investigates. |
| "A few regexes will normalize the countries" | Unmaintainable and silent on new values. Use a mapping dict; unmapped → NA → caught by the gate. |
"pd.to_datetime figures out the format" | Ambiguous dates parse silently wrong. State format=, errors="coerce", then assert the NaT count. |
| "Validation passed, so we're good" | A gate that never fails is a no-op. Feed it a known-bad row and confirm it rejects. |
| "It's slow, rewrite everything in Polars" | Switch the engine, not the discipline — profile→normalize→dedupe→validate still applies. |
scripts/verify.sh runs from anywhere, no network. It does static structure checks on this skill
(frontmatter keys, references present) always, and — when pandas + pandera are installed — extracts the
documented pattern, feeds it one clearly-good row and one clearly-bad row, and asserts the good row PASSES
validation while the bad row is FLAGGED/quarantined, proving the gate is not a no-op. Without
pandas/pandera it prints SKIP for the runtime check and still passes the static checks.
In a project with a 02-DOCS/ layer (the harness wiki), record this dataset's
cleaning decisions — the schema/contract, the category mapping tables, the dedupe key, the quarantine
location, version pins — in 02-DOCS/wiki/data/<dataset>.md, link it from the root CLAUDE.md
## Knowledge map, and read it first on every re-run so the contract stays consistent. No 02-DOCS/? Skip
silently. Conventions are recorded, never gated.
© ericrisco, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 5 other files (scripts, references) in skills/data-cleaning of ericrisco/rsc-harness.
Open the folder on GitHubat commit 92fde8f
Data Cleaning next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Data Cleaning this skillericrisco/rsc-harness | 156 | — | ~3.6k | Automated safety check: Pass | MIT | |
| CSV and Excel MergerOneWave-AI/claude-skills | 322 | — | ~1.6k | Automated safety check: Pass | MIT | |
| Sn Da Excel WorkflowMichaelYang-lyx/AIDABench | 111 | 1 repos | ~2.5k | Automated safety check: Pass | None | |
| Invalid Data CleaningMichaelYang-lyx/AIDABench | 111 | 1 repos | ~410 | Automated safety check: Pass | None | |
| XLSX Spreadsheet ToolkitXiaomiMiMo/MiMo-Code | 14k | — | ~2.9k | Automated safety check: Pass | Apache-2.0 | |
| Verified Data Analysis with pandaspipeshub-ai/pipeshub-ai | 3.8k | — | ~1.2k | Automated safety check: Pass | Apache-2.0 |
OneWave-AI/claude-skills
Combines CSV, TSV and Excel files into one verified table with pandas, by stacking or joining, mapping columns, normalizing keys and removing duplicates.
MichaelYang-lyx/AIDABench
Excel 数据分析多步编排器。覆盖:(1) 读取多 Sheet Excel 文件并统计行数,(2) 大文件检测(≥10k 行自动 Parquet 优化),(3) 数据清洗(缺失值、文本标准化、无效字符),(4) 条件筛选与分类提取,(5) 跨 Sheet 统计聚合,(6) 导出 Excel/CSV 并提供下载链接。覆盖从数据读取到报告生成全流程,按步骤编排 capability 子…
MichaelYang-lyx/AIDABench
用于大规模Excel数据的预处理,通过统计总行数判断是否转换为Parquet格式以提升读写效率,并使用正则表达式清洗指定文本列(如仅保留中文字符),最后导出清洗后的文件并提供下载链接。
XiaomiMiMo/MiMo-Code
Builds, edits, cleans, recalculates and reads Excel workbooks and CSV files with openpyxl and pandas, plus LibreOffice for recalculation and PDF export.
pipeshub-ai/pipeshub-ai
Loads, cleans, aggregates and joins tabular data with pandas under a verification rule: every number reported must be one that the code actually printed.
oaustegard/claude-skills
Exploratory data analysis. An agent skill from oaustegard/claude-skills.
ericrisco/rsc-harness
A skill your agent uses when designing or analyzing a controlled experiment — falsifiable hypothesis, sample size from an MDE, reading significance/CI/power, CUPED, or rescuing tests that won't go…
ericrisco/rsc-harness
A skill your agent uses when making a web UI conform to WCAG 2.2 Level AA — axe-core or Lighthouse a11y violations, keyboard operability, focus management, ARIA roles/names/live regions, contrast…
ericrisco/rsc-harness
A skill your agent uses when running or fixing paid acquisition on Google or Meta — campaign structure (Performance Max, Demand Gen, Search, Advantage+), platform-fit creative, budget/scaling rules…
ericrisco/rsc-harness
A skill your agent uses when measuring whether an LLM or agent system actually got better and gating merges on it: golden sets, fixing an inflated LLM-as-judge, scoring RAG (faithfulness, contextual…
ericrisco/rsc-harness
A skill your agent uses when a creative goal must become a finished media file: pick and order generative-media models per modality — AI voiceover, image-to-video clips, score — then glue them with…
ericrisco/rsc-harness
A skill your agent uses when instrumenting product or web analytics — GA4/PostHog SDK wiring, event taxonomy, funnels, double-counted events, consent gating, PII scrubbing.
Works with
Categories
A skill your agent uses when a raw table is too dirty to trust — nulls, sentinels, duplicate rows, category sprawl, mixed types, bad dates — and you need a re-runnable clean() plus a schema gate…. Data Cleaning is an agent skill from ericrisco/rsc-harness. Use when a raw table is too dirty to trust — nulls, sentinels, duplicate rows, category sprawl, mixed types, bad dates — and you need a re-runnable clean() plus a schema gate that fails loud.
Data Cleaning fits situations like: A raw table is too dirty to trust — nulls; category sprawl; bad dates — and you need a re-runnable clean() plus a schema gate that fails loud.
Run `npx skills add ericrisco/rsc-harness --skill data-cleaning -a claude-code`. Or copy the skill folder (skills/data-cleaning in ericrisco/rsc-harness) into .claude/skills/data-cleaning in your project. Claude Code loads it when a task matches its description.
Run `npx skills add ericrisco/rsc-harness --skill data-cleaning -a codex`. Or copy the skill folder (skills/data-cleaning in ericrisco/rsc-harness) into .agents/skills/data-cleaning in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ericrisco/rsc-harness --skill data-cleaning -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/data-cleaning, .gemini/skills/data-cleaning, .github/skills/data-cleaning and .opencode/skills/data-cleaning in your project.
Going by SKILL.md and its folder, Data Cleaning needs a shell for the scripts in its folder and the command-line tools its instructions call (pip). Our summary lists: Python 3; A Bash shell.
SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Data Cleaning is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 3.6k tokens (SKILL.md is roughly 15k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Data Cleaning: CSV and Excel Merger (OneWave-AI/claude-skills, 322 stars), Sn Da Excel Workflow (MichaelYang-lyx/AIDABench, 111 stars), Invalid Data Cleaning (MichaelYang-lyx/AIDABench, 111 stars) and XLSX Spreadsheet Toolkit (XiaomiMiMo/MiMo-Code, 14k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
ericrisco (a GitHub user) maintains it in ericrisco/rsc-harness, which has 156 GitHub stars. The repository holds 229 skills in this directory. The repository was last updated on October 6, 2026.
Source: ericrisco/rsc-harness on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.