Agent skill

Data Cleaning

by magnus919 in magnus919/agent-skills

Clean, profile, validate, reshape, and document messy tabular, text, JSON, and relational data through an evidence-first, reproducible workflow.

MITAuto-check passedData & Analytics

Install Data Cleaning

skills CLI
$ npx skills add magnus919/agent-skills --skill data-cleaning -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install magnus919/agent-skills data-cleaning --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/magnus919/agent-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/data-cleaning .claude/skills/data-cleaning && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
data-cleaning
GitHub stars
116
Token cost
~2.1k tokens
SKILL.md length
888 words
Files
23 (incl. scripts, references)
Skills in repo
130
Repo updated
First seen
Licence
MIT

At a glance

Clean, profile, validate, reshape, and document messy tabular, text, JSON, and relational data through an evidence-first, reproducible workflow.

  • Works in 7 steps: Frame: identify the decision, owner,… → Freeze evidence: record source path/URI,… → Profile before changing: inspect… → …
  • Preparing data for analysis
  • SKILL.md covers Route by task, Available Scripts, Default workflow and Non-negotiable controls, plus 4 more sections
  • Runs Python scripts from its folder; calls python3

What it does

Data Cleaning is an agent skill from magnus919/agent-skills. Clean, profile, validate, reshape, and document messy tabular, text, JSON, and relational data through an evidence-first, reproducible workflow. Use when preparing data for analysis, reporting, modeling, ingestion, migration, or matching, including AI-suggested repair, entity-match review, score calibration limits, and reversible repair ledgers. Do not use for statistical modeling, dashboard design, or operating a named data platform; route those tasks to data-scientist, data-engineering, or the relevant tool…

Its SKILL.md is about 2.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 26 other files, including scripts and reference files (for example `README.md`, `evals/evals.json` and `evals/trigger-cases.json`). Compatibility notes: Works with any Agent Skills client. The bundled profiler requires Python 3.9+ and the standard library; ecosystem tools are optional.

It sits in Data & Analytics, covering Data cleaning, Data pipelines and ETL and Performance reviews. It works with Python. The repository describes itself as: Curated collection of AI agent skills for Hermes and other agent frameworks. The licence is MIT.

When your agent uses it

  • Preparing data for analysis
  • Including AI-suggested repair
  • Entity-match review
  • Score calibration limits

Example prompts

  • “/data-cleaning”

Requirements

  • Python 3
  • Compatibility (from SKILL.md): Works with any Agent Skills client. The bundled profiler requires Python 3.9+ and the standard library; ecosystem tools are optional.

Workflow steps

7 steps, taken from the first numbered list in SKILL.md.

  1. Frame: identify the decision, owner, source, privacy constraints, unit of observation, keys, expected grain, time window, and acceptance…
  2. Freeze evidence: record source path/URI, retrieval time, file size/hash where feasible, encoding, delimiter, schema, row/column counts…
  3. Profile before changing: inspect missingness, sentinel values, duplicates, cardinality, type candidates, ranges, invalid dates…
  4. Design decisions: classify each finding as preserve, standardize, repair, impute, quarantine, reject, or escalate. Record rationale, rule…
  5. Transform in layers: prefer deterministic named steps: parse → canonicalize → type/coerce → validate → deduplicate → resolve entities →…
  6. Validate twice: run structural checks before and after transformation. Validate row/grain preservation, key uniqueness, referential…
  7. Review and release: compare before/after metrics, inspect samples of every changed class, obtain domain approval for semantic or lossy…

What it can do on your machine

Read from SKILL.md and the folder at commit c545c2b. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 4 files in scripts/ (Python, from the files we listed), which the agent can run.

    Shell commands in SKILL.md call:

    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Works with any Agent Skills client. The bundled profiler requires Python 3.9+ and the standard library; ecosystem tools are optional.

    From compatibility in the SKILL.md frontmatter.

Context cost

Data Cleaning loads about 2.1k tokens when it runs, and up to ~6k if it reads all its reference files. Until then it costs about 134 tokens; SKILL.md has 888 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~134
When it runs · the whole SKILL.md, loaded when a task matches
~2.1k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from magnus919/agent-skills at commit c545c2b, republished under its MIT licence (© magnus919). 888 words, ~2,066 tokens.

Download SKILL.mdSave it as .claude/skills/data-cleaning/SKILL.md (or your agent's skills folder). This skill also uses 22 other files; get the full folder from GitHub.
name
data-cleaning
description
Clean, profile, validate, reshape, and document messy tabular, text, JSON, and relational data through an evidence-first, reproducible workflow. Use when preparing data for analysis, reporting, modeling, ingestion, migration, or matching, including AI-suggested repair, entity-match review, score calibration limits, and reversible repair ledgers. Do not use for statistical modeling, dashboard design, or operating a named data platform; route those tasks to data-scientist, data-engineering, or the relevant tool skill.
compatibility
Works with any Agent Skills client. The bundled profiler requires Python 3.9+ and the standard library; ecosystem tools are optional.
license
MIT
metadata.domain
data-quality-and-cleaning
metadata.source
primary-docs-plus-orientation-article

Data cleaning

Treat cleaning as a controlled transformation of an observed dataset, not cosmetic editing. Preserve raw input, state the target use and grain, make every lossy decision explicit, and prove that the cleaned output satisfies a contract.

When a model proposes data transformations, load the AI boundary workflow and use the companion record.

Route by task

NeedRead next
End-to-end method, scope, and stopping rulesreferences/methodology.md
Choose a library or platformreferences/tool-selection.md
Missingness, duplicates, types, ranges, categories, dates, joinsreferences/operations.md
Text, identifiers, Unicode, and entity resolutionreferences/text-and-entity.md
Schemas, contracts, validation, drift, scalereferences/validation-and-scale.md
CLI, OpenRefine, monitoring, and interactive remediationreferences/cli-and-interactive-tools.md
Source claims and version-sensitive caveatsreferences/sources.md
Plan, logs, exceptions, contracts, or reportstemplates/cleaning-plan.md, templates/transformation-log.jsonl, templates/exception-register.csv, templates/schema-contract.yml, templates/quality-report.md
Lightweight profile or reconciliationRun python3 scripts/profile_dataset.py --help or python3 scripts/reconcile_dataset.py --help

Available Scripts

ScriptPurposeInvocation
scripts/profile_dataset.pyDependency-free first-pass profiling of a CSV, TSV, or JSONL input without modifying it: missingness, cardinality, type candidates, duplicates, ranges, and value anomalies. Run it at workflow step 3 (Profile before changing) as the evidence-gathering pass before designing any cleaning decision.python3 scripts/profile_dataset.py data.csv --output profile.json
scripts/reconcile_dataset.pyReconciliation between a before and after delimited dataset: row counts, key uniqueness/overlap, and per-column sums (--sum), keyed by --key, writing a machine-readable report. Run it during Validate twice / Review to prove grain preservation and quantify exactly what a transformation changed.python3 scripts/reconcile_dataset.py raw.csv cleaned.csv --key id --sum amount --output reconciliation.json
scripts/test_profile_dataset.pyPytest suite covering the profiler's behavior on representative inputs. Run it after modifying the profiler or when auditing its output; CI discovers it automatically.python3 -m pytest scripts/test_profile_dataset.py
scripts/test_reconcile_dataset.pyPytest suite covering the reconciler's keying, summing, and reporting behavior. Run it after modifying the reconciler or when auditing its output; CI discovers it automatically.python3 -m pytest scripts/test_reconcile_dataset.py

Default workflow

  1. Frame: identify the decision, owner, source, privacy constraints, unit of observation, keys, expected grain, time window, and acceptance threshold. Do not silently infer a business rule from a suspicious value.
  2. Freeze evidence: record source path/URI, retrieval time, file size/hash where feasible, encoding, delimiter, schema, row/column counts, and software versions. Keep raw data read-only and write to a new output.
  3. Profile before changing: inspect missingness, sentinel values, duplicates, cardinality, type candidates, ranges, invalid dates, whitespace/Unicode anomalies, cross-field relationships, and drift. Use the bundled profiler for a dependency-free first pass.
  4. Design decisions: classify each finding as preserve, standardize, repair, impute, quarantine, reject, or escalate. Record rationale, rule, affected rows, confidence, reversibility, and owner.
  5. Transform in layers: prefer deterministic named steps: parse → canonicalize → type/coerce → validate → deduplicate → resolve entities → impute/quarantine → reshape. Keep raw, staged, rejected, and final datasets distinct.
  6. Validate twice: run structural checks before and after transformation. Validate row/grain preservation, key uniqueness, referential integrity, allowed values, units, bounds, null policy, and expected distributions. Tests should identify failing records.
  7. Review and release: compare before/after metrics, inspect samples of every changed class, obtain domain approval for semantic or lossy changes, publish the report and provenance, and make the run reproducible.
Show full SKILL.md (393 more words)Show less

Non-negotiable controls

  • Never overwrite raw data or silently drop rows, columns, categories, outliers, or unmatched entities.
  • Separate invalid, missing, not applicable, not collected, and withheld when the domain distinguishes them.
  • Parse dates and numbers with an explicit locale, timezone, unit, and error policy. Count parse failures; do not silently turn them into nulls.
  • Normalize text conservatively. Retain original and normalized values plus confidence when matching or repairing.
  • Fit imputers, encoders, normalization parameters, and deduplication rules only on the permitted training/reference partition. Avoid leakage across time or evaluation boundaries.
  • Treat profiling as evidence for investigation, not permission to auto-fix. An anomaly can be a real event.
  • Use quarantine for records that cannot be repaired safely. “Clean” means accepted by a stated contract, not “no rows remain.”

Completion gate

A cleaning task is complete only when the output, transformation/decision log, validation evidence, provenance, and unresolved issues exist; raw data remains intact; acceptance checks pass; and a reviewer can reproduce or audit the result. If semantic ambiguity remains, stop at quarantine or escalation rather than inventing a value.

When not to use

Do not use this skill for inferential statistics or model selection, which belong to data-scientist; for ETL orchestration, storage, or production data-quality operations, route to data-engineering; or for operating a named validation or database platform, route to that tool's skill. This skill supplies cleaning judgment and artifacts those workflows consume.

Prerequisites

  • Python 3.9+ with the standard library only for both bundled scripts (per compatibility); ecosystem tools (OpenRefine, pandas-backed tooling) are optional accelerators covered in references/cli-and-interactive-tools.md.
  • A raw input you can keep read-only plus write access to a separate output location — every script reads without modifying its input.
  • The templates above when the task warrants formal artifacts: a cleaning plan, transformation log, exception register, schema contract, or quality report.
  • pytest only when running the bundled test suites.

Limitations

  • The bundled profiler and reconciler are first-pass evidence tools: they surface anomalies and quantify deltas but do not decide preserve/repair/impute/quarantine — those classifications stay with the workflow's decision step.
  • Both scripts handle delimited text and JSONL; binary formats, relational databases, and nested document stores need other tooling.
  • Profiling output is evidence for investigation, never permission to auto-fix; an anomaly can be a real event.
  • A passing reconciliation proves structural preservation on the checked keys and sums only — semantic correctness of values still requires the review and release gate.

© magnus919, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 22 other files (scripts, references) in data-cleaning of magnus919/agent-skills.

  • SKILL.md
  • README.md
  • evals/evals.json
  • evals/trigger-cases.json
  • references/ai-repair-review.md
  • references/cli-and-interactive-tools.md
  • references/methodology.md
  • references/operations.md
  • references/sources.md
  • references/text-and-entity.md
  • references/tool-selection.md
  • references/validation-and-scale.md
  • scripts/profile_dataset.py
  • scripts/reconcile_dataset.py
  • scripts/test_profile_dataset.py
  • scripts/test_reconcile_dataset.py
  • templates/ai-repair-ledger.csv
  • … and 6 more

Open the folder on GitHubat commit c545c2b

Compare with similar skills

Data Cleaning next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Data Cleaning compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Data Cleaning this skillmagnus919/agent-skills116—~2.1kAutomated safety check: PassMIT
Credit Risk Data Cleaninggithub/awesome-copilot40k1 repos~1.5kAutomated safety check: PassMIT
Authoritative Data Harvesteryushui2022/MathModel-Skill4531 repos~1.1kAutomated safety check: PassMIT
Data Quality Frameworkswshobson/agents40k11 repos~1.1kAutomated safety check: PassMIT
Bio Batch ProcessingGPTomics/bioSkills1.2k1 repos~3kAutomated safety check: PassMIT
Preprocessing Data With Automated Pipelinesjeremylongshore/tons-of-skills-marketplace2.8k—~1kAutomated safety check: PassMIT

Similar skills

  • Credit Risk Data Cleaning

    github/awesome-copilot

    Official

    Cleans raw credit data and screens variables before loan modeling, dropping unstable, noisy or redundant features and writing an Excel report of every step.

    40k GitHub starsUsed in 1 repo~1.5k tokens
    Data & AnalyticsAuto-check passed
  • Authoritative Data Harvester

    yushui2022/MathModel-Skill

    Finds authoritative public data sources for modeling tasks, prefers official APIs and bulk downloads, and outputs a reproducible fetch and cleaning plan with citations.

    453 GitHub starsUsed in 1 repo~1.1k tokens
    Data & AnalyticsAuto-check passed
  • Sets up data quality checks with Great Expectations, dbt tests and data contracts, with checkpoints and pass-fail reports for pipelines.

    40k GitHub starsUsed in 11 repos~1.1k tokens
    Data & AnalyticsAuto-check passed
  • Bio Batch Processing

    GPTomics/bioSkills

    Process many sequence files in batch (count, merge, split, convert, summarize) with memory-safe streaming and on-disk indexing using Biopython, pysam, or pyfastx.

    1.2k GitHub starsUsed in 1 repo~3k tokens
    Data & AnalyticsAuto-check passed
  • Preprocessing Data With Automated Pipelines

    jeremylongshore/tons-of-skills-marketplace

    Process automate data cleaning, transformation, and validation for ML tasks.

    2.8k GitHub stars~1k tokensUpdated today
    Data & AnalyticsAuto-check passed
  • Crawl4AI Web Scraping

    smallnest/goclaw

    Scrapes sites, handles JavaScript-heavy pages and extracts structured data with Crawl4AI, through its crwl CLI or Python SDK, including schema-based extraction without an LLM.

    599 GitHub starsUsed in 1 repo~2.5k tokens
    Data & AnalyticsAuto-check passed

More from magnus919/agent-skills

All 130 skills in this repo
  • Artifact Pyramids

    magnus919/agent-skills

    Organize durable agent research outputs as summaries, analysis, and evidence dossiers.

    116 GitHub stars~2.7k tokensUpdated today
    Auto-check passed
  • Ascii City Engine

    magnus919/agent-skills

    Build portable, first-person colored ASCII city engines and small GIS-derived city packs.

    116 GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Color Management

    magnus919/agent-skills

    Manage color workflows with ICC profiles, working spaces, gamut mapping, and color science.

    116 GitHub stars~2.6k tokensUpdated today
    Auto-check: notes
  • Data Scientist

    magnus919/agent-skills

    A skill your agent uses for PhD-level expertise in data science, statistics, and machine learning: rigorous statistical analysis, experimental design, causal inference, advanced modeling, research…

    116 GitHub stars~4.1k tokensUpdated today
    Auto-check passed
  • Docker Compose

    magnus919/agent-skills

    Use Docker Compose to define, run, debug, and harden multi-container applications.

    116 GitHub stars~2k tokensUpdated today
    Auto-check: notes
  • Fpga Development

    magnus919/agent-skills

    Design, review, simulate, and verify FPGA logic using explicit RTL contracts, clock and reset models, CDC analysis, timing constraints, and reproducible implementation evidence.

    116 GitHub stars~2.7k tokensUpdated today
    Auto-check passed

Works with

Questions about Data Cleaning

What does Data Cleaning do?

Clean, profile, validate, reshape, and document messy tabular, text, JSON, and relational data through an evidence-first, reproducible workflow. Data Cleaning is an agent skill from magnus919/agent-skills. Clean, profile, validate, reshape, and document messy tabular, text, JSON, and relational data through an evidence-first, reproducible workflow.

When should I use Data Cleaning?

Data Cleaning fits situations like: preparing data for analysis; including AI-suggested repair; entity-match review; score calibration limits.

How do I install Data Cleaning in Claude Code?

Run `npx skills add magnus919/agent-skills --skill data-cleaning -a claude-code`. Or copy the skill folder (data-cleaning in magnus919/agent-skills) into .claude/skills/data-cleaning in your project. Claude Code loads it when a task matches its description.

How do I install Data Cleaning in Codex?

Run `npx skills add magnus919/agent-skills --skill data-cleaning -a codex`. Or copy the skill folder (data-cleaning in magnus919/agent-skills) into .agents/skills/data-cleaning in your project. Codex loads it when a task matches its description.

Can I use Data Cleaning in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add magnus919/agent-skills --skill data-cleaning -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/data-cleaning, .gemini/skills/data-cleaning, .github/skills/data-cleaning and .opencode/skills/data-cleaning in your project.

What does Data Cleaning need to run?

Going by SKILL.md and its folder, Data Cleaning needs Python for the scripts in its folder and the command-line tools its instructions call (python3). Our summary lists: Python 3. Compatibility (from SKILL.md): Works with any Agent Skills client. The bundled profiler requires Python 3.9+ and the standard library; ecosystem tools are optional..

Does Data Cleaning access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Data Cleaning safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Data Cleaning use?

Data Cleaning is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Data Cleaning use?

About 2.1k tokens (SKILL.md is roughly 8.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 4k tokens, read only when the agent opens those files.

What are the alternatives to Data Cleaning?

Skills that share tags, products or a category with Data Cleaning: Credit Risk Data Cleaning (github/awesome-copilot, 40k stars), Authoritative Data Harvester (yushui2022/MathModel-Skill, 453 stars), Data Quality Frameworks (wshobson/agents, 40k stars) and Bio Batch Processing (GPTomics/bioSkills, 1.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Data Cleaning?

magnus919 (a GitHub user) maintains it in magnus919/agent-skills, which has 116 GitHub stars. The repository holds 130 skills in this directory. The repository was last updated on October 8, 2026.

Source: magnus919/agent-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.