Agent skill

Clean Data

by explorium-ai in explorium-ai/gtm-skills

Data cleaning, entity matching, and deduplication skill for Claude Code and Codex: triage, standardize, and validate a CSV, Excel, or JSON list of B2B companies or contacts before enrichment.

MITAuto-check passedData & Analytics

Install Clean Data

skills CLI
$ npx skills add explorium-ai/gtm-skills --skill clean-data -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install explorium-ai/gtm-skills clean-data --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/explorium-ai/gtm-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/clean-data .claude/skills/clean-data && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
clean-data
GitHub stars
160
Token cost
~2k tokens
SKILL.md length
953 words
Files
1
Skills in repo
17
Repo updated
First seen
Licence
MIT

At a glance

Data cleaning, entity matching, and deduplication skill for Claude Code and Codex: triage, standardize, and validate a CSV, Excel, or JSON list of B2B companies or contacts before enrichment.

  • Works in 5 steps: Copy raw input. Before any transform,… → Profile the file. Compute fill rate,… → Standardize string fields. Run… → …
  • CRM enrichment prep
  • SKILL.md covers Input, Workflow, Output Format and Limitations
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Clean Data is an agent skill from explorium-ai/gtm-skills. Data cleaning, entity matching, and deduplication skill for Claude Code and Codex: triage, standardize, and validate a CSV, Excel, or JSON list of B2B companies or contacts before enrichment. Normalizes company names and domains, validates emails and phone numbers, matches and deduplicates records, and tags invalid rows non-destructively. Run before enrichment or CRM import to avoid paying for noisy records. Use for CRM enrichment prep, company name and website matching, entity matching, and contact…

Its SKILL.md is about 2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Data & Analytics, covering Data cleaning and Excel spreadsheets. It works with Microsoft Excel, Model Context Protocol and n8n. The repository describes itself as: GTM Skills for Claude & Codex. The licence is MIT.

When your agent uses it

  • CRM enrichment prep
  • Company name and website matching
  • Entity matching
  • Contact deduplication

Example prompts

  • “/clean-data”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Copy raw input. Before any transform, copy the input to ./00_raw/ and never write back. All transforms write to numbered phase folders…
  2. Profile the file. Compute fill rate, cardinality, top values, length distribution, and format-pattern frequency for every column. Save the…
  3. Standardize string fields. Run standardization BEFORE validation: a valid email like JOHN@ACME.COM fails naive regex without…
  4. Validate field-by-field. Per field, add a boolean _valid and a _reason text column when invalid. Tag invalid rows; never delete them.
  5. Hand off to entity resolution (optional). This skill cleans rows in isolation; it cannot tell you that Starbucks EMEA and Starbucks…

What it can do on your machine

Read from SKILL.md and the folder at commit f0efa6b. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Clean Data loads about 2k tokens when it runs. Until then it costs about 144 tokens; SKILL.md has 953 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~144
When it runs · the whole SKILL.md, loaded when a task matches
~2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from explorium-ai/gtm-skills at commit f0efa6b, republished under its MIT licence (© explorium-ai). 953 words, ~2,005 tokens.

Download SKILL.mdSave it as .claude/skills/clean-data/SKILL.md (or your agent's skills folder).
name
clean-data
description
Data cleaning, entity matching, and deduplication skill for Claude Code and Codex: triage, standardize, and validate a CSV, Excel, or JSON list of B2B companies or contacts before enrichment. Normalizes company names and domains, validates emails and phone numbers, matches and deduplicates records, and tags invalid rows non-destructively. Run before enrichment or CRM import to avoid paying for noisy records. Use for CRM enrichment prep, company name and website matching, entity matching, and contact deduplication. Works in Claude Code, Codex, and n8n via MCP.

Clean Data

Per-row cleanup on a GTM list before any match or enrich call. Profile, standardize, validate. Never destructive: raw input is preserved, invalid rows are tagged with reason codes rather than deleted. Out of scope: deduplication and canonical entity resolution.

Input

$ARGUMENTS is a path to a CSV, Excel, or JSON file. Parse the user message for optional sub-inputs:

  • Schema hints if column names are ambiguous: which column is company name, domain, email, phone, country.
  • Entity type: companies, contacts, or both. Default: infer from columns.
  • Whether to run an MX-record check on email domains (slower, requires DNS). Default: off.
  • Whether to produce a match-ready subset for downstream entity resolution. Default: no.

Example phrasings:

  • "Clean this leads CSV before I import it to HubSpot."
  • "Normalize the company names and domains in /path/accounts.xlsx."
  • "Validate the emails in this file and tag the bad rows."
  • "Prep this list for prospect matching, only keep contactable corporate rows."
  • "Why does my country column have 200 different spellings."

Workflow

  1. Copy raw input. Before any transform, copy the input to ./00_raw/<filename> and never write back. All transforms write to numbered phase folders: 01_profiled/, 02_standardized/, 03_validated/. The single most common failure mode in cleanup work is destructive transforms with no path back.

  2. Profile the file. Compute fill rate, cardinality, top values, length distribution, and format-pattern frequency for every column. Save the snapshot. Read it before deciding what to clean.

    python
    import pandas as pd
    df = pd.read_csv(input_path)
    profile = pd.DataFrame({
        "fill_rate_pct": (df.notna().mean() * 100).round(1),
        "cardinality": df.nunique(),
        "top_value": df.apply(lambda c: c.dropna().astype(str).mode().iloc[0] if c.dropna().size else None),
        "avg_len": df.apply(lambda c: c.dropna().astype(str).str.len().mean()),
    })

    Things to look for: country columns with 200+ distinct values (standardization problem, build ISO Alpha-2 lookup); phone columns where under 50% parse as E.164 (need country hint); company-name 99th-percentile length above 100 chars (pasted addresses, quarantine); free-email providers in the top 5 of email column (decide policy now); fields under 10% fill (probably not worth normalizing); literal strings "NA", "N/A", "None", "null", "-" (collapse to real nulls before validating).

  3. Standardize string fields. Run standardization BEFORE validation: a valid email like JOHN@ACME.COM fails naive regex without trim+lowercase first. For every string column do Unicode NFKC, trim, collapse internal whitespace, strip leading and trailing punctuation, collapse null-token strings to real nulls. Then field-specific:

    • Company name. Strip legal suffixes (Inc, LLC, Ltd, GmbH, S.A., 株式会社) at end of string only. Use cleanco if available. Keep BOTH raw and normalized columns.
    • Domain. Strip protocol and www. Fold to the eTLD+1 via tldextract. Flag free-email providers and disposable domains separately.
    • Person name. Parse with nameparser: honorifics, generational suffixes, credentials, particles. If confidence is low, store the raw string with a low-confidence flag.
    • Phone. Format to E.164 with phonenumbers. Hint country from the country column when available.
    • Country. Map free-text to ISO Alpha-2 codes (United States to US, UK to GB, Deutschland to DE). Reusable downstream for country filters.
    • Address. Use libpostal if installed. Country-aware parsing.

    Common mistake: overwriting the display column with the normalized version. Always keep raw alongside normalized.

  4. Validate field-by-field. Per field, add a boolean <field>_valid and a <field>_reason text column when invalid. Tag invalid rows; never delete them.

    • Emails: RFC 5322 syntax via email-validator; role-address detection (info@, sales@, noreply@, support@, hello@); disposable-domain check; free-provider flag (gmail, yahoo, qq); optional MX-record check (off by default).
    • Phones: parse + format via phonenumbers. Tag invalid_too_short, invalid_country, invalid_format.
    • Domains: valid eTLD, no IP literals, optional MX check.
    • Country codes: valid ISO Alpha-2 after normalization.
  5. Hand off to entity resolution (optional). This skill cleans rows in isolation; it cannot tell you that Starbucks EMEA and Starbucks Corporation point to the same company. If the user wants the handoff, produce a match-ready subset and route rows by available signal:

    • Rows with normalized company name + domain: route to match a business (name + website, falls back to domain-only on mismatch).
    • Rows with a corporate (not role / free-provider / disposable) email: route to match a prospect via email.
    • Rows with parsed person name + company name: route to match a prospect via name + company.
    • Rows with a validated LinkedIn URL: route to match a prospect via LinkedIn.

    Filter out tagged-invalid rows before the handoff so you do not spend credits matching noreply@example.com or disposable addresses. The returned IDs become the join keys for any later enrich a business or enrich a prospect call.

Show full SKILL.md (268 more words)Show less

Output Format

Profile Snapshot

Per column: fill_rate_pct, cardinality, top_value, avg_len. Markdown table. After cleanup, re-run the profile and show before vs after on touched columns.

Standardization Map

Per normalized field: raw column name, normalized column name, 3 to 5 example transformations (" ACME, Inc. " to acme, "WWW.Acme.COM" to acme.com, "+1 (415) 555 1212" to +14155551212).

Validation Verdict

Per validated field: counts of valid, invalid, risky. Frequency table of reason codes (e.g. role_address: 42, disposable_domain: 18, invalid_syntax: 6).

Cleaned File

A single CSV at ./03_validated/<input_name>_clean.csv with all original columns plus <field>_norm, <field>_valid, and <field>_reason columns.

Match-Ready Subset (only if requested)

A second CSV at ./04_match_ready/<input_name>_for_match.csv containing only rows that passed validation, plus a one-line summary of which match path each row subset should route to (count by path).

Limitations

  • Does NOT deduplicate or resolve to canonical entities. Two normalized strings can still refer to the same real-world business. Hand off to a match step for that.
  • Person-name parsing is best-effort. Non-Western order, hyphenated families, and missing separators are flagged with a low-confidence marker; raw string is always preserved.
  • The MX-record check requires DNS and adds 50 to 200ms per unique domain. Off by default.
  • Free-email providers (gmail, yahoo) are flagged but cannot be linked to a corporate identity from this skill alone. Pair with a prospect match (email + company) when needed.
  • Holding-company and subsidiary pitfalls are out of scope (meta.com vs instagram.com vs whatsapp.com). Resolve via company-hierarchies enrichment after matching.
  • If the input lacks a country column entirely, phone normalization defaults to a permissive parser and may mis-format short numbers. Provide a default country in the user message when possible.

© explorium-ai, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/clean-data of explorium-ai/gtm-skills.

Open the folder on GitHubat commit f0efa6b

Compare with similar skills

Clean Data next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Clean Data compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Clean Data this skillexplorium-ai/gtm-skills160—~2kAutomated safety check: PassMIT
Clean DataAperivue/medsci-skills329—~2kAutomated safety check: PassMIT
Dataset Quality Auditzebbern/claude-code-guide4.6k—~996Automated safety check: PassMIT
Outlier Detection And Quality AssessmentMichaelYang-lyx/AIDABench1111 repos~1kAutomated safety check: PassNone
Visual Skillsnpc-live/clawfirm156—~7.4kAutomated safety check: PassNone
Invalid Data CleaningMichaelYang-lyx/AIDABench1111 repos~410Automated safety check: PassNone

Similar skills

  • Clean Data

    Aperivue/medsci-skills

    A skill your agent uses when a clinical CSV/Excel dataset needs profiling and cleaning before analysis (missing values, outliers, duplicates, type mismatches).

    329 GitHub stars~2k tokensUpdated 2 days ago
    Data & AnalyticsAuto-check passed
  • Dataset Quality Audit

    zebbern/claude-code-guide

    Run comprehensive quality checks on tabular data (CSV/Excel/TSV/JSON), detecting missing values, duplicates, outliers, format issues, and type inconsistencies to produce an overall score, grade, and…

    4.6k GitHub stars~996 tokensUpdated today
    Data & AnalyticsAuto-check passed
  • 执行全面的异常值检测与数据质量评估,利用 IQR 方法识别异常值并结合偏度、峰度分析数据分布特征,适用于非正态分布数据的预处理阶段。

    111 GitHub starsUsed in 1 repo~1k tokens
    Data & AnalyticsAuto-check passed
  • Visual Skills

    npc-live/clawfirm

    A skill your agent uses whenever the user provides data (CSV, JSON, table, pasted numbers, or any structured dataset) and expects a visual output — even if they don't say 'chart' or 'visualize'.

    156 GitHub stars~7.4k tokensUpdated 3 mo ago
    Data & AnalyticsAuto-check passed
  • Invalid Data Cleaning

    MichaelYang-lyx/AIDABench

    用于大规模Excel数据的预处理,通过统计总行数判断是否转换为Parquet格式以提升读写效率,并使用正则表达式清洗指定文本列(如仅保留中文字符),最后导出清洗后的文件并提供下载链接。

    111 GitHub starsUsed in 1 repo~410 tokens
    Data & AnalyticsAuto-check passed
  • Data Cleaning

    ericrisco/rsc-harness

    A skill your agent uses when a raw table is too dirty to trust — nulls, sentinels, duplicate rows, category sprawl, mixed types, bad dates — and you need a re-runnable clean() plus a schema gate…

    156 GitHub stars~3.6k tokensUpdated yesterday
    Data & AnalyticsAuto-check passed

More from explorium-ai/gtm-skills

All 17 skills in this repo
  • Browser Extension Builder

    explorium-ai/gtm-skills

    Browser extension builder skill for Claude Code and Codex: scaffolds a local, unpacked Chrome extension that reveals verified B2B contact info (email, phone, job title, company) directly on a…

    160 GitHub stars~2.6k tokensUpdated 13 days ago
    Auto-check passed
  • Lead Gen Tool Builder

    explorium-ai/gtm-skills

    Lead generation tool builder skill for Claude Code and Codex: scaffolds a complete, self-hostable, ZoomInfo-style B2B lead-generation web app — company & contact search UI, firmographic and…

    160 GitHub stars~1.8k tokensUpdated 13 days ago
    Auto-check: notes
  • Abm Diy Campaign

    explorium-ai/gtm-skills

    ABM campaign skill for Claude Code: run a full Account-Based Marketing campaign end-to-end — from ICP definition to live LinkedIn Ads.

    160 GitHub stars~2.4k tokensUpdated 13 days ago
    Auto-check passed
  • Account Contact Shortlist

    explorium-ai/gtm-skills

    Contact data skill for Claude Code and Codex: build a ranked shortlist of decision-makers and contacts at a target company for outbound prospecting, deal acceleration, or renewal/expansion plays.

    160 GitHub stars~1.6k tokensUpdated 13 days ago
    Auto-check passed
  • Account Fit Rank

    explorium-ai/gtm-skills

    Lead scoring and buying signals skill for Claude Code and Codex: rank a list of accounts by ICP fit, buying intent, real-time trigger events, and workforce momentum.

    160 GitHub stars~2k tokensUpdated 13 days ago
    Auto-check passed
  • Account Research

    explorium-ai/gtm-skills

    Account research skill for Claude Code and Codex: generate a high-signal company intelligence brief including firmographics, technographics, funding history, hiring signals, business events, recent…

    160 GitHub stars~2.8k tokensUpdated 13 days ago
    Auto-check passed

Questions about Clean Data

What does Clean Data do?

Data cleaning, entity matching, and deduplication skill for Claude Code and Codex: triage, standardize, and validate a CSV, Excel, or JSON list of B2B companies or contacts before enrichment. Clean Data is an agent skill from explorium-ai/gtm-skills. Data cleaning, entity matching, and deduplication skill for Claude Code and Codex: triage, standardize, and validate a CSV, Excel, or JSON list of B2B companies or contacts before enrichment.

When should I use Clean Data?

Clean Data fits situations like: CRM enrichment prep; company name and website matching; entity matching; contact deduplication.

How do I install Clean Data in Claude Code?

Run `npx skills add explorium-ai/gtm-skills --skill clean-data -a claude-code`. Or copy the skill folder (skills/clean-data in explorium-ai/gtm-skills) into .claude/skills/clean-data in your project. Claude Code loads it when a task matches its description.

How do I install Clean Data in Codex?

Run `npx skills add explorium-ai/gtm-skills --skill clean-data -a codex`. Or copy the skill folder (skills/clean-data in explorium-ai/gtm-skills) into .agents/skills/clean-data in your project. Codex loads it when a task matches its description.

Can I use Clean Data in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add explorium-ai/gtm-skills --skill clean-data -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/clean-data, .gemini/skills/clean-data, .github/skills/clean-data and .opencode/skills/clean-data in your project.

What does Clean Data need to run?

SKILL.md names no scripts, command-line tools or credentials: Clean Data is instructions for the agent only. Our summary lists: Python 3.

Does Clean Data access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Clean Data safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Clean Data use?

Clean Data is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Clean Data use?

About 2k tokens (SKILL.md is roughly 8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Clean Data?

Skills that share tags, products or a category with Clean Data: Clean Data (Aperivue/medsci-skills, 329 stars), Dataset Quality Audit (zebbern/claude-code-guide, 4.6k stars), Outlier Detection And Quality Assessment (MichaelYang-lyx/AIDABench, 111 stars) and Visual Skills (npc-live/clawfirm, 156 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Clean Data?

explorium-ai (a GitHub organization) maintains it in explorium-ai/gtm-skills, which has 160 GitHub stars. The repository holds 17 skills in this directory. The repository was last updated on September 24, 2026.

Source: explorium-ai/gtm-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.