CRM Data Cleaner
LeoYeAI/openclaw-master-skills
Deduplicate, normalize, and enrich CRM contacts and companies.
Finds duplicate and junk records in a CRM CSV export with fuzzy matching, normalizes fields and writes a reviewable merge plan plus import-ready files without touching the live CRM.
$ npx skills add OneWave-AI/claude-skills --skill crm-data-cleanup -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install OneWave-AI/claude-skills crm-data-cleanup --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/OneWave-AI/claude-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/crm-data-cleanup .claude/skills/crm-data-cleanup && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "crm-data-cleanup" agent skill from https://github.com/OneWave-AI/claude-skills/tree/main/crm-data-cleanup into .claude/skills/crm-data-cleanup/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "crm-data-cleanup", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/OneWave-AI/claude-skills/tree/main/crm-data-cleanupType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add OneWave-AI/claude-skills --skill crm-data-cleanup -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install OneWave-AI/claude-skills crm-data-cleanup --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/OneWave-AI/claude-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/crm-data-cleanup .agents/skills/crm-data-cleanup && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "crm-data-cleanup" agent skill from https://github.com/OneWave-AI/claude-skills/tree/main/crm-data-cleanup into .agents/skills/crm-data-cleanup/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "crm-data-cleanup", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add OneWave-AI/claude-skills --skill crm-data-cleanup -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install OneWave-AI/claude-skills crm-data-cleanup --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/OneWave-AI/claude-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/crm-data-cleanup .cursor/skills/crm-data-cleanup && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "crm-data-cleanup" agent skill from https://github.com/OneWave-AI/claude-skills/tree/main/crm-data-cleanup into .cursor/skills/crm-data-cleanup/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "crm-data-cleanup", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/OneWave-AI/claude-skills.git --path crm-data-cleanup--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add OneWave-AI/claude-skills --skill crm-data-cleanup -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install OneWave-AI/claude-skills crm-data-cleanup --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/OneWave-AI/claude-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/crm-data-cleanup .gemini/skills/crm-data-cleanup && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "crm-data-cleanup" agent skill from https://github.com/OneWave-AI/claude-skills/tree/main/crm-data-cleanup into .gemini/skills/crm-data-cleanup/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "crm-data-cleanup", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install OneWave-AI/claude-skills crm-data-cleanupInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add OneWave-AI/claude-skills --skill crm-data-cleanup -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/OneWave-AI/claude-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/crm-data-cleanup .github/skills/crm-data-cleanup && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "crm-data-cleanup" agent skill from https://github.com/OneWave-AI/claude-skills/tree/main/crm-data-cleanup into .github/skills/crm-data-cleanup/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "crm-data-cleanup", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add OneWave-AI/claude-skills --skill crm-data-cleanup -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install OneWave-AI/claude-skills crm-data-cleanup --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/OneWave-AI/claude-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/crm-data-cleanup .opencode/skills/crm-data-cleanup && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "crm-data-cleanup" agent skill from https://github.com/OneWave-AI/claude-skills/tree/main/crm-data-cleanup into .opencode/skills/crm-data-cleanup/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "crm-data-cleanup", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
crm-data-cleanupFinds duplicate and junk records in a CRM CSV export with fuzzy matching, normalizes fields and writes a reviewable merge plan plus import-ready files without touching the live CRM.
Because every CRM can export CSV, the skill works the same for HubSpot, Salesforce, Pipedrive, Zoho, Close, Attio or any other. It addresses three ways duplicate cleanup goes wrong: exact-email matching misses most duplicates, merging on shared inboxes like `info@` fuses different people, and a merge can overwrite good values with blanks. Most CRM merges cannot be undone, so the skill never merges anything: it plans, a human approves and the CRM's own merge tool does the work. It also normalizes emails, phones to E.164, names, domains, states and countries, and flags stale, invalid, role and test records and broken contact-company links.
`scripts/dedupe.py` is the dry-run engine that normalizes, blocks, scores, clusters, picks a survivor and merges fields, and needs `pandas` and `rapidfuzz`, with `phonenumbers` used when installed. `scripts/evaluate.py` scores a run against a labeled truth file, and a fixture holds 41 synthetic contacts and 14 companies with planted duplicates. The workflow gets full exports with record IDs, confirms a dated backup, runs the script and checks `summary.json` for the column map, phone engine and warnings before trusting any number. Reference notes cover field standards, vendor merge behavior and safe rollout.
7 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit fc5b785. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 2 files in scripts/ (Python), which the agent can run.
Shell commands in SKILL.md call:
pythonpipFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
CRM Data Cleanup loads about 2.5k tokens when it runs, and up to ~9.2k if it reads all its reference files. Until then it costs about 203 tokens; SKILL.md has 1,182 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from OneWave-AI/claude-skills at commit fc5b785, republished under its MIT licence (© OneWave-AI). 1,182 words, ~2,468 tokens.
.claude/skills/crm-data-cleanup/SKILL.md (or your agent's skills folder). This skill also uses 8 other files; get the full folder from GitHub.This skill turns a messy CRM export into a merge plan a human can check, plus files that are ready to import. It works the same way for any vendor, because every CRM exports CSV.
Duplicate cleanup usually fails in one of three ways. Exact-email matching misses most
duplicates. Merging on shared inboxes like info@ fuses different people into one. And a
merge overwrites good values with blanks, losing history that can't be recovered. Most
CRM merges can't be undone, so this skill never merges anything itself. It plans, a
human approves, and the CRM's own merge tool does the merging.
scripts/dedupe.py: the dry-run engine: normalize, block, score, cluster, pick a
survivor, merge fields. Needs pandas and rapidfuzz. Uses phonenumbers if it's
installed.scripts/evaluate.py: scores a run against a labeled truth file (pairwise precision
and recall).assets/fixture/: 41 synthetic contacts and 14 companies with planted duplicates,
plus truth.csv.references/field-standards.md: normalization rules, state and country codes,
lifecycle stage mapping, and flag meanings.references/vendor-notes.md: what merge keeps and destroys in each CRM, import
quirks, and links to the official docs.references/safe-rollout.md: backup, sample merge, batch execution, and rollback.safe-rollout.md, Phase 0). A backup is the only real rollback.pip install pandas rapidfuzz phonenumbers # phonenumbers is optional but preferred
python scripts/dedupe.py --contacts contacts.csv --companies companies.csv \
--out plan/ --as-of 2026-09-22 [--default-region GB] [--gmail-canonical]--col email="E-mail 1" (or --company-col for the companies file).summary.json before trusting any number. Confirm column_map (a deal
Stage column mapped as lifecycle skews everything), phone_engine, and warnings,
such as a block that was skipped for being too large.merge or keep in the
decision column. Then look at the kept_survivor_conflict rows in
merge_plan.csv: those are the values that will disappear.safe-rollout.md, and check vendor-notes.md for the specific CRM.info@, sales@, a household
inbox, or an office main line shows that records belong to the same organization, not
that they're the same person. A shared identifier goes to review at most.--max-cluster (5) members, to review.additional_emails and additional_phones, so they aren't dropped.most_recent_activity, then
most_complete, then oldest_created. Add business_email first when the work
address should be the primary email. In HubSpot and Pipedrive the primary's values win
every conflict, so the survivor choice decides what's kept.--gmail-canonical). Google says dots
don't matter for personal Gmail, but they do matter for Workspace domains. Applied by
default, it would merge real, distinct mailboxes.record_flags.csv. The
data owner decides whether to archive or delete them.| File | Contents |
|---|---|
merge_plan.csv | One row per cluster per field: cluster_id, survivor_id, loser_ids, survivor_rule, field, winning_value, source_id, action, other_values. The action is kept_survivor, kept_survivor_conflict, filled_from_loser, earliest, latest, furthest_stage, or preserved_secondary |
review_queue.csv | Pairs needing a human: both records side by side, confidence, reasons, suggested_survivor, and blank decision and decided_by columns |
cleaned_contacts.csv | The original columns with normalized values. Survivors carry merged values, and losers are removed. Also additional_emails, additional_phones, _cleanup_action, and _flags |
record_flags.csv | invalid_email, email_domain_typo, role_email, shared_phone, no_contact_method, phone_unparsed, junk_or_test, missing_name, stale |
companies_*.csv, cleaned_companies.csv | The same outputs for companies (domain-first matching) |
associations.csv | company_merged (the ID was re-pointed to the survivor), orphan_company_id, missing_association (the email domain matches a company) |
summary.json | Counts, flag totals, column map, settings, warnings |
When reporting to the user, give: records in and out, auto clusters, review pairs, flag
counts, the three riskiest review pairs, and the next gate from safe-rollout.md.
python scripts/dedupe.py --contacts assets/fixture/contacts.csv \
--companies assets/fixture/companies.csv --out plan/ --as-of 2026-09-22
python scripts/evaluate.py --plan-dir plan/ --truth assets/fixture/truth.csvThe fixture is 41 contacts holding 14 true duplicate pairs. It includes case and
whitespace email variants, a name typo, a nickname with a reformatted phone, the same
person with a work and a Gmail address, Inc versus no Inc, a hyphenated married name,
a 3-record cluster, and a Gmail dot/plus variant. It also plants traps that must not
merge: three people sharing info@stark.com, a couple sharing a household inbox, three
colleagues on one office line, and two different people named David Kim.
| Run | Auto pairs | Wrong auto | Auto precision | Auto recall | Recall incl. review |
|---|---|---|---|---|---|
| Naive: exact lowercase email only | 11 | 4 (the info@ trio and the household) | 63.6% | 50.0% | n/a |
| Default | 11 | 0 | 100% | 78.6% | 100% |
--gmail-canonical | 12 | 0 | 100% | 85.7% | 100% |
In the default run, 7 contact clusters absorb 9 records, and 6 pairs go to review: 3 true
duplicates plus the household and info@ traps. For companies, Acme Inc/ACME and
Tyrell Corp/The Tyrell Corporation auto-merge by domain. Globex LLC/Globex Corporation (no domain) and Umbrella Ltd/Umbrella Corporation (different domains) go
to review. Three association issues surface, including an orphaned CO99 reference.
Cluster MC0001 (Jane Doe x3) shows the field rules at work. Create date 2022-03-14
comes from the oldest record. Lifecycle stage customer is the furthest along. The blank
state on the survivor is filled from C001, and the work email is kept in
additional_emails.
--auto-name-sim 85, --review-name-sim 88, --conflict-name-sim 60: raise the auto
threshold for large, name-dense databases.--stale-days 365, --as-of: set the staleness window. Always pass --as-of so reruns
are reproducible.--max-block 500: blocks bigger than this are skipped with a warning. Keeping blocks
small keeps pair scoring fast on 100k-row exports.NICKNAMES, ROLE_LOCALS, LIFECYCLE_MAP, or LEGAL_SUFFIXES in the script
for a specific region or portal. Custom HubSpot lifecycle stages have numeric internal
IDs.© OneWave-AI, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 8 other files (scripts, references, assets) in crm-data-cleanup of OneWave-AI/claude-skills.
Open the folder on GitHubat commit fc5b785
CRM Data Cleanup next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| CRM Data Cleanup this skillOneWave-AI/claude-skills | 336 | — | ~2.5k | Automated safety check: Pass | MIT | |
| CRM Data CleanerLeoYeAI/openclaw-master-skills | 2.2k | — | ~7.3k | Automated safety check: Pass | MIT | |
| Platform Validation Rule Generateforcedotcom/sf-skills | 1.1k | — | ~979 | Automated safety check: Pass | Apache-2.0 | |
| Lead Gen Tool Builderexplorium-ai/gtm-skills | 185 | — | ~1.8k | Automated safety check: Notes | MIT | |
| Google Maps Exportgmapsscraper/google-maps-agent-skills | 132 | — | ~1.2k | Automated safety check: Pass | MIT | |
| Hubspot Integrationdavila7/claude-code-templates | 33k | 4 repos | ~265 | Automated safety check: Pass | MIT |
LeoYeAI/openclaw-master-skills
Deduplicate, normalize, and enrich CRM contacts and companies.
forcedotcom/sf-skills
A skill your agent uses when users need to create, modify, or validate Salesforce Validation Rules.
explorium-ai/gtm-skills
Lead generation tool builder skill for Claude Code and Codex: scaffolds a complete, self-hostable, ZoomInfo-style B2B lead-generation web app — company & contact search UI, firmographic and…
gmapsscraper/google-maps-agent-skills
Export Google Maps business data to CSV, JSON, or CRM format (HubSpot, Pipedrive, Salesforce).
davila7/claude-code-templates
Expert patterns for HubSpot CRM integration including OAuth authentication, CRM objects, associations, batch operations, webhooks, and custom objects.
databricks/databricks-agent-skills
Build managed ingestion pipelines into Databricks using Lakeflow Connect.
OneWave-AI/claude-skills
Repairs broken decks and PDFs exported from Claude Design or similar AI deck generators: clipped text, wrong fonts and corrupted .pptx package structure.
OneWave-AI/claude-skills
Writes, explains, debugs, and optimizes BI calculations - Power BI / Fabric DAX measures and calculated columns, Tableau calculated fields (FIXED/INCLUDE/EXCLUDE LOD expressions, table…
OneWave-AI/claude-skills
Categorizes transactions, reconciles bank and card statements to the ledger, works a month-end checklist and prepares a close package, without ever forcing a balance.
OneWave-AI/claude-skills
Combines CSV, TSV and Excel files into one verified table with pandas, by stacking or joining, mapping columns, normalizing keys and removing duplicates.
OneWave-AI/claude-skills
Pulls financial statement numbers for US public companies straight from SEC EDGAR's free official XBRL APIs (companyfacts, companyconcept, frames, submissions) into a cited table.
OneWave-AI/claude-skills
Answers business questions about the user's own spreadsheet or data export (CSV, TSV, XLSX from a CRM, Shopify, Stripe, QuickBooks, ad platforms, HR or payroll systems) correctly and auditably.
Works with
Categories
Finds duplicate and junk records in a CRM CSV export with fuzzy matching, normalizes fields and writes a reviewable merge plan plus import-ready files without touching the live CRM. Because every CRM can export CSV, the skill works the same for HubSpot, Salesforce, Pipedrive, Zoho, Close, Attio or any other. It addresses three ways duplicate cleanup goes wrong: exact-email matching misses most duplicates, merging on shared inboxes like `info@` fuses different people, and a merge can overwrite good values with blanks.
CRM Data Cleanup fits situations like: deduplicating contacts or companies in a CRM export; cleaning a CRM before a migration or re-import; finding invalid, role-based, stale or test records in a lead list; checking contact-company associations that do not match.
Run `npx skills add OneWave-AI/claude-skills --skill crm-data-cleanup -a claude-code`. Or copy the skill folder (crm-data-cleanup in OneWave-AI/claude-skills) into .claude/skills/crm-data-cleanup in your project. Claude Code loads it when a task matches its description.
Run `npx skills add OneWave-AI/claude-skills --skill crm-data-cleanup -a codex`. Or copy the skill folder (crm-data-cleanup in OneWave-AI/claude-skills) into .agents/skills/crm-data-cleanup in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add OneWave-AI/claude-skills --skill crm-data-cleanup -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/crm-data-cleanup, .gemini/skills/crm-data-cleanup, .github/skills/crm-data-cleanup and .opencode/skills/crm-data-cleanup in your project.
Going by SKILL.md and its folder, CRM Data Cleanup needs Python for the scripts in its folder and the command-line tools its instructions call (python and pip). Our summary lists: Python with pandas and rapidfuzz (phonenumbers optional); CSV exports of contacts and companies with record IDs.
SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
CRM Data Cleanup is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.5k tokens (SKILL.md is roughly 9.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 6.8k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with CRM Data Cleanup: CRM Data Cleaner (LeoYeAI/openclaw-master-skills, 2.2k stars), Platform Validation Rule Generate (forcedotcom/sf-skills, 1.1k stars), Lead Gen Tool Builder (explorium-ai/gtm-skills, 185 stars) and Google Maps Export (gmapsscraper/google-maps-agent-skills, 132 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
OneWave-AI (a GitHub organization) maintains it in OneWave-AI/claude-skills, which has 336 GitHub stars. The repository holds 69 skills in this directory. The repository was last updated on October 2, 2026.
Source: OneWave-AI/claude-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.