Agent skill

CSV and Excel Merger

by OneWave-AI in OneWave-AI/claude-skills

Combines CSV, TSV and Excel files into one verified table with pandas, by stacking or joining, mapping columns, normalizing keys and removing duplicates.

MITAuto-check passedDocuments & Office

Install CSV and Excel Merger

skills CLI
$ npx skills add OneWave-AI/claude-skills --skill csv-excel-merger -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install OneWave-AI/claude-skills csv-excel-merger --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/OneWave-AI/claude-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/csv-excel-merger .claude/skills/csv-excel-merger && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
csv-excel-merger
GitHub stars
336
Token cost
~1.6k tokens
SKILL.md length
556 words
Files
4 (incl. scripts, references)
Skills in repo
69
Repo updated
First seen
Licence
MIT

At a glance

Combines CSV, TSV and Excel files into one verified table with pandas, by stacking or joining, mapping columns, normalizing keys and removing duplicates.

  • Works in 6 steps: Profile inputs. Run the bundled profiler… → Choose the operation. This is the… → Map columns and normalize keys. Build an… → …
  • Stacking monthly exports into a single spreadsheet
  • SKILL.md covers Workflow, pandas version notes and Failure modes to check for
  • Runs Python scripts from its folder; calls python

What it does

The workflow starts with a bundled `scripts/profile_inputs.py` that reports each file's encoding, delimiter, row count, headers, candidate keys and header overlap without changing anything. Excel workbooks are profiled sheet by sheet, and the agent confirms with you which sheets count. It then decides between appending, for the same kind of records from different periods or sources, and joining, for different facts about the same entities.

Columns are mapped through an explicit rename map shown to you when matches are fuzzy, and keys such as emails and phone numbers are normalized before deduplication. Files are read as strings to keep leading zeros, and joins use pandas validation so duplicate keys raise an error. Before reporting, the agent checks row counts, tracks which file each row came from, and follows conflict strategies and an output template in the references folder.

When your agent uses it

  • Stacking monthly exports into a single spreadsheet
  • Consolidating contact or lead lists from several sources and removing duplicates
  • Joining two sheets on an ID or email column, like a VLOOKUP
  • Merging multi-sheet Excel workbooks whose column names do not match

Example prompts

  • “Merge the January, February and March sales CSVs in ./exports into one file and drop duplicate orders.”
  • “Join contacts.xlsx with deals.csv on email and list the contacts with no matching deal.”
  • “Put these three lead lists together and tell me which file each row came from.”
  • “Clean up the two vendor spreadsheets into one list with a single set of column names.”

Requirements

  • Python 3 with pandas

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Profile inputs. Run the bundled profiler first; it reports encoding, delimiter, rows, headers, candidate keys, and header overlap without…
  2. Choose the operation. This is the decision that most often goes wrong
  3. Map columns and normalize keys. Build an explicit {original: unified} rename map per file (see references/merge_strategies.md for common…
  4. Merge. Read every file with dtype=str so IDs, ZIP codes, and phone numbers keep leading zeros, then convert specific columns afterward.
  5. Verify before reporting. Never hand back a merge without checking it
  6. Write output and report. Use the layout in references/output_template.md.

What it can do on your machine

Read from SKILL.md and the folder at commit fc5b785. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

CSV and Excel Merger loads about 1.6k tokens when it runs, and up to ~3k if it reads all its reference files. Until then it costs about 145 tokens; SKILL.md has 556 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~145
When it runs · the whole SKILL.md, loaded when a task matches
~1.6k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from OneWave-AI/claude-skills at commit fc5b785, republished under its MIT licence (© OneWave-AI). 556 words, ~1,638 tokens.

Download SKILL.mdSave it as .claude/skills/csv-excel-merger/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
csv-excel-merger
description
Merges, appends, joins, and deduplicates multiple CSV, TSV, and Excel files (including multi-sheet workbooks) with pandas - mapping mismatched column names to one schema, normalizing keys, resolving conflicting values, tracking which file each row came from, and verifying the row math. Use when the user wants to combine spreadsheets or exports, stack monthly files, consolidate contact or lead lists, VLOOKUP-style join two sheets on an ID or email, or dedupe records across sources, even if they only say "put these together" or "clean up these lists into one".

CSV/Excel Merger

Combine tabular files into one clean output without silently losing, duplicating, or corrupting rows.

Workflow

Copy this checklist and track progress:

- [ ] 1. Profile inputs
- [ ] 2. Choose append vs join
- [ ] 3. Map columns and normalize keys
- [ ] 4. Merge and resolve conflicts
- [ ] 5. Verify row math
- [ ] 6. Write output and report
  1. Profile inputs. Run the bundled profiler first; it reports encoding, delimiter, rows, headers, candidate keys, and header overlap without changing anything:

    bash
    python scripts/profile_inputs.py file1.csv file2.xlsx

    Excel files are profiled per sheet. Confirm with the user which sheets count if a workbook has more than one.

  2. Choose the operation. This is the decision that most often goes wrong:

    • Append (stack) - same kind of records from different sources or periods (Jan + Feb exports, three lead lists). Use pd.concat, then dedupe.
    • Join (enrich) - different facts about the same entities (contacts + their deal values). Use pd.merge on a key.
    • Unsure? If the files share most columns, append. If they share only an ID column, join.
  3. Map columns and normalize keys. Build an explicit {original: unified} rename map per file (see references/merge_strategies.md for common variants) and show it to the user when any match is fuzzy. Normalize key columns before dedupe or join: strip whitespace, lowercase emails, strip non-digits from phones, unify date formats. Without this, A@x.com and a@x.com survive as two people.

  4. Merge. Read every file with dtype=str so IDs, ZIP codes, and phone numbers keep leading zeros, then convert specific columns afterward.

    python
    import pandas as pd
    
    frames = []
    for path, rename in [("jan.csv", {"E-mail": "email"}), ("feb.xlsx", {"Email Address": "email"})]:
        df = (pd.read_excel(path, dtype=str) if path.endswith((".xlsx", ".xls"))
              else pd.read_csv(path, dtype=str, encoding="utf-8-sig",  # use the profiler's encoding
                            keep_default_na=False))
        df = df.rename(columns=rename)
        df["email"] = df["email"].str.strip().str.lower()
        df["source_file"] = path          # lineage for every row
        frames.append(df)
    
    combined = pd.concat(frames, ignore_index=True, sort=False)
    # Blank keys are not duplicates of each other: set them aside before deduping.
    has_key = combined["email"].fillna("") != ""
    # Later files win: list the most recent source last, then keep="last".
    deduped = combined[has_key].drop_duplicates(subset=["email"], keep="last")
    no_key = combined[~has_key]
    merged = pd.concat([deduped, no_key], ignore_index=True)

    For a join, make pandas enforce the relationship you expect so a duplicate key raises instead of multiplying rows:

    python
    out = pd.merge(contacts, deals, on="email", how="left",
                   validate="one_to_one", indicator=True)
    unmatched = out[out["_merge"] == "left_only"]

    Conflict strategies (keep first/last/most complete, combine fields, flag for review) are in references/merge_strategies.md.

  5. Verify before reporting. Never hand back a merge without checking it:

    python
    rows_in = sum(len(f) for f in frames)
    assert len(merged) > 0, "merge produced an empty frame"
    assert len(merged) <= rows_in, "more rows out than in: check the join keys"
    assert deduped["email"].is_unique, "duplicate keys remain after dedupe"
    print(f"in={rows_in} out={len(merged)} removed={rows_in - len(merged)} blank_keys={len(no_key)}")
    print(merged["source_file"].value_counts())

    Spot-check three removed duplicates by hand against the source files; the asserts prove the math, not that the right row won.

  6. Write output and report. Use the layout in references/output_template.md.

    • CSV for Excel users: to_csv(path, index=False, encoding="utf-8-sig") (the BOM makes Excel read accents correctly).
    • Excel: to_excel(path, index=False) with openpyxl installed. A sheet holds at most 1,048,576 rows; split or use CSV/Parquet beyond that.
    • Also write conflicts_review.csv or unmatched.csv when those sets are non-empty.
Show full SKILL.md (226 more words)Show less

pandas version notes

Current pandas is 3.x (Python 3.11+). Differences that affect merges:

  • Text columns default to the str dtype, not object. Check pd.api.types.is_string_dtype(col) instead of dtype == object.
  • Copy-on-Write is always on. Chained assignment such as df[col][mask] = x never updates df (pandas only warns); use df.loc[mask, col] = x.
  • Parsed datetimes default to microsecond resolution. Call .dt.as_unit("ns") before casting to integers if something downstream expects nanoseconds.
  • pd.read_excel(..., engine="calamine") (needs python-calamine) reads large workbooks much faster than openpyxl.

The code in this skill also runs on pandas 2.2.

Failure modes to check for

  • Row explosion on join - duplicate keys on both sides multiply rows. validate= catches it.
  • Leading zeros lost - reading without dtype=str turns 01234 into 1234.
  • Excel-mangled values - long IDs already shown as 1.23E+15 or dates already reformatted in the source file cannot be recovered by pandas; flag them.
  • Header rows not on line 1 - exports with a title block need skiprows= or header=.
  • Mixed encodings - one file in cp1252 among UTF-8 files shows up as é artifacts. The profiler reports the encoding per file.
  • Silent column drops - a column present in only one file becomes mostly empty after append. Keep it and report its completeness; never drop data without saying so.
  • Large files (over a few hundred MB) - read with chunksize= or use Polars/DuckDB, and dedupe with a key set instead of loading everything into memory.

© OneWave-AI, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files (scripts, references) in csv-excel-merger of OneWave-AI/claude-skills.

  • SKILL.md
  • references/merge_strategies.md
  • references/output_template.md
  • scripts/profile_inputs.py

Open the folder on GitHubat commit fc5b785

Compare with similar skills

CSV and Excel Merger next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

CSV and Excel Merger compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
CSV and Excel Merger this skillOneWave-AI/claude-skills336—~1.6kAutomated safety check: PassMIT
XLSX Spreadsheet ToolkitXiaomiMiMo/MiMo-Code14k—~2.9kAutomated safety check: PassApache-2.0
Excel Spreadsheet Creation and Editinganthropics/skills180k4 repos~2.1kAutomated safety check: PassProprietary
Excel Spreadsheet Builderagentscope-ai/QwenPaw36k—~1.8kAutomated safety check: PassProprietary
Sn Da Excel WorkflowMichaelYang-lyx/AIDABench1111 repos~2.5kAutomated safety check: PassNone
Codebookbrycewang-stanford/Auto-Empirical-Research-Skills4.6k—~527Automated safety check: NotesCustom licence

Similar skills

  • XLSX Spreadsheet Toolkit

    XiaomiMiMo/MiMo-Code

    Builds, edits, cleans, recalculates and reads Excel workbooks and CSV files with openpyxl and pandas, plus LibreOffice for recalculation and PDF export.

    14k GitHub stars~2.9k tokensUpdated today
    Documents & OfficeAuto-check passed
  • Official

    Creates, edits and analyzes spreadsheets (.xlsx, .xlsm, .csv, .tsv) with openpyxl and pandas, writing live formulas and recalculating to confirm zero formula errors.

    180k GitHub starsUsed in 4 repos~2.1k tokens
    Documents & OfficeAuto-check passed
  • Excel Spreadsheet Builder

    agentscope-ai/QwenPaw

    Creates, edits, cleans and analyzes Excel and CSV files with openpyxl and pandas, recalculating formulas through LibreOffice so files are delivered without formula errors.

    36k GitHub stars~1.8k tokensUpdated today
    Documents & OfficeAuto-check passed
  • Sn Da Excel Workflow

    MichaelYang-lyx/AIDABench

    Excel 数据分析多步编排器。覆盖:(1) 读取多 Sheet Excel 文件并统计行数,(2) 大文件检测(≥10k 行自动 Parquet 优化),(3) 数据清洗(缺失值、文本标准化、无效字符),(4) 条件筛选与分类提取,(5) 跨 Sheet 统计聚合,(6) 导出 Excel/CSV 并提供下载链接。覆盖从数据读取到报告生成全流程,按步骤编排 capability 子…

    111 GitHub starsUsed in 1 repo~2.5k tokens
    Documents & OfficeAuto-check passed
  • Codebook

    brycewang-stanford/Auto-Empirical-Research-Skills

    Auto-generates a Markdown codebook from a dataset (CSV, DTA, Excel, Parquet) with types and summary statistics.

    4.6k GitHub stars~527 tokensUpdated 4 days ago
    Documents & OfficeAuto-check: notes
  • Excel Parser

    Harryoung/efka

    Smart Excel/CSV file parsing with intelligent routing based on file complexity analysis.

    104 GitHub stars~2.3k tokensUpdated 6 mo ago
    Documents & OfficeAuto-check passed

More from OneWave-AI/claude-skills

All 69 skills in this repo
  • CRM Data Cleanup

    OneWave-AI/claude-skills

    Finds duplicate and junk records in a CRM CSV export with fuzzy matching, normalizes fields and writes a reviewable merge plan plus import-ready files without touching the live CRM.

    336 GitHub stars~2.5k tokensUpdated 8 days ago
    Auto-check passed
  • Design Export Repair

    OneWave-AI/claude-skills

    Repairs broken decks and PDFs exported from Claude Design or similar AI deck generators: clipped text, wrong fonts and corrupted .pptx package structure.

    336 GitHub stars~2.6k tokensUpdated 8 days ago
    Auto-check passed
  • Bi Measure Builder

    OneWave-AI/claude-skills

    Writes, explains, debugs, and optimizes BI calculations - Power BI / Fabric DAX measures and calculated columns, Tableau calculated fields (FIXED/INCLUDE/EXCLUDE LOD expressions, table…

    336 GitHub stars~2.2k tokensUpdated 8 days ago
    Auto-check passed
  • Bookkeeping Close

    OneWave-AI/claude-skills

    Categorizes transactions, reconciles bank and card statements to the ledger, works a month-end checklist and prepares a close package, without ever forcing a balance.

    336 GitHub stars~2k tokensUpdated 8 days ago
    Auto-check passed
  • Sec Filing Puller

    OneWave-AI/claude-skills

    Pulls financial statement numbers for US public companies straight from SEC EDGAR's free official XBRL APIs (companyfacts, companyconcept, frames, submissions) into a cited table.

    336 GitHub stars~2.1k tokensUpdated 8 days ago
    Auto-check passed
  • Spreadsheet QA

    OneWave-AI/claude-skills

    Answers business questions about the user's own spreadsheet or data export (CSV, TSV, XLSX from a CRM, Shopify, Stripe, QuickBooks, ad platforms, HR or payroll systems) correctly and auditably.

    336 GitHub stars~2.2k tokensUpdated 8 days ago
    Auto-check passed

Questions about CSV and Excel Merger

What does CSV and Excel Merger do?

Combines CSV, TSV and Excel files into one verified table with pandas, by stacking or joining, mapping columns, normalizing keys and removing duplicates. py` that reports each file's encoding, delimiter, row count, headers, candidate keys and header overlap without changing anything. Excel workbooks are profiled sheet by sheet, and the agent confirms with you which sheets count.

When should I use CSV and Excel Merger?

CSV and Excel Merger fits situations like: stacking monthly exports into a single spreadsheet; consolidating contact or lead lists from several sources and removing duplicates; joining two sheets on an ID or email column, like a VLOOKUP; merging multi-sheet Excel workbooks whose column names do not match.

How do I install CSV and Excel Merger in Claude Code?

Run `npx skills add OneWave-AI/claude-skills --skill csv-excel-merger -a claude-code`. Or copy the skill folder (csv-excel-merger in OneWave-AI/claude-skills) into .claude/skills/csv-excel-merger in your project. Claude Code loads it when a task matches its description.

How do I install CSV and Excel Merger in Codex?

Run `npx skills add OneWave-AI/claude-skills --skill csv-excel-merger -a codex`. Or copy the skill folder (csv-excel-merger in OneWave-AI/claude-skills) into .agents/skills/csv-excel-merger in your project. Codex loads it when a task matches its description.

Can I use CSV and Excel Merger in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add OneWave-AI/claude-skills --skill csv-excel-merger -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/csv-excel-merger, .gemini/skills/csv-excel-merger, .github/skills/csv-excel-merger and .opencode/skills/csv-excel-merger in your project.

What does CSV and Excel Merger need to run?

Going by SKILL.md and its folder, CSV and Excel Merger needs Python for the scripts in its folder and the command-line tools its instructions call (python). Our summary lists: Python 3 with pandas.

Does CSV and Excel Merger access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is CSV and Excel Merger safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does CSV and Excel Merger use?

CSV and Excel Merger is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does CSV and Excel Merger use?

About 1.6k tokens (SKILL.md is roughly 6.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.3k tokens, read only when the agent opens those files.

What are the alternatives to CSV and Excel Merger?

Skills that share tags, products or a category with CSV and Excel Merger: XLSX Spreadsheet Toolkit (XiaomiMiMo/MiMo-Code, 14k stars), Excel Spreadsheet Creation and Editing (anthropics/skills, 180k stars), Excel Spreadsheet Builder (agentscope-ai/QwenPaw, 36k stars) and Sn Da Excel Workflow (MichaelYang-lyx/AIDABench, 111 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains CSV and Excel Merger?

OneWave-AI (a GitHub organization) maintains it in OneWave-AI/claude-skills, which has 336 GitHub stars. The repository holds 69 skills in this directory. The repository was last updated on October 2, 2026.

Source: OneWave-AI/claude-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.