Agent skill

CSV Data Cleaner

by FerroxLabs in FerroxLabs/wayland

Quick techniques for cleaning CSV files - deduplication, normalization, encoding fixes, merging, and common data quality repairs using command-line tools and scripting.

Apache-2.0Auto-check passedData & Analytics

Install CSV Data Cleaner

skills CLI
$ npx skills add FerroxLabs/wayland --skill csv-data-cleaner -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install FerroxLabs/wayland csv-data-cleaner --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/FerroxLabs/wayland.git skills-src && mkdir -p .claude/skills && cp -r skills-src/src/process/resources/skills-library/bodies/skills/data-analysis/csv-data-cleaner .claude/skills/csv-data-cleaner && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
csv-data-cleaner
GitHub stars
608
Token cost
~2.5k tokens
SKILL.md length
367 words
Files
1
Skills in repo
1,194
Repo updated
First seen
Licence
Apache-2.0

At a glance

Quick techniques for cleaning CSV files - deduplication, normalization, encoding fixes, merging, and common data quality repairs using command-line tools and scripting.

  • Works in 5 steps: Gather information. Ask the user… → Analyze context. Review the information… → Develop recommendations. Apply domain… → …
  • The user asks about csv data cleaner
  • SKILL.md covers When to Use, Quick Diagnosis, Encoding Fixes and Deduplication, plus 9 more sections
  • Calls python3

What it does

CSV Data Cleaner is an agent skill from FerroxLabs/wayland. Quick techniques for cleaning CSV files - deduplication, normalization, encoding fixes, merging, and common data quality repairs using command-line tools and scripting. Use when the user asks about csv data cleaner, related techniques, best practices, or needs guidance in this domain. Do NOT use when the request is outside the scope of csv data cleaner or requires a different specialized skill.

Its SKILL.md is about 2.5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Data & Analytics, covering CSV and tabular files, Data cleaning and Database schema design. The repository describes itself as: Wayland - The AI Agent That Perceives. Reasons. Acts. Evolves. The licence is Apache-2.0.

When your agent uses it

  • The user asks about csv data cleaner
  • Related techniques
  • Needs guidance in this domain
  • The request is outside the scope of csv data cleaner

Example prompts

  • “/csv-data-cleaner”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Gather information. Ask the user clarifying questions to understand their specific situation, goals, and constraints
  2. Analyze context. Review the information provided and identify key factors relevant to csv data cleaner
  3. Develop recommendations. Apply domain expertise to create actionable guidance tailored to the user's needs
  4. Present structured output. Deliver findings in the output format below with clear next steps
  5. Address follow-ups. Answer additional questions and refine recommendations based on feedback

What it can do on your machine

Read from SKILL.md and the folder at commit 4c030c7. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

CSV Data Cleaner loads about 2.5k tokens when it runs. Until then it costs about 104 tokens; SKILL.md has 367 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~104
When it runs · the whole SKILL.md, loaded when a task matches
~2.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from FerroxLabs/wayland at commit 4c030c7, republished under its Apache-2.0 licence (© FerroxLabs). 367 words, ~2,473 tokens.

Download SKILL.mdSave it as .claude/skills/csv-data-cleaner/SKILL.md (or your agent's skills folder).
name
csv-data-cleaner
description
Quick techniques for cleaning CSV files - deduplication, normalization, encoding fixes, merging, and common data quality repairs using command-line tools and scripting. Use when the user asks about csv data cleaner, related techniques, best practices, or needs guidance in this domain. Do NOT use when the request is outside the scope of csv data cleaner or requires a different specialized skill.
license
Apache-2.0
metadata.author
foundry-skills
metadata.version
1.0.0
metadata.tags
quickstart data-science python email
metadata.category
data-analysis
metadata.subcategory
statistics-modeling
metadata.disclaimer
none
metadata.difficulty
intermediate

CSV Data Cleaner

You are a data cleaning specialist. Help the user fix messy CSV files quickly using the most appropriate tool. Provide exact commands and scripts. Prioritize one-liners for simple tasks, scripts for complex ones.

When to Use

Use this skill when:

  • User asks about csv data cleaner techniques or best practices
  • User needs guidance on csv data cleaner concepts
  • User wants to implement or improve their approach to csv data cleaner

Do NOT use when:

  • The request falls outside the scope of csv data cleaner
  • User needs a different specialized skill for their specific situation
  • The topic requires professional consultation beyond general guidance

Quick Diagnosis

shell
# Preview file structure
head -5 data.csv

# Count rows (excluding header)
wc -l data.csv

# Check encoding
file -i data.csv                   # Linux/Mac
# Windows PowerShell:
[System.IO.File]::ReadAllBytes("data.csv")[0..2]   # check BOM

# Count columns (assuming comma delimiter)
head -1 data.csv | awk -F',' '{print NF}'

# Check for inconsistent column counts
awk -F',' '{print NF}' data.csv | sort | uniq -c

Encoding Fixes

shell
# Convert to UTF-8 from unknown encoding
iconv -f ISO-8859-1 -t UTF-8 input.csv > output.csv

# Remove BOM (byte order mark)
sed '1s/^\xEF\xBB\xBF//' input.csv > output.csv

# Fix Windows line endings (CRLF -> LF)
sed 's/\r$//' input.csv > output.csv
# Or:
dos2unix input.csv

# Python (handles any encoding)
python3 -c "
import csv, codecs
with codecs.open('input.csv','r','latin-1') as f:
    data = f.read()
with codecs.open('output.csv','w','utf-8') as f:
    f.write(data)
"

Deduplication

shell
# Remove exact duplicate rows (keeps first occurrence)
awk '!seen[$0]++' data.csv > deduped.csv

# Remove duplicates based on specific column (column 1)
awk -F',' '!seen[$1]++' data.csv > deduped.csv

# Python: deduplicate with more control
python3 << 'EOF'
import csv
seen = set()
with open('data.csv') as f, open('deduped.csv', 'w', newline='') as out:
    reader = csv.reader(f)
    writer = csv.writer(out)
    header = next(reader)
    writer.writerow(header)
    key_col = 0  # column index to deduplicate on
    for row in reader:
        key = row[key_col].strip().lower()  # normalize before checking
        if key not in seen:
            seen.add(key)
            writer.writerow(row)
EOF

Normalization

Whitespace Cleanup
shell
# Trim whitespace from all fields
python3 -c "
import csv, sys
reader = csv.reader(open('data.csv'))
writer = csv.writer(sys.stdout)
for row in reader:
    writer.writerow([cell.strip() for cell in row])
" > cleaned.csv
Case Normalization
python
# Python: normalize specific columns
import csv
with open('data.csv') as f, open('out.csv', 'w', newline='') as out:
    reader = csv.DictReader(f)
    writer = csv.DictWriter(out, fieldnames=reader.fieldnames)
    writer.writeheader()
    for row in reader:
        row['email'] = row['email'].strip().lower()
        row['name'] = row['name'].strip().title()
        row['state'] = row['state'].strip().upper()
        writer.writerow(row)
Date Normalization
python
from datetime import datetime
import csv

formats_to_try = ['%m/%d/%Y', '%d-%m-%Y', '%Y-%m-%d', '%B %d, %Y', '%m/%d/%y']

def normalize_date(value, target_format='%Y-%m-%d'):
    for fmt in formats_to_try:
        try:
            return datetime.strptime(value.strip(), fmt).strftime(target_format)
        except ValueError:
            continue
    return value  # return original if no format matches

# Apply to column index 3
with open('data.csv') as f, open('out.csv', 'w', newline='') as out:
    reader = csv.reader(f)
    writer = csv.writer(out)
    writer.writerow(next(reader))  # header
    for row in reader:
        row[3] = normalize_date(row[3])
        writer.writerow(row)
Phone Number Normalization
python
import re, csv

def normalize_phone(phone):
    digits = re.sub(r'\D', '', phone)
    if len(digits) == 11 and digits[0] == '1':
        digits = digits[1:]
    if len(digits) == 10:
        return f"({digits[:3]}) {digits[3:6]}-{digits[6:]}"
    return phone  # return original if unexpected format

Merging CSV Files

shell
# Stack files with same columns (skip header on 2nd+ files)
head -1 file1.csv > merged.csv
tail -n +2 -q file1.csv file2.csv file3.csv >> merged.csv

# Python: merge/join on a key column
python3 << 'EOF'
import csv

# Load lookup data
lookup = {}
with open('lookup.csv') as f:
    for row in csv.DictReader(f):
        lookup[row['id']] = row

# Merge with main data
with open('main.csv') as f, open('merged.csv', 'w', newline='') as out:
    reader = csv.DictReader(f)
    extra_fields = ['extra_col1', 'extra_col2']  # fields from lookup
    writer = csv.DictWriter(out, fieldnames=reader.fieldnames + extra_fields)
    writer.writeheader()
    for row in reader:
        match = lookup.get(row['id'], {})
        for field in extra_fields:
            row[field] = match.get(field, '')
        writer.writerow(row)
EOF

Common Data Quality Fixes

Remove Empty Rows
shell
awk -F',' 'NF && $0 !~ /^[,\s]*$/' data.csv > cleaned.csv
Fill Missing Values
python
import csv
default_values = {'status': 'unknown', 'count': '0', 'category': 'other'}

with open('data.csv') as f, open('filled.csv', 'w', newline='') as out:
    reader = csv.DictReader(f)
    writer = csv.DictWriter(out, fieldnames=reader.fieldnames)
    writer.writeheader()
    for row in reader:
        for field, default in default_values.items():
            if not row.get(field, '').strip():
                row[field] = default
        writer.writerow(row)
Split Column into Multiple
python
# Split "Full Name" into "First" and "Last"
import csv
with open('data.csv') as f, open('out.csv', 'w', newline='') as out:
    reader = csv.DictReader(f)
    fields = [fn for fn in reader.fieldnames if fn != 'full_name'] + ['first_name', 'last_name']
    writer = csv.DictWriter(out, fieldnames=fields)
    writer.writeheader()
    for row in reader:
        parts = row.pop('full_name', '').strip().split(None, 1)
        row['first_name'] = parts[0] if parts else ''
        row['last_name'] = parts[1] if len(parts) > 1 else ''
        writer.writerow(row)

Quick Validation Checks

shell
# Find rows with wrong column count
expected=5
awk -F',' -v exp="$expected" 'NF != exp {print NR": "NF" cols - "$0}' data.csv

# Find rows with empty required fields (column 1 and 3)
awk -F',' '$1=="" || $3=="" {print NR": "$0}' data.csv

# Summary statistics for a numeric column (column 2)
awk -F',' 'NR>1 {sum+=$2; count++; if($2>max||NR==2)max=$2; if($2<min||NR==2)min=$2} END {print "count:"count, "sum:"sum, "avg:"sum/count, "min:"min, "max:"max}' data.csv

Tool Recommendations

TaskBest Tool
Simple column extractioncut -d',' -f1,3 data.csv
Complex transformsPython csv module
Large files (GB+)csvkit, xsv, or miller
Quick explorationcsvlook (from csvkit)
SQL on CSVcsvsql or q
Excel interoppandas or openpyxl
shell
# csvkit essentials
install the package via pip csvkit
csvlook data.csv              # pretty print
csvstat data.csv              # column statistics
csvsort -c 2 data.csv         # sort by column 2
csvgrep -c 3 -m "value" data.csv  # filter rows
csvjoin -c id file1.csv file2.csv  # join files

Process

  1. Gather information. Ask the user clarifying questions to understand their specific situation, goals, and constraints
  2. Analyze context. Review the information provided and identify key factors relevant to csv data cleaner
  3. Develop recommendations. Apply domain expertise to create actionable guidance tailored to the user's needs
  4. Present structured output. Deliver findings in the output format below with clear next steps
  5. Address follow-ups. Answer additional questions and refine recommendations based on feedback
Show full SKILL.md (112 more words)Show less

Output Format

template
## Csv Data Cleaner Analysis

### Assessment
[Key findings and observations]

### Recommendations
1. [Primary recommendation]
2. [Secondary recommendation]
3. [Additional suggestions]

### Action Items
- [ ] [First action step]
- [ ] [Second action step]
- [ ] [Follow-up task]

Edge Cases

  • Incomplete information: Ask clarifying questions before proceeding with recommendations
  • Conflicting requirements: Prioritize the most critical constraint and note trade-offs
  • Out of scope requests: Redirect to appropriate specialized skill or professional resource
  • Beginner vs advanced: Adjust depth and terminology based on user's experience level

Example

Input: "Help me with csv data cleaner for my current situation"

Output:

Based on your situation, here is a structured approach to csv data cleaner:

  1. Assessment: Evaluate your current state and identify key areas for improvement
  2. Strategy: Develop a targeted plan based on best practices
  3. Implementation: Execute the plan with specific, measurable steps
  4. Review: Monitor progress and adjust as needed

© FerroxLabs, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in src/process/resources/skills-library/bodies/skills/data-analysis/csv-data-cleaner of FerroxLabs/wayland.

Open the folder on GitHubat commit 4c030c7

Compare with similar skills

CSV Data Cleaner next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

CSV Data Cleaner compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
CSV Data Cleaner this skillFerroxLabs/wayland608—~2.5kAutomated safety check: PassApache-2.0
Metabolomics QuantificationTianGzlab/OmicsClaw161—~1.1kAutomated safety check: PassMIT
Data Table AnalysisNVIDIA-AI-Blueprints/deep-researcher-agent883—~2.5kAutomated safety check: PassApache-2.0
Dataset Quality Auditzebbern/claude-code-guide4.6k—~996Automated safety check: PassMIT
Portaljs Check Data Qualitydatopian/portaljs2.4k—~1.5kAutomated safety check: PassMIT
Education Cloud Course Catalog Migrateforcedotcom/sf-skills1.1k—~5.4kAutomated safety check: PassApache-2.0

Similar skills

  • Metabolomics Quantification

    TianGzlab/OmicsClaw

    Load when imputing missing values (min / median / KNN) and normalising (TIC / median / log) a feature × sample metabolomics CSV.

    161 GitHub stars~1.1k tokensUpdated 2 mo ago
    Data & AnalyticsAuto-check passed
  • Data Table Analysis

    NVIDIA-AI-Blueprints/deep-researcher-agent

    A skill your agent uses for converting researched facts or user-provided data into structured tables by writing code, then running Python/pandas calculations in the job-scoped sandbox.

    883 GitHub stars~2.5k tokensUpdated yesterday
    Data & AnalyticsAuto-check passed
  • Dataset Quality Audit

    zebbern/claude-code-guide

    Run comprehensive quality checks on tabular data (CSV/Excel/TSV/JSON), detecting missing values, duplicates, outliers, format issues, and type inconsistencies to produce an overall score, grade, and…

    4.6k GitHub stars~996 tokensUpdated yesterday
    Data & AnalyticsAuto-check passed
  • Audit a local or remote tabular file (CSV/TSV) for common data quality issues — schema, nulls, types, duplicates.

    2.4k GitHub stars~1.5k tokensUpdated 1 mo ago
    Data & AnalyticsAuto-check passed
  • A skill your agent uses to migrate course catalog data from external sources (CSV, PDF, website) and bulk-create Learning and LearningCourse records in Education Cloud.

    1.1k GitHub stars~5.4k tokensUpdated 5 days ago
    Data & AnalyticsAuto-check passed
  • CSV Processing

    benchflow-ai/skillsbench

    A skill your agent uses when reading sensor data from CSV files, writing simulation results to CSV, processing time-series data with pandas, or handling missing values in datasets.

    1.8k GitHub stars~455 tokensUpdated 2 mo ago
    Data & AnalyticsAuto-check passed

More from FerroxLabs/wayland

All 1,194 skills in this repo
  • Star Office Helper

    FerroxLabs/wayland

    Install, start, connect, and troubleshoot visualization companion projects for Aion/OpenClaw, with Star-Office-UI as the default recommendation.

    608 GitHub stars~2.2k tokensUpdated yesterday
    Auto-check: notes
  • Openclaw Setup

    FerroxLabs/wayland

    OpenClaw usage expert: Helps you install, deploy, configure, and use OpenClaw personal AI assistant.

    608 GitHub stars~1.9k tokensUpdated yesterday
    Auto-check passed
  • Tvcontrol Setup

    FerroxLabs/wayland

    Set up TVControl end to end: install the connector, start TradingView Desktop with its control port open, load a watchlist export, add the indicators they use, and leave a working chart.

    608 GitHub stars~5.7k tokensUpdated yesterday
    Auto-check passed
  • Ab Testing Specialist

    FerroxLabs/wayland

    End-to-end guide for designing, running, and analyzing A/B tests including experiment design, statistical significance, sample size calculation, common pitfalls, and advanced testing patterns.

    608 GitHub stars~3.7k tokensUpdated yesterday
    Auto-check passed
  • Academic Writer

    FerroxLabs/wayland

    Complete academic writing guide covering thesis and dissertation structure, journal article format using IMRaD, literature review methodology, citation management, the peer review process, and…

    608 GitHub stars~4.5k tokensUpdated yesterday
    Auto-check passed
  • Accessibility Auditor

    FerroxLabs/wayland

    Web accessibility expertise covering WCAG 2.2 conformance, audit methodology, ARIA patterns, keyboard navigation, screen reader testing, focus management, form accessibility, and automated vs manual…

    608 GitHub stars~4.1k tokensUpdated yesterday
    Auto-check passed

Questions about CSV Data Cleaner

What does CSV Data Cleaner do?

Quick techniques for cleaning CSV files - deduplication, normalization, encoding fixes, merging, and common data quality repairs using command-line tools and scripting. CSV Data Cleaner is an agent skill from FerroxLabs/wayland. Quick techniques for cleaning CSV files - deduplication, normalization, encoding fixes, merging, and common data quality repairs using command-line tools and scripting.

When should I use CSV Data Cleaner?

CSV Data Cleaner fits situations like: the user asks about csv data cleaner; related techniques; needs guidance in this domain; the request is outside the scope of csv data cleaner.

How do I install CSV Data Cleaner in Claude Code?

Run `npx skills add FerroxLabs/wayland --skill csv-data-cleaner -a claude-code`. Or copy the skill folder (src/process/resources/skills-library/bodies/skills/data-analysis/csv-data-cleaner in FerroxLabs/wayland) into .claude/skills/csv-data-cleaner in your project. Claude Code loads it when a task matches its description.

How do I install CSV Data Cleaner in Codex?

Run `npx skills add FerroxLabs/wayland --skill csv-data-cleaner -a codex`. Or copy the skill folder (src/process/resources/skills-library/bodies/skills/data-analysis/csv-data-cleaner in FerroxLabs/wayland) into .agents/skills/csv-data-cleaner in your project. Codex loads it when a task matches its description.

Can I use CSV Data Cleaner in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add FerroxLabs/wayland --skill csv-data-cleaner -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/csv-data-cleaner, .gemini/skills/csv-data-cleaner, .github/skills/csv-data-cleaner and .opencode/skills/csv-data-cleaner in your project.

What does CSV Data Cleaner need to run?

Going by SKILL.md and its folder, CSV Data Cleaner needs the command-line tools its instructions call (python3). Our summary lists: Python 3.

Does CSV Data Cleaner access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is CSV Data Cleaner safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does CSV Data Cleaner use?

CSV Data Cleaner is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does CSV Data Cleaner use?

About 2.5k tokens (SKILL.md is roughly 9.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to CSV Data Cleaner?

Skills that share tags, products or a category with CSV Data Cleaner: Metabolomics Quantification (TianGzlab/OmicsClaw, 161 stars), Data Table Analysis (NVIDIA-AI-Blueprints/deep-researcher-agent, 883 stars), Dataset Quality Audit (zebbern/claude-code-guide, 4.6k stars) and Portaljs Check Data Quality (datopian/portaljs, 2.4k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains CSV Data Cleaner?

FerroxLabs (a GitHub user) maintains it in FerroxLabs/wayland, which has 608 GitHub stars. The repository holds 1,194 skills in this directory. The repository was last updated on October 6, 2026.

Source: FerroxLabs/wayland on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.