Metabolomics Quantification
TianGzlab/OmicsClaw
Load when imputing missing values (min / median / KNN) and normalising (TIC / median / log) a feature × sample metabolomics CSV.
Quick techniques for cleaning CSV files - deduplication, normalization, encoding fixes, merging, and common data quality repairs using command-line tools and scripting.
$ npx skills add FerroxLabs/wayland --skill csv-data-cleaner -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install FerroxLabs/wayland csv-data-cleaner --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/FerroxLabs/wayland.git skills-src && mkdir -p .claude/skills && cp -r skills-src/src/process/resources/skills-library/bodies/skills/data-analysis/csv-data-cleaner .claude/skills/csv-data-cleaner && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "csv-data-cleaner" agent skill from https://github.com/FerroxLabs/wayland/tree/main/src/process/resources/skills-library/bodies/skills/data-analysis/csv-data-cleaner into .claude/skills/csv-data-cleaner/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "csv-data-cleaner", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/FerroxLabs/wayland/tree/main/src/process/resources/skills-library/bodies/skills/data-analysis/csv-data-cleanerType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add FerroxLabs/wayland --skill csv-data-cleaner -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install FerroxLabs/wayland csv-data-cleaner --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/FerroxLabs/wayland.git skills-src && mkdir -p .agents/skills && cp -r skills-src/src/process/resources/skills-library/bodies/skills/data-analysis/csv-data-cleaner .agents/skills/csv-data-cleaner && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "csv-data-cleaner" agent skill from https://github.com/FerroxLabs/wayland/tree/main/src/process/resources/skills-library/bodies/skills/data-analysis/csv-data-cleaner into .agents/skills/csv-data-cleaner/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "csv-data-cleaner", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add FerroxLabs/wayland --skill csv-data-cleaner -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install FerroxLabs/wayland csv-data-cleaner --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/FerroxLabs/wayland.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/src/process/resources/skills-library/bodies/skills/data-analysis/csv-data-cleaner .cursor/skills/csv-data-cleaner && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "csv-data-cleaner" agent skill from https://github.com/FerroxLabs/wayland/tree/main/src/process/resources/skills-library/bodies/skills/data-analysis/csv-data-cleaner into .cursor/skills/csv-data-cleaner/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "csv-data-cleaner", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/FerroxLabs/wayland.git --path src/process/resources/skills-library/bodies/skills/data-analysis/csv-data-cleaner--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add FerroxLabs/wayland --skill csv-data-cleaner -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install FerroxLabs/wayland csv-data-cleaner --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/FerroxLabs/wayland.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/src/process/resources/skills-library/bodies/skills/data-analysis/csv-data-cleaner .gemini/skills/csv-data-cleaner && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "csv-data-cleaner" agent skill from https://github.com/FerroxLabs/wayland/tree/main/src/process/resources/skills-library/bodies/skills/data-analysis/csv-data-cleaner into .gemini/skills/csv-data-cleaner/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "csv-data-cleaner", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install FerroxLabs/wayland csv-data-cleanerInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add FerroxLabs/wayland --skill csv-data-cleaner -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/FerroxLabs/wayland.git skills-src && mkdir -p .github/skills && cp -r skills-src/src/process/resources/skills-library/bodies/skills/data-analysis/csv-data-cleaner .github/skills/csv-data-cleaner && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "csv-data-cleaner" agent skill from https://github.com/FerroxLabs/wayland/tree/main/src/process/resources/skills-library/bodies/skills/data-analysis/csv-data-cleaner into .github/skills/csv-data-cleaner/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "csv-data-cleaner", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add FerroxLabs/wayland --skill csv-data-cleaner -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install FerroxLabs/wayland csv-data-cleaner --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/FerroxLabs/wayland.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/src/process/resources/skills-library/bodies/skills/data-analysis/csv-data-cleaner .opencode/skills/csv-data-cleaner && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "csv-data-cleaner" agent skill from https://github.com/FerroxLabs/wayland/tree/main/src/process/resources/skills-library/bodies/skills/data-analysis/csv-data-cleaner into .opencode/skills/csv-data-cleaner/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "csv-data-cleaner", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
csv-data-cleanerQuick techniques for cleaning CSV files - deduplication, normalization, encoding fixes, merging, and common data quality repairs using command-line tools and scripting.
CSV Data Cleaner is an agent skill from FerroxLabs/wayland. Quick techniques for cleaning CSV files - deduplication, normalization, encoding fixes, merging, and common data quality repairs using command-line tools and scripting. Use when the user asks about csv data cleaner, related techniques, best practices, or needs guidance in this domain. Do NOT use when the request is outside the scope of csv data cleaner or requires a different specialized skill.
Its SKILL.md is about 2.5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Data & Analytics, covering CSV and tabular files, Data cleaning and Database schema design. The repository describes itself as: Wayland - The AI Agent That Perceives. Reasons. Acts. Evolves. The licence is Apache-2.0.
5 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 4c030c7. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
python3From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
CSV Data Cleaner loads about 2.5k tokens when it runs. Until then it costs about 104 tokens; SKILL.md has 367 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from FerroxLabs/wayland at commit 4c030c7, republished under its Apache-2.0 licence (© FerroxLabs). 367 words, ~2,473 tokens.
.claude/skills/csv-data-cleaner/SKILL.md (or your agent's skills folder).You are a data cleaning specialist. Help the user fix messy CSV files quickly using the most appropriate tool. Provide exact commands and scripts. Prioritize one-liners for simple tasks, scripts for complex ones.
Use this skill when:
Do NOT use when:
# Preview file structure
head -5 data.csv
# Count rows (excluding header)
wc -l data.csv
# Check encoding
file -i data.csv # Linux/Mac
# Windows PowerShell:
[System.IO.File]::ReadAllBytes("data.csv")[0..2] # check BOM
# Count columns (assuming comma delimiter)
head -1 data.csv | awk -F',' '{print NF}'
# Check for inconsistent column counts
awk -F',' '{print NF}' data.csv | sort | uniq -c# Convert to UTF-8 from unknown encoding
iconv -f ISO-8859-1 -t UTF-8 input.csv > output.csv
# Remove BOM (byte order mark)
sed '1s/^\xEF\xBB\xBF//' input.csv > output.csv
# Fix Windows line endings (CRLF -> LF)
sed 's/\r$//' input.csv > output.csv
# Or:
dos2unix input.csv
# Python (handles any encoding)
python3 -c "
import csv, codecs
with codecs.open('input.csv','r','latin-1') as f:
data = f.read()
with codecs.open('output.csv','w','utf-8') as f:
f.write(data)
"# Remove exact duplicate rows (keeps first occurrence)
awk '!seen[$0]++' data.csv > deduped.csv
# Remove duplicates based on specific column (column 1)
awk -F',' '!seen[$1]++' data.csv > deduped.csv
# Python: deduplicate with more control
python3 << 'EOF'
import csv
seen = set()
with open('data.csv') as f, open('deduped.csv', 'w', newline='') as out:
reader = csv.reader(f)
writer = csv.writer(out)
header = next(reader)
writer.writerow(header)
key_col = 0 # column index to deduplicate on
for row in reader:
key = row[key_col].strip().lower() # normalize before checking
if key not in seen:
seen.add(key)
writer.writerow(row)
EOF# Trim whitespace from all fields
python3 -c "
import csv, sys
reader = csv.reader(open('data.csv'))
writer = csv.writer(sys.stdout)
for row in reader:
writer.writerow([cell.strip() for cell in row])
" > cleaned.csv# Python: normalize specific columns
import csv
with open('data.csv') as f, open('out.csv', 'w', newline='') as out:
reader = csv.DictReader(f)
writer = csv.DictWriter(out, fieldnames=reader.fieldnames)
writer.writeheader()
for row in reader:
row['email'] = row['email'].strip().lower()
row['name'] = row['name'].strip().title()
row['state'] = row['state'].strip().upper()
writer.writerow(row)from datetime import datetime
import csv
formats_to_try = ['%m/%d/%Y', '%d-%m-%Y', '%Y-%m-%d', '%B %d, %Y', '%m/%d/%y']
def normalize_date(value, target_format='%Y-%m-%d'):
for fmt in formats_to_try:
try:
return datetime.strptime(value.strip(), fmt).strftime(target_format)
except ValueError:
continue
return value # return original if no format matches
# Apply to column index 3
with open('data.csv') as f, open('out.csv', 'w', newline='') as out:
reader = csv.reader(f)
writer = csv.writer(out)
writer.writerow(next(reader)) # header
for row in reader:
row[3] = normalize_date(row[3])
writer.writerow(row)import re, csv
def normalize_phone(phone):
digits = re.sub(r'\D', '', phone)
if len(digits) == 11 and digits[0] == '1':
digits = digits[1:]
if len(digits) == 10:
return f"({digits[:3]}) {digits[3:6]}-{digits[6:]}"
return phone # return original if unexpected format# Stack files with same columns (skip header on 2nd+ files)
head -1 file1.csv > merged.csv
tail -n +2 -q file1.csv file2.csv file3.csv >> merged.csv
# Python: merge/join on a key column
python3 << 'EOF'
import csv
# Load lookup data
lookup = {}
with open('lookup.csv') as f:
for row in csv.DictReader(f):
lookup[row['id']] = row
# Merge with main data
with open('main.csv') as f, open('merged.csv', 'w', newline='') as out:
reader = csv.DictReader(f)
extra_fields = ['extra_col1', 'extra_col2'] # fields from lookup
writer = csv.DictWriter(out, fieldnames=reader.fieldnames + extra_fields)
writer.writeheader()
for row in reader:
match = lookup.get(row['id'], {})
for field in extra_fields:
row[field] = match.get(field, '')
writer.writerow(row)
EOFawk -F',' 'NF && $0 !~ /^[,\s]*$/' data.csv > cleaned.csvimport csv
default_values = {'status': 'unknown', 'count': '0', 'category': 'other'}
with open('data.csv') as f, open('filled.csv', 'w', newline='') as out:
reader = csv.DictReader(f)
writer = csv.DictWriter(out, fieldnames=reader.fieldnames)
writer.writeheader()
for row in reader:
for field, default in default_values.items():
if not row.get(field, '').strip():
row[field] = default
writer.writerow(row)# Split "Full Name" into "First" and "Last"
import csv
with open('data.csv') as f, open('out.csv', 'w', newline='') as out:
reader = csv.DictReader(f)
fields = [fn for fn in reader.fieldnames if fn != 'full_name'] + ['first_name', 'last_name']
writer = csv.DictWriter(out, fieldnames=fields)
writer.writeheader()
for row in reader:
parts = row.pop('full_name', '').strip().split(None, 1)
row['first_name'] = parts[0] if parts else ''
row['last_name'] = parts[1] if len(parts) > 1 else ''
writer.writerow(row)# Find rows with wrong column count
expected=5
awk -F',' -v exp="$expected" 'NF != exp {print NR": "NF" cols - "$0}' data.csv
# Find rows with empty required fields (column 1 and 3)
awk -F',' '$1=="" || $3=="" {print NR": "$0}' data.csv
# Summary statistics for a numeric column (column 2)
awk -F',' 'NR>1 {sum+=$2; count++; if($2>max||NR==2)max=$2; if($2<min||NR==2)min=$2} END {print "count:"count, "sum:"sum, "avg:"sum/count, "min:"min, "max:"max}' data.csv| Task | Best Tool |
|---|---|
| Simple column extraction | cut -d',' -f1,3 data.csv |
| Complex transforms | Python csv module |
| Large files (GB+) | csvkit, xsv, or miller |
| Quick exploration | csvlook (from csvkit) |
| SQL on CSV | csvsql or q |
| Excel interop | pandas or openpyxl |
# csvkit essentials
install the package via pip csvkit
csvlook data.csv # pretty print
csvstat data.csv # column statistics
csvsort -c 2 data.csv # sort by column 2
csvgrep -c 3 -m "value" data.csv # filter rows
csvjoin -c id file1.csv file2.csv # join files## Csv Data Cleaner Analysis
### Assessment
[Key findings and observations]
### Recommendations
1. [Primary recommendation]
2. [Secondary recommendation]
3. [Additional suggestions]
### Action Items
- [ ] [First action step]
- [ ] [Second action step]
- [ ] [Follow-up task]Input: "Help me with csv data cleaner for my current situation"
Output:
Based on your situation, here is a structured approach to csv data cleaner:
© FerroxLabs, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in src/process/resources/skills-library/bodies/skills/data-analysis/csv-data-cleaner of FerroxLabs/wayland.
Open the folder on GitHubat commit 4c030c7
CSV Data Cleaner next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| CSV Data Cleaner this skillFerroxLabs/wayland | 608 | — | ~2.5k | Automated safety check: Pass | Apache-2.0 | |
| Metabolomics QuantificationTianGzlab/OmicsClaw | 161 | — | ~1.1k | Automated safety check: Pass | MIT | |
| Data Table AnalysisNVIDIA-AI-Blueprints/deep-researcher-agent | 883 | — | ~2.5k | Automated safety check: Pass | Apache-2.0 | |
| Dataset Quality Auditzebbern/claude-code-guide | 4.6k | — | ~996 | Automated safety check: Pass | MIT | |
| Portaljs Check Data Qualitydatopian/portaljs | 2.4k | — | ~1.5k | Automated safety check: Pass | MIT | |
| Education Cloud Course Catalog Migrateforcedotcom/sf-skills | 1.1k | — | ~5.4k | Automated safety check: Pass | Apache-2.0 |
TianGzlab/OmicsClaw
Load when imputing missing values (min / median / KNN) and normalising (TIC / median / log) a feature × sample metabolomics CSV.
NVIDIA-AI-Blueprints/deep-researcher-agent
A skill your agent uses for converting researched facts or user-provided data into structured tables by writing code, then running Python/pandas calculations in the job-scoped sandbox.
zebbern/claude-code-guide
Run comprehensive quality checks on tabular data (CSV/Excel/TSV/JSON), detecting missing values, duplicates, outliers, format issues, and type inconsistencies to produce an overall score, grade, and…
datopian/portaljs
Audit a local or remote tabular file (CSV/TSV) for common data quality issues — schema, nulls, types, duplicates.
forcedotcom/sf-skills
A skill your agent uses to migrate course catalog data from external sources (CSV, PDF, website) and bulk-create Learning and LearningCourse records in Education Cloud.
benchflow-ai/skillsbench
A skill your agent uses when reading sensor data from CSV files, writing simulation results to CSV, processing time-series data with pandas, or handling missing values in datasets.
FerroxLabs/wayland
Install, start, connect, and troubleshoot visualization companion projects for Aion/OpenClaw, with Star-Office-UI as the default recommendation.
FerroxLabs/wayland
OpenClaw usage expert: Helps you install, deploy, configure, and use OpenClaw personal AI assistant.
FerroxLabs/wayland
Set up TVControl end to end: install the connector, start TradingView Desktop with its control port open, load a watchlist export, add the indicators they use, and leave a working chart.
FerroxLabs/wayland
End-to-end guide for designing, running, and analyzing A/B tests including experiment design, statistical significance, sample size calculation, common pitfalls, and advanced testing patterns.
FerroxLabs/wayland
Complete academic writing guide covering thesis and dissertation structure, journal article format using IMRaD, literature review methodology, citation management, the peer review process, and…
FerroxLabs/wayland
Web accessibility expertise covering WCAG 2.2 conformance, audit methodology, ARIA patterns, keyboard navigation, screen reader testing, focus management, form accessibility, and automated vs manual…
Categories
Quick techniques for cleaning CSV files - deduplication, normalization, encoding fixes, merging, and common data quality repairs using command-line tools and scripting. CSV Data Cleaner is an agent skill from FerroxLabs/wayland. Quick techniques for cleaning CSV files - deduplication, normalization, encoding fixes, merging, and common data quality repairs using command-line tools and scripting.
CSV Data Cleaner fits situations like: the user asks about csv data cleaner; related techniques; needs guidance in this domain; the request is outside the scope of csv data cleaner.
Run `npx skills add FerroxLabs/wayland --skill csv-data-cleaner -a claude-code`. Or copy the skill folder (src/process/resources/skills-library/bodies/skills/data-analysis/csv-data-cleaner in FerroxLabs/wayland) into .claude/skills/csv-data-cleaner in your project. Claude Code loads it when a task matches its description.
Run `npx skills add FerroxLabs/wayland --skill csv-data-cleaner -a codex`. Or copy the skill folder (src/process/resources/skills-library/bodies/skills/data-analysis/csv-data-cleaner in FerroxLabs/wayland) into .agents/skills/csv-data-cleaner in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add FerroxLabs/wayland --skill csv-data-cleaner -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/csv-data-cleaner, .gemini/skills/csv-data-cleaner, .github/skills/csv-data-cleaner and .opencode/skills/csv-data-cleaner in your project.
Going by SKILL.md and its folder, CSV Data Cleaner needs the command-line tools its instructions call (python3). Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
CSV Data Cleaner is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.5k tokens (SKILL.md is roughly 9.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with CSV Data Cleaner: Metabolomics Quantification (TianGzlab/OmicsClaw, 161 stars), Data Table Analysis (NVIDIA-AI-Blueprints/deep-researcher-agent, 883 stars), Dataset Quality Audit (zebbern/claude-code-guide, 4.6k stars) and Portaljs Check Data Quality (datopian/portaljs, 2.4k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
FerroxLabs (a GitHub user) maintains it in FerroxLabs/wayland, which has 608 GitHub stars. The repository holds 1,194 skills in this directory. The repository was last updated on October 6, 2026.
Source: FerroxLabs/wayland on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.