Agent skill

Invalid Data Cleaning

by MichaelYang-lyx in MichaelYang-lyx/AIDABench

用于大规模Excel数据的预处理,通过统计总行数判断是否转换为Parquet格式以提升读写效率,并使用正则表达式清洗指定文本列(如仅保留中文字符),最后导出清洗后的文件并提供下载链接。

No licenceAuto-check passedData & Analytics

Install Invalid Data Cleaning

skills CLI
$ npx skills add MichaelYang-lyx/AIDABench --skill invalid-data-cleaning -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install MichaelYang-lyx/AIDABench invalid-data-cleaning --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/MichaelYang-lyx/AIDABench.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/sn-da-excel-workflow/capability/excel-data-cleaning/invalid-data-cleaning .claude/skills/invalid-data-cleaning && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
invalid-data-cleaning
GitHub stars
111
Used in
1 other repo
Token cost
~410 tokens
SKILL.md length
31 words
Files
1
Skills in repo
39
Repo updated
First seen
Licence
None found

At a glance

用于大规模Excel数据的预处理,通过统计总行数判断是否转换为Parquet格式以提升读写效率,并使用正则表达式清洗指定文本列(如仅保留中文字符),最后导出清洗后的文件并提供下载链接。

  • Tasks that involve Data cleaning
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • Tasks that involve Excel spreadsheets
  • Tasks that involve DataFrames

What it does

Invalid Data Cleaning is an agent skill from MichaelYang-lyx/AIDABench. 用于大规模Excel数据的预处理,通过统计总行数判断是否转换为Parquet格式以提升读写效率,并使用正则表达式清洗指定文本列(如仅保留中文字符),最后导出清洗后的文件并提供下载链接。

Its SKILL.md is about 410 tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Data & Analytics, covering Data cleaning, Excel spreadsheets and DataFrames. It works with Microsoft Excel. The repository describes itself as: Code for paper AIDABench: AI Data Analytics Benchmark.

When your agent uses it

  • Tasks that involve Data cleaning
  • Tasks that involve Excel spreadsheets
  • Tasks that involve DataFrames

Example prompts

  • “/invalid-data-cleaning”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit 6dd4206. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Invalid Data Cleaning loads about 410 tokens when it runs. Until then it costs about 29 tokens; SKILL.md has 31 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~29
When it runs · the whole SKILL.md, loaded when a task matches
~410

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

Without a licence we can't republish the file, so here is its outline and opening line. It has 31 words (~410 tokens).

“Step1 根据总行数判断是否数据量过大,若满足条件,则将 Excel 文件转换为 Parquet 格式提升读写效率,再读取数据进行后续分析。”

— opening of SKILL.md by MichaelYang-lyx
name
invalid-data-cleaning

Read the full SKILL.md on GitHub

Files

Just SKILL.md in skills/sn-da-excel-workflow/capability/excel-data-cleaning/invalid-data-cleaning of MichaelYang-lyx/AIDABench.

Open the folder on GitHubat commit 6dd4206

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in MichaelYang-lyx/AIDABench, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Invalid Data Cleaning next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Invalid Data Cleaning compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Invalid Data Cleaning this skillMichaelYang-lyx/AIDABench1111 repos~410Automated safety check: PassNone
Data Cleaningericrisco/rsc-harness167—~3.6kAutomated safety check: PassMIT
CSV and Excel MergerOneWave-AI/claude-skills323—~1.6kAutomated safety check: PassMIT
Dataset Quality Auditzebbern/claude-code-guide4.7k—~996Automated safety check: PassMIT
Clean DataAperivue/medsci-skills329—~2kAutomated safety check: PassMIT
Clean Dataexplorium-ai/gtm-skills163—~2kAutomated safety check: PassMIT

Similar skills

  • Data Cleaning

    ericrisco/rsc-harness

    A skill your agent uses when a raw table is too dirty to trust — nulls, sentinels, duplicate rows, category sprawl, mixed types, bad dates — and you need a re-runnable clean() plus a schema gate…

    167 GitHub stars~3.6k tokensUpdated today
    Data & AnalyticsAuto-check passed
  • CSV and Excel Merger

    OneWave-AI/claude-skills

    Combines CSV, TSV and Excel files into one verified table with pandas, by stacking or joining, mapping columns, normalizing keys and removing duplicates.

    323 GitHub stars~1.6k tokensUpdated 6 days ago
    Documents & OfficeAuto-check passed
  • Dataset Quality Audit

    zebbern/claude-code-guide

    Run comprehensive quality checks on tabular data (CSV/Excel/TSV/JSON), detecting missing values, duplicates, outliers, format issues, and type inconsistencies to produce an overall score, grade, and…

    4.7k GitHub stars~996 tokensUpdated today
    Data & AnalyticsAuto-check passed
  • Clean Data

    Aperivue/medsci-skills

    A skill your agent uses when a clinical CSV/Excel dataset needs profiling and cleaning before analysis (missing values, outliers, duplicates, type mismatches).

    329 GitHub stars~2k tokensUpdated 2 days ago
    Data & AnalyticsAuto-check passed
  • Clean Data

    explorium-ai/gtm-skills

    Data cleaning, entity matching, and deduplication skill for Claude Code and Codex: triage, standardize, and validate a CSV, Excel, or JSON list of B2B companies or contacts before enrichment.

    163 GitHub stars~2k tokensUpdated 13 days ago
    Data & AnalyticsAuto-check passed
  • Visual Skills

    npc-live/clawfirm

    A skill your agent uses whenever the user provides data (CSV, JSON, table, pasted numbers, or any structured dataset) and expects a visual output — even if they don't say 'chart' or 'visualize'.

    156 GitHub stars~7.4k tokensUpdated 3 mo ago
    Data & AnalyticsAuto-check passed

More from MichaelYang-lyx/AIDABench

All 39 skills in this repo
  • Sn Da Image Caption

    MichaelYang-lyx/AIDABench

    图片理解与数据提取 skill。当图片文件(.png/.jpg/.jpeg/.gif/.webp/.bmp)是主要输入且用户需要理解、提取数据或分析图片内容时使用。提供预配置的 caption 脚本(scripts/caption.py),通过 vision 模型将图片转为文本描述,无需额外配置 API Key。覆盖:(1) 通过 scripts/caption.py…

    111 GitHub starsUsed in 1 repo~2k tokens
    Auto-check passed
  • 对Excel数据进行自定义分类统计、交叉分析与可视化,并基于多维度指标(如文本长度、术语密度、正则匹配等)进行综合评分与分级,适用于多类别数据分布统计及文本内容难度/质量评估场景。

    111 GitHub starsUsed in 2 repos~1.6k tokens
    Auto-check passed
  • Excel Bar Chart Visualization

    MichaelYang-lyx/AIDABench

    读取多工作表Excel文件,自动处理合并单元格与数据清洗,进行交叉分组统计并生成带总计行的结果表,最后绘制支持中英文字体的美化柱状图,适用于多维度数据汇总与可视化分析。

    111 GitHub starsUsed in 1 repo~991 tokens
    Auto-check passed
  • 执行全面的异常值检测与数据质量评估,利用 IQR 方法识别异常值并结合偏度、峰度分析数据分布特征,适用于非正态分布数据的预处理阶段。

    111 GitHub starsUsed in 1 repo~1k tokens
    Auto-check passed
  • Sn Da Excel Workflow

    MichaelYang-lyx/AIDABench

    Excel 数据分析多步编排器。覆盖:(1) 读取多 Sheet Excel 文件并统计行数,(2) 大文件检测(≥10k 行自动 Parquet 优化),(3) 数据清洗(缺失值、文本标准化、无效字符),(4) 条件筛选与分类提取,(5) 跨 Sheet 统计聚合,(6) 导出 Excel/CSV 并提供下载链接。覆盖从数据读取到报告生成全流程,按步骤编排 capability 子…

    111 GitHub starsUsed in 1 repo~2.5k tokens
    Auto-check passed
  • Sn Da Large File Analysis

    MichaelYang-lyx/AIDABench

    万行以上 Excel 数据集的高性能分析引擎。提供 openpyxl readonly 流式读取(iterrows 支持 10 万行以上)、Parquet 转换加速、内存优化、分块处理和大文件写入模式。遇到以下任一情况就主动使用本 skill:①数据行数 ≥ 10k(由 sn-da-excel-workflow 的行数评估步骤触发);②用户出现触发词:大文件 / 大数据量 / 性能优化 /…

    111 GitHub starsUsed in 1 repo~3.1k tokens
    Auto-check passed

Works with

Questions about Invalid Data Cleaning

What does Invalid Data Cleaning do?

用于大规模Excel数据的预处理,通过统计总行数判断是否转换为Parquet格式以提升读写效率,并使用正则表达式清洗指定文本列(如仅保留中文字符),最后导出清洗后的文件并提供下载链接。. Invalid Data Cleaning is an agent skill from MichaelYang-lyx/AIDABench.

When should I use Invalid Data Cleaning?

Invalid Data Cleaning fits situations like: tasks that involve Data cleaning; tasks that involve Excel spreadsheets; tasks that involve DataFrames.

How do I install Invalid Data Cleaning in Claude Code?

Run `npx skills add MichaelYang-lyx/AIDABench --skill invalid-data-cleaning -a claude-code`. Or copy the skill folder (skills/sn-da-excel-workflow/capability/excel-data-cleaning/invalid-data-cleaning in MichaelYang-lyx/AIDABench) into .claude/skills/invalid-data-cleaning in your project. Claude Code loads it when a task matches its description.

How do I install Invalid Data Cleaning in Codex?

Run `npx skills add MichaelYang-lyx/AIDABench --skill invalid-data-cleaning -a codex`. Or copy the skill folder (skills/sn-da-excel-workflow/capability/excel-data-cleaning/invalid-data-cleaning in MichaelYang-lyx/AIDABench) into .agents/skills/invalid-data-cleaning in your project. Codex loads it when a task matches its description.

Can I use Invalid Data Cleaning in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add MichaelYang-lyx/AIDABench --skill invalid-data-cleaning -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/invalid-data-cleaning, .gemini/skills/invalid-data-cleaning, .github/skills/invalid-data-cleaning and .opencode/skills/invalid-data-cleaning in your project.

What does Invalid Data Cleaning need to run?

SKILL.md names no scripts, command-line tools or credentials: Invalid Data Cleaning is instructions for the agent only. Our summary lists: Python 3.

Does Invalid Data Cleaning access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Invalid Data Cleaning safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Invalid Data Cleaning use?

No licence was found for Invalid Data Cleaning or its repository. Without one, default copyright applies: ask the author before reusing or redistributing it.

How many tokens does Invalid Data Cleaning use?

About 410 tokens (SKILL.md is roughly 1.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Invalid Data Cleaning?

Skills that share tags, products or a category with Invalid Data Cleaning: Data Cleaning (ericrisco/rsc-harness, 167 stars), CSV and Excel Merger (OneWave-AI/claude-skills, 323 stars), Dataset Quality Audit (zebbern/claude-code-guide, 4.7k stars) and Clean Data (Aperivue/medsci-skills, 329 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Invalid Data Cleaning?

MichaelYang-lyx (a GitHub user) maintains it in MichaelYang-lyx/AIDABench, which has 111 GitHub stars. The repository holds 39 skills in this directory. The repository was last updated on September 28, 2026.

Source: MichaelYang-lyx/AIDABench on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.