Agent skill

Data Auditor Cleaner

by zhnnky329 in zhnnky329/MathModeling-skills

Map contest attachments to subquestions, audit and clean raw data, and emit one reusable data profile with quality, coverage, imbalance, concentration, and method-readiness evidence for downstream…

MITAuto-check passed

Install Data Auditor Cleaner

skills CLI
$ npx skills add zhnnky329/MathModeling-skills --skill data-auditor-cleaner -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install zhnnky329/MathModeling-skills data-auditor-cleaner --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/zhnnky329/MathModeling-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.codex/skills/data-auditor-cleaner .claude/skills/data-auditor-cleaner && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
data-auditor-cleaner
GitHub stars
1.1k
Token cost
~1.1k tokens
SKILL.md length
400 words
Files
1
Skills in repo
29
Repo updated
First seen
Licence
MIT

At a glance

Map contest attachments to subquestions, audit and clean raw data, and emit one reusable data profile with quality, coverage, imbalance, concentration, and method-readiness evidence for downstream…

  • Works in 6 steps: Map attachments before cleaning. → Preserve raw data. → Audit structure and semantics. → …
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Data Auditor Cleaner is an agent skill from zhnnky329/MathModeling-skills. Map contest attachments to subquestions, audit and clean raw data, and emit one reusable data profile with quality, coverage, imbalance, concentration, and method-readiness evidence for downstream risk screening.

Its SKILL.md is about 1.1k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

The repository describes itself as: 面向数学建模竞赛的 Claude Code / Codex Skills ,支持分阶段建模流程与 Python、MATLAB/北太天元代码分支。 The licence is MIT.

Example prompts

  • “/data-auditor-cleaner”

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Map attachments before cleaning.
  2. Preserve raw data.
  3. Audit structure and semantics.
  4. Compute reusable risk-profile statistics.
  5. Plan and apply cleaning.
  6. Assess readiness per Qx.

What it can do on your machine

Read from SKILL.md and the folder at commit 0b46e9c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are json).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Data Auditor Cleaner loads about 1.1k tokens when it runs. Until then it costs about 58 tokens; SKILL.md has 400 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~58
When it runs · the whole SKILL.md, loaded when a task matches
~1.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from zhnnky329/MathModeling-skills at commit 0b46e9c, republished under its MIT licence (© zhnnky329). 400 words, ~1,055 tokens.

Download SKILL.mdSave it as .claude/skills/data-auditor-cleaner/SKILL.md (or your agent's skills folder).
name
data-auditor-cleaner
description
Map contest attachments to subquestions, audit and clean raw data, and emit one reusable data profile with quality, coverage, imbalance, concentration, and method-readiness evidence for downstream risk screening.

Purpose

Create traceable cleaned data and one reusable profile. Do not repeat the same data inspection separately for every candidate method.

Preconditions

  • Problem parse and subquestion IDs exist.
  • Raw files are available under workspace/data_raw/ or the workspace's documented legacy raw-data path.
  • Required outputs and known field needs are available.

Stop rather than fabricate a missing attachment, unit, field meaning, or label.

Workflow

  1. Map attachments before cleaning.

    • List each attachment with name, size, sheet names, headers, and a small preview.
    • Map it to Qx or mark it shared.
    • Ask the user only when two mappings remain materially plausible.
  2. Preserve raw data.

    • Treat raw files as read-only.
    • Record hashes or stable file metadata when practical.
    • Write cleaned copies under workspace/data_clean/.
  3. Audit structure and semantics.

    • Rows, columns, keys, types, units, categories, time granularity, and encoding.
    • Missing values, duplicates, impossible values, outliers, discontinuities, and leakage risks.
    • Field-to-subquestion and field-to-required-output mapping.
  4. Compute reusable risk-profile statistics.

    • Effective sample size and rows usable per Qx.
    • Missingness by field and row.
    • Numeric distribution summaries and extreme-value rates.
    • Category/class counts, imbalance ratios, rare levels, and cardinality.
    • Time coverage, gaps, sampling interval, and chronological split constraints.
    • Correlation/redundancy warnings where relevant.
    • Target or score concentration indicators when a target exists.
    • Record facts; do not convert them into a final method verdict.
  5. Plan and apply cleaning.

    • Separate safe normalization of representation from assumption-bearing imputations or removals.
    • Explain and record every assumption-bearing operation.
    • Keep reproducible cleaning code only when transformations are nontrivial.
  6. Assess readiness per Qx.

    • ready, ready_with_warnings, or blocked.
    • Name missing fields and risks precisely.
    • Hand the profile to method-selector for method-specific risk probes.
Show full SKILL.md (132 more words)Show less

Canonical Outputs

text
workspace/data/data_report.md
workspace/data/data_profile.json
workspace/data_clean/<cleaned files>
workspace/code/scripts/<cleaning script>   # only when needed

Accept legacy workspace/data/data_clean/ as an input/output location during migration.

Data Profile Contract

data_profile.json contains:

json
{
  "schema_version": 1,
  "raw_files": [],
  "attachment_mapping": [],
  "fields": [],
  "quality": {
    "missingness": {},
    "duplicates": {},
    "impossible_values": {},
    "outliers": {}
  },
  "coverage": {
    "rows": 0,
    "effective_sample_size": null,
    "time_range": null,
    "time_gaps": null
  },
  "distribution_risks": {
    "class_imbalance": null,
    "rare_categories": [],
    "high_cardinality": [],
    "redundancy_warnings": [],
    "concentration_metrics": {}
  },
  "per_question_readiness": {},
  "cleaned_files": [],
  "unresolved_risks": []
}

Use null with an explanation when a field is not applicable; do not invent a value to fill the schema.

Rules

  • Do not select the model.
  • Do not overwrite raw data.
  • Do not silently delete, impute, winsorize, rescale, or recode.
  • Do not produce decorative EDA.
  • Reuse one profile downstream instead of regenerating statistics.
  • Store detailed row-level change logs only when changes occurred; successful no-op checks need only summary counts.

Verification

  • Attachment mapping is unambiguous or human-confirmed.
  • Raw files remain untouched.
  • Cleaned files trace to raw sources and transformation rules.
  • Profile includes effective sample size, imbalance/cardinality, and concentration evidence when applicable.
  • Readiness is reported per subquestion.
  • Downstream handoff points to paths rather than pasting the full report.

© zhnnky329, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .codex/skills/data-auditor-cleaner of zhnnky329/MathModeling-skills.

Open the folder on GitHubat commit 0b46e9c

Compare with similar skills

Data Auditor Cleaner next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Data Auditor Cleaner compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Data Auditor Cleaner this skillzhnnky329/MathModeling-skills1.1k—~1.1kAutomated safety check: PassMIT
Token Mapnexu-io/open-design100k—~1.4kAutomated safety check: PassApache-2.0
Maps Geographyasgeirtj/system_prompts_leaks69k—~717Automated safety check: PassCC0-1.0
Gh Attachsickn33/agentic-awesome-skills47k1 repos~1.6kAutomated safety check: PassMIT
Gh Attachgithub/awesome-copilot40k—~529Automated safety check: PassMIT
Android Clean Architectureaffaan-m/ECC277k4 repos~2.2kAutomated safety check: PassMIT

Similar skills

  • Token Map

    nexu-io/open-design

    Map an extracted Figma / source-code token bag onto the active OD design system, producing a deterministic mapping the generate stage can consume.

    100k GitHub stars~1.4k tokensUpdated yesterday
    Frontend & DesignAuto-check passed
  • Maps Geography

    asgeirtj/system_prompts_leaks

    Accurate maps from real geo data — use for any map, or whenever geography would make a good graphic for a deliverable

    69k GitHub stars~717 tokensUpdated yesterday
    Auto-check passed
  • Gh Attach

    sickn33/agentic-awesome-skills

    Upload and download GitHub user-attachments (screenshots, PDFs, zips, videos) from the terminal; use when asked to attach or embed a file in a PR, issue, or comment, or download an attachment URL.

    47k GitHub starsUsed in 1 repo~1.6k tokens
    Documents & OfficeAuto-check passed
  • Gh Attach

    github/awesome-copilot

    Official

    Uploads a local file (screenshot, image, PDF, zip, video) to GitHub user-attachments, downloads GitHub user-attachments, and embeds local files in a PR, issue, or comment.

    40k GitHub stars~529 tokensUpdated 2 days ago
    Documents & OfficeAuto-check passed
  • Applies Clean Architecture to Android and Kotlin Multiplatform projects: module layout, dependency rules, UseCases, Repositories and data layer patterns.

    277k GitHub starsUsed in 4 repos~2.2k tokens
    DevelopmentAuto-check passed
  • Feature Map

    onyx-dot-app/onyx

    Use the Onyx feature map (.agents/feature-map/) to learn what a product surface does, the code behind it, and what a change can break.

    32k GitHub stars~459 tokensUpdated yesterday
    Auto-check passed

More from zhnnky329/MathModeling-skills

All 29 skills in this repo
  • Method Selector

    zhnnky329/MathModeling-skills

    Build and risk-screen a compact role-based method shortlist for a mathematical-modeling subquestion.

    1.1k GitHub stars~1.6k tokensUpdated 17 days ago
    Auto-check passed
  • Problem Classifier

    zhnnky329/MathModeling-skills

    Classify each parsed mathematical-modeling subquestion by required output and structure, surface ambiguous framing trade-offs for human choice, and record primary/secondary task types without…

    1.1k GitHub stars~611 tokensUpdated 17 days ago
    Auto-check passed
  • Decision Prompt Builder

    zhnnky329/MathModeling-skills

    Build one compact choice card at a genuine mathematical-modeling judgment point.

    1.1k GitHub stars~837 tokensUpdated 17 days ago
    Auto-check passed
  • Matlab Model Code Generator

    zhnnky329/MathModeling-skills

    Generate and run minimal reproducible MATLAB or Beita Tianyuan compatible code for the human-approved main method and usable baseline, with compact experiment artifacts and a canonical run summary.

    1.1k GitHub stars~685 tokensUpdated 17 days ago
    Auto-check passed
  • Model Code Analyzer

    zhnnky329/MathModeling-skills

    Translate a human-approved main method and usable baseline into a minimal language-neutral implementation and experiment contract.

    1.1k GitHub stars~881 tokensUpdated 17 days ago
    Auto-check passed
  • Modeler Decision Logger

    zhnnky329/MathModeling-skills

    Faithfully append a human modeler's choice and rationale to one canonical per-subquestion JSONL decision ledger.

    1.1k GitHub stars~783 tokensUpdated 17 days ago
    Auto-check passed

Questions about Data Auditor Cleaner

What does Data Auditor Cleaner do?

Map contest attachments to subquestions, audit and clean raw data, and emit one reusable data profile with quality, coverage, imbalance, concentration, and method-readiness evidence for downstream…. Data Auditor Cleaner is an agent skill from zhnnky329/MathModeling-skills. Map contest attachments to subquestions, audit and clean raw data, and emit one reusable data profile with quality, coverage, imbalance, concentration, and method-readiness evidence for downstream risk screening.

How do I install Data Auditor Cleaner in Claude Code?

Run `npx skills add zhnnky329/MathModeling-skills --skill data-auditor-cleaner -a claude-code`. Or copy the skill folder (.codex/skills/data-auditor-cleaner in zhnnky329/MathModeling-skills) into .claude/skills/data-auditor-cleaner in your project. Claude Code loads it when a task matches its description.

How do I install Data Auditor Cleaner in Codex?

Run `npx skills add zhnnky329/MathModeling-skills --skill data-auditor-cleaner -a codex`. Or copy the skill folder (.codex/skills/data-auditor-cleaner in zhnnky329/MathModeling-skills) into .agents/skills/data-auditor-cleaner in your project. Codex loads it when a task matches its description.

Can I use Data Auditor Cleaner in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add zhnnky329/MathModeling-skills --skill data-auditor-cleaner -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/data-auditor-cleaner, .gemini/skills/data-auditor-cleaner, .github/skills/data-auditor-cleaner and .opencode/skills/data-auditor-cleaner in your project.

What does Data Auditor Cleaner need to run?

SKILL.md names no scripts, command-line tools or credentials: Data Auditor Cleaner is instructions for the agent only.

Does Data Auditor Cleaner access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Data Auditor Cleaner safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Data Auditor Cleaner use?

Data Auditor Cleaner is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Data Auditor Cleaner use?

About 1.1k tokens (SKILL.md is roughly 4.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Data Auditor Cleaner?

Skills that share tags, products or a category with Data Auditor Cleaner: Token Map (nexu-io/open-design, 100k stars), Maps Geography (asgeirtj/system_prompts_leaks, 69k stars), Gh Attach (sickn33/agentic-awesome-skills, 47k stars) and Gh Attach (github/awesome-copilot, 40k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Data Auditor Cleaner?

zhnnky329 (a GitHub user) maintains it in zhnnky329/MathModeling-skills, which has 1,060 GitHub stars. The repository holds 29 skills in this directory. The repository was last updated on September 24, 2026.

Source: zhnnky329/MathModeling-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.