Agent skill

Audit Dataset

by PKU-YuanGroup in PKU-YuanGroup/OpenAI4S

Audit tabular datasets before analysis or training for schema drift, missing values, duplicate rows or IDs, target imbalance, and entity or group leakage across splits using pure-stdlib helpers.

MITAuto-check passedData & Analytics

Install Audit Dataset

skills CLI
$ npx skills add PKU-YuanGroup/OpenAI4S --skill audit-dataset -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install PKU-YuanGroup/OpenAI4S audit-dataset --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/PKU-YuanGroup/OpenAI4S.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/audit-dataset .claude/skills/audit-dataset && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
audit-dataset
GitHub stars
622
Token cost
~637 tokens
SKILL.md length
265 words
Files
4
Skills in repo
17
Repo updated
First seen
Licence
MIT

At a glance

Audit tabular datasets before analysis or training for schema drift, missing values, duplicate rows or IDs, target imbalance, and entity or group leakage across splits using pure-stdlib helpers.

  • Works in 6 steps: Load records without silently coercing… → Call audit_rows on a representative or… → Inspect missingness and observed type… → …
  • Tasks that involve Data cleaning
  • SKILL.md covers Workflow, Import and run, Interpretation and Required output
  • Runs Python scripts from its folder

What it does

Audit Dataset is an agent skill from PKU-YuanGroup/OpenAI4S. Audit tabular datasets before analysis or training for schema drift, missing values, duplicate rows or IDs, target imbalance, and entity or group leakage across splits using pure-stdlib helpers.

Its SKILL.md is about 640 tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files (for example `README.md`, `README_zh.md` and `kernel.py`).

It sits in Data & Analytics, covering Data cleaning. The repository describes itself as: Open-source AI agent for scientific research. Analyze data in Python/R with Claude, GPT, Gemini, and more. The licence is MIT.

When your agent uses it

  • Tasks that involve Data cleaning

Example prompts

  • “/audit-dataset”

Requirements

  • Python 3

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Load records without silently coercing values. Preserve source row IDs.
  2. Call audit_rows on a representative or complete list of row mappings.
  3. Inspect missingness and observed type mixtures column by column.
  4. Resolve duplicate records and non-unique identifiers deliberately.
  5. If a split column exists, check both stable IDs and grouping entities for
  6. Record accepted exceptions, then rerun the audit and save the JSON result

What it can do on your machine

Read from SKILL.md and the folder at commit 4a72e87. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Audit Dataset loads about 637 tokens when it runs. Until then it costs about 52 tokens; SKILL.md has 265 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~52
When it runs · the whole SKILL.md, loaded when a task matches
~637

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from PKU-YuanGroup/OpenAI4S at commit 4a72e87, republished under its MIT licence (© PKU-YuanGroup). 265 words, ~637 tokens.

Download SKILL.mdSave it as .claude/skills/audit-dataset/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
audit-dataset
description
Audit tabular datasets before analysis or training for schema drift, missing values, duplicate rows or IDs, target imbalance, and entity or group leakage across splits using pure-stdlib helpers.
origin
openai4s
category
data-quality

Audit a dataset

Use this skill before statistics, model training, or external publication. The goal is a compact, machine-readable audit plus explicit decisions about every issue that could invalidate downstream results.

Workflow

  1. Load records without silently coercing values. Preserve source row IDs.
  2. Call audit_rows on a representative or complete list of row mappings.
  3. Inspect missingness and observed type mixtures column by column.
  4. Resolve duplicate records and non-unique identifiers deliberately.
  5. If a split column exists, check both stable IDs and grouping entities for train/validation/test overlap.
  6. Record accepted exceptions, then rerun the audit and save the JSON result next to the cleaned dataset.

Import and run

Hyphenated Skill directories are loaded with importlib:

python
from importlib import import_module

audit_rows = import_module("audit-dataset.kernel").audit_rows
report = audit_rows(
    rows,
    target="label",
    id_columns=("sample_id",),
    group_columns=("patient_id",),
    split_column="split",
)

rows must be a sequence of mappings. The report contains row and column counts, per-column missing/type/unique summaries, duplicate row and ID counts, target frequencies, and split-leakage examples.

Interpretation

  • Mixed numeric/string types usually indicate parsing or sentinel-value bugs.
  • Missingness is a property of both the data and the collection process; do not impute before checking whether it correlates with label, site, or time.
  • Duplicate IDs are not automatically duplicate observations. Decide whether repeated measures are expected and group them during splitting.
  • Any patient, molecule scaffold, time series, or near-duplicate entity shared across evaluation boundaries can inflate performance even when row IDs differ.
  • A clean structural audit does not establish representativeness, label validity, causal identifiability, or ethical suitability.

Required output

Report the checks performed, blocking findings, accepted exceptions, and the exact source artifact/version. Never describe a dataset as clean without naming the leakage keys and missing-value policy that were checked.

© PKU-YuanGroup, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files in skills/audit-dataset of PKU-YuanGroup/OpenAI4S.

  • SKILL.md
  • README.md
  • README_zh.md
  • kernel.py

Open the folder on GitHubat commit 4a72e87

Compare with similar skills

Audit Dataset next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Audit Dataset compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Audit Dataset this skillPKU-YuanGroup/OpenAI4S622—~637Automated safety check: PassMIT
Question2reportrefraction-ray/xalpha2.7k—~3.2kAutomated safety check: PassMIT
Dingo VerifyMigoXLab/dingo757—~741Automated safety check: NotesApache-2.0
Data Validationplatonai/Browser41.2k—~896Automated safety check: PassApache-2.0
Pandas ProJeffallan/claude-skills12k1 repos~1.5kAutomated safety check: PassMIT
Issues DeduplicationJetBrains/ideavim10k—~1.3kAutomated safety check: PassMIT

Similar skills

  • Question2report

    refraction-ray/xalpha

    Turn a natural-language financial question into a polished, self-contained HTML report.

    2.7k GitHub stars~3.2k tokensUpdated 2 mo ago
    Data & AnalyticsAuto-check passed
  • Dingo Verify

    MigoXLab/dingo

    A skill your agent uses when the user wants to fact-check an article or verify factual claims in a document.

    757 GitHub stars~741 tokensUpdated 2 days ago
    Data & AnalyticsAuto-check: notes
  • Data Validation

    platonai/Browser4

    Validates data against common and custom rules (required fields, formats, ranges).

    1.2k GitHub stars~896 tokensUpdated yesterday
    Data & AnalyticsAuto-check passed
  • Pandas Pro

    Jeffallan/claude-skills

    Handles pandas DataFrame work: cleaning, merging, groupby aggregation, pivots, time-series resampling and memory tuning, with checks on dtypes, shapes and nulls.

    12k GitHub starsUsed in 1 repo~1.5k tokens
    Data & AnalyticsAuto-check passed
  • Issues Deduplication

    JetBrains/ideavim

    Official

    Handles deduplication of YouTrack issues. An agent skill from JetBrains/ideavim.

    10k GitHub stars~1.3k tokensUpdated today
    Data & AnalyticsAuto-check passed
  • Openbb Data Fetcher

    monarchjuno/vibe-investing

    Fetch financial, market, economic, fundamental, news, options, crypto, ETF, index, and macro data through the OpenBB Python interface instead of the OpenBB MCP server.

    299 GitHub stars~2.9k tokensUpdated 5 mo ago
    Data & AnalyticsAuto-check: notes

More from PKU-YuanGroup/OpenAI4S

All 17 skills in this repo
  • Single Cell Rna Analysis

    PKU-YuanGroup/OpenAI4S

    Reproducible Scanpy workflow for human or mouse 10x scRNA-seq and snRNA-seq count matrices: single-sample descriptive QC, clustering and annotation, or comparative donor-aware pseudobulk DE and Milo…

    622 GitHub stars~1.3k tokensUpdated yesterday
    Auto-check passed
  • Bioprobench

    PKU-YuanGroup/OpenAI4S

    Score an LLM's biological-protocol reasoning on the BioProBench benchmark: protocol QA, step ordering, error detection, protocol generation, and LLM-judged error reasoning; or generate the responses.

    622 GitHub stars~2.3k tokensUpdated yesterday
    Auto-check passed
  • Reaction Atom Mapping

    PKU-YuanGroup/OpenAI4S

    Map atoms and changed bonds for a complete reaction with RXNMapper.

    622 GitHub stars~1.2k tokensUpdated yesterday
    Auto-check passed
  • Reaction Forward Prediction

    PKU-YuanGroup/OpenAI4S

    Predict ranked products from reactants and reagents with ReactionT5v2-forward; use for outcome prediction or round-trip recovery.

    622 GitHub stars~2k tokensUpdated yesterday
    Auto-check passed
  • Reaction Yield Estimation

    PKU-YuanGroup/OpenAI4S

    Estimate yield for a fully specified reactant/reagent/product record with ReactionT5v2-yield.

    622 GitHub stars~2.3k tokensUpdated yesterday
    Auto-check passed
  • Rfdiffusion

    PKU-YuanGroup/OpenAI4S

    Generate de novo protein backbones with RFdiffusion for protein-target binders, hotspot-conditioned interfaces, motif scaffolding, partial diffusion, or symmetric assemblies.

    622 GitHub stars~2.2k tokensUpdated yesterday
    Auto-check passed

Questions about Audit Dataset

What does Audit Dataset do?

Audit tabular datasets before analysis or training for schema drift, missing values, duplicate rows or IDs, target imbalance, and entity or group leakage across splits using pure-stdlib helpers. Audit Dataset is an agent skill from PKU-YuanGroup/OpenAI4S. Audit tabular datasets before analysis or training for schema drift, missing values, duplicate rows or IDs, target imbalance, and entity or group leakage across splits using pure-stdlib helpers.

When should I use Audit Dataset?

Audit Dataset fits situations like: tasks that involve Data cleaning.

How do I install Audit Dataset in Claude Code?

Run `npx skills add PKU-YuanGroup/OpenAI4S --skill audit-dataset -a claude-code`. Or copy the skill folder (skills/audit-dataset in PKU-YuanGroup/OpenAI4S) into .claude/skills/audit-dataset in your project. Claude Code loads it when a task matches its description.

How do I install Audit Dataset in Codex?

Run `npx skills add PKU-YuanGroup/OpenAI4S --skill audit-dataset -a codex`. Or copy the skill folder (skills/audit-dataset in PKU-YuanGroup/OpenAI4S) into .agents/skills/audit-dataset in your project. Codex loads it when a task matches its description.

Can I use Audit Dataset in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add PKU-YuanGroup/OpenAI4S --skill audit-dataset -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/audit-dataset, .gemini/skills/audit-dataset, .github/skills/audit-dataset and .opencode/skills/audit-dataset in your project.

What does Audit Dataset need to run?

Going by SKILL.md and its folder, Audit Dataset needs Python for the scripts in its folder. Our summary lists: Python 3.

Does Audit Dataset access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Audit Dataset safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Audit Dataset use?

Audit Dataset is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Audit Dataset use?

About 637 tokens (SKILL.md is roughly 2.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Audit Dataset?

Skills that share tags, products or a category with Audit Dataset: Question2report (refraction-ray/xalpha, 2.7k stars), Dingo Verify (MigoXLab/dingo, 757 stars), Data Validation (platonai/Browser4, 1.2k stars) and Pandas Pro (Jeffallan/claude-skills, 12k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Audit Dataset?

PKU-YuanGroup (a GitHub organization) maintains it in PKU-YuanGroup/OpenAI4S, which has 622 GitHub stars. The repository holds 17 skills in this directory. The repository was last updated on October 9, 2026.

Source: PKU-YuanGroup/OpenAI4S on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.