Agent skill

Audit Data Integrity

by NeuroAIHub in NeuroAIHub/BrainPilot

Audit scientific data semantics, sample and label alignment, leakage, group splits, preprocessing boundaries, and train-to-inference transforms.

AGPL-3.0Auto-check passed

Install Audit Data Integrity

skills CLI
$ npx skills add NeuroAIHub/BrainPilot --skill audit-data-integrity -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install NeuroAIHub/BrainPilot audit-data-integrity --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/NeuroAIHub/BrainPilot.git skills-src && mkdir -p .claude/skills && cp -r skills-src/packages/plugin-auditor/skills/audit-data-integrity .claude/skills/audit-data-integrity && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
audit-data-integrity
GitHub stars
1.1k
Token cost
~433 tokens
SKILL.md length
186 words
Files
2
Skills in repo
59
Repo updated
First seen
Licence
AGPL-3.0

At a glance

Audit scientific data semantics, sample and label alignment, leakage, group splits, preprocessing boundaries, and train-to-inference transforms.

  • Works in 5 steps: Establish the semantic meaning and order… → Verify that every feature row remains… → Verify train/test independence at the… → …
  • Any result based on datasets
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • Feature matrices

What it does

Audit Data Integrity is an agent skill from NeuroAIHub/BrainPilot. Audit scientific data semantics, sample and label alignment, leakage, group splits, preprocessing boundaries, and train-to-inference transforms. Use for any result based on datasets, feature matrices, tensors, repeated observations, or learned preprocessing.

Its SKILL.md is about 430 tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files (for example `agents/openai.yaml`).

The repository describes itself as: BrainPilot: Automating Brain Discovery with Agentic Research. The licence is AGPL-3.0.

When your agent uses it

  • Any result based on datasets
  • Feature matrices
  • Repeated observations
  • Learned preprocessing

Example prompts

  • “/audit-data-integrity”

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Establish the semantic meaning and order of every source axis. Trace each
  2. Verify that every feature row remains aligned with its label, subject,
  3. Verify train/test independence at the real sampling unit. Check subject,
  4. Verify that imputation, scaling, feature selection, dimensionality reduction,
  5. Establish the value domain and transform contract. Confirm that training,

What it can do on your machine

Read from SKILL.md and the folder at commit 93f6855. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Audit Data Integrity loads about 433 tokens when it runs. Until then it costs about 70 tokens; SKILL.md has 186 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~70
When it runs · the whole SKILL.md, loaded when a task matches
~433

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from NeuroAIHub/BrainPilot at commit 93f6855, republished under its AGPL-3.0 licence (© NeuroAIHub). 186 words, ~433 tokens.

Download SKILL.mdSave it as .claude/skills/audit-data-integrity/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
audit-data-integrity
description
Audit scientific data semantics, sample and label alignment, leakage, group splits, preprocessing boundaries, and train-to-inference transforms. Use for any result based on datasets, feature matrices, tensors, repeated observations, or learned preprocessing.

Audit Data Integrity

Inspect existing code, contracts, manifests, and logs; do not recompute missing results. Report each applicable check as pass, flaw, or unverified, with a specific evidence path or exact missing evidence.

Checks

  1. Establish the semantic meaning and order of every source axis. Trace each transpose, reshape, ravel, flatten, stack, concatenation, cache, and reload that can change sample identity. Shape equality alone is insufficient.
  2. Verify that every feature row remains aligned with its label, subject, condition, session, bin, and other grouping identifiers. Look for explicit value-level assertions on representative samples.
  3. Verify train/test independence at the real sampling unit. Check subject, session, site, family, temporal, repeated-measure, and augmentation leakage as applicable.
  4. Verify that imputation, scaling, feature selection, dimensionality reduction, resampling, threshold selection, and model selection are fitted only within training folds.
  5. Establish the value domain and transform contract. Confirm that training, exported artifacts, manifests, and inference use the same clipping, missing-value handling, transform, and feature order.

For a bounded parallel review, give a method-reviewer the data contract, pipeline files, and validation evidence. Ask for evidence and candidate findings, not a verdict.

© NeuroAIHub, AGPL-3.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in packages/plugin-auditor/skills/audit-data-integrity of NeuroAIHub/BrainPilot.

  • SKILL.md
  • agents/openai.yaml

Open the folder on GitHubat commit 93f6855

Compare with similar skills

Audit Data Integrity next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Audit Data Integrity compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Audit Data Integrity this skillNeuroAIHub/BrainPilot1.1k—~433Automated safety check: PassAGPL-3.0
Integration Testingthedaviddias/Front-End-Checklist74k—~514Automated safety check: PassMIT
API Integrationsickn33/agentic-awesome-skills47k1 repos~1.3kAutomated safety check: PassMIT
Labarchive Integrationdavila7/claude-code-templates32k10 repos~2.3kAutomated safety check: PassMIT
MAUI Integration Test Runnerdotnet/maui23k—~2.1kAutomated safety check: PassMIT
Protocolsio Integrationdavila7/claude-code-templates32k10 repos~3.7kAutomated safety check: PassMIT

Similar skills

  • Integration Testing

    thedaviddias/Front-End-Checklist

    A skill your agent uses when reviewing CI coverage, automated checks, or test strategy related to Write integration tests for key workflows.

    74k GitHub stars~514 tokensUpdated 3 days ago
    Testing & QAAuto-check passed
  • API Integration

    sickn33/agentic-awesome-skills

    Designs event-driven architectures, webhook systems, API chaining flows, ETL pipelines, and integration patterns between services.

    47k GitHub starsUsed in 1 repo~1.3k tokens
    Backend & APIsAuto-check passed
  • Labarchive Integration

    davila7/claude-code-templates

    Electronic lab notebook API integration. An agent skill from davila7/claude-code-templates.

    32k GitHub starsUsed in 10 repos~2.3k tokens
    Backend & APIsAuto-check passed
  • Official

    Builds and packs .NET MAUI, installs the local workloads and runs integration tests for templates, samples and end-to-end scenarios by category.

    23k GitHub stars~2.1k tokensUpdated today
    Testing & QAAuto-check passed
  • Protocolsio Integration

    davila7/claude-code-templates

    Integration with protocols.io API for managing scientific protocols.

    32k GitHub starsUsed in 10 repos~3.7k tokens
    Backend & APIsAuto-check passed
  • Robius Matrix Integration

    sickn33/agentic-awesome-skills

    CRITICAL: Use for Matrix SDK integration with Makepad. An agent skill from sickn33/agentic-awesome-skills.

    47k GitHub starsUsed in 2 repos~3.6k tokens
    Backend & APIsAuto-check passed

More from NeuroAIHub/BrainPilot

All 59 skills in this repo
  • Deeplabcut

    NeuroAIHub/BrainPilot

    Toolbox for markerless animal pose estimation with DeepLabCut.

    1.1k GitHub stars~1.7k tokensUpdated 7 days ago
    Auto-check passed
  • Fmriprep

    NeuroAIHub/BrainPilot

    Preprocess task-based or resting-state fMRI data with fMRIPrep — a robust, BIDS-App preprocessing pipeline built on FSL, ANTs, FreeSurfer, AFNI, and Nilearn.

    1.1k GitHub stars~4.1k tokensUpdated 7 days ago
    Auto-check passed
  • Mne Python Guide

    NeuroAIHub/BrainPilot

    Domain-validated pipeline guidance for EEG/MEG data analysis using MNE-Python: data loading, preprocessing (filtering, ICA, re-referencing), epoching, ERP/ERF computation, time-frequency…

    1.1k GitHub stars~2.3k tokensUpdated 7 days ago
    Auto-check passed
  • Netneurotools Guide

    NeuroAIHub/BrainPilot

    Domain-validated guidance for network neuroscience analysis using netneurotools: datasets, brain network metrics, connectivity consensus, modularity, spatial statistics, null models, and cortical…

    1.1k GitHub stars~2.6k tokensUpdated 7 days ago
    Auto-check passed
  • Nature Figure

    NeuroAIHub/BrainPilot

    Submission-grade Nature/high-impact journal figure workflow for Python or R.

    1.1k GitHub starsUsed in 1 repo~1.3k tokens
    Auto-check passed
  • Pycortex Guide

    NeuroAIHub/BrainPilot

    Domain-validated guidance for cortical surface visualization and brain surface rendering of fMRI data using pycortex: data types (Volume, Vertex, Dataset), 2D cortical flatmaps, 3D WebGL brain…

    1.1k GitHub stars~1.6k tokensUpdated 7 days ago
    Auto-check passed

Questions about Audit Data Integrity

What does Audit Data Integrity do?

Audit scientific data semantics, sample and label alignment, leakage, group splits, preprocessing boundaries, and train-to-inference transforms. Audit Data Integrity is an agent skill from NeuroAIHub/BrainPilot. Audit scientific data semantics, sample and label alignment, leakage, group splits, preprocessing boundaries, and train-to-inference transforms.

When should I use Audit Data Integrity?

Audit Data Integrity fits situations like: any result based on datasets; feature matrices; repeated observations; learned preprocessing.

How do I install Audit Data Integrity in Claude Code?

Run `npx skills add NeuroAIHub/BrainPilot --skill audit-data-integrity -a claude-code`. Or copy the skill folder (packages/plugin-auditor/skills/audit-data-integrity in NeuroAIHub/BrainPilot) into .claude/skills/audit-data-integrity in your project. Claude Code loads it when a task matches its description.

How do I install Audit Data Integrity in Codex?

Run `npx skills add NeuroAIHub/BrainPilot --skill audit-data-integrity -a codex`. Or copy the skill folder (packages/plugin-auditor/skills/audit-data-integrity in NeuroAIHub/BrainPilot) into .agents/skills/audit-data-integrity in your project. Codex loads it when a task matches its description.

Can I use Audit Data Integrity in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NeuroAIHub/BrainPilot --skill audit-data-integrity -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/audit-data-integrity, .gemini/skills/audit-data-integrity, .github/skills/audit-data-integrity and .opencode/skills/audit-data-integrity in your project.

What does Audit Data Integrity need to run?

SKILL.md names no scripts, command-line tools or credentials: Audit Data Integrity is instructions for the agent only.

Does Audit Data Integrity access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Audit Data Integrity safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Audit Data Integrity use?

Audit Data Integrity is published under the AGPL-3.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Audit Data Integrity use?

About 433 tokens (SKILL.md is roughly 1.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Audit Data Integrity?

Skills that share tags, products or a category with Audit Data Integrity: Integration Testing (thedaviddias/Front-End-Checklist, 74k stars), API Integration (sickn33/agentic-awesome-skills, 47k stars), Labarchive Integration (davila7/claude-code-templates, 32k stars) and MAUI Integration Test Runner (dotnet/maui, 23k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Audit Data Integrity?

NeuroAIHub (a GitHub organization) maintains it in NeuroAIHub/BrainPilot, which has 1,062 GitHub stars. The repository holds 59 skills in this directory. The repository was last updated on October 2, 2026.

Source: NeuroAIHub/BrainPilot on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.