Agent skill

Audit Model Validation

by NeuroAIHub in NeuroAIHub/BrainPilot

Audit whether method discovery, comparison, representative real-data validation, collapse diagnostics, pruning, and selection evidence support claims of suitability or superiority.

AGPL-3.0Auto-check passed

Install Audit Model Validation

skills CLI
$ npx skills add NeuroAIHub/BrainPilot --skill audit-model-validation -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install NeuroAIHub/BrainPilot audit-model-validation --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/NeuroAIHub/BrainPilot.git skills-src && mkdir -p .claude/skills && cp -r skills-src/packages/plugin-auditor/skills/audit-model-validation .claude/skills/audit-model-validation && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
audit-model-validation
GitHub stars
1.1k
Token cost
~1.7k tokens
SKILL.md length
838 words
Files
2
Skills in repo
59
Repo updated
First seen
Licence
AGPL-3.0

At a glance

Audit whether method discovery, comparison, representative real-data validation, collapse diagnostics, pruning, and selection evidence support claims of suitability or superiority.

  • Works in 10 steps: Identify the intended use and claim.… → Inventory the evidence by source:… → Verify that discovery covered a… → …
  • Research-method selection
  • SKILL.md covers Separate operational and…, Checks and Audit techniques
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Audit Model Validation is an agent skill from NeuroAIHub/BrainPilot. Audit whether method discovery, comparison, representative real-data validation, collapse diagnostics, pruning, and selection evidence support claims of suitability or superiority. Use for research-method selection, empirical evaluation, benchmarking, modelling, prediction, or other conclusions that depend on choosing among alternatives.

Its SKILL.md is about 1.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files (for example `agents/openai.yaml`).

The repository describes itself as: BrainPilot: Automating Brain Discovery with Agentic Research. The licence is AGPL-3.0.

When your agent uses it

  • Research-method selection
  • Empirical evaluation
  • Other conclusions that depend on choosing among alternatives

Example prompts

  • “/audit-model-validation”

Workflow steps

10 steps, taken from the first numbered list in SKILL.md.

  1. Identify the intended use and claim. Determine whether the evidence source,
  2. Inventory the evidence by source: representative real data,
  3. Verify that discovery covered a sufficiently broad set of credible,
  4. Distinguish operational correctness and feasibility from evidential
  5. Check that comparisons are fair and that screening, pruning, adaptation, and
  6. Check applicable baselines, chance levels, negative controls, grouped metrics,
  7. Check for degenerate predictions and hidden training failure before accepting
  8. Verify that implementation parameters preserve their scientific meaning.
  9. For each material iteration, verify that its representative real-data
  10. Bound every conclusion to the alternatives, evidence, conditions, and checks

What it can do on your machine

Read from SKILL.md and the folder at commit 93f6855. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Audit Model Validation loads about 1.7k tokens when it runs. Until then it costs about 91 tokens; SKILL.md has 838 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~91
When it runs · the whole SKILL.md, loaded when a task matches
~1.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from NeuroAIHub/BrainPilot at commit 93f6855, republished under its AGPL-3.0 licence (© NeuroAIHub). 838 words, ~1,711 tokens.

Download SKILL.mdSave it as .claude/skills/audit-model-validation/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
audit-model-validation
description
Audit whether method discovery, comparison, representative real-data validation, collapse diagnostics, pruning, and selection evidence support claims of suitability or superiority. Use for research-method selection, empirical evaluation, benchmarking, modelling, prediction, or other conclusions that depend on choosing among alternatives.

Audit Method Validation

Judge whether the alternatives considered and evidence gathered support the stated conclusion; do not choose a method or redesign the work.

Separate operational and empirical evidence

Classify every supplied check before judging the claim:

  • Operational evidence covers importability, tensor shapes, finite outputs, parameter counts, deterministic evaluation, packaging, gradient flow, and memorization of synthetic or random data. It can establish that an artifact runs, but normally cannot establish task suitability or generalization.
  • Empirical evidence measures decision-relevant behavior on representative real data with labels, splits, metrics, and conditions matched closely enough to the intended use. It supports performance and selection claims only within those observed conditions.

Do not let the volume, precision, or independence of operational checks substitute for missing empirical evidence. If the intended claim is that a method is suitable, effective, robust, or preferred and representative real-data evidence is absent, mark empirical adequacy unverified. This is a material finding requiring REVISE or BLOCK, even when every implementation and protocol-compliance check passes. A report may separately confirm the narrower claim that the artifact is operationally valid.

Checks

  1. Identify the intended use and claim. Determine whether the evidence source, comparison unit, conditions, measures, and time horizon represent that use.
  2. Inventory the evidence by source: representative real data, non-representative real data, simulated data, random tensors, literature, or implementation inspection. Verify that the strongest conclusion does not exceed the strongest applicable evidence class.
  3. Verify that discovery covered a sufficiently broad set of credible, substantively different alternatives for the claim. Check established baselines and require justification for material omissions; a long list of minor variants is not breadth. The previous working baseline is an essential comparison when a proposed replacement claims improvement or suitability.
  4. Distinguish operational correctness and feasibility from evidential adequacy. Synthetic, self-consistency, or convenience evidence establishes real-world suitability only when the task makes it representative.
  5. Check that comparisons are fair and that screening, pruning, adaptation, and resource-driven reductions follow declared rules without discarding essential breadth merely for implementation convenience. Treat a missing reader, dependency, accelerator, or preprocessing artifact as a blocked evidence path, not as justification for replacing real-data validation with synthetic checks.
  6. Check applicable baselines, chance levels, negative controls, grouped metrics, uncertainty, multiplicity, evidence reuse, and anomalously optimistic results. Treat nuisance, proxy, or setting-specific signals that can mimic the intended target as material risks.
  7. Check for degenerate predictions and hidden training failure before accepting aggregate scores. Require class coverage and per-group diagnostics when class collapse is plausible.
  8. Verify that implementation parameters preserve their scientific meaning. Recompute durations, frequencies, window sizes, and other physical quantities after resampling or unit conversion; copying sample counts across sampling rates is not protocol fidelity.
  9. For each material iteration, verify that its representative real-data validation ran on the same or later code and configuration revision than the implementation diff under review. Preserve the valid iterative sequence in which one result motivates the next decision; final validation of an older revision cannot validate a newer candidate.
  10. Bound every conclusion to the alternatives, evidence, conditions, and checks actually completed.
Show full SKILL.md (326 more words)Show less

Audit techniques

Use the smallest read-only checks that can expose an invalid conclusion:

  1. Build a claim-to-evidence table with separate rows for operational validity, within-condition performance, transfer performance, and method superiority. Never merge these into one generic validated status.
  2. For a balanced K-class problem, compare results with constant prediction: accuracy 1/K, kappa 0, and macro-F1 2 / (K * (K + 1)) under the usual zero-division convention. Exact or near-exact agreement is a collapse warning, not proof. Confirm with per-class prediction counts, class coverage, normalized prediction entropy, confusion matrices, or per-group macro-F1. If those artifacts are unavailable, report collapse as suspected and the diagnosis as unverified rather than asserting certainty.
  3. Inspect train/validation loss, selected epoch, logits or probability ranges, and prediction histograms per subject/fold/group. Aggregate means can hide constant predictors, failed subjects, or cancellation across groups.
  4. Match the validation split to the deployment shift. Random validation within one session does not validate cross-session transfer. Look for session/run/time blocks, held-out conditions, external data, or explicit shift stress tests; otherwise narrow the claim to same-condition performance.
  5. Recalculate architecture time scales from the actual sample rate. Check normalization statistics, nonlinearities that amplify shift, and numerical guards such as epsilon/clamping when their failure could create uniform class bias or non-finite values. Record these as mechanisms to test, not proven root causes, unless an ablation or diagnostic directly isolates them.
  6. Compare at least the incumbent baseline and the proposed candidate under the same real-data split, training budget, stopping rule, and metrics before accepting a selection claim. Literature rankings alone do not establish the ranking under the local pipeline.
  7. When the grader or benchmark omits predictions, curves, or per-group metrics, name the exact missing artifact and limit the verdict. Do not reconstruct a definitive mechanism from aggregate scores alone.

For a bounded parallel review, give a method-reviewer the method survey, protocol, comparison evidence, validation outputs, prediction diagnostics, and stated claims. Ask for evidence and candidate findings, not a verdict.

© NeuroAIHub, AGPL-3.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in packages/plugin-auditor/skills/audit-model-validation of NeuroAIHub/BrainPilot.

  • SKILL.md
  • agents/openai.yaml

Open the folder on GitHubat commit 93f6855

Compare with similar skills

Audit Model Validation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Audit Model Validation compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Audit Model Validation this skillNeuroAIHub/BrainPilot1.1k—~1.7kAutomated safety check: PassAGPL-3.0
Santa Methodaffaan-m/ECC276k3 repos~3.1kAutomated safety check: PassMIT
Santa Methodaffaan-m/ECC276k—~2.1kAutomated safety check: PassMIT
Santa Methodaffaan-m/ECC276k—~1.9kAutomated safety check: PassMIT
Gaia Architecture Comparisonruvnet/ruflo74k—~1.3kAutomated safety check: NotesMIT
Modern Array Methodsthedaviddias/Front-End-Checklist74k—~494Automated safety check: PassMIT

Similar skills

  • Santa Method

    affaan-m/ECC

    Multi-agent adversarial verification: two independent reviewers with the same rubric must both pass before output ships, with a fix-and-re-review convergence loop and human escalation cap.

    276k GitHub starsUsed in 3 repos~3.1k tokens
    EducationAuto-check passed
  • Santa Method

    affaan-m/ECC

    収束ループを持つマルチエージェント敵対的検証。2つの独立したレビューエージェントが両方合格して初めて出力を出荷できます。

    276k GitHub stars~2.1k tokensUpdated 4 days ago
    Auto-check passed
  • Santa Method

    affaan-m/ECC

    具有收敛循环的多智能体对抗验证。两个独立的审查代理必须都通过,输出才能发送。

    276k GitHub stars~1.9k tokensUpdated 4 days ago
    Auto-check passed
  • Side-by-side comparison of ruflo vs HAL vs other GAIA harnesses — capability gaps, design decisions, and improvement roadmap

    74k GitHub stars~1.3k tokensUpdated today
    DevelopmentAuto-check: notes
  • Modern Array Methods

    thedaviddias/Front-End-Checklist

    A skill your agent uses when reviewing scripts, client components, bundles, or runtime behavior related to Use modern array and object methods.

    74k GitHub stars~494 tokensUpdated 3 days ago
    Auto-check passed
  • Refactor Method Complexity Reduce

    github/awesome-copilot

    Official

    Refactor given method ${input:methodName} to reduce its cognitive complexity to ${input:complexityThreshold} or below, by extracting helper methods.

    40k GitHub starsUsed in 1 repo~1.1k tokens
    DevelopmentAuto-check passed

More from NeuroAIHub/BrainPilot

All 59 skills in this repo
  • Deeplabcut

    NeuroAIHub/BrainPilot

    Toolbox for markerless animal pose estimation with DeepLabCut.

    1.1k GitHub stars~1.7k tokensUpdated 7 days ago
    Auto-check passed
  • Fmriprep

    NeuroAIHub/BrainPilot

    Preprocess task-based or resting-state fMRI data with fMRIPrep — a robust, BIDS-App preprocessing pipeline built on FSL, ANTs, FreeSurfer, AFNI, and Nilearn.

    1.1k GitHub stars~4.1k tokensUpdated 7 days ago
    Auto-check passed
  • Mne Python Guide

    NeuroAIHub/BrainPilot

    Domain-validated pipeline guidance for EEG/MEG data analysis using MNE-Python: data loading, preprocessing (filtering, ICA, re-referencing), epoching, ERP/ERF computation, time-frequency…

    1.1k GitHub stars~2.3k tokensUpdated 7 days ago
    Auto-check passed
  • Netneurotools Guide

    NeuroAIHub/BrainPilot

    Domain-validated guidance for network neuroscience analysis using netneurotools: datasets, brain network metrics, connectivity consensus, modularity, spatial statistics, null models, and cortical…

    1.1k GitHub stars~2.6k tokensUpdated 7 days ago
    Auto-check passed
  • Nature Figure

    NeuroAIHub/BrainPilot

    Submission-grade Nature/high-impact journal figure workflow for Python or R.

    1.1k GitHub starsUsed in 1 repo~1.3k tokens
    Auto-check passed
  • Pycortex Guide

    NeuroAIHub/BrainPilot

    Domain-validated guidance for cortical surface visualization and brain surface rendering of fMRI data using pycortex: data types (Volume, Vertex, Dataset), 2D cortical flatmaps, 3D WebGL brain…

    1.1k GitHub stars~1.6k tokensUpdated 7 days ago
    Auto-check passed

Questions about Audit Model Validation

What does Audit Model Validation do?

Audit whether method discovery, comparison, representative real-data validation, collapse diagnostics, pruning, and selection evidence support claims of suitability or superiority. Audit Model Validation is an agent skill from NeuroAIHub/BrainPilot. Audit whether method discovery, comparison, representative real-data validation, collapse diagnostics, pruning, and selection evidence support claims of suitability or superiority.

When should I use Audit Model Validation?

Audit Model Validation fits situations like: research-method selection; empirical evaluation; other conclusions that depend on choosing among alternatives.

How do I install Audit Model Validation in Claude Code?

Run `npx skills add NeuroAIHub/BrainPilot --skill audit-model-validation -a claude-code`. Or copy the skill folder (packages/plugin-auditor/skills/audit-model-validation in NeuroAIHub/BrainPilot) into .claude/skills/audit-model-validation in your project. Claude Code loads it when a task matches its description.

How do I install Audit Model Validation in Codex?

Run `npx skills add NeuroAIHub/BrainPilot --skill audit-model-validation -a codex`. Or copy the skill folder (packages/plugin-auditor/skills/audit-model-validation in NeuroAIHub/BrainPilot) into .agents/skills/audit-model-validation in your project. Codex loads it when a task matches its description.

Can I use Audit Model Validation in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NeuroAIHub/BrainPilot --skill audit-model-validation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/audit-model-validation, .gemini/skills/audit-model-validation, .github/skills/audit-model-validation and .opencode/skills/audit-model-validation in your project.

What does Audit Model Validation need to run?

SKILL.md names no scripts, command-line tools or credentials: Audit Model Validation is instructions for the agent only.

Does Audit Model Validation access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Audit Model Validation safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Audit Model Validation use?

Audit Model Validation is published under the AGPL-3.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Audit Model Validation use?

About 1.7k tokens (SKILL.md is roughly 6.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Audit Model Validation?

Skills that share tags, products or a category with Audit Model Validation: Santa Method (affaan-m/ECC, 276k stars), Santa Method (affaan-m/ECC, 276k stars), Santa Method (affaan-m/ECC, 276k stars) and Gaia Architecture Comparison (ruvnet/ruflo, 74k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Audit Model Validation?

NeuroAIHub (a GitHub organization) maintains it in NeuroAIHub/BrainPilot, which has 1,062 GitHub stars. The repository holds 59 skills in this directory. The repository was last updated on October 2, 2026.

Source: NeuroAIHub/BrainPilot on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.