Agent skill

Experiment Audit

by appleweiping in appleweiping/WEIPING_WIKI

Rigorous audit of experiment results before they become paper claims.

MITAuto-check passedResearch & Science

Install Experiment Audit

skills CLI
$ npx skills add appleweiping/WEIPING_WIKI --skill experiment-audit -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install appleweiping/WEIPING_WIKI experiment-audit --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/appleweiping/WEIPING_WIKI.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/aris/skills/experiment-audit .claude/skills/experiment-audit && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
experiment-audit
GitHub stars
119
Token cost
~999 tokens
SKILL.md length
431 words
Files
1
Skills in repo
51
Repo updated
First seen
Licence
MIT

At a glance

Rigorous audit of experiment results before they become paper claims.

  • Works in 5 steps: Statistical Validity → Reproducibility Check → Fair Comparison → …
  • User says experiment-audit
  • SKILL.md covers Decision Gate, Phase 1 — Statistical Validity, Phase 2 — Reproducibility Check and Phase 3 — Fair Comparison, plus 4 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Experiment Audit is an agent skill from appleweiping/WEIPING_WIKI. Rigorous audit of experiment results before they become paper claims. Checks statistical validity, reproducibility, fair comparison, and evidence quality. Use when user says "experiment-audit", "审核实验", "audit results", "check my results", or after run-experiment produces results that need validation.

Its SKILL.md is about 1000 tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Research & Science, covering Reproducible research. The repository describes itself as: knowledge base managed with an LLM workflow. The licence is MIT.

When your agent uses it

  • User says experiment-audit
  • Check my results
  • After run-experiment produces results that need validation

Example prompts

  • “experiment-audit”
  • “audit results”
  • “check my results”
  • “/experiment-audit”

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Statistical Validity
  2. Reproducibility Check
  3. Fair Comparison
  4. Evidence Classification
  5. Audit Report

What it can do on your machine

Read from SKILL.md and the folder at commit 76fdc42. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are markdown).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Experiment Audit loads about 999 tokens when it runs. Until then it costs about 80 tokens; SKILL.md has 431 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~80
When it runs · the whole SKILL.md, loaded when a task matches
~999

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from appleweiping/WEIPING_WIKI at commit 76fdc42, republished under its MIT licence (© appleweiping). 431 words, ~999 tokens.

Download SKILL.mdSave it as .claude/skills/experiment-audit/SKILL.md (or your agent's skills folder).
name
experiment-audit
description
Rigorous audit of experiment results before they become paper claims. Checks statistical validity, reproducibility, fair comparison, and evidence quality. Use when user says "experiment-audit", "审核实验", "audit results", "check my results", or after run-experiment produces results that need validation.

Experiment Audit

Audit experiment results with the rigor of a hostile reviewer. Your job is to find problems BEFORE submission.

Decision Gate

Before running:

  • Results exist in results/ or refine-logs/EXPERIMENT_TRACKER.md shows completed blocks
  • At least one block has full seed runs (20+ seeds for paper claims)
  • You have access to the experiment configs and code

Phase 1 — Statistical Validity

For each reported result:

  1. Sample size check: Are there 20+ seeds? (Required for paper evidence)
  2. Significance test: Run paired t-test or Wilcoxon signed-rank between our method and each baseline
  3. Effect size: Report Cohen's d or similar — is the improvement meaningful, not just significant?
  4. Confidence intervals: Report 95% CI for all primary metrics
  5. Multiple comparison correction: If testing against 8+ baselines, apply Bonferroni or Holm-Bonferroni

Verdict per comparison: PASS (p<0.05, meaningful effect) / MARGINAL (p<0.1) / FAIL (not significant)

Phase 2 — Reproducibility Check

  1. Config audit: Can you reproduce the exact run from config alone?
  2. Seed sensitivity: Is variance across seeds reasonable? (CV < 20% for stable metrics)
  3. Hardware sensitivity: Would different GPU/batch size change results?
  4. Code-result alignment: Does the code actually implement what the paper claims?

Red flags:

  • Results only work with specific seeds → cherry-picking
  • Variance is huge → unstable method
  • Config doesn't match paper description → misrepresentation

Phase 3 — Fair Comparison

For each baseline:

  1. Same data splits? (Must be identical)
  2. Same preprocessing? (Must be identical)
  3. Same compute budget? (Comparable training time/FLOPs)
  4. Best hyperparameters? (Did you tune baselines fairly, or use defaults while tuning yours?)
  5. Official numbers match? (If using official implementation, do you reproduce their reported numbers?)

Red flags:

  • Our method gets 10x more compute → unfair
  • Baselines use default hyperparams while ours is tuned → unfair
  • Different data splits → incomparable
Show full SKILL.md (144 more words)Show less

Phase 4 — Evidence Classification

Label every result:

LabelMeaningCan use in paper?
paper_resultFull seeds, statistically valid, fair comparisonYES
officialFull seeds but needs one more checkAlmost
diagnosticPartial seeds or preliminaryNO (supplementary only)
pilotQuick test, not rigorousNO

Only paper_result labeled evidence goes into the main paper.

Phase 5 — Audit Report

Produce a structured audit report:

markdown
# Experiment Audit Report
Date: YYYY-MM-DD
Auditor: Codex (GPT-5.5)

## Summary Verdict: PASS / CONDITIONAL PASS / FAIL

## Per-Block Results
| Block | Seeds | Significance | Effect Size | Fair? | Evidence Label |
|-------|-------|-------------|-------------|-------|----------------|

## Issues Found
1. [CRITICAL/MAJOR/MINOR] Description...

## Recommendations
1. ...

Handoff

  • Output: refine-logs/AUDIT_REPORT.md
  • Update refine-logs/EXPERIMENT_TRACKER.md with audit status
  • If PASS → Next ARIS step: paper-plan
  • If CONDITIONAL PASS → Fix issues, re-audit
  • If FAIL → Back to experiment-bridge or pivot

Hard Rules

  • Never approve results with <20 seeds as paper evidence
  • Never approve unfair comparisons (compute, tuning, data asymmetry)
  • Always run significance tests — "looks better" is not evidence
  • Audit must be done by a different agent than the one who ran experiments (cross-model review)
  • Be adversarial — your job is to find problems, not confirm success

© appleweiping, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/aris/skills/experiment-audit of appleweiping/WEIPING_WIKI.

Open the folder on GitHubat commit 76fdc42

Compare with similar skills

Experiment Audit next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Experiment Audit compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Experiment Audit this skillappleweiping/WEIPING_WIKI119—~999Automated safety check: PassMIT
Peer ReviewK-Dense-AI/claude-scientific-writer2.4k2 repos~3.1kAutomated safety check: NotesMIT
Compute Environment Setupaipoch/open-science5.5k1 repos~2.6kAutomated safety check: PassApache-2.0
CHARLS Paper Reproduction Guidexjtulyc/MedgeClaw6171 repos~1.8kAutomated safety check: PassNone
Figure Styleaipoch/open-science5.5k—~5.1kAutomated safety check: PassApache-2.0
Add Bactopia Toolbactopia/bactopia522—~4.1kAutomated safety check: PassMIT

Similar skills

  • Peer Review

    K-Dense-AI/claude-scientific-writer

    Prepare evidence-bounded, constructive peer-review drafts and structured manuscript assessments.

    2.4k GitHub starsUsed in 2 repos~3.1k tokens
    Research & ScienceAuto-check: notes
  • Compute Environment Setup

    aipoch/open-science

    Prepares setup instructions and a named activation file for a user-managed software environment on an Open-Science SSH or Slurm compute host.

    5.5k GitHub starsUsed in 1 repo~2.6k tokens
    Research & ScienceAuto-check passed
  • Guides an agent through reproducing papers built on the CHARLS health and retirement survey, from variable mapping to cognition, depression and isolation scores.

    617 GitHub starsUsed in 1 repo~1.8k tokens
    Research & ScienceAuto-check passed
  • Figure Style

    aipoch/open-science

    Publication-grade correctness and legibility rules for final-deliverable scientific figures, not exploratory plots.

    5.5k GitHub stars~5.1k tokensUpdated today
    Research & ScienceAuto-check passed
  • Add Bactopia Tool

    bactopia/bactopia

    Scaffold a complete Bactopia Tool across all three tiers -- module, subworkflow, and workflow entry point under workflows/bactopia-tools/.

    522 GitHub stars~4.1k tokensUpdated 2 mo ago
    Research & ScienceAuto-check passed
  • Modeling Code and Result Contracts

    yushui2022/MathModel-Skill

    Generates result-evidence contracts, tables and runnable q1 to q3 modeling code scaffolds for a math modeling paper from a model route, a data plan and cleaned data.

    452 GitHub stars~1.4k tokensUpdated yesterday
    Research & ScienceAuto-check passed

More from appleweiping/WEIPING_WIKI

All 51 skills in this repo
  • Communication Assistant

    appleweiping/WEIPING_WIKI

    Unified lazy-mode communication assistant for Vipin across WhatsApp, WeChat, QQ, Feishu/Lark, and email.

    119 GitHub stars~1.1k tokensUpdated 1 mo ago
    Auto-check passed
  • Content Refinement Agent

    appleweiping/WEIPING_WIKI

    Step 5 of the PaperOrchestra pipeline (arXiv:2604.05018). An agent skill from appleweiping/WEIPING_WIKI.

    119 GitHub stars~3k tokensUpdated 1 mo ago
    Auto-check passed
  • Chrome Automation

    appleweiping/WEIPING_WIKI

    Connect to and control Google Chrome browser using agent-browser with CDP (Chrome DevTools Protocol).

    119 GitHub starsUsed in 1 repo~5.3k tokens
    Auto-check: warnings
  • Email Assistant

    appleweiping/WEIPING_WIKI

    Personal Gmail and Google Workspace email assistant for Vipin.

    119 GitHub stars~1.3k tokensUpdated 1 mo ago
    Auto-check passed
  • Wechat Video Channel Publish

    appleweiping/WEIPING_WIKI

    A skill your agent uses when the user wants to log into 微信视频号, validate cookie state, upload videos, set scheduled publish time, fill long description, set a cover image, or save drafts through a…

    119 GitHub stars~765 tokensUpdated 1 mo ago
    Auto-check passed
  • Feishu Bridge

    appleweiping/WEIPING_WIKI

    Route Feishu/Lark content access for Codex. An agent skill from appleweiping/WEIPING_WIKI.

    119 GitHub stars~906 tokensUpdated 1 mo ago
    Auto-check passed

Questions about Experiment Audit

What does Experiment Audit do?

Rigorous audit of experiment results before they become paper claims. Experiment Audit is an agent skill from appleweiping/WEIPING_WIKI. Rigorous audit of experiment results before they become paper claims.

When should I use Experiment Audit?

Experiment Audit fits situations like: user says experiment-audit; check my results; after run-experiment produces results that need validation.

How do I install Experiment Audit in Claude Code?

Run `npx skills add appleweiping/WEIPING_WIKI --skill experiment-audit -a claude-code`. Or copy the skill folder (.claude/skills/aris/skills/experiment-audit in appleweiping/WEIPING_WIKI) into .claude/skills/experiment-audit in your project. Claude Code loads it when a task matches its description.

How do I install Experiment Audit in Codex?

Run `npx skills add appleweiping/WEIPING_WIKI --skill experiment-audit -a codex`. Or copy the skill folder (.claude/skills/aris/skills/experiment-audit in appleweiping/WEIPING_WIKI) into .agents/skills/experiment-audit in your project. Codex loads it when a task matches its description.

Can I use Experiment Audit in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add appleweiping/WEIPING_WIKI --skill experiment-audit -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/experiment-audit, .gemini/skills/experiment-audit, .github/skills/experiment-audit and .opencode/skills/experiment-audit in your project.

What does Experiment Audit need to run?

SKILL.md names no scripts, command-line tools or credentials: Experiment Audit is instructions for the agent only.

Does Experiment Audit access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Experiment Audit safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Experiment Audit use?

Experiment Audit is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Experiment Audit use?

About 999 tokens (SKILL.md is roughly 4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Experiment Audit?

Skills that share tags, products or a category with Experiment Audit: Peer Review (K-Dense-AI/claude-scientific-writer, 2.4k stars), Compute Environment Setup (aipoch/open-science, 5.5k stars), CHARLS Paper Reproduction Guide (xjtulyc/MedgeClaw, 617 stars) and Figure Style (aipoch/open-science, 5.5k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Experiment Audit?

appleweiping (a GitHub user) maintains it in appleweiping/WEIPING_WIKI, which has 119 GitHub stars. The repository holds 51 skills in this directory. The repository was last updated on August 26, 2026.

Source: appleweiping/WEIPING_WIKI on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.