Agent skill

Aris Experiment Audit

by appleweiping in appleweiping/WEIPING_WIKI

Full statistical audit of experiment results. An agent skill from appleweiping/WEIPING_WIKI.

MITAuto-check passedResearch & Science

Install Aris Experiment Audit

skills CLI
$ npx skills add appleweiping/WEIPING_WIKI --skill aris-experiment-audit -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install appleweiping/WEIPING_WIKI aris-experiment-audit --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/appleweiping/WEIPING_WIKI.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.codex/skills/aris-experiment-audit .claude/skills/aris-experiment-audit && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
aris-experiment-audit
GitHub stars
119
Token cost
~2.3k tokens
SKILL.md length
938 words
Files
1
Skills in repo
51
Repo updated
First seen
Licence
MIT

At a glance

Full statistical audit of experiment results. An agent skill from appleweiping/WEIPING_WIKI.

  • Works in 5 steps: Statistical Validity → Reproducibility Audit → Fair Comparison Audit → …
  • Tasks that involve Reproducible research
  • SKILL.md covers Your Mandate, Phase 1: Statistical Validity, Phase 2: Reproducibility Audit and Phase 3: Fair Comparison Audit, plus 3 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Aris Experiment Audit is an agent skill from appleweiping/WEIPING_WIKI. Full statistical audit of experiment results. This is Codex's PRIMARY role in ARIS. Check significance, effect sizes, fair comparison, reproducibility, evidence labeling. Triggers: "experiment-audit", "audit results", "check my results", "审核实验", "statistical review", "are these results valid", "can I publish this"

Its SKILL.md is about 2.3k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Research & Science, covering Reproducible research and Statistics. The repository describes itself as: knowledge base managed with an LLM workflow. The licence is MIT.

When your agent uses it

  • Tasks that involve Reproducible research
  • Tasks that involve Statistics

Example prompts

  • “experiment-audit”
  • “audit results”
  • “check my results”
  • “/aris-experiment-audit”

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Statistical Validity
  2. Reproducibility Audit
  3. Fair Comparison Audit
  4. Evidence Labeling
  5. Structured Audit Report

What it can do on your machine

Read from SKILL.md and the folder at commit 76fdc42. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Aris Experiment Audit loads about 2.3k tokens when it runs. Until then it costs about 84 tokens; SKILL.md has 938 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~84
When it runs · the whole SKILL.md, loaded when a task matches
~2.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from appleweiping/WEIPING_WIKI at commit 76fdc42, republished under its MIT licence (© appleweiping). 938 words, ~2,321 tokens.

Download SKILL.mdSave it as .claude/skills/aris-experiment-audit/SKILL.md (or your agent's skills folder).
name
aris-experiment-audit
description
Full statistical audit of experiment results. This is Codex's PRIMARY role in ARIS. Check significance, effect sizes, fair comparison, reproducibility, evidence labeling. Triggers: "experiment-audit", "audit results", "check my results", "审核实验", "statistical review", "are these results valid", "can I publish this"
role
auditor
stage
experiment-audit

ARIS Experiment-Audit: Full Statistical Audit

This is your PRIMARY skill. When results come in, you are the last line of defense before claims enter a paper. You determine whether evidence is real or noise, whether comparisons are fair, and whether the results deserve publication.

Your Mandate

  • NO result enters a paper without your audit stamp
  • You protect the researcher from publishing false claims
  • You protect the field from irreproducible results
  • A "PASS" from you means: "I would defend these results under cross-examination"

Phase 1: Statistical Validity

1.1 Significance Testing

For EACH claimed improvement, verify:

ClaimTest Usedp-valueThresholdSignificant?Appropriate Test?
[claim][test][p][alpha]Yes/NoYes/No
Required Checks
  • Correct test selected for data distribution
    • Normal data → paired t-test or ANOVA
    • Non-normal → Wilcoxon signed-rank or Mann-Whitney U
    • Multiple comparisons → Bonferroni/Holm correction applied
  • Two-tailed test used (unless one-tailed pre-registered)
  • Sample size sufficient for claimed effect (power analysis)
  • Independence assumption holds (seeds are truly independent runs)
  • Multiple comparison correction applied if >3 comparisons
Significance Red Flags
FlagSeverityAction
p = 0.04-0.05 with no correctionHIGHRequire more seeds or correction
Only best seed reportedCRITICALInvalidate result
Significance claimed without testCRITICALCannot publish
Different N across methodsHIGHExplain or equalize
p-hacking pattern (many metrics, report best)CRITICALRequire pre-registration
1.2 Effect Size

Raw p-values are insufficient. For each significant result:

ComparisonEffect Size MetricValueInterpretation
[A vs B]Cohen's d / eta^2 / CLES[value]Negligible/Small/Medium/Large
Effect Size Interpretation (Cohen's d)
  • |d| < 0.2: Negligible — not practically meaningful even if significant
  • 0.2 <= |d| < 0.5: Small — real but may not matter in practice
  • 0.5 <= |d| < 0.8: Medium — meaningful improvement
  • |d| >= 0.8: Large — substantial improvement

Hard rule: If effect size is "Negligible" → result cannot be a main claim.

1.3 Confidence Intervals
  • 95% CI reported for all main results
  • CIs do not overlap for claimed improvements (or overlap is minimal)
  • CI width is reasonable (not so wide as to be uninformative)
  • Bootstrap CIs used if distribution is unclear

Phase 2: Reproducibility Audit

2.1 Seed Variance Analysis
MethodMeanStdCV (%)MinMaxOutlier Seeds?
Proposed
Baseline 1
Baseline 2
Variance Checks
  • CV < 5% for stable methods (flag if higher)
  • No single seed dominates the mean (remove-one-seed sensitivity)
  • Variance is similar across methods (heteroscedasticity check)
  • Number of seeds >= 20 for paper results, >= 5 for pilot
  • All seeds completed (no silent failures)
Variance Red Flags
FlagSeverityImplication
CV > 10%HIGHResults unstable, need more seeds
One seed is 2+ std from meanMEDIUMInvestigate outlier
Proposed method has lower variance than baselinesMEDIUMMay indicate overfitting to eval
Some seeds failed/missingHIGHSurvivorship bias
2.2 Config Alignment Verification
  • Reported config matches actual run config (check logs)
  • No manual overrides during execution
  • Hardware matches what's reported
  • Library versions are pinned and recorded

Phase 3: Fair Comparison Audit

3.1 Condition Equality

For each pair of methods being compared:

ConditionMethod AMethod BEqual?
Training data
Preprocessing
Compute budget (FLOPs or time)
Hyperparameter tuning budget
Number of parameters
Evaluation data
Evaluation metric implementation
Random seeds used
Hardware
3.2 Unfair Advantage Detection

Check for these common sources of unfair advantage:

  • Proposed method does NOT get more tuning iterations
  • Proposed method does NOT use additional data (pre-training, augmentation)
  • Proposed method does NOT have more parameters without acknowledgment
  • Baselines are NOT using outdated implementations
  • Baselines are NOT using known-bad hyperparameters
  • Evaluation metric is NOT custom-designed to favor proposed method
  • Test set was NOT used for any development decisions
Show full SKILL.md (370 more words)Show less
3.3 Comparison Verdict
COMPARISON FAIRNESS: [FAIR / MOSTLY_FAIR / UNFAIR]
ADVANTAGES DETECTED: [list]
DISADVANTAGES TO BASELINES: [list]
RECOMMENDATION: [proceed / equalize / re-run]

Phase 4: Evidence Labeling

Every result must be labeled with its evidence tier. This determines where it can appear.

Evidence Tier Definitions
TierLabelRequirementsCan Appear In
T1paper_result>=20 seeds, significance test passes, effect size reported, fair comparisonMain paper claims
T2official>=10 seeds, significance test passes, fair comparisonPaper with caveat
T3diagnostic>=5 seeds, informative but not rigorous enough for claimsAppendix only
T4pilot<5 seeds or exploratoryInternal use only
Labeling Checklist

For each result, assign tier:

ResultSeedsSig. TestEffect SizeFairTier
[result 1]Pass/FailReported?Yes/NoT?
[result 2]Pass/FailReported?Yes/NoT?
Labeling Rules
  • A result CANNOT be labeled higher than its weakest dimension allows
  • If seeds < 20 → maximum T2
  • If no significance test → maximum T3
  • If comparison is unfair → maximum T3
  • If effect size is negligible → maximum T3 regardless of other factors
  • Pilot results (T4) must NEVER appear in paper, even in appendix

Phase 5: Structured Audit Report

Summary Scores
DimensionScoreStatus
Statistical Validity/10
Reproducibility/10
Fair Comparison/10
Evidence Quality/10
Paper Readiness/10
Scoring Rubric
  • 1-3: Results are unreliable. Cannot be published in any form.
  • 4-5: Results have significant gaps. Major revisions needed.
  • 6-7: Results are acceptable with noted limitations.
  • 8-9: Results are strong. Minor notes only.
  • 10: Results are exemplary. Would survive any review.
Verdict Decision Matrix
ConditionVerdict
All >= 8, all T1 results have full evidencePASS — ready for paper
All >= 6, most results T1/T2CONDITIONAL PASS — address notes
Any dimension <= 4FAIL — fix before any publication
Statistical validity <= 5FAIL — fundamental stats issues
Fair comparison <= 5FAIL — re-run with equalized conditions
Only T3/T4 resultsINSUFFICIENT — need more seeds/rigor
Final Verdict Format
========================================
EXPERIMENT AUDIT VERDICT
========================================
VERDICT: [PASS / CONDITIONAL PASS / FAIL / INSUFFICIENT]
CONFIDENCE: [Low / Medium / High]

SCORES:
  Statistical Validity: X/10
  Reproducibility:      X/10
  Fair Comparison:      X/10
  Evidence Quality:     X/10
  Paper Readiness:      X/10
  AVERAGE:              X.X/10

EVIDENCE TIERS:
  T1 (paper_result): [count] results
  T2 (official):     [count] results
  T3 (diagnostic):   [count] results
  T4 (pilot):        [count] results

CRITICAL ISSUES: [count]
[list each]

HIGH ISSUES: [count]
[list each]

BLOCKING FOR PUBLICATION: [Yes/No]
BLOCKING ISSUE: [one-line or "None"]

RECOMMENDED ACTIONS:
1. [most important fix]
2. [second fix]
3. [third fix]
========================================

Interaction Rules

  • This is your most important role. Be thorough. Miss nothing.
  • Never rubber-stamp results. Even good results deserve full audit.
  • If you cannot verify something (e.g., you don't have access to logs), state it explicitly.
  • Distinguish between "I verified this is correct" and "I cannot check this."
  • If results are genuinely strong, say so clearly. Researchers need confidence too.
  • Track audit history. If results were previously FAIL and are resubmitted, verify fixes.
  • When in doubt, request more seeds. Seeds are cheap; retracted papers are expensive.

© appleweiping, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .codex/skills/aris-experiment-audit of appleweiping/WEIPING_WIKI.

Open the folder on GitHubat commit 76fdc42

Compare with similar skills

Aris Experiment Audit next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Aris Experiment Audit compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Aris Experiment Audit this skillappleweiping/WEIPING_WIKI119—~2.3kAutomated safety check: PassMIT
Peer ReviewK-Dense-AI/claude-scientific-writer2.4k2 repos~3.1kAutomated safety check: NotesMIT
Academic Paper Reproduction Methodologyxjtulyc/MedgeClaw6171 repos~1.3kAutomated safety check: PassNone
Bio Experimental Design Multiple TestingGPTomics/bioSkills1.2k1 repos~3.5kAutomated safety check: PassMIT
Scientific Critical Thinkingjaechang-hits/SciAgent-Skills3711 repos~4.7kAutomated safety check: PassCC-BY-4.0
Fcr Revision And Rebuttalfranklee16/academic-research-skills2231 repos~1.1kAutomated safety check: PassNone

Similar skills

  • Peer Review

    K-Dense-AI/claude-scientific-writer

    Prepare evidence-bounded, constructive peer-review drafts and structured manuscript assessments.

    2.4k GitHub starsUsed in 2 repos~3.1k tokens
    Research & ScienceAuto-check: notes
  • Six-phase process for reproducing a published paper's results from provided data, from variable mapping and sample filtering through regression tables and a written report.

    617 GitHub starsUsed in 1 repo~1.3k tokens
    Research & ScienceAuto-check passed
  • Controls error rates across thousands of simultaneous tests in genomics discovery using false-discovery-rate methods (Benjamini-Hochberg 1995; Benjamini-Yekutieli 2001 for arbitrary dependence…

    1.2k GitHub starsUsed in 1 repo~3.5k tokens
    Research & ScienceAuto-check passed
  • Scientific Critical Thinking

    jaechang-hits/SciAgent-Skills

    Evaluating scientific evidence and claims. An agent skill from jaechang-hits/SciAgent-Skills.

    371 GitHub starsUsed in 1 repo~4.7k tokens
    Research & ScienceAuto-check passed
  • Fcr Revision And Rebuttal

    franklee16/academic-research-skills

    A skill your agent uses when writing the response to a Field Crops Research (FCR) revision decision (major or minor) and revising the manuscript.

    223 GitHub starsUsed in 1 repo~1.1k tokens
    Research & ScienceAuto-check passed
  • Aistats Writing Style

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when revising an AISTATS paper for concise AI-statistics framing, theorem-and-experiment clarity, 8-page two-column compression, double-blind wording, reproducibility…

    1.2k GitHub stars~865 tokensUpdated 12 days ago
    Research & ScienceAuto-check passed

More from appleweiping/WEIPING_WIKI

All 51 skills in this repo
  • Communication Assistant

    appleweiping/WEIPING_WIKI

    Unified lazy-mode communication assistant for Vipin across WhatsApp, WeChat, QQ, Feishu/Lark, and email.

    119 GitHub stars~1.1k tokensUpdated 1 mo ago
    Auto-check passed
  • Content Refinement Agent

    appleweiping/WEIPING_WIKI

    Step 5 of the PaperOrchestra pipeline (arXiv:2604.05018). An agent skill from appleweiping/WEIPING_WIKI.

    119 GitHub stars~3k tokensUpdated 1 mo ago
    Auto-check passed
  • Chrome Automation

    appleweiping/WEIPING_WIKI

    Connect to and control Google Chrome browser using agent-browser with CDP (Chrome DevTools Protocol).

    119 GitHub starsUsed in 1 repo~5.3k tokens
    Auto-check: warnings
  • Email Assistant

    appleweiping/WEIPING_WIKI

    Personal Gmail and Google Workspace email assistant for Vipin.

    119 GitHub stars~1.3k tokensUpdated 1 mo ago
    Auto-check passed
  • Wechat Video Channel Publish

    appleweiping/WEIPING_WIKI

    A skill your agent uses when the user wants to log into 微信视频号, validate cookie state, upload videos, set scheduled publish time, fill long description, set a cover image, or save drafts through a…

    119 GitHub stars~765 tokensUpdated 1 mo ago
    Auto-check passed
  • Feishu Bridge

    appleweiping/WEIPING_WIKI

    Route Feishu/Lark content access for Codex. An agent skill from appleweiping/WEIPING_WIKI.

    119 GitHub stars~906 tokensUpdated 1 mo ago
    Auto-check passed

Questions about Aris Experiment Audit

What does Aris Experiment Audit do?

Full statistical audit of experiment results. An agent skill from appleweiping/WEIPING_WIKI. Aris Experiment Audit is an agent skill from appleweiping/WEIPING_WIKI. Full statistical audit of experiment results.

When should I use Aris Experiment Audit?

Aris Experiment Audit fits situations like: tasks that involve Reproducible research; tasks that involve Statistics.

How do I install Aris Experiment Audit in Claude Code?

Run `npx skills add appleweiping/WEIPING_WIKI --skill aris-experiment-audit -a claude-code`. Or copy the skill folder (.codex/skills/aris-experiment-audit in appleweiping/WEIPING_WIKI) into .claude/skills/aris-experiment-audit in your project. Claude Code loads it when a task matches its description.

How do I install Aris Experiment Audit in Codex?

Run `npx skills add appleweiping/WEIPING_WIKI --skill aris-experiment-audit -a codex`. Or copy the skill folder (.codex/skills/aris-experiment-audit in appleweiping/WEIPING_WIKI) into .agents/skills/aris-experiment-audit in your project. Codex loads it when a task matches its description.

Can I use Aris Experiment Audit in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add appleweiping/WEIPING_WIKI --skill aris-experiment-audit -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/aris-experiment-audit, .gemini/skills/aris-experiment-audit, .github/skills/aris-experiment-audit and .opencode/skills/aris-experiment-audit in your project.

What does Aris Experiment Audit need to run?

SKILL.md names no scripts, command-line tools or credentials: Aris Experiment Audit is instructions for the agent only.

Does Aris Experiment Audit access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Aris Experiment Audit safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Aris Experiment Audit use?

Aris Experiment Audit is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Aris Experiment Audit use?

About 2.3k tokens (SKILL.md is roughly 9.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Aris Experiment Audit?

Skills that share tags, products or a category with Aris Experiment Audit: Peer Review (K-Dense-AI/claude-scientific-writer, 2.4k stars), Academic Paper Reproduction Methodology (xjtulyc/MedgeClaw, 617 stars), Bio Experimental Design Multiple Testing (GPTomics/bioSkills, 1.2k stars) and Scientific Critical Thinking (jaechang-hits/SciAgent-Skills, 371 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Aris Experiment Audit?

appleweiping (a GitHub user) maintains it in appleweiping/WEIPING_WIKI, which has 119 GitHub stars. The repository holds 51 skills in this directory. The repository was last updated on August 26, 2026.

Source: appleweiping/WEIPING_WIKI on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.