Agent skill

Aris Experiment Plan

by appleweiping in appleweiping/WEIPING_WIKI

Review and score experiment plans produced by Claude Code. An agent skill from appleweiping/WEIPING_WIKI.

MITAuto-check passedResearch & Science

Install Aris Experiment Plan

skills CLI
$ npx skills add appleweiping/WEIPING_WIKI --skill aris-experiment-plan -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install appleweiping/WEIPING_WIKI aris-experiment-plan --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/appleweiping/WEIPING_WIKI.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.codex/skills/aris-experiment-plan .claude/skills/aris-experiment-plan && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
aris-experiment-plan
GitHub stars
119
Token cost
~1.4k tokens
SKILL.md length
619 words
Files
1
Skills in repo
51
Repo updated
First seen
Licence
MIT

At a glance

Review and score experiment plans produced by Claude Code. An agent skill from appleweiping/WEIPING_WIKI.

  • Works in 4 steps: Baseline Completeness Check → Decision Gate Quality → Compute Feasibility → …
  • Tasks that involve Experimental design
  • SKILL.md covers Your Mandate, Phase 1: Baseline Completeness…, Phase 2: Decision Gate Quality and Phase 3: Compute Feasibility, plus 2 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Aris Experiment Plan is an agent skill from appleweiping/WEIPING_WIKI. Review and score experiment plans produced by Claude Code. Evaluate evidence quality, rigor, decision gates, feasibility, and paper potential. Triggers: "review experiment plan", "audit plan", "score plan", "check my experiment design", "is this plan rigorous enough"

Its SKILL.md is about 1.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Research & Science, covering Experimental design. The repository describes itself as: knowledge base managed with an LLM workflow. The licence is MIT.

When your agent uses it

  • Tasks that involve Experimental design

Example prompts

  • “review experiment plan”
  • “audit plan”
  • “score plan”
  • “/aris-experiment-plan”

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Baseline Completeness Check
  2. Decision Gate Quality
  3. Compute Feasibility
  4. Score and Verdict

What it can do on your machine

Read from SKILL.md and the folder at commit 76fdc42. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Aris Experiment Plan loads about 1.4k tokens when it runs. Until then it costs about 72 tokens; SKILL.md has 619 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~72
When it runs · the whole SKILL.md, loaded when a task matches
~1.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from appleweiping/WEIPING_WIKI at commit 76fdc42, republished under its MIT licence (© appleweiping). 619 words, ~1,430 tokens.

Download SKILL.mdSave it as .claude/skills/aris-experiment-plan/SKILL.md (or your agent's skills folder).
name
aris-experiment-plan
description
Review and score experiment plans produced by Claude Code. Evaluate evidence quality, rigor, decision gates, feasibility, and paper potential. Triggers: "review experiment plan", "audit plan", "score plan", "check my experiment design", "is this plan rigorous enough"
role
auditor
stage
experiment-plan

ARIS Experiment-Plan Auditor

You are the AUDITOR for experiment plans. You review the experimental design produced by the implementation agent. You check whether the plan, if executed perfectly, would produce publishable evidence. You do NOT design experiments — you judge them.

Your Mandate

  • A plan that passes your audit should survive peer review methodology critique
  • Catch missing baselines, weak gates, and unrealistic compute assumptions NOW
  • Every "proceed" means: "If they follow this plan exactly, the results are trustworthy"

Phase 1: Baseline Completeness Check

Required Baselines Checklist
  • Random baseline included (sanity check)
  • Current SOTA baseline identified and included
  • At least one "simple but strong" baseline (e.g., well-tuned linear model)
  • Ablation baselines for each proposed component
  • Total baseline count >= 3 (excluding ablations)
Baseline Quality Criteria
CriterionPass/FailNotes
Baselines use same data splits
Baselines use same preprocessing
Baselines are tuned (not strawman)
Baseline implementations are cited/verified
Compute budget for baselines is allocated

Hard rule: If baselines < 3 or any baseline is a strawman → automatic "iterate" verdict.

Seed and Repetition Requirements
  • Number of random seeds specified (minimum: 5 for pilot, 20 for paper results)
  • Seed selection strategy defined (sequential from 0, or pre-registered list)
  • Variance reporting method specified (std dev, CI, IQR)
  • Seeds apply to ALL methods equally (not just proposed method)

Hard rule: If seeds < 20 for paper-targeted experiments → flag as insufficient.

Phase 2: Decision Gate Quality

For each decision gate in the plan, verify:

GateMetricThresholdAction if FailedVerdict
Gate N[specific metric][numeric threshold][specific action]Pass/Fail
Gate Quality Criteria
  • Every gate has a NUMERIC threshold (not "significant improvement")
  • Every gate specifies what happens on failure (pivot/iterate/stop)
  • Gates are ordered — later gates depend on earlier gates passing
  • At least one early "kill gate" (cheap experiment that validates core assumption)
  • No gate relies solely on visual inspection or subjective judgment
Common Gate Failures
  • "If results look promising" → FAIL (not numeric)
  • "If better than baseline" → FAIL (by how much? significance level?)
  • "If compute allows" → FAIL (compute should be pre-allocated)
  • Gate threshold set at exactly SOTA → SUSPICIOUS (should exceed by margin)

Hard rule: If any gate lacks a numeric threshold → automatic "iterate" verdict.

Phase 3: Compute Feasibility

Resource Audit
ResourceRequiredAvailableMarginStatus
GPU hoursOK/Tight/Infeasible
StorageOK/Tight/Infeasible
Wall-clock timeOK/Tight/Infeasible
API costs (if any)OK/Tight/Infeasible
Show full SKILL.md (244 more words)Show less
Feasibility Checks
  • Total compute estimated for ALL runs (method * baselines * seeds * gates)
  • 30% buffer included for failed runs and debugging
  • Longest single run fits within available wall-clock
  • Data download/preprocessing time accounted for
  • Checkpoint strategy defined (can resume from failure)

Hard rule: If any resource is "Infeasible" → automatic "pivot" verdict.

Phase 4: Score and Verdict

Scoring
DimensionScoreJustification
Evidence Quality/10Will the results be believed?
Rigor/10Are comparisons fair and complete?
Gates/10Will bad directions be caught early?
Feasibility/10Can this actually be executed?
Paper Potential/10Does this plan produce a publishable story?
Scoring Rubric
  • 1-3: Plan has structural flaws. Cannot produce trustworthy results.
  • 4-5: Plan is incomplete. Key elements missing.
  • 6-7: Plan is workable but has gaps that could weaken the paper.
  • 8-9: Plan is solid. Minor suggestions only.
  • 10: Plan is exemplary. Would use as a template.
Decision Matrix
ConditionVerdict
All dimensions >= 7, all hard rules passPROCEED
Any hard rule violatedITERATE (fix specific issue)
Any dimension <= 3PIVOT (plan needs fundamental redesign)
Average >= 7 but one dimension is 5-6ITERATE (targeted fix)
Average < 5PIVOT
Verdict Format
VERDICT: [PROCEED / ITERATE / PIVOT]
CONFIDENCE: [Low / Medium / High]
SCORES: E={evidence} R={rigor} G={gates} F={feasibility} P={paper} AVG={avg}
HARD RULE VIOLATIONS: [list or "None"]
BLOCKING ISSUE: [one-line summary or "None"]
NEXT ACTION: [specific fix the implementation agent should make]

Interaction Rules

  • Never design experiments. Your job is to judge plans, not create them.
  • Be specific: "Gate 3 needs a threshold" not "gates need work."
  • If a plan is genuinely good, say so. Don't manufacture criticism.
  • Track iteration count. If a plan has been through 3+ iterations without passing, recommend stepping back to re-examine the research question.

© appleweiping, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .codex/skills/aris-experiment-plan of appleweiping/WEIPING_WIKI.

Open the folder on GitHubat commit 76fdc42

Compare with similar skills

Aris Experiment Plan next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Aris Experiment Plan compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Aris Experiment Plan this skillappleweiping/WEIPING_WIKI119—~1.4kAutomated safety check: PassMIT
Scientific Critical Thinkingweapp-tailwindcss/weapp-tailwindcss1.9k22 repos~5.9kAutomated safety check: NotesMIT
Benchmark Paper TemplateHKUSTDial/Supervisor-Skills8.8k—~2.8kAutomated safety check: PassCC-BY-4.0
Claim-Driven Experiment PlannerzjYao36/Auto-Research-Refine1286 repos~2.3kAutomated safety check: NotesNone
Research Refine PipelinezjYao36/Auto-Research-Refine1285 repos~1.4kAutomated safety check: NotesNone
Metabolic Study Planneraiming-lab/AutoResearchClaw15k—~1.9kAutomated safety check: PassMIT

Similar skills

  • Scientific Critical Thinking

    weapp-tailwindcss/weapp-tailwindcss

    Evaluate research rigor. An agent skill from weapp-tailwindcss/weapp-tailwindcss.

    1.9k GitHub starsUsed in 22 repos~5.9k tokens
    Research & ScienceAuto-check: notes
  • Benchmark Paper Template

    HKUSTDial/Supervisor-Skills

    Structures benchmark and evaluation papers around five pillars, with a completeness audit, an Introduction logic chain, a section skeleton and a pre-submission checklist.

    8.8k GitHub stars~2.8k tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed
  • Claim-Driven Experiment Planner

    zjYao36/Auto-Research-Refine

    Turns a refined research proposal into a claim-to-evidence-to-run-order roadmap instead of a sprawling benchmark wishlist.

    128 GitHub starsUsed in 6 repos~2.3k tokens
    Research & ScienceAuto-check: notes
  • Research Refine Pipeline

    zjYao36/Auto-Research-Refine

    Chains research-refine and experiment-plan to turn a vague research direction into a focused proposal and a claim-driven experiment roadmap.

    128 GitHub starsUsed in 5 repos~1.4k tokens
    Research & ScienceAuto-check: notes
  • Metabolic Study Planner

    aiming-lab/AutoResearchClaw

    Turns a broad metabolic modelling topic into a concrete, paper-shaped plan with organism, model, perturbations, metrics and figures before any FBA code is written.

    15k GitHub stars~1.9k tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed
  • Analytical Method Validation Planner

    K-Dense-AI/scientific-agent-skills

    Plans, runs, and documents analytical method validation, verification, or transfer studies under ICH Q2(R2)/Q14, USP, ICH M10, CLSI EP, or ISO/IEC 17025.

    48k GitHub starsUsed in 1 repo~4.9k tokens
    Research & ScienceAuto-check: notes

More from appleweiping/WEIPING_WIKI

All 51 skills in this repo
  • Communication Assistant

    appleweiping/WEIPING_WIKI

    Unified lazy-mode communication assistant for Vipin across WhatsApp, WeChat, QQ, Feishu/Lark, and email.

    119 GitHub stars~1.1k tokensUpdated 1 mo ago
    Auto-check passed
  • Content Refinement Agent

    appleweiping/WEIPING_WIKI

    Step 5 of the PaperOrchestra pipeline (arXiv:2604.05018). An agent skill from appleweiping/WEIPING_WIKI.

    119 GitHub stars~3k tokensUpdated 1 mo ago
    Auto-check passed
  • Chrome Automation

    appleweiping/WEIPING_WIKI

    Connect to and control Google Chrome browser using agent-browser with CDP (Chrome DevTools Protocol).

    119 GitHub starsUsed in 1 repo~5.3k tokens
    Auto-check: warnings
  • Email Assistant

    appleweiping/WEIPING_WIKI

    Personal Gmail and Google Workspace email assistant for Vipin.

    119 GitHub stars~1.3k tokensUpdated 1 mo ago
    Auto-check passed
  • Wechat Video Channel Publish

    appleweiping/WEIPING_WIKI

    A skill your agent uses when the user wants to log into 微信视频号, validate cookie state, upload videos, set scheduled publish time, fill long description, set a cover image, or save drafts through a…

    119 GitHub stars~765 tokensUpdated 1 mo ago
    Auto-check passed
  • Feishu Bridge

    appleweiping/WEIPING_WIKI

    Route Feishu/Lark content access for Codex. An agent skill from appleweiping/WEIPING_WIKI.

    119 GitHub stars~906 tokensUpdated 1 mo ago
    Auto-check passed

Questions about Aris Experiment Plan

What does Aris Experiment Plan do?

Review and score experiment plans produced by Claude Code. An agent skill from appleweiping/WEIPING_WIKI. Aris Experiment Plan is an agent skill from appleweiping/WEIPING_WIKI. Review and score experiment plans produced by Claude Code.

When should I use Aris Experiment Plan?

Aris Experiment Plan fits situations like: tasks that involve Experimental design.

How do I install Aris Experiment Plan in Claude Code?

Run `npx skills add appleweiping/WEIPING_WIKI --skill aris-experiment-plan -a claude-code`. Or copy the skill folder (.codex/skills/aris-experiment-plan in appleweiping/WEIPING_WIKI) into .claude/skills/aris-experiment-plan in your project. Claude Code loads it when a task matches its description.

How do I install Aris Experiment Plan in Codex?

Run `npx skills add appleweiping/WEIPING_WIKI --skill aris-experiment-plan -a codex`. Or copy the skill folder (.codex/skills/aris-experiment-plan in appleweiping/WEIPING_WIKI) into .agents/skills/aris-experiment-plan in your project. Codex loads it when a task matches its description.

Can I use Aris Experiment Plan in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add appleweiping/WEIPING_WIKI --skill aris-experiment-plan -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/aris-experiment-plan, .gemini/skills/aris-experiment-plan, .github/skills/aris-experiment-plan and .opencode/skills/aris-experiment-plan in your project.

What does Aris Experiment Plan need to run?

SKILL.md names no scripts, command-line tools or credentials: Aris Experiment Plan is instructions for the agent only.

Does Aris Experiment Plan access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Aris Experiment Plan safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Aris Experiment Plan use?

Aris Experiment Plan is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Aris Experiment Plan use?

About 1.4k tokens (SKILL.md is roughly 5.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Aris Experiment Plan?

Skills that share tags, products or a category with Aris Experiment Plan: Scientific Critical Thinking (weapp-tailwindcss/weapp-tailwindcss, 1.9k stars), Benchmark Paper Template (HKUSTDial/Supervisor-Skills, 8.8k stars), Claim-Driven Experiment Planner (zjYao36/Auto-Research-Refine, 128 stars) and Research Refine Pipeline (zjYao36/Auto-Research-Refine, 128 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Aris Experiment Plan?

appleweiping (a GitHub user) maintains it in appleweiping/WEIPING_WIKI, which has 119 GitHub stars. The repository holds 51 skills in this directory. The repository was last updated on August 26, 2026.

Source: appleweiping/WEIPING_WIKI on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.