Agent skill

Experiment

by ai-analyst-lab in ai-analyst-lab/ai-analyst

The analysis and lifecycle owner for experiments. An agent skill from ai-analyst-lab/ai-analyst.

MITAuto-check passedMarketing & SEO

Install Experiment

skills CLI
$ npx skills add ai-analyst-lab/ai-analyst --skill experiment -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install ai-analyst-lab/ai-analyst experiment --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/ai-analyst-lab/ai-analyst.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/experiment .claude/skills/experiment && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
experiment
GitHub stars
304
Token cost
~2k tokens
SKILL.md length
625 words
Files
1
Skills in repo
43
Repo updated
First seen
Licence
MIT

At a glance

The analysis and lifecycle owner for experiments. An agent skill from ai-analyst-lab/ai-analyst.

  • Works in 3 steps: Run Experiment Brief skill to capture… → Invoke Experiment Designer agent → Output:…
  • Treatment vs control
  • SKILL.md covers Purpose, When to Use, Modes and State Management, plus 2 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Experiment is an agent skill from ai-analyst-lab/ai-analyst. The analysis and lifecycle owner for experiments. Full experiment lifecycle: design, power analysis, statistical analysis, interpretation, reporting, and monitoring of A/B tests. Invoke as /experiment. Trigger on "A/B test", "experiment", "treatment vs control", "sample size", "MDE", "statistical significance", "ship decision", "test readout", "is this result significant?". Runs the SRM gate first.

Its SKILL.md is about 2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Marketing & SEO, covering A/B testing, Experimental design and Statistics. The repository describes itself as: AI Product Analyst — Claude Code-powered data analysis toolkit. The licence is MIT.

When your agent uses it

  • Treatment vs control
  • Statistical significance
  • Is this result significant?

Example prompts

  • “A/B test”
  • “experiment”
  • “treatment vs control”
  • “/experiment”

Requirements

  • Python 3

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. Run Experiment Brief skill to capture hypothesis, north star, guardrails
  2. Invoke Experiment Designer agent
  3. Output: experiments/{slug}/experiment.yaml (from templates/experiment.yaml)

What it can do on your machine

Read from SKILL.md and the folder at commit 52c0744. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Experiment loads about 2k tokens when it runs. Until then it costs about 103 tokens; SKILL.md has 625 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~103
When it runs · the whole SKILL.md, loaded when a task matches
~2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from ai-analyst-lab/ai-analyst at commit 52c0744, republished under its MIT licence (© ai-analyst-lab). 625 words, ~1,979 tokens.

Download SKILL.mdSave it as .claude/skills/experiment/SKILL.md (or your agent's skills folder).
name
experiment
description
The analysis and lifecycle owner for experiments. Full experiment lifecycle: design, power analysis, statistical analysis, interpretation, reporting, and monitoring of A/B tests. Invoke as /experiment. Trigger on "A/B test", "experiment", "treatment vs control", "sample size", "MDE", "statistical significance", "ship decision", "test readout", "is this result significant?". Runs the SRM gate first.

Skill: /experiment — OpenXP Experimentation Platform

Purpose

Multi-mode skill for the full experiment lifecycle — from design through analysis to ship/no-ship decision. Orchestrates experiment agents and calls coded statistical helpers from helpers/stats/experiment_stats/ instead of improvising Python.

When to Use

Invoke as /experiment [mode] or trigger on experiment-related intents:

  • "I want to run an experiment"
  • "Analyze this A/B test"
  • "Did this experiment work?"
  • "What's the power for this test?"

Modes

/experiment design

Purpose: Create a pre-registered experiment config. Agent: agents/experiments/experiment-designer.md Flow:

  1. Run Experiment Brief skill to capture hypothesis, north star, guardrails
  2. Invoke Experiment Designer agent
  3. Output: experiments/{slug}/experiment.yaml (from templates/experiment.yaml) Checkpoint: Config review (Type B — skippable with --just-do-it)
/experiment power

Purpose: Power analysis + duration estimation. Flow:

  1. Read experiments/{slug}/experiment.yaml for metric type, baseline, MDE
  2. Call helpers/stats/experiment_stats/power.py:
    • Proportion metric → power_proportion(baseline_rate, mde)
    • Continuous metric → power_mean(baseline_mean, baseline_std, mde)
  3. Call duration_estimate(total_sample, daily_traffic, allocation)
  4. Update experiment.yaml with computed values (sample_size, duration, viable)
  5. If NOT_VIABLE → suggest /causal select as alternative Checkpoint: Power viability (Type C — NOT_VIABLE fires mandatory checkpoint)
/experiment analyze

Purpose: Run statistical tests on experiment data. Agent: agents/experiments/experiment-analyzer.md Flow:

  1. Read experiments/{slug}/experiment.yaml for pre-registered config
  2. SRM Gate (mandatory first step):
    python
    from helpers.stats.experiment_stats import srm_check
    # Positional lists ONLY — do not pass dicts.
    # First arg: observed counts per variant (order must match expected_ratios).
    # Second arg: expected allocation ratios, summing to 1.0.
    result = srm_check([4218, 4196], [0.5, 0.5])
    # result = {"chi2_stat": 0.058, "p_value": 0.81, "verdict": "PASS", ...}
    if result["verdict"] == "BLOCK":
        # HALT — do not proceed to treatment effect analysis
  3. Treatment effect analysis using coded helpers:
    python
    from helpers.stats.experiment_stats import welch_test, proportion_test, ratio_metric_test
    # Select based on metric type from experiment.yaml
    if metric_type == "proportion":
        result = proportion_test(c_success, c_n, t_success, t_n)
    elif metric_type == "continuous":
        result = welch_test(control_values, treatment_values)
    elif metric_type == "ratio":
        result = ratio_metric_test(num_c, den_c, num_t, den_t)
  4. Effect size: cohens_d(control, treatment)
  5. Multiple comparisons: adjust_pvalues(all_p_values, method="holm")
  6. Guardrail checks against thresholds from experiment.yaml
  7. Segment analysis (Simpson's paradox check)
  8. Output: experiments/{slug}/working/analysis_results.json Checkpoint: SRM gate (Type C — BLOCK halts everything)
/experiment interpret

Purpose: Walk the Result Interpretation Tree and classify the outcome. Agent: agents/experiments/experiment-interpreter.md Flow:

  1. Read analysis results from experiments/{slug}/working/analysis_results.json
  2. Walk the Result Interpretation Tree:
    • Positive result + clean guardrails → SHIP
    • Positive result + degraded guardrails → INVESTIGATE (Mixed Results Framework)
    • Null result (powered) → ABORT (no evidence of benefit)
    • Null result (underpowered) → LEARN (extend or re-design)
    • Negative result → ABORT
    • SRM or data quality issue → INVALID
  3. Apply Spotify's EwL classification: Ship / Abort / Learn / Invalid
  4. Reference pre-registered decision rules from experiment.yaml
  5. Output: classification + rationale Checkpoint: Ship decision (Type C — always fires); INVALID → refuse to proceed
Show full SKILL.md (299 more words)Show less
/experiment report

Purpose: Generate markdown report from analysis results. Agent: agents/experiments/experiment-readout.md Flow:

  1. Read analysis results (structured JSON, not re-computing)
  2. Read experiment.yaml for context
  3. Fill report template (templates/experiment-report.md)
  4. Adapt to audience (executive/technical/cross-functional)
  5. Output: experiments/{slug}/reports/experiment_report_{{DATE}}.md
/experiment monitor

Purpose: SRM check + guardrail status + sample tracking during a running experiment. Agent: agents/experiments/experiment-monitor.md Flow:

  1. Read experiment.yaml for expected allocation and guardrail thresholds
  2. Run srm_check() with p < 0.0005 threshold (Microsoft production standard)
  3. Run guardrail tests (one-sided where appropriate)
  4. Track sample accumulation vs. required sample size
  5. Output: experiments/{slug}/working/monitoring_update.md
    • Traffic light status: GREEN (on track) / YELLOW (watch) / RED (halt) Checkpoint: RED guardrail (Type C — triggers halt)
/experiment status

Purpose: Show experiment lifecycle state. Flow:

  1. Read experiments/{slug}/experiment.yaml
  2. Display: current status, key metrics, timeline, any blockers
  3. No agent needed — direct YAML read and format
/experiment full

Purpose: End-to-end: design → power → analyze → interpret → report. Flow: Runs design, power, analyze, interpret, report in sequence. Checkpoints: All Type C checkpoints fire. Type B skipped with --just-do-it.

State Management

experiments/{slug}/
├── experiment.yaml          # Pre-registered config (tracked)
├── working/                 # Intermediates (gitignored)
│   ├── analysis_results.json
│   ├── monitoring_update.md
│   └── ...
└── reports/                 # Final reports (tracked)
    └── experiment_report_{{DATE}}.md

Helper Function Reference

All statistical work uses coded helpers from helpers/stats/experiment_stats/:

FunctionModuleUse For
welch_test()ab_testsContinuous metric A/B test
proportion_test()ab_testsBinary metric A/B test
ratio_metric_test()ab_testsRatio metric (delta method)
winsorize()ab_testsOutlier-robust pre-processing
power_proportion()powerSample size for proportions
power_mean()powerSample size for means
detectable_effect()powerMDE from fixed sample
duration_estimate()powerTimeline planning
srm_check()srmSample ratio mismatch
srm_diagnose()srmSegmented SRM root cause
cohens_d()effect_sizeStandardized effect size
relative_lift()effect_sizePercentage change
adjust_pvalues()correctionsMultiple comparison correction
cuped_adjust()variance_reductionCUPED variance reduction
confidence_sequence()sequentialAlways-valid CI (peeking ok)
bayesian_proportion()bayesianBayesian A/B (proportions)
bayesian_mean()bayesianBayesian A/B (means)

Cross-Product Handoffs

  • /experiment power → NOT_VIABLE → suggest /causal select (quasi-experimental)
  • /causal select → "Can you randomize? YES" → suggest /experiment design
  • /experiment analyze → SRM BLOCK → suggest investigating assignment logic

© ai-analyst-lab, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/experiment of ai-analyst-lab/ai-analyst.

Open the folder on GitHubat commit 52c0744

Compare with similar skills

Experiment next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Experiment compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Experiment this skillai-analyst-lab/ai-analyst304—~2kAutomated safety check: PassMIT
Statistical Analystalirezarezvani/claude-skills28k1 repos~2.5kAutomated safety check: PassMIT
Ab Test Analyzeririnabuht12-oss/marketing-skills3.9k—~1.4kAutomated safety check: PassNone
Define Hypothesisproduct-on-purpose/pm-skills715—~966Automated safety check: PassApache-2.0
A B Test DesignOwl-Listener/designer-skills2.9k1 repos~472Automated safety check: PassMIT
Data Scientistmagnus919/hermes-profiles281—~3.3kAutomated safety check: PassMIT

Similar skills

  • Statistical Analyst

    alirezarezvani/claude-skills

    Run hypothesis tests, analyze A/B experiment results, calculate sample sizes, and interpret statistical significance with effect sizes.

    28k GitHub starsUsed in 1 repo~2.5k tokens
    Data & AnalyticsAuto-check passed
  • Ab Test Analyzer

    irinabuht12-oss/marketing-skills

    Statistical significance calculator for A/B test results with sample size requirements, segment breakdowns, and hypothesis generation.

    3.9k GitHub stars~1.4k tokensUpdated 14 days ago
    Marketing & SEOAuto-check passed
  • Define Hypothesis

    product-on-purpose/pm-skills

    Defines a testable hypothesis with clear success metrics and a validation approach.

    715 GitHub stars~966 tokensUpdated 3 days ago
    Marketing & SEOAuto-check passed
  • A B Test Design

    Owl-Listener/designer-skills

    Design an A/B experiment — hypothesis, variants, primary metric, and sample size.

    2.9k GitHub starsUsed in 1 repo~472 tokens
    Marketing & SEOAuto-check passed
  • Data Scientist

    magnus919/hermes-profiles

    PhD-level expertise in data science, statistics, and machine learning.

    281 GitHub stars~3.3k tokensUpdated 3 mo ago
    Research & ScienceAuto-check passed
  • Data Scientist

    magnus919/agent-skills

    A skill your agent uses for PhD-level expertise in data science, statistics, and machine learning: rigorous statistical analysis, experimental design, causal inference, advanced modeling, research…

    113 GitHub stars~4.1k tokensUpdated 2 days ago
    Research & ScienceAuto-check passed

More from ai-analyst-lab/ai-analyst

All 43 skills in this repo
  • Always Compare

    ai-analyst-lab/ai-analyst

    Never present a metric or number in isolation; anchor every number to a comparison (prior period, benchmark, or another segment) or state that none is available.

    304 GitHub stars~1.4k tokensUpdated 7 days ago
    Auto-check passed
  • Archaeology

    ai-analyst-lab/ai-analyst

    Retrieve proven SQL patterns, table cheatsheets, and join patterns from .knowledge/query-archaeology/ so past work gets reused.

    304 GitHub stars~1.3k tokensUpdated 7 days ago
    Auto-check passed
  • Archive Analysis

    ai-analyst-lab/ai-analyst

    Save completed analyses to the knowledge system's analysis archive for future reference.

    304 GitHub stars~2.7k tokensUpdated 7 days ago
    Auto-check passed
  • Auth Preflight

    ai-analyst-lab/ai-analyst

    Verify Google Workspace MCP authentication at the start of any session that needs Google APIs (Docs, Slides, Drive).

    304 GitHub stars~3.1k tokensUpdated 7 days ago
    Auto-check passed
  • Causal

    ai-analyst-lab/ai-analyst

    Causal inference toolkit for when experiments are not possible: estimate treatment effects from observational data with assumption checks and mandatory caveats.

    304 GitHub stars~1.8k tokensUpdated 7 days ago
    Auto-check passed
  • Chart To Drive

    ai-analyst-lab/ai-analyst

    Standardized workflow for uploading local chart PNGs to Google Drive and making them available for insertion into Google Docs and Slides.

    304 GitHub stars~1.4k tokensUpdated 7 days ago
    Auto-check passed

Questions about Experiment

What does Experiment do?

The analysis and lifecycle owner for experiments. An agent skill from ai-analyst-lab/ai-analyst. Experiment is an agent skill from ai-analyst-lab/ai-analyst. The analysis and lifecycle owner for experiments.

When should I use Experiment?

Experiment fits situations like: treatment vs control; statistical significance; is this result significant?.

How do I install Experiment in Claude Code?

Run `npx skills add ai-analyst-lab/ai-analyst --skill experiment -a claude-code`. Or copy the skill folder (.claude/skills/experiment in ai-analyst-lab/ai-analyst) into .claude/skills/experiment in your project. Claude Code loads it when a task matches its description.

How do I install Experiment in Codex?

Run `npx skills add ai-analyst-lab/ai-analyst --skill experiment -a codex`. Or copy the skill folder (.claude/skills/experiment in ai-analyst-lab/ai-analyst) into .agents/skills/experiment in your project. Codex loads it when a task matches its description.

Can I use Experiment in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ai-analyst-lab/ai-analyst --skill experiment -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/experiment, .gemini/skills/experiment, .github/skills/experiment and .opencode/skills/experiment in your project.

What does Experiment need to run?

SKILL.md names no scripts, command-line tools or credentials: Experiment is instructions for the agent only. Our summary lists: Python 3.

Does Experiment access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Experiment safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Experiment use?

Experiment is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Experiment use?

About 2k tokens (SKILL.md is roughly 7.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Experiment?

Skills that share tags, products or a category with Experiment: Statistical Analyst (alirezarezvani/claude-skills, 28k stars), Ab Test Analyzer (irinabuht12-oss/marketing-skills, 3.9k stars), Define Hypothesis (product-on-purpose/pm-skills, 715 stars) and A B Test Design (Owl-Listener/designer-skills, 2.9k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Experiment?

ai-analyst-lab (a GitHub organization) maintains it in ai-analyst-lab/ai-analyst, which has 304 GitHub stars. The repository holds 43 skills in this directory. The repository was last updated on September 30, 2026.

Source: ai-analyst-lab/ai-analyst on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.