A/B test design — produce an experiment spec with hypothesis, primary metric, MDE, sample size, run time, and decision rule.

MITAuto-check: notesMarketing & SEO

Install Lumen Abtest

skills CLI
$ npx skills add jeremylongshore/tons-of-skills-marketplace --skill lumen-abtest -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install jeremylongshore/tons-of-skills-marketplace lumen-abtest --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/jeremylongshore/tons-of-skills-marketplace.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/ai-agency/tonone/skills/lumen-abtest .claude/skills/lumen-abtest && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
lumen-abtest
GitHub stars
2.8k
Token cost
~2.3k tokens
SKILL.md length
815 words
Files
2
Skills in repo
3,342
Repo updated
First seen
Licence
MIT

At a glance

A/B test design — produce an experiment spec with hypothesis, primary metric, MDE, sample size, run time, and decision rule.

  • Works in 7 steps: Make the Call — Test or Don't Test → Write the Hypothesis → Define the Metrics → …
  • Asked to design an A/B test
  • SKILL.md covers Step 0: Make the Call — Test…, Step 1: Write the Hypothesis, Step 2: Define the Metrics and Step 3: Calculate Sample Size, plus 5 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Lumen Abtest is an agent skill from jeremylongshore/tons-of-skills-marketplace. A/B test design — produce an experiment spec with hypothesis, primary metric, MDE, sample size, run time, and decision rule. Also determines when NOT to A/B test and what to do instead. Use when asked to "design an A/B test", "should we test this", "experiment design", "how do we know if this works", "what's the sample size", or "set up an experiment".

Its SKILL.md is about 2.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files (for example `.claude-plugin/plugin.json`).

It sits in Marketing & SEO, covering Experimental design and A/B testing. The repository describes itself as: Model-agnostic agent-skills platform with a harness-free canonical layer, verified adapters, and the ccpi package manager. Explore at tonsofskills.com. The licence is MIT.

When your agent uses it

  • Asked to design an A/B test
  • Should we test this
  • Experiment design
  • How do we know if this works

Example prompts

  • “design an A/B test”
  • “should we test this”
  • “experiment design”
  • “/lumen-abtest”

Requirements

  • Pre-approved tools (allowed-tools): Read, Write, Edit, Bash, Glob, Grep, WebFetch, WebSearch, Task, TodoWrite, AskUserQuestion

Workflow steps

7 steps, taken from the step headings in SKILL.md.

  1. Make the Call — Test or Don't Test
  2. Write the Hypothesis
  3. Define the Metrics
  4. Calculate Sample Size
  5. Calculate Run Time
  6. Write the Decision Rule
  7. Pre-Launch Checklist

What it can do on your machine

Read from SKILL.md and the folder at commit cfae287. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Write
    • Edit
    • Bash
    • Glob
    • Grep
    • WebFetch
    • WebSearch
    • Task
    • TodoWrite

    …and 1 more on the same allowed-tools line.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Lumen Abtest loads about 2.3k tokens when it runs. Until then it costs about 92 tokens; SKILL.md has 815 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~92
When it runs · the whole SKILL.md, loaded when a task matches
~2.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Read, Write, Edit, Bash, Glob, Grep, WebFetch, WebSearch, Task, TodoWrite, AskUserQuestion

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from jeremylongshore/tons-of-skills-marketplace at commit cfae287, republished under its MIT licence (© jeremylongshore). 815 words, ~2,321 tokens.

Download SKILL.mdSave it as .claude/skills/lumen-abtest/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
lumen-abtest
description
A/B test design — produce an experiment spec with hypothesis, primary metric, MDE, sample size, run time, and decision rule. Also determines when NOT to A/B test and what to do instead. Use when asked to "design an A/B test", "should we test this", "experiment design", "how do we know if this works", "what's the sample size", or "set up an experiment".
allowed-tools
Read, Write, Edit, Bash, Glob, Grep, WebFetch, WebSearch, Task, TodoWrite, AskUserQuestion
version
0.6.4
author
tonone-ai <hello@tonone.ai>
license
MIT

Lumen A/B Test

You are Lumen — the product analyst on the Product Team. Given a change to test, produce a complete experiment spec with decision rule. Or tell the team this is not the right tool — and say what to do instead.

Follow the output format defined in docs/output-kit.md — 40-line CLI max, box-drawing skeleton, unified severity indicators, compressed prose.

Step 0: Make the Call — Test or Don't Test

Before writing any spec, answer three questions. If any answer is NO, do not design an A/B test. Prescribe the right alternative instead.

Question 1: Do you have enough traffic?

Minimum viable traffic for a standard A/B test:

  • 500+ conversions per week on the metric you're testing
  • Enough to reach required sample size in ≤6 weeks
  • If below this: don't test. Use qualitative methods.

Question 2: Is this a tactical question or a strategic one?

A/B tests answer tactical questions: "Does button copy A or B convert better?" They do not answer strategic questions: "Should we build this feature at all?" or "Are we solving the right problem?"

  • Tactical (copy, layout, flow step, UI element) → A/B test
  • Strategic (positioning, core value prop, major feature direction) → user research, not an experiment

Question 3: Is the change big enough to detect?

If testing a change you believe will move primary metric by <5% relative, and baseline rate is below 20%, you will need tens of thousands of users per variant. Be honest about whether this is worth running.

When NOT to A/B Test — and What to Do Instead
SituationDon't TestDo This Instead
<500 conversions/weekUnderpowered — results are noiseSession recordings, user interviews (Echo)
Strategic questionTest won't answer itUser research, Jobs-to-Be-Done with Echo
One-time irreversible changeNo rollback pathStaged rollout with monitoring, not a test
Change is qualitative (tone, brand)No clean metricExpert review + user feedback
Pre-PMF, <1k usersToo few to segmentTalk to users. Don't build dashboards.

Make the call explicitly. If this shouldn't be an A/B test, say so, say why, and prescribe the alternative. Don't design a bad experiment because someone asked for one.


Step 1: Write the Hypothesis

If we [specific change],
then [primary metric] will [increase / decrease] by [X%],
because [mechanism — why this change produces this effect].

We will know this is true if [primary metric] moves by [MDE] or more
with 95% statistical confidence within [N] days.

The "because" is not optional. It forces a causal theory, not a hope. A hypothesis without a mechanism is a guess dressed up as a test.


Step 2: Define the Metrics

Primary metric — one only. This single metric decides the test. If it moves by MDE or more, the variant wins. Do not change this metric after the test starts.

Secondary metrics — 2–4 metrics that help explain why the primary moved. Directional only — they don't decide the outcome.

Guardrail metrics — 1–2 metrics that must not degrade. A test that wins on primary but tanks a guardrail is a failed test. Ship nothing until guardrails pass.

TypeMetricDirectionThreshold
Primary[metric]↑≥[MDE]% lift
Secondary[metric]↑/↓directional
Secondary[metric]↑/↓directional
Guardrail[metric]→must not drop >5%
Guardrail[metric]→must not drop >5%

Show full SKILL.md (330 more words)Show less

Step 3: Calculate Sample Size

n = (Zα/2 + Zβ)² × 2 × p × (1 - p) / MDE²

Where:
  Zα/2 = 1.96  (95% confidence, two-tailed)
  Zβ   = 0.84  (80% power) — standard default
         1.28  (90% power) — use for high-stakes decisions
  p    = baseline conversion rate (decimal)
  MDE  = minimum detectable effect (decimal, e.g. 0.02 for 2pp)

Lookup table (80% power, 95% confidence, two-tailed):

Baseline RateMDE (relative)MDE (absolute)Users per variant
5%20% relative1pp~3,700
10%10% relative1pp~14,800
20%10% relative2pp~14,800
20%5% relative1pp~59,200
50%5% relative2.5pp~62,900

State: "We need [N] users per variant — [2N] total across control and variant."

If required sample size implies run time >6 weeks at current traffic volume, this test is not viable as designed. Options: increase the MDE (test a bolder change), segment to a higher-traffic subpopulation, or don't test.


Step 4: Calculate Run Time

Run time (days) = (users per variant × number of variants) / daily eligible users

Minimum: 14 days — captures weekly seasonality patterns
Maximum: 42 days (6 weeks) — beyond this, novelty effects and seasonal drift contaminate results

If run time < 14 days even with required sample size: run full 14 days anyway. Novelty effects in first few days will inflate variant's early numbers.

If run time > 42 days: do not run this test. MDE is too small or traffic too thin. See Step 0.


Step 5: Write the Decision Rule

State this before the test launches. Do not revise after seeing interim results.

DECISION RULE — [test name]

WIN: primary metric lifts ≥ [MDE] with p < 0.05 AND all guardrails pass
  → Ship variant to 100%. Rollout plan: [staged / immediate / feature flag].

GUARDRAIL FAIL: primary wins but a guardrail metric drops >5%
  → Do NOT ship. Investigate guardrail failure before any decision.
     Root cause question: [what does the guardrail failure tell us?]

NULL: primary metric does not lift by MDE
  → Keep control. Document the learning:
     [what does this null result tell us about the hypothesis/mechanism?]

EARLY STOP: test stopped before planned end date
  → Default to control. Early stopping inflates false positive rate.
     No winner can be declared from a stopped test.

Peeking at results and stopping early is the most common way teams deceive themselves. Decision rule must be written down and shared before Day 1.


Step 6: Pre-Launch Checklist

Complete before starting the test clock:

  • Experiment framework configured (feature flag, split testing tool)
  • Randomization unit defined — user ID (preferred), session, or device
  • Sticky assignment confirmed — same user always sees same variant
  • All metrics instrumented and verified firing correctly in both variants
  • Control and variant verified functionally (QA pass)
  • Split defined: [50/50] or [90/10 for risky changes]
  • Start date and hard end date set
  • Decision rule documented and shared with stakeholders
  • Interim check-in date set — for guardrail monitoring only, not winner declaration

Output Format

┌─────────────────────────────────────────────────────┐
│  EXPERIMENT SPEC — [Test Name]                      │
└─────────────────────────────────────────────────────┘

HYPOTHESIS
  If [change], then [metric] will [direction] by [X%]
  because [mechanism].

METRICS
  Primary:   [metric] — need ≥[MDE]% lift to declare win
  Secondary: [metric], [metric]
  Guardrail: [metric] must not drop >5%

SIZING
  Baseline rate:        [X]%
  MDE:                  [X]% relative ([Xpp] absolute)
  Users per variant:    [N]
  Daily eligible users: [N]
  Run time:             [N] days
  Start date:           [date]
  Decision date:        [date]

DECISION RULE
  WIN  → ship if primary ≥ MDE and guardrails pass
  FAIL → revert if guardrail fails regardless of primary
  NULL → keep control; learning: [what this tells us]
  STOP → default to control; no winner declared

CHECKLIST
  [ ] Feature flag configured
  [ ] Randomization unit: [user ID / session]
  [ ] All metrics verified firing
  [ ] Decision rule shared with stakeholders

Deliver this spec. The team ships the experiment, not more deliberation.

Delivery

If output exceeds the 40-line CLI budget, invoke /atlas-report with the full findings. The HTML report is the output. CLI is the receipt — box header, one-line verdict, top 3 findings, and the report path. Never dump analysis to CLI.

© jeremylongshore, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in plugins/ai-agency/tonone/skills/lumen-abtest of jeremylongshore/tons-of-skills-marketplace.

  • SKILL.md
  • .claude-plugin/plugin.json

Open the folder on GitHubat commit cfae287

Compare with similar skills

Lumen Abtest next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Lumen Abtest compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Lumen Abtest this skilljeremylongshore/tons-of-skills-marketplace2.8k—~2.3kAutomated safety check: NotesMIT
Ab Test Analyzeririnabuht12-oss/marketing-skills4.1k—~1.4kAutomated safety check: PassNone
Define Hypothesisproduct-on-purpose/pm-skills716—~966Automated safety check: PassApache-2.0
A B Test DesignOwl-Listener/designer-skills2.9k1 repos~472Automated safety check: PassMIT
Send Experiment Designeraaron-he-zhu/aaron-marketing-skills2.9k2 repos~4.1kAutomated safety check: PassApache-2.0
Content Experimentation Best Practicessanity-io/agent-toolkit1881 repos~469Automated safety check: PassMIT

Similar skills

  • Ab Test Analyzer

    irinabuht12-oss/marketing-skills

    Statistical significance calculator for A/B test results with sample size requirements, segment breakdowns, and hypothesis generation.

    4.1k GitHub stars~1.4k tokensUpdated 17 days ago
    Marketing & SEOAuto-check passed
  • Define Hypothesis

    product-on-purpose/pm-skills

    Defines a testable hypothesis with clear success metrics and a validation approach.

    716 GitHub stars~966 tokensUpdated 3 days ago
    Marketing & SEOAuto-check passed
  • A B Test Design

    Owl-Listener/designer-skills

    Design an A/B experiment — hypothesis, variants, primary metric, and sample size.

    2.9k GitHub starsUsed in 1 repo~472 tokens
    Marketing & SEOAuto-check passed
  • Send Experiment Designer

    aaron-he-zhu/aaron-marketing-skills

    A skill your agent uses when the user asks to "design an email A/B test", "set up a multivariate subject/CTA test", "run a send-time test", "build a hold-out group", or "is this email result…

    2.9k GitHub starsUsed in 2 repos~4.1k tokens
    Marketing & SEOAuto-check passed
  • Official

    Content experimentation and A/B testing guidance covering experiment design, hypotheses, metrics, sample size, statistical foundations, CMS-managed variants, and common analysis pitfalls.

    188 GitHub starsUsed in 1 repo~469 tokens
    Marketing & SEOAuto-check passed
  • Ads Test

    AgriciDaniel/claude-ads

    Design and evaluate paid-ad experiments with hypotheses, randomization units, sample-size and duration assumptions, guardrails, platform experiment tools, analysis, and decision rules.

    9.9k GitHub stars~312 tokensUpdated 3 days ago
    Marketing & SEOAuto-check passed

More from jeremylongshore/tons-of-skills-marketplace

All 3,342 skills in this repo
  • Performing Security Code Review

    jeremylongshore/tons-of-skills-marketplace

    Execute this skill enables AI assistant to conduct a security-focused code review using the security-agent plugin.

    2.8k GitHub starsUsed in 2 repos~1.3k tokens
    Auto-check: notes
  • Adapting Transfer Learning Models

    jeremylongshore/tons-of-skills-marketplace

    Build this skill automates the adaptation of pre-trained machine learning models using transfer learning techniques.

    2.8k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Agent Context Loader

    jeremylongshore/tons-of-skills-marketplace

    Execute proactive auto-loading: automatically detects and loads agents.md files.

    2.8k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Aggregating Performance Metrics

    jeremylongshore/tons-of-skills-marketplace

    Aggregate and centralize performance metrics from applications, systems, databases, caches, and services.

    2.8k GitHub stars~1.2k tokensUpdated today
    Auto-check passed
  • Analyzing Capacity Planning

    jeremylongshore/tons-of-skills-marketplace

    Execute this skill enables AI assistant to analyze capacity requirements and plan for future growth.

    2.8k GitHub stars~947 tokensUpdated today
    Auto-check passed
  • Analyzing Database Indexes

    jeremylongshore/tons-of-skills-marketplace

    Process use when you need to work with database indexing. An agent skill from jeremylongshore/tons-of-skills-marketplace.

    2.8k GitHub stars~2k tokensUpdated today
    Auto-check passed

Questions about Lumen Abtest

What does Lumen Abtest do?

A/B test design — produce an experiment spec with hypothesis, primary metric, MDE, sample size, run time, and decision rule. Lumen Abtest is an agent skill from jeremylongshore/tons-of-skills-marketplace. A/B test design — produce an experiment spec with hypothesis, primary metric, MDE, sample size, run time, and decision rule.

When should I use Lumen Abtest?

Lumen Abtest fits situations like: asked to design an A/B test; should we test this; experiment design; how do we know if this works.

How do I install Lumen Abtest in Claude Code?

Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill lumen-abtest -a claude-code`. Or copy the skill folder (plugins/ai-agency/tonone/skills/lumen-abtest in jeremylongshore/tons-of-skills-marketplace) into .claude/skills/lumen-abtest in your project. Claude Code loads it when a task matches its description.

How do I install Lumen Abtest in Codex?

Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill lumen-abtest -a codex`. Or copy the skill folder (plugins/ai-agency/tonone/skills/lumen-abtest in jeremylongshore/tons-of-skills-marketplace) into .agents/skills/lumen-abtest in your project. Codex loads it when a task matches its description.

Can I use Lumen Abtest in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill lumen-abtest -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/lumen-abtest, .gemini/skills/lumen-abtest, .github/skills/lumen-abtest and .opencode/skills/lumen-abtest in your project.

What does Lumen Abtest need to run?

SKILL.md names no scripts, command-line tools or credentials: Lumen Abtest is instructions for the agent only. Its frontmatter pre-approves these tools: Read, Write, Edit, Bash, Glob, Grep, WebFetch, WebSearch, Task, TodoWrite, AskUserQuestion.

Does Lumen Abtest access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Lumen Abtest safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Lumen Abtest use?

Lumen Abtest is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Lumen Abtest use?

About 2.3k tokens (SKILL.md is roughly 9.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Lumen Abtest?

Skills that share tags, products or a category with Lumen Abtest: Ab Test Analyzer (irinabuht12-oss/marketing-skills, 4.1k stars), Define Hypothesis (product-on-purpose/pm-skills, 716 stars), A B Test Design (Owl-Listener/designer-skills, 2.9k stars) and Send Experiment Designer (aaron-he-zhu/aaron-marketing-skills, 2.9k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Lumen Abtest?

jeremylongshore (a GitHub user) maintains it in jeremylongshore/tons-of-skills-marketplace, which has 2,827 GitHub stars. The repository holds 3,342 skills in this directory. The repository was last updated on October 10, 2026.

Source: jeremylongshore/tons-of-skills-marketplace on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.