Agent skill

Ab Test Setup

by borghei in borghei/Claude-Skills

Design and run statistically rigorous A/B tests and experiments.

MITAuto-check passedMarketing & SEO

Install Ab Test Setup

skills CLI
$ npx skills add borghei/Claude-Skills --skill ab-test-setup -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install borghei/Claude-Skills ab-test-setup --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/borghei/Claude-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/product-team/ab-test-setup .claude/skills/ab-test-setup && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
ab-test-setup
GitHub stars
874
Token cost
~5.3k tokens
SKILL.md length
2,168 words
Files
8 (incl. scripts)
Skills in repo
364
Repo updated
First seen
Licence
MIT

At a glance

Design and run statistically rigorous A/B tests and experiments.

  • Works in 7 steps: Hypothesis Formulation → Test Design → Sample Size Calculation → …
  • Planning experiments
  • SKILL.md covers Overview, Clarify First, The Experiment Lifecycle and Step 1: Hypothesis Formulation, plus 6 more sections
  • Runs Python scripts from its folder; calls python

What it does

Ab Test Setup is an agent skill from borghei/Claude-Skills. Design and run statistically rigorous A/B tests and experiments. Use when planning experiments, calculating sample sizes, designing test variants, selecting metrics, analyzing results, or when someone says "let's test that."

Its SKILL.md is about 5.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 9 other files, including scripts (for example `evals/README.md`, `evals/grader.py` and `evals/runner.py`).

It sits in Marketing & SEO, covering A/B testing and Experimental design. The repository describes itself as: 385 AI skills, 77 expert agents, and 900 stdlib Python tools for every team: engineering, PM, marketing, C-level, compliance, business ops, research, and a LinkedIn toolkit… The licence is MIT.

When your agent uses it

  • Planning experiments
  • Calculating sample sizes
  • Designing test variants
  • Selecting metrics

Example prompts

  • “s test that.”
  • “/ab-test-setup”

Requirements

  • Python 3

Workflow steps

7 steps, taken from the step headings in SKILL.md.

  1. Hypothesis Formulation
  2. Test Design
  3. Sample Size Calculation
  4. Implementation
  5. Running the Test
  6. Analysis
  7. Documentation

What it can do on your machine

Read from SKILL.md and the folder at commit c9a1487. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 3 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Ab Test Setup loads about 5.3k tokens when it runs. Until then it costs about 60 tokens; SKILL.md has 2,168 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~60
When it runs · the whole SKILL.md, loaded when a task matches
~5.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from borghei/Claude-Skills at commit c9a1487, republished under its MIT licence (© borghei). 2,168 words, ~5,270 tokens.

Download SKILL.mdSave it as .claude/skills/ab-test-setup/SKILL.md (or your agent's skills folder). This skill also uses 7 other files; get the full folder from GitHub.
name
ab-test-setup
description
Design and run statistically rigorous A/B tests and experiments. Use when planning experiments, calculating sample sizes, designing test variants, selecting metrics, analyzing results, or when someone says "let's test that."
license
MIT + Commons Clause
metadata.version
1.0.0
metadata.author
borghei
metadata.category
product-team
metadata.domain
experimentation
metadata.updated
2026-03-09
metadata.tags
ab-testing, experimentation, hypothesis, statistical-significance
metadata.frameworks
hypothesis-testing, statistical-significance, feature-flags

A/B Test Setup - Experimentation Design & Analysis

Category: Product Team Tags: A/B testing, experiments, statistical significance, sample size, feature flags, hypothesis testing

Overview

A/B Test Setup provides the complete framework for designing experiments that produce statistically valid, actionable results. Most A/B tests fail not because the variant was wrong, but because the test was poorly designed: wrong sample size, wrong metric, or someone peeked at results and stopped early. This skill prevents those mistakes.


Clarify First

Before designing the experiment, confirm these inputs. If any is unknown or vague, ASK — do not assume:

  • Primary metric + minimum detectable effect — the one success metric and smallest lift worth detecting (drives sample size, duration, and metric selection)
  • Baseline conversion rate — current rate for the primary metric (sets required sample size per variant)
  • Available traffic to the test surface — daily eligible visitors (decides whether the test is feasible or needs a bolder change / qualitative method)
  • The change and its rationale — what varies and the data behind it (drives the hypothesis and variant design)

Stop rule: ask only the 2-3 that most change the output. If the user says "just draft it," proceed and list your assumptions at the top of the artifact.

The Experiment Lifecycle

1. HYPOTHESIZE  →  2. DESIGN  →  3. CALCULATE  →  4. IMPLEMENT
       ↑                                                    │
       │                                                    ▼
7. ITERATE  ←  6. DOCUMENT  ←  5. ANALYZE  ←  [Run to completion]

Step 1: Hypothesis Formulation

The Hypothesis Template
Because [observation or data point],
we believe [specific change]
will cause [measurable outcome]
for [defined audience segment].

We'll know this is true when [primary metric] changes by [minimum detectable effect].
We'll watch [guardrail metrics] to ensure no negative impact.
Good vs Bad Hypotheses
QualityHypothesisProblem
Bad"Changing the button color might increase clicks"No data basis, no target, no measurement plan
Mediocre"A green button will get more clicks than blue"No "why", no target size, no guardrails
Good"Because heatmaps show 40% of users don't notice our CTA, making the button 2x larger with contrasting color will increase CTA clicks by 15%+ for new visitors. Guardrail: page load time stays under 2s."Data-backed, specific change, measurable outcome, defined audience, guardrail
Hypothesis Sources (Where to Find Test Ideas)
SourceWhat to Look ForExample
Analytics dataDrop-off points, low-performing pages"80% of users drop off at step 3 of onboarding"
User researchConfusion, frustration, unmet needs"Users don't understand what the product does from the homepage"
Heatmaps/session recordingsIgnored elements, rage clicks"Nobody scrolls past the fold on pricing page"
Support ticketsRecurring complaints, feature confusion"Users constantly ask how to invite team members"
Competitor analysisDifferent approaches to same problem"Competitor uses a wizard; we use a form"
Sales objectionsCommon reasons prospects don't convert"Prospects want to see pricing before signing up"

Step 2: Test Design

Test Types
TypeVariantsTraffic NeedBest For
A/B2 (control + 1 variant)ModerateSingle change validation
A/B/n3+ variantsHighComparing multiple approaches
Multivariate (MVT)Combinations of changesVery highOptimizing multiple elements
Split URLDifferent pagesModerateMajor redesigns
BanditDynamic allocationLow-moderateRevenue optimization

Default recommendation: Standard A/B test. Only use A/B/n or MVT when you have enough traffic and a specific need.

What to Test (By Impact)
CategoryHigh ImpactMedium ImpactLow Impact
CopyHeadline/value prop, CTA textBody copy, social proofMicrocopy, labels
DesignPage layout, above-fold contentVisual hierarchy, imageryColor, font size
UXNumber of steps, form fieldsButton placement, navigationAnimations, transitions
PricingPrice point, plan namesFeature packaging, anchoringBilling frequency display
Social ProofTestimonials vs none, logosTestimonial format, placementTestimonial count
Metric Selection

Every test needs three types of metrics:

Primary Metric (1 only)

  • The single metric that determines success
  • Directly tied to the hypothesis
  • Must be measurable within the test duration
  • Examples: signup rate, click-through rate, purchase rate

Secondary Metrics (2-3)

  • Explain why the primary metric moved
  • Provide context for decision-making
  • Examples: time on page, scroll depth, feature adoption rate

Guardrail Metrics (1-3)

  • Things that must NOT get worse
  • Stop the test if significantly negative
  • Examples: error rate, support ticket volume, page load time, refund rate

Step 3: Sample Size Calculation

Quick Reference Table

Minimum visitors PER VARIANT needed (95% confidence, 80% power):

Baseline Rate5% Lift10% Lift15% Lift20% Lift50% Lift
1%620,000156,00070,00039,0006,400
2%305,00077,00034,00019,5003,200
3%200,00051,00023,00012,8002,100
5%116,00029,50013,2007,5001,250
10%54,00013,8006,2003,500600
20%24,0006,2002,8001,600280
50%6,1001,60072041075
Duration Calculation
Duration (days) = (Sample size per variant * Number of variants) / Daily traffic to test page

Minimum duration: 7 days (to capture day-of-week effects) Maximum recommended: 6 weeks (beyond this, external factors contaminate results)

What If You Don't Have Enough Traffic?
SituationSolution
Need 100K visitors, get 5K/weekIncrease minimum detectable effect (test bolder changes)
Very low traffic (<1K/week)Use qualitative testing (user testing, surveys) instead
Medium traffic (5-20K/week)Run for 4-6 weeks, test big changes only
High traffic (50K+/week)You can test subtle changes, run multiple tests

Step 4: Implementation

Client-Side Implementation

JavaScript modifies the page after initial render.

Pros: Quick to implement, no deploy needed Cons: Can cause flicker (flash of original content), blocked by ad blockers Tools: PostHog, Optimizely, VWO, Google Optimize

Anti-flicker pattern:

javascript
// Add to <head> before any rendering
<style>.ab-test-hide { opacity: 0 !important; }</style>
<script>document.documentElement.classList.add('ab-test-hide');</script>

// In your test script (runs after variant assignment):
document.documentElement.classList.remove('ab-test-hide');
Server-Side Implementation

Variant determined before page renders. No flicker, no client-side dependency.

Pros: No flicker, not blocked by ad blockers, works for logged-in features Cons: Requires engineering work, deploy needed Tools: PostHog, LaunchDarkly, Split, Unleash, custom feature flags

Basic feature flag pattern:

python
# Server-side variant assignment
def get_variant(user_id: str, experiment: str) -> str:
    # Deterministic hash ensures same user always sees same variant
    hash_input = f"{user_id}:{experiment}"
    hash_value = hashlib.md5(hash_input.encode()).hexdigest()
    bucket = int(hash_value[:8], 16) % 100

    if bucket < 50:
        return "control"
    else:
        return "variant"
Traffic Allocation
StrategySplitWhen to Use
Standard50/50Default. Maximum statistical power.
Conservative90/10 or 80/20Risky changes, revenue-impacting tests
RampedStart 95/5, increase to 50/50New infrastructure, technical risk

Critical rules:

  • Users must see the same variant on every visit (sticky assignment by user ID or cookie)
  • Allocation must be balanced across time of day and day of week
  • Never change allocation mid-test

Step 5: Running the Test

Pre-Launch Checklist
  • Hypothesis documented with primary metric and minimum detectable effect
  • Sample size calculated, expected duration estimated
  • Both variants implemented and QA'd on all device types
  • Tracking verified (events fire correctly for both variants)
  • No other tests running on the same page/feature
  • Stakeholders informed of test duration and "no peeking" rule
  • External factor calendar checked (no major launches, holidays, press)
During the Test

DO:

  • Monitor for technical errors (variant not rendering, tracking broken)
  • Check that traffic split is balanced daily
  • Document any external events that might affect results

DO NOT:

  • Look at results before reaching sample size ("peeking problem")
  • Make changes to either variant
  • Add traffic from new sources mid-test
  • Stop the test early because one variant "looks like it's winning"
The Peeking Problem (Critical)

Looking at results before reaching the planned sample size and stopping because one variant looks better leads to a 25-40% false positive rate (vs the intended 5%).

Why: Statistical significance fluctuates wildly with small samples. A variant can show p < 0.05 at 20% of planned sample size and p > 0.30 at full sample.

Solutions:

  1. Pre-commit to sample size and do not check results until reached
  2. If you must monitor: use sequential testing methods (group sequential design, always-valid p-values)
  3. Set calendar reminder for expected completion date -- that is when you look

Step 6: Analysis

Analysis Checklist
  1. Did we reach planned sample size? If not, results are preliminary only.
  2. Is it statistically significant? p < 0.05 = 95% confidence the difference is real.
  3. What's the confidence interval? Tells you the range of likely true effect.
  4. Is the effect size meaningful? A 0.1% lift that's "significant" may not be worth implementing.
  5. Are secondary metrics consistent? Do they support the primary result?
  6. Any guardrail violations? Did anything get worse?
  7. Segment analysis: Different results for mobile vs desktop? New vs returning?
Interpreting Results
ResultPrimary MetricConfidenceAction
Clear winnerVariant +15%, p < 0.01HighImplement variant
Modest winnerVariant +5%, p < 0.05MediumImplement if easy, else run longer
Flat< 2% difference, p > 0.20High (no effect)Keep control, test something bolder
LoserVariant -10%, p < 0.05HighKeep control, investigate why
Inconclusive5% difference, p = 0.08LowNeed more traffic or bolder test
Mixed signalsPrimary up, guardrail downInvestigateDig into segments, do not ship blindly
Show full SKILL.md (869 more words)Show less
Common Analysis Mistakes
MistakeConsequencePrevention
Stopping at first significance25-40% false positive rateCommit to sample size
Cherry-picking segmentsFinding "winners" that don't replicatePre-register segments of interest
Ignoring confidence intervalsOverestimating effect sizeAlways report CI alongside p-value
Multiple comparisonsInflated Type I errorBonferroni correction for A/B/n
Survivorship biasOnly analyzing users who completed flowInclude all users from assignment point
Simpson's paradoxAggregate hides segment reversalAlways check key segments

Step 7: Documentation

Every test must be documented, regardless of outcome.

Test Documentation Template
EXPERIMENT: [Name]
DATE: [Start] to [End]
OWNER: [Name]

HYPOTHESIS:
Because [observation], we believed [change] would cause [outcome] for [audience].

VARIANTS:
- Control: [description]
- Variant: [description + screenshot]

METRICS:
- Primary: [metric] (baseline: [X]%, MDE: [Y]%)
- Secondary: [metrics]
- Guardrails: [metrics]

RESULTS:
- Sample size: [actual] / [planned]
- Duration: [X] days
- Primary metric: Control [X]% vs Variant [Y]% (p = [Z], CI: [range])
- Secondary metrics: [results]
- Guardrails: [all clear / violation noted]

DECISION: [Ship variant / Keep control / Iterate]

LEARNINGS:
- [What we learned about our users]
- [What we'd do differently next time]

Experiment Prioritization Framework

ICE Scoring
FactorScore (1-10)Question
ImpactHow much will this move the metric?Big change to primary KPI = 10
ConfidenceHow sure are we it will work?Strong data supporting hypothesis = 10
EaseHow easy is it to implement and measure?Can ship in a day = 10

ICE Score = (Impact + Confidence + Ease) / 3

Rank all test ideas by ICE score. Run highest first.

Test Backlog Template
#HypothesisPrimary MetricICEEst. DurationStatus
1Larger CTA increases signupsSignup rate8.32 weeksReady
2Social proof on pricing increases conversionPlan selection rate7.03 weeksNeeds design
3Shorter onboarding increases activationFeature activation6.74 weeksIn backlog

Proactive Triggers

  • Someone debates between two design options: propose an A/B test instead of opinionating
  • Conversion rate mentioned as underperforming: offer to design a test, not guess at solutions
  • Pricing page changes discussed: always test pricing changes with guardrail metrics
  • Post-launch of any feature: propose follow-up experiment to optimize
  • "Let's just try it and see": redirect to structured hypothesis before implementation

SkillUse When
analytics-trackingSetting up event tracking that feeds experiment metrics
campaign-analyticsFolding experiment results into broader attribution
launch-strategyTesting within a product launch sequence
prompt-engineer-toolkitA/B testing AI prompts in production

Tool Reference

sample_size_calculator.py

Calculates required sample size per variant using the normal approximation to the two-proportion z-test. Includes Bonferroni correction for multi-variant tests and duration estimation.

FlagTypeDefaultDescription
--baseline, -bfloat(required)Baseline conversion rate (e.g. 0.05 for 5%)
--mde, -mfloat(required)Minimum detectable effect as relative lift (e.g. 0.10 for 10%)
--alpha, -afloat0.05Significance level
--power, -pfloat0.80Statistical power
--variants, -vint2Number of variants including control
--daily-traffic, -dint0Daily eligible traffic for duration estimation
--one-tailedflagFalseUse one-tailed test instead of two-tailed
--jsonflagFalseOutput as JSON
bash
python scripts/sample_size_calculator.py --baseline 0.05 --mde 0.10
python scripts/sample_size_calculator.py --baseline 0.12 --mde 0.15 --power 0.9 --daily-traffic 5000
python scripts/sample_size_calculator.py --baseline 0.05 --mde 0.10 --variants 3 --json
experiment_analyzer.py

Analyzes A/B test results using the two-proportion z-test with confidence intervals and segment breakdown.

FlagTypeDefaultDescription
inputpositional(required)CSV file with results or "sample" to create sample
--alpha, -afloat0.05Significance level
--jsonflagFalseOutput as JSON

CSV format: variant,visitors,conversions,segment

bash
python scripts/experiment_analyzer.py sample
python scripts/experiment_analyzer.py results.csv
python scripts/experiment_analyzer.py results.csv --alpha 0.01 --json
experiment_planner.py

Generates a structured experiment plan from a hypothesis text, including metric selection, sample size, timeline, risks, and documentation template.

FlagTypeDefaultDescription
--hypothesis, -Hstring(required)Experiment hypothesis text
--baseline, -bfloat0.05Baseline conversion rate
--mde, -mfloat0.10Minimum detectable effect as relative lift
--daily-traffic, -dint0Daily eligible traffic
--variants, -vint2Number of variants including control
--jsonflagFalseOutput as JSON
bash
python scripts/experiment_planner.py --hypothesis "Larger CTA will increase signups by 15%"
python scripts/experiment_planner.py -H "Simplified checkout boosts conversions" -b 0.08 -m 0.15 -d 3000
python scripts/experiment_planner.py -H "New pricing page" --json

Troubleshooting

ProblemCauseSolution
Sample size is unrealistically largeMDE too small or baseline too lowIncrease MDE (test bolder changes) or target a higher-traffic page
Test duration exceeds 6 weeksInsufficient daily trafficConsider qualitative methods, test bigger changes, or combine traffic from multiple pages
p-value hovers around 0.05Borderline significanceDo not stop early; run to planned sample size or extend 20%
Results significant but lift is tiny (<1%)Overpowered testCheck practical significance alongside statistical significance
Segment results contradict overallSimpson's paradoxInvestigate segment composition; report both overall and segment results
Variant performs differently on mobile vs desktopDevice-specific UX issuesDesign device-specific variants; increase per-segment sample size
Calculator produces negative CIVery small samples or extreme ratesEnsure sufficient sample size; check data integrity

Success Criteria

CriterionTargetHow to Measure
Tests reach planned sample size100% of testsCompare actual vs planned sample at conclusion
False positive rate<5%Track post-implementation lift vs test prediction
Test velocity2+ tests per team per monthCount experiments documented per sprint
Documentation completeness100% of tests documentedAudit experiment records quarterly
Average test duration<4 weeksMeasure start-to-conclusion calendar days
Decision quality>80% of shipped variants hold gains at 90 daysPost-ship metric tracking

Scope & Limitations

In scope:

  • Hypothesis formulation and validation
  • Sample size and power calculations
  • Frequentist two-proportion z-tests
  • A/B, A/B/n, and split URL test planning
  • Segment-level analysis
  • Pre/post test documentation

Out of scope:

  • Bayesian A/B testing methods (use dedicated Bayesian tools)
  • Multi-armed bandit algorithms (require real-time allocation infrastructure)
  • Multivariate testing (MVT) analysis (combinatorial explosion requires specialized tools)
  • Server-side feature flag implementation (see engineering skills)
  • Revenue-based metrics requiring transaction-level data
  • Sequential testing / always-valid p-values (use Optimizely Stats Engine or similar)

Integration Points

Tool / PlatformIntegration MethodUse Case
PostHog / AmplitudeJSON export from experiment_analyzerFeed results into product analytics
Jira / Linearexperiment_planner JSON outputCreate experiment tickets with metadata
Google SheetsCSV export from experiment_analyzerShare results with non-technical stakeholders
LaunchDarkly / Unleashexperiment_planner checklistPre-launch validation before feature flag rollout
Slack / NotionCopy human-readable outputAsync experiment status updates
CI/CD pipelines--json flag on all scriptsAutomated experiment health checks

© borghei, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 7 other files (scripts) in product-team/ab-test-setup of borghei/Claude-Skills.

  • SKILL.md
  • evals/README.md
  • evals/grader.py
  • evals/runner.py
  • evals/test_cases.json
  • scripts/experiment_analyzer.py
  • scripts/experiment_planner.py
  • scripts/sample_size_calculator.py

Open the folder on GitHubat commit c9a1487

Compare with similar skills

Ab Test Setup next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Ab Test Setup compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Ab Test Setup this skillborghei/Claude-Skills874—~5.3kAutomated safety check: PassMIT
Ad Test Designeraaron-he-zhu/aaron-marketing-skills2.9k2 repos~2.8kAutomated safety check: PassApache-2.0
Ab Test Analyzeririnabuht12-oss/marketing-skills3.8k—~1.4kAutomated safety check: PassNone
Define Hypothesisproduct-on-purpose/pm-skills713—~966Automated safety check: PassApache-2.0
A B Test DesignOwl-Listener/designer-skills2.9k1 repos~472Automated safety check: PassMIT
Ab Test Planindranilbanerjee/digital-marketing-pro8541 repos~2kAutomated safety check: PassMIT

Similar skills

  • Ad Test Designer

    aaron-he-zhu/aaron-marketing-skills

    A skill your agent uses when the user asks to "design an A/B test", "set up a creative/landing test", "run an incrementality test", or "is this result statistically and practically material?"…

    2.9k GitHub starsUsed in 2 repos~2.8k tokens
    Marketing & SEOAuto-check passed
  • Ab Test Analyzer

    irinabuht12-oss/marketing-skills

    Statistical significance calculator for A/B test results with sample size requirements, segment breakdowns, and hypothesis generation.

    3.8k GitHub stars~1.4k tokensUpdated 13 days ago
    Marketing & SEOAuto-check passed
  • Define Hypothesis

    product-on-purpose/pm-skills

    Defines a testable hypothesis with clear success metrics and a validation approach.

    713 GitHub stars~966 tokensUpdated 3 days ago
    Marketing & SEOAuto-check passed
  • A B Test Design

    Owl-Listener/designer-skills

    Design an A/B experiment — hypothesis, variants, primary metric, and sample size.

    2.9k GitHub starsUsed in 1 repo~472 tokens
    Marketing & SEOAuto-check passed
  • Ab Test Plan

    indranilbanerjee/digital-marketing-pro

    Design a statistically rigorous A/B or multivariate test plan — If/Then/Because hypothesis, control and variant specs, required sample size per variant (absolute vs relative MDE via…

    854 GitHub starsUsed in 1 repo~2k tokens
    Marketing & SEOAuto-check passed
  • Ab Testing

    ericrisco/rsc-harness

    A skill your agent uses when designing or analyzing a controlled experiment — falsifiable hypothesis, sample size from an MDE, reading significance/CI/power, CUPED, or rescuing tests that won't go…

    156 GitHub stars~2.4k tokensUpdated yesterday
    Marketing & SEOAuto-check passed

More from borghei/Claude-Skills

All 364 skills in this repo
  • Agents In The Team

    borghei/Claude-Skills

    Run delivery when AI coding and ops agents take tickets. An agent skill from borghei/Claude-Skills.

    874 GitHub stars~4.2k tokensUpdated today
    Auto-check passed
  • AI Content Disclosure

    borghei/Claude-Skills

    Check AI-generated marketing content and reviews for required disclosures under the EU AI Act, FTC rules and platform AI-label policies.

    874 GitHub stars~3.4k tokensUpdated today
    Auto-check passed
  • AI Prototyping

    borghei/Claude-Skills

    Idea to AI-generated prototype to customer validation to engineering handoff.

    874 GitHub stars~3.6k tokensUpdated today
    Auto-check passed
  • Analytics Engineer

    borghei/Claude-Skills

    Analytics engineering across data modeling, dbt, transformation, and semantic layers.

    874 GitHub stars~3.4k tokensUpdated today
    Auto-check passed
  • Ansoff Matrix

    borghei/Claude-Skills

    Ansoff Matrix — 4-quadrant framework for growth options: market penetration, market/product development, and diversification.

    874 GitHub stars~2.2k tokensUpdated today
    Auto-check passed
  • Brainstorm Okrs

    borghei/Claude-Skills

    OKR brainstorming and validation using the Radical Focus framework — outcome objectives, measurable key results, counter-metrics.

    874 GitHub stars~1.4k tokensUpdated today
    Auto-check passed

Categories

Questions about Ab Test Setup

What does Ab Test Setup do?

Design and run statistically rigorous A/B tests and experiments. Ab Test Setup is an agent skill from borghei/Claude-Skills. Design and run statistically rigorous A/B tests and experiments.

When should I use Ab Test Setup?

Ab Test Setup fits situations like: planning experiments; calculating sample sizes; designing test variants; selecting metrics.

How do I install Ab Test Setup in Claude Code?

Run `npx skills add borghei/Claude-Skills --skill ab-test-setup -a claude-code`. Or copy the skill folder (product-team/ab-test-setup in borghei/Claude-Skills) into .claude/skills/ab-test-setup in your project. Claude Code loads it when a task matches its description.

How do I install Ab Test Setup in Codex?

Run `npx skills add borghei/Claude-Skills --skill ab-test-setup -a codex`. Or copy the skill folder (product-team/ab-test-setup in borghei/Claude-Skills) into .agents/skills/ab-test-setup in your project. Codex loads it when a task matches its description.

Can I use Ab Test Setup in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add borghei/Claude-Skills --skill ab-test-setup -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ab-test-setup, .gemini/skills/ab-test-setup, .github/skills/ab-test-setup and .opencode/skills/ab-test-setup in your project.

What does Ab Test Setup need to run?

Going by SKILL.md and its folder, Ab Test Setup needs Python for the scripts in its folder and the command-line tools its instructions call (python). Our summary lists: Python 3.

Does Ab Test Setup access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Ab Test Setup safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Ab Test Setup use?

Ab Test Setup is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Ab Test Setup use?

About 5.3k tokens (SKILL.md is roughly 21k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Ab Test Setup?

Skills that share tags, products or a category with Ab Test Setup: Ad Test Designer (aaron-he-zhu/aaron-marketing-skills, 2.9k stars), Ab Test Analyzer (irinabuht12-oss/marketing-skills, 3.8k stars), Define Hypothesis (product-on-purpose/pm-skills, 713 stars) and A B Test Design (Owl-Listener/designer-skills, 2.9k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Ab Test Setup?

borghei (a GitHub user) maintains it in borghei/Claude-Skills, which has 874 GitHub stars. The repository holds 364 skills in this directory. The repository was last updated on October 7, 2026.

Source: borghei/Claude-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.