Agent skill

Ab Test Setup

by TerminalSkills in TerminalSkills/skills

Plans a controlled experiment (A/B test) so its result can be trusted: writes the hypothesis, picks one primary metric and the guardrails, computes sample size and run time, specifies how visitors…

Apache-2.0Auto-check passedMarketing & SEO

Install Ab Test Setup

skills CLI
$ npx skills add TerminalSkills/skills --skill ab-test-setup -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install TerminalSkills/skills ab-test-setup --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/TerminalSkills/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/ab-test-setup .claude/skills/ab-test-setup && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
ab-test-setup
GitHub stars
163
Token cost
~4k tokens
SKILL.md length
1,706 words
Files
2
Skills in repo
10
Repo updated
First seen
Licence
Apache-2.0

At a glance

Plans a controlled experiment (A/B test) so its result can be trusted: writes the hypothesis, picks one primary metric and the guardrails, computes sample size and run time, specifies how visitors…

  • Works in 6 steps: Gather the inputs → Do the arithmetic before anything else → Write the plan and commit it → …
  • Someone says set up an A/B test
  • SKILL.md covers Overview, Instructions, Examples and Guidelines
  • Calls python3

What it does

Ab Test Setup is an agent skill from TerminalSkills/skills. Plans a controlled experiment (A/B test) so its result can be trusted: writes the hypothesis, picks one primary metric and the guardrails, computes sample size and run time, specifies how visitors are assigned and when exposure is logged, and reads out the result with a confidence interval. Use when someone says "set up an A/B test", "split test this page", "how many visitors do I need", "how long should the experiment run", "is this result significant", "can I stop the test early", or wants to test a headline…

Its SKILL.md is about 4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 1 other file (for example `_scores.json`). Compatibility notes: Python 3.8+ (standard library only) for the calculator, Node.js 18+ for the assignment snippet. Works with any assignment tool: feature flags, GrowthBook…

It sits in Marketing & SEO, covering A/B testing. The repository describes itself as: Open-source library of AI agent skills — SKILL.md files for Claude Code, Codex, Gemini CLI, Cursor. The licence is Apache-2.0.

When your agent uses it

  • Someone says set up an A/B test
  • Split test this page
  • How many visitors do I need
  • How long should the experiment run

Example prompts

  • “set up an A/B test”
  • “split test this page”
  • “how many visitors do I need”
  • “/ab-test-setup”

Requirements

  • Python 3
  • Node.js
  • Compatibility (from SKILL.md): Python 3.8+ (standard library only) for the calculator, Node.js 18+ for the assignment snippet. Works with any assignment tool: feature flags, GrowthBook, PostHog, Optimizely, Statsig, or a hash split in your own code.

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Gather the inputs
  2. Do the arithmetic before anything else
  3. Write the plan and commit it
  4. Specify assignment and exposure logging
  5. Launch checks and the stopping rule
  6. Read out the result

What it can do on your machine

Read from SKILL.md and the folder at commit a021875. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Python 3.8+ (standard library only) for the calculator, Node.js 18+ for the assignment snippet. Works with any assignment tool: feature flags, GrowthBook, PostHog, Optimizely, Statsig, or a hash split in your own code.

    From compatibility in the SKILL.md frontmatter.

Context cost

Ab Test Setup loads about 4k tokens when it runs. Until then it costs about 149 tokens; SKILL.md has 1,706 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~149
When it runs · the whole SKILL.md, loaded when a task matches
~4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from TerminalSkills/skills at commit a021875, republished under its Apache-2.0 licence (© TerminalSkills). 1,706 words, ~3,999 tokens.

Download SKILL.mdSave it as .claude/skills/ab-test-setup/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
ab-test-setup
description
Plans a controlled experiment (A/B test) so its result can be trusted: writes the hypothesis, picks one primary metric and the guardrails, computes sample size and run time, specifies how visitors are assigned and when exposure is logged, and reads out the result with a confidence interval. Use when someone says "set up an A/B test", "split test this page", "how many visitors do I need", "how long should the experiment run", "is this result significant", "can I stop the test early", or wants to test a headline, price, layout or onboarding change against the current version.
compatibility
Python 3.8+ (standard library only) for the calculator, Node.js 18+ for the assignment snippet. Works with any assignment tool: feature flags, GrowthBook, PostHog, Optimizely, Statsig, or a hash split in your own code.
license
Apache-2.0
metadata.author
terminal-skills
metadata.version
2.0.0
metadata.category
business
metadata.tags
ab-testing, experimentation, statistics, conversion, sample-size

A/B Test Setup

Overview

An A/B test answers one question: did this change move this metric, or was it noise? The answer is only worth something when the sample size, the metric and the stopping rule were fixed before the first visitor arrived. This skill produces three things: a written experiment plan committed before launch, an assignment and logging spec a developer can implement, and a readout that states the effect as an interval instead of a "winner" badge.

Most failed tests fail on arithmetic that could have been done on day zero: the site does not have enough traffic for the effect the team hopes for. Do that arithmetic first.

Instructions

1. Gather the inputs

Ask for what is missing, and look in the project for the rest (feature-flag SDK, analytics events, an experiments/ folder with earlier plans):

  • The change and the evidence behind it. Session recordings, support tickets, a funnel drop, a survey. "We want to try it" is allowed, but record that the prior is weak.
  • Unit of assignment. Logged-in user, account, or anonymous visitor ID. Accounts when users inside one company would otherwise see different prices or features.
  • Primary metric and its baseline, measured over the last 4 full weeks on the same population that will enter the test (same pages, same devices, same traffic sources).
  • Eligible units per week — people who actually reach the changed element, not total site traffic.
  • Smallest lift worth shipping. This is a business question: below what improvement would nobody bother to keep the variant?
  • What else is running on the same pages, and any sale, launch or campaign planned during the run.
2. Do the arithmetic before anything else

For a conversion rate, with control rate p1, variant rate p2 = p1 × (1 + lift) and p̄ their mean, the visitors needed in each arm are:

text
n = ( z_a × sqrt(2 × p̄ × (1 − p̄)) + z_b × sqrt(p1 × (1 − p1) + p2 × (1 − p2)) )² ÷ (p2 − p1)²

z_a = 1.960   two-sided significance level 0.05
z_b = 0.8416  power 0.80 (an effect of this size is detected 4 times out of 5)

Save the calculator as abtest.py. It also covers the two checks used later.

python
"""Sample size, sample-ratio check and readout for a two-arm conversion test. Standard library only."""
import math, sys
from statistics import NormalDist

Z = NormalDist()

def size(base, rel_lift, alpha=0.05, power=0.80, comparisons=1):
    """Visitors per arm to detect base -> base*(1+rel_lift), two-sided."""
    p1, p2 = base, base * (1 + rel_lift)
    za = Z.inv_cdf(1 - alpha / comparisons / 2)
    zb = Z.inv_cdf(power)
    pbar = (p1 + p2) / 2
    root = za * math.sqrt(2 * pbar * (1 - pbar)) + zb * math.sqrt(p1 * (1 - p1) + p2 * (1 - p2))
    return math.ceil(root ** 2 / (p2 - p1) ** 2)

def srm(n_a, n_b, share_a=0.5):
    """Chi-square p-value that the observed split matches the planned one."""
    total = n_a + n_b
    exp_a, exp_b = total * share_a, total * (1 - share_a)
    chi2 = (n_a - exp_a) ** 2 / exp_a + (n_b - exp_b) ** 2 / exp_b
    return math.erfc(math.sqrt(chi2 / 2))

def result(n_a, conv_a, n_b, conv_b, alpha=0.05):
    """Two-proportion z-test plus a confidence interval for the absolute difference."""
    pa, pb = conv_a / n_a, conv_b / n_b
    pooled = (conv_a + conv_b) / (n_a + n_b)
    z = (pb - pa) / math.sqrt(pooled * (1 - pooled) * (1 / n_a + 1 / n_b))
    p_value = 2 * (1 - Z.cdf(abs(z)))
    half = Z.inv_cdf(1 - alpha / 2) * math.sqrt(pa * (1 - pa) / n_a + pb * (1 - pb) / n_b)
    return pa, pb, pb - pa, (pb - pa - half, pb - pa + half), p_value

if __name__ == "__main__":
    cmd, *a = sys.argv[1:]
    if cmd == "size":
        print(size(float(a[0]), float(a[1]), comparisons=int(a[2]) if len(a) > 2 else 1), "visitors per arm")
    elif cmd == "srm":
        p = srm(int(a[0]), int(a[1]))
        print(f"SRM p = {p:.4g} ->", "STOP: assignment is broken" if p < 0.001 else "split is fine")
    elif cmd == "result":
        pa, pb, diff, ci, p = result(*map(int, a))
        print(f"control {pa:.2%}  variant {pb:.2%}  diff {diff:+.2%} points "
              f"(95% CI {ci[0]:+.2%} to {ci[1]:+.2%})  relative {diff / pa:+.1%}  p = {p:.4f}")

Then: weeks = ceil(arms × n ÷ eligible units per week), rounded up to whole weeks, never fewer than 2. Whole weeks matter because weekday and weekend visitors behave differently.

Weeks neededWhat to do
2–6Run it as planned.
7–8Acceptable only for a decision that matters; cookie loss and returning visitors blur the arms over time.
More than 8Do not run this test. Test a bolder change, measure a step with a higher baseline, pool several similar pages into one experiment, or ship and watch the trend.

Other cases:

  • Several variants. Each comparison against control gets 0.05 ÷ number of variants (Bonferroni). Pass the count as the third argument: size 0.04 0.15 2.
  • Unequal split. Total sample grows by 1 ÷ (4 × w × (1 − w)) where w is the control share: 70/30 costs 1.19×, 80/20 costs 1.56×, 90/10 costs 2.78×.
  • Averages (revenue per visitor, order value): n = 2 × σ² × (z_a + z_b)² ÷ δ², about 16 × σ² ÷ δ², with σ the standard deviation and δ the absolute difference to detect. Revenue is heavy-tailed, so state in the plan how extreme orders are capped.
3. Write the plan and commit it

Create experiments/YYYY-MM-DD-short-name.md before launch. Example 1 shows a filled one. It must contain:

  1. Hypothesis in one sentence: the change, who sees it, the metric, the direction, and the evidence.
  2. Primary metric — exactly one, with numerator, denominator and the window in which a conversion counts (for example "paid within 14 days of first exposure").
  3. Guardrails — two to four metrics that must not get worse (refund rate, page errors, support contacts, revenue per visitor), each with the drop that would stop the test.
  4. Sample size, split, start and end dates.
  5. Stopping rule (section 5) and the decision rule for each outcome.
  6. Segments that will be reported, named in advance. Anything else found later is a new hypothesis, not a finding.
4. Specify assignment and exposure logging
javascript
import { createHash } from "node:crypto";

// Same unit + same experiment key -> same arm, on every server and every visit.
export function assign(experimentKey, unitId, arms = ["control", "variant"]) {
  const digest = createHash("sha256").update(`${experimentKey}:${unitId}`).digest();
  const bucket = digest.readUInt32BE(0) / 2 ** 32; // uniform in [0, 1)
  return arms[Math.floor(bucket * arms.length)];
}
  • Hash a stable identifier. A new random draw per page view puts one person in both arms.
  • Put the experiment key in the hash so two experiments split independently of each other.
  • Log an exposure event (experiment_key, arm, unit_id, timestamp) at the moment the person reaches the changed element, not at assignment. Analyse only exposed units, and analyse by the same unit that was randomised.
  • Render the variant on the server or before first paint. A page that flashes the control and then swaps is a different treatment from the one in the plan.
  • For tests that split by URL, follow Google Search's guidance for site tests: no different content for the crawler than for people, rel="canonical" from each variant URL to the original, a 302 redirect rather than a 301, and remove the test when it ends.
5. Launch checks and the stopping rule

Before launch, force each arm with an override and confirm on desktop and mobile that the page renders, the exposure event fires once, and the conversion event carries the arm.

Daily during the run, look only at: python3 abtest.py srm on exposed counts, guardrails, and error rates. A sample-ratio mismatch (p below 0.001) means units are being lost in one arm (redirects, bot filtering, a crash, late logging). The result cannot be read: stop, fix, restart with a new experiment key.

Do not act on the primary metric early. Checking it repeatedly and stopping at the first p < 0.05 produces "winners" from identical pages:

Looks at the primary metricChance an A/A test shows p < 0.05 at some look
1 (at the planned end)5%
514%
1019%
14 (daily for two weeks)22%
28 (daily for four weeks)27%

Pick one rule and write it in the plan:

  • Fixed horizon (default): read the primary metric once, at the planned sample size.
  • Planned interim looks: stop early only if an interim p-value is below 0.001; the final look still uses 0.05. This keeps the overall false-positive rate at about 5%.
  • Sequential or Bayesian engine in the testing tool: follow that tool's own stopping rule and do not mix it with the fixed-horizon numbers above. Sequential methods allow continuous monitoring at the price of wider intervals.
Show full SKILL.md (702 more words)Show less
6. Read out the result

Run python3 abtest.py result with exposed units and conversions per arm. Report the interval first.

Interval for the differenceDecision
Entirely above zero, guardrails healthyShip. Quote the range, not only the point estimate.
Includes zeroNot detected at this sample size. This is not proof of "no effect". Keep control unless the variant is cheaper to maintain and the lower bound is tolerable.
Entirely below zeroKeep control and record what was learned.

Append the readout to the plan file: dates, exposed counts, SRM p-value, the interval, guardrails, the decision, and what to test next.

Examples

Example 1: Signup page headline, with the arithmetic

Prompt: "Our signup page for Tidewater Invoicing converts 4.0% of about 9,000 weekly visitors. We want to test a headline about getting paid faster. A 15% relative lift would be worth it. How long do we run?"

p1 = 0.040, p2 = 0.040 × 1.15 = 0.046, difference 0.006, p̄ = 0.043.

text
z_a × sqrt(2 × 0.043 × 0.957)            = 1.960 × 0.28688 = 0.56228
z_b × sqrt(0.040 × 0.960 + 0.046 × 0.954) = 0.8416 × 0.28685 = 0.24141
n = (0.56228 + 0.24141)² ÷ 0.006² = 0.64592 ÷ 0.000036 = 17,942.2 -> 17,943 per arm

total = 2 × 17,943 = 35,886      35,886 ÷ 9,000 per week = 3.99  ->  4 full weeks

python3 abtest.py size 0.04 0.15 prints 17943 visitors per arm. The plan the agent writes:

markdown
# experiments/2026-10-05-signup-headline.md
Hypothesis: Replacing "Invoicing made simple" with "Get paid 9 days sooner" will raise
  signup completion for new visitors, because 31 of 50 interviewed customers named late
  payment as the reason they went looking.
Unit: anonymous visitor ID (first-party cookie), 50/50 hash split, key signup-headline-2026-10
Primary metric: signups completed ÷ visitors exposed to /signup, same session
Baseline: 4.0% (7 Sep – 4 Oct). Smallest lift worth shipping: +15% relative (4.0% -> 4.6%)
Sample: 17,943 per arm, alpha 0.05 two-sided, power 0.80. Run 5 Oct – 1 Nov (4 full weeks)
Guardrails: activation within 7 days (stop if down more than 10% relative), JS error rate (stop if it doubles)
Stopping rule: fixed horizon. SRM and guardrails checked daily; primary metric read on 2 Nov
Decision: ship if the 95% interval is above zero and activation holds; otherwise keep control
Segments reported: device type, paid vs organic

Four weeks later the counts are 17,990 exposed with 716 signups in control and 17,954 with 811 in the variant.

text
$ python3 abtest.py srm 17990 17954
SRM p = 0.8494 -> split is fine
$ python3 abtest.py result 17990 716 17954 811
control 3.98%  variant 4.52%  diff +0.54% points (95% CI +0.12% to +0.95%)  relative +13.5%  p = 0.0116

Readout: the headline raised signup completion by between 0.12 and 0.95 percentage points (roughly +3% to +24% relative). Ship it, and expect the long-run gain to sit below the +13.5% point estimate.

Example 2: A store without the traffic

Prompt: "Harbor Kiln sells ceramics. The product page gets 1,400 visitors a week and 6% add to cart. Can we A/B test a new 'free shipping over $60' badge? I'd be happy with 10% more add-to-carts."

text
$ python3 abtest.py size 0.06 0.10
25740 visitors per arm        -> 51,480 total ÷ 1,400 per week = 36.8 weeks. Not testable.
$ python3 abtest.py size 0.06 0.30
3112 visitors per arm         -> 6,224 total ÷ 1,400 per week = 4.4 -> 5 full weeks.

The agent answers that a badge alone is unlikely to move add-to-cart by 30%, and that at this traffic only an effect of that size can be seen. It offers two honest routes: bundle the badge with a rebuilt page (new photos, reviews above the fold) and test the bundle for 5 weeks, accepting that the test cannot say which part worked; or add the badge to every product page at once and compare the four weeks after with the four before, labelled as a before/after observation rather than an experiment.

Example 3: A split that is not 50/50

After one week of a redirect test the exposure counts are 18,102 and 17,402.

text
$ python3 abtest.py srm 18102 17402
SRM p = 0.0002032 -> STOP: assignment is broken

The variant lost about 4% of its visitors, most likely during the redirect. The agent stops the test, moves exposure logging before the redirect, and restarts under a new key. The first week's data is discarded, not merged.

Guidelines

  • Stopping at half the planned sample because "it is already significant" halves more than the wait: at 8,972 per arm the test in Example 1 has 51% power, and effects found that way are overstated.
  • Do not extend a test that ended without a detected effect "until it gets there". That is peeking with extra steps. Plan a new test with a larger sample or a bolder change.
  • Relative and absolute lifts are different numbers. "15%" on a 4.0% baseline means 4.6%, not 19%. Write both in the plan.
  • One primary metric. With five metrics and no correction, the chance that at least one shows p < 0.05 by luck is about 23%.
  • A click on a button is a cheap metric with a high baseline, and it often rises while revenue does not. Use it as primary only when the plan says why it predicts the outcome that matters.
  • Changing the variant, the split or the audience during the run starts a new experiment. Restart with a new key.
  • Surprisingly large wins (a +40% lift from a copy change) are more often logging bugs than discoveries. Check SRM, duplicated events and bot traffic before celebrating.
  • The formula assumes independent units. If one account contains many users, randomise and analyse by account.
  • Google Optimize was shut down on 30 September 2023; tutorials built on it no longer apply.
  • When not to test: fewer than about 400 conversions a month on the primary metric (at that volume four weeks can only detect lifts of 30% or more), legal or security fixes, changes that cannot be shown to only part of the audience, and decisions that would be made the same way regardless of the result.

© TerminalSkills, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in skills/ab-test-setup of TerminalSkills/skills.

  • SKILL.md
  • _scores.json

Open the folder on GitHubat commit a021875

Compare with similar skills

Ab Test Setup next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Ab Test Setup compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Ab Test Setup this skillTerminalSkills/skills163—~4kAutomated safety check: PassApache-2.0
Ab Testingcoreyhaines31/marketingskills54k3 repos~2.8kAutomated safety check: PassMIT
AnalyticsNexus-JPF/note-companion8707 repos~2.2kAutomated safety check: PassMIT
Ab Test Setupfreekmurze/dotfiles1k15 repos~1.8kAutomated safety check: PassNone
Ad Test Designeraaron-he-zhu/aaron-marketing-skills2.9k2 repos~2.8kAutomated safety check: PassApache-2.0
Meta Tags Optimizernowork-studio/notfair-plugin3.9k1 repos~2.7kAutomated safety check: PassMIT

Similar skills

  • Ab Testing

    coreyhaines31/marketingskills

    When the user wants to plan, design, or implement an A/B test or experiment, or build a growth experimentation program.

    54k GitHub starsUsed in 3 repos~2.8k tokens
    Marketing & SEOAuto-check passed
  • Analytics

    Nexus-JPF/note-companion

    When the user wants to set up, improve, or audit analytics tracking and measurement.

    870 GitHub starsUsed in 7 repos~2.2k tokens
    Marketing & SEOAuto-check passed
  • Ab Test Setup

    freekmurze/dotfiles

    When the user wants to plan, design, or implement an A/B test or experiment.

    1k GitHub starsUsed in 15 repos~1.8k tokens
    Marketing & SEOAuto-check passed
  • Ad Test Designer

    aaron-he-zhu/aaron-marketing-skills

    A skill your agent uses when the user asks to "design an A/B test", "set up a creative/landing test", "run an incrementality test", or "is this result statistically and practically material?"…

    2.9k GitHub starsUsed in 2 repos~2.8k tokens
    Marketing & SEOAuto-check passed
  • Meta Tags Optimizer

    nowork-studio/notfair-plugin

    Writes and improves title tags, meta descriptions, Open Graph and Twitter card tags for click-through, with character counts and A/B test variants.

    3.9k GitHub starsUsed in 1 repo~2.7k tokens
    Marketing & SEOAuto-check passed
  • Ab Test Analyzer

    irinabuht12-oss/marketing-skills

    Statistical significance calculator for A/B test results with sample size requirements, segment breakdowns, and hypothesis generation.

    3.9k GitHub stars~1.4k tokensUpdated 14 days ago
    Marketing & SEOAuto-check passed

More from TerminalSkills/skills

All 10 skills in this repo
  • Blender Addon Dev

    TerminalSkills/skills

    Build custom Blender add-ons with Python. An agent skill from TerminalSkills/skills.

    163 GitHub stars~2k tokensUpdated 4 days ago
    Auto-check passed
  • Seog

    TerminalSkills/skills

    Run local SEO from your terminal via the SEOG MCP server — Google Business Profile management, map-pack rank tracking and geo-grid scans, review sync and replies published to Google, competitor…

    163 GitHub stars~3.4k tokensUpdated 4 days ago
    Auto-check passed
  • Stripe Connect

    TerminalSkills/skills

    Build marketplace and platform payment flows with Stripe Connect.

    163 GitHub stars~2.9k tokensUpdated 4 days ago
    Auto-check passed
  • 3dsmax Rendering

    TerminalSkills/skills

    Scripts and configures production rendering in Autodesk 3ds Max with the V-Ray and Corona renderers: output size and files, render elements, denoising, light mix, batch and command-line rendering…

    163 GitHub stars~3.8k tokensUpdated 4 days ago
    Auto-check passed
  • 3dsmax Scripting

    TerminalSkills/skills

    Covers scripting Autodesk 3ds Max, the 3D modeling and rendering application, with MAXScript and Python (pymxs): scene manipulation, object creation, material assignment, camera and light setup…

    163 GitHub stars~3.2k tokensUpdated 4 days ago
    Auto-check passed
  • 3proxy

    TerminalSkills/skills

    3proxy is a small open-source proxy server that runs HTTP/HTTPS, SOCKS4/5, SNI and TCP/UDP port-mapping proxies from one config file.

    163 GitHub stars~3.5k tokensUpdated 4 days ago
    Auto-check: notes

Categories

Questions about Ab Test Setup

What does Ab Test Setup do?

Plans a controlled experiment (A/B test) so its result can be trusted: writes the hypothesis, picks one primary metric and the guardrails, computes sample size and run time, specifies how visitors…. Ab Test Setup is an agent skill from TerminalSkills/skills. Plans a controlled experiment (A/B test) so its result can be trusted: writes the hypothesis, picks one primary metric and the guardrails, computes sample size and run time, specifies how visitors are assigned and when exposure is logged, and reads out the result with a confidence interval.

When should I use Ab Test Setup?

Ab Test Setup fits situations like: someone says set up an A/B test; split test this page; how many visitors do I need; how long should the experiment run.

How do I install Ab Test Setup in Claude Code?

Run `npx skills add TerminalSkills/skills --skill ab-test-setup -a claude-code`. Or copy the skill folder (skills/ab-test-setup in TerminalSkills/skills) into .claude/skills/ab-test-setup in your project. Claude Code loads it when a task matches its description.

How do I install Ab Test Setup in Codex?

Run `npx skills add TerminalSkills/skills --skill ab-test-setup -a codex`. Or copy the skill folder (skills/ab-test-setup in TerminalSkills/skills) into .agents/skills/ab-test-setup in your project. Codex loads it when a task matches its description.

Can I use Ab Test Setup in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add TerminalSkills/skills --skill ab-test-setup -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ab-test-setup, .gemini/skills/ab-test-setup, .github/skills/ab-test-setup and .opencode/skills/ab-test-setup in your project.

What does Ab Test Setup need to run?

Going by SKILL.md and its folder, Ab Test Setup needs the command-line tools its instructions call (python3). Our summary lists: Python 3; Node.js. Compatibility (from SKILL.md): Python 3.8+ (standard library only) for the calculator, Node.js 18+ for the assignment snippet. Works with any assignment tool: feature flags, GrowthBook, PostHog, Optimizely, Statsig, or a hash split in your own code..

Does Ab Test Setup access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Ab Test Setup safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Ab Test Setup use?

Ab Test Setup is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Ab Test Setup use?

About 4k tokens (SKILL.md is roughly 16k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Ab Test Setup?

Skills that share tags, products or a category with Ab Test Setup: Ab Testing (coreyhaines31/marketingskills, 54k stars), Analytics (Nexus-JPF/note-companion, 870 stars), Ab Test Setup (freekmurze/dotfiles, 1k stars) and Ad Test Designer (aaron-he-zhu/aaron-marketing-skills, 2.9k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Ab Test Setup?

TerminalSkills (a GitHub organization) maintains it in TerminalSkills/skills, which has 163 GitHub stars. The repository holds 10 skills in this directory. The repository was last updated on October 3, 2026.

Source: TerminalSkills/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.