Ab Testing
coreyhaines31/marketingskills
When the user wants to plan, design, or implement an A/B test or experiment, or build a growth experimentation program.
Plans a controlled experiment (A/B test) so its result can be trusted: writes the hypothesis, picks one primary metric and the guardrails, computes sample size and run time, specifies how visitors…
$ npx skills add TerminalSkills/skills --skill ab-test-setup -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install TerminalSkills/skills ab-test-setup --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/TerminalSkills/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/ab-test-setup .claude/skills/ab-test-setup && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "ab-test-setup" agent skill from https://github.com/TerminalSkills/skills/tree/main/skills/ab-test-setup into .claude/skills/ab-test-setup/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ab-test-setup", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/TerminalSkills/skills/tree/main/skills/ab-test-setupType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add TerminalSkills/skills --skill ab-test-setup -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install TerminalSkills/skills ab-test-setup --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/TerminalSkills/skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/ab-test-setup .agents/skills/ab-test-setup && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "ab-test-setup" agent skill from https://github.com/TerminalSkills/skills/tree/main/skills/ab-test-setup into .agents/skills/ab-test-setup/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ab-test-setup", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add TerminalSkills/skills --skill ab-test-setup -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install TerminalSkills/skills ab-test-setup --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/TerminalSkills/skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/ab-test-setup .cursor/skills/ab-test-setup && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "ab-test-setup" agent skill from https://github.com/TerminalSkills/skills/tree/main/skills/ab-test-setup into .cursor/skills/ab-test-setup/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ab-test-setup", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/TerminalSkills/skills.git --path skills/ab-test-setup--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add TerminalSkills/skills --skill ab-test-setup -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install TerminalSkills/skills ab-test-setup --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/TerminalSkills/skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/ab-test-setup .gemini/skills/ab-test-setup && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "ab-test-setup" agent skill from https://github.com/TerminalSkills/skills/tree/main/skills/ab-test-setup into .gemini/skills/ab-test-setup/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ab-test-setup", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install TerminalSkills/skills ab-test-setupInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add TerminalSkills/skills --skill ab-test-setup -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/TerminalSkills/skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/ab-test-setup .github/skills/ab-test-setup && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "ab-test-setup" agent skill from https://github.com/TerminalSkills/skills/tree/main/skills/ab-test-setup into .github/skills/ab-test-setup/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ab-test-setup", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add TerminalSkills/skills --skill ab-test-setup -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install TerminalSkills/skills ab-test-setup --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/TerminalSkills/skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/ab-test-setup .opencode/skills/ab-test-setup && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "ab-test-setup" agent skill from https://github.com/TerminalSkills/skills/tree/main/skills/ab-test-setup into .opencode/skills/ab-test-setup/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ab-test-setup", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
ab-test-setupPlans a controlled experiment (A/B test) so its result can be trusted: writes the hypothesis, picks one primary metric and the guardrails, computes sample size and run time, specifies how visitors…
Ab Test Setup is an agent skill from TerminalSkills/skills. Plans a controlled experiment (A/B test) so its result can be trusted: writes the hypothesis, picks one primary metric and the guardrails, computes sample size and run time, specifies how visitors are assigned and when exposure is logged, and reads out the result with a confidence interval. Use when someone says "set up an A/B test", "split test this page", "how many visitors do I need", "how long should the experiment run", "is this result significant", "can I stop the test early", or wants to test a headline…
Its SKILL.md is about 4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 1 other file (for example `_scores.json`). Compatibility notes: Python 3.8+ (standard library only) for the calculator, Node.js 18+ for the assignment snippet. Works with any assignment tool: feature flags, GrowthBook…
It sits in Marketing & SEO, covering A/B testing. The repository describes itself as: Open-source library of AI agent skills — SKILL.md files for Claude Code, Codex, Gemini CLI, Cursor. The licence is Apache-2.0.
6 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit a021875. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
python3From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Python 3.8+ (standard library only) for the calculator, Node.js 18+ for the assignment snippet. Works with any assignment tool: feature flags, GrowthBook, PostHog, Optimizely, Statsig, or a hash split in your own code.
From compatibility in the SKILL.md frontmatter.
Ab Test Setup loads about 4k tokens when it runs. Until then it costs about 149 tokens; SKILL.md has 1,706 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from TerminalSkills/skills at commit a021875, republished under its Apache-2.0 licence (© TerminalSkills). 1,706 words, ~3,999 tokens.
.claude/skills/ab-test-setup/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.An A/B test answers one question: did this change move this metric, or was it noise? The answer is only worth something when the sample size, the metric and the stopping rule were fixed before the first visitor arrived. This skill produces three things: a written experiment plan committed before launch, an assignment and logging spec a developer can implement, and a readout that states the effect as an interval instead of a "winner" badge.
Most failed tests fail on arithmetic that could have been done on day zero: the site does not have enough traffic for the effect the team hopes for. Do that arithmetic first.
Ask for what is missing, and look in the project for the rest (feature-flag SDK, analytics events, an experiments/ folder with earlier plans):
For a conversion rate, with control rate p1, variant rate p2 = p1 × (1 + lift) and p̄ their mean, the visitors needed in each arm are:
n = ( z_a × sqrt(2 × p̄ × (1 − p̄)) + z_b × sqrt(p1 × (1 − p1) + p2 × (1 − p2)) )² ÷ (p2 − p1)²
z_a = 1.960 two-sided significance level 0.05
z_b = 0.8416 power 0.80 (an effect of this size is detected 4 times out of 5)Save the calculator as abtest.py. It also covers the two checks used later.
"""Sample size, sample-ratio check and readout for a two-arm conversion test. Standard library only."""
import math, sys
from statistics import NormalDist
Z = NormalDist()
def size(base, rel_lift, alpha=0.05, power=0.80, comparisons=1):
"""Visitors per arm to detect base -> base*(1+rel_lift), two-sided."""
p1, p2 = base, base * (1 + rel_lift)
za = Z.inv_cdf(1 - alpha / comparisons / 2)
zb = Z.inv_cdf(power)
pbar = (p1 + p2) / 2
root = za * math.sqrt(2 * pbar * (1 - pbar)) + zb * math.sqrt(p1 * (1 - p1) + p2 * (1 - p2))
return math.ceil(root ** 2 / (p2 - p1) ** 2)
def srm(n_a, n_b, share_a=0.5):
"""Chi-square p-value that the observed split matches the planned one."""
total = n_a + n_b
exp_a, exp_b = total * share_a, total * (1 - share_a)
chi2 = (n_a - exp_a) ** 2 / exp_a + (n_b - exp_b) ** 2 / exp_b
return math.erfc(math.sqrt(chi2 / 2))
def result(n_a, conv_a, n_b, conv_b, alpha=0.05):
"""Two-proportion z-test plus a confidence interval for the absolute difference."""
pa, pb = conv_a / n_a, conv_b / n_b
pooled = (conv_a + conv_b) / (n_a + n_b)
z = (pb - pa) / math.sqrt(pooled * (1 - pooled) * (1 / n_a + 1 / n_b))
p_value = 2 * (1 - Z.cdf(abs(z)))
half = Z.inv_cdf(1 - alpha / 2) * math.sqrt(pa * (1 - pa) / n_a + pb * (1 - pb) / n_b)
return pa, pb, pb - pa, (pb - pa - half, pb - pa + half), p_value
if __name__ == "__main__":
cmd, *a = sys.argv[1:]
if cmd == "size":
print(size(float(a[0]), float(a[1]), comparisons=int(a[2]) if len(a) > 2 else 1), "visitors per arm")
elif cmd == "srm":
p = srm(int(a[0]), int(a[1]))
print(f"SRM p = {p:.4g} ->", "STOP: assignment is broken" if p < 0.001 else "split is fine")
elif cmd == "result":
pa, pb, diff, ci, p = result(*map(int, a))
print(f"control {pa:.2%} variant {pb:.2%} diff {diff:+.2%} points "
f"(95% CI {ci[0]:+.2%} to {ci[1]:+.2%}) relative {diff / pa:+.1%} p = {p:.4f}")Then: weeks = ceil(arms × n ÷ eligible units per week), rounded up to whole weeks, never fewer than 2. Whole weeks matter because weekday and weekend visitors behave differently.
| Weeks needed | What to do |
|---|---|
| 2–6 | Run it as planned. |
| 7–8 | Acceptable only for a decision that matters; cookie loss and returning visitors blur the arms over time. |
| More than 8 | Do not run this test. Test a bolder change, measure a step with a higher baseline, pool several similar pages into one experiment, or ship and watch the trend. |
Other cases:
0.05 ÷ number of variants (Bonferroni). Pass the count as the third argument: size 0.04 0.15 2.1 ÷ (4 × w × (1 − w)) where w is the control share: 70/30 costs 1.19×, 80/20 costs 1.56×, 90/10 costs 2.78×.n = 2 × σ² × (z_a + z_b)² ÷ δ², about 16 × σ² ÷ δ², with σ the standard deviation and δ the absolute difference to detect. Revenue is heavy-tailed, so state in the plan how extreme orders are capped.Create experiments/YYYY-MM-DD-short-name.md before launch. Example 1 shows a filled one. It must contain:
import { createHash } from "node:crypto";
// Same unit + same experiment key -> same arm, on every server and every visit.
export function assign(experimentKey, unitId, arms = ["control", "variant"]) {
const digest = createHash("sha256").update(`${experimentKey}:${unitId}`).digest();
const bucket = digest.readUInt32BE(0) / 2 ** 32; // uniform in [0, 1)
return arms[Math.floor(bucket * arms.length)];
}experiment_key, arm, unit_id, timestamp) at the moment the person reaches the changed element, not at assignment. Analyse only exposed units, and analyse by the same unit that was randomised.rel="canonical" from each variant URL to the original, a 302 redirect rather than a 301, and remove the test when it ends.Before launch, force each arm with an override and confirm on desktop and mobile that the page renders, the exposure event fires once, and the conversion event carries the arm.
Daily during the run, look only at: python3 abtest.py srm on exposed counts, guardrails, and error rates. A sample-ratio mismatch (p below 0.001) means units are being lost in one arm (redirects, bot filtering, a crash, late logging). The result cannot be read: stop, fix, restart with a new experiment key.
Do not act on the primary metric early. Checking it repeatedly and stopping at the first p < 0.05 produces "winners" from identical pages:
| Looks at the primary metric | Chance an A/A test shows p < 0.05 at some look |
|---|---|
| 1 (at the planned end) | 5% |
| 5 | 14% |
| 10 | 19% |
| 14 (daily for two weeks) | 22% |
| 28 (daily for four weeks) | 27% |
Pick one rule and write it in the plan:
Run python3 abtest.py result with exposed units and conversions per arm. Report the interval first.
| Interval for the difference | Decision |
|---|---|
| Entirely above zero, guardrails healthy | Ship. Quote the range, not only the point estimate. |
| Includes zero | Not detected at this sample size. This is not proof of "no effect". Keep control unless the variant is cheaper to maintain and the lower bound is tolerable. |
| Entirely below zero | Keep control and record what was learned. |
Append the readout to the plan file: dates, exposed counts, SRM p-value, the interval, guardrails, the decision, and what to test next.
Prompt: "Our signup page for Tidewater Invoicing converts 4.0% of about 9,000 weekly visitors. We want to test a headline about getting paid faster. A 15% relative lift would be worth it. How long do we run?"
p1 = 0.040, p2 = 0.040 × 1.15 = 0.046, difference 0.006, p̄ = 0.043.
z_a × sqrt(2 × 0.043 × 0.957) = 1.960 × 0.28688 = 0.56228
z_b × sqrt(0.040 × 0.960 + 0.046 × 0.954) = 0.8416 × 0.28685 = 0.24141
n = (0.56228 + 0.24141)² ÷ 0.006² = 0.64592 ÷ 0.000036 = 17,942.2 -> 17,943 per arm
total = 2 × 17,943 = 35,886 35,886 ÷ 9,000 per week = 3.99 -> 4 full weekspython3 abtest.py size 0.04 0.15 prints 17943 visitors per arm. The plan the agent writes:
# experiments/2026-10-05-signup-headline.md
Hypothesis: Replacing "Invoicing made simple" with "Get paid 9 days sooner" will raise
signup completion for new visitors, because 31 of 50 interviewed customers named late
payment as the reason they went looking.
Unit: anonymous visitor ID (first-party cookie), 50/50 hash split, key signup-headline-2026-10
Primary metric: signups completed ÷ visitors exposed to /signup, same session
Baseline: 4.0% (7 Sep – 4 Oct). Smallest lift worth shipping: +15% relative (4.0% -> 4.6%)
Sample: 17,943 per arm, alpha 0.05 two-sided, power 0.80. Run 5 Oct – 1 Nov (4 full weeks)
Guardrails: activation within 7 days (stop if down more than 10% relative), JS error rate (stop if it doubles)
Stopping rule: fixed horizon. SRM and guardrails checked daily; primary metric read on 2 Nov
Decision: ship if the 95% interval is above zero and activation holds; otherwise keep control
Segments reported: device type, paid vs organicFour weeks later the counts are 17,990 exposed with 716 signups in control and 17,954 with 811 in the variant.
$ python3 abtest.py srm 17990 17954
SRM p = 0.8494 -> split is fine
$ python3 abtest.py result 17990 716 17954 811
control 3.98% variant 4.52% diff +0.54% points (95% CI +0.12% to +0.95%) relative +13.5% p = 0.0116Readout: the headline raised signup completion by between 0.12 and 0.95 percentage points (roughly +3% to +24% relative). Ship it, and expect the long-run gain to sit below the +13.5% point estimate.
Prompt: "Harbor Kiln sells ceramics. The product page gets 1,400 visitors a week and 6% add to cart. Can we A/B test a new 'free shipping over $60' badge? I'd be happy with 10% more add-to-carts."
$ python3 abtest.py size 0.06 0.10
25740 visitors per arm -> 51,480 total ÷ 1,400 per week = 36.8 weeks. Not testable.
$ python3 abtest.py size 0.06 0.30
3112 visitors per arm -> 6,224 total ÷ 1,400 per week = 4.4 -> 5 full weeks.The agent answers that a badge alone is unlikely to move add-to-cart by 30%, and that at this traffic only an effect of that size can be seen. It offers two honest routes: bundle the badge with a rebuilt page (new photos, reviews above the fold) and test the bundle for 5 weeks, accepting that the test cannot say which part worked; or add the badge to every product page at once and compare the four weeks after with the four before, labelled as a before/after observation rather than an experiment.
After one week of a redirect test the exposure counts are 18,102 and 17,402.
$ python3 abtest.py srm 18102 17402
SRM p = 0.0002032 -> STOP: assignment is brokenThe variant lost about 4% of its visitors, most likely during the redirect. The agent stops the test, moves exposure logging before the redirect, and restarts under a new key. The first week's data is discarded, not merged.
© TerminalSkills, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 1 other file in skills/ab-test-setup of TerminalSkills/skills.
Open the folder on GitHubat commit a021875
Ab Test Setup next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Ab Test Setup this skillTerminalSkills/skills | 163 | — | ~4k | Automated safety check: Pass | Apache-2.0 | |
| Ab Testingcoreyhaines31/marketingskills | 54k | 3 repos | ~2.8k | Automated safety check: Pass | MIT | |
| AnalyticsNexus-JPF/note-companion | 870 | 7 repos | ~2.2k | Automated safety check: Pass | MIT | |
| Ab Test Setupfreekmurze/dotfiles | 1k | 15 repos | ~1.8k | Automated safety check: Pass | None | |
| Ad Test Designeraaron-he-zhu/aaron-marketing-skills | 2.9k | 2 repos | ~2.8k | Automated safety check: Pass | Apache-2.0 | |
| Meta Tags Optimizernowork-studio/notfair-plugin | 3.9k | 1 repos | ~2.7k | Automated safety check: Pass | MIT |
coreyhaines31/marketingskills
When the user wants to plan, design, or implement an A/B test or experiment, or build a growth experimentation program.
Nexus-JPF/note-companion
When the user wants to set up, improve, or audit analytics tracking and measurement.
freekmurze/dotfiles
When the user wants to plan, design, or implement an A/B test or experiment.
aaron-he-zhu/aaron-marketing-skills
A skill your agent uses when the user asks to "design an A/B test", "set up a creative/landing test", "run an incrementality test", or "is this result statistically and practically material?"…
nowork-studio/notfair-plugin
Writes and improves title tags, meta descriptions, Open Graph and Twitter card tags for click-through, with character counts and A/B test variants.
irinabuht12-oss/marketing-skills
Statistical significance calculator for A/B test results with sample size requirements, segment breakdowns, and hypothesis generation.
TerminalSkills/skills
Build custom Blender add-ons with Python. An agent skill from TerminalSkills/skills.
TerminalSkills/skills
Run local SEO from your terminal via the SEOG MCP server — Google Business Profile management, map-pack rank tracking and geo-grid scans, review sync and replies published to Google, competitor…
TerminalSkills/skills
Build marketplace and platform payment flows with Stripe Connect.
TerminalSkills/skills
Scripts and configures production rendering in Autodesk 3ds Max with the V-Ray and Corona renderers: output size and files, render elements, denoising, light mix, batch and command-line rendering…
TerminalSkills/skills
Covers scripting Autodesk 3ds Max, the 3D modeling and rendering application, with MAXScript and Python (pymxs): scene manipulation, object creation, material assignment, camera and light setup…
TerminalSkills/skills
3proxy is a small open-source proxy server that runs HTTP/HTTPS, SOCKS4/5, SNI and TCP/UDP port-mapping proxies from one config file.
Categories
Plans a controlled experiment (A/B test) so its result can be trusted: writes the hypothesis, picks one primary metric and the guardrails, computes sample size and run time, specifies how visitors…. Ab Test Setup is an agent skill from TerminalSkills/skills. Plans a controlled experiment (A/B test) so its result can be trusted: writes the hypothesis, picks one primary metric and the guardrails, computes sample size and run time, specifies how visitors are assigned and when exposure is logged, and reads out the result with a confidence interval.
Ab Test Setup fits situations like: someone says set up an A/B test; split test this page; how many visitors do I need; how long should the experiment run.
Run `npx skills add TerminalSkills/skills --skill ab-test-setup -a claude-code`. Or copy the skill folder (skills/ab-test-setup in TerminalSkills/skills) into .claude/skills/ab-test-setup in your project. Claude Code loads it when a task matches its description.
Run `npx skills add TerminalSkills/skills --skill ab-test-setup -a codex`. Or copy the skill folder (skills/ab-test-setup in TerminalSkills/skills) into .agents/skills/ab-test-setup in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add TerminalSkills/skills --skill ab-test-setup -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ab-test-setup, .gemini/skills/ab-test-setup, .github/skills/ab-test-setup and .opencode/skills/ab-test-setup in your project.
Going by SKILL.md and its folder, Ab Test Setup needs the command-line tools its instructions call (python3). Our summary lists: Python 3; Node.js. Compatibility (from SKILL.md): Python 3.8+ (standard library only) for the calculator, Node.js 18+ for the assignment snippet. Works with any assignment tool: feature flags, GrowthBook, PostHog, Optimizely, Statsig, or a hash split in your own code..
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Ab Test Setup is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 4k tokens (SKILL.md is roughly 16k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Ab Test Setup: Ab Testing (coreyhaines31/marketingskills, 54k stars), Analytics (Nexus-JPF/note-companion, 870 stars), Ab Test Setup (freekmurze/dotfiles, 1k stars) and Ad Test Designer (aaron-he-zhu/aaron-marketing-skills, 2.9k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
TerminalSkills (a GitHub organization) maintains it in TerminalSkills/skills, which has 163 GitHub stars. The repository holds 10 skills in this directory. The repository was last updated on October 3, 2026.
Source: TerminalSkills/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.