Ad Test Designer
aaron-he-zhu/aaron-marketing-skills
A skill your agent uses when the user asks to "design an A/B test", "set up a creative/landing test", "run an incrementality test", or "is this result statistically and practically material?"…
Design and run statistically rigorous A/B tests and experiments.
$ npx skills add borghei/Claude-Skills --skill ab-test-setup -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install borghei/Claude-Skills ab-test-setup --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/borghei/Claude-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/product-team/ab-test-setup .claude/skills/ab-test-setup && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "ab-test-setup" agent skill from https://github.com/borghei/Claude-Skills/tree/main/product-team/ab-test-setup into .claude/skills/ab-test-setup/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ab-test-setup", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/borghei/Claude-Skills/tree/main/product-team/ab-test-setupType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add borghei/Claude-Skills --skill ab-test-setup -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install borghei/Claude-Skills ab-test-setup --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/borghei/Claude-Skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/product-team/ab-test-setup .agents/skills/ab-test-setup && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "ab-test-setup" agent skill from https://github.com/borghei/Claude-Skills/tree/main/product-team/ab-test-setup into .agents/skills/ab-test-setup/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ab-test-setup", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add borghei/Claude-Skills --skill ab-test-setup -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install borghei/Claude-Skills ab-test-setup --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/borghei/Claude-Skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/product-team/ab-test-setup .cursor/skills/ab-test-setup && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "ab-test-setup" agent skill from https://github.com/borghei/Claude-Skills/tree/main/product-team/ab-test-setup into .cursor/skills/ab-test-setup/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ab-test-setup", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/borghei/Claude-Skills.git --path product-team/ab-test-setup--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add borghei/Claude-Skills --skill ab-test-setup -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install borghei/Claude-Skills ab-test-setup --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/borghei/Claude-Skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/product-team/ab-test-setup .gemini/skills/ab-test-setup && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "ab-test-setup" agent skill from https://github.com/borghei/Claude-Skills/tree/main/product-team/ab-test-setup into .gemini/skills/ab-test-setup/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ab-test-setup", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install borghei/Claude-Skills ab-test-setupInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add borghei/Claude-Skills --skill ab-test-setup -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/borghei/Claude-Skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/product-team/ab-test-setup .github/skills/ab-test-setup && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "ab-test-setup" agent skill from https://github.com/borghei/Claude-Skills/tree/main/product-team/ab-test-setup into .github/skills/ab-test-setup/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ab-test-setup", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add borghei/Claude-Skills --skill ab-test-setup -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install borghei/Claude-Skills ab-test-setup --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/borghei/Claude-Skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/product-team/ab-test-setup .opencode/skills/ab-test-setup && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "ab-test-setup" agent skill from https://github.com/borghei/Claude-Skills/tree/main/product-team/ab-test-setup into .opencode/skills/ab-test-setup/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ab-test-setup", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
ab-test-setupDesign and run statistically rigorous A/B tests and experiments.
Ab Test Setup is an agent skill from borghei/Claude-Skills. Design and run statistically rigorous A/B tests and experiments. Use when planning experiments, calculating sample sizes, designing test variants, selecting metrics, analyzing results, or when someone says "let's test that."
Its SKILL.md is about 5.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 9 other files, including scripts (for example `evals/README.md`, `evals/grader.py` and `evals/runner.py`).
It sits in Marketing & SEO, covering A/B testing and Experimental design. The repository describes itself as: 385 AI skills, 77 expert agents, and 900 stdlib Python tools for every team: engineering, PM, marketing, C-level, compliance, business ops, research, and a LinkedIn toolkit… The licence is MIT.
7 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit c9a1487. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 3 files in scripts/ (Python), which the agent can run.
Shell commands in SKILL.md call:
pythonFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Ab Test Setup loads about 5.3k tokens when it runs. Until then it costs about 60 tokens; SKILL.md has 2,168 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from borghei/Claude-Skills at commit c9a1487, republished under its MIT licence (© borghei). 2,168 words, ~5,270 tokens.
.claude/skills/ab-test-setup/SKILL.md (or your agent's skills folder). This skill also uses 7 other files; get the full folder from GitHub.Category: Product Team Tags: A/B testing, experiments, statistical significance, sample size, feature flags, hypothesis testing
A/B Test Setup provides the complete framework for designing experiments that produce statistically valid, actionable results. Most A/B tests fail not because the variant was wrong, but because the test was poorly designed: wrong sample size, wrong metric, or someone peeked at results and stopped early. This skill prevents those mistakes.
Before designing the experiment, confirm these inputs. If any is unknown or vague, ASK — do not assume:
Stop rule: ask only the 2-3 that most change the output. If the user says "just draft it," proceed and list your assumptions at the top of the artifact.
1. HYPOTHESIZE → 2. DESIGN → 3. CALCULATE → 4. IMPLEMENT
↑ │
│ ▼
7. ITERATE ← 6. DOCUMENT ← 5. ANALYZE ← [Run to completion]Because [observation or data point],
we believe [specific change]
will cause [measurable outcome]
for [defined audience segment].
We'll know this is true when [primary metric] changes by [minimum detectable effect].
We'll watch [guardrail metrics] to ensure no negative impact.| Quality | Hypothesis | Problem |
|---|---|---|
| Bad | "Changing the button color might increase clicks" | No data basis, no target, no measurement plan |
| Mediocre | "A green button will get more clicks than blue" | No "why", no target size, no guardrails |
| Good | "Because heatmaps show 40% of users don't notice our CTA, making the button 2x larger with contrasting color will increase CTA clicks by 15%+ for new visitors. Guardrail: page load time stays under 2s." | Data-backed, specific change, measurable outcome, defined audience, guardrail |
| Source | What to Look For | Example |
|---|---|---|
| Analytics data | Drop-off points, low-performing pages | "80% of users drop off at step 3 of onboarding" |
| User research | Confusion, frustration, unmet needs | "Users don't understand what the product does from the homepage" |
| Heatmaps/session recordings | Ignored elements, rage clicks | "Nobody scrolls past the fold on pricing page" |
| Support tickets | Recurring complaints, feature confusion | "Users constantly ask how to invite team members" |
| Competitor analysis | Different approaches to same problem | "Competitor uses a wizard; we use a form" |
| Sales objections | Common reasons prospects don't convert | "Prospects want to see pricing before signing up" |
| Type | Variants | Traffic Need | Best For |
|---|---|---|---|
| A/B | 2 (control + 1 variant) | Moderate | Single change validation |
| A/B/n | 3+ variants | High | Comparing multiple approaches |
| Multivariate (MVT) | Combinations of changes | Very high | Optimizing multiple elements |
| Split URL | Different pages | Moderate | Major redesigns |
| Bandit | Dynamic allocation | Low-moderate | Revenue optimization |
Default recommendation: Standard A/B test. Only use A/B/n or MVT when you have enough traffic and a specific need.
| Category | High Impact | Medium Impact | Low Impact |
|---|---|---|---|
| Copy | Headline/value prop, CTA text | Body copy, social proof | Microcopy, labels |
| Design | Page layout, above-fold content | Visual hierarchy, imagery | Color, font size |
| UX | Number of steps, form fields | Button placement, navigation | Animations, transitions |
| Pricing | Price point, plan names | Feature packaging, anchoring | Billing frequency display |
| Social Proof | Testimonials vs none, logos | Testimonial format, placement | Testimonial count |
Every test needs three types of metrics:
Primary Metric (1 only)
Secondary Metrics (2-3)
Guardrail Metrics (1-3)
Minimum visitors PER VARIANT needed (95% confidence, 80% power):
| Baseline Rate | 5% Lift | 10% Lift | 15% Lift | 20% Lift | 50% Lift |
|---|---|---|---|---|---|
| 1% | 620,000 | 156,000 | 70,000 | 39,000 | 6,400 |
| 2% | 305,000 | 77,000 | 34,000 | 19,500 | 3,200 |
| 3% | 200,000 | 51,000 | 23,000 | 12,800 | 2,100 |
| 5% | 116,000 | 29,500 | 13,200 | 7,500 | 1,250 |
| 10% | 54,000 | 13,800 | 6,200 | 3,500 | 600 |
| 20% | 24,000 | 6,200 | 2,800 | 1,600 | 280 |
| 50% | 6,100 | 1,600 | 720 | 410 | 75 |
Duration (days) = (Sample size per variant * Number of variants) / Daily traffic to test pageMinimum duration: 7 days (to capture day-of-week effects) Maximum recommended: 6 weeks (beyond this, external factors contaminate results)
| Situation | Solution |
|---|---|
| Need 100K visitors, get 5K/week | Increase minimum detectable effect (test bolder changes) |
| Very low traffic (<1K/week) | Use qualitative testing (user testing, surveys) instead |
| Medium traffic (5-20K/week) | Run for 4-6 weeks, test big changes only |
| High traffic (50K+/week) | You can test subtle changes, run multiple tests |
JavaScript modifies the page after initial render.
Pros: Quick to implement, no deploy needed Cons: Can cause flicker (flash of original content), blocked by ad blockers Tools: PostHog, Optimizely, VWO, Google Optimize
Anti-flicker pattern:
// Add to <head> before any rendering
<style>.ab-test-hide { opacity: 0 !important; }</style>
<script>document.documentElement.classList.add('ab-test-hide');</script>
// In your test script (runs after variant assignment):
document.documentElement.classList.remove('ab-test-hide');Variant determined before page renders. No flicker, no client-side dependency.
Pros: No flicker, not blocked by ad blockers, works for logged-in features Cons: Requires engineering work, deploy needed Tools: PostHog, LaunchDarkly, Split, Unleash, custom feature flags
Basic feature flag pattern:
# Server-side variant assignment
def get_variant(user_id: str, experiment: str) -> str:
# Deterministic hash ensures same user always sees same variant
hash_input = f"{user_id}:{experiment}"
hash_value = hashlib.md5(hash_input.encode()).hexdigest()
bucket = int(hash_value[:8], 16) % 100
if bucket < 50:
return "control"
else:
return "variant"| Strategy | Split | When to Use |
|---|---|---|
| Standard | 50/50 | Default. Maximum statistical power. |
| Conservative | 90/10 or 80/20 | Risky changes, revenue-impacting tests |
| Ramped | Start 95/5, increase to 50/50 | New infrastructure, technical risk |
Critical rules:
DO:
DO NOT:
Looking at results before reaching the planned sample size and stopping because one variant looks better leads to a 25-40% false positive rate (vs the intended 5%).
Why: Statistical significance fluctuates wildly with small samples. A variant can show p < 0.05 at 20% of planned sample size and p > 0.30 at full sample.
Solutions:
| Result | Primary Metric | Confidence | Action |
|---|---|---|---|
| Clear winner | Variant +15%, p < 0.01 | High | Implement variant |
| Modest winner | Variant +5%, p < 0.05 | Medium | Implement if easy, else run longer |
| Flat | < 2% difference, p > 0.20 | High (no effect) | Keep control, test something bolder |
| Loser | Variant -10%, p < 0.05 | High | Keep control, investigate why |
| Inconclusive | 5% difference, p = 0.08 | Low | Need more traffic or bolder test |
| Mixed signals | Primary up, guardrail down | Investigate | Dig into segments, do not ship blindly |
| Mistake | Consequence | Prevention |
|---|---|---|
| Stopping at first significance | 25-40% false positive rate | Commit to sample size |
| Cherry-picking segments | Finding "winners" that don't replicate | Pre-register segments of interest |
| Ignoring confidence intervals | Overestimating effect size | Always report CI alongside p-value |
| Multiple comparisons | Inflated Type I error | Bonferroni correction for A/B/n |
| Survivorship bias | Only analyzing users who completed flow | Include all users from assignment point |
| Simpson's paradox | Aggregate hides segment reversal | Always check key segments |
Every test must be documented, regardless of outcome.
EXPERIMENT: [Name]
DATE: [Start] to [End]
OWNER: [Name]
HYPOTHESIS:
Because [observation], we believed [change] would cause [outcome] for [audience].
VARIANTS:
- Control: [description]
- Variant: [description + screenshot]
METRICS:
- Primary: [metric] (baseline: [X]%, MDE: [Y]%)
- Secondary: [metrics]
- Guardrails: [metrics]
RESULTS:
- Sample size: [actual] / [planned]
- Duration: [X] days
- Primary metric: Control [X]% vs Variant [Y]% (p = [Z], CI: [range])
- Secondary metrics: [results]
- Guardrails: [all clear / violation noted]
DECISION: [Ship variant / Keep control / Iterate]
LEARNINGS:
- [What we learned about our users]
- [What we'd do differently next time]| Factor | Score (1-10) | Question |
|---|---|---|
| Impact | How much will this move the metric? | Big change to primary KPI = 10 |
| Confidence | How sure are we it will work? | Strong data supporting hypothesis = 10 |
| Ease | How easy is it to implement and measure? | Can ship in a day = 10 |
ICE Score = (Impact + Confidence + Ease) / 3
Rank all test ideas by ICE score. Run highest first.
| # | Hypothesis | Primary Metric | ICE | Est. Duration | Status |
|---|---|---|---|---|---|
| 1 | Larger CTA increases signups | Signup rate | 8.3 | 2 weeks | Ready |
| 2 | Social proof on pricing increases conversion | Plan selection rate | 7.0 | 3 weeks | Needs design |
| 3 | Shorter onboarding increases activation | Feature activation | 6.7 | 4 weeks | In backlog |
| Skill | Use When |
|---|---|
| analytics-tracking | Setting up event tracking that feeds experiment metrics |
| campaign-analytics | Folding experiment results into broader attribution |
| launch-strategy | Testing within a product launch sequence |
| prompt-engineer-toolkit | A/B testing AI prompts in production |
Calculates required sample size per variant using the normal approximation to the two-proportion z-test. Includes Bonferroni correction for multi-variant tests and duration estimation.
| Flag | Type | Default | Description |
|---|---|---|---|
--baseline, -b | float | (required) | Baseline conversion rate (e.g. 0.05 for 5%) |
--mde, -m | float | (required) | Minimum detectable effect as relative lift (e.g. 0.10 for 10%) |
--alpha, -a | float | 0.05 | Significance level |
--power, -p | float | 0.80 | Statistical power |
--variants, -v | int | 2 | Number of variants including control |
--daily-traffic, -d | int | 0 | Daily eligible traffic for duration estimation |
--one-tailed | flag | False | Use one-tailed test instead of two-tailed |
--json | flag | False | Output as JSON |
python scripts/sample_size_calculator.py --baseline 0.05 --mde 0.10
python scripts/sample_size_calculator.py --baseline 0.12 --mde 0.15 --power 0.9 --daily-traffic 5000
python scripts/sample_size_calculator.py --baseline 0.05 --mde 0.10 --variants 3 --jsonAnalyzes A/B test results using the two-proportion z-test with confidence intervals and segment breakdown.
| Flag | Type | Default | Description |
|---|---|---|---|
input | positional | (required) | CSV file with results or "sample" to create sample |
--alpha, -a | float | 0.05 | Significance level |
--json | flag | False | Output as JSON |
CSV format: variant,visitors,conversions,segment
python scripts/experiment_analyzer.py sample
python scripts/experiment_analyzer.py results.csv
python scripts/experiment_analyzer.py results.csv --alpha 0.01 --jsonGenerates a structured experiment plan from a hypothesis text, including metric selection, sample size, timeline, risks, and documentation template.
| Flag | Type | Default | Description |
|---|---|---|---|
--hypothesis, -H | string | (required) | Experiment hypothesis text |
--baseline, -b | float | 0.05 | Baseline conversion rate |
--mde, -m | float | 0.10 | Minimum detectable effect as relative lift |
--daily-traffic, -d | int | 0 | Daily eligible traffic |
--variants, -v | int | 2 | Number of variants including control |
--json | flag | False | Output as JSON |
python scripts/experiment_planner.py --hypothesis "Larger CTA will increase signups by 15%"
python scripts/experiment_planner.py -H "Simplified checkout boosts conversions" -b 0.08 -m 0.15 -d 3000
python scripts/experiment_planner.py -H "New pricing page" --json| Problem | Cause | Solution |
|---|---|---|
| Sample size is unrealistically large | MDE too small or baseline too low | Increase MDE (test bolder changes) or target a higher-traffic page |
| Test duration exceeds 6 weeks | Insufficient daily traffic | Consider qualitative methods, test bigger changes, or combine traffic from multiple pages |
| p-value hovers around 0.05 | Borderline significance | Do not stop early; run to planned sample size or extend 20% |
| Results significant but lift is tiny (<1%) | Overpowered test | Check practical significance alongside statistical significance |
| Segment results contradict overall | Simpson's paradox | Investigate segment composition; report both overall and segment results |
| Variant performs differently on mobile vs desktop | Device-specific UX issues | Design device-specific variants; increase per-segment sample size |
| Calculator produces negative CI | Very small samples or extreme rates | Ensure sufficient sample size; check data integrity |
| Criterion | Target | How to Measure |
|---|---|---|
| Tests reach planned sample size | 100% of tests | Compare actual vs planned sample at conclusion |
| False positive rate | <5% | Track post-implementation lift vs test prediction |
| Test velocity | 2+ tests per team per month | Count experiments documented per sprint |
| Documentation completeness | 100% of tests documented | Audit experiment records quarterly |
| Average test duration | <4 weeks | Measure start-to-conclusion calendar days |
| Decision quality | >80% of shipped variants hold gains at 90 days | Post-ship metric tracking |
In scope:
Out of scope:
| Tool / Platform | Integration Method | Use Case |
|---|---|---|
| PostHog / Amplitude | JSON export from experiment_analyzer | Feed results into product analytics |
| Jira / Linear | experiment_planner JSON output | Create experiment tickets with metadata |
| Google Sheets | CSV export from experiment_analyzer | Share results with non-technical stakeholders |
| LaunchDarkly / Unleash | experiment_planner checklist | Pre-launch validation before feature flag rollout |
| Slack / Notion | Copy human-readable output | Async experiment status updates |
| CI/CD pipelines | --json flag on all scripts | Automated experiment health checks |
© borghei, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 7 other files (scripts) in product-team/ab-test-setup of borghei/Claude-Skills.
Open the folder on GitHubat commit c9a1487
Ab Test Setup next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Ab Test Setup this skillborghei/Claude-Skills | 874 | — | ~5.3k | Automated safety check: Pass | MIT | |
| Ad Test Designeraaron-he-zhu/aaron-marketing-skills | 2.9k | 2 repos | ~2.8k | Automated safety check: Pass | Apache-2.0 | |
| Ab Test Analyzeririnabuht12-oss/marketing-skills | 3.8k | — | ~1.4k | Automated safety check: Pass | None | |
| Define Hypothesisproduct-on-purpose/pm-skills | 713 | — | ~966 | Automated safety check: Pass | Apache-2.0 | |
| A B Test DesignOwl-Listener/designer-skills | 2.9k | 1 repos | ~472 | Automated safety check: Pass | MIT | |
| Ab Test Planindranilbanerjee/digital-marketing-pro | 854 | 1 repos | ~2k | Automated safety check: Pass | MIT |
aaron-he-zhu/aaron-marketing-skills
A skill your agent uses when the user asks to "design an A/B test", "set up a creative/landing test", "run an incrementality test", or "is this result statistically and practically material?"…
irinabuht12-oss/marketing-skills
Statistical significance calculator for A/B test results with sample size requirements, segment breakdowns, and hypothesis generation.
product-on-purpose/pm-skills
Defines a testable hypothesis with clear success metrics and a validation approach.
Owl-Listener/designer-skills
Design an A/B experiment — hypothesis, variants, primary metric, and sample size.
indranilbanerjee/digital-marketing-pro
Design a statistically rigorous A/B or multivariate test plan — If/Then/Because hypothesis, control and variant specs, required sample size per variant (absolute vs relative MDE via…
ericrisco/rsc-harness
A skill your agent uses when designing or analyzing a controlled experiment — falsifiable hypothesis, sample size from an MDE, reading significance/CI/power, CUPED, or rescuing tests that won't go…
borghei/Claude-Skills
Run delivery when AI coding and ops agents take tickets. An agent skill from borghei/Claude-Skills.
borghei/Claude-Skills
Check AI-generated marketing content and reviews for required disclosures under the EU AI Act, FTC rules and platform AI-label policies.
borghei/Claude-Skills
Idea to AI-generated prototype to customer validation to engineering handoff.
borghei/Claude-Skills
Analytics engineering across data modeling, dbt, transformation, and semantic layers.
borghei/Claude-Skills
Ansoff Matrix — 4-quadrant framework for growth options: market penetration, market/product development, and diversification.
borghei/Claude-Skills
OKR brainstorming and validation using the Radical Focus framework — outcome objectives, measurable key results, counter-metrics.
Categories
Design and run statistically rigorous A/B tests and experiments. Ab Test Setup is an agent skill from borghei/Claude-Skills. Design and run statistically rigorous A/B tests and experiments.
Ab Test Setup fits situations like: planning experiments; calculating sample sizes; designing test variants; selecting metrics.
Run `npx skills add borghei/Claude-Skills --skill ab-test-setup -a claude-code`. Or copy the skill folder (product-team/ab-test-setup in borghei/Claude-Skills) into .claude/skills/ab-test-setup in your project. Claude Code loads it when a task matches its description.
Run `npx skills add borghei/Claude-Skills --skill ab-test-setup -a codex`. Or copy the skill folder (product-team/ab-test-setup in borghei/Claude-Skills) into .agents/skills/ab-test-setup in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add borghei/Claude-Skills --skill ab-test-setup -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ab-test-setup, .gemini/skills/ab-test-setup, .github/skills/ab-test-setup and .opencode/skills/ab-test-setup in your project.
Going by SKILL.md and its folder, Ab Test Setup needs Python for the scripts in its folder and the command-line tools its instructions call (python). Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Ab Test Setup is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 5.3k tokens (SKILL.md is roughly 21k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Ab Test Setup: Ad Test Designer (aaron-he-zhu/aaron-marketing-skills, 2.9k stars), Ab Test Analyzer (irinabuht12-oss/marketing-skills, 3.8k stars), Define Hypothesis (product-on-purpose/pm-skills, 713 stars) and A B Test Design (Owl-Listener/designer-skills, 2.9k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
borghei (a GitHub user) maintains it in borghei/Claude-Skills, which has 874 GitHub stars. The repository holds 364 skills in this directory. The repository was last updated on October 7, 2026.
Source: borghei/Claude-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.