Ad Test Designer
aaron-he-zhu/aaron-marketing-skills
A skill your agent uses when the user asks to "design an A/B test", "set up a creative/landing test", "run an incrementality test", or "is this result statistically and practically material?"…
End-to-end guide for designing, running, and analyzing A/B tests including experiment design, statistical significance, sample size calculation, common pitfalls, and advanced testing patterns.
$ npx skills add FerroxLabs/wayland --skill ab-testing-specialist -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install FerroxLabs/wayland ab-testing-specialist --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/FerroxLabs/wayland.git skills-src && mkdir -p .claude/skills && cp -r skills-src/src/process/resources/skills-library/bodies/skills/data-analysis/ab-testing-specialist .claude/skills/ab-testing-specialist && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "ab-testing-specialist" agent skill from https://github.com/FerroxLabs/wayland/tree/main/src/process/resources/skills-library/bodies/skills/data-analysis/ab-testing-specialist into .claude/skills/ab-testing-specialist/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ab-testing-specialist", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/FerroxLabs/wayland/tree/main/src/process/resources/skills-library/bodies/skills/data-analysis/ab-testing-specialistType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add FerroxLabs/wayland --skill ab-testing-specialist -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install FerroxLabs/wayland ab-testing-specialist --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/FerroxLabs/wayland.git skills-src && mkdir -p .agents/skills && cp -r skills-src/src/process/resources/skills-library/bodies/skills/data-analysis/ab-testing-specialist .agents/skills/ab-testing-specialist && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "ab-testing-specialist" agent skill from https://github.com/FerroxLabs/wayland/tree/main/src/process/resources/skills-library/bodies/skills/data-analysis/ab-testing-specialist into .agents/skills/ab-testing-specialist/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ab-testing-specialist", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add FerroxLabs/wayland --skill ab-testing-specialist -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install FerroxLabs/wayland ab-testing-specialist --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/FerroxLabs/wayland.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/src/process/resources/skills-library/bodies/skills/data-analysis/ab-testing-specialist .cursor/skills/ab-testing-specialist && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "ab-testing-specialist" agent skill from https://github.com/FerroxLabs/wayland/tree/main/src/process/resources/skills-library/bodies/skills/data-analysis/ab-testing-specialist into .cursor/skills/ab-testing-specialist/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ab-testing-specialist", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/FerroxLabs/wayland.git --path src/process/resources/skills-library/bodies/skills/data-analysis/ab-testing-specialist--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add FerroxLabs/wayland --skill ab-testing-specialist -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install FerroxLabs/wayland ab-testing-specialist --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/FerroxLabs/wayland.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/src/process/resources/skills-library/bodies/skills/data-analysis/ab-testing-specialist .gemini/skills/ab-testing-specialist && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "ab-testing-specialist" agent skill from https://github.com/FerroxLabs/wayland/tree/main/src/process/resources/skills-library/bodies/skills/data-analysis/ab-testing-specialist into .gemini/skills/ab-testing-specialist/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ab-testing-specialist", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install FerroxLabs/wayland ab-testing-specialistInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add FerroxLabs/wayland --skill ab-testing-specialist -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/FerroxLabs/wayland.git skills-src && mkdir -p .github/skills && cp -r skills-src/src/process/resources/skills-library/bodies/skills/data-analysis/ab-testing-specialist .github/skills/ab-testing-specialist && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "ab-testing-specialist" agent skill from https://github.com/FerroxLabs/wayland/tree/main/src/process/resources/skills-library/bodies/skills/data-analysis/ab-testing-specialist into .github/skills/ab-testing-specialist/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ab-testing-specialist", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add FerroxLabs/wayland --skill ab-testing-specialist -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install FerroxLabs/wayland ab-testing-specialist --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/FerroxLabs/wayland.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/src/process/resources/skills-library/bodies/skills/data-analysis/ab-testing-specialist .opencode/skills/ab-testing-specialist && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "ab-testing-specialist" agent skill from https://github.com/FerroxLabs/wayland/tree/main/src/process/resources/skills-library/bodies/skills/data-analysis/ab-testing-specialist into .opencode/skills/ab-testing-specialist/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ab-testing-specialist", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
ab-testing-specialistEnd-to-end guide for designing, running, and analyzing A/B tests including experiment design, statistical significance, sample size calculation, common pitfalls, and advanced testing patterns.
Ab Testing Specialist is an agent skill from FerroxLabs/wayland. End-to-end guide for designing, running, and analyzing A/B tests including experiment design, statistical significance, sample size calculation, common pitfalls, and advanced testing patterns. Use when the user asks about ab testing specialist, related techniques, best practices, or needs guidance in this domain. Do NOT use when the request is outside the scope of ab testing specialist or requires a different specialized skill.
Its SKILL.md is about 3.7k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Marketing & SEO, covering A/B testing and Experimental design. The repository describes itself as: Wayland - The AI Agent That Perceives. Reasons. Acts. Evolves. The licence is Apache-2.0.
6 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 4c030c7. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are python, markdown and template).
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Ab Testing Specialist loads about 3.7k tokens when it runs. Until then it costs about 113 tokens; SKILL.md has 501 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from FerroxLabs/wayland at commit 4c030c7, republished under its Apache-2.0 licence (© FerroxLabs). 501 words, ~3,748 tokens.
.claude/skills/ab-testing-specialist/SKILL.md (or your agent's skills folder).You are an expert in online experimentation and A/B testing who designs rigorous experiments, avoids common statistical traps, and translates test results into confident product decisions.
Use this skill when:
Do NOT use when:
Population: [Who are we testing?]
Treatment: [What change are we making?]
Metric: [What are we measuring?]
Direction: [Do we expect increase or decrease?]
Magnitude: [What is the minimum detectable effect?]
Timeline: [How long will we run the test?]
Example:
Population: All logged-in users on the checkout page
Treatment: Single-page checkout vs. current multi-step checkout
Metric: Checkout completion rate (primary), revenue per visitor (secondary)
Direction: Increase
Magnitude: 2 percentage points (from 35% to 37%)
Timeline: 14 days minimumfrom statsmodels.stats.power import NormalIndPower
import numpy as np
def calculate_sample_size(
baseline_rate: float,
minimum_detectable_effect: float, # Absolute difference
alpha: float = 0.05,
power: float = 0.80,
two_sided: bool = True,
):
"""Calculate required sample size per group for a proportion test."""
effect_size = minimum_detectable_effect / np.sqrt(
baseline_rate * (1 - baseline_rate)
)
analysis = NormalIndPower()
n = analysis.solve_power(
effect_size=effect_size,
alpha=alpha,
power=power,
alternative='two-sided' if two_sided else 'larger',
)
return int(np.ceil(n))
# Example: Detect 2pp lift from 10% baseline
n = calculate_sample_size(baseline_rate=0.10, minimum_detectable_effect=0.02)
print(f"Sample size per group: {n:,}")
# With 100k visitors/day, test duration = 2 * n / 100_000from statsmodels.stats.power import TTestIndPower
def sample_size_continuous(
baseline_mean: float,
baseline_std: float,
minimum_detectable_effect: float, # Absolute difference in means
alpha: float = 0.05,
power: float = 0.80,
):
"""Sample size for continuous metric (e.g., revenue per user)."""
cohens_d = minimum_detectable_effect / baseline_std
analysis = TTestIndPower()
n = analysis.solve_power(effect_size=cohens_d, alpha=alpha, power=power)
return int(np.ceil(n))
# Detect $2 lift in average order value (mean=$50, std=$30)
n = sample_size_continuous(50, 30, 2.0)
print(f"Sample size per group: {n:,}")| Baseline Rate | MDE | Required n (per group) | Power |
|---|---|---|---|
| 5% | 0.5pp | 30,424 | 80% |
| 5% | 1.0pp | 7,724 | 80% |
| 10% | 1.0pp | 14,314 | 80% |
| 10% | 2.0pp | 3,623 | 80% |
| 20% | 2.0pp | 6,280 | 80% |
| 20% | 5.0pp | 1,030 | 80% |
| 50% | 5.0pp | 1,571 | 80% |
from scipy import stats
import numpy as np
def analyze_ab_test(
control_visitors: int,
control_conversions: int,
treatment_visitors: int,
treatment_conversions: int,
alpha: float = 0.05,
):
"""Complete frequentist analysis of an A/B test."""
# Conversion rates
p_control = control_conversions / control_visitors
p_treatment = treatment_conversions / treatment_visitors
lift = (p_treatment - p_control) / p_control
# Pooled proportion for z-test
p_pool = (control_conversions + treatment_conversions) / (
control_visitors + treatment_visitors
)
se = np.sqrt(p_pool * (1 - p_pool) * (1/control_visitors + 1/treatment_visitors))
z_stat = (p_treatment - p_control) / se
p_value = 2 * (1 - stats.norm.cdf(abs(z_stat)))
# Confidence interval for the difference
se_diff = np.sqrt(
p_control * (1 - p_control) / control_visitors +
p_treatment * (1 - p_treatment) / treatment_visitors
)
z_crit = stats.norm.ppf(1 - alpha / 2)
ci_low = (p_treatment - p_control) - z_crit * se_diff
ci_high = (p_treatment - p_control) + z_crit * se_diff
return {
'control_rate': p_control,
'treatment_rate': p_treatment,
'absolute_lift': p_treatment - p_control,
'relative_lift': lift,
'z_statistic': z_stat,
'p_value': p_value,
'significant': p_value < alpha,
'ci_low': ci_low,
'ci_high': ci_high,
}
results = analyze_ab_test(
control_visitors=50000,
control_conversions=5000,
treatment_visitors=50000,
treatment_conversions=5400,
)from scipy import stats
import numpy as np
def bayesian_ab_test(
control_conversions: int,
control_visitors: int,
treatment_conversions: int,
treatment_visitors: int,
n_simulations: int = 100_000,
prior_alpha: float = 1,
prior_beta: float = 1,
):
"""Bayesian analysis using Beta-Binomial model."""
# Posterior distributions (Beta)
control_posterior = stats.beta(
prior_alpha + control_conversions,
prior_beta + control_visitors - control_conversions,
)
treatment_posterior = stats.beta(
prior_alpha + treatment_conversions,
prior_beta + treatment_visitors - treatment_conversions,
)
# Monte Carlo simulation
control_samples = control_posterior.rvs(n_simulations)
treatment_samples = treatment_posterior.rvs(n_simulations)
# Probability that treatment is better
prob_treatment_better = np.mean(treatment_samples > control_samples)
# Expected lift distribution
lift_samples = (treatment_samples - control_samples) / control_samples
expected_lift = np.mean(lift_samples)
lift_ci = np.percentile(lift_samples, [2.5, 97.5])
# Expected loss (risk of choosing treatment if it is worse)
loss_if_treatment = np.mean(np.maximum(control_samples - treatment_samples, 0))
loss_if_control = np.mean(np.maximum(treatment_samples - control_samples, 0))
return {
'prob_treatment_better': prob_treatment_better,
'expected_lift': expected_lift,
'lift_ci_95': lift_ci,
'expected_loss_treatment': loss_if_treatment,
'expected_loss_control': loss_if_control,
}def sequential_test_boundary(n_looks: int, alpha: float = 0.05):
"""Calculate adjusted significance thresholds for sequential testing."""
# O'Brien-Fleming spending function
from scipy.stats import norm
import numpy as np
info_fractions = np.linspace(1/n_looks, 1.0, n_looks)
boundaries = []
for t in info_fractions:
# O'Brien-Fleming boundary
z_boundary = norm.ppf(1 - alpha / 2) / np.sqrt(t)
p_boundary = 2 * (1 - norm.cdf(z_boundary))
boundaries.append({
'look': int(t * n_looks),
'info_fraction': t,
'z_boundary': z_boundary,
'p_threshold': p_boundary,
})
return pd.DataFrame(boundaries)
# Plan 5 interim analyses
boundaries = sequential_test_boundary(n_looks=5)
print(boundaries)
# Early looks require very strong evidence; final look is near alpha=0.05Problem: Checking results daily and stopping when significant
Impact: Inflated false positive rate (up to 30% instead of 5%)
Fix: Pre-commit to sample size, or use sequential testing methodsProblem: Testing 20 metrics and highlighting the one that is significant
Impact: 1 - (1 - 0.05)^20 = 64% chance of at least one false positive
Fix: Designate one primary metric; apply Bonferroni or FDR correctionProblem: Overall result differs from every segment's result
Example: Treatment wins overall but loses in mobile AND desktop
(because treatment got more high-converting desktop traffic)
Fix: Check results across key segments; use stratified analysisProblem: New UI gets more clicks initially due to curiosity
Impact: Overstated lift that decays over time
Fix: Run test for 2+ weeks; analyze new-user cohort separately
from those who switched mid-experimentProblem: Control users are affected by treatment users (e.g., social features)
Impact: Understated or biased treatment effect
Fix: Cluster randomization (randomize by region, team, or network cluster)Problem: Test cannot detect realistic effect sizes
Impact: Many "no result" tests that waste time
Fix: Calculate sample size beforehand; accept larger MDE or run longerdef segment_analysis(df, metric_col, treatment_col, segment_col, alpha=0.05):
"""Analyze A/B test results across segments."""
results = []
for segment in df[segment_col].unique():
seg_data = df[df[segment_col] == segment]
control = seg_data[seg_data[treatment_col] == 'control'][metric_col]
treatment = seg_data[seg_data[treatment_col] == 'treatment'][metric_col]
t_stat, p_val = stats.ttest_ind(control, treatment)
lift = treatment.mean() - control.mean()
results.append({
'segment': segment,
'n_control': len(control),
'n_treatment': len(treatment),
'control_mean': control.mean(),
'treatment_mean': treatment.mean(),
'absolute_lift': lift,
'relative_lift': lift / control.mean(),
'p_value': p_val,
'significant': p_val < alpha,
})
return pd.DataFrame(results).sort_values('relative_lift', ascending=False)def thompson_sampling_step(arms_data):
"""One step of Thompson Sampling for multi-armed bandit."""
samples = {}
for arm_name, data in arms_data.items():
alpha = 1 + data['successes']
beta = 1 + data['failures']
samples[arm_name] = np.random.beta(alpha, beta)
chosen_arm = max(samples, key=samples.get)
return chosen_arm, samples
# When to use Bandits vs. A/B Tests
# Bandits: Optimizing during the test (minimize regret)
# A/B: Need clean causal measurement (maximize learning)| Result | Powered? | Effect Meaningful? | Decision |
|---|---|---|---|
| Significant | Yes | Yes | Ship it |
| Significant | Yes | No | Consider cost to implement |
| Not significant | Yes | N/A | No meaningful effect exists |
| Not significant | No | N/A | Run longer or accept larger MDE |
## Experiment: [Name]
**Hypothesis:** [What we expected]
**Duration:** [Start] to [End] ([X] days)
**Traffic:** [N control] / [N treatment]
### Primary Metric: [Metric Name]
| Group | Value | 95% CI |
|-------|-------|--------|
| Control | X.XX% | [X.XX%, X.XX%] |
| Treatment | X.XX% | [X.XX%, X.XX%] |
| **Lift** | **+X.XX%** | **[X.XX%, X.XX%]** |
**P-value:** X.XXXX | **Significant:** Yes/No
**Power:** XX% | **Effect size:** X.XX
### Guard-rail Metrics
| Metric | Control | Treatment | Change | Status |
|--------|---------|-----------|--------|--------|
| Revenue/user | $X.XX | $X.XX | +X.X% | OK |
| Page load | X.Xs | X.Xs | -X.X% | OK |
### Recommendation
[Ship / Iterate / Kill] - [Reasoning]## Ab Testing Specialist Analysis
### Assessment
[Key findings and observations]
### Recommendations
1. [Primary recommendation]
2. [Secondary recommendation]
3. [Additional suggestions]
### Action Items
- [ ] [First action step]
- [ ] [Second action step]
- [ ] [Follow-up task]Input: "Help me with ab testing specialist for my current situation"
Output:
Based on your situation, here is a structured approach to ab testing specialist:
© FerroxLabs, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in src/process/resources/skills-library/bodies/skills/data-analysis/ab-testing-specialist of FerroxLabs/wayland.
Open the folder on GitHubat commit 4c030c7
Ab Testing Specialist next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Ab Testing Specialist this skillFerroxLabs/wayland | 608 | — | ~3.7k | Automated safety check: Pass | Apache-2.0 | |
| Ad Test Designeraaron-he-zhu/aaron-marketing-skills | 2.9k | 2 repos | ~2.8k | Automated safety check: Pass | Apache-2.0 | |
| Ab Test Analyzeririnabuht12-oss/marketing-skills | 3.9k | — | ~1.4k | Automated safety check: Pass | None | |
| Define Hypothesisproduct-on-purpose/pm-skills | 715 | — | ~966 | Automated safety check: Pass | Apache-2.0 | |
| A B Test DesignOwl-Listener/designer-skills | 2.9k | 1 repos | ~472 | Automated safety check: Pass | MIT | |
| Ab Test Planindranilbanerjee/digital-marketing-pro | 855 | 1 repos | ~2k | Automated safety check: Pass | MIT |
aaron-he-zhu/aaron-marketing-skills
A skill your agent uses when the user asks to "design an A/B test", "set up a creative/landing test", "run an incrementality test", or "is this result statistically and practically material?"…
irinabuht12-oss/marketing-skills
Statistical significance calculator for A/B test results with sample size requirements, segment breakdowns, and hypothesis generation.
product-on-purpose/pm-skills
Defines a testable hypothesis with clear success metrics and a validation approach.
Owl-Listener/designer-skills
Design an A/B experiment — hypothesis, variants, primary metric, and sample size.
indranilbanerjee/digital-marketing-pro
Design a statistically rigorous A/B or multivariate test plan — If/Then/Because hypothesis, control and variant specs, required sample size per variant (absolute vs relative MDE via…
ericrisco/rsc-harness
A skill your agent uses when designing or analyzing a controlled experiment — falsifiable hypothesis, sample size from an MDE, reading significance/CI/power, CUPED, or rescuing tests that won't go…
FerroxLabs/wayland
Install, start, connect, and troubleshoot visualization companion projects for Aion/OpenClaw, with Star-Office-UI as the default recommendation.
FerroxLabs/wayland
OpenClaw usage expert: Helps you install, deploy, configure, and use OpenClaw personal AI assistant.
FerroxLabs/wayland
Set up TVControl end to end: install the connector, start TradingView Desktop with its control port open, load a watchlist export, add the indicators they use, and leave a working chart.
FerroxLabs/wayland
Complete academic writing guide covering thesis and dissertation structure, journal article format using IMRaD, literature review methodology, citation management, the peer review process, and…
FerroxLabs/wayland
Web accessibility expertise covering WCAG 2.2 conformance, audit methodology, ARIA patterns, keyboard navigation, screen reader testing, focus management, form accessibility, and automated vs manual…
FerroxLabs/wayland
Organizational accessibility culture building expertise covering champion network design, accessibility training curricula, audit cadence and methodology, KPI definition and tracking, executive…
Categories
End-to-end guide for designing, running, and analyzing A/B tests including experiment design, statistical significance, sample size calculation, common pitfalls, and advanced testing patterns. Ab Testing Specialist is an agent skill from FerroxLabs/wayland. End-to-end guide for designing, running, and analyzing A/B tests including experiment design, statistical significance, sample size calculation, common pitfalls, and advanced testing patterns.
Ab Testing Specialist fits situations like: the user asks about ab testing specialist; related techniques; needs guidance in this domain; the request is outside the scope of ab testing specialist.
Run `npx skills add FerroxLabs/wayland --skill ab-testing-specialist -a claude-code`. Or copy the skill folder (src/process/resources/skills-library/bodies/skills/data-analysis/ab-testing-specialist in FerroxLabs/wayland) into .claude/skills/ab-testing-specialist in your project. Claude Code loads it when a task matches its description.
Run `npx skills add FerroxLabs/wayland --skill ab-testing-specialist -a codex`. Or copy the skill folder (src/process/resources/skills-library/bodies/skills/data-analysis/ab-testing-specialist in FerroxLabs/wayland) into .agents/skills/ab-testing-specialist in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add FerroxLabs/wayland --skill ab-testing-specialist -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ab-testing-specialist, .gemini/skills/ab-testing-specialist, .github/skills/ab-testing-specialist and .opencode/skills/ab-testing-specialist in your project.
SKILL.md names no scripts, command-line tools or credentials: Ab Testing Specialist is instructions for the agent only. Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Ab Testing Specialist is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 3.7k tokens (SKILL.md is roughly 15k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Ab Testing Specialist: Ad Test Designer (aaron-he-zhu/aaron-marketing-skills, 2.9k stars), Ab Test Analyzer (irinabuht12-oss/marketing-skills, 3.9k stars), Define Hypothesis (product-on-purpose/pm-skills, 715 stars) and A B Test Design (Owl-Listener/designer-skills, 2.9k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
FerroxLabs (a GitHub user) maintains it in FerroxLabs/wayland, which has 608 GitHub stars. The repository holds 22 skills in this directory. The repository was last updated on October 6, 2026.
Source: FerroxLabs/wayland on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.