Agent skill

Ab Testing Specialist

by FerroxLabs in FerroxLabs/wayland

End-to-end guide for designing, running, and analyzing A/B tests including experiment design, statistical significance, sample size calculation, common pitfalls, and advanced testing patterns.

Apache-2.0Auto-check passedMarketing & SEO

Install Ab Testing Specialist

skills CLI
$ npx skills add FerroxLabs/wayland --skill ab-testing-specialist -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install FerroxLabs/wayland ab-testing-specialist --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/FerroxLabs/wayland.git skills-src && mkdir -p .claude/skills && cp -r skills-src/src/process/resources/skills-library/bodies/skills/data-analysis/ab-testing-specialist .claude/skills/ab-testing-specialist && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
ab-testing-specialist
GitHub stars
608
Token cost
~3.7k tokens
SKILL.md length
501 words
Files
1
Skills in repo
22
Repo updated
First seen
Licence
Apache-2.0

At a glance

End-to-end guide for designing, running, and analyzing A/B tests including experiment design, statistical significance, sample size calculation, common pitfalls, and advanced testing patterns.

  • Works in 6 steps: Peeking Problem → Multiple Testing → Simpson's Paradox → …
  • The user asks about ab testing specialist
  • SKILL.md covers When to Use, Experiment Design Framework, Sample Size Calculation and Running the Analysis, plus 9 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Ab Testing Specialist is an agent skill from FerroxLabs/wayland. End-to-end guide for designing, running, and analyzing A/B tests including experiment design, statistical significance, sample size calculation, common pitfalls, and advanced testing patterns. Use when the user asks about ab testing specialist, related techniques, best practices, or needs guidance in this domain. Do NOT use when the request is outside the scope of ab testing specialist or requires a different specialized skill.

Its SKILL.md is about 3.7k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Marketing & SEO, covering A/B testing and Experimental design. The repository describes itself as: Wayland - The AI Agent That Perceives. Reasons. Acts. Evolves. The licence is Apache-2.0.

When your agent uses it

  • The user asks about ab testing specialist
  • Related techniques
  • Needs guidance in this domain
  • The request is outside the scope of ab testing specialist

Example prompts

  • “/ab-testing-specialist”

Requirements

  • Python 3

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Peeking Problem
  2. Multiple Testing
  3. Simpson's Paradox
  4. Novelty and Primacy Effects
  5. Network Effects and Interference
  6. Low Power

What it can do on your machine

Read from SKILL.md and the folder at commit 4c030c7. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python, markdown and template).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Ab Testing Specialist loads about 3.7k tokens when it runs. Until then it costs about 113 tokens; SKILL.md has 501 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~113
When it runs · the whole SKILL.md, loaded when a task matches
~3.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from FerroxLabs/wayland at commit 4c030c7, republished under its Apache-2.0 licence (© FerroxLabs). 501 words, ~3,748 tokens.

Download SKILL.mdSave it as .claude/skills/ab-testing-specialist/SKILL.md (or your agent's skills folder).
name
ab-testing-specialist
description
End-to-end guide for designing, running, and analyzing A/B tests including experiment design, statistical significance, sample size calculation, common pitfalls, and advanced testing patterns. Use when the user asks about ab testing specialist, related techniques, best practices, or needs guidance in this domain. Do NOT use when the request is outside the scope of ab testing specialist or requires a different specialized skill.
license
Apache-2.0
metadata.author
foundry-skills
metadata.version
1.0.0
metadata.tags
data-science statistics checklist template quick-reference python testing analysis
metadata.category
data-analysis
metadata.subcategory
statistics-modeling
metadata.disclaimer
none
metadata.difficulty
intermediate

A/B Testing Specialist

You are an expert in online experimentation and A/B testing who designs rigorous experiments, avoids common statistical traps, and translates test results into confident product decisions.

When to Use

Use this skill when:

  • User asks about ab testing specialist techniques or best practices
  • User needs guidance on ab testing specialist concepts
  • User wants to implement or improve their approach to ab testing specialist

Do NOT use when:

  • The request falls outside the scope of ab testing specialist
  • User needs a different specialized skill for their specific situation
  • The topic requires professional consultation beyond general guidance

Experiment Design Framework

Pre-Experiment Checklist
  1. Define the hypothesis - "Changing X will improve Y by Z%"
  2. Select primary metric - One clear success metric
  3. Define guard-rail metrics - Metrics that must not degrade
  4. Calculate sample size - Based on MDE, power, and significance level
  5. Determine test duration - Account for weekly cycles (minimum 1-2 weeks)
  6. Define segments - Mobile/desktop, new/returning, geography
  7. Document exclusion criteria - Bots, internal users, edge cases
  8. Get stakeholder alignment - Agree on decision criteria before launching
Hypothesis Template
Population:  [Who are we testing?]
Treatment:   [What change are we making?]
Metric:      [What are we measuring?]
Direction:   [Do we expect increase or decrease?]
Magnitude:   [What is the minimum detectable effect?]
Timeline:    [How long will we run the test?]

Example:
Population:  All logged-in users on the checkout page
Treatment:   Single-page checkout vs. current multi-step checkout
Metric:      Checkout completion rate (primary), revenue per visitor (secondary)
Direction:   Increase
Magnitude:   2 percentage points (from 35% to 37%)
Timeline:    14 days minimum

Sample Size Calculation

For Conversion Rate Tests
python
from statsmodels.stats.power import NormalIndPower
import numpy as np

def calculate_sample_size(
    baseline_rate: float,
    minimum_detectable_effect: float,  # Absolute difference
    alpha: float = 0.05,
    power: float = 0.80,
    two_sided: bool = True,
):
    """Calculate required sample size per group for a proportion test."""
    effect_size = minimum_detectable_effect / np.sqrt(
        baseline_rate * (1 - baseline_rate)
    )
    analysis = NormalIndPower()
    n = analysis.solve_power(
        effect_size=effect_size,
        alpha=alpha,
        power=power,
        alternative='two-sided' if two_sided else 'larger',
    )
    return int(np.ceil(n))

# Example: Detect 2pp lift from 10% baseline
n = calculate_sample_size(baseline_rate=0.10, minimum_detectable_effect=0.02)
print(f"Sample size per group: {n:,}")
# With 100k visitors/day, test duration = 2 * n / 100_000
For Revenue / Continuous Metrics
python
from statsmodels.stats.power import TTestIndPower

def sample_size_continuous(
    baseline_mean: float,
    baseline_std: float,
    minimum_detectable_effect: float,  # Absolute difference in means
    alpha: float = 0.05,
    power: float = 0.80,
):
    """Sample size for continuous metric (e.g., revenue per user)."""
    cohens_d = minimum_detectable_effect / baseline_std
    analysis = TTestIndPower()
    n = analysis.solve_power(effect_size=cohens_d, alpha=alpha, power=power)
    return int(np.ceil(n))

# Detect $2 lift in average order value (mean=$50, std=$30)
n = sample_size_continuous(50, 30, 2.0)
print(f"Sample size per group: {n:,}")
Sample Size Quick Reference
Baseline RateMDERequired n (per group)Power
5%0.5pp30,42480%
5%1.0pp7,72480%
10%1.0pp14,31480%
10%2.0pp3,62380%
20%2.0pp6,28080%
20%5.0pp1,03080%
50%5.0pp1,57180%

Running the Analysis

Frequentist Approach
python
from scipy import stats
import numpy as np

def analyze_ab_test(
    control_visitors: int,
    control_conversions: int,
    treatment_visitors: int,
    treatment_conversions: int,
    alpha: float = 0.05,
):
    """Complete frequentist analysis of an A/B test."""
    # Conversion rates
    p_control = control_conversions / control_visitors
    p_treatment = treatment_conversions / treatment_visitors
    lift = (p_treatment - p_control) / p_control

    # Pooled proportion for z-test
    p_pool = (control_conversions + treatment_conversions) / (
        control_visitors + treatment_visitors
    )
    se = np.sqrt(p_pool * (1 - p_pool) * (1/control_visitors + 1/treatment_visitors))
    z_stat = (p_treatment - p_control) / se
    p_value = 2 * (1 - stats.norm.cdf(abs(z_stat)))

    # Confidence interval for the difference
    se_diff = np.sqrt(
        p_control * (1 - p_control) / control_visitors +
        p_treatment * (1 - p_treatment) / treatment_visitors
    )
    z_crit = stats.norm.ppf(1 - alpha / 2)
    ci_low = (p_treatment - p_control) - z_crit * se_diff
    ci_high = (p_treatment - p_control) + z_crit * se_diff

    return {
        'control_rate': p_control,
        'treatment_rate': p_treatment,
        'absolute_lift': p_treatment - p_control,
        'relative_lift': lift,
        'z_statistic': z_stat,
        'p_value': p_value,
        'significant': p_value < alpha,
        'ci_low': ci_low,
        'ci_high': ci_high,
    }

results = analyze_ab_test(
    control_visitors=50000,
    control_conversions=5000,
    treatment_visitors=50000,
    treatment_conversions=5400,
)
Bayesian Approach
python
from scipy import stats
import numpy as np

def bayesian_ab_test(
    control_conversions: int,
    control_visitors: int,
    treatment_conversions: int,
    treatment_visitors: int,
    n_simulations: int = 100_000,
    prior_alpha: float = 1,
    prior_beta: float = 1,
):
    """Bayesian analysis using Beta-Binomial model."""
    # Posterior distributions (Beta)
    control_posterior = stats.beta(
        prior_alpha + control_conversions,
        prior_beta + control_visitors - control_conversions,
    )
    treatment_posterior = stats.beta(
        prior_alpha + treatment_conversions,
        prior_beta + treatment_visitors - treatment_conversions,
    )

    # Monte Carlo simulation
    control_samples = control_posterior.rvs(n_simulations)
    treatment_samples = treatment_posterior.rvs(n_simulations)

    # Probability that treatment is better
    prob_treatment_better = np.mean(treatment_samples > control_samples)

    # Expected lift distribution
    lift_samples = (treatment_samples - control_samples) / control_samples
    expected_lift = np.mean(lift_samples)
    lift_ci = np.percentile(lift_samples, [2.5, 97.5])

    # Expected loss (risk of choosing treatment if it is worse)
    loss_if_treatment = np.mean(np.maximum(control_samples - treatment_samples, 0))
    loss_if_control = np.mean(np.maximum(treatment_samples - control_samples, 0))

    return {
        'prob_treatment_better': prob_treatment_better,
        'expected_lift': expected_lift,
        'lift_ci_95': lift_ci,
        'expected_loss_treatment': loss_if_treatment,
        'expected_loss_control': loss_if_control,
    }

Sequential Testing

Avoiding Peeking Problems
python
def sequential_test_boundary(n_looks: int, alpha: float = 0.05):
    """Calculate adjusted significance thresholds for sequential testing."""
    # O'Brien-Fleming spending function
    from scipy.stats import norm
    import numpy as np

    info_fractions = np.linspace(1/n_looks, 1.0, n_looks)
    boundaries = []

    for t in info_fractions:
        # O'Brien-Fleming boundary
        z_boundary = norm.ppf(1 - alpha / 2) / np.sqrt(t)
        p_boundary = 2 * (1 - norm.cdf(z_boundary))
        boundaries.append({
            'look': int(t * n_looks),
            'info_fraction': t,
            'z_boundary': z_boundary,
            'p_threshold': p_boundary,
        })

    return pd.DataFrame(boundaries)

# Plan 5 interim analyses
boundaries = sequential_test_boundary(n_looks=5)
print(boundaries)
# Early looks require very strong evidence; final look is near alpha=0.05

Common Pitfalls

1. Peeking Problem
Problem: Checking results daily and stopping when significant
Impact:  Inflated false positive rate (up to 30% instead of 5%)
Fix:     Pre-commit to sample size, or use sequential testing methods
2. Multiple Testing
Problem: Testing 20 metrics and highlighting the one that is significant
Impact:  1 - (1 - 0.05)^20 = 64% chance of at least one false positive
Fix:     Designate one primary metric; apply Bonferroni or FDR correction
3. Simpson's Paradox
Problem: Overall result differs from every segment's result
Example: Treatment wins overall but loses in mobile AND desktop
         (because treatment got more high-converting desktop traffic)
Fix:     Check results across key segments; use stratified analysis
4. Novelty and Primacy Effects
Problem: New UI gets more clicks initially due to curiosity
Impact:  Overstated lift that decays over time
Fix:     Run test for 2+ weeks; analyze new-user cohort separately
         from those who switched mid-experiment
5. Network Effects and Interference
Problem: Control users are affected by treatment users (e.g., social features)
Impact:  Understated or biased treatment effect
Fix:     Cluster randomization (randomize by region, team, or network cluster)
6. Low Power
Problem: Test cannot detect realistic effect sizes
Impact:  Many "no result" tests that waste time
Fix:     Calculate sample size beforehand; accept larger MDE or run longer

Segmentation Analysis

python
def segment_analysis(df, metric_col, treatment_col, segment_col, alpha=0.05):
    """Analyze A/B test results across segments."""
    results = []

    for segment in df[segment_col].unique():
        seg_data = df[df[segment_col] == segment]
        control = seg_data[seg_data[treatment_col] == 'control'][metric_col]
        treatment = seg_data[seg_data[treatment_col] == 'treatment'][metric_col]

        t_stat, p_val = stats.ttest_ind(control, treatment)
        lift = treatment.mean() - control.mean()

        results.append({
            'segment': segment,
            'n_control': len(control),
            'n_treatment': len(treatment),
            'control_mean': control.mean(),
            'treatment_mean': treatment.mean(),
            'absolute_lift': lift,
            'relative_lift': lift / control.mean(),
            'p_value': p_val,
            'significant': p_val < alpha,
        })

    return pd.DataFrame(results).sort_values('relative_lift', ascending=False)

Multi-Armed Bandit Alternative

python
def thompson_sampling_step(arms_data):
    """One step of Thompson Sampling for multi-armed bandit."""
    samples = {}
    for arm_name, data in arms_data.items():
        alpha = 1 + data['successes']
        beta = 1 + data['failures']
        samples[arm_name] = np.random.beta(alpha, beta)

    chosen_arm = max(samples, key=samples.get)
    return chosen_arm, samples

# When to use Bandits vs. A/B Tests
# Bandits: Optimizing during the test (minimize regret)
# A/B:     Need clean causal measurement (maximize learning)

Decision Framework

After the Test
ResultPowered?Effect Meaningful?Decision
SignificantYesYesShip it
SignificantYesNoConsider cost to implement
Not significantYesN/ANo meaningful effect exists
Not significantNoN/ARun longer or accept larger MDE
Show full SKILL.md (189 more words)Show less
Reporting Template
markdown
## Experiment: [Name]
**Hypothesis:** [What we expected]
**Duration:** [Start] to [End] ([X] days)
**Traffic:** [N control] / [N treatment]

### Primary Metric: [Metric Name]
| Group | Value | 95% CI |
|-------|-------|--------|
| Control | X.XX% | [X.XX%, X.XX%] |
| Treatment | X.XX% | [X.XX%, X.XX%] |
| **Lift** | **+X.XX%** | **[X.XX%, X.XX%]** |

**P-value:** X.XXXX | **Significant:** Yes/No
**Power:** XX% | **Effect size:** X.XX

### Guard-rail Metrics
| Metric | Control | Treatment | Change | Status |
|--------|---------|-----------|--------|--------|
| Revenue/user | $X.XX | $X.XX | +X.X% | OK |
| Page load | X.Xs | X.Xs | -X.X% | OK |

### Recommendation
[Ship / Iterate / Kill] - [Reasoning]

Process

  1. Gather information. Ask the user clarifying questions to understand their specific situation, goals, and constraints
  2. Analyze context. Review the information provided and identify key factors relevant to ab testing specialist
  3. Develop recommendations. Apply domain expertise to create actionable guidance tailored to the user's needs
  4. Present structured output. Deliver findings in the output format below with clear next steps
  5. Address follow-ups. Answer additional questions and refine recommendations based on feedback

Output Format

template
## Ab Testing Specialist Analysis

### Assessment
[Key findings and observations]

### Recommendations
1. [Primary recommendation]
2. [Secondary recommendation]
3. [Additional suggestions]

### Action Items
- [ ] [First action step]
- [ ] [Second action step]
- [ ] [Follow-up task]

Edge Cases

  • Incomplete information: Ask clarifying questions before proceeding with recommendations
  • Conflicting requirements: Prioritize the most critical constraint and note trade-offs
  • Out of scope requests: Redirect to appropriate specialized skill or professional resource
  • Beginner vs advanced: Adjust depth and terminology based on user's experience level

Example

Input: "Help me with ab testing specialist for my current situation"

Output:

Based on your situation, here is a structured approach to ab testing specialist:

  1. Assessment: Evaluate your current state and identify key areas for improvement
  2. Strategy: Develop a targeted plan based on best practices
  3. Implementation: Execute the plan with specific, measurable steps
  4. Review: Monitor progress and adjust as needed

© FerroxLabs, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in src/process/resources/skills-library/bodies/skills/data-analysis/ab-testing-specialist of FerroxLabs/wayland.

Open the folder on GitHubat commit 4c030c7

Compare with similar skills

Ab Testing Specialist next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Ab Testing Specialist compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Ab Testing Specialist this skillFerroxLabs/wayland608—~3.7kAutomated safety check: PassApache-2.0
Ad Test Designeraaron-he-zhu/aaron-marketing-skills2.9k2 repos~2.8kAutomated safety check: PassApache-2.0
Ab Test Analyzeririnabuht12-oss/marketing-skills3.9k—~1.4kAutomated safety check: PassNone
Define Hypothesisproduct-on-purpose/pm-skills715—~966Automated safety check: PassApache-2.0
A B Test DesignOwl-Listener/designer-skills2.9k1 repos~472Automated safety check: PassMIT
Ab Test Planindranilbanerjee/digital-marketing-pro8551 repos~2kAutomated safety check: PassMIT

Similar skills

  • Ad Test Designer

    aaron-he-zhu/aaron-marketing-skills

    A skill your agent uses when the user asks to "design an A/B test", "set up a creative/landing test", "run an incrementality test", or "is this result statistically and practically material?"…

    2.9k GitHub starsUsed in 2 repos~2.8k tokens
    Marketing & SEOAuto-check passed
  • Ab Test Analyzer

    irinabuht12-oss/marketing-skills

    Statistical significance calculator for A/B test results with sample size requirements, segment breakdowns, and hypothesis generation.

    3.9k GitHub stars~1.4k tokensUpdated 15 days ago
    Marketing & SEOAuto-check passed
  • Define Hypothesis

    product-on-purpose/pm-skills

    Defines a testable hypothesis with clear success metrics and a validation approach.

    715 GitHub stars~966 tokensUpdated 4 days ago
    Marketing & SEOAuto-check passed
  • A B Test Design

    Owl-Listener/designer-skills

    Design an A/B experiment — hypothesis, variants, primary metric, and sample size.

    2.9k GitHub starsUsed in 1 repo~472 tokens
    Marketing & SEOAuto-check passed
  • Ab Test Plan

    indranilbanerjee/digital-marketing-pro

    Design a statistically rigorous A/B or multivariate test plan — If/Then/Because hypothesis, control and variant specs, required sample size per variant (absolute vs relative MDE via…

    855 GitHub starsUsed in 1 repo~2k tokens
    Marketing & SEOAuto-check passed
  • Ab Testing

    ericrisco/rsc-harness

    A skill your agent uses when designing or analyzing a controlled experiment — falsifiable hypothesis, sample size from an MDE, reading significance/CI/power, CUPED, or rescuing tests that won't go…

    167 GitHub stars~2.4k tokensUpdated yesterday
    Marketing & SEOAuto-check passed

More from FerroxLabs/wayland

All 22 skills in this repo
  • Star Office Helper

    FerroxLabs/wayland

    Install, start, connect, and troubleshoot visualization companion projects for Aion/OpenClaw, with Star-Office-UI as the default recommendation.

    608 GitHub stars~2.2k tokensUpdated 2 days ago
    Auto-check: notes
  • Openclaw Setup

    FerroxLabs/wayland

    OpenClaw usage expert: Helps you install, deploy, configure, and use OpenClaw personal AI assistant.

    608 GitHub stars~1.9k tokensUpdated 2 days ago
    Auto-check passed
  • Tvcontrol Setup

    FerroxLabs/wayland

    Set up TVControl end to end: install the connector, start TradingView Desktop with its control port open, load a watchlist export, add the indicators they use, and leave a working chart.

    608 GitHub stars~5.7k tokensUpdated 2 days ago
    Auto-check passed
  • Academic Writer

    FerroxLabs/wayland

    Complete academic writing guide covering thesis and dissertation structure, journal article format using IMRaD, literature review methodology, citation management, the peer review process, and…

    608 GitHub stars~4.5k tokensUpdated 2 days ago
    Auto-check passed
  • Accessibility Auditor

    FerroxLabs/wayland

    Web accessibility expertise covering WCAG 2.2 conformance, audit methodology, ARIA patterns, keyboard navigation, screen reader testing, focus management, form accessibility, and automated vs manual…

    608 GitHub stars~4.1k tokensUpdated 2 days ago
    Auto-check passed
  • Organizational accessibility culture building expertise covering champion network design, accessibility training curricula, audit cadence and methodology, KPI definition and tracking, executive…

    608 GitHub stars~3.9k tokensUpdated 2 days ago
    Auto-check passed

Categories

Questions about Ab Testing Specialist

What does Ab Testing Specialist do?

End-to-end guide for designing, running, and analyzing A/B tests including experiment design, statistical significance, sample size calculation, common pitfalls, and advanced testing patterns. Ab Testing Specialist is an agent skill from FerroxLabs/wayland. End-to-end guide for designing, running, and analyzing A/B tests including experiment design, statistical significance, sample size calculation, common pitfalls, and advanced testing patterns.

When should I use Ab Testing Specialist?

Ab Testing Specialist fits situations like: the user asks about ab testing specialist; related techniques; needs guidance in this domain; the request is outside the scope of ab testing specialist.

How do I install Ab Testing Specialist in Claude Code?

Run `npx skills add FerroxLabs/wayland --skill ab-testing-specialist -a claude-code`. Or copy the skill folder (src/process/resources/skills-library/bodies/skills/data-analysis/ab-testing-specialist in FerroxLabs/wayland) into .claude/skills/ab-testing-specialist in your project. Claude Code loads it when a task matches its description.

How do I install Ab Testing Specialist in Codex?

Run `npx skills add FerroxLabs/wayland --skill ab-testing-specialist -a codex`. Or copy the skill folder (src/process/resources/skills-library/bodies/skills/data-analysis/ab-testing-specialist in FerroxLabs/wayland) into .agents/skills/ab-testing-specialist in your project. Codex loads it when a task matches its description.

Can I use Ab Testing Specialist in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add FerroxLabs/wayland --skill ab-testing-specialist -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ab-testing-specialist, .gemini/skills/ab-testing-specialist, .github/skills/ab-testing-specialist and .opencode/skills/ab-testing-specialist in your project.

What does Ab Testing Specialist need to run?

SKILL.md names no scripts, command-line tools or credentials: Ab Testing Specialist is instructions for the agent only. Our summary lists: Python 3.

Does Ab Testing Specialist access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Ab Testing Specialist safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Ab Testing Specialist use?

Ab Testing Specialist is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Ab Testing Specialist use?

About 3.7k tokens (SKILL.md is roughly 15k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Ab Testing Specialist?

Skills that share tags, products or a category with Ab Testing Specialist: Ad Test Designer (aaron-he-zhu/aaron-marketing-skills, 2.9k stars), Ab Test Analyzer (irinabuht12-oss/marketing-skills, 3.9k stars), Define Hypothesis (product-on-purpose/pm-skills, 715 stars) and A B Test Design (Owl-Listener/designer-skills, 2.9k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Ab Testing Specialist?

FerroxLabs (a GitHub user) maintains it in FerroxLabs/wayland, which has 608 GitHub stars. The repository holds 22 skills in this directory. The repository was last updated on October 6, 2026.

Source: FerroxLabs/wayland on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.