Topic · Marketing & SEO

Best A/B testing skills for Claude Code, Codex and other agents.

Skills that design, run and read A/B tests and growth experiments.
skills
267
official
9

A/B testing skills, ranked

Ranked by score. Sort bymost stars,trending,newest,recently updated

A/B testing skills, ranked
#SkillRepositoryStarsUsed inTokensAuto-checkLicenceUpdated
1

When the user wants to set up, improve, or audit analytics tracking and measurement.

Nexus-JPF/note-companion8696 repos~2.2kAutomated safety check: PassMIT4 days ago
2

Pre-pipeline aggregator that scans AI agent cache directories (.claude, .cursor, .antigravity, .openclaw) or any user-specified directory for experimentation logs, extracts insights and numeric…

Ar9av/PaperOrchestra6762 repos~3.5kAutomated safety check: PassUnknown16 days ago
3

When the user wants to plan, design, or implement an A/B test or experiment.

freekmurze/dotfiles1k15 repos~1.8kAutomated safety check: PassNo licence2 days ago
4

Rigor Improve / Rigor Explore run leaf skill for bounded exploratory evidence in deep learning research repositories.

lllllllama/RigorPilot-Skills4972 repos~833Automated safety check: PassMIT14 days ago
5

A skill your agent uses when the user asks to "design an A/B test", "set up a creative/landing test", "run an incrementality test", or "is this result statistically and practically material?"…

aaron-he-zhu/aaron-marketing-skills2.9k2 repos~2.8kAutomated safety check: PassApache-2.0today
6

When the user wants to plan, design, or implement an A/B test or experiment, or build a growth experimentation program.

Cesarjoquin/Marketing-Skills1992 repos~2.8kAutomated safety check: PassMIT23 days ago
7

World-class data science skill for statistical modeling, experimentation, causal inference, and advanced analytics.

Raidriar7170/hermes-skilleval1256 repos~1.4kAutomated safety check: PassMIT11 days ago
8

Writes and improves title tags, meta descriptions, Open Graph and Twitter card tags for click-through, with character counts and A/B test variants.

nowork-studio/notfair-plugin3.9k1 repo~2.7kAutomated safety check: PassMIT6 days ago
9

Statistical significance calculator for A/B test results with sample size requirements, segment breakdowns, and hypothesis generation.

irinabuht12-oss/marketing-skills3.8k—~1.4kAutomated safety check: PassNo licence14 days ago
10

Generate high-CTR YouTube thumbnails using Nano Banana 2 via the Arcads external API.

krusemediallc/arcads-claude-code1.6k—~2.4kAutomated safety check: NotesMIT15 days ago
11

When the user wants to A/B test App Store product page elements to improve conversion rate.

appeeky/aso-skills2.1k—~1.8kAutomated safety check: PassMITyesterday
12

Benchmark and profile the Dynamo frontend (dynamo.frontend HTTP + tokenizer + KV router) against mock workers (dynamo.mocker).

ai-dynamo/dynamo8.2k—~3.5kAutomated safety check: NotesApache-2.0today
13

Translate an existing Remotion (React-based) video composition into a HyperFrames HTML composition.

boraoztunc/skills393—~2.2kAutomated safety check: PassApache-2.01 mo ago
14

Hack, modify, and translate ColecoVision and Super Game Module ROMs using the Gearcoleco emulator MCP server.

drhelius/Gearcoleco141—~3.9kAutomated safety check: PassGPL-3.0yesterday
15

Transform Claude Code into an AI Scientist that orchestrates research workflows using tree-based hypothesis exploration.

sundial-org/skills152—~2.5kAutomated safety check: PassNo licence2 mo ago
16

Defines a testable hypothesis with clear success metrics and a validation approach.

product-on-purpose/pm-skills713—~966Automated safety check: PassApache-2.03 days ago
17

A skill your agent uses when the user has changed a prompt (system prompt, RAG template, agent instruction, etc.) and wants to know whether the candidate is better or worse than the baseline.

agentscope-ai/OpenJudge867—~2.8kAutomated safety check: PassApache-2.026 days ago
18

Ultimate Claude Code skill creator and architect. An agent skill from AgriciDaniel/skill-forge.

AgriciDaniel/skill-forge177—~1.9kAutomated safety check: NotesMIT6 mo ago
19

When the user wants to set up, improve, or audit analytics tracking and measurement.

freekmurze/dotfiles1k12 repos~2kAutomated safety check: PassNo licence2 days ago
20

When the user wants to set up, interpret, or improve their app analytics and tracking.

appeeky/aso-skills2.1k—~1.6kAutomated safety check: PassMITyesterday
21

Rigor Explore compatible skill slug for meaningful and potentially novel deep learning research candidates.

lllllllama/RigorPilot-Skills4971 repo~1.7kAutomated safety check: PassMIT14 days ago
22

Optimize code using KAPSO (Knowledge-Grounded Optimization).

Leeroo-AI/kapso120—~642Automated safety check: PassMITyesterday
23

Validates an experiment's setup, works out lift, p-value and confidence interval from A/B test data, and recommends whether to ship, extend or stop.

phuryn/pm-skills27k—~893Automated safety check: PassMIT23 days ago
24

A skill your agent uses when the user asks to "write the email", "draft subject lines", or "build email creative"; produces the pre-click unit — subject-line variants + preheader, body copy, one…

aaron-he-zhu/aaron-marketing-skills2.9k2 repos~3.1kAutomated safety check: PassApache-2.0today
25

Add full builder API support (@tag, @parse, @split) for a TTIR op.

tenstorrent/tt-mlir311—~1.6kAutomated safety check: PassApache-2.0today
26

Design an A/B experiment — hypothesis, variants, primary metric, and sample size.

Owl-Listener/designer-skills2.9k1 repo~472Automated safety check: PassMIT1 mo ago
27

Design a statistically rigorous A/B or multivariate test plan — If/Then/Because hypothesis, control and variant specs, required sample size per variant (absolute vs relative MDE via…

indranilbanerjee/digital-marketing-pro8541 repo~2kAutomated safety check: PassMIT3 days ago
28

Validate multi-line log stitching behavior for an ama-logs image change.

microsoft/Docker-Provider173—~2.5kAutomated safety check: PassUnknownyesterday
29

The standard way to run a listening test in HOT-Step - a local HTML score sheet next to the renders where Rob plays each track, scores it 1-5 on named criteria, and the page charts the two score…

scragnog/HOT-Step-CPP170—~1.9kAutomated safety check: PassMIT2 days ago
30

PhD-level expertise in data science, statistics, and machine learning.

magnus919/hermes-profiles278—~3.3kAutomated safety check: PassMIT3 mo ago
31

When the user wants to design, test, or improve their app icon to increase tap-through rate and conversions in App Store search and browse.

appeeky/aso-skills2.1k—~1.5kAutomated safety check: PassMITyesterday
32

Benchmark Claude Code skill performance with variance analysis, tracking pass rate, execution time, and token usage across iterations.

AgriciDaniel/skill-forge177—~1.4kAutomated safety check: PassMIT6 mo ago
33

Rigorous A/B test statistical analysis. An agent skill from nimrodfisher/data-analytics-skills.

nimrodfisher/data-analytics-skills465—~708Automated safety check: PassMIT12 days ago
34

Forces heavy internal computation (thinking tokens) before each response.

Lomnus-ai/TokenBurner173—~9.1kAutomated safety check: PassMIT5 mo ago
35

A skill your agent uses when user wants to optimize or tune something measurable through repeated experiments — "make X faster", "improve [metric]", "find best config", "iterate overnight", "run…

krzysztofdudek/ResearcherSkill265—~5.5kAutomated safety check: WarnMIT2 days ago
36

Craft persuasive, goal-driven emails using AI. An agent skill from ILoveDotNet/ilovedotnet.

ILoveDotNet/ilovedotnet155—~3kAutomated safety check: PassCC0-1.0yesterday
37

Interviews you about objective, audience and context, then designs an end-of-article call to action with copy, placement, A/B test plan and accessibility check.

samber/cc-skills228—~3.1kAutomated safety check: PassMIT6 days ago
38

Make browser automation reliable — anchor every click, type, and scroll to stable DOM selectors and semantic roles instead of screen coordinates, so actions survive layout shifts, A/B tests, and…

breakstageaxe61/genspark-claw125—~1kAutomated safety check: PassMIT11 days ago
39

Run a structured 5-day process to prototype, test, and validate product ideas with real users.

wondelai/skills2.4k—~3.8kAutomated safety check: PassMIT27 days ago
40

A skill your agent uses when designing or analyzing a controlled experiment — falsifiable hypothesis, sample size from an MDE, reading significance/CI/power, CUPED, or rescuing tests that won't go…

ericrisco/rsc-harness156—~2.4kAutomated safety check: PassMITyesterday
41

How to size, write and A/B-test the text of a tool — its description and its inputSchema parameter prose — so that cutting it does not cost call quality, and so that a tool that IS getting called…

DitriXNew/EDT-MCP295—~3kAutomated safety check: PassAGPL-3.0today
42

When the user wants to plan, script, produce, or optimize App Store Preview videos or Google Play promo videos — the autoplay videos that show in App Store/Play Store search and product pages.

appeeky/aso-skills2.1k—~1.6kAutomated safety check: PassMITyesterday
43

Edge middleware for EdgeOne Makers — request interception, redirects, rewrites, auth guards, A/B testing, and header injection at the edge (V8 runtime).

TencentEdgeOne/edgeone-makers-tools1.9k1 repo~1.2kAutomated safety check: PassMIT14 days ago
44

Calculate A/B test statistical significance. An agent skill from guia-matthieu/clawfu-skills.

guia-matthieu/clawfu-skills150—~1kAutomated safety check: PassMIT6 days ago
45

Master comprehensive evaluation strategies for LLM applications, from automated metrics to human evaluation and A/B testing.

davila7/claude-code-templates32k12 repos~3.5kAutomated safety check: PassMITtoday
46

Guides development with SAP AI Core and SAP AI Launchpad for enterprise AI/ML workloads on SAP BTP.

secondsky/sap-skills460—~3.3kAutomated safety check: PassGPL-3.02 days ago
47

Run Karpathy-style autoresearch optimization on any content.

ericosiu/ai-marketing-skills3.6k2 repos~2.2kAutomated safety check: PassMIT15 days ago
48

Run an explicitly requested Zephyr control/treatment benchmark on the same pre-normalized sample and compare Finelog stage metrics.

marin-community/marin3.9k—~3.3kAutomated safety check: PassApache-2.0today

Questions, answered from the data.

What is the best A/B testing skill?

Analytics from Nexus-JPF/note-companion ranks first of the 267 A/B testing skills listed here, with the highest score: its repository has 869 GitHub stars, 6 other GitHub owners carry a copy, its SKILL.md loads about 2.2k tokens and it passes the automated safety check with no findings. Next come Agent Research Aggregator and Ab Test Setup.

Which A/B testing skills are official?

9 of the 267 A/B testing skills are official, published by the vendor's own GitHub organization: Multiline Validation, Autoresearch, Content Experimentation Best Practices, Arize Experiment, Flags SDK and 4 more.

How are these skills ranked?

By Skill Navigator score, which combines the GitHub stars of the skill's repository (shared across that repo's skills and discounted for large collections), how many other GitHub owners carry a copy of the skill, and automated SKILL.md quality checks, minus penalties for safety-check warnings and for each further skill from the same repository. Skills that fail the safety check are not listed.