Agent skill

Experimentation

by cbrock84 in cbrock84/headcount

Designs, runs, and reads A/B tests and growth experiments — hypothesis, sample size, duration, and honest interpretation.

MITAuto-check passedMarketing & SEO

Install Experimentation

skills CLI
$ npx skills add cbrock84/headcount --skill experimentation -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install cbrock84/headcount experimentation --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/cbrock84/headcount.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/demand-generation/skills/experimentation .claude/skills/experimentation && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
experimentation
GitHub stars
2k
Token cost
~971 tokens
SKILL.md length
543 words
Files
2 (incl. references)
Skills in repo
175
Repo updated
First seen
Licence
MIT

At a glance

Designs, runs, and reads A/B tests and growth experiments — hypothesis, sample size, duration, and honest interpretation.

  • Tasks that involve A/B testing
  • SKILL.md covers Before running, While running, Reading and Program level, plus 3 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Experimentation is an agent skill from cbrock84/headcount. Designs, runs, and reads A/B tests and growth experiments — hypothesis, sample size, duration, and honest interpretation. Use this to plan a test, judge whether a result is real, build an experimentation program, decide what to test next, or diagnose why tests keep producing inconclusive or non-replicating results.

Its SKILL.md is about 970 tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including reference files (for example `references/sources.md`).

It sits in Marketing & SEO, covering A/B testing. The repository describes itself as: An agent organization structured as a company — 15+ departments, 125+ skills, each independently installable, citing the standards and regulators that settle the question. Runs… The licence is MIT.

When your agent uses it

  • Tasks that involve A/B testing

Example prompts

  • “/experimentation”

What it can do on your machine

Read from SKILL.md and the folder at commit 98d1c17. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Experimentation loads about 971 tokens when it runs, and up to ~1.3k if it reads all its reference files. Until then it costs about 83 tokens; SKILL.md has 543 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~83
When it runs · the whole SKILL.md, loaded when a task matches
~971
With references · SKILL.md plus every file in references/, read only if the agent opens them
~1.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from cbrock84/headcount at commit 98d1c17, republished under its MIT licence (© cbrock84). 543 words, ~971 tokens.

Download SKILL.mdSave it as .claude/skills/experimentation/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
experimentation
description
Designs, runs, and reads A/B tests and growth experiments — hypothesis, sample size, duration, and honest interpretation. Use this to plan a test, judge whether a result is real, build an experimentation program, decide what to test next, or diagnose why tests keep producing inconclusive or non-replicating results.

Experimentation

Most A/B testing programs produce confident conclusions from insufficient data. The discipline is almost entirely in what you do before launch.

Before running

  • Hypothesis with a mechanism. "Moving the pricing table above the fold will raise trial starts, because visitors currently leave before seeing pricing." Not "let's try a green button."
  • One primary metric, chosen in advance. Secondary metrics are context, never the verdict.
  • Sample size calculated in advance, from your baseline rate and the smallest lift that would change a decision. If the required sample is unreachable, do not run the test — decide by judgment and say so.
  • Duration set in advance, covering at least one full weekly cycle, and two if the buying cycle is long.
  • Guardrail metrics that would make you reject a win: refunds, support volume, downstream retention.

While running

Do not look at results and act on them mid-flight. Peeking and stopping at significance is the single most common way to generate false positives, and it is very effective at it.

Check only that the test is running correctly — even split, no broken variant, tracking firing.

Reading

  • At the pre-set duration, not before, and not extended because it is nearly significant. Extending until significance manufactures it.
  • Significance is not size. A statistically significant 0.3% lift may not be worth shipping.
  • Inconclusive is a real result and the most common one. It means the change did not matter enough to detect, which is useful.
  • Check the guardrails before declaring a win.
  • Segment afterward for hypotheses only, never for verdicts. Slice enough ways and something is always significant.

Program level

Test where the traffic and the leverage are. Most sites can only run a handful of adequately powered tests a year — spend them on structural questions, not button colors.

Keep a log of every test: hypothesis, result, decision. Without it, teams re-run the same tests every eighteen months and re-learn the same things.

Show full SKILL.md (226 more words)Show less

Sources

references/sources.md in this skill lists the outside authorities that settle the questions here — what each one is authoritative for, and what you may do with it. Check them before answering on anything they cover, and cite what you used. Most are free to read and not free to reproduce; the use note on each is binding.

Tooling

Client-side and web testing: Optimizely, VWO, AB Tasty, and similar. Warehouse- or product-native: GrowthBook, Statsig, Eppo, PostHog, and similar — these compute against your own event data, which is what you want once the metric definitions matter.

Feature flags are the server-side path to the same thing: LaunchDarkly, Unleash, Split, and similar. Running an experiment behind a flag you already use for release control is cheaper than adding a second system, and technology:release-and-deployment covers the release side of it.

No tool fixes an underpowered test. The platform reports a result either way, which is exactly the risk.

Never

  • Stop a test because it reached significance early. Peeking until it looks conclusive manufactures the result.
  • Run a test that cannot reach adequate sample size in a reasonable window. Ship the change on judgment instead and say so.
  • Change more than one variable and attribute the outcome to the one you liked.
  • Count a flat result as a failure. A well-run test that rules out a plausible idea has bought information.

© cbrock84, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (references) in plugins/demand-generation/skills/experimentation of cbrock84/headcount.

  • SKILL.md
  • references/sources.md

Open the folder on GitHubat commit 98d1c17

Compare with similar skills

Experimentation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Experimentation compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Experimentation this skillcbrock84/headcount2k—~971Automated safety check: PassMIT
Ab Testingcoreyhaines31/marketingskills54k3 repos~2.8kAutomated safety check: PassMIT
AnalyticsNexus-JPF/note-companion8707 repos~2.2kAutomated safety check: PassMIT
Ab Test Setupfreekmurze/dotfiles1k15 repos~1.8kAutomated safety check: PassNone
Ad Test Designeraaron-he-zhu/aaron-marketing-skills2.9k2 repos~2.8kAutomated safety check: PassApache-2.0
Meta Tags Optimizernowork-studio/notfair-plugin3.9k1 repos~2.7kAutomated safety check: PassMIT

Similar skills

  • Ab Testing

    coreyhaines31/marketingskills

    When the user wants to plan, design, or implement an A/B test or experiment, or build a growth experimentation program.

    54k GitHub starsUsed in 3 repos~2.8k tokens
    Marketing & SEOAuto-check passed
  • Analytics

    Nexus-JPF/note-companion

    When the user wants to set up, improve, or audit analytics tracking and measurement.

    870 GitHub starsUsed in 7 repos~2.2k tokens
    Marketing & SEOAuto-check passed
  • Ab Test Setup

    freekmurze/dotfiles

    When the user wants to plan, design, or implement an A/B test or experiment.

    1k GitHub starsUsed in 15 repos~1.8k tokens
    Marketing & SEOAuto-check passed
  • Ad Test Designer

    aaron-he-zhu/aaron-marketing-skills

    A skill your agent uses when the user asks to "design an A/B test", "set up a creative/landing test", "run an incrementality test", or "is this result statistically and practically material?"…

    2.9k GitHub starsUsed in 2 repos~2.8k tokens
    Marketing & SEOAuto-check passed
  • Meta Tags Optimizer

    nowork-studio/notfair-plugin

    Writes and improves title tags, meta descriptions, Open Graph and Twitter card tags for click-through, with character counts and A/B test variants.

    3.9k GitHub starsUsed in 1 repo~2.7k tokens
    Marketing & SEOAuto-check passed
  • Ab Test Analyzer

    irinabuht12-oss/marketing-skills

    Statistical significance calculator for A/B test results with sample size requirements, segment breakdowns, and hypothesis generation.

    3.9k GitHub stars~1.4k tokensUpdated 14 days ago
    Marketing & SEOAuto-check passed

More from cbrock84/headcount

All 175 skills in this repo
  • Agent Hierarchy

    cbrock84/headcount

    Designs orchestrator-and-subagent hierarchies for a repository — splitting agents by exclusive write surface, pairing every producer with an independent auditor, and enforcing the split with a…

    2k GitHub stars~1.2k tokensUpdated 21 days ago
    Auto-check passed
  • Access And Identity

    cbrock84/headcount

    Designs and audits who can reach what — authentication, authorization models, privileged access, service credentials, and joiner-mover-leaver process.

    2k GitHub stars~1.1k tokensUpdated 21 days ago
    Auto-check passed
  • Account Based Marketing

    cbrock84/headcount

    Concentrates marketing and sales effort on a named set of accounts rather than on volume — qualifying whether the model fits your economics at all, building the account list and the buying group…

    2k GitHub stars~1.2k tokensUpdated 21 days ago
    Auto-check passed
  • Activation

    cbrock84/headcount

    Gets new users from signup to first real value — signup flow, onboarding, time-to-value, and the early experience that determines whether someone becomes a user or a lapsed account.

    2k GitHub stars~865 tokensUpdated 21 days ago
    Auto-check passed
  • AI Research Analyst

    cbrock84/headcount

    Produces executive-level research — market sizing, competitor mapping, trend analysis, and strategic intelligence — grounded in cited sources with the confidence in each claim made explicit.

    2k GitHub stars~916 tokensUpdated 21 days ago
    Auto-check passed
  • AI Search Optimization

    cbrock84/headcount

    Optimizes for AI assistants and AI-generated answers — being retrievable, being cited, and being represented accurately when a model answers on your behalf.

    2k GitHub stars~829 tokensUpdated 21 days ago
    Auto-check passed

Categories

Questions about Experimentation

What does Experimentation do?

Designs, runs, and reads A/B tests and growth experiments — hypothesis, sample size, duration, and honest interpretation. Experimentation is an agent skill from cbrock84/headcount. Designs, runs, and reads A/B tests and growth experiments — hypothesis, sample size, duration, and honest interpretation.

When should I use Experimentation?

Experimentation fits situations like: tasks that involve A/B testing.

How do I install Experimentation in Claude Code?

Run `npx skills add cbrock84/headcount --skill experimentation -a claude-code`. Or copy the skill folder (plugins/demand-generation/skills/experimentation in cbrock84/headcount) into .claude/skills/experimentation in your project. Claude Code loads it when a task matches its description.

How do I install Experimentation in Codex?

Run `npx skills add cbrock84/headcount --skill experimentation -a codex`. Or copy the skill folder (plugins/demand-generation/skills/experimentation in cbrock84/headcount) into .agents/skills/experimentation in your project. Codex loads it when a task matches its description.

Can I use Experimentation in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add cbrock84/headcount --skill experimentation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/experimentation, .gemini/skills/experimentation, .github/skills/experimentation and .opencode/skills/experimentation in your project.

What does Experimentation need to run?

SKILL.md names no scripts, command-line tools or credentials: Experimentation is instructions for the agent only.

Does Experimentation access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Experimentation safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Experimentation use?

Experimentation is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Experimentation use?

About 971 tokens (SKILL.md is roughly 3.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 366 tokens, read only when the agent opens those files.

What are the alternatives to Experimentation?

Skills that share tags, products or a category with Experimentation: Ab Testing (coreyhaines31/marketingskills, 54k stars), Analytics (Nexus-JPF/note-companion, 870 stars), Ab Test Setup (freekmurze/dotfiles, 1k stars) and Ad Test Designer (aaron-he-zhu/aaron-marketing-skills, 2.9k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Experimentation?

cbrock84 (a GitHub user) maintains it in cbrock84/headcount, which has 2,007 GitHub stars. The repository holds 175 skills in this directory. The repository was last updated on September 17, 2026.

Source: cbrock84/headcount on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.