Framework for running single-variable cold email experiments.

MITAuto-check passedResearch & Science

Install Experiment Design

skills CLI
$ npx skills add growthenginenowoslawski/coldoutboundskills --skill experiment-design -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install growthenginenowoslawski/coldoutboundskills experiment-design --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/growthenginenowoslawski/coldoutboundskills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/experiment-design .claude/skills/experiment-design && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
experiment-design
GitHub stars
753
Token cost
~2.5k tokens
SKILL.md length
1,147 words
Files
1
Skills in repo
49
Repo updated
First seen
Licence
MIT

At a glance

Framework for running single-variable cold email experiments.

  • Works in 8 steps: Name your hypothesis → Identify the single variable → Calculate minimum sample size → …
  • The user wants to improve a campaign
  • SKILL.md covers Why this exists, The three experiment types, The Framework and What NOT to experiment on (at…, plus 5 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Experiment Design is an agent skill from growthenginenowoslawski/coldoutboundskills. Framework for running single-variable cold email experiments. Defines experiment types (list-only, copy-only, combined), confidence weighting, minimum sample sizes, and success criteria. Use when the user wants to improve a campaign, test a new list vs old one, or compare copy variants. Prevents the

Its SKILL.md is about 2.5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Research & Science, covering Experimental design and Cold outreach. The repository describes itself as: Open-source Claude Code skills for cold email and outbound sales. Grade campaigns, export Prospeo searches, scrape Google Maps — all from Claude Code. The licence is MIT.

When your agent uses it

  • The user wants to improve a campaign
  • Test a new list vs old one
  • Compare copy variants

Example prompts

  • “/experiment-design”

Workflow steps

8 steps, taken from the step headings in SKILL.md.

  1. Name your hypothesis
  2. Identify the single variable
  3. Calculate minimum sample size
  4. Build the success criteria up front
  5. Launch both arms simultaneously
  6. Measure at day 21
  7. Weight the learnings
  8. Decide what to do with the result

What it can do on your machine

Read from SKILL.md and the folder at commit 25c5d85. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are yaml).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Experiment Design loads about 2.5k tokens when it runs. Until then it costs about 80 tokens; SKILL.md has 1,147 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~80
When it runs · the whole SKILL.md, loaded when a task matches
~2.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from growthenginenowoslawski/coldoutboundskills at commit 25c5d85, republished under its MIT licence (© growthenginenowoslawski). 1,147 words, ~2,495 tokens.

Download SKILL.mdSave it as .claude/skills/experiment-design/SKILL.md (or your agent's skills folder).
name
experiment-design
description
Framework for running single-variable cold email experiments. Defines experiment types (list-only, copy-only, combined), confidence weighting, minimum sample sizes, and success criteria. Use when the user wants to improve a campaign, test a new list vs old one, or compare copy variants. Prevents the

Experiment Design

If you change your list, your copy, and your offer at the same time, you learn nothing. This skill forces you to isolate one variable per experiment so you actually learn what's working.

Why this exists

Most cold email operators run "throw-everything" experiments. Campaign 1 gets a new list, new copy, and a new offer. It works better. They declare victory. But they can't tell you WHY — was it the list? The copy? The offer?

Then campaign 2 changes all three again. Regression. Nobody knows why.

This skill is the antidote: plan each experiment around ONE variable, keep everything else constant, and confidence-weight the results.

The three experiment types

A. List-only experiment
  • What varies: the list (targeting criteria)
  • What stays fixed: copy, offer, sending infrastructure, sequence timing
  • What you learn: whether this segment is a better fit than the baseline
  • Confidence on learnings: HIGH on targeting, LOW on copy (because copy wasn't tested)
B. Copy-only experiment
  • What varies: the copy (subject, body, sequence, or A/B variant)
  • What stays fixed: list, offer, infrastructure
  • What you learn: whether this copy resonates with this audience
  • Confidence on learnings: HIGH on copy, LOW on targeting
C. Combined experiment (use sparingly)
  • What varies: list AND copy (and sometimes offer)
  • What stays fixed: only infrastructure
  • When to use: launching a whole new campaign for a new ICP. You can't isolate because everything is new.
  • Confidence on learnings: MEDIUM on everything. Use as hypothesis-generation, not conclusion.

The Framework

Step 1: Name your hypothesis

Every experiment starts with a one-sentence hypothesis:

"Targeting Heads of Marketing at 50-200 person B2B SaaS companies will get a higher positive reply rate than our current VP Sales baseline, because [reason]."

Or:

"Leading with a question about their recent product launch will get a higher reply rate than our current benefit-focused opener, because [reason]."

If you can't write the hypothesis in one sentence, you don't understand the experiment yet. Go back.

Step 2: Identify the single variable

Write down exactly what changes and what stays the same.

Variable: Target job title
Change: "VP Sales" → "Head of Marketing"

Constants:
- Industry filter: unchanged
- Headcount: unchanged
- Geography: unchanged
- Copy: unchanged (same 4-step sequence)
- Offer: unchanged (same lead magnet)
- Sending infrastructure: unchanged (same 20 domains, 40 inboxes)
- Send schedule: unchanged

If ANY constant is actually changing, stop. Either lock it down, or reclassify as a combined experiment.

Baseline sanity check — the 1% rule

Before running any experiment, confirm your baseline is healthy: overall reply rate ≥1% after 200+ sends. If your baseline is below 1% after 200 sends, the problem isn't your experiment — your infrastructure or copy is already broken. Run /email-deliverability-audit first.

Running an experiment on a broken baseline is wasted effort: you'll learn that "both arms are bad," not "which arm wins."

Step 3: Calculate minimum sample size

The smaller your effect, the more leads you need. Use these rough rules for cold email:

Current baselineExpected liftMinimum sends per arm
1% positive reply rate2x (1% → 2%)~500
1% positive reply rate1.5x (1% → 1.5%)~2,000
1% positive reply rate1.2x (1% → 1.2%)~10,000
2% positive reply rate2x (2% → 4%)~250
2% positive reply rate1.5x (2% → 3%)~1,000

Rule of thumb: if your test has fewer than 500 sends per arm, you can't tell signal from noise.

For most beginners, 2,000 sends per arm is the right default.

Step 4: Build the success criteria up front

Before launching, write:

Success = positive reply rate > X% (our current baseline is Y%)
Failure = positive reply rate < Z%
Inconclusive = between X and Z
Required sample: at least N sends per arm, reported after day 21 of sequence

Decide now — not after seeing the data. This prevents "oh we learned something else instead" rationalization.

Step 5: Launch both arms simultaneously

Same day, same sending infrastructure split, same sequence. If your control arm sends Monday and your variant arm sends Thursday, day-of-week effects will confound the test.

Best practice in Smartlead/Instantly: create two campaigns, assign each half of your inboxes, launch at the exact same time, same schedule.

Step 6: Measure at day 21

Wait until the full sequence (typically Day 0, 3, 7, 11 + reply grace period) has finished for ALL leads. Measuring earlier biases toward the first email's reply rate.

Pull metrics via /positive-reply-scoring skill:

  • Total sent (per arm)
  • Total replies (per arm)
  • Positive replies (per arm, classified by Claude)
  • Positive reply rate = positive replies / total sent

Secondary metrics (report but don't optimize for):

  • Overall reply rate (positive replies / total sent is primary, but this shows raw engagement)
  • Open rate (if available)
  • Bounce rate (sanity check — if one arm bounces more, your list is bad, not your copy)
Show full SKILL.md (445 more words)Show less
Step 7: Weight the learnings

Use this framework when reporting:

Experiment: <name>
Type: List-only | Copy-only | Combined
Variable: <what changed>

Result: <winner name> at <positive reply rate>% vs <baseline>%
Confidence: HIGH | MEDIUM | LOW (based on experiment type + sample size)

Learnings (by confidence):
  HIGH confidence:
    - <thing you can trust>
  MEDIUM confidence:
    - <thing that looks good but needs replication>
  LOW confidence:
    - <thing you're speculating about>

HIGH only if: experiment type isolates the variable AND sample size meets the minimum.

Step 8: Decide what to do with the result
  • Winner by ≥20% lift, HIGH confidence: adopt as new baseline. Document. Move to next experiment.
  • Winner by 10-20% lift, HIGH confidence: run a replication experiment with fresh leads. If it wins again, adopt.
  • Winner by <10% lift: inconclusive. Run bigger next time or drop.
  • Loser: document WHY you think it lost. Don't just move on — the loss is a learning.
  • Combined experiment winner: do NOT adopt as a new baseline. Instead, split into single-variable follow-ups to figure out which part actually drove the lift.

What NOT to experiment on (at first)

If you're just starting out, don't experiment at all until you have a baseline from a single shipped campaign running for 3 weeks. You need a control before you can run tests.

Once you have a baseline, the priority order of experiments is usually:

  1. List (biggest impact — bad list kills any copy)
  2. Offer / lead magnet (second biggest — "book a call" vs a real magnet)
  3. Subject line (cheap to test, drives open rate)
  4. Opener / first line (after subject, the big lever)
  5. CTA (how you end the email)
  6. Sequence timing (Day 3 vs Day 2 follow-up)
  7. Sequence length (4-step vs 6-step)

Don't jump to step 6 when step 1 is broken.

Common mistakes

  • A/B testing inside one campaign. Smartlead's A/B variant feature mixes the data — fine for small copy tweaks, terrible for hypothesis testing. Use TWO campaigns for real isolation.
  • "I'll test 3 things at once." You'll learn nothing.
  • Calling it early. Wait 21 days minimum. Cold email replies trickle in over weeks.
  • Changing infrastructure mid-test. If one arm uses new domains and one uses old, deliverability skews everything.
  • Ignoring bounce rate. If variant's bounce rate is 2x control, the list is bad — not the copy. Disqualify the test.

Output: experiment plan file

At the end of planning, write to:

~/cold-email-ai-skills/profiles/<business-slug>/experiments/YYYY-MM-DD-<name>.yaml

Schema:

yaml
experiment:
  name: <short name>
  hypothesis: <one sentence>
  type: list-only | copy-only | combined
  variable: <what changes>
  constants: <list of what stays fixed>

success_criteria:
  positive_reply_rate_target: <float>
  baseline: <float>
  minimum_sends_per_arm: <int>
  measurement_date: <YYYY-MM-DD>

arms:
  control:
    smartlead_campaign_id: <tbd until launch>
    description: <what's in the control>
  variant:
    smartlead_campaign_id: <tbd until launch>
    description: <what's in the variant>

results: <empty until day 21>
  control_positive_reply_rate: null
  variant_positive_reply_rate: null
  winner: null
  confidence: null
  decision: null

References

  • references/sample-size-calculator.md — longer math for power calculations
  • references/example-experiments/ — 3 worked examples (list, copy, combined)
  • /positive-reply-scoring skill — how to actually measure the outcome

What to do next

Launch the planned experiment via /smartlead-campaign-upload-public (manual) or /auto-research-public (automated). Use the variants.yaml from /campaign-copywriting.

Then wait 21 days before evaluating — reply rate needs that long to stabilize. After 21 days, /positive-reply-scoring on each arm.

Or wait: if you don't have 2,000+ leads per experiment arm, you can't detect normal-sized effects. Build a bigger list (/prospeo-full-export, /disco-like) first.

  • /campaign-copywriting — produces the copy variants this experiment tests
  • /smartlead-campaign-upload-public — launches each arm
  • /positive-reply-scoring — measures the outcome after 21 days

© growthenginenowoslawski, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/experiment-design of growthenginenowoslawski/coldoutboundskills.

Open the folder on GitHubat commit 25c5d85

Compare with similar skills

Experiment Design next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Experiment Design compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Experiment Design this skillgrowthenginenowoslawski/coldoutboundskills753—~2.5kAutomated safety check: PassMIT
Academic Researchvoidful/academic-skills135—~887Automated safety check: PassMIT
Scientific Critical Thinkingweapp-tailwindcss/weapp-tailwindcss1.9k22 repos~5.9kAutomated safety check: NotesMIT
Benchmark Paper TemplateHKUSTDial/Supervisor-Skills8.8k—~2.8kAutomated safety check: PassCC-BY-4.0
Claim-Driven Experiment PlannerzjYao36/Auto-Research-Refine1286 repos~2.3kAutomated safety check: NotesNone
Research Refine PipelinezjYao36/Auto-Research-Refine1285 repos~1.4kAutomated safety check: NotesNone

Similar skills

  • Academic Research

    voidful/academic-skills

    Complete academic research skill suite covering the full pipeline: paper reading (read/explain papers with storytelling), idea generation (brainstorm research directions), experiment design (plan…

    135 GitHub stars~887 tokensUpdated 6 mo ago
    Research & ScienceAuto-check passed
  • Scientific Critical Thinking

    weapp-tailwindcss/weapp-tailwindcss

    Evaluate research rigor. An agent skill from weapp-tailwindcss/weapp-tailwindcss.

    1.9k GitHub starsUsed in 22 repos~5.9k tokens
    Research & ScienceAuto-check: notes
  • Benchmark Paper Template

    HKUSTDial/Supervisor-Skills

    Structures benchmark and evaluation papers around five pillars, with a completeness audit, an Introduction logic chain, a section skeleton and a pre-submission checklist.

    8.8k GitHub stars~2.8k tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed
  • Claim-Driven Experiment Planner

    zjYao36/Auto-Research-Refine

    Turns a refined research proposal into a claim-to-evidence-to-run-order roadmap instead of a sprawling benchmark wishlist.

    128 GitHub starsUsed in 6 repos~2.3k tokens
    Research & ScienceAuto-check: notes
  • Research Refine Pipeline

    zjYao36/Auto-Research-Refine

    Chains research-refine and experiment-plan to turn a vague research direction into a focused proposal and a claim-driven experiment roadmap.

    128 GitHub starsUsed in 5 repos~1.4k tokens
    Research & ScienceAuto-check: notes
  • Metabolic Study Planner

    aiming-lab/AutoResearchClaw

    Turns a broad metabolic modelling topic into a concrete, paper-shaped plan with organism, model, perturbations, metrics and figures before any FBA code is written.

    15k GitHub stars~1.9k tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed

More from growthenginenowoslawski/coldoutboundskills

All 49 skills in this repo
  • Email Deliverability Audit

    growthenginenowoslawski/coldoutboundskills

    Diagnostic audit for a running cold email program. An agent skill from growthenginenowoslawski/coldoutboundskills.

    753 GitHub stars~3.4k tokensUpdated 5 days ago
    Auto-check passed
  • Icp Onboarding

    growthenginenowoslawski/coldoutboundskills

    Conversational intake for cold email campaigns. An agent skill from growthenginenowoslawski/coldoutboundskills.

    753 GitHub stars~1.9k tokensUpdated 5 days ago
    Auto-check passed
  • List Builder

    growthenginenowoslawski/coldoutboundskills

    META skill — build the largest possible qualified lead list for any request, end to end.

    753 GitHub stars~4.7k tokensUpdated 5 days ago
    Auto-check: notes
  • Auto Research Public

    growthenginenowoslawski/coldoutboundskills

    Autonomous cold email campaign launcher. An agent skill from growthenginenowoslawski/coldoutboundskills.

    753 GitHub stars~2.9k tokensUpdated 5 days ago
    Auto-check passed
  • Blitz List Builder

    growthenginenowoslawski/coldoutboundskills

    Use the Blitz API to find decision-makers at specific companies when you already have a list of company domains.

    753 GitHub stars~2k tokensUpdated 5 days ago
    Auto-check passed
  • Deliverability Test Public

    growthenginenowoslawski/coldoutboundskills

    Compare reply rates, bounce rates, and positive reply rates broken down by inbox type (SMTP / Gmail / Outlook) for a Smartlead account.

    753 GitHub stars~1k tokensUpdated 5 days ago
    Auto-check passed

Questions about Experiment Design

What does Experiment Design do?

Framework for running single-variable cold email experiments. Experiment Design is an agent skill from growthenginenowoslawski/coldoutboundskills. Framework for running single-variable cold email experiments.

When should I use Experiment Design?

Experiment Design fits situations like: the user wants to improve a campaign; test a new list vs old one; compare copy variants.

How do I install Experiment Design in Claude Code?

Run `npx skills add growthenginenowoslawski/coldoutboundskills --skill experiment-design -a claude-code`. Or copy the skill folder (skills/experiment-design in growthenginenowoslawski/coldoutboundskills) into .claude/skills/experiment-design in your project. Claude Code loads it when a task matches its description.

How do I install Experiment Design in Codex?

Run `npx skills add growthenginenowoslawski/coldoutboundskills --skill experiment-design -a codex`. Or copy the skill folder (skills/experiment-design in growthenginenowoslawski/coldoutboundskills) into .agents/skills/experiment-design in your project. Codex loads it when a task matches its description.

Can I use Experiment Design in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add growthenginenowoslawski/coldoutboundskills --skill experiment-design -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/experiment-design, .gemini/skills/experiment-design, .github/skills/experiment-design and .opencode/skills/experiment-design in your project.

What does Experiment Design need to run?

SKILL.md names no scripts, command-line tools or credentials: Experiment Design is instructions for the agent only.

Does Experiment Design access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Experiment Design safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Experiment Design use?

Experiment Design is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Experiment Design use?

About 2.5k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Experiment Design?

Skills that share tags, products or a category with Experiment Design: Academic Research (voidful/academic-skills, 135 stars), Scientific Critical Thinking (weapp-tailwindcss/weapp-tailwindcss, 1.9k stars), Benchmark Paper Template (HKUSTDial/Supervisor-Skills, 8.8k stars) and Claim-Driven Experiment Planner (zjYao36/Auto-Research-Refine, 128 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Experiment Design?

growthenginenowoslawski (a GitHub user) maintains it in growthenginenowoslawski/coldoutboundskills, which has 753 GitHub stars. The repository holds 49 skills in this directory. The repository was last updated on October 5, 2026.

Source: growthenginenowoslawski/coldoutboundskills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.