Growth experiment design — structure a growth hypothesis, define metric, baseline, expected lift, and kill condition for a single experiment.

MITAuto-check: notesResearch & Science

Install Surge Experiment

skills CLI
$ npx skills add jeremylongshore/tons-of-skills-marketplace --skill surge-experiment -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install jeremylongshore/tons-of-skills-marketplace surge-experiment --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/jeremylongshore/tons-of-skills-marketplace.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/ai-agency/tonone/skills/surge-experiment .claude/skills/surge-experiment && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
surge-experiment
GitHub stars
2.8k
Token cost
~1.2k tokens
SKILL.md length
348 words
Files
2
Skills in repo
3,342
Repo updated
First seen
Licence
MIT

At a glance

Growth experiment design — structure a growth hypothesis, define metric, baseline, expected lift, and kill condition for a single experiment.

  • Works in 8 steps: State the Growth Lever → Write the Growth Hypothesis → Define the Experiment → …
  • Asked to design a growth experiment
  • SKILL.md covers Steps and Delivery
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Surge Experiment is an agent skill from jeremylongshore/tons-of-skills-marketplace. Growth experiment design — structure a growth hypothesis, define metric, baseline, expected lift, and kill condition for a single experiment. Use when asked to "design a growth experiment", "test this growth idea", "experiment framework", "how do we test if this works", or "growth hypothesis".

Its SKILL.md is about 1.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files (for example `.claude-plugin/plugin.json`).

It sits in Research & Science, covering Experimental design. The repository describes itself as: Model-agnostic agent-skills platform with a harness-free canonical layer, verified adapters, and the ccpi package manager. Explore at tonsofskills.com. The licence is MIT.

When your agent uses it

  • Asked to design a growth experiment
  • Test this growth idea
  • Experiment framework
  • How do we test if this works

Example prompts

  • “design a growth experiment”
  • “test this growth idea”
  • “experiment framework”
  • “/surge-experiment”

Requirements

  • Pre-approved tools (allowed-tools): Read, Write, Edit, Bash, Glob, Grep, WebFetch, WebSearch, Task, TodoWrite, AskUserQuestion

Workflow steps

8 steps, taken from the step headings in SKILL.md.

  1. State the Growth Lever
  2. Write the Growth Hypothesis
  3. Define the Experiment
  4. Define Metrics
  5. Size and Timeline
  6. Define the Decision Playbook
  7. Implementation Checklist
  8. Present Experiment Design

What it can do on your machine

Read from SKILL.md and the folder at commit cfae287. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Write
    • Edit
    • Bash
    • Glob
    • Grep
    • WebFetch
    • WebSearch
    • Task
    • TodoWrite

    …and 1 more on the same allowed-tools line.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Surge Experiment loads about 1.2k tokens when it runs. Until then it costs about 78 tokens; SKILL.md has 348 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~78
When it runs · the whole SKILL.md, loaded when a task matches
~1.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Read, Write, Edit, Bash, Glob, Grep, WebFetch, WebSearch, Task, TodoWrite, AskUserQuestion

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from jeremylongshore/tons-of-skills-marketplace at commit cfae287, republished under its MIT licence (© jeremylongshore). 348 words, ~1,219 tokens.

Download SKILL.mdSave it as .claude/skills/surge-experiment/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
surge-experiment
description
Growth experiment design — structure a growth hypothesis, define metric, baseline, expected lift, and kill condition for a single experiment. Use when asked to "design a growth experiment", "test this growth idea", "experiment framework", "how do we test if this works", or "growth hypothesis".
allowed-tools
Read, Write, Edit, Bash, Glob, Grep, WebFetch, WebSearch, Task, TodoWrite, AskUserQuestion
version
0.6.4
author
tonone-ai <hello@tonone.ai>
license
MIT

Growth Experiment Design

You are Surge — the growth engineer on the Product Team. Design the experiment before you build anything.

Follow the output format defined in docs/output-kit.md — 40-line CLI max, box-drawing skeleton, unified severity indicators, compressed prose.

Steps

Step 1: State the Growth Lever

Identify which part of the funnel this experiment targets:

Funnel StageExamples
AcquisitionSEO, paid ads, referral, partner integrations, content
ActivationOnboarding flow, time-to-value, setup wizard, templates
RetentionHabit loops, notifications, win-back emails, feature discovery
RevenueUpgrade triggers, paywall design, pricing page, trial length
ReferralInvite mechanics, share flows, virality coefficient

State: "This experiment targets [stage] and specifically [the lever]."

Step 2: Write the Growth Hypothesis

Use this format:

Hypothesis: If we [specific change], then [primary metric] will [increase/decrease]
            by [X%], because [mechanism — the causal theory].

We believe this because: [evidence — past experiment, user research, competitor observation,
                           or first-principles reasoning]

Kill condition: If [primary metric] does not move by [MDE] within [N days], we stop.

The mechanism is mandatory. Without it, you're guessing and won't learn from the result.

Step 3: Define the Experiment
Experiment name: [short, memorable]
Type: A/B test / Multi-variate / Phased rollout / Qualitative test

Control: [what the current experience is]
Variant: [exactly what changes — be specific enough to implement]

Target population: [who is included — new users / existing / paid / all?]
Exclusions: [who is excluded — why]
Traffic split: [50/50 / 90/10 / staged rollout — and why]
Step 4: Define Metrics

Primary metric (one only — the decision metric):

  • Metric: [name]
  • Baseline: [current value]
  • MDE: [minimum detectable effect — the smallest lift worth shipping for]
  • Direction: [increase / decrease]

Secondary metrics (directional, not decision):

  • [metric 1] — expected direction
  • [metric 2] — expected direction

Guardrail metrics (must not regress):

  • [metric] — must not drop more than [X%]
Step 5: Size and Timeline
Required users per variant: [N] — (use lumen-abtest for precise calculation)
Daily eligible traffic: [N]
Minimum run time: 14 days (for weekly seasonality)
Estimated run time: [N] days
Decision date: [date]

If run time exceeds 6 weeks, the experiment is too ambitious for available traffic. Options:

  • Increase MDE (accept a smaller win threshold)
  • Narrow the target population (run on power users only)
  • Run a qualitative test instead (5-user session, directional signal only)
Step 6: Define the Decision Playbook

What happens in each outcome:

WIN (primary metric ≥ MDE, p < 0.05, guardrails pass):
  → Ship to 100%. Timeline: [N days]. Owner: [eng]
  → Document: what we learned, why we think it worked

LOSS (null result — no significant movement):
  → Revert. Do NOT re-run without changing the hypothesis.
  → Document: what the null tells us about the mechanism

GUARDRAIL FAIL (primary wins but guardrail regresses):
  → Revert. Investigate the guardrail failure before re-running.

EARLY STOP (inconclusive after N days):
  → Default to control. Do not call a winner early.
Step 7: Implementation Checklist
  • Feature flag or experiment tool configured
  • All metrics instrumented (verify with lumen-instrument if needed)
  • Control and variant tested end-to-end in staging
  • Randomization unit set (user ID recommended — not session)
  • Holdout logged and reproducible
  • Stakeholders aware of timeline and decision criteria
  • Calendar reminder set for decision date
Step 8: Present Experiment Design

Output the complete experiment spec using the CLI skeleton format.

Delivery

If output exceeds the 40-line CLI budget, invoke /atlas-report with the full findings. The HTML report is the output. CLI is the receipt — box header, one-line verdict, top 3 findings, and the report path. Never dump analysis to CLI.

© jeremylongshore, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in plugins/ai-agency/tonone/skills/surge-experiment of jeremylongshore/tons-of-skills-marketplace.

  • SKILL.md
  • .claude-plugin/plugin.json

Open the folder on GitHubat commit cfae287

Compare with similar skills

Surge Experiment next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Surge Experiment compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Surge Experiment this skilljeremylongshore/tons-of-skills-marketplace2.8k—~1.2kAutomated safety check: NotesMIT
Scientific Critical Thinkingweapp-tailwindcss/weapp-tailwindcss1.9k22 repos~5.9kAutomated safety check: NotesMIT
Benchmark Paper TemplateHKUSTDial/Supervisor-Skills8.8k—~2.8kAutomated safety check: PassCC-BY-4.0
Claim-Driven Experiment PlannerzjYao36/Auto-Research-Refine1286 repos~2.3kAutomated safety check: NotesNone
Research Refine PipelinezjYao36/Auto-Research-Refine1285 repos~1.4kAutomated safety check: NotesNone
Metabolic Study Planneraiming-lab/AutoResearchClaw15k—~1.9kAutomated safety check: PassMIT

Similar skills

  • Scientific Critical Thinking

    weapp-tailwindcss/weapp-tailwindcss

    Evaluate research rigor. An agent skill from weapp-tailwindcss/weapp-tailwindcss.

    1.9k GitHub starsUsed in 22 repos~5.9k tokens
    Research & ScienceAuto-check: notes
  • Benchmark Paper Template

    HKUSTDial/Supervisor-Skills

    Structures benchmark and evaluation papers around five pillars, with a completeness audit, an Introduction logic chain, a section skeleton and a pre-submission checklist.

    8.8k GitHub stars~2.8k tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed
  • Claim-Driven Experiment Planner

    zjYao36/Auto-Research-Refine

    Turns a refined research proposal into a claim-to-evidence-to-run-order roadmap instead of a sprawling benchmark wishlist.

    128 GitHub starsUsed in 6 repos~2.3k tokens
    Research & ScienceAuto-check: notes
  • Research Refine Pipeline

    zjYao36/Auto-Research-Refine

    Chains research-refine and experiment-plan to turn a vague research direction into a focused proposal and a claim-driven experiment roadmap.

    128 GitHub starsUsed in 5 repos~1.4k tokens
    Research & ScienceAuto-check: notes
  • Metabolic Study Planner

    aiming-lab/AutoResearchClaw

    Turns a broad metabolic modelling topic into a concrete, paper-shaped plan with organism, model, perturbations, metrics and figures before any FBA code is written.

    15k GitHub stars~1.9k tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed
  • Analytical Method Validation Planner

    K-Dense-AI/scientific-agent-skills

    Plans, runs, and documents analytical method validation, verification, or transfer studies under ICH Q2(R2)/Q14, USP, ICH M10, CLSI EP, or ISO/IEC 17025.

    48k GitHub starsUsed in 1 repo~4.9k tokens
    Research & ScienceAuto-check: notes

More from jeremylongshore/tons-of-skills-marketplace

All 3,342 skills in this repo
  • Performing Security Code Review

    jeremylongshore/tons-of-skills-marketplace

    Execute this skill enables AI assistant to conduct a security-focused code review using the security-agent plugin.

    2.8k GitHub starsUsed in 2 repos~1.3k tokens
    Auto-check: notes
  • Adapting Transfer Learning Models

    jeremylongshore/tons-of-skills-marketplace

    Build this skill automates the adaptation of pre-trained machine learning models using transfer learning techniques.

    2.8k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Agent Context Loader

    jeremylongshore/tons-of-skills-marketplace

    Execute proactive auto-loading: automatically detects and loads agents.md files.

    2.8k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Aggregating Performance Metrics

    jeremylongshore/tons-of-skills-marketplace

    Aggregate and centralize performance metrics from applications, systems, databases, caches, and services.

    2.8k GitHub stars~1.2k tokensUpdated today
    Auto-check passed
  • Analyzing Capacity Planning

    jeremylongshore/tons-of-skills-marketplace

    Execute this skill enables AI assistant to analyze capacity requirements and plan for future growth.

    2.8k GitHub stars~947 tokensUpdated today
    Auto-check passed
  • Analyzing Database Indexes

    jeremylongshore/tons-of-skills-marketplace

    Process use when you need to work with database indexing. An agent skill from jeremylongshore/tons-of-skills-marketplace.

    2.8k GitHub stars~2k tokensUpdated today
    Auto-check passed

Questions about Surge Experiment

What does Surge Experiment do?

Growth experiment design — structure a growth hypothesis, define metric, baseline, expected lift, and kill condition for a single experiment. Surge Experiment is an agent skill from jeremylongshore/tons-of-skills-marketplace. Growth experiment design — structure a growth hypothesis, define metric, baseline, expected lift, and kill condition for a single experiment.

When should I use Surge Experiment?

Surge Experiment fits situations like: asked to design a growth experiment; test this growth idea; experiment framework; how do we test if this works.

How do I install Surge Experiment in Claude Code?

Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill surge-experiment -a claude-code`. Or copy the skill folder (plugins/ai-agency/tonone/skills/surge-experiment in jeremylongshore/tons-of-skills-marketplace) into .claude/skills/surge-experiment in your project. Claude Code loads it when a task matches its description.

How do I install Surge Experiment in Codex?

Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill surge-experiment -a codex`. Or copy the skill folder (plugins/ai-agency/tonone/skills/surge-experiment in jeremylongshore/tons-of-skills-marketplace) into .agents/skills/surge-experiment in your project. Codex loads it when a task matches its description.

Can I use Surge Experiment in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill surge-experiment -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/surge-experiment, .gemini/skills/surge-experiment, .github/skills/surge-experiment and .opencode/skills/surge-experiment in your project.

What does Surge Experiment need to run?

SKILL.md names no scripts, command-line tools or credentials: Surge Experiment is instructions for the agent only. Its frontmatter pre-approves these tools: Read, Write, Edit, Bash, Glob, Grep, WebFetch, WebSearch, Task, TodoWrite, AskUserQuestion.

Does Surge Experiment access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Surge Experiment safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Surge Experiment use?

Surge Experiment is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Surge Experiment use?

About 1.2k tokens (SKILL.md is roughly 4.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Surge Experiment?

Skills that share tags, products or a category with Surge Experiment: Scientific Critical Thinking (weapp-tailwindcss/weapp-tailwindcss, 1.9k stars), Benchmark Paper Template (HKUSTDial/Supervisor-Skills, 8.8k stars), Claim-Driven Experiment Planner (zjYao36/Auto-Research-Refine, 128 stars) and Research Refine Pipeline (zjYao36/Auto-Research-Refine, 128 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Surge Experiment?

jeremylongshore (a GitHub user) maintains it in jeremylongshore/tons-of-skills-marketplace, which has 2,827 GitHub stars. The repository holds 3,342 skills in this directory. The repository was last updated on October 10, 2026.

Source: jeremylongshore/tons-of-skills-marketplace on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.