Craft Experiment Design
amplitude/builder-skills
Write a hypothesis, define success metrics, and plan a holdout strategy.
Structured hypothesis formulation, experiment design, and results interpretation for Product Managers.
$ npx skills add jeremylongshore/tons-of-skills-marketplace --skill hypothesis-tester -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install jeremylongshore/tons-of-skills-marketplace hypothesis-tester --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/jeremylongshore/tons-of-skills-marketplace.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/.curated/hypothesis-tester .claude/skills/hypothesis-tester && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "hypothesis-tester" agent skill from https://github.com/jeremylongshore/tons-of-skills-marketplace/tree/main/skills/.curated/hypothesis-tester into .claude/skills/hypothesis-tester/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "hypothesis-tester", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/jeremylongshore/tons-of-skills-marketplace/tree/main/skills/.curated/hypothesis-testerType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add jeremylongshore/tons-of-skills-marketplace --skill hypothesis-tester -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install jeremylongshore/tons-of-skills-marketplace hypothesis-tester --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jeremylongshore/tons-of-skills-marketplace.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/.curated/hypothesis-tester .agents/skills/hypothesis-tester && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "hypothesis-tester" agent skill from https://github.com/jeremylongshore/tons-of-skills-marketplace/tree/main/skills/.curated/hypothesis-tester into .agents/skills/hypothesis-tester/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "hypothesis-tester", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add jeremylongshore/tons-of-skills-marketplace --skill hypothesis-tester -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install jeremylongshore/tons-of-skills-marketplace hypothesis-tester --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jeremylongshore/tons-of-skills-marketplace.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/.curated/hypothesis-tester .cursor/skills/hypothesis-tester && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "hypothesis-tester" agent skill from https://github.com/jeremylongshore/tons-of-skills-marketplace/tree/main/skills/.curated/hypothesis-tester into .cursor/skills/hypothesis-tester/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "hypothesis-tester", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/jeremylongshore/tons-of-skills-marketplace.git --path skills/.curated/hypothesis-tester--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add jeremylongshore/tons-of-skills-marketplace --skill hypothesis-tester -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install jeremylongshore/tons-of-skills-marketplace hypothesis-tester --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jeremylongshore/tons-of-skills-marketplace.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/.curated/hypothesis-tester .gemini/skills/hypothesis-tester && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "hypothesis-tester" agent skill from https://github.com/jeremylongshore/tons-of-skills-marketplace/tree/main/skills/.curated/hypothesis-tester into .gemini/skills/hypothesis-tester/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "hypothesis-tester", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install jeremylongshore/tons-of-skills-marketplace hypothesis-testerInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add jeremylongshore/tons-of-skills-marketplace --skill hypothesis-tester -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/jeremylongshore/tons-of-skills-marketplace.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/.curated/hypothesis-tester .github/skills/hypothesis-tester && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "hypothesis-tester" agent skill from https://github.com/jeremylongshore/tons-of-skills-marketplace/tree/main/skills/.curated/hypothesis-tester into .github/skills/hypothesis-tester/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "hypothesis-tester", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add jeremylongshore/tons-of-skills-marketplace --skill hypothesis-tester -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install jeremylongshore/tons-of-skills-marketplace hypothesis-tester --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jeremylongshore/tons-of-skills-marketplace.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/.curated/hypothesis-tester .opencode/skills/hypothesis-tester && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "hypothesis-tester" agent skill from https://github.com/jeremylongshore/tons-of-skills-marketplace/tree/main/skills/.curated/hypothesis-tester into .opencode/skills/hypothesis-tester/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "hypothesis-tester", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
hypothesis-testerStructured hypothesis formulation, experiment design, and results interpretation for Product Managers.
Hypothesis Tester is an agent skill from jeremylongshore/tons-of-skills-marketplace. Structured hypothesis formulation, experiment design, and results interpretation for Product Managers. Use when the user needs to validate an assumption, design an A/B test, or evaluate results. Trigger with "hypothesis", "A/B test", "experiment", "validate assumption", "test this", or "should we ship".
Its SKILL.md is about 2.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including reference files (for example `references/evidence-and-review.md`). Compatibility notes: Designed for Claude Code
It sits in Research & Science, covering Experimental design and A/B testing. The repository describes itself as: Model-agnostic agent-skills platform with a harness-free canonical layer, verified adapters, and the ccpi package manager. Explore at tonsofskills.com. The licence is MIT.
5 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit cfae287. It shows what the files ask for, not the result of running them.
Pre-approves these tools, so the agent can use them without asking each time:
ReadGlobGrepFrom allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md.
From the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
evanmiller.orgen.wikipedia.orgFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Designed for Claude Code
From compatibility in the SKILL.md frontmatter.
Hypothesis Tester loads about 2.6k tokens when it runs, and up to ~2.9k if it reads all its reference files. Until then it costs about 81 tokens; SKILL.md has 1,353 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from jeremylongshore/tons-of-skills-marketplace at commit cfae287, republished under its MIT licence (© jeremylongshore). 1,353 words, ~2,649 tokens.
.claude/skills/hypothesis-tester/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.Convert assumptions into falsifiable hypotheses, design proportionate tests, and interpret results without overstating causality. Use the evidence and review checklist before making a ship decision.
Use Glob to locate experiment artifacts, Grep to trace metrics, and Read to verify the design.
Act as an experiment design partner for a Product Manager. Your role is to help formulate testable hypotheses, design rigorous experiments, and interpret results honestly — including when the data says "don't ship."
Structure experiment work as:
User: We want to test whether adding a "remind me" button to push notifications improves engagement.
Good Response:
Let me help you design this properly.
Hypothesis
Users who see a "Remind me" action on push notifications will have a higher notification action rate than users without it, because re-scheduling gives users control over timing rather than forcing an immediate decision.
Falsifiable prediction: Treatment group notification action rate > Control group by ≥3 percentage points within 14 days.
Design
| Element | Detail |
|---|---|
| Type | A/B test, user-level randomization |
| Treatment | Push notifications include "Remind me" quick action (1hr / Tonight / Tomorrow) |
| Control | Current notifications (no remind action) |
| Split | 50/50 |
| Duration | 14 days minimum (to capture novelty decay) |
| Targeting | All users on iOS 16+ and Android 13+ (quick actions require these OS versions) |
Metrics
| Type | Metric | Current | Target |
|---|---|---|---|
| Primary | Notification action rate (any positive action) | 12% | 15% (+3pp) |
| Secondary | Reminded notification action rate | n/a | 40% |
| Secondary | Notification dismissal rate | 35% | ≤35% |
| Guardrail | App open rate | 45% | ≥43% (no more than 2pp drop) |
| Guardrail | Notification opt-out rate | 2.1% | ≤2.3% |
Sample Size
To detect a 3pp lift (12% → 15%) at 95% confidence and 80% power:
Pre-commitment
Risks to watch:
Want me to draft the experiment spec for engineering?
User: Our A/B test on the new onboarding flow is done. Treatment group had 38.7% D7 retention vs. 34.2% for control. Should we ship?
Good Response:
Let me look at this carefully before recommending.
What the headline says: +4.5pp retention lift (13% relative improvement). Sounds great.
What I need to check before recommending ship:
Statistical significance?
Novelty decay?
Segment consistency?
Guardrail metrics?
Cohort effects?
Interpretation:
The retention lift is real and statistically significant, but there are two yellow flags:
Recommendation: Ship with monitoring.
The core signal is strong (+4.5pp, significant, consistent across platforms). The decay and marketing concerns don't invalidate the result — they mean D30 data matters.
Do NOT declare "13% retention improvement" to leadership. Say: "4.5pp D7 lift, monitoring for D30 confirmation. Early signs positive but effect may moderate."
Structured experiment documentation including falsifiable hypotheses, test designs with sample size calculations, metric definitions (primary, secondary, guardrail), pre-commitment criteria, and honest ship/iterate/kill recommendations.
When traffic is insufficient for the desired minimum detectable effect, recommend alternative validation methods (user interviews, fake door tests, or qualitative signals). If experiment results are ambiguous, recommend extending rather than forcing a conclusion. When guardrail metrics are breached, flag this prominently even if the primary metric shows a lift.
© jeremylongshore, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 1 other file (references) in skills/.curated/hypothesis-tester of jeremylongshore/tons-of-skills-marketplace.
Open the folder on GitHubat commit cfae287
Hypothesis Tester next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Hypothesis Tester this skilljeremylongshore/tons-of-skills-marketplace | 2.8k | — | ~2.6k | Automated safety check: Pass | MIT | |
| Craft Experiment Designamplitude/builder-skills | 159 | — | ~522 | Automated safety check: Pass | None | |
| Experiment Designermohitagw15856/pm-claude-skills | 1.4k | — | ~1.1k | Automated safety check: Pass | MIT | |
| Data Scientistmagnus919/hermes-profiles | 289 | — | ~3.3k | Automated safety check: Pass | MIT | |
| Data Scientistmagnus919/agent-skills | 119 | — | ~4.1k | Automated safety check: Pass | MIT | |
| Measure Experiment Designproduct-on-purpose/pm-skills | 716 | — | ~1.1k | Automated safety check: Pass | Apache-2.0 |
amplitude/builder-skills
Write a hypothesis, define success metrics, and plan a holdout strategy.
mohitagw15856/pm-claude-skills
Design statistically rigorous A/B tests and interpret experiment results.
magnus919/hermes-profiles
PhD-level expertise in data science, statistics, and machine learning.
magnus919/agent-skills
A skill your agent uses for PhD-level expertise in data science, statistics, and machine learning: rigorous statistical analysis, experimental design, causal inference, advanced modeling, research…
product-on-purpose/pm-skills
Designs an A/B test or experiment with variants, success metrics, sample size, and duration for an existing hypothesis.
irinabuht12-oss/marketing-skills
Statistical significance calculator for A/B test results with sample size requirements, segment breakdowns, and hypothesis generation.
jeremylongshore/tons-of-skills-marketplace
Execute this skill enables AI assistant to conduct a security-focused code review using the security-agent plugin.
jeremylongshore/tons-of-skills-marketplace
Build this skill automates the adaptation of pre-trained machine learning models using transfer learning techniques.
jeremylongshore/tons-of-skills-marketplace
Execute proactive auto-loading: automatically detects and loads agents.md files.
jeremylongshore/tons-of-skills-marketplace
Aggregate and centralize performance metrics from applications, systems, databases, caches, and services.
jeremylongshore/tons-of-skills-marketplace
Execute this skill enables AI assistant to analyze capacity requirements and plan for future growth.
jeremylongshore/tons-of-skills-marketplace
Process use when you need to work with database indexing. An agent skill from jeremylongshore/tons-of-skills-marketplace.
Categories
Structured hypothesis formulation, experiment design, and results interpretation for Product Managers. Hypothesis Tester is an agent skill from jeremylongshore/tons-of-skills-marketplace. Structured hypothesis formulation, experiment design, and results interpretation for Product Managers.
Hypothesis Tester fits situations like: the user needs to validate an assumption; design an A/B test; evaluate results; with hypothesis.
Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill hypothesis-tester -a claude-code`. Or copy the skill folder (skills/.curated/hypothesis-tester in jeremylongshore/tons-of-skills-marketplace) into .claude/skills/hypothesis-tester in your project. Claude Code loads it when a task matches its description.
Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill hypothesis-tester -a codex`. Or copy the skill folder (skills/.curated/hypothesis-tester in jeremylongshore/tons-of-skills-marketplace) into .agents/skills/hypothesis-tester in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill hypothesis-tester -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/hypothesis-tester, .gemini/skills/hypothesis-tester, .github/skills/hypothesis-tester and .opencode/skills/hypothesis-tester in your project.
SKILL.md names no scripts, command-line tools or credentials: Hypothesis Tester is instructions for the agent only. Its frontmatter pre-approves these tools: Read, Glob, Grep. Compatibility (from SKILL.md): Designed for Claude Code.
SKILL.md names 2 domains. As links in the text: evanmiller.org and en.wikipedia.org. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Hypothesis Tester is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.6k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 239 tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Hypothesis Tester: Craft Experiment Design (amplitude/builder-skills, 159 stars), Experiment Designer (mohitagw15856/pm-claude-skills, 1.4k stars), Data Scientist (magnus919/hermes-profiles, 289 stars) and Data Scientist (magnus919/agent-skills, 119 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
jeremylongshore (a GitHub user) maintains it in jeremylongshore/tons-of-skills-marketplace, which has 2,827 GitHub stars. The repository holds 3,342 skills in this directory. The repository was last updated on October 10, 2026.
Source: jeremylongshore/tons-of-skills-marketplace on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.