Agent skill

Product Experiments

by RefoundAI in RefoundAI/lenny-skills

Help users design, execute, and analyze product experiments to validate hypotheses and measure true incremental impact while avoiding common statistical pitfalls.

MITAuto-check passedResearch & Science

Install Product Experiments

skills CLI
$ npx skills add RefoundAI/lenny-skills --skill product-experiments -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install RefoundAI/lenny-skills product-experiments --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/RefoundAI/lenny-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/product-experiments .claude/skills/product-experiments && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
product-experiments
GitHub stars
1.4k
Token cost
~2k tokens
SKILL.md length
1,166 words
Files
3 (incl. references)
Skills in repo
76
Repo updated
First seen
Licence
MIT

At a glance

Help users design, execute, and analyze product experiments to validate hypotheses and measure true incremental impact while avoiding common statistical pitfalls.

  • Works in 4 steps: Hypothesis Definition - Guide the user… → Experimental Design - Help determine the… → Statistical Analysis - Support the… → …
  • Research & Science work in your project
  • SKILL.md covers How to Help, Core Principles, Templates & Frameworks and Questions to Help Users, plus 3 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Product Experiments is an agent skill from RefoundAI/lenny-skills. Help users design, execute, and analyze product experiments to validate hypotheses and measure true incremental impact while avoiding common statistical pitfalls.

Its SKILL.md is about 2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files, including reference files (for example `references/artifacts.md` and `references/guest-insights.md`).

It sits in Research & Science. The repository describes itself as: 86 product management skills from Lenny's Podcast for Claude Code and AI agents. Hiring, user research, strategy, shipping, and more. The licence is MIT.

When your agent uses it

  • Research & Science work in your project

Example prompts

  • “/product-experiments”

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Hypothesis Definition - Guide the user in drafting clear, falsifiable hypotheses based on user behavior theories.
  2. Experimental Design - Help determine the right metrics, sample sizes, and guardrail metrics for a clean test.
  3. Statistical Analysis - Support the interpretation of p-values, confidence intervals, and potential sample ratio mismatches.
  4. Strategic Evaluation - Assist in deciding whether to ship, iterate, or kill a feature based on experiment results and long-term business…

What it can do on your machine

Read from SKILL.md and the folder at commit 13598cc. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Product Experiments loads about 2k tokens when it runs, and up to ~15k if it reads all its reference files. Until then it costs about 46 tokens; SKILL.md has 1,166 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~46
When it runs · the whole SKILL.md, loaded when a task matches
~2k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~15k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from RefoundAI/lenny-skills at commit 13598cc, republished under its MIT licence (© RefoundAI). 1,166 words, ~2,011 tokens.

Download SKILL.mdSave it as .claude/skills/product-experiments/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
product-experiments
description
Help users design, execute, and analyze product experiments to validate hypotheses and measure true incremental impact while avoiding common statistical pitfalls.

Product Experimentation Excellence

Drive measurable growth and mitigate risk through rigorous A/B testing and data-driven learning.

Help the user with product experimentation excellence using insights from 9 guests and posts across Lenny's Podcast and Newsletter.

How to Help

  1. Hypothesis Definition - Guide the user in drafting clear, falsifiable hypotheses based on user behavior theories.
  2. Experimental Design - Help determine the right metrics, sample sizes, and guardrail metrics for a clean test.
  3. Statistical Analysis - Support the interpretation of p-values, confidence intervals, and potential sample ratio mismatches.
  4. Strategic Evaluation - Assist in deciding whether to ship, iterate, or kill a feature based on experiment results and long-term business impact.

Core Principles

Use long-term holdouts for true incrementality

Archie Abrams: "So we constantly will relook at an experiment a year later, see that the way the GMV curve for the distribution was different than we might've originally thought. And that'll actually change what we do from that previous experiment. And so there's a lot of longterm monitoring of experiments over these very long time horizons to both inform what those input metrics are and more importantly hold ourselves accountable to, did we actually move what we cared about, which is that longterm GMV, in the right way?"

Implement holdout groups for one or more years to distinguish between immediate growth and short-term pull-forward effects. This ensures you are measuring the genuine downstream business impact of changes.

Focus experiments on risk mitigation

Lauryn Isford: "So, with all that said, generally my advice is to experiment when you need to and to primarily see it as a risk mitigation tactic when you're making dramatic changes and to let the product development process do more work. So, spend more time with customers, be more rigorous in understanding precisely what problem you're solving, get mocks in front of people and see how they react, and hopefully have more conviction than you otherwise would when you ship something that it's okay if every customer sees it tomorrow and that the experiment doesn't actually matter as much."

Prioritize A/B testing for high-stakes, dramatic product changes rather than using it solely for precise metric attribution. Invest in qualitative research first to build conviction before launching high-risk tests.

Balance success metrics with guardrails

Ronny Kohavi: "It's very easy to increase revenue by doing theatrics. Displaying more ads is a trivial way to raise revenue, but it hurts the user experience. And we've done the experiments to show that. In this case, this was just a home run that improved revenue, didn't significantly hurt the guardrail metrics."

Develop a robust Overall Evaluation Criterion (OEC) that includes guardrail metrics. This prevents short-term wins from inadvertently degrading the long-term user experience or retention.

Normalize a high failure rate

Ronny Kohavi: "At Bing, which is a much more optimized domain after we've been optimizing it for a while, the failure rate was around 85%. So it's harder to improve something that you've been optimizing for a while. And then at Airbnb, this 92% number is the highest failure rate that I've observed."

Expect that 80 percent to 92 percent of experiments in optimized domains will fail. Calibrating team expectations around these industry standards prevents discouragement and maintains high testing volume.

Skip testing for low-risk best practices

From "When NOT to run an experiment – Issue 54": "If you can run experiments quickly and easily (e.g. a few hours), this decision is generally easy: run the experiment. If running experiments is a pain in the butt, and the changes are relatively benign, you can probably skip the experiment."

Shipping directly is often superior to experimenting when the time required for statistical significance outweighs the data value. Avoid formal tests for standard industry practices where downside risk is minimal.

Apply Twyman's Law to surprising wins

Ronny Kohavi: "We can talk later about Wyman's law, but that was the first reaction, which is, 'This is too good to be true. Let's find a bug.' And we did. And we looked for several times, and we replicated the experiment several times, and there was nothing wrong with it."

Treat any result that looks too good to be true with immediate skepticism. Conduct rigorous bug-hunting and replicate surprising results multiple times to ensure they are not technical flukes.

Show full SKILL.md (458 more words)Show less

Templates & Frameworks

  • Three Reasons to Skip an Experiment (When NOT to run an experiment – Issue 54) - A decision framework with three distinct scenarios where shipping without an experiment is the right call
  • Whole Pie Impact Calculation (Product Amdahl's Law) (The secret to Duolingo’s exponential growth) - Framework for calculating total experiment impact by factoring in what percentage of users will actually see the change
  • 5 Pillars of Experimentation Culture (Fostering a culture of experimentation) - A framework of five key areas to build a strong culture of experimentation, derived from Airbnb's approach
  • Risk Assessment Questions Before Running an Experiment (When NOT to run an experiment – Issue 54) - Five questions to evaluate whether the risk/reward of running an experiment is worth it
  • Long-Term Holdout Experiment Method (How Duolingo builds product) - A technique for measuring long-term effects of features like social by maintaining a control group for 3+ months
  • Sample Size Calculation Example for Sign-up Flow (When NOT to run an experiment – Issue 54) - A concrete example showing how many users are needed to detect a 5% change on a step converting at 10%
  • Decision Framework for Failed Experiments: Kill, Iterate, or Ship (Communicating bad news - Issue 26) - Three options to evaluate when a project experiment shows negative results
  • Sample Ratio Mismatch (SRM) check (Ronny Kohavi) - A statistical test to verify that the ratio of users in control vs. treatment matches the designed ratio. The single most important validity check for any A/B t

See references/artifacts.md for the full list with details.

Questions to Help Users

  • "What is the core psychological hypothesis you are testing with this change?"
  • "What are the guardrail metrics we need to monitor to ensure we don't hurt the long-term experience?"
  • "Do we have enough traffic to reach statistical significance within a reasonable timeframe?"
  • "Is the potential lift from this experiment worth the engineering and analysis time compared to shipping a known best practice?"
  • "How does this experiment fit into our balance of low-effort wins versus high-variance big bets?"
  • "If this experiment fails, what specific user behavior insight will we gain?"

Common Mistakes to Flag

  • The search for significance - Making decisions based on real-time P-values leads to high false positive rates and invalid conclusions.
  • Ignoring Sample Ratio Mismatch (SRM) - Failing to verify that the control and treatment counts match the intended ratio can hide fundamental technical bugs that invalidate results.
  • Over-optimizing for short-term theatrics - Focusing purely on immediate conversion metrics can mask long-term damage to the user experience or revenue quality.
  • Losing institutional memory - Failing to document past wins and failures leads to teams repeating the same experiments or losing valuable product patterns to organizational churn.

Deep Dive

For all 13 sourced insights from 9 guests, see references/guest-insights.md

  • Customer Interviews
  • Continuous Discovery
  • Idea Validation
  • Defining Icp

© RefoundAI, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (references) in skills/product-experiments of RefoundAI/lenny-skills.

  • SKILL.md
  • references/artifacts.md
  • references/guest-insights.md

Open the folder on GitHubat commit 13598cc

Compare with similar skills

Product Experiments next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Product Experiments compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Product Experiments this skillRefoundAI/lenny-skills1.4k—~2kAutomated safety check: PassMIT
Hypothesis Generationspacering-net/codeg3.8k15 repos~3.6kAutomated safety check: NotesMIT
GitHub Deep Researchbytedance/deer-flow83k5 repos~1.3kAutomated safety check: PassMIT
Nature Paper CardYuan1z0825/nature-skills46k2 repos~2.1kAutomated safety check: PassApache-2.0
Read arXiv Paperkarpathy/nanochat58k2 repos~494Automated safety check: PassMIT
Content Research Writerweapp-tailwindcss/weapp-tailwindcss1.9k25 repos~3.5kAutomated safety check: PassMIT

Similar skills

  • Hypothesis Generation

    spacering-net/codeg

    Structured hypothesis formulation from observations. An agent skill from spacering-net/codeg.

    3.8k GitHub starsUsed in 15 repos~3.6k tokens
    Research & ScienceAuto-check: notes
  • GitHub Deep Research

    bytedance/deer-flow

    Researches a GitHub repository over four rounds using the GitHub API and web search, then writes a structured markdown report with timeline, metrics and Mermaid diagrams.

    83k GitHub starsUsed in 5 repos~1.3k tokens
    Research & ScienceAuto-check passed
  • Nature Paper Card

    Yuan1z0825/nature-skills

    Builds a structured deep-reading card for one scientific paper, covering methods, how experiments support claims, limitations and research ideas, with a script to prepare the source.

    46k GitHub starsUsed in 2 repos~2.1k tokens
    Research & ScienceAuto-check passed
  • Read arXiv Paper

    karpathy/nanochat

    Fetches the TeX source of an arXiv paper from its URL, reads it and writes a markdown summary tied to the nanochat project.

    58k GitHub starsUsed in 2 repos~494 tokens
    Research & ScienceAuto-check passed
  • Content Research Writer

    weapp-tailwindcss/weapp-tailwindcss

    Assists in writing high-quality content by conducting research, adding citations, improving hooks, iterating on outlines, and providing real-time feedback on each section.

    1.9k GitHub starsUsed in 25 repos~3.5k tokens
    Research & ScienceAuto-check passed
  • Peer Review

    spacering-net/codeg

    Structured manuscript/grant review with checklist-based evaluation.

    3.8k GitHub starsUsed in 18 repos~5.9k tokens
    Research & ScienceAuto-check: notes

More from RefoundAI/lenny-skills

All 76 skills in this repo
  • Customer Interviews

    RefoundAI/lenny-skills

    Help users conduct high-impact customer interviews that move beyond surface-level feature requests to identify root emotional frustrations and specific causal triggers.

    1.4k GitHub stars~1.7k tokensUpdated 2 mo ago
    Auto-check passed
  • Time Energy Management

    RefoundAI/lenny-skills

    Help users master their personal output by shifting from reactive scheduling to intentional energy management, internal trigger mastery, and proactive boundary setting.

    1.4k GitHub stars~1.6k tokensUpdated 2 mo ago
    Auto-check passed
  • User Onboarding Activation

    RefoundAI/lenny-skills

    Help users reach their first moment of core value by optimizing the first-run experience, removing friction, and aligning product design with psychological triggers.

    1.4k GitHub stars~2.1k tokensUpdated 2 mo ago
    Auto-check passed
  • Acquisition Channels

    RefoundAI/lenny-skills

    Help users identify unique distribution advantages and master the lifecycle of acquisition channels to build a sustainable engine for growth and retention.

    1.4k GitHub stars~2.2k tokensUpdated 2 mo ago
    Auto-check passed
  • AI Assisted Prototyping

    RefoundAI/lenny-skills

    Help users build functional product prototypes from natural language or visual mocks using AI coding tools.

    1.4k GitHub stars~1.2k tokensUpdated 2 mo ago
    Auto-check passed
  • AI Evals

    RefoundAI/lenny-skills

    Help users build robust infrastructure for measuring, monitoring, and iterating on AI product performance using human, code-based, and LLM-as-a-judge methodologies.

    1.4k GitHub stars~1.7k tokensUpdated 2 mo ago
    Auto-check passed

Questions about Product Experiments

What does Product Experiments do?

Help users design, execute, and analyze product experiments to validate hypotheses and measure true incremental impact while avoiding common statistical pitfalls. Product Experiments is an agent skill from RefoundAI/lenny-skills. Help users design, execute, and analyze product experiments to validate hypotheses and measure true incremental impact while avoiding common statistical pitfalls.

When should I use Product Experiments?

Product Experiments fits situations like: research & Science work in your project.

How do I install Product Experiments in Claude Code?

Run `npx skills add RefoundAI/lenny-skills --skill product-experiments -a claude-code`. Or copy the skill folder (skills/product-experiments in RefoundAI/lenny-skills) into .claude/skills/product-experiments in your project. Claude Code loads it when a task matches its description.

How do I install Product Experiments in Codex?

Run `npx skills add RefoundAI/lenny-skills --skill product-experiments -a codex`. Or copy the skill folder (skills/product-experiments in RefoundAI/lenny-skills) into .agents/skills/product-experiments in your project. Codex loads it when a task matches its description.

Can I use Product Experiments in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add RefoundAI/lenny-skills --skill product-experiments -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/product-experiments, .gemini/skills/product-experiments, .github/skills/product-experiments and .opencode/skills/product-experiments in your project.

What does Product Experiments need to run?

SKILL.md names no scripts, command-line tools or credentials: Product Experiments is instructions for the agent only.

Does Product Experiments access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Product Experiments safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Product Experiments use?

Product Experiments is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Product Experiments use?

About 2k tokens (SKILL.md is roughly 8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 13k tokens, read only when the agent opens those files.

What are the alternatives to Product Experiments?

Skills that share tags, products or a category with Product Experiments: Hypothesis Generation (spacering-net/codeg, 3.8k stars), GitHub Deep Research (bytedance/deer-flow, 83k stars), Nature Paper Card (Yuan1z0825/nature-skills, 46k stars) and Read arXiv Paper (karpathy/nanochat, 58k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Product Experiments?

RefoundAI (a GitHub organization) maintains it in RefoundAI/lenny-skills, which has 1,377 GitHub stars. The repository holds 76 skills in this directory. The repository was last updated on July 16, 2026.

Source: RefoundAI/lenny-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.