Experimentation
andreaskelm/pm-brain
Design and run product experiments at a practical PM level — A/B tests, hypothesis tests, rollouts, feature flags, and reading results without pretending to be a statistician.
Run end-to-end product experiments from assumption to decision: translate assumptions into testable hypotheses and experiment briefs, select the right method among qualitative interviews…
$ npx skills add magnus919/agent-skills --skill product-experimentation -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install magnus919/agent-skills product-experimentation --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/magnus919/agent-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/product-experimentation .claude/skills/product-experimentation && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "product-experimentation" agent skill from https://github.com/magnus919/agent-skills/tree/main/product-experimentation into .claude/skills/product-experimentation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "product-experimentation", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/magnus919/agent-skills/tree/main/product-experimentationType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add magnus919/agent-skills --skill product-experimentation -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install magnus919/agent-skills product-experimentation --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/magnus919/agent-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/product-experimentation .agents/skills/product-experimentation && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "product-experimentation" agent skill from https://github.com/magnus919/agent-skills/tree/main/product-experimentation into .agents/skills/product-experimentation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "product-experimentation", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add magnus919/agent-skills --skill product-experimentation -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install magnus919/agent-skills product-experimentation --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/magnus919/agent-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/product-experimentation .cursor/skills/product-experimentation && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "product-experimentation" agent skill from https://github.com/magnus919/agent-skills/tree/main/product-experimentation into .cursor/skills/product-experimentation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "product-experimentation", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/magnus919/agent-skills.git --path product-experimentation--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add magnus919/agent-skills --skill product-experimentation -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install magnus919/agent-skills product-experimentation --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/magnus919/agent-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/product-experimentation .gemini/skills/product-experimentation && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "product-experimentation" agent skill from https://github.com/magnus919/agent-skills/tree/main/product-experimentation into .gemini/skills/product-experimentation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "product-experimentation", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install magnus919/agent-skills product-experimentationInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add magnus919/agent-skills --skill product-experimentation -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/magnus919/agent-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/product-experimentation .github/skills/product-experimentation && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "product-experimentation" agent skill from https://github.com/magnus919/agent-skills/tree/main/product-experimentation into .github/skills/product-experimentation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "product-experimentation", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add magnus919/agent-skills --skill product-experimentation -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install magnus919/agent-skills product-experimentation --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/magnus919/agent-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/product-experimentation .opencode/skills/product-experimentation && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "product-experimentation" agent skill from https://github.com/magnus919/agent-skills/tree/main/product-experimentation into .opencode/skills/product-experimentation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "product-experimentation", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
product-experimentationRun end-to-end product experiments from assumption to decision: translate assumptions into testable hypotheses and experiment briefs, select the right method among qualitative interviews…
Product Experimentation is an agent skill from magnus919/agent-skills. Run end-to-end product experiments from assumption to decision: translate assumptions into testable hypotheses and experiment briefs, select the right method among qualitative interviews, prototypes, concierge tests, fake doors, feature flags, and A/B tests, and produce readouts that update the roadmap and decision record. Do not use when a qualitative or prototype test is the clearly right answer without statistical measurement; do not prescribe A/B testing by default; do not treat statistical significance as…
Its SKILL.md is about 2.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 13 other files, including reference files (for example `README.md`, `evals/evals.json` and `references/discovery-brief.md`). Compatibility notes: Agent-agnostic — works with any agent framework supporting the Agent Skills format. No external services, proprietary tools, or runtime dependencies required.
It sits in Marketing & SEO, covering A/B testing. The repository describes itself as: Curated collection of AI agent skills for Hermes and other agent frameworks. The licence is MIT.
8 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 22b4723. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md.
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Agent-agnostic — works with any agent framework supporting the Agent Skills format. No external services, proprietary tools, or runtime dependencies required.
From compatibility in the SKILL.md frontmatter.
Product Experimentation loads about 2.6k tokens when it runs, and up to ~7.3k if it reads all its reference files. Until then it costs about 153 tokens; SKILL.md has 925 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from magnus919/agent-skills at commit 22b4723, republished under its MIT licence (© magnus919). 925 words, ~2,571 tokens.
.claude/skills/product-experimentation/SKILL.md (or your agent's skills folder). This skill also uses 10 other files; get the full folder from GitHub.End-to-end product experimentation: from assumption mapping through method selection, instrumentation, guardrail enforcement, and decision-readout that updates the product roadmap. Owns the complete experiment workflow; routes statistical design and rollout mechanics to specialist skills.
ASSUMPTIONS → [HYPOTHESIS] → [METHOD SELECT] → [INSTRUMENT] → [RUN] → [DECIDE] → [RECORD]
| | | | | |
Experiment Qualitative Tracking Guardrail Decision Readout
brief Prototype plan monitor rules learning
Operational
QuantitativeLoad only the reference or template relevant to the task. Do not load every file at once.
| File | Load when |
|---|---|
| references/discovery-brief.md | You need to understand how experimentation concepts map across skills and where this skill's boundaries are |
| references/method-selection.md | Choosing among qualitative, prototype, operational, and quantitative test methods |
| references/guardrails-and-ethics.md | Defining guardrail metrics, ethical boundaries, stopping rules, and decision ownership |
| references/experiment-readout.md | Producing a decision-impact readout that updates the roadmap or decision record |
| templates/experiment-brief.md | Filling out a structured experiment brief from an assumption |
| templates/assumption-map.md | Mapping assumptions to risk, evidence, and testability before designing experiments |
| templates/guardrail-and-decision-rule.md | Recording guardrails, stopping rules, and decision criteria for an experiment |
| templates/readout-learning-entry.md | Documenting experiment outcome and updating the roadmap, decision log, or lifecycle evidence |
Surface the assumptions driving the proposed change. Classify each by risk (what breaks if it is wrong), evidence strength (what evidence already exists), and testability (can it be tested, and how cheaply). Use templates/assumption-map.md.
Convert the riskiest, least-evidenced assumptions into falsifiable hypotheses. Each hypothesis names the independent variable (what changes), the dependent variable (what outcome is measured), the predicted direction, and the smallest effect that matters. Use templates/experiment-brief.md.
Choose the lightest-weight method that can falsify the hypothesis with sufficient confidence. The method ladder, from lightest to heaviest:
| Method | Best for | Cost | Statistical rigor |
|---|---|---|---|
| Qualitative interviews | Uncovering unknown unknowns, mental models, problem validation | Lowest | None (descriptive) |
| Prototype tests | Interaction flow, usability, concept validation | Low | None (observational) |
| Concierge tests | Value delivery, willingness to pay, operational feasibility | Low-Medium | None (manual) |
| Fake doors | Demand signals, willingness to click/commit | Medium | Low (conversion rate only) |
| Feature flags | Operational safety, incremental rollout, kill-switch | Medium | Medium (controlled rollout) |
| A/B tests | Causal attribution of a specific change to a metric | High | High (randomized controlled) |
Do not default to A/B testing. Start at the top of the ladder and only move down when the question cannot be answered at the current level. A qualitative interview or prototype test is often the right answer. Full method selection guidance is in references/method-selection.md.
Before running the experiment, define:
Use templates/guardrail-and-decision-rule.md to record these.
Define the target population, allocation, and minimum detectable effect. Route statistical design (power analysis, sample-size calculation, estimator selection) to ../data-scientist/SKILL.md. An underpowered experiment — one that cannot detect the smallest effect that matters — is a validity failure; do not ship based on a null result from an underpowered test.
Execute the experiment. Monitor guardrails continuously. Route production rollout mechanics (feature flags, canary stages, progressive delivery) to ../release-engineering/SKILL.md.
Make the ship/no-ship decision using multiple criteria, never statistical significance alone:
| Criterion | Weight | Source |
|---|---|---|
| Statistical evidence | Required | data-scientist |
| Practical significance | Required | Is the effect large enough to matter? |
| Guardrail evidence | Blocking | All guardrails must pass |
| Qualitative evidence | Informative | User feedback, support tickets |
| Reversibility | Informative | Can we undo this if wrong? |
| Opportunity cost | Informative | What else could we build instead? |
A statistically significant result with a failing guardrail is a no-ship. A statistically significant result that exceeds authority boundaries (e.g., safety, compliance, ethics) is a no-ship. Record the decision and its rationale.
Document what was learned and what changed as a result. The readout updates the product roadmap, backlog, decision log, or lifecycle evidence. Routing: feeds product-roadmapping-and-portfolio (roadmap updates), product-adoption (adoption evidence), and product-lifecycle-learning (retained learning). Use templates/readout-learning-entry.md and references/experiment-readout.md.
Load this skill when:
This skill is intentionally host-neutral. It requires no profile system, output format, scripts, or external services. Load references and templates directly by path using the host agent's normal file-loading mechanism.
© magnus919, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 10 other files (references) in product-experimentation of magnus919/agent-skills.
Open the folder on GitHubat commit 22b4723
Product Experimentation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Product Experimentation this skillmagnus919/agent-skills | 115 | — | ~2.6k | Automated safety check: Pass | MIT | |
| Experimentationandreaskelm/pm-brain | 234 | — | ~2.3k | Automated safety check: Pass | Custom licence | |
| Experiment Designrampstackco/claude-skills | 945 | — | ~7.9k | Automated safety check: Pass | MIT | |
| Ab Test Results Readouthashgraph-online/awesome-codex-plugins | 1.3k | — | ~859 | Automated safety check: Pass | MIT | |
| Ab Testingcoreyhaines31/marketingskills | 54k | 3 repos | ~3.1k | Automated safety check: Pass | MIT | |
| AnalyticsNexus-JPF/note-companion | 870 | 7 repos | ~2.2k | Automated safety check: Pass | MIT |
andreaskelm/pm-brain
Design and run product experiments at a practical PM level — A/B tests, hypothesis tests, rollouts, feature flags, and reading results without pretending to be a statistician.
rampstackco/claude-skills
A discipline for designing experiments (A/B tests, multivariate, holdouts) so the results actually answer the question you asked.
hashgraph-online/awesome-codex-plugins
Analyze and communicate A/B test results with metric readouts, subgroup analysis, data-quality checks, ad hoc investigation, visualization, and launch recommendations.
coreyhaines31/marketingskills
When the user wants to plan, design, or implement an A/B test or experiment, or build a growth experimentation program.
Nexus-JPF/note-companion
When the user wants to set up, improve, or audit analytics tracking and measurement.
aaron-he-zhu/aaron-marketing-skills
A skill your agent uses when the user asks to "design an A/B test", "set up a creative/landing test", "run an incrementality test", or "is this result statistically and practically material?"…
magnus919/agent-skills
Organize durable agent research outputs as summaries, analysis, and evidence dossiers.
magnus919/agent-skills
Build portable, first-person colored ASCII city engines and small GIS-derived city packs.
magnus919/agent-skills
Manage color workflows with ICC profiles, working spaces, gamut mapping, and color science.
magnus919/agent-skills
A skill your agent uses for PhD-level expertise in data science, statistics, and machine learning: rigorous statistical analysis, experimental design, causal inference, advanced modeling, research…
magnus919/agent-skills
Use Docker Compose to define, run, debug, and harden multi-container applications.
magnus919/agent-skills
Design, review, simulate, and verify FPGA logic using explicit RTL contracts, clock and reset models, CDC analysis, timing constraints, and reproducible implementation evidence.
Categories
Run end-to-end product experiments from assumption to decision: translate assumptions into testable hypotheses and experiment briefs, select the right method among qualitative interviews…. Product Experimentation is an agent skill from magnus919/agent-skills. Run end-to-end product experiments from assumption to decision: translate assumptions into testable hypotheses and experiment briefs, select the right method among qualitative interviews, prototypes, concierge tests, fake doors, feature flags, and A/B tests, and produce readouts that update the roadmap and decision record.
Product Experimentation fits situations like: prototype test is the clearly right answer without statistical measurement; do not prescribe A/B testing by default; do not treat statistical significance as the only decision criterion; hide ethical and guardrail considerations.
Run `npx skills add magnus919/agent-skills --skill product-experimentation -a claude-code`. Or copy the skill folder (product-experimentation in magnus919/agent-skills) into .claude/skills/product-experimentation in your project. Claude Code loads it when a task matches its description.
Run `npx skills add magnus919/agent-skills --skill product-experimentation -a codex`. Or copy the skill folder (product-experimentation in magnus919/agent-skills) into .agents/skills/product-experimentation in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add magnus919/agent-skills --skill product-experimentation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/product-experimentation, .gemini/skills/product-experimentation, .github/skills/product-experimentation and .opencode/skills/product-experimentation in your project.
SKILL.md names no scripts, command-line tools or credentials: Product Experimentation is instructions for the agent only. Compatibility (from SKILL.md): Agent-agnostic — works with any agent framework supporting the Agent Skills format. No external services, proprietary tools, or runtime dependencies required..
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Product Experimentation is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.6k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 4.7k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Product Experimentation: Experimentation (andreaskelm/pm-brain, 234 stars), Experiment Design (rampstackco/claude-skills, 945 stars), Ab Test Results Readout (hashgraph-online/awesome-codex-plugins, 1.3k stars) and Ab Testing (coreyhaines31/marketingskills, 54k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
magnus919 (a GitHub user) maintains it in magnus919/agent-skills, which has 115 GitHub stars. The repository holds 131 skills in this directory. The repository was last updated on October 10, 2026.
Source: magnus919/agent-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.