Ab Testing
coreyhaines31/marketingskills
When the user wants to plan, design, or implement an A/B test or experiment, or build a growth experimentation program.
Create, update (metrics, dates, description, hypothesis, baseline/comparison treatments, status), or delete Harness FME feature-flag experiments.
$ npx skills add harness/harness-skills --skill manage-experiments -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install harness/harness-skills manage-experiments --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/harness/harness-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/manage-experiments .claude/skills/manage-experiments && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "manage-experiments" agent skill from https://github.com/harness/harness-skills/tree/main/skills/manage-experiments into .claude/skills/manage-experiments/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "manage-experiments", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/harness/harness-skills/tree/main/skills/manage-experimentsType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add harness/harness-skills --skill manage-experiments -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install harness/harness-skills manage-experiments --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/harness/harness-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/manage-experiments .agents/skills/manage-experiments && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "manage-experiments" agent skill from https://github.com/harness/harness-skills/tree/main/skills/manage-experiments into .agents/skills/manage-experiments/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "manage-experiments", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add harness/harness-skills --skill manage-experiments -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install harness/harness-skills manage-experiments --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/harness/harness-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/manage-experiments .cursor/skills/manage-experiments && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "manage-experiments" agent skill from https://github.com/harness/harness-skills/tree/main/skills/manage-experiments into .cursor/skills/manage-experiments/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "manage-experiments", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/harness/harness-skills.git --path skills/manage-experiments--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add harness/harness-skills --skill manage-experiments -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install harness/harness-skills manage-experiments --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/harness/harness-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/manage-experiments .gemini/skills/manage-experiments && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "manage-experiments" agent skill from https://github.com/harness/harness-skills/tree/main/skills/manage-experiments into .gemini/skills/manage-experiments/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "manage-experiments", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install harness/harness-skills manage-experimentsInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add harness/harness-skills --skill manage-experiments -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/harness/harness-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/manage-experiments .github/skills/manage-experiments && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "manage-experiments" agent skill from https://github.com/harness/harness-skills/tree/main/skills/manage-experiments into .github/skills/manage-experiments/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "manage-experiments", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add harness/harness-skills --skill manage-experiments -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install harness/harness-skills manage-experiments --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/harness/harness-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/manage-experiments .opencode/skills/manage-experiments && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "manage-experiments" agent skill from https://github.com/harness/harness-skills/tree/main/skills/manage-experiments into .opencode/skills/manage-experiments/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "manage-experiments", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
manage-experimentsCreate, update (metrics, dates, description, hypothesis, baseline/comparison treatments, status), or delete Harness FME feature-flag experiments.
Manage Experiments is an agent skill from harness/harness-skills. Create, update (metrics, dates, description, hypothesis, baseline/comparison treatments, status), or delete Harness FME feature-flag experiments. Delegates metric selection to choose-metric, metric creation to create-metric, event wiring to instrument-metric, results to review-experiment-results, flag/definition prep to create-feature-flag and update-flag-targeting. Use when asked to create, modify, pause, resume, complete, or delete an experiment. Do not use for explaining results (review-experiment-results)…
Its SKILL.md is about 3.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including reference files (for example `references/stop-conditions.md`). Compatibility notes: Requires the Harness MCP server or the Harness CLI
It sits in Marketing & SEO, covering A/B testing. The repository describes itself as: A collection of structured AI agent skills that enable Claude Code, Cursor, GitHub Copilot, and other AI coding assistants to create, operate, debug, and govern Harness CI/CD… The licence is Apache-2.0.
2 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit c25faee. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md.
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Requires the Harness MCP server or the Harness CLI
From compatibility in the SKILL.md frontmatter.
Manage Experiments loads about 3.2k tokens when it runs, and up to ~3.6k if it reads all its reference files. Until then it costs about 173 tokens; SKILL.md has 1,299 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from harness/harness-skills at commit c25faee, republished under its Apache-2.0 licence (© harness). 1,299 words, ~3,184 tokens.
.claude/skills/manage-experiments/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.Create, update, or delete FEATURE_FLAG experiments. AI_CONFIG mutation workflows are not implemented here; use read-only results inspection or a separately supported workflow. Orchestrates decisions and delegates metric selection, flag setup, and results interpretation to specialist skills.
Related: choose-metric (metric selection), create-metric (metric creation), instrument-metric (event wiring), review-experiment-results (results), create-feature-flag, update-flag-targeting.
Works through the Harness MCP server or the Harness CLI; names are from tool-map.md.
| Operation | MCP | CLI |
|---|---|---|
| List experiments | harness_list · fme_experiment · filters: { parent_type, parent_name?, environment_id?, status?: ["ACTIVE", "PAUSED"], … } · compact: false | harness list experiment --parent-type FEATURE_FLAG [--search <name>] [--status ACTIVE], then again with --status PAUSED |
| Get experiment | harness_get · fme_experiment · params: { experiment_id } | harness get experiment <experiment-id> |
| Create experiment | harness_create · fme_experiment · params: { environment_id } · body: { parent, name, startAt, endAt, baselineTreatment, comparisonTreatments, rule: "default rule", owners?, … } | harness create experiment <name> --env <env-id> -f experiment.json --json |
| Update experiment | harness_update · fme_experiment · params: { experiment_id } · body: { description?, hypothesis?, startAt?, endAt?, keyMetrics?, status?, rule?, … } | harness update experiment <experiment-id> --set description=foo --set status=PAUSED |
| Delete experiment | harness_delete · fme_experiment · params: { experiment_id } | harness delete experiment <experiment-id> |
| List environments | harness_list · fme_environment · compact: false | harness list fme_environment --json |
| Get definition | harness_get · fme_feature_flag_definition · params: { feature_flag_name, environment_id } | harness get feature_flag:definition <flag-name> --env <env-id> |
Follow scope-establishment.md.
Create: new experiment on a feature flag.
Update: change description, hypothesis, dates, metrics, baseline/comparison
treatments, or status (pause/resume/complete/archive).
Delete: hard delete (permanent, no archive/restore).
Default parent type: FEATURE_FLAG. If flag doesn't exist, route to create-feature-flag, then return here. List environments; ask which environment if not stated.
Stop for AI_CONFIG or CONFIG mutations. The declared tools cannot verify their parent definitions, treatment names or traffic readiness. Do not substitute a feature-flag lookup or user guesses for that check. Read-only experiment/results inspection remains possible; ask for a supported AI Config workflow before writing.
Get definition for the flag in the chosen environment. 404 = no definition; route to update-flag-targeting to initialize it, then return here.
Read treatments[].name. Experiment's baselineTreatment and comparisonTreatments must match real treatment names. Default baseline = definition's baselineTreatment, but confirm with user. Fewer than 2 treatments → stop (see stop-conditions.md).
Check traffic readiness: isKilled: false, trafficAllocation > 0, baseline and each comparison treatment have size > 0 in defaultRule buckets (or in the targeting rule the experiment will use). If any check fails, warn and ask whether to proceed. Fixing it requires update-flag-targeting.
List experiments for the parent flag and environment, with ACTIVE and PAUSED status. Apply the experiment check: ACTIVE means confirm the user wants a second experiment; PAUSED means warn and require explicit acknowledgement.
Get hypothesis: "if [change], then [metric] will increase/decrease because [reasoning]" (max 500 chars). Link concepts.md. Missing or non-causal hypothesis → stop (see stop-conditions.md).
Hand off to choose-metric to select key and supporting metrics. If no suitable metric exists, route to create-metric (and instrument-metric if event missing). Return here with metric IDs once complete.
startAt/endAt are required ISO-8601 timestamps. Ask explicitly; no default duration.
Present full payload: parent (type, name), name (2-250 chars, unique), description, hypothesis, startAt/endAt (ISO-8601), baselineTreatment, comparisonTreatments, keyMetrics, supportingMetrics, rule (set to "default rule" unless the user wants results from a specific targeting rule label; results only count impressions whose label matches the experiment's rule), owners, tags. Link write-safety.md. Do not send assignmentSource (400 if present).
CLI only: --env <env-id> is a required flag even when using -f - it sets the environment_id query param, which the request body never carries. rule and owners have no dedicated create flags (only --parent-type, --parent-name/--parent-id, --description, --hypothesis, --start-at, --end-at, --baseline-treatment, --comparison-treatment, --key-metric, --supporting-metric exist) - put them in the -f experiment.json body instead; don't invent flags for them. STOP HERE. Wait for explicit confirmation.
Create experiment with confirmed payload. 409 = duplicate name; report the conflict and ask the user for a new name (never rename silently). 404 = parent doesn't exist in environment; back to Step 1.
Get experiment by returned id. Compare name, hypothesis, startAt/endAt, treatments, key/supporting metric IDs, rule, parent/environment, owners and tags with the approved draft. Stop and report any substantive mismatch; never automatically recreate/delete to correct immutable scope. Report status (typically ACTIVE on create). Hand off to review-experiment-results once data collected.
If given an exact ID, Get experiment directly. Otherwise fully paginate name discovery across ACTIVE, PAUSED, COMPLETED and ARCHIVED (one CLI status per call), then ask on ambiguity. Require parent type FEATURE_FLAG before mutation. Show current status, description, hypothesis, startAt, endAt, baselineTreatment, comparisonTreatments, keyMetrics, supportingMetrics, rule.
Ask what to update: description, hypothesis, dates, baseline/comparison treatments, metrics, status, rule, owners, tags. For metrics, route to choose-metric if selecting new ones.
When updating baselineTreatment or comparisonTreatments, verify the treatments exist in the flag definition for the experiment's environment (Get definition for that environment and check treatments[].name, same as create Step 2). See stop-conditions.md.
Valid status values: ACTIVE, PAUSED, COMPLETED, ARCHIVED (null never allowed). The backend enforces which transitions are actually legal - don't assume a fixed chain (e.g. ACTIVE → PAUSED → COMPLETED); send the requested target status and let a 400 reveal an invalid transition. ARCHIVED is a status value here, not a substitute for delete. Changing metrics or dates on ACTIVE experiment → explicit warning.
Clearable (via API/MCP merge patch): description, hypothesis, rule, keyMetrics, supportingMetrics, tags. Non-clearable: name, startAt, endAt, baselineTreatment, comparisonTreatments, status.
CLI gap: the CLI's update experiment has no -f/file-body option and no mutable field for name or rule - scalar fields use their declared snake_case --set IDs; collections such as tags/owners/metric references use their declared --add/--del handlers, not arbitrary arrays. To change rule or name, stop or propose available MCP with explicit approval; never silently drop the requested field.
Show diff (current vs. new). Link write-safety.md. If ACTIVE experiment and changing key fields, add explicit warning.
Update experiment with merge patch body. 409 = duplicate name. 400 = invalid field, null on non-clearable, or an illegal status transition.
Get experiment again. Compare updated fields, including rule if it was part of this update. Report new status if changed.
Resolve exact ID or fully paginated name/status discovery as in Update Step 1. Require FEATURE_FLAG before mutation. Show name, status, parent, environment, dates.
Hard delete = permanent, no archive/restore. Stricter confirmation required. Offer status update to COMPLETED instead (update operation). Link write-safety.md.
Delete experiment. Confirm deletion; parent flag not affected.
| Issue | Resolution |
|---|---|
| 404 on create even though flag exists | Parent must exist in chosen environment; check Get definition for that environment |
| User wants to change treatments | Route to update-flag-targeting to modify definition, then return here |
| Experiment needs workspace guardrail metric | Guardrails apply automatically; never set via keyMetrics/supportingMetrics |
| 400 on update with null | Field is non-clearable; omit it to leave unchanged |
| Empty keyMetrics on create | Allowed but no winner criterion; warn and confirm before Step 6 draft |
| ACTIVE experiment blocks targeting change | See write-safety.md; require explicit acknowledgement |
User wants to change rule or name via CLI | Not supported - CLI update experiment has no field/-f for either; use MCP or say so |
Parent is AI_CONFIG | Mutation validation is unsupported here; stop and request a supported AI Config workflow. Do not use the feature-flag definition route |
| 400 on status update | Transition not allowed by the backend for the experiment's current status; report the error, don't retry with a guessed intermediate status |
© harness, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 1 other file (references) in skills/manage-experiments of harness/harness-skills.
Open the folder on GitHubat commit c25faee
Manage Experiments next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Manage Experiments this skillharness/harness-skills | 115 | — | ~3.2k | Automated safety check: Pass | Apache-2.0 | |
| Ab Testingcoreyhaines31/marketingskills | 54k | 3 repos | ~2.8k | Automated safety check: Pass | MIT | |
| AnalyticsNexus-JPF/note-companion | 870 | 7 repos | ~2.2k | Automated safety check: Pass | MIT | |
| Ab Test Setupfreekmurze/dotfiles | 1k | 15 repos | ~1.8k | Automated safety check: Pass | None | |
| Ad Test Designeraaron-he-zhu/aaron-marketing-skills | 2.9k | 2 repos | ~2.8k | Automated safety check: Pass | Apache-2.0 | |
| Meta Tags Optimizernowork-studio/notfair-plugin | 3.9k | 1 repos | ~2.7k | Automated safety check: Pass | MIT |
coreyhaines31/marketingskills
When the user wants to plan, design, or implement an A/B test or experiment, or build a growth experimentation program.
Nexus-JPF/note-companion
When the user wants to set up, improve, or audit analytics tracking and measurement.
freekmurze/dotfiles
When the user wants to plan, design, or implement an A/B test or experiment.
aaron-he-zhu/aaron-marketing-skills
A skill your agent uses when the user asks to "design an A/B test", "set up a creative/landing test", "run an incrementality test", or "is this result statistically and practically material?"…
nowork-studio/notfair-plugin
Writes and improves title tags, meta descriptions, Open Graph and Twitter card tags for click-through, with character counts and A/B test variants.
irinabuht12-oss/marketing-skills
Statistical significance calculator for A/B test results with sample size requirements, segment breakdowns, and hypothesis generation.
harness/harness-skills
Generate audit reports and compliance trails using Harness audit trail data via MCP v2 tools.
harness/harness-skills
A skill your agent uses when working with Chaos Engineering steps inside a Harness pipeline.
harness/harness-skills
A skill your agent uses when the user asks to create, edit, update, design, or configure a Harness Chaos Experiment — including faults, probes, actions, experiment YAML, fault injection, pod-delete…
harness/harness-skills
Remove a launched Harness FME feature flag from application code, keeping the treatment FME serves today, and open a pull request.
harness/harness-skills
Configure code scanning in Harness pipelines using STO security scanners.
harness/harness-skills
Generate Harness Agent Template files for AI-powered automation agents.
Categories
Create, update (metrics, dates, description, hypothesis, baseline/comparison treatments, status), or delete Harness FME feature-flag experiments. Manage Experiments is an agent skill from harness/harness-skills. Create, update (metrics, dates, description, hypothesis, baseline/comparison treatments, status), or delete Harness FME feature-flag experiments.
Manage Experiments fits situations like: asked to create; delete an experiment; explaining results (review-experiment-results); phrases: create experiment.
Run `npx skills add harness/harness-skills --skill manage-experiments -a claude-code`. Or copy the skill folder (skills/manage-experiments in harness/harness-skills) into .claude/skills/manage-experiments in your project. Claude Code loads it when a task matches its description.
Run `npx skills add harness/harness-skills --skill manage-experiments -a codex`. Or copy the skill folder (skills/manage-experiments in harness/harness-skills) into .agents/skills/manage-experiments in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add harness/harness-skills --skill manage-experiments -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/manage-experiments, .gemini/skills/manage-experiments, .github/skills/manage-experiments and .opencode/skills/manage-experiments in your project.
SKILL.md names no scripts, command-line tools or credentials: Manage Experiments is instructions for the agent only. Compatibility (from SKILL.md): Requires the Harness MCP server or the Harness CLI.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Manage Experiments is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 3.2k tokens (SKILL.md is roughly 13k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 389 tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Manage Experiments: Ab Testing (coreyhaines31/marketingskills, 54k stars), Analytics (Nexus-JPF/note-companion, 870 stars), Ab Test Setup (freekmurze/dotfiles, 1k stars) and Ad Test Designer (aaron-he-zhu/aaron-marketing-skills, 2.9k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
harness (a GitHub organization) maintains it in harness/harness-skills, which has 115 GitHub stars. The repository holds 24 skills in this directory. The repository was last updated on October 6, 2026.
Source: harness/harness-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.