Agent skill

Manage Experiments

by harness in harness/harness-skills

Create, update (metrics, dates, description, hypothesis, baseline/comparison treatments, status), or delete Harness FME feature-flag experiments.

Apache-2.0Auto-check passedMarketing & SEO

Install Manage Experiments

skills CLI
$ npx skills add harness/harness-skills --skill manage-experiments -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install harness/harness-skills manage-experiments --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/harness/harness-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/manage-experiments .claude/skills/manage-experiments && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
manage-experiments
GitHub stars
115
Token cost
~3.2k tokens
SKILL.md length
1,299 words
Files
2 (incl. references)
Skills in repo
24
Repo updated
First seen
Licence
Apache-2.0

At a glance

Create, update (metrics, dates, description, hypothesis, baseline/comparison treatments, status), or delete Harness FME feature-flag experiments.

  • Works in 2 steps: Establish scope → Choose operation
  • Asked to create
  • SKILL.md covers Tools, Instructions, Examples and Performance Notes, plus 1 more section
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Manage Experiments is an agent skill from harness/harness-skills. Create, update (metrics, dates, description, hypothesis, baseline/comparison treatments, status), or delete Harness FME feature-flag experiments. Delegates metric selection to choose-metric, metric creation to create-metric, event wiring to instrument-metric, results to review-experiment-results, flag/definition prep to create-feature-flag and update-flag-targeting. Use when asked to create, modify, pause, resume, complete, or delete an experiment. Do not use for explaining results (review-experiment-results)…

Its SKILL.md is about 3.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including reference files (for example `references/stop-conditions.md`). Compatibility notes: Requires the Harness MCP server or the Harness CLI

It sits in Marketing & SEO, covering A/B testing. The repository describes itself as: A collection of structured AI agent skills that enable Claude Code, Cursor, GitHub Copilot, and other AI coding assistants to create, operate, debug, and govern Harness CI/CD… The licence is Apache-2.0.

When your agent uses it

  • Asked to create
  • Delete an experiment
  • Explaining results (review-experiment-results)
  • Phrases: create experiment

Example prompts

  • “/manage-experiments”

Requirements

  • Compatibility (from SKILL.md): Requires the Harness MCP server or the Harness CLI

Workflow steps

2 steps, taken from the step headings in SKILL.md.

  1. Establish scope
  2. Choose operation

What it can do on your machine

Read from SKILL.md and the folder at commit c25faee. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Requires the Harness MCP server or the Harness CLI

    From compatibility in the SKILL.md frontmatter.

Context cost

Manage Experiments loads about 3.2k tokens when it runs, and up to ~3.6k if it reads all its reference files. Until then it costs about 173 tokens; SKILL.md has 1,299 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~173
When it runs · the whole SKILL.md, loaded when a task matches
~3.2k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~3.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from harness/harness-skills at commit c25faee, republished under its Apache-2.0 licence (© harness). 1,299 words, ~3,184 tokens.

Download SKILL.mdSave it as .claude/skills/manage-experiments/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
manage-experiments
description
Create, update (metrics, dates, description, hypothesis, baseline/comparison treatments, status), or delete Harness FME feature-flag experiments. Delegates metric selection to choose-metric, metric creation to create-metric, event wiring to instrument-metric, results to review-experiment-results, flag/definition prep to create-feature-flag and update-flag-targeting. Use when asked to create, modify, pause, resume, complete, or delete an experiment. Do not use for explaining results (review-experiment-results). Trigger phrases: create experiment, set up A/B test, design experiment, launch test, update experiment, pause/resume/complete experiment, delete experiment.
compatibility
Requires the Harness MCP server or the Harness CLI
metadata.author
Harness
metadata.version
2.1.0
metadata.mcp-server
harness-mcp
license
Apache-2.0

Manage Experiments

Create, update, or delete FEATURE_FLAG experiments. AI_CONFIG mutation workflows are not implemented here; use read-only results inspection or a separately supported workflow. Orchestrates decisions and delegates metric selection, flag setup, and results interpretation to specialist skills.

Related: choose-metric (metric selection), create-metric (metric creation), instrument-metric (event wiring), review-experiment-results (results), create-feature-flag, update-flag-targeting.

Tools

Works through the Harness MCP server or the Harness CLI; names are from tool-map.md.

OperationMCPCLI
List experimentsharness_list · fme_experiment · filters: { parent_type, parent_name?, environment_id?, status?: ["ACTIVE", "PAUSED"], … } · compact: falseharness list experiment --parent-type FEATURE_FLAG [--search <name>] [--status ACTIVE], then again with --status PAUSED
Get experimentharness_get · fme_experiment · params: { experiment_id }harness get experiment <experiment-id>
Create experimentharness_create · fme_experiment · params: { environment_id } · body: { parent, name, startAt, endAt, baselineTreatment, comparisonTreatments, rule: "default rule", owners?, … }harness create experiment <name> --env <env-id> -f experiment.json --json
Update experimentharness_update · fme_experiment · params: { experiment_id } · body: { description?, hypothesis?, startAt?, endAt?, keyMetrics?, status?, rule?, … }harness update experiment <experiment-id> --set description=foo --set status=PAUSED
Delete experimentharness_delete · fme_experiment · params: { experiment_id }harness delete experiment <experiment-id>
List environmentsharness_list · fme_environment · compact: falseharness list fme_environment --json
Get definitionharness_get · fme_feature_flag_definition · params: { feature_flag_name, environment_id }harness get feature_flag:definition <flag-name> --env <env-id>

Instructions

Phase 1: Establish scope

Follow scope-establishment.md.

Phase 2: Choose operation

Create: new experiment on a feature flag. Update: change description, hypothesis, dates, metrics, baseline/comparison treatments, or status (pause/resume/complete/archive).
Delete: hard delete (permanent, no archive/restore).


Create experiment
Step 1: Resolve parent and environment

Default parent type: FEATURE_FLAG. If flag doesn't exist, route to create-feature-flag, then return here. List environments; ask which environment if not stated.

Stop for AI_CONFIG or CONFIG mutations. The declared tools cannot verify their parent definitions, treatment names or traffic readiness. Do not substitute a feature-flag lookup or user guesses for that check. Read-only experiment/results inspection remains possible; ask for a supported AI Config workflow before writing.

Step 2: Confirm treatments against live definition

Get definition for the flag in the chosen environment. 404 = no definition; route to update-flag-targeting to initialize it, then return here.

Read treatments[].name. Experiment's baselineTreatment and comparisonTreatments must match real treatment names. Default baseline = definition's baselineTreatment, but confirm with user. Fewer than 2 treatments → stop (see stop-conditions.md).

Check traffic readiness: isKilled: false, trafficAllocation > 0, baseline and each comparison treatment have size > 0 in defaultRule buckets (or in the targeting rule the experiment will use). If any check fails, warn and ask whether to proceed. Fixing it requires update-flag-targeting.

List experiments for the parent flag and environment, with ACTIVE and PAUSED status. Apply the experiment check: ACTIVE means confirm the user wants a second experiment; PAUSED means warn and require explicit acknowledgement.

Step 3: Hypothesis

Get hypothesis: "if [change], then [metric] will increase/decrease because [reasoning]" (max 500 chars). Link concepts.md. Missing or non-causal hypothesis → stop (see stop-conditions.md).

Step 4: Metrics

Hand off to choose-metric to select key and supporting metrics. If no suitable metric exists, route to create-metric (and instrument-metric if event missing). Return here with metric IDs once complete.

Step 5: Experiment window

startAt/endAt are required ISO-8601 timestamps. Ask explicitly; no default duration.

Step 6: Draft and confirm

Present full payload: parent (type, name), name (2-250 chars, unique), description, hypothesis, startAt/endAt (ISO-8601), baselineTreatment, comparisonTreatments, keyMetrics, supportingMetrics, rule (set to "default rule" unless the user wants results from a specific targeting rule label; results only count impressions whose label matches the experiment's rule), owners, tags. Link write-safety.md. Do not send assignmentSource (400 if present).

CLI only: --env <env-id> is a required flag even when using -f - it sets the environment_id query param, which the request body never carries. rule and owners have no dedicated create flags (only --parent-type, --parent-name/--parent-id, --description, --hypothesis, --start-at, --end-at, --baseline-treatment, --comparison-treatment, --key-metric, --supporting-metric exist) - put them in the -f experiment.json body instead; don't invent flags for them. STOP HERE. Wait for explicit confirmation.

Step 7: Create

Create experiment with confirmed payload. 409 = duplicate name; report the conflict and ask the user for a new name (never rename silently). 404 = parent doesn't exist in environment; back to Step 1.

Step 8: Verify

Get experiment by returned id. Compare name, hypothesis, startAt/endAt, treatments, key/supporting metric IDs, rule, parent/environment, owners and tags with the approved draft. Stop and report any substantive mismatch; never automatically recreate/delete to correct immutable scope. Report status (typically ACTIVE on create). Hand off to review-experiment-results once data collected.


Update experiment
Step 1: Get current experiment

If given an exact ID, Get experiment directly. Otherwise fully paginate name discovery across ACTIVE, PAUSED, COMPLETED and ARCHIVED (one CLI status per call), then ask on ambiguity. Require parent type FEATURE_FLAG before mutation. Show current status, description, hypothesis, startAt, endAt, baselineTreatment, comparisonTreatments, keyMetrics, supportingMetrics, rule.

Show full SKILL.md (539 more words)Show less
Step 2: Draft changes

Ask what to update: description, hypothesis, dates, baseline/comparison treatments, metrics, status, rule, owners, tags. For metrics, route to choose-metric if selecting new ones.

When updating baselineTreatment or comparisonTreatments, verify the treatments exist in the flag definition for the experiment's environment (Get definition for that environment and check treatments[].name, same as create Step 2). See stop-conditions.md.

Valid status values: ACTIVE, PAUSED, COMPLETED, ARCHIVED (null never allowed). The backend enforces which transitions are actually legal - don't assume a fixed chain (e.g. ACTIVE → PAUSED → COMPLETED); send the requested target status and let a 400 reveal an invalid transition. ARCHIVED is a status value here, not a substitute for delete. Changing metrics or dates on ACTIVE experiment → explicit warning.

Clearable (via API/MCP merge patch): description, hypothesis, rule, keyMetrics, supportingMetrics, tags. Non-clearable: name, startAt, endAt, baselineTreatment, comparisonTreatments, status.

CLI gap: the CLI's update experiment has no -f/file-body option and no mutable field for name or rule - scalar fields use their declared snake_case --set IDs; collections such as tags/owners/metric references use their declared --add/--del handlers, not arbitrary arrays. To change rule or name, stop or propose available MCP with explicit approval; never silently drop the requested field.

Step 3: Confirm and update

Show diff (current vs. new). Link write-safety.md. If ACTIVE experiment and changing key fields, add explicit warning.

Update experiment with merge patch body. 409 = duplicate name. 400 = invalid field, null on non-clearable, or an illegal status transition.

Step 4: Verify

Get experiment again. Compare updated fields, including rule if it was part of this update. Report new status if changed.


Delete experiment
Step 1: Get experiment

Resolve exact ID or fully paginated name/status discovery as in Update Step 1. Require FEATURE_FLAG before mutation. Show name, status, parent, environment, dates.

Step 2: Confirm delete

Hard delete = permanent, no archive/restore. Stricter confirmation required. Offer status update to COMPLETED instead (update operation). Link write-safety.md.

Step 3: Delete

Delete experiment. Confirm deletion; parent flag not affected.

Examples

  • "Set up A/B test on new-checkout flag" → create flow
  • "Pause experiment pricing_test_2" → update status
  • "Update experiment abc-123 to end next Friday" → update endAt
  • "Delete experiment old_test" → delete flow with confirmation

Performance Notes

  • Resolve environment and definition once (create Step 1-2), not per metric
  • Metric selection is choose-metric's scope
  • Update uses merge patch: only changed fields in body

Troubleshooting

IssueResolution
404 on create even though flag existsParent must exist in chosen environment; check Get definition for that environment
User wants to change treatmentsRoute to update-flag-targeting to modify definition, then return here
Experiment needs workspace guardrail metricGuardrails apply automatically; never set via keyMetrics/supportingMetrics
400 on update with nullField is non-clearable; omit it to leave unchanged
Empty keyMetrics on createAllowed but no winner criterion; warn and confirm before Step 6 draft
ACTIVE experiment blocks targeting changeSee write-safety.md; require explicit acknowledgement
User wants to change rule or name via CLINot supported - CLI update experiment has no field/-f for either; use MCP or say so
Parent is AI_CONFIGMutation validation is unsupported here; stop and request a supported AI Config workflow. Do not use the feature-flag definition route
400 on status updateTransition not allowed by the backend for the experiment's current status; report the error, don't retry with a guessed intermediate status

© harness, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (references) in skills/manage-experiments of harness/harness-skills.

  • SKILL.md
  • references/stop-conditions.md

Open the folder on GitHubat commit c25faee

Compare with similar skills

Manage Experiments next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Manage Experiments compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Manage Experiments this skillharness/harness-skills115—~3.2kAutomated safety check: PassApache-2.0
Ab Testingcoreyhaines31/marketingskills54k3 repos~2.8kAutomated safety check: PassMIT
AnalyticsNexus-JPF/note-companion8707 repos~2.2kAutomated safety check: PassMIT
Ab Test Setupfreekmurze/dotfiles1k15 repos~1.8kAutomated safety check: PassNone
Ad Test Designeraaron-he-zhu/aaron-marketing-skills2.9k2 repos~2.8kAutomated safety check: PassApache-2.0
Meta Tags Optimizernowork-studio/notfair-plugin3.9k1 repos~2.7kAutomated safety check: PassMIT

Similar skills

  • Ab Testing

    coreyhaines31/marketingskills

    When the user wants to plan, design, or implement an A/B test or experiment, or build a growth experimentation program.

    54k GitHub starsUsed in 3 repos~2.8k tokens
    Marketing & SEOAuto-check passed
  • Analytics

    Nexus-JPF/note-companion

    When the user wants to set up, improve, or audit analytics tracking and measurement.

    870 GitHub starsUsed in 7 repos~2.2k tokens
    Marketing & SEOAuto-check passed
  • Ab Test Setup

    freekmurze/dotfiles

    When the user wants to plan, design, or implement an A/B test or experiment.

    1k GitHub starsUsed in 15 repos~1.8k tokens
    Marketing & SEOAuto-check passed
  • Ad Test Designer

    aaron-he-zhu/aaron-marketing-skills

    A skill your agent uses when the user asks to "design an A/B test", "set up a creative/landing test", "run an incrementality test", or "is this result statistically and practically material?"…

    2.9k GitHub starsUsed in 2 repos~2.8k tokens
    Marketing & SEOAuto-check passed
  • Meta Tags Optimizer

    nowork-studio/notfair-plugin

    Writes and improves title tags, meta descriptions, Open Graph and Twitter card tags for click-through, with character counts and A/B test variants.

    3.9k GitHub starsUsed in 1 repo~2.7k tokens
    Marketing & SEOAuto-check passed
  • Ab Test Analyzer

    irinabuht12-oss/marketing-skills

    Statistical significance calculator for A/B test results with sample size requirements, segment breakdowns, and hypothesis generation.

    3.9k GitHub stars~1.4k tokensUpdated 14 days ago
    Marketing & SEOAuto-check passed

More from harness/harness-skills

All 24 skills in this repo
  • Audit Report

    harness/harness-skills

    Generate audit reports and compliance trails using Harness audit trail data via MCP v2 tools.

    115 GitHub stars~1.3k tokensUpdated yesterday
    Auto-check passed
  • Chaos Dr Test

    harness/harness-skills

    A skill your agent uses when working with Chaos Engineering steps inside a Harness pipeline.

    115 GitHub stars~2.6k tokensUpdated yesterday
    Auto-check passed
  • Chaos Experiment

    harness/harness-skills

    A skill your agent uses when the user asks to create, edit, update, design, or configure a Harness Chaos Experiment — including faults, probes, actions, experiment YAML, fault injection, pod-delete…

    115 GitHub stars~1.6k tokensUpdated yesterday
    Auto-check passed
  • Cleanup Feature Flags

    harness/harness-skills

    Remove a launched Harness FME feature flag from application code, keeping the treatment FME serves today, and open a pull request.

    115 GitHub stars~2.4k tokensUpdated yesterday
    Auto-check passed
  • Configure Repo Scan

    harness/harness-skills

    Configure code scanning in Harness pipelines using STO security scanners.

    115 GitHub stars~2.2k tokensUpdated yesterday
    Auto-check passed
  • Create Agent Template

    harness/harness-skills

    Generate Harness Agent Template files for AI-powered automation agents.

    115 GitHub stars~2.2k tokensUpdated yesterday
    Auto-check passed

Categories

Questions about Manage Experiments

What does Manage Experiments do?

Create, update (metrics, dates, description, hypothesis, baseline/comparison treatments, status), or delete Harness FME feature-flag experiments. Manage Experiments is an agent skill from harness/harness-skills. Create, update (metrics, dates, description, hypothesis, baseline/comparison treatments, status), or delete Harness FME feature-flag experiments.

When should I use Manage Experiments?

Manage Experiments fits situations like: asked to create; delete an experiment; explaining results (review-experiment-results); phrases: create experiment.

How do I install Manage Experiments in Claude Code?

Run `npx skills add harness/harness-skills --skill manage-experiments -a claude-code`. Or copy the skill folder (skills/manage-experiments in harness/harness-skills) into .claude/skills/manage-experiments in your project. Claude Code loads it when a task matches its description.

How do I install Manage Experiments in Codex?

Run `npx skills add harness/harness-skills --skill manage-experiments -a codex`. Or copy the skill folder (skills/manage-experiments in harness/harness-skills) into .agents/skills/manage-experiments in your project. Codex loads it when a task matches its description.

Can I use Manage Experiments in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add harness/harness-skills --skill manage-experiments -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/manage-experiments, .gemini/skills/manage-experiments, .github/skills/manage-experiments and .opencode/skills/manage-experiments in your project.

What does Manage Experiments need to run?

SKILL.md names no scripts, command-line tools or credentials: Manage Experiments is instructions for the agent only. Compatibility (from SKILL.md): Requires the Harness MCP server or the Harness CLI.

Does Manage Experiments access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Manage Experiments safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Manage Experiments use?

Manage Experiments is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Manage Experiments use?

About 3.2k tokens (SKILL.md is roughly 13k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 389 tokens, read only when the agent opens those files.

What are the alternatives to Manage Experiments?

Skills that share tags, products or a category with Manage Experiments: Ab Testing (coreyhaines31/marketingskills, 54k stars), Analytics (Nexus-JPF/note-companion, 870 stars), Ab Test Setup (freekmurze/dotfiles, 1k stars) and Ad Test Designer (aaron-he-zhu/aaron-marketing-skills, 2.9k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Manage Experiments?

harness (a GitHub organization) maintains it in harness/harness-skills, which has 115 GitHub stars. The repository holds 24 skills in this directory. The repository was last updated on October 6, 2026.

Source: harness/harness-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.