Agent skill

Review Experiment Results

by harness in harness/harness-skills

Explain Harness FME experiment results: winner determination, statistical significance, guardrail metric impact, and data-quality caveats (low sample size, missing data).

Apache-2.0Auto-check passedData & Analytics

Install Review Experiment Results

skills CLI
$ npx skills add harness/harness-skills --skill review-experiment-results -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install harness/harness-skills review-experiment-results --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/harness/harness-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/review-experiment-results .claude/skills/review-experiment-results && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
review-experiment-results
GitHub stars
115
Token cost
~3.9k tokens
SKILL.md length
1,757 words
Files
2 (incl. references)
Skills in repo
24
Repo updated
First seen
Licence
Apache-2.0

At a glance

Explain Harness FME experiment results: winner determination, statistical significance, guardrail metric impact, and data-quality caveats (low sample size, missing data).

  • Works in 9 steps: Establish scope → Resolve the experiment → Fetch experiment definition → …
  • Asked to explain
  • SKILL.md covers Tools, Instructions, Output Format and Examples, plus 2 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Review Experiment Results is an agent skill from harness/harness-skills. Explain Harness FME experiment results: winner determination, statistical significance, guardrail metric impact, and data-quality caveats (low sample size, missing data). Classifies each metric's outcome and produces plain-language readout by default, surfacing underlying stats (p-value, CI, sample size) on request. Use when asked to explain, review, or interpret experiment results; "did treatment X win"; "is this experiment significant"; "what happened to metric Y in this experiment"; or "should we ship this…

Its SKILL.md is about 3.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including reference files (for example `references/state-classification.md`). Compatibility notes: Requires the Harness MCP server or the Harness CLI

It sits in Data & Analytics, covering A/B testing, Experimental design and Data cleaning. The repository describes itself as: A collection of structured AI agent skills that enable Claude Code, Cursor, GitHub Copilot, and other AI coding assistants to create, operate, debug, and govern Harness CI/CD… The licence is Apache-2.0.

When your agent uses it

  • Asked to explain
  • Interpret experiment results
  • Did treatment X win
  • Is this experiment significant

Example prompts

  • “did treatment X win”
  • “is this experiment significant”
  • “what happened to metric Y in this experiment”
  • “/review-experiment-results”

Requirements

  • Compatibility (from SKILL.md): Requires the Harness MCP server or the Harness CLI

Workflow steps

9 steps, taken from the step headings in SKILL.md.

  1. Establish scope
  2. Resolve the experiment
  3. Fetch experiment definition
  4. Fetch experiment settings
  5. Fetch evaluated results
  6. Resolve metric names and descriptions
  7. Classify each result
  8. Determine the verdict
  9. Explain the results

What it can do on your machine

Read from SKILL.md and the folder at commit c25faee. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Requires the Harness MCP server or the Harness CLI

    From compatibility in the SKILL.md frontmatter.

Context cost

Review Experiment Results loads about 3.9k tokens when it runs, and up to ~4.5k if it reads all its reference files. Until then it costs about 208 tokens; SKILL.md has 1,757 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~208
When it runs · the whole SKILL.md, loaded when a task matches
~3.9k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~4.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from harness/harness-skills at commit c25faee, republished under its Apache-2.0 licence (© harness). 1,757 words, ~3,901 tokens.

Download SKILL.mdSave it as .claude/skills/review-experiment-results/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
review-experiment-results
description
Explain Harness FME experiment results: winner determination, statistical significance, guardrail metric impact, and data-quality caveats (low sample size, missing data). Classifies each metric's outcome and produces plain-language readout by default, surfacing underlying stats (p-value, CI, sample size) on request. Use when asked to explain, review, or interpret experiment results; "did treatment X win"; "is this experiment significant"; "what happened to metric Y in this experiment"; or "should we ship this experiment". Do NOT use for changing experiments (manage-experiments), choosing metrics (choose-metric), or diagnosing bias (SRM root cause, peeking, Simpson's paradox). Trigger phrases: experiment results, readout, winner, statistical significance, guardrail metric, ship this experiment.
compatibility
Requires the Harness MCP server or the Harness CLI
metadata.author
Harness
metadata.version
2.1.2
metadata.mcp-server
harness-mcp
license
Apache-2.0

Review Experiment Results

Explain an FME experiment's outcome in plain language: whether there's a winner, why or why not, how guardrail metrics were affected, and any data-quality caveats - with underlying statistics available on request.

Related: manage-experiments (experiment changes), choose-metric (metric selection).

Tools

Works through the Harness MCP server or the Harness CLI; names are from tool-map.md.

OperationMCPCLI
List experimentsharness_list · fme_experiment · filters: { parent_type, name?, match_type: "contains", status?: ["ACTIVE", "PAUSED", "COMPLETED", "ARCHIVED"], offset: 0, limit: 100 } · compact: falseharness list experiment --parent-type FEATURE_FLAG [--search <name>] --status ACTIVE, then repeat with --status PAUSED, --status COMPLETED, --status ARCHIVED (CLI --status is single-value; API defaults to ACTIVE if omitted)
Get experimentharness_get · fme_experiment · params: { experiment_id }harness get experiment <experiment-id>
Get settingsharness_get · fme_experiment_settings · params: { experiment_id }harness get experiment:settings <experiment-id>
List resultsharness_list · fme_experiment_result · filters: { experiment_id, comparisons? }harness list experiment:results <experiment-id> --raw
Get/list metricsharness_get · fme_metric · params.metric_id (one call per id), or harness_list · fme_metric · filters: { ids } · compact: falseharness get metric <id> once per id, or harness list metric --id <id> once per id - the CLI's --id flag is repeatable but only the first value reaches the API (spec bug), so repeating --id id1 --id id2 on one call silently drops id2

Instructions

Phase 1: Establish scope

Follow scope-establishment.md.

Step 1: Resolve the experiment

If the user gives an exact experiment ID, Get experiment directly - skip list. Otherwise List experiments for parent type FEATURE_FLAG (or AI_CONFIG if user specified), with name substring match if provided (MCP: match_type: "contains"). Default parent type: FEATURE_FLAG.

If user didn't specify status, list ACTIVE, PAUSED, COMPLETED, and ARCHIVED - a finished or archived experiment is exactly the kind a results question is usually about, so never default to active-only here. Fully paginate all selected statuses per pagination before deciding uniqueness or absence; CLI needs a separate paginated query per status. More than one match or no name given → ask which experiment. If a scan is incomplete, report that instead of silently selecting a first-page match.

Step 2: Fetch experiment definition

Get experiment. Key fields: status, environment, startAt/endAt, baselineTreatment, comparisonTreatments[], keyMetrics[], supportingMetrics[].

keyMetrics/supportingMetrics are the only metric lists on the experiment. GUARDRAIL and ALERT metrics are workspace-wide categories applying to every experiment, so they never appear here - read role from each result's category (Step 4). Empty keyMetrics → no winner criterion; ask which metric(s) should drive verdict.

More than one comparisonTreatments and user didn't name one → report every treatment (one verdict each, see Output Format). If question presumes a single treatment ("did it win?") → ask which treatment or confirm per-treatment readout.

Step 3: Fetch experiment settings

Get settings. Always returns applied settings (404 only if experiment doesn't exist). source: "EXPERIMENT_OVERRIDE" means experiment has its own override; other values mean it inherits org defaults.

Fields to carry forward: significanceThreshold (not always 0.05), multipleComparisonCorrection (NONE or GROUPWISE_HOCHBERG; applies only to KEY/SUPPORTING, never GUARDRAIL/ALERT), minimumSampleSize (per treatment), statisticalTestType (FIXED_HORIZON/SEQUENTIAL), reviewPeriod (ISO-8601 duration), varianceReduction.method (NONE/CUPED).

Step 4: Fetch evaluated results

List results for the experiment, optionally narrowed to specific comparison treatments. List-only (no single-result get, no per-environment filter). Returns latest calculation run only, as one row per (metric, comparison treatment) pair, plus top-level calculatedAt (null if never calculated). If narrowing to specific treatments, filter the returned rows on each item's comparison (see tool-map.md).

Per-row fields: category (KEY, SUPPORTING, GUARDRAIL, or ALERT), comparison (comparison treatment this row is for), metricId.id / metricId.name (name can be null; resolve in Step 5), metricResultState (stats engine verdict; null if none yet), pvalue, impactLower/impactUpper, value, errorMargin, baselineMean/comparisonMean (each with Lower/Upper), baselineSampleSize/comparisonSampleSize, varianceReduction, positive (metric's configured isPositive, not observed outcome; non-null even when numeric fields are null).

No SRM field in response, so can't confirm observed treatment allocation matched configured treatment allocation - Step 8 says so whenever it reports a winner. Report metricResultState/pvalue as given; never recompute significance or claim to correct for peeking.

Reconcile before scoring. The expected row set is the Cartesian product of every configured metric (keyMetrics ∪ supportingMetrics from Step 2) × every comparison treatment in scope (all of them, or the one(s) picked in Step 2); it must be non-empty, since Step 2 already stops on empty keyMetrics. Match returned rows against it:

  • Two or more rows for the same (metricId.id, comparison) pair → duplicate/conflicting data; don't pick one arbitrarily - treat that combination as needs_more_data and note the conflict in Step 8.
  • An expected (metric, comparison) combination has no row at all → treat as needs_more_data, same as a row with metricResultState: null; never treat a missing combination as "nothing to report."
  • A row's comparison isn't in the experiment's current comparisonTreatments → drop it (stale data from a removed treatment).
  • calculatedAt is null, or the whole result list is empty → no calculation has run; every KEY result is needs_more_data, so Step 7's verdict is NEED_MORE_DATA - never report WINNER from an empty or uncalculated result set.

GUARDRAIL/ALERT rows aren't part of the expected set (they're workspace-wide, not listed on the experiment) but must still flow through to Step 7's DATA_QUALITY_CONCERN modifier when present with no data - don't drop them while reconciling KEY/SUPPORTING coverage.

Step 5: Resolve metric names and descriptions

Resolve every distinct metric ID from returned rows and configured metrics missing from those rows, so gaps can be named. MCP may List metrics by IDs, fully paginating until every requested ID is resolved or the complete inventory establishes a missing ID; alternatively get each ID directly. CLI must Get metric separately for each ID (or one fully paginated single-ID list per call). An unresolved ID stays unverified, not absent from a first-page miss. Include returned GUARDRAIL/ALERT IDs too. Descriptions help judge severity: a 2% dip on "leading indicator, noisy" reads differently from same dip on "primary revenue guardrail".

Step 6: Classify each result

Map each metricResultState to desired, undesired, inconclusive, or needs_more_data using state-classification.md. For needs_more_data results, compare baselineSampleSize/comparisonSampleSize to minimumSampleSize so readout can say how far along collection is.

Show full SKILL.md (789 more words)Show less
Step 7: Determine the verdict

One verdict per selected comparison. Completeness gate runs first: empty/uncalculated results, missing or duplicate configured metric/comparison pairs, or mismatched expected categories force NEED_MORE_DATA + DATA_QUALITY_CONCERN for the affected comparison. Show observed regressions/guardrails separately; don't hide them. Only a non-empty reconciled set with every configured KEY and SUPPORTING pair present exactly once enters the table below. A missing workspace-wide guardrail inventory remains unverified, not evidence of no breaches.

Then evaluate top to bottom; first match wins:

VerdictCondition
MIXEDAt least one KEY result is desired and at least one result of any category is undesired
WINNERCompleteness gate passed; at least one configured KEY result exists and every configured KEY result is desired
REGRESSIONAt least one KEY result is undesired
NEED_MORE_DATAAt least one KEY result is needs_more_data (including missing or conflicting combinations caught by Step 4's reconciliation)
NO_WINNEROtherwise (KEY results are inconclusive, or mix of desired and inconclusive)

Modifiers (attach to any verdict):

  • GUARDRAIL_BREACH - any category: GUARDRAIL result is undesired. An undesired SUPPORTING or ALERT result affects the verdict but is not a guardrail breach. By row order, WINNER never carries this modifier.
  • DATA_QUALITY_CONCERN - more than half of KEY results are in a NO_DATA_* state; any result is NO_DATA_SERVER_ERROR/FAILED_METRIC; or any GUARDRAIL/ALERT result is in a NO_DATA_* state at all. A guardrail that never received data is a monitoring gap, not just a data point to omit.
Step 8: Explain the results

Default to plain language. Link concepts.md for SRM and peeking caveats. Lead with verdict, then why. Describe trade-offs, not business decision. DATA_QUALITY_CONCERN: say so plainly, and call out any duplicate/conflicting rows or missing metric×treatment combinations Step 4 found. WINNER or MIXED: add one line: SRM isn't exposed through API; check experiment's results page in Harness UI before acting. Fixed-horizon + before review period: say readout is preliminary - if the verdict is WINNER, report it as WINNER (preliminary) in the Output Format and don't call it actionable until the review period ends or the test type is SEQUENTIAL. Inconclusive but not significance-tested: say why (e.g. ACROSS metric). calculatedAt non-null: include timestamp. pvalue null (e.g. WAITING_NORMALITY): don't quote value/impact as confident. GUARDRAIL_BREACH + MIXED: lead with breach. Guardrail or supporting metrics with data (any category present in Step 4's reconciled rows) must appear in the Metric Impact table - never omit a row just because it isn't KEY. If multipleComparisonCorrection is NONE and the experiment has more than one key metric, add one line flagging the false-positive risk. Stats detail: add p-value, CI, sample sizes, significanceThreshold, multipleComparisonCorrection, settings source when question uses statistical terms or user asks. No stats detail: close with one-line note that stats available on request.

Output Format

For a single comparison treatment:

## Experiment Readout
- Experiment: <name> (<status>)
- Comparing: <treatment> vs <baseline>
- Calculated: <calculatedAt>

## Verdict
**<VERDICT>** [+ `(preliminary)` if fixed-horizon and before the review period ends] [+ modifiers if any]
<2-4 sentence explanation>

## Metric Impact
| Metric | Role | Result |
|---|---|---|
| <name> | Key / Supporting / Guardrail / Alert | <desired/undesired/inconclusive/needs more data, in plain words> |

## Notes
<data-quality caveats, if any; otherwise omit>

For more than one comparison treatment, keep one shared experiment header and Notes section, but repeat verdict + table block per treatment under its own ### <treatment name> subheading (independent verdict per treatment).

Add p-value/CI/sample-size columns to Metric Impact only when Step 8's stats condition is met.

Examples

  • "Explain the results of checkout-redesign experiment" → resolve by name, Steps 2-8, plain-language output
  • "Did treatment B win in exp_8f2a1c?" → resolve by ID, verdict for treatment B only
  • "Is the signup-flow experiment statistically significant?" → statistical wording triggers stats detail
  • "Should we ship the onboarding experiment?" → verdict + trade-offs; no ship/no-ship recommendation

Performance Notes

  • Resolve concrete experiment id before Step 4; never fetch results by name
  • Fetch settings (Step 3) and metric names (Step 5) once per experiment, not per treatment
  • Narrow Step 4 to specific comparison treatments once user picks a treatment instead of fetching every comparison

Troubleshooting

IssueResolution
An experiment operation failsReport error and stop
Step 2 404s but Steps 3/4 succeedWithout Step 2 there are no treatment or key-metric definitions; tell user full readout isn't possible for this experiment
Experiment still ACTIVE or PAUSEDIf verdict is NEED_MORE_DATA or NO_WINNER, ask whether user wants preliminary readout (caveated) or would rather wait; use endAt (and reviewPeriod) to say how much longer it's configured to run
Unrecognized or null metricResultStateTreat as needs_more_data; don't guess a direction; see state-classification.md
Duplicate or conflicting rows for the same metric + comparisonTreat that combination as needs_more_data; note the conflict; never pick one row arbitrarily
Expected metric×treatment combination missing from resultsTreat as needs_more_data; never report WINNER from partial coverage
Results are empty and experiment's rule is default or another label with no trafficResults only count impressions whose label matches the experiment's rule; offer to update the experiment's rule to "default rule" via manage-experiments
User asks about SRMResults API exposes no SRM value; say so and point to experiment's results page in Harness UI; don't diagnose causes
User wants to change experimentRoute to manage-experiments
User asks which metric to useRoute to choose-metric

© harness, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (references) in skills/review-experiment-results of harness/harness-skills.

  • SKILL.md
  • references/state-classification.md

Open the folder on GitHubat commit c25faee

Compare with similar skills

Review Experiment Results next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Review Experiment Results compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Review Experiment Results this skillharness/harness-skills115—~3.9kAutomated safety check: PassApache-2.0
Statistical Analystalirezarezvani/claude-skills28k—~2.5kAutomated safety check: PassMIT
Experimentation Analyticsrampstackco/claude-skills945—~8.9kAutomated safety check: PassMIT
Power Analysisgaasher/Agent-Loop-Skills174—~2.2kAutomated safety check: PassMIT
Data Scientistmagnus919/hermes-profiles289—~3.3kAutomated safety check: PassMIT
Data Scientistmagnus919/agent-skills119—~4.1kAutomated safety check: PassMIT

Similar skills

  • Statistical Analyst

    alirezarezvani/claude-skills

    Run hypothesis tests, analyze A/B experiment results, calculate sample sizes, and interpret statistical significance with effect sizes.

    28k GitHub stars~2.5k tokensUpdated 1 mo ago
    Data & AnalyticsAuto-check passed
  • Experimentation Analytics

    rampstackco/claude-skills

    How to read experiment results without fooling yourself. An agent skill from rampstackco/claude-skills.

    945 GitHub stars~8.9k tokensUpdated 4 days ago
    Data & AnalyticsAuto-check passed
  • Power Analysis

    gaasher/Agent-Loop-Skills

    A skill your agent uses when the user is planning a two-arm comparison (an A/B test, a simple RCT, a behavioral study, or a two-model/two-config evaluation) and needs to size it and preregister it…

    174 GitHub stars~2.2k tokensUpdated 3 mo ago
    Data & AnalyticsAuto-check passed
  • Data Scientist

    magnus919/hermes-profiles

    PhD-level expertise in data science, statistics, and machine learning.

    289 GitHub stars~3.3k tokensUpdated 3 mo ago
    Research & ScienceAuto-check passed
  • Data Scientist

    magnus919/agent-skills

    A skill your agent uses for PhD-level expertise in data science, statistics, and machine learning: rigorous statistical analysis, experimental design, causal inference, advanced modeling, research…

    119 GitHub stars~4.1k tokensUpdated today
    Research & ScienceAuto-check passed
  • Designs QC, corrects signal drift, removes batch effects, filters features, normalizes samples, and imputes missing values for untargeted LC-MS/GC-MS metabolomics, framing each step as a measurement…

    1.2k GitHub starsUsed in 1 repo~5k tokens
    DatabasesAuto-check passed

More from harness/harness-skills

All 24 skills in this repo
  • Audit Report

    harness/harness-skills

    Generate audit reports and compliance trails using Harness audit trail data via MCP v2 tools.

    115 GitHub stars~1.3k tokensUpdated 4 days ago
    Auto-check passed
  • Chaos Dr Test

    harness/harness-skills

    A skill your agent uses when working with Chaos Engineering steps inside a Harness pipeline.

    115 GitHub stars~2.6k tokensUpdated 4 days ago
    Auto-check passed
  • Chaos Experiment

    harness/harness-skills

    A skill your agent uses when the user asks to create, edit, update, design, or configure a Harness Chaos Experiment — including faults, probes, actions, experiment YAML, fault injection, pod-delete…

    115 GitHub stars~1.6k tokensUpdated 4 days ago
    Auto-check passed
  • Cleanup Feature Flags

    harness/harness-skills

    Remove a launched Harness FME feature flag from application code, keeping the treatment FME serves today, and open a pull request.

    115 GitHub stars~2.4k tokensUpdated 4 days ago
    Auto-check passed
  • Configure Repo Scan

    harness/harness-skills

    Configure code scanning in Harness pipelines using STO security scanners.

    115 GitHub stars~2.2k tokensUpdated 4 days ago
    Auto-check passed
  • Create Agent Template

    harness/harness-skills

    Generate Harness Agent Template files for AI-powered automation agents.

    115 GitHub stars~2.2k tokensUpdated 4 days ago
    Auto-check passed

Questions about Review Experiment Results

What does Review Experiment Results do?

Explain Harness FME experiment results: winner determination, statistical significance, guardrail metric impact, and data-quality caveats (low sample size, missing data). Review Experiment Results is an agent skill from harness/harness-skills. Explain Harness FME experiment results: winner determination, statistical significance, guardrail metric impact, and data-quality caveats (low sample size, missing data).

When should I use Review Experiment Results?

Review Experiment Results fits situations like: asked to explain; interpret experiment results; did treatment X win; is this experiment significant.

How do I install Review Experiment Results in Claude Code?

Run `npx skills add harness/harness-skills --skill review-experiment-results -a claude-code`. Or copy the skill folder (skills/review-experiment-results in harness/harness-skills) into .claude/skills/review-experiment-results in your project. Claude Code loads it when a task matches its description.

How do I install Review Experiment Results in Codex?

Run `npx skills add harness/harness-skills --skill review-experiment-results -a codex`. Or copy the skill folder (skills/review-experiment-results in harness/harness-skills) into .agents/skills/review-experiment-results in your project. Codex loads it when a task matches its description.

Can I use Review Experiment Results in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add harness/harness-skills --skill review-experiment-results -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/review-experiment-results, .gemini/skills/review-experiment-results, .github/skills/review-experiment-results and .opencode/skills/review-experiment-results in your project.

What does Review Experiment Results need to run?

SKILL.md names no scripts, command-line tools or credentials: Review Experiment Results is instructions for the agent only. Compatibility (from SKILL.md): Requires the Harness MCP server or the Harness CLI.

Does Review Experiment Results access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Review Experiment Results safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Review Experiment Results use?

Review Experiment Results is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Review Experiment Results use?

About 3.9k tokens (SKILL.md is roughly 16k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 555 tokens, read only when the agent opens those files.

What are the alternatives to Review Experiment Results?

Skills that share tags, products or a category with Review Experiment Results: Statistical Analyst (alirezarezvani/claude-skills, 28k stars), Experimentation Analytics (rampstackco/claude-skills, 945 stars), Power Analysis (gaasher/Agent-Loop-Skills, 174 stars) and Data Scientist (magnus919/hermes-profiles, 289 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Review Experiment Results?

harness (a GitHub organization) maintains it in harness/harness-skills, which has 115 GitHub stars. The repository holds 24 skills in this directory. The repository was last updated on October 6, 2026.

Source: harness/harness-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.