Statistical Analyst
alirezarezvani/claude-skills
Run hypothesis tests, analyze A/B experiment results, calculate sample sizes, and interpret statistical significance with effect sizes.
Explain Harness FME experiment results: winner determination, statistical significance, guardrail metric impact, and data-quality caveats (low sample size, missing data).
$ npx skills add harness/harness-skills --skill review-experiment-results -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install harness/harness-skills review-experiment-results --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/harness/harness-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/review-experiment-results .claude/skills/review-experiment-results && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "review-experiment-results" agent skill from https://github.com/harness/harness-skills/tree/main/skills/review-experiment-results into .claude/skills/review-experiment-results/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "review-experiment-results", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/harness/harness-skills/tree/main/skills/review-experiment-resultsType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add harness/harness-skills --skill review-experiment-results -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install harness/harness-skills review-experiment-results --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/harness/harness-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/review-experiment-results .agents/skills/review-experiment-results && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "review-experiment-results" agent skill from https://github.com/harness/harness-skills/tree/main/skills/review-experiment-results into .agents/skills/review-experiment-results/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "review-experiment-results", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add harness/harness-skills --skill review-experiment-results -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install harness/harness-skills review-experiment-results --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/harness/harness-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/review-experiment-results .cursor/skills/review-experiment-results && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "review-experiment-results" agent skill from https://github.com/harness/harness-skills/tree/main/skills/review-experiment-results into .cursor/skills/review-experiment-results/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "review-experiment-results", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/harness/harness-skills.git --path skills/review-experiment-results--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add harness/harness-skills --skill review-experiment-results -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install harness/harness-skills review-experiment-results --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/harness/harness-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/review-experiment-results .gemini/skills/review-experiment-results && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "review-experiment-results" agent skill from https://github.com/harness/harness-skills/tree/main/skills/review-experiment-results into .gemini/skills/review-experiment-results/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "review-experiment-results", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install harness/harness-skills review-experiment-resultsInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add harness/harness-skills --skill review-experiment-results -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/harness/harness-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/review-experiment-results .github/skills/review-experiment-results && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "review-experiment-results" agent skill from https://github.com/harness/harness-skills/tree/main/skills/review-experiment-results into .github/skills/review-experiment-results/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "review-experiment-results", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add harness/harness-skills --skill review-experiment-results -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install harness/harness-skills review-experiment-results --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/harness/harness-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/review-experiment-results .opencode/skills/review-experiment-results && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "review-experiment-results" agent skill from https://github.com/harness/harness-skills/tree/main/skills/review-experiment-results into .opencode/skills/review-experiment-results/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "review-experiment-results", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
review-experiment-resultsExplain Harness FME experiment results: winner determination, statistical significance, guardrail metric impact, and data-quality caveats (low sample size, missing data).
Review Experiment Results is an agent skill from harness/harness-skills. Explain Harness FME experiment results: winner determination, statistical significance, guardrail metric impact, and data-quality caveats (low sample size, missing data). Classifies each metric's outcome and produces plain-language readout by default, surfacing underlying stats (p-value, CI, sample size) on request. Use when asked to explain, review, or interpret experiment results; "did treatment X win"; "is this experiment significant"; "what happened to metric Y in this experiment"; or "should we ship this…
Its SKILL.md is about 3.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including reference files (for example `references/state-classification.md`). Compatibility notes: Requires the Harness MCP server or the Harness CLI
It sits in Data & Analytics, covering A/B testing, Experimental design and Data cleaning. The repository describes itself as: A collection of structured AI agent skills that enable Claude Code, Cursor, GitHub Copilot, and other AI coding assistants to create, operate, debug, and govern Harness CI/CD… The licence is Apache-2.0.
9 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit c25faee. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md.
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Requires the Harness MCP server or the Harness CLI
From compatibility in the SKILL.md frontmatter.
Review Experiment Results loads about 3.9k tokens when it runs, and up to ~4.5k if it reads all its reference files. Until then it costs about 208 tokens; SKILL.md has 1,757 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from harness/harness-skills at commit c25faee, republished under its Apache-2.0 licence (© harness). 1,757 words, ~3,901 tokens.
.claude/skills/review-experiment-results/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.Explain an FME experiment's outcome in plain language: whether there's a winner, why or why not, how guardrail metrics were affected, and any data-quality caveats - with underlying statistics available on request.
Related: manage-experiments (experiment changes), choose-metric (metric selection).
Works through the Harness MCP server or the Harness CLI; names are from tool-map.md.
| Operation | MCP | CLI |
|---|---|---|
| List experiments | harness_list · fme_experiment · filters: { parent_type, name?, match_type: "contains", status?: ["ACTIVE", "PAUSED", "COMPLETED", "ARCHIVED"], offset: 0, limit: 100 } · compact: false | harness list experiment --parent-type FEATURE_FLAG [--search <name>] --status ACTIVE, then repeat with --status PAUSED, --status COMPLETED, --status ARCHIVED (CLI --status is single-value; API defaults to ACTIVE if omitted) |
| Get experiment | harness_get · fme_experiment · params: { experiment_id } | harness get experiment <experiment-id> |
| Get settings | harness_get · fme_experiment_settings · params: { experiment_id } | harness get experiment:settings <experiment-id> |
| List results | harness_list · fme_experiment_result · filters: { experiment_id, comparisons? } | harness list experiment:results <experiment-id> --raw |
| Get/list metrics | harness_get · fme_metric · params.metric_id (one call per id), or harness_list · fme_metric · filters: { ids } · compact: false | harness get metric <id> once per id, or harness list metric --id <id> once per id - the CLI's --id flag is repeatable but only the first value reaches the API (spec bug), so repeating --id id1 --id id2 on one call silently drops id2 |
Follow scope-establishment.md.
If the user gives an exact experiment ID, Get experiment directly - skip list. Otherwise List experiments for parent type FEATURE_FLAG (or AI_CONFIG if user specified), with name substring match if provided (MCP: match_type: "contains"). Default parent type: FEATURE_FLAG.
If user didn't specify status, list ACTIVE, PAUSED, COMPLETED, and ARCHIVED - a finished or archived experiment is exactly the kind a results question is usually about, so never default to active-only here. Fully paginate all selected statuses per pagination before deciding uniqueness or absence; CLI needs a separate paginated query per status. More than one match or no name given → ask which experiment. If a scan is incomplete, report that instead of silently selecting a first-page match.
Get experiment. Key fields: status, environment, startAt/endAt, baselineTreatment, comparisonTreatments[], keyMetrics[], supportingMetrics[].
keyMetrics/supportingMetrics are the only metric lists on the experiment. GUARDRAIL and ALERT metrics are workspace-wide categories applying to every experiment, so they never appear here - read role from each result's category (Step 4). Empty keyMetrics → no winner criterion; ask which metric(s) should drive verdict.
More than one comparisonTreatments and user didn't name one → report every treatment (one verdict each, see Output Format). If question presumes a single treatment ("did it win?") → ask which treatment or confirm per-treatment readout.
Get settings. Always returns applied settings (404 only if experiment doesn't exist). source: "EXPERIMENT_OVERRIDE" means experiment has its own override; other values mean it inherits org defaults.
Fields to carry forward: significanceThreshold (not always 0.05), multipleComparisonCorrection (NONE or GROUPWISE_HOCHBERG; applies only to KEY/SUPPORTING, never GUARDRAIL/ALERT), minimumSampleSize (per treatment), statisticalTestType (FIXED_HORIZON/SEQUENTIAL), reviewPeriod (ISO-8601 duration), varianceReduction.method (NONE/CUPED).
List results for the experiment, optionally narrowed to specific comparison treatments. List-only (no single-result get, no per-environment filter). Returns latest calculation run only, as one row per (metric, comparison treatment) pair, plus top-level calculatedAt (null if never calculated). If narrowing to specific treatments, filter the returned rows on each item's comparison (see tool-map.md).
Per-row fields: category (KEY, SUPPORTING, GUARDRAIL, or ALERT), comparison (comparison treatment this row is for), metricId.id / metricId.name (name can be null; resolve in Step 5), metricResultState (stats engine verdict; null if none yet), pvalue, impactLower/impactUpper, value, errorMargin, baselineMean/comparisonMean (each with Lower/Upper), baselineSampleSize/comparisonSampleSize, varianceReduction, positive (metric's configured isPositive, not observed outcome; non-null even when numeric fields are null).
No SRM field in response, so can't confirm observed treatment allocation matched configured treatment allocation - Step 8 says so whenever it reports a winner. Report metricResultState/pvalue as given; never recompute significance or claim to correct for peeking.
Reconcile before scoring. The expected row set is the Cartesian product of every configured metric (keyMetrics ∪ supportingMetrics from Step 2) × every comparison treatment in scope (all of them, or the one(s) picked in Step 2); it must be non-empty, since Step 2 already stops on empty keyMetrics. Match returned rows against it:
metricId.id, comparison) pair → duplicate/conflicting data; don't pick one arbitrarily - treat that combination as needs_more_data and note the conflict in Step 8.needs_more_data, same as a row with metricResultState: null; never treat a missing combination as "nothing to report."comparison isn't in the experiment's current comparisonTreatments → drop it (stale data from a removed treatment).calculatedAt is null, or the whole result list is empty → no calculation has run; every KEY result is needs_more_data, so Step 7's verdict is NEED_MORE_DATA - never report WINNER from an empty or uncalculated result set.GUARDRAIL/ALERT rows aren't part of the expected set (they're workspace-wide, not listed on the experiment) but must still flow through to Step 7's DATA_QUALITY_CONCERN modifier when present with no data - don't drop them while reconciling KEY/SUPPORTING coverage.
Resolve every distinct metric ID from returned rows and configured metrics missing from those rows, so gaps can be named. MCP may List metrics by IDs, fully paginating until every requested ID is resolved or the complete inventory establishes a missing ID; alternatively get each ID directly. CLI must Get metric separately for each ID (or one fully paginated single-ID list per call). An unresolved ID stays unverified, not absent from a first-page miss. Include returned GUARDRAIL/ALERT IDs too. Descriptions help judge severity: a 2% dip on "leading indicator, noisy" reads differently from same dip on "primary revenue guardrail".
Map each metricResultState to desired, undesired, inconclusive, or needs_more_data using state-classification.md. For needs_more_data results, compare baselineSampleSize/comparisonSampleSize to minimumSampleSize so readout can say how far along collection is.
One verdict per selected comparison. Completeness gate runs first: empty/uncalculated results, missing or duplicate configured metric/comparison pairs, or mismatched expected categories force NEED_MORE_DATA + DATA_QUALITY_CONCERN for the affected comparison. Show observed regressions/guardrails separately; don't hide them. Only a non-empty reconciled set with every configured KEY and SUPPORTING pair present exactly once enters the table below. A missing workspace-wide guardrail inventory remains unverified, not evidence of no breaches.
Then evaluate top to bottom; first match wins:
| Verdict | Condition |
|---|---|
MIXED | At least one KEY result is desired and at least one result of any category is undesired |
WINNER | Completeness gate passed; at least one configured KEY result exists and every configured KEY result is desired |
REGRESSION | At least one KEY result is undesired |
NEED_MORE_DATA | At least one KEY result is needs_more_data (including missing or conflicting combinations caught by Step 4's reconciliation) |
NO_WINNER | Otherwise (KEY results are inconclusive, or mix of desired and inconclusive) |
Modifiers (attach to any verdict):
GUARDRAIL_BREACH - any category: GUARDRAIL result is undesired. An
undesired SUPPORTING or ALERT result affects the verdict but is not a
guardrail breach. By row order, WINNER never carries this modifier.DATA_QUALITY_CONCERN - more than half of KEY results are in a NO_DATA_*
state; any result is NO_DATA_SERVER_ERROR/FAILED_METRIC; or any
GUARDRAIL/ALERT result is in a NO_DATA_* state at all. A guardrail that
never received data is a monitoring gap, not just a data point to omit.Default to plain language. Link concepts.md for SRM and peeking caveats. Lead with verdict, then why. Describe trade-offs, not business decision. DATA_QUALITY_CONCERN: say so plainly, and call out any duplicate/conflicting rows or missing metric×treatment combinations Step 4 found. WINNER or MIXED: add one line: SRM isn't exposed through API; check experiment's results page in Harness UI before acting. Fixed-horizon + before review period: say readout is preliminary - if the verdict is WINNER, report it as WINNER (preliminary) in the Output Format and don't call it actionable until the review period ends or the test type is SEQUENTIAL. Inconclusive but not significance-tested: say why (e.g. ACROSS metric). calculatedAt non-null: include timestamp. pvalue null (e.g. WAITING_NORMALITY): don't quote value/impact as confident. GUARDRAIL_BREACH + MIXED: lead with breach. Guardrail or supporting metrics with data (any category present in Step 4's reconciled rows) must appear in the Metric Impact table - never omit a row just because it isn't KEY. If multipleComparisonCorrection is NONE and the experiment has more than one key metric, add one line flagging the false-positive risk. Stats detail: add p-value, CI, sample sizes, significanceThreshold, multipleComparisonCorrection, settings source when question uses statistical terms or user asks. No stats detail: close with one-line note that stats available on request.
For a single comparison treatment:
## Experiment Readout
- Experiment: <name> (<status>)
- Comparing: <treatment> vs <baseline>
- Calculated: <calculatedAt>
## Verdict
**<VERDICT>** [+ `(preliminary)` if fixed-horizon and before the review period ends] [+ modifiers if any]
<2-4 sentence explanation>
## Metric Impact
| Metric | Role | Result |
|---|---|---|
| <name> | Key / Supporting / Guardrail / Alert | <desired/undesired/inconclusive/needs more data, in plain words> |
## Notes
<data-quality caveats, if any; otherwise omit>For more than one comparison treatment, keep one shared experiment header and
Notes section, but repeat verdict + table block per treatment under its own
### <treatment name> subheading (independent verdict per treatment).
Add p-value/CI/sample-size columns to Metric Impact only when Step 8's stats condition is met.
| Issue | Resolution |
|---|---|
| An experiment operation fails | Report error and stop |
| Step 2 404s but Steps 3/4 succeed | Without Step 2 there are no treatment or key-metric definitions; tell user full readout isn't possible for this experiment |
| Experiment still ACTIVE or PAUSED | If verdict is NEED_MORE_DATA or NO_WINNER, ask whether user wants preliminary readout (caveated) or would rather wait; use endAt (and reviewPeriod) to say how much longer it's configured to run |
Unrecognized or null metricResultState | Treat as needs_more_data; don't guess a direction; see state-classification.md |
| Duplicate or conflicting rows for the same metric + comparison | Treat that combination as needs_more_data; note the conflict; never pick one row arbitrarily |
| Expected metric×treatment combination missing from results | Treat as needs_more_data; never report WINNER from partial coverage |
Results are empty and experiment's rule is default or another label with no traffic | Results only count impressions whose label matches the experiment's rule; offer to update the experiment's rule to "default rule" via manage-experiments |
| User asks about SRM | Results API exposes no SRM value; say so and point to experiment's results page in Harness UI; don't diagnose causes |
| User wants to change experiment | Route to manage-experiments |
| User asks which metric to use | Route to choose-metric |
© harness, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 1 other file (references) in skills/review-experiment-results of harness/harness-skills.
Open the folder on GitHubat commit c25faee
Review Experiment Results next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Review Experiment Results this skillharness/harness-skills | 115 | — | ~3.9k | Automated safety check: Pass | Apache-2.0 | |
| Statistical Analystalirezarezvani/claude-skills | 28k | — | ~2.5k | Automated safety check: Pass | MIT | |
| Experimentation Analyticsrampstackco/claude-skills | 945 | — | ~8.9k | Automated safety check: Pass | MIT | |
| Power Analysisgaasher/Agent-Loop-Skills | 174 | — | ~2.2k | Automated safety check: Pass | MIT | |
| Data Scientistmagnus919/hermes-profiles | 289 | — | ~3.3k | Automated safety check: Pass | MIT | |
| Data Scientistmagnus919/agent-skills | 119 | — | ~4.1k | Automated safety check: Pass | MIT |
alirezarezvani/claude-skills
Run hypothesis tests, analyze A/B experiment results, calculate sample sizes, and interpret statistical significance with effect sizes.
rampstackco/claude-skills
How to read experiment results without fooling yourself. An agent skill from rampstackco/claude-skills.
gaasher/Agent-Loop-Skills
A skill your agent uses when the user is planning a two-arm comparison (an A/B test, a simple RCT, a behavioral study, or a two-model/two-config evaluation) and needs to size it and preregister it…
magnus919/hermes-profiles
PhD-level expertise in data science, statistics, and machine learning.
magnus919/agent-skills
A skill your agent uses for PhD-level expertise in data science, statistics, and machine learning: rigorous statistical analysis, experimental design, causal inference, advanced modeling, research…
GPTomics/bioSkills
Designs QC, corrects signal drift, removes batch effects, filters features, normalizes samples, and imputes missing values for untargeted LC-MS/GC-MS metabolomics, framing each step as a measurement…
harness/harness-skills
Generate audit reports and compliance trails using Harness audit trail data via MCP v2 tools.
harness/harness-skills
A skill your agent uses when working with Chaos Engineering steps inside a Harness pipeline.
harness/harness-skills
A skill your agent uses when the user asks to create, edit, update, design, or configure a Harness Chaos Experiment — including faults, probes, actions, experiment YAML, fault injection, pod-delete…
harness/harness-skills
Remove a launched Harness FME feature flag from application code, keeping the treatment FME serves today, and open a pull request.
harness/harness-skills
Configure code scanning in Harness pipelines using STO security scanners.
harness/harness-skills
Generate Harness Agent Template files for AI-powered automation agents.
Explain Harness FME experiment results: winner determination, statistical significance, guardrail metric impact, and data-quality caveats (low sample size, missing data). Review Experiment Results is an agent skill from harness/harness-skills. Explain Harness FME experiment results: winner determination, statistical significance, guardrail metric impact, and data-quality caveats (low sample size, missing data).
Review Experiment Results fits situations like: asked to explain; interpret experiment results; did treatment X win; is this experiment significant.
Run `npx skills add harness/harness-skills --skill review-experiment-results -a claude-code`. Or copy the skill folder (skills/review-experiment-results in harness/harness-skills) into .claude/skills/review-experiment-results in your project. Claude Code loads it when a task matches its description.
Run `npx skills add harness/harness-skills --skill review-experiment-results -a codex`. Or copy the skill folder (skills/review-experiment-results in harness/harness-skills) into .agents/skills/review-experiment-results in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add harness/harness-skills --skill review-experiment-results -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/review-experiment-results, .gemini/skills/review-experiment-results, .github/skills/review-experiment-results and .opencode/skills/review-experiment-results in your project.
SKILL.md names no scripts, command-line tools or credentials: Review Experiment Results is instructions for the agent only. Compatibility (from SKILL.md): Requires the Harness MCP server or the Harness CLI.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Review Experiment Results is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 3.9k tokens (SKILL.md is roughly 16k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 555 tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Review Experiment Results: Statistical Analyst (alirezarezvani/claude-skills, 28k stars), Experimentation Analytics (rampstackco/claude-skills, 945 stars), Power Analysis (gaasher/Agent-Loop-Skills, 174 stars) and Data Scientist (magnus919/hermes-profiles, 289 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
harness (a GitHub organization) maintains it in harness/harness-skills, which has 115 GitHub stars. The repository holds 24 skills in this directory. The repository was last updated on October 6, 2026.
Source: harness/harness-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.