Plan real experiments and analyze supplied continuous measurements in Lab Bench.

MITAuto-check passedDocuments & Office

Install Bench

skills CLI
$ npx skills add autonomous-ai/openharness --skill bench -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install autonomous-ai/openharness bench --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/autonomous-ai/openharness.git skills-src && mkdir -p .claude/skills && cp -r skills-src/store/agents/lab-bench/skills/bench .claude/skills/bench && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
bench
GitHub stars
1.1k
Token cost
~900 tokens
SKILL.md length
448 words
Files
3 (incl. scripts, references)
Skills in repo
99
Repo updated
First seen
Licence
MIT

At a glance

Plan real experiments and analyze supplied continuous measurements in Lab Bench.

  • Works in 8 steps: Inspect bench/project.json and… → Generate a balanced factorial with… → Build with node tools/build.mjs and give… → …
  • Defining factors and independent runs
  • Runs Shell scripts from its folder; calls node
  • Randomized collection sheets

What it does

Bench is an agent skill from autonomous-ai/openharness. Plan real experiments and analyze supplied continuous measurements in Lab Bench. Use for defining factors and independent runs, randomized collection sheets, measurement CSV imports, experimental uncertainty and diagnostics, comparisons and follow-up confirmation runs with a reproducible report.

Its SKILL.md is about 900 tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files, including scripts and reference files (for example `references/project.md` and `scripts/seed-verdict.sh`).

It sits in Documents & Office, covering CSV and tabular files. The repository describes itself as: The ultimate harness for coding agents and beyond. All your agents. All your machines. One command center. Start with code, then follow your curiosity and build across… The licence is MIT.

When your agent uses it

  • Defining factors and independent runs
  • Randomized collection sheets
  • Measurement CSV imports
  • Experimental uncertainty and diagnostics

Example prompts

  • “/bench”

Requirements

  • Python 3
  • A Bash shell

Workflow steps

8 steps, taken from the first numbered list in SKILL.md.

  1. Inspect bench/project.json and bench/DESIGN.md. Preserve existing collected runs, approved
  2. Generate a balanced factorial with independent repetitions and complete blocks. If adding
  3. Build with node tools/build.mjs and give the person a real collection sheet. Do not invent
  4. Import original bytes with prepareImport / commitImport, or use recordMeasurement with a
  5. Use model terms consistent with the question and design. Inspect residual structure and
  6. Compare settings within the planned ranges. A model mean interval differs from a new-observation
  7. Freeze predictions before collecting a confirmation phase. Keep its new data out of the model
  8. Run node tools/check.mjs, exercise real browser inputs and export with

What it can do on your machine

Read from SKILL.md and the folder at commit 54a1f1b. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Shell), which the agent can run.

    Shell commands in SKILL.md call:

    • node

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Bench loads about 900 tokens when it runs, and up to ~2.9k if it reads all its reference files. Until then it costs about 76 tokens; SKILL.md has 448 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~76
When it runs · the whole SKILL.md, loaded when a task matches
~900
With references · SKILL.md plus every file in references/, read only if the agent opens them
~2.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from autonomous-ai/openharness at commit 54a1f1b, republished under its MIT licence (© autonomous-ai). 448 words, ~900 tokens.

Download SKILL.mdSave it as .claude/skills/bench/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
bench
description
Plan real experiments and analyze supplied continuous measurements in Lab Bench. Use for defining factors and independent runs, randomized collection sheets, measurement CSV imports, experimental uncertainty and diagnostics, comparisons and follow-up confirmation runs with a reproducible report.

bench

Build a usable experiment from the person's question. Read the project contract when authoring or revising the source. Keep the procedure, independent unit, factors, levels, response unit and intended decision explicit. Use the studio to make the work inspectable.

  1. Inspect bench/project.json and bench/DESIGN.md. Preserve existing collected runs, approved protocol, original files and change history. A new brief is open ended; the coffee plan is only an example. Record necessary assumptions and choose an achievable initial run budget.
  2. Generate a balanced factorial with independent repetitions and complete blocks. If adding controls, bracket and intersperse them. A two-level design plus shared centers cannot separately identify multiple quadratic effects; inspect matrix rank before collection.
  3. Build with node tools/build.mjs and give the person a real collection sheet. Do not invent observations. Use blank responses for unfinished runs. Labels, units and ids must survive CSV round trips. A source edit must leave the browser's source-save bridge operational.
  4. Import original bytes with prepareImport / commitImport, or use recordMeasurement with a reason for a direct observation or correction. Do not quietly change units, sample identity, settings or response values. Exclude only with a defensible recorded reason; keep the value and compare the sensitivity fit using every recorded training measurement.
  5. Use model terms consistent with the question and design. Inspect residual structure and independent units, not only the fit statistic. Post-data term changes are exploratory. Explain numeric coding, category references, block effects and pointwise interval assumptions.
  6. Compare settings within the planned ranges. A model mean interval differs from a new-observation prediction interval. Preserve their covariance when comparing two predictions. Do not call the best observed or modeled candidate a proven optimum.
  7. Freeze predictions before collecting a confirmation phase. Keep its new data out of the model it evaluates. Append an extension explicitly when new data should instead improve the model. A protocol change requires a separate experiment while retaining the previous source and data.
  8. Run node tools/check.mjs, exercise real browser inputs and export with node tools/export.mjs delivery. Reopen HTML/source/ZIP, verify raw data, independently reproduce calculations and visually inspect PDFs. Helpers never establish empirical validity or set ready.
Show full SKILL.md (91 more words)Show less

Update bench/DESIGN.md and .harness/verdict.json with actual evidence and the remaining work. A verified collection plan may be ready for collection even when no measured result exists. Never turn an acceptance fixture or a successful build into a claim of real experimental evidence.

The browser needs no account, cloud computation or Python. Managed Node and package-local build utilities are installed by setup; LAB_DSH_DIR resolves them in materialized workspaces. The optional independent Python reader has pinned requirements in the exported kit. Retain complete numerical-library licenses and the original Signal logo/icon in portable deliveries.

© autonomous-ai, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (scripts, references) in store/agents/lab-bench/skills/bench of autonomous-ai/openharness.

  • SKILL.md
  • references/project.md
  • scripts/seed-verdict.sh

Open the folder on GitHubat commit 54a1f1b

Compare with similar skills

Bench next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Bench compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Bench this skillautonomous-ai/openharness1.1k—~900Automated safety check: PassMIT
Instrument Data To Allotropeaws-samples/amazon-bedrock-agents-healthcare-lifesciences2742 repos~2.7kAutomated safety check: PassApache-2.0
Sector Analysttradermonty/claude-trading-skills3k1 repos~2.3kAutomated safety check: PassMIT
Cliare Artifact Reviewmodiqo/cliare469—~2.4kAutomated safety check: PassApache-2.0
Convert Fileduckdb/duckdb-skills5991 repos~720Automated safety check: NotesMIT
Research Integrity Auditxuzhougeng/wisp-science1k—~2.6kAutomated safety check: PassAGPL-3.0

Similar skills

  • Instrument Data To Allotrope

    aws-samples/amazon-bedrock-agents-healthcare-lifesciences

    Official

    Convert laboratory instrument output files (PDF, CSV, Excel, TXT) to Allotrope Simple Model (ASM) JSON format or flattened 2D CSV.

    274 GitHub starsUsed in 2 repos~2.7k tokens
    Documents & OfficeAuto-check passed
  • Sector Analyst

    tradermonty/claude-trading-skills

    This skill should be used when analyzing sector rotation patterns and market cycle positioning.

    3k GitHub starsUsed in 1 repo~2.3k tokens
    Documents & OfficeAuto-check passed
  • A skill your agent uses when reviewing a CLIARE measurement artifact directory, explaining score changes, triaging issues, finding evidence, or proposing CLI remediation work from artifact-map.json…

    469 GitHub stars~2.4k tokensUpdated 2 mo ago
    Documents & OfficeAuto-check passed
  • Convert File

    duckdb/duckdb-skills

    Official

    Convert any data file to another format: CSV, Parquet, JSON, Excel, GeoJSON, and more.

    599 GitHub starsUsed in 1 repo~720 tokens
    Documents & OfficeAuto-check: notes
  • Research Integrity Audit

    xuzhougeng/wisp-science

    学术审查 / research-integrity screening of a manuscript's figures and reported numbers.

    1k GitHub stars~2.6k tokensUpdated today
    Documents & OfficeAuto-check passed
  • Best Skills

    LinklyAI/best-skills

    Daily cross-platform rankings of AI agent skills (skills.sh, ClawHub, Tencent SkillHub, GitHub, X/HN/Bluesky).

    630 GitHub stars~575 tokensUpdated yesterday
    Documents & OfficeAuto-check passed

More from autonomous-ai/openharness

All 99 skills in this repo
  • G-code Slicer Tool

    autonomous-ai/openharness

    Slices 3D mesh files into printer-profiled plain G-code through real slicer CLIs, with backend discovery, input inspection, dry runs and static validation.

    1.1k GitHub stars~1.5k tokensUpdated today
    Auto-check passed
  • Home Assistant Automation Builder

    autonomous-ai/openharness

    Turns a home-automation request into standard, testable automations.yaml, run against Home Assistant Core's real triggers and verified with its own trace tool.

    1.1k GitHub stars~1.2k tokensUpdated today
    Auto-check passed
  • Score Music Composer

    autonomous-ai/openharness

    Turns a musical brief into LilyPond concert-pitch music, checked parts for each instrument and a playable practice pack.

    1.1k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • OrcaSlicer 3MF and G-code Workflow

    autonomous-ai/openharness

    Turns an STL and explicit printer and material requirements into compared OrcaSlicer plans, an editable 3MF project, checked G-code and a portable handoff.

    1.1k GitHub stars~1.5k tokensUpdated today
    Auto-check passed
  • Sheets and Docs Report Builder

    autonomous-ai/openharness

    Builds an editable DOCX report, a formula-driven XLSX workbook and a fresh LibreOffice PDF preview from one structured source file, then checks them together.

    1.1k GitHub stars~708 tokensUpdated today
    Auto-check passed
  • Bambu Labs

    autonomous-ai/openharness

    Dry-run, upload, and cautiously initiate local Bambu Lab print jobs from validated plain .gcode, using Bambu LAN FTPS/MQTT handoffs.

    1.1k GitHub starsUsed in 1 repo~2.1k tokens
    Auto-check: warnings

Questions about Bench

What does Bench do?

Plan real experiments and analyze supplied continuous measurements in Lab Bench. Bench is an agent skill from autonomous-ai/openharness. Plan real experiments and analyze supplied continuous measurements in Lab Bench.

When should I use Bench?

Bench fits situations like: defining factors and independent runs; randomized collection sheets; measurement CSV imports; experimental uncertainty and diagnostics.

How do I install Bench in Claude Code?

Run `npx skills add autonomous-ai/openharness --skill bench -a claude-code`. Or copy the skill folder (store/agents/lab-bench/skills/bench in autonomous-ai/openharness) into .claude/skills/bench in your project. Claude Code loads it when a task matches its description.

How do I install Bench in Codex?

Run `npx skills add autonomous-ai/openharness --skill bench -a codex`. Or copy the skill folder (store/agents/lab-bench/skills/bench in autonomous-ai/openharness) into .agents/skills/bench in your project. Codex loads it when a task matches its description.

Can I use Bench in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add autonomous-ai/openharness --skill bench -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bench, .gemini/skills/bench, .github/skills/bench and .opencode/skills/bench in your project.

What does Bench need to run?

Going by SKILL.md and its folder, Bench needs a shell for the scripts in its folder and the command-line tools its instructions call (node). Our summary lists: Python 3; A Bash shell.

Does Bench access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Bench safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Bench use?

Bench is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Bench use?

About 900 tokens (SKILL.md is roughly 3.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2k tokens, read only when the agent opens those files.

What are the alternatives to Bench?

Skills that share tags, products or a category with Bench: Instrument Data To Allotrope (aws-samples/amazon-bedrock-agents-healthcare-lifesciences, 274 stars), Sector Analyst (tradermonty/claude-trading-skills, 3k stars), Cliare Artifact Review (modiqo/cliare, 469 stars) and Convert File (duckdb/duckdb-skills, 599 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Bench?

autonomous-ai (a GitHub organization) maintains it in autonomous-ai/openharness, which has 1,137 GitHub stars. The repository holds 99 skills in this directory. The repository was last updated on October 7, 2026.

Source: autonomous-ai/openharness on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.