Agent skill

Weavebench Cua Reproduce

by AMAP-ML in AMAP-ML/LongHorizon-Harness

Reproduce CUA-Harness experiments on WeaveBench from a GitHub checkout.

MITAuto-check passedTesting & QA

Install Weavebench Cua Reproduce

skills CLI
$ npx skills add AMAP-ML/LongHorizon-Harness --skill weavebench-cua-reproduce -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install AMAP-ML/LongHorizon-Harness weavebench-cua-reproduce --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/AMAP-ML/LongHorizon-Harness.git skills-src && mkdir -p .claude/skills && cp -r skills-src/eval/WeaveBench-harness/skills/weavebench-cua-reproduce .claude/skills/weavebench-cua-reproduce && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
weavebench-cua-reproduce
GitHub stars
1.7k
Token cost
~1.6k tokens
SKILL.md length
481 words
Files
7 (incl. scripts, references)
Skills in repo
1
Repo updated
First seen
Licence
MIT

At a glance

Reproduce CUA-Harness experiments on WeaveBench from a GitHub checkout.

  • Works in 7 steps: Locate the project root. It must contain… → Read references/configuration.md when… → Read references/assets.md before… → …
  • The user wants an AI coding agent to set up dependencies
  • SKILL.md covers Supported Scope, Start With Intake, Workflow and Helper Script, plus 4 more sections
  • Runs Shell scripts from its folder; needs WEAVEBENCH_LITELLM_KEY

What it does

Weavebench Cua Reproduce is an agent skill from AMAP-ML/LongHorizon-Harness. Reproduce CUA-Harness experiments on WeaveBench from a GitHub checkout. Use when the user wants an AI coding agent to set up dependencies, download WeaveBench assets, prepare the 120G VM, configure Qwen/Anthropic-compatible APIs, run smoke tests, launch full or subset evaluations, inspect logs, or summarize scores for this repository.

Its SKILL.md is about 1.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 9 other files, including scripts and reference files (for example `agents/openai.yaml`, `references/assets.md` and `references/configuration.md`).

It sits in Testing & QA, covering QA and bug reports. It works with Qwen, GitHub, DeepSeek and Docker. The repository describes itself as: The long-horizon computer-use harness. Run AI agents across desktop apps and the CLI for extended periods while preserving task state and making reliable progress on complex… The licence is MIT.

When your agent uses it

  • The user wants an AI coding agent to set up dependencies
  • Download WeaveBench assets
  • Prepare the 120G VM
  • Configure Qwen/Anthropic-compatible APIs

Example prompts

  • “/weavebench-cua-reproduce”

Requirements

  • A Bash shell
  • Docker
  • A credential in WEAVEBENCH_LITELLM_KEY
  • A credential in YOUR_API_KEY

Workflow steps

7 steps, taken from the first numbered list in SKILL.md.

  1. Locate the project root. It must contain both WeaveBench/ and cua-harness/.
  2. Read references/configuration.md when you need exact defaults, environment variables, or paths.
  3. Read references/assets.md before downloading or validating WeaveBench tasks, runtime assets, judge templates, or VM files.
  4. Read references/verify.md before reporting whether setup or a run is actually verified.
  5. Read references/troubleshooting.md when setup, VM, API, image proxy, warmup, or judge errors occur.
  6. Ask the user for missing API details only when they are required and not discoverable from the environment.
  7. Run commands from the project root unless a command explicitly changes into WeaveBench/.

What it can do on your machine

Read from SKILL.md and the folder at commit a1dd930. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Shell), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • WEAVEBENCH_LITELLM_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Weavebench Cua Reproduce loads about 1.6k tokens when it runs, and up to ~4.7k if it reads all its reference files. Until then it costs about 90 tokens; SKILL.md has 481 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~90
When it runs · the whole SKILL.md, loaded when a task matches
~1.6k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~4.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from AMAP-ML/LongHorizon-Harness at commit a1dd930, republished under its MIT licence (© AMAP-ML). 481 words, ~1,635 tokens.

Download SKILL.mdSave it as .claude/skills/weavebench-cua-reproduce/SKILL.md (or your agent's skills folder). This skill also uses 6 other files; get the full folder from GitHub.
name
weavebench-cua-reproduce
description
Reproduce CUA-Harness experiments on WeaveBench from a GitHub checkout. Use when the user wants an AI coding agent to set up dependencies, download WeaveBench assets, prepare the 120G VM, configure Qwen/Anthropic-compatible APIs, run smoke tests, launch full or subset evaluations, inspect logs, or summarize scores for this repository.

WeaveBench CUA-Harness Reproduction

Use this skill to guide an agent through a complete, reproducible WeaveBench CUA-Harness run. The expected repository layout is:

text
WeaveBench-harness/
  WeaveBench/
  cua-harness/
  skills/weavebench-cua-reproduce/

Do not invent alternate launch commands. Prefer the bundled helper script and the project scripts under WeaveBench/scripts/.

Supported Scope

Supported automation:

  • Local Docker/KVM WeaveBench evaluation.
  • Qwen 3.7-Plus through an Anthropic-compatible endpoint.
  • Claude Code execution backend through cua_harness_claudecode.
  • OpenClaw judge setup and verification.
  • Official WeaveBench assets downloaded from the HuggingFace dataset.
  • A 120G copy of the official Ubuntu qcow2 plus per-task rootfs growth.

Not automated:

  • Cloud VM provisioning.
  • Public image-hosting service deployment.
  • New model-provider adapters beyond the existing Anthropic-compatible path.
  • Non-Claude-Code execution backend reproduction.
  • Deleting or pruning existing results.

If the user requests unsupported automation, explain the boundary and provide the closest supported local Docker/KVM path.

Start With Intake

Before doing setup work on a new machine, run or mentally follow:

bash
./skills/weavebench-cua-reproduce/scripts/reproduce.sh intake

Confirm:

  • Linux Docker/KVM is available or installable.
  • The user accepts large downloads and creation of Ubuntu_120G.qcow2.
  • The model endpoint and API key are available.
  • Image proxy should be enabled with public upload/show URLs, or disabled because the endpoint accepts large base64 screenshots.
  • The desired scope is doctor, smoke, subset, or full 114-task evaluation.
  • Long-running/full evaluation is allowed, especially if other experiments may already be active.

Workflow

  1. Locate the project root. It must contain both WeaveBench/ and cua-harness/.
  2. Read references/configuration.md when you need exact defaults, environment variables, or paths.
  3. Read references/assets.md before downloading or validating WeaveBench tasks, runtime assets, judge templates, or VM files.
  4. Read references/verify.md before reporting whether setup or a run is actually verified.
  5. Read references/troubleshooting.md when setup, VM, API, image proxy, warmup, or judge errors occur.
  6. Ask the user for missing API details only when they are required and not discoverable from the environment.
  7. Run commands from the project root unless a command explicitly changes into WeaveBench/.
Show full SKILL.md (173 more words)Show less

Helper Script

Use:

bash
./skills/weavebench-cua-reproduce/scripts/reproduce.sh <command>

Commands:

text
install     Install local WeaveBench package and OpenClaw if needed.
download    Download task files, Claude Code runtime, judge template, and VM.
vm120g      Create cache/vm/Ubuntu_120G.qcow2 from the official Ubuntu.qcow2.
doctor      Run the read-only environment checker.
intake      Print setup questions for a new machine.
status      Print a non-destructive setup status report.
smoke       Run a one-task, one-round smoke test.
full        Launch the full 114-task evaluation in tmux.
stats       Summarize scores for a result directory.
plan        Print the minimal manual command sequence.

For a fresh machine, the usual sequence is:

bash
./skills/weavebench-cua-reproduce/scripts/reproduce.sh install
./skills/weavebench-cua-reproduce/scripts/reproduce.sh download
./skills/weavebench-cua-reproduce/scripts/reproduce.sh vm120g
./skills/weavebench-cua-reproduce/scripts/reproduce.sh status
./skills/weavebench-cua-reproduce/scripts/reproduce.sh doctor
./skills/weavebench-cua-reproduce/scripts/reproduce.sh smoke
./skills/weavebench-cua-reproduce/scripts/reproduce.sh full

Required User Configuration

Before doctor, smoke, or full, ensure these are set:

bash
export WEAVEBENCH_LITELLM_KEY="YOUR_API_KEY"
export WEAVEBENCH_LITELLM_BASE_URL="https://YOUR_ANTHROPIC_COMPATIBLE_ENDPOINT/v1"

If the provider cannot accept large base64 image payloads, enable image URL proxy:

bash
export WEAVEBENCH_IMAGE_PROXY=1
export WEAVEBENCH_IMAGE_PROXY_UPLOAD_URL_TPL="https://YOUR_IMAGE_UPLOAD_ENDPOINT/{id}"
export WEAVEBENCH_IMAGE_PROXY_SHOW_URL_TPL="https://YOUR_PUBLIC_IMAGE_URL/{id}.png"
export WEAVEBENCH_IMAGE_PROXY_UPLOAD_MODE="raw"

If the provider can directly accept large base64 screenshots:

bash
export WEAVEBENCH_IMAGE_PROXY=0

Launch Rules

  • Use the 120G VM copy by default: WeaveBench/cache/vm/Ubuntu_120G.qcow2.
  • Keep WEAVEBENCH_GROW_ROOTFS=1 unless the user explicitly disables it.
  • Use qwen3.7-plus, cua_harness_claudecode, GUI mode, 25 CUA rounds, and 5 concurrent VM environments for full evaluation unless the user asks for a subset.
  • Use smoke before full on a new machine.
  • Run long evaluations in tmux.
  • Never delete cache/, results/, logs/, Docker volumes, or judge workspaces unless the user explicitly asks.
  • When reporting API keys, redact them unless the user explicitly requests full values.
  • At the end of setup, classify each surface as configured and verified, configured but not smoke-tested, not configured, or blocked awaiting user action.

Common Subsets

Run one domain:

bash
export WEAVEBENCH_DOMAINS="WEB"
./skills/weavebench-cua-reproduce/scripts/reproduce.sh full

Run one task:

bash
export WEAVEBENCH_DOMAINS="WEB"
export WEAVEBENCH_TASK_FILTER="WEB_task_10_lighthouse"
export WEAVEBENCH_NUM_ENVS=1
./skills/weavebench-cua-reproduce/scripts/reproduce.sh full

Summarize a run:

bash
./skills/weavebench-cua-reproduce/scripts/reproduce.sh stats \
  WeaveBench/results/<run_name>/gui/qwen3.7-plus

Final Report Template

End setup or launch tasks with this shape:

text
Environment: configured and verified / configured but not smoke-tested / not configured / blocked awaiting user action
Assets: ...
120G VM: ...
Model API: ...
Image proxy: ...
Judge: ...
Smoke test: passed / failed / skipped
Full run: launched in tmux <name> / not launched
Provenance: run_provenance.json written / not applicable
Remaining blockers: ...

© AMAP-ML, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 6 other files (scripts, references) in eval/WeaveBench-harness/skills/weavebench-cua-reproduce of AMAP-ML/LongHorizon-Harness.

  • SKILL.md
  • agents/openai.yaml
  • references/assets.md
  • references/configuration.md
  • references/troubleshooting.md
  • references/verify.md
  • scripts/reproduce.sh

Open the folder on GitHubat commit a1dd930

Compare with similar skills

Weavebench Cua Reproduce next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Weavebench Cua Reproduce compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Weavebench Cua Reproduce this skillAMAP-ML/LongHorizon-Harness1.7k—~1.6kAutomated safety check: PassMIT
Issue WriterNVIDIA/container-canary308—~1.1kAutomated safety check: PassApache-2.0
Plugin TestNanmiCoder/dsh-auto-mode1641 repos~2.9kAutomated safety check: PassMIT
Weave Router Local Testingweave-os/router5.6k—~3.1kAutomated safety check: NotesApache-2.0
Serving LLMs On Instinctamd/skills408—~4kAutomated safety check: NotesMIT
Vllm Daily PR Issue Trackerascend-ai-coding/awesome-ascend-skills174—~731Automated safety check: PassNone

Similar skills

  • Issue Writer

    NVIDIA/container-canary

    Official

    Draft and revise concise, human-focused GitHub issues for pytest-kind-ng.

    308 GitHub stars~1.1k tokensUpdated 1 mo ago
    Testing & QAAuto-check passed
  • Plugin Test

    NanmiCoder/dsh-auto-mode

    A skill your agent uses when writing or reviewing tests for DeepSeek Harness plugins, external DSH plugin packages, or package changes in the deepseek-harness repository.

    164 GitHub starsUsed in 1 repo~2.9k tokens
    Testing & QAAuto-check passed
  • Stands up the Weave model router in Docker Compose and drives it with claude -p against a real or mocked upstream to reproduce and verify routing and streaming behavior.

    5.6k GitHub stars~3.1k tokensUpdated today
    DevOps & CloudAuto-check: notes
  • Serves AI models on AMD Instinct GPU hardware using vLLM. An agent skill from amd/skills.

    408 GitHub stars~4k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check: notes
  • Vllm Daily PR Issue Tracker

    ascend-ai-coding/awesome-ascend-skills

    Track daily PRs and Issues from vllm-project/vllm and vllm-project/vllm-ascend, filter by model (DeepSeek/Qwen/GLM/MiniMax/Kimi) and tech topics (PD disaggregation, MTP, quantization, graph mode…

    174 GitHub stars~731 tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • DeerFlow Smoke Test

    bytedance/deer-flow

    Walks through an end-to-end smoke test of a DeerFlow deployment: pull the latest code, deploy with Docker or locally, verify services, run health checks and write a report.

    84k GitHub stars~2.5k tokensUpdated today
    Testing & QAAuto-check: notes

Categories

Questions about Weavebench Cua Reproduce

What does Weavebench Cua Reproduce do?

Reproduce CUA-Harness experiments on WeaveBench from a GitHub checkout. Weavebench Cua Reproduce is an agent skill from AMAP-ML/LongHorizon-Harness. Reproduce CUA-Harness experiments on WeaveBench from a GitHub checkout.

When should I use Weavebench Cua Reproduce?

Weavebench Cua Reproduce fits situations like: the user wants an AI coding agent to set up dependencies; download WeaveBench assets; prepare the 120G VM; configure Qwen/Anthropic-compatible APIs.

How do I install Weavebench Cua Reproduce in Claude Code?

Run `npx skills add AMAP-ML/LongHorizon-Harness --skill weavebench-cua-reproduce -a claude-code`. Or copy the skill folder (eval/WeaveBench-harness/skills/weavebench-cua-reproduce in AMAP-ML/LongHorizon-Harness) into .claude/skills/weavebench-cua-reproduce in your project. Claude Code loads it when a task matches its description.

How do I install Weavebench Cua Reproduce in Codex?

Run `npx skills add AMAP-ML/LongHorizon-Harness --skill weavebench-cua-reproduce -a codex`. Or copy the skill folder (eval/WeaveBench-harness/skills/weavebench-cua-reproduce in AMAP-ML/LongHorizon-Harness) into .agents/skills/weavebench-cua-reproduce in your project. Codex loads it when a task matches its description.

Can I use Weavebench Cua Reproduce in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add AMAP-ML/LongHorizon-Harness --skill weavebench-cua-reproduce -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/weavebench-cua-reproduce, .gemini/skills/weavebench-cua-reproduce, .github/skills/weavebench-cua-reproduce and .opencode/skills/weavebench-cua-reproduce in your project.

What does Weavebench Cua Reproduce need to run?

Going by SKILL.md and its folder, Weavebench Cua Reproduce needs a shell for the scripts in its folder and credentials named WEAVEBENCH_LITELLM_KEY. Our summary lists: A Bash shell; Docker; A credential in WEAVEBENCH_LITELLM_KEY; A credential in YOUR_API_KEY.

Does Weavebench Cua Reproduce access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Weavebench Cua Reproduce safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Weavebench Cua Reproduce use?

Weavebench Cua Reproduce is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Weavebench Cua Reproduce use?

About 1.6k tokens (SKILL.md is roughly 6.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3.1k tokens, read only when the agent opens those files.

What are the alternatives to Weavebench Cua Reproduce?

Skills that share tags, products or a category with Weavebench Cua Reproduce: Issue Writer (NVIDIA/container-canary, 308 stars), Plugin Test (NanmiCoder/dsh-auto-mode, 164 stars), Weave Router Local Testing (weave-os/router, 5.6k stars) and Serving LLMs On Instinct (amd/skills, 408 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Weavebench Cua Reproduce?

AMAP-ML (a GitHub organization) maintains it in AMAP-ML/LongHorizon-Harness, which has 1,715 GitHub stars. The repository was last updated on August 20, 2026.

Source: AMAP-ML/LongHorizon-Harness on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.