Issue Writer
NVIDIA/container-canary
Draft and revise concise, human-focused GitHub issues for pytest-kind-ng.
Reproduce CUA-Harness experiments on WeaveBench from a GitHub checkout.
$ npx skills add AMAP-ML/LongHorizon-Harness --skill weavebench-cua-reproduce -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install AMAP-ML/LongHorizon-Harness weavebench-cua-reproduce --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/AMAP-ML/LongHorizon-Harness.git skills-src && mkdir -p .claude/skills && cp -r skills-src/eval/WeaveBench-harness/skills/weavebench-cua-reproduce .claude/skills/weavebench-cua-reproduce && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "weavebench-cua-reproduce" agent skill from https://github.com/AMAP-ML/LongHorizon-Harness/tree/main/eval/WeaveBench-harness/skills/weavebench-cua-reproduce into .claude/skills/weavebench-cua-reproduce/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "weavebench-cua-reproduce", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/AMAP-ML/LongHorizon-Harness/tree/main/eval/WeaveBench-harness/skills/weavebench-cua-reproduceType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add AMAP-ML/LongHorizon-Harness --skill weavebench-cua-reproduce -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install AMAP-ML/LongHorizon-Harness weavebench-cua-reproduce --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/AMAP-ML/LongHorizon-Harness.git skills-src && mkdir -p .agents/skills && cp -r skills-src/eval/WeaveBench-harness/skills/weavebench-cua-reproduce .agents/skills/weavebench-cua-reproduce && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "weavebench-cua-reproduce" agent skill from https://github.com/AMAP-ML/LongHorizon-Harness/tree/main/eval/WeaveBench-harness/skills/weavebench-cua-reproduce into .agents/skills/weavebench-cua-reproduce/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "weavebench-cua-reproduce", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add AMAP-ML/LongHorizon-Harness --skill weavebench-cua-reproduce -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install AMAP-ML/LongHorizon-Harness weavebench-cua-reproduce --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/AMAP-ML/LongHorizon-Harness.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/eval/WeaveBench-harness/skills/weavebench-cua-reproduce .cursor/skills/weavebench-cua-reproduce && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "weavebench-cua-reproduce" agent skill from https://github.com/AMAP-ML/LongHorizon-Harness/tree/main/eval/WeaveBench-harness/skills/weavebench-cua-reproduce into .cursor/skills/weavebench-cua-reproduce/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "weavebench-cua-reproduce", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/AMAP-ML/LongHorizon-Harness.git --path eval/WeaveBench-harness/skills/weavebench-cua-reproduce--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add AMAP-ML/LongHorizon-Harness --skill weavebench-cua-reproduce -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install AMAP-ML/LongHorizon-Harness weavebench-cua-reproduce --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/AMAP-ML/LongHorizon-Harness.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/eval/WeaveBench-harness/skills/weavebench-cua-reproduce .gemini/skills/weavebench-cua-reproduce && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "weavebench-cua-reproduce" agent skill from https://github.com/AMAP-ML/LongHorizon-Harness/tree/main/eval/WeaveBench-harness/skills/weavebench-cua-reproduce into .gemini/skills/weavebench-cua-reproduce/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "weavebench-cua-reproduce", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install AMAP-ML/LongHorizon-Harness weavebench-cua-reproduceInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add AMAP-ML/LongHorizon-Harness --skill weavebench-cua-reproduce -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/AMAP-ML/LongHorizon-Harness.git skills-src && mkdir -p .github/skills && cp -r skills-src/eval/WeaveBench-harness/skills/weavebench-cua-reproduce .github/skills/weavebench-cua-reproduce && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "weavebench-cua-reproduce" agent skill from https://github.com/AMAP-ML/LongHorizon-Harness/tree/main/eval/WeaveBench-harness/skills/weavebench-cua-reproduce into .github/skills/weavebench-cua-reproduce/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "weavebench-cua-reproduce", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add AMAP-ML/LongHorizon-Harness --skill weavebench-cua-reproduce -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install AMAP-ML/LongHorizon-Harness weavebench-cua-reproduce --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/AMAP-ML/LongHorizon-Harness.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/eval/WeaveBench-harness/skills/weavebench-cua-reproduce .opencode/skills/weavebench-cua-reproduce && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "weavebench-cua-reproduce" agent skill from https://github.com/AMAP-ML/LongHorizon-Harness/tree/main/eval/WeaveBench-harness/skills/weavebench-cua-reproduce into .opencode/skills/weavebench-cua-reproduce/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "weavebench-cua-reproduce", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
weavebench-cua-reproduceReproduce CUA-Harness experiments on WeaveBench from a GitHub checkout.
Weavebench Cua Reproduce is an agent skill from AMAP-ML/LongHorizon-Harness. Reproduce CUA-Harness experiments on WeaveBench from a GitHub checkout. Use when the user wants an AI coding agent to set up dependencies, download WeaveBench assets, prepare the 120G VM, configure Qwen/Anthropic-compatible APIs, run smoke tests, launch full or subset evaluations, inspect logs, or summarize scores for this repository.
Its SKILL.md is about 1.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 9 other files, including scripts and reference files (for example `agents/openai.yaml`, `references/assets.md` and `references/configuration.md`).
It sits in Testing & QA, covering QA and bug reports. It works with Qwen, GitHub, DeepSeek and Docker. The repository describes itself as: The long-horizon computer-use harness. Run AI agents across desktop apps and the CLI for extended periods while preserving task state and making reliable progress on complex… The licence is MIT.
7 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit a1dd930. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 1 file in scripts/ (Shell), which the agent can run.
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
WEAVEBENCH_LITELLM_KEYFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Weavebench Cua Reproduce loads about 1.6k tokens when it runs, and up to ~4.7k if it reads all its reference files. Until then it costs about 90 tokens; SKILL.md has 481 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from AMAP-ML/LongHorizon-Harness at commit a1dd930, republished under its MIT licence (© AMAP-ML). 481 words, ~1,635 tokens.
.claude/skills/weavebench-cua-reproduce/SKILL.md (or your agent's skills folder). This skill also uses 6 other files; get the full folder from GitHub.Use this skill to guide an agent through a complete, reproducible WeaveBench CUA-Harness run. The expected repository layout is:
WeaveBench-harness/
WeaveBench/
cua-harness/
skills/weavebench-cua-reproduce/Do not invent alternate launch commands. Prefer the bundled helper script and the project scripts under WeaveBench/scripts/.
Supported automation:
cua_harness_claudecode.Not automated:
If the user requests unsupported automation, explain the boundary and provide the closest supported local Docker/KVM path.
Before doing setup work on a new machine, run or mentally follow:
./skills/weavebench-cua-reproduce/scripts/reproduce.sh intakeConfirm:
Ubuntu_120G.qcow2.doctor, smoke, subset, or full 114-task evaluation.WeaveBench/ and cua-harness/.references/configuration.md when you need exact defaults, environment variables, or paths.references/assets.md before downloading or validating WeaveBench tasks, runtime assets, judge templates, or VM files.references/verify.md before reporting whether setup or a run is actually verified.references/troubleshooting.md when setup, VM, API, image proxy, warmup, or judge errors occur.WeaveBench/.Use:
./skills/weavebench-cua-reproduce/scripts/reproduce.sh <command>Commands:
install Install local WeaveBench package and OpenClaw if needed.
download Download task files, Claude Code runtime, judge template, and VM.
vm120g Create cache/vm/Ubuntu_120G.qcow2 from the official Ubuntu.qcow2.
doctor Run the read-only environment checker.
intake Print setup questions for a new machine.
status Print a non-destructive setup status report.
smoke Run a one-task, one-round smoke test.
full Launch the full 114-task evaluation in tmux.
stats Summarize scores for a result directory.
plan Print the minimal manual command sequence.For a fresh machine, the usual sequence is:
./skills/weavebench-cua-reproduce/scripts/reproduce.sh install
./skills/weavebench-cua-reproduce/scripts/reproduce.sh download
./skills/weavebench-cua-reproduce/scripts/reproduce.sh vm120g
./skills/weavebench-cua-reproduce/scripts/reproduce.sh status
./skills/weavebench-cua-reproduce/scripts/reproduce.sh doctor
./skills/weavebench-cua-reproduce/scripts/reproduce.sh smoke
./skills/weavebench-cua-reproduce/scripts/reproduce.sh fullBefore doctor, smoke, or full, ensure these are set:
export WEAVEBENCH_LITELLM_KEY="YOUR_API_KEY"
export WEAVEBENCH_LITELLM_BASE_URL="https://YOUR_ANTHROPIC_COMPATIBLE_ENDPOINT/v1"If the provider cannot accept large base64 image payloads, enable image URL proxy:
export WEAVEBENCH_IMAGE_PROXY=1
export WEAVEBENCH_IMAGE_PROXY_UPLOAD_URL_TPL="https://YOUR_IMAGE_UPLOAD_ENDPOINT/{id}"
export WEAVEBENCH_IMAGE_PROXY_SHOW_URL_TPL="https://YOUR_PUBLIC_IMAGE_URL/{id}.png"
export WEAVEBENCH_IMAGE_PROXY_UPLOAD_MODE="raw"If the provider can directly accept large base64 screenshots:
export WEAVEBENCH_IMAGE_PROXY=0WeaveBench/cache/vm/Ubuntu_120G.qcow2.WEAVEBENCH_GROW_ROOTFS=1 unless the user explicitly disables it.qwen3.7-plus, cua_harness_claudecode, GUI mode, 25 CUA rounds, and 5 concurrent VM environments for full evaluation unless the user asks for a subset.smoke before full on a new machine.cache/, results/, logs/, Docker volumes, or judge workspaces unless the user explicitly asks.configured and verified, configured but not smoke-tested, not configured, or blocked awaiting user action.Run one domain:
export WEAVEBENCH_DOMAINS="WEB"
./skills/weavebench-cua-reproduce/scripts/reproduce.sh fullRun one task:
export WEAVEBENCH_DOMAINS="WEB"
export WEAVEBENCH_TASK_FILTER="WEB_task_10_lighthouse"
export WEAVEBENCH_NUM_ENVS=1
./skills/weavebench-cua-reproduce/scripts/reproduce.sh fullSummarize a run:
./skills/weavebench-cua-reproduce/scripts/reproduce.sh stats \
WeaveBench/results/<run_name>/gui/qwen3.7-plusEnd setup or launch tasks with this shape:
Environment: configured and verified / configured but not smoke-tested / not configured / blocked awaiting user action
Assets: ...
120G VM: ...
Model API: ...
Image proxy: ...
Judge: ...
Smoke test: passed / failed / skipped
Full run: launched in tmux <name> / not launched
Provenance: run_provenance.json written / not applicable
Remaining blockers: ...© AMAP-ML, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 6 other files (scripts, references) in eval/WeaveBench-harness/skills/weavebench-cua-reproduce of AMAP-ML/LongHorizon-Harness.
Open the folder on GitHubat commit a1dd930
Weavebench Cua Reproduce next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Weavebench Cua Reproduce this skillAMAP-ML/LongHorizon-Harness | 1.7k | — | ~1.6k | Automated safety check: Pass | MIT | |
| Issue WriterNVIDIA/container-canary | 308 | — | ~1.1k | Automated safety check: Pass | Apache-2.0 | |
| Plugin TestNanmiCoder/dsh-auto-mode | 164 | 1 repos | ~2.9k | Automated safety check: Pass | MIT | |
| Weave Router Local Testingweave-os/router | 5.6k | — | ~3.1k | Automated safety check: Notes | Apache-2.0 | |
| Serving LLMs On Instinctamd/skills | 408 | — | ~4k | Automated safety check: Notes | MIT | |
| Vllm Daily PR Issue Trackerascend-ai-coding/awesome-ascend-skills | 174 | — | ~731 | Automated safety check: Pass | None |
NVIDIA/container-canary
Draft and revise concise, human-focused GitHub issues for pytest-kind-ng.
NanmiCoder/dsh-auto-mode
A skill your agent uses when writing or reviewing tests for DeepSeek Harness plugins, external DSH plugin packages, or package changes in the deepseek-harness repository.
weave-os/router
Stands up the Weave model router in Docker Compose and drives it with claude -p against a real or mocked upstream to reproduce and verify routing and streaming behavior.
amd/skills
Serves AI models on AMD Instinct GPU hardware using vLLM. An agent skill from amd/skills.
ascend-ai-coding/awesome-ascend-skills
Track daily PRs and Issues from vllm-project/vllm and vllm-project/vllm-ascend, filter by model (DeepSeek/Qwen/GLM/MiniMax/Kimi) and tech topics (PD disaggregation, MTP, quantization, graph mode…
bytedance/deer-flow
Walks through an end-to-end smoke test of a DeerFlow deployment: pull the latest code, deploy with Docker or locally, verify services, run health checks and write a report.
Categories
Reproduce CUA-Harness experiments on WeaveBench from a GitHub checkout. Weavebench Cua Reproduce is an agent skill from AMAP-ML/LongHorizon-Harness. Reproduce CUA-Harness experiments on WeaveBench from a GitHub checkout.
Weavebench Cua Reproduce fits situations like: the user wants an AI coding agent to set up dependencies; download WeaveBench assets; prepare the 120G VM; configure Qwen/Anthropic-compatible APIs.
Run `npx skills add AMAP-ML/LongHorizon-Harness --skill weavebench-cua-reproduce -a claude-code`. Or copy the skill folder (eval/WeaveBench-harness/skills/weavebench-cua-reproduce in AMAP-ML/LongHorizon-Harness) into .claude/skills/weavebench-cua-reproduce in your project. Claude Code loads it when a task matches its description.
Run `npx skills add AMAP-ML/LongHorizon-Harness --skill weavebench-cua-reproduce -a codex`. Or copy the skill folder (eval/WeaveBench-harness/skills/weavebench-cua-reproduce in AMAP-ML/LongHorizon-Harness) into .agents/skills/weavebench-cua-reproduce in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add AMAP-ML/LongHorizon-Harness --skill weavebench-cua-reproduce -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/weavebench-cua-reproduce, .gemini/skills/weavebench-cua-reproduce, .github/skills/weavebench-cua-reproduce and .opencode/skills/weavebench-cua-reproduce in your project.
Going by SKILL.md and its folder, Weavebench Cua Reproduce needs a shell for the scripts in its folder and credentials named WEAVEBENCH_LITELLM_KEY. Our summary lists: A Bash shell; Docker; A credential in WEAVEBENCH_LITELLM_KEY; A credential in YOUR_API_KEY.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Weavebench Cua Reproduce is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.6k tokens (SKILL.md is roughly 6.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3.1k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Weavebench Cua Reproduce: Issue Writer (NVIDIA/container-canary, 308 stars), Plugin Test (NanmiCoder/dsh-auto-mode, 164 stars), Weave Router Local Testing (weave-os/router, 5.6k stars) and Serving LLMs On Instinct (amd/skills, 408 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
AMAP-ML (a GitHub organization) maintains it in AMAP-ML/LongHorizon-Harness, which has 1,715 GitHub stars. The repository was last updated on August 20, 2026.
Source: AMAP-ML/LongHorizon-Harness on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.