Agent skill

Verify Openhands

by OpenHands in OpenHands/OpenHands

This skill should be used to "verify OpenHands features", "test the Canvas UI like a user", "drive Agent Canvas", "check a UI change in the real app", "create or update the feature map", "run the…

MITAuto-check passed

Install Verify Openhands

skills CLI
$ npx skills add OpenHands/OpenHands --skill verify-openhands -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install OpenHands/OpenHands verify-openhands --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/OpenHands/OpenHands.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/verify-openhands .claude/skills/verify-openhands && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
verify-openhands
GitHub stars
90k
Token cost
~2.9k tokens
SKILL.md length
1,440 words
Files
50 (incl. scripts, references)
Skills in repo
8
Repo updated
First seen
Licence
MIT

At a glance

This skill should be used to "verify OpenHands features", "test the Canvas UI like a user", "drive Agent Canvas", "check a UI change in the real app", "create or update the feature map", "run the…

  • Works in 3 steps: control-openhands (scripts/) is the… → The feature map (references/feature-map/) → Evidence: screenshots, ARIA snapshots…
  • SKILL.md covers The CLI comes first, Launch → doctor → drive →…, LLM budget and Choose the job, plus 1 more section
  • Calls npm and node; needs DEEPSEEK_API_KEY

What it does

Verify Openhands is an agent skill from OpenHands/OpenHands. This skill should be used to "verify OpenHands features", "test the Canvas UI like a user", "drive Agent Canvas", "check a UI change in the real app", "create or update the feature map", "run the daily pass" or "run the weekly feature audit". Ships control-openhands (launch, doctor, browser, LLM profiles, conversations, evidence, map tooling, cleanup) and the maintained map of every user-facing feature.

Its SKILL.md is about 2.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 51 other files, including scripts and reference files (for example `references/adaptation.md`, `references/daily.md` and `references/feature-map/F01-first-run-and-sign-in.md`).

It works with OpenAI. The repository describes itself as: 🙌 OpenHands: AI-Driven Development. The licence is MIT.

Example prompts

  • “verify OpenHands features”
  • “test the Canvas UI like a user”
  • “drive Agent Canvas”
  • “/verify-openhands”

Requirements

  • A credential in DEEPSEEK_API_KEY

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. control-openhands (scripts/) is the lever. It launches this
  2. The feature map (references/feature-map/)
  3. Evidence: screenshots, ARIA snapshots and a pass/fail/blocked/not-run

What it can do on your machine

Read from SKILL.md and the folder at commit 7ea83ba. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/, which the agent can run.

    Shell commands in SKILL.md call:

    • npm
    • node

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use npm, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • DEEPSEEK_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Verify Openhands loads about 2.9k tokens when it runs, and up to ~252k if it reads all its reference files. Until then it costs about 106 tokens; SKILL.md has 1,440 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~106
When it runs · the whole SKILL.md, loaded when a task matches
~2.9k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~252k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from OpenHands/OpenHands at commit 7ea83ba, republished under its MIT licence (© OpenHands). 1,440 words, ~2,949 tokens.

Download SKILL.mdSave it as .claude/skills/verify-openhands/SKILL.md (or your agent's skills folder). This skill also uses 49 other files; get the full folder from GitHub.
name
verify-openhands
description
This skill should be used to "verify OpenHands features", "test the Canvas UI like a user", "drive Agent Canvas", "check a UI change in the real app", "create or update the feature map", "run the daily pass" or "run the weekly feature audit". Ships control-openhands (launch, doctor, browser, LLM profiles, conversations, evidence, map tooling, cleanup) and the maintained map of every user-facing feature.
triggers
/verify-openhands, feature map, verify feature

Verify OpenHands through the real app

Prove what a user sees, not that a route rendered or CI passed. Three parts work together:

  1. control-openhands (scripts/) is the lever. It launches this checkout as an isolated real stack (Agent Server, automation, static frontend, ingress), keeps one browser alive between commands, and turns every user step into a command you can rerun.
  2. The feature map (references/feature-map/) lists every user-facing behavior with stable IDs, user entry points, exact control-openhands recipes, observable results and gotchas.
  3. Evidence: screenshots, ARIA snapshots and a pass/fail/blocked/not-run ledger that survive cleanup.

Neither this skill nor the map authorizes external writes, paid models beyond the budget you were given, or product fixes the user did not ask for.

The CLI comes first

Put it on PATH and read its help before anything else:

sh
export PATH="$PWD/.agents/skills/verify-openhands/scripts:$PATH"
control-openhands --help                # then: control-openhands <command> --help

Requirements: Node >=24 (the launcher's engine), npm dependencies installed (npm ci --ignore-scripts), uv/uvx, and a Chromium Playwright can launch (set CONTROL_OPENHANDS_BROWSER=/path/to/chrome when the pinned browser is not installed; launch then reports the running stack and browser start picks it up). Every command prints one JSON object; exit 0 ok, 1 action failed, 2 usage, 3 environment.

If a user path cannot be driven with the CLI, that is a harness gap: extend scripts/control-openhands.mjs (keep it executable, document the verb in --help and below), prove the new verb live, then write the recipe. Never work around a gap with an untracked one-off script the next agent cannot rerun. Other agents on the machine may be running the same script: edit a copy, run node --check and the tests on it, then move it into place in one step.

Launch → doctor → drive → evidence → cleanup

sh
export OH_VERIFY_RUN=$(control-openhands launch --new --print-run)   # builds if needed; isolated run
control-openhands doctor                       # read-only; must be ok before driving
control-openhands llm preset deepseek          # deepseek-flash (active) + deepseek-pro; key from $DEEPSEEK_API_KEY or --api-key-file
control-openhands onboard --skip               # consent + onboarding (walk it instead when F01 is under test)
control-openhands browser goto /settings/secrets
control-openhands browser testids              # discover handles on the current page
control-openhands browser click 'testid=add-secret-button'
control-openhands browser screenshot --feature F14.create --name form
control-openhands evidence add --feature F14.create --result pass --entry "Settings > Secrets > Add" \
  --expected "add form" --actual "add form" --artifact evidence/F14.create/form.png
control-openhands stop                         # stops only this run; evidence stays
  • Which run. Commands use --run DIR (anywhere on the line), else $OH_VERIFY_RUN, else the only live run. With several live runs they refuse to guess, so export OH_VERIFY_RUN in every shell command when other agents share the machine: an agent that drives someone else's run corrupts both evidence ledgers. For the same reason never pkill -f or killall by pattern; control-openhands stop ends only your run.
  • Launch refuses to start when less than about 2 GB of memory is free: each run holds an Agent Server, automation, a static frontend and Chromium (about 1.5 GB together). Stop runs you are done with; several agents on one machine should each launch --new, export their own OH_VERIFY_RUN, and never touch another agent's run.
  • Launch starts bin/agent-canvas.mjs from this checkout with a private HOME, state, session key and free port block, so it never touches a user's ~/.openhands or another run. launch --new starts a second independent run (for a baseline, or to drive two backends); --public exercises the API-key login screen (control-openhands login). It still binds to 127.0.0.1; it only stops injecting the session key into the page. --sdk-version, --sdk-ref or --sdk-path (and the --automation-* equivalents) choose other backends; record them. Version variables exported in your shell are not forwarded. The block includes the VS Code editor port (ports.vscode) only when the checkout's launcher sets VS Code up (the port is reserved even if the agent-server has no editor binary and nothing listens there). Where that is opt-in (#17660), --vscode turns it on; where the editor is bundled (before #17660, or since #18048), it is always included and --vscode changes nothing.
  • Doctor checks the launcher's process group, ports, served build revision, unauthenticated rejection, authenticated settings, Agent Server pin and the UI's minimum Agent Server version, automation health and a throwaway-tab UI probe. Run it first, after every surprising failure, and before blaming the product. A wedged UI on a healthy stack: capture evidence, browser reload or goto /, retry once.
  • Drive through the real UI with control-openhands browser .... Selectors are testid=, role=button[name="Save"], label=, text=, chained with >>. Use browser testids and browser snapshot to find handles; prefer scoped test IDs and accessible names over CSS. Failures return a hint and a screenshot path; read the screenshot before retrying. A click returns the URL from before any client-side navigation: use click ... --expect-url '<regex>' (or browser wait-url) after every navigating click, because many routes redirect (/settings → /settings/agents, /customize → /mcp).
  • Arrange, don't fake. llm, fixture and api ... --write exist to set up preconditions (a configured profile, a git repo, a dummy secret). They never count as proof that the UI path works; the map says which steps are UI proof. Never intercept routes or add mock LLM responses to make a live check pass.
  • Essential pathways are single commands so recipes can start from a known state: onboard, llm preset|set, conversation start --prompt ... --wait (add --workspace qa-repo to run it in a fixture git-repo), conversation events <id>, workspace open.
  • State control reaches states a happy path never shows, without mocks: browser reset (fresh browser profile: first run again), service stop automation|agent-server (backend-down UI), restart (same state after a backend restart: persistence and reconnection) and restart --rotate-key (stale session key). restart keeps the browser on its page. While a service is stopped, a reload replaces the page with the backend-unavailable screen, so drive backend-down states on the page that was already loaded. browser network, browser toasts and browser media observe requests by origin, toasts and sound without changing anything. Check --help before calling a state unreachable: most "can't be driven" claims predate a verb.
  • Evidence goes under <run>/evidence/<feature-id>/ and the append-only ledger <run>/evidence/ledger.jsonl. Nothing is overwritten: a repeated screenshot name is saved as <name>-2.png, so cite the path the command prints; evidence report renders the table from the report contract, fail and blocked rows first with their --note (a blocked row names its missing prerequisite there), and --baseline <yesterday's run> lists what changed since that ledger. Keys, logs, browser profile and downloads stay in <run>/private/. Evidence is not automatically public: review every image before publishing it. The CLI masks password fields in snapshot, value and testids; a screenshot of a visible key field is still a leak.
  • Cleanup with control-openhands stop (add --purge-private to delete keys and state once you have checked the evidence). It only signals the process group it launched and verifies the ports closed. Delete run-owned fixtures through the UI when deletion is the path under test.
Show full SKILL.md (420 more words)Show less

LLM budget

Features that start a real agent with automation or debugging skills (for example "Debug with OpenHands") can act on other fixtures by name: stop those conversations when their recipe is done.

Use deepseek-flash for everything that needs a model; switch to deepseek-pro only for checks that need a second profile or a stronger model. Keep prompts small and confined to the run workspace. Without a key, run every credential-free recipe and record model-dependent ones as blocked with the missing prerequisite; never substitute a mock and call it a pass.

Choose the job

  • Verify a change or a PR: control-openhands map affected --base <ref> (or --paths with the PR's file list) maps the changed paths to the families whose Source: lines own them, and separates shared code to widen, src/ paths no family owns (a map gap) and non-user-facing paths. Drive every entry point those features list at desktop and phone viewports, and report with the report contract.
  • Run the daily pass: follow references/daily.md: static checks (map check, map coverage, map testids), the changed families since the Maintenance baseline line of the map index (control-openhands map baseline; map affected starts there by default), a smoke row per family, and a date-derived rotation that gives every family a full live pass once a week. A pass that finished its changed families proposes TARGET as the next baseline with map baseline --set "$TARGET" in its PR; compare ledgers with evidence report --baseline when yesterday's is at hand.
  • Create or extend the map: follow references/mapping.md. It teaches how to discover features, write entries against the CLI and prove each one live. control-openhands map coverage measures what is still unmapped.
  • Maintain the map (periodic or weekly): follow references/maintenance.md: index hygiene, a source wave, one live pass over every feature, PR-intent reconciliation for the week's changes, and at most one PR of proven corrections.

Triage what you find

  • Map drift: the map describes something the app no longer does by design. Fix the entry with current source and live evidence.
  • Harness gap: the app works but the CLI cannot drive it. Fix the CLI, re-drive.
  • Product bug: the app is broken. Keep the evidence, search existing issues, report it separately (right repository: Canvas UI here, Agent Server/SDK in OpenHands/software-agent-sdk, scheduling/dispatch in OpenHands/automation). Never rewrite an expected result to bless broken behavior.
  • Blocked: name the missing prerequisite (account, entitlement, OS, binary) and the route you attempted. Unreachable is never a pass.

See references/adaptation.md for where these ideas come from and what was deliberately left out.

© OpenHands, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 49 other files (scripts, references) in .agents/skills/verify-openhands of OpenHands/OpenHands.

  • SKILL.md
  • references/adaptation.md
  • references/daily.md
  • references/feature-map/F01-first-run-and-sign-in.md
  • references/feature-map/F02-app-shell.md
  • references/feature-map/F03-home.md
  • references/feature-map/F04-conversation-list.md
  • references/feature-map/F05-composer.md
  • references/feature-map/F06-agent-activity.md
  • references/feature-map/F07-conversation-page.md
  • references/feature-map/F08-workspace-files-and-changes.md
  • references/feature-map/F09-settings-shell.md
  • references/feature-map/F10-llm-profiles.md
  • references/feature-map/F11-provider-connections.md
  • references/feature-map/F12-model-router.md
  • references/feature-map/F13-agent-profiles.md
  • references/feature-map/F14-secrets.md
  • references/feature-map/F15-agent-behavior-settings.md
  • references/feature-map/F16-application-settings.md
  • … and 31 more

Open the folder on GitHubat commit 7ea83ba

Compare with similar skills

Verify Openhands next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Verify Openhands compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Verify Openhands this skillOpenHands/OpenHands90k—~2.9kAutomated safety check: PassMIT
Geo Fundamentalswasp-lang/wasp19k8 repos~861Automated safety check: PassMIT
AI SDKvercel-labs/ai-facts16821 repos~1.2kAutomated safety check: PassNone
AI Image Generation and Editingzhayujie/CowAgent47k—~1.3kAutomated safety check: PassMIT
SEO GeoReScienceLab/opc-skills1.8k4 repos~2.1kAutomated safety check: PassApache-2.0
Get API Docs with chubandrewyng/context-hub14k2 repos~775Automated safety check: PassMIT

Similar skills

  • Geo Fundamentals

    wasp-lang/wasp

    Generative Engine Optimization for AI search engines (ChatGPT, Claude, Perplexity).

    19k GitHub starsUsed in 8 repos~861 tokens
    Marketing & SEOAuto-check passed
  • AI SDK

    vercel-labs/ai-facts

    Official

    Answer questions about the AI SDK and help build AI-powered features.

    168 GitHub starsUsed in 21 repos~1.2k tokens
    AI & LLM EngineeringAuto-check passed
  • Generates or edits images from text prompts through a Python script that picks an image backend based on which API keys are configured.

    47k GitHub stars~1.3k tokensUpdated yesterday
    Media & CreativeAuto-check passed
  • SEO Geo

    ReScienceLab/opc-skills

    SEO & GEO (Generative Engine Optimization) for websites. An agent skill from ReScienceLab/opc-skills.

    1.8k GitHub starsUsed in 4 repos~2.1k tokens
    Marketing & SEOAuto-check passed
  • Get API Docs with chub

    andrewyng/context-hub

    Fetches current documentation for third-party APIs and SDKs with the chub CLI before the agent writes code against them, instead of relying on remembered API shapes.

    14k GitHub starsUsed in 2 repos~775 tokens
    DevelopmentAuto-check passed
  • Cognee Docker Setup

    topoteretes/cognee

    Runs the Cognee AI memory platform in Docker, from a one-file prebuilt image to a full compose stack with UI, MCP server, Postgres and Neo4j.

    32k GitHub starsUsed in 1 repo~901 tokens
    DevOps & CloudAuto-check: notes

More from OpenHands/OpenHands

All 8 skills in this repo
  • PR Design Doc

    OpenHands/OpenHands

    For a non-trivial pull request, write a self-contained HTML design doc under the temporary .pr/ directory and link a visibility-appropriate preview in the PR description, so maintainers grasp the…

    90k GitHub stars~2.4k tokensUpdated today
    Auto-check passed
  • Desktop Electron

    OpenHands/OpenHands

    This skill should be used when the user asks to "change the desktop app", "package Electron", "build a universal macOS app", "bundle Node or uv", "fix Electron startup", or changes electron/…

    90k GitHub stars~302 tokensUpdated today
    Auto-check passed
  • E2E Testing

    OpenHands/OpenHands

    This skill should be used when the user asks to "add an E2E test", "run live E2E", "run mock-LLM tests", "debug Playwright CI", "test the Docker image", or changes tests/e2e, Playwright configs, E2E…

    90k GitHub stars~308 tokensUpdated today
    Auto-check passed
  • Frontend API Contracts

    OpenHands/OpenHands

    This skill should be used when the user asks to "add an API call", "change a backend", "update settings persistence", "change conversation events", "fix backend auth", "add Agent Server support", or…

    90k GitHub stars~390 tokensUpdated today
    Auto-check passed
  • Frontend Development

    OpenHands/OpenHands

    This skill should be used when the user asks to "add UI copy", "add a translation", "optimize the frontend bundle", "change onboarding", "change conversation UI", "add a query key", "change MSW…

    90k GitHub stars~324 tokensUpdated today
    Auto-check passed
  • Local Stack Runtime

    OpenHands/OpenHands

    This skill should be used when the user asks to "change the dev stack", "add a runtime service", "change the launcher", "update Docker", "bump Agent Server", "change ingress routing", or changes…

    90k GitHub stars~375 tokensUpdated today
    Auto-check passed

Works with

Questions about Verify Openhands

What does Verify Openhands do?

This skill should be used to "verify OpenHands features", "test the Canvas UI like a user", "drive Agent Canvas", "check a UI change in the real app", "create or update the feature map", "run the…. Verify Openhands is an agent skill from OpenHands/OpenHands. This skill should be used to "verify OpenHands features", "test the Canvas UI like a user", "drive Agent Canvas", "check a UI change in the real app", "create or update the feature map", "run the daily pass" or "run the weekly feature audit".

How do I install Verify Openhands in Claude Code?

Run `npx skills add OpenHands/OpenHands --skill verify-openhands -a claude-code`. Or copy the skill folder (.agents/skills/verify-openhands in OpenHands/OpenHands) into .claude/skills/verify-openhands in your project. Claude Code loads it when a task matches its description.

How do I install Verify Openhands in Codex?

Run `npx skills add OpenHands/OpenHands --skill verify-openhands -a codex`. Or copy the skill folder (.agents/skills/verify-openhands in OpenHands/OpenHands) into .agents/skills/verify-openhands in your project. Codex loads it when a task matches its description.

Can I use Verify Openhands in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add OpenHands/OpenHands --skill verify-openhands -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/verify-openhands, .gemini/skills/verify-openhands, .github/skills/verify-openhands and .opencode/skills/verify-openhands in your project.

What does Verify Openhands need to run?

Going by SKILL.md and its folder, Verify Openhands needs the command-line tools its instructions call (npm and node) and credentials named DEEPSEEK_API_KEY. Our summary lists: A credential in DEEPSEEK_API_KEY.

Does Verify Openhands access the network?

SKILL.md contains no URLs. Its commands use npm, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Verify Openhands safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Verify Openhands use?

Verify Openhands is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Verify Openhands use?

About 2.9k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 249k tokens, read only when the agent opens those files.

What are the alternatives to Verify Openhands?

Skills that share tags, products or a category with Verify Openhands: Geo Fundamentals (wasp-lang/wasp, 19k stars), AI SDK (vercel-labs/ai-facts, 168 stars), AI Image Generation and Editing (zhayujie/CowAgent, 47k stars) and SEO Geo (ReScienceLab/opc-skills, 1.8k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Verify Openhands?

OpenHands (a GitHub organization) maintains it in OpenHands/OpenHands, which has 90,141 GitHub stars. The repository holds 8 skills in this directory. The repository was last updated on October 7, 2026.

Source: OpenHands/OpenHands on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.