Agent skill

Computer Use

by zai-org in zai-org/ZCode

A skill your agent uses when a task needs a native desktop app's own UI or the OS.

Apache-2.0Auto-check passedProductivity & Automation

Install Computer Use

skills CLI
$ npx skills add zai-org/ZCode --skill computer-use -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install zai-org/ZCode computer-use --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/zai-org/ZCode.git skills-src && mkdir -p .claude/skills && cp -r skills-src/apps/zcode-cli/packages/zcode-cua-plugin/skills/computer-use .claude/skills/computer-use && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
computer-use
GitHub stars
7.7k
Token cost
~3.6k tokens
SKILL.md length
1,500 words
Files
1
Skills in repo
5
Repo updated
First seen
Licence
Apache-2.0

At a glance

A skill your agent uses when a task needs a native desktop app's own UI or the OS.

  • Works in 4 steps: Observe, then search the returned tree… → When a matching element exists, act on… → For a settable element prefer setValue… → …
  • A task needs a native desktop apps own UI
  • SKILL.md covers Bootstrap every call, Accessibility first, API and The loop, plus 6 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Computer Use is an agent skill from zai-org/ZCode. Use when a task needs a native desktop app's own UI or the OS. For anything inside a web page, use Browser Use. Main agent only.

Its SKILL.md is about 3.6k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Productivity & Automation, covering Desktop control and Browser automation. The repository describes itself as: Z.ai's coding agent harness. Powerful, intelligent, extensible. The licence is Apache-2.0.

When your agent uses it

  • A task needs a native desktop apps own UI
  • Tasks that involve Desktop control
  • Tasks that involve Browser automation

Example prompts

  • “/computer-use”

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Observe, then search the returned tree for the target by its role, name,
  2. When a matching element exists, act on it by index. For a control that
  3. For a settable element prefer setValue over typing or pasting. Reach for
  4. Keyboard input is the fallback: use pressKey only when no element expresses

What it can do on your machine

Read from SKILL.md and the folder at commit aac4755. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are javascript and typescript).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Computer Use loads about 3.6k tokens when it runs. Until then it costs about 35 tokens; SKILL.md has 1,500 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~35
When it runs · the whole SKILL.md, loaded when a task matches
~3.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from zai-org/ZCode at commit aac4755, republished under its Apache-2.0 licence (© zai-org). 1,500 words, ~3,558 tokens.

Download SKILL.mdSave it as .claude/skills/computer-use/SKILL.md (or your agent's skills folder).
name
computer-use
description
Use when a task needs a native desktop app's own UI or the OS. For anything inside a web page, use Browser Use. Main agent only.

Computer Use

Read or operate the UI of native apps on the user's computer.

  • Prefer a dedicated connector, API, CLI or skill when one can complete the task.
  • For browser and web tasks, use Browser Use instead.
  • Do not use AppleScript, osascript, JXA, System Events, shell commands, or any other UI-automation technology unless the user explicitly asks for that technology.
  • Main agent only. Never delegate Computer Use to a subagent.

Bootstrap every call

Each mcp__node_repl__js call runs in a fresh Worker. Globals, imports, module cache and any binding are gone by the next call, so const app does not survive the end of the cell. The UI state does survive, so re-binding in the next cell is cheap.

Put the bootstrap and the actions in the same cell, bootstrap first:

js
const root =
  process.env.ZCODE_CUA_PLUGIN_ROOT ??
  process.env.ZCODE_PLUGIN_ROOT ??
  process.env.CLAUDE_PLUGIN_ROOT;
const { join } = await import("node:path");
const { pathToFileURL } = await import("node:url");
const { setupComputerUseRuntime } = await import(
  pathToFileURL(join(root, "scripts", "computer-use-client.mjs")).href,
);
await setupComputerUseRuntime({ globals: globalThis });

Accessibility first

Accessibility is the primary action path. It is semantic, precise, works on a background app and does not steal the user's focus.

  1. Observe, then search the returned tree for the target by its role, name, title, value or other visible identity.
  2. When a matching element exists, act on it by index. For a control that advertises a semantic action, performSecondaryAction is also correct.
  3. For a settable element prefer setValue over typing or pasting. Reach for paste only when the target is not settable or the content is rich text.
  4. Keyboard input is the fallback: use pressKey only when no element expresses the operation, or the user asked for keyboard interaction. Coordinates are the last fallback, for canvas, games and Electron content accessibility cannot see.

Do not replace an available element action with a keyboard shortcut just because the shortcut is shorter. Do not run both the accessibility and visual paths for the same action.

Observation works on a background app, screenshots included — on any display, including windows at negative coordinates. has_image: false means the call did not ask for pixels (getAXState), never that capture failed.

Success means the API accepted an action, not that the app acted. Typing into a web-content editor can be accepted and change nothing, so re-observe to confirm the text landed.

API

typescript
type Vec2 = [x: number, y: number];
type ObservationOptions = { emit?: boolean };
type StateOptions = ObservationOptions & { disableDiffing?: boolean };
type StateAndScreenshot = { state: string; screenshot?: Uint8Array };
type Direction = "up" | "down" | "left" | "right" | "u" | "d" | "l" | "r";
type MouseButton = "left" | "right" | "middle" | "l" | "r" | "m";
type SelectionType = "text" | "cursor_before" | "cursor_after";
type Strategy = "auto" | "a11y" | "event";

type ClickOptions = {
  mouseButton?: MouseButton;
  clickCount?: number;
  modifiers?: string;
  strategy?: Strategy;
};
type SelectTextOptions = { prefix?: string; suffix?: string; selectionType?: SelectionType };
type PasteOptions = { format?: "text" | "md" | "html" };
type PressKeyOptions = { holdSeconds?: number; strategy?: Strategy };

interface Target {
  getAXState(options?: StateOptions): Promise<string>;
  getScreenshot(options?: ObservationOptions): Promise<Uint8Array>;
  getAXStateAndScreenshot(options?: StateOptions): Promise<StateAndScreenshot>;
  elements(): Promise<{ index: number; kind: string; title: string | null; value: string | null; actions: string[] }[]>;

  paste(text: string, options?: PasteOptions): Promise<void>;
  click(target: number | Vec2, options?: ClickOptions): Promise<void>;
  drag(from: number | Vec2, to: number | Vec2, options?: { modifiers?: string }): Promise<void>;
  pressKey(key: string, options?: PressKeyOptions): Promise<void>;
  scroll(target: number | Vec2, direction: Direction, pages?: number): Promise<void>;
  selectText(elementIndex: number, text: string, options?: SelectTextOptions): Promise<void>;
  setValue(elementIndex: number, value: string): Promise<void>;
  typeText(text: string): Promise<void>;
  performSecondaryAction(elementIndex: number, action: string): Promise<void>;
}

interface App extends Target {}

type AppRef = { name?: string; bundle_id?: string; pid?: number; window_id?: number };
type AppInfo = { pid: number; name: string | null; bundle_id: string | null; active: boolean };
type State = { apps: AppInfo[] };

declare const agent: {
  computerUse: {
    getState(options?: ObservationOptions): Promise<State>;
    getApp(target: string | AppRef): Promise<App>;
    listApps(options?: ObservationOptions): Promise<AppInfo[]>;
    computer: Record<string, (args: object) => Promise<unknown>>;

    requestAccess(capabilities?: string[]): Promise<Record<string, unknown>>;
    stop(reason?: string): Promise<void>;
  };
};

The full reference is agent.documentation.get("computer-use"), fetched on demand; read it only for an argument shape or response field this page does not give.

The loop

Observe once, act, then observe again before deciding the next step.

agent.computerUse.getApp(...) binds an app and shows nothing; call getAXState when you need to see its state. Binding also launches it if not running; there is no separate launch tool. Its argument is a display name or a bundle identifier — the same strings listApps() returns — or an AppRef.

When the user names an application, copy it character-for-character into the app identifier. Do not translate, localize, normalize, shorten, or remove a suffix: {"name":"网易云音乐app"} is not {"name":"网易云音乐"}; {"name":"日历"} is not {"name":"Calendar"}. A rewritten name resolves to a different app or to nothing, and the failure reads "app not found". If the exact string does not resolve, call agent.computerUse.listApps() once and pick the matching identifier.

Batch related actions and a single closing observation into one cell:

js
const app = await agent.computerUse.getApp("Notes");
await app.click(box);
await app.typeText("hello");
await app.pressKey("Return");
await app.getAXState();

An element index addresses that app's latest observation, so several actions may reuse one index without re-observing between them — click an index, then type into it, in the same cell. Observing renumbers the tree, so take indices from the newest one. A vanished element fails closed with ELEMENT_UNAVAILABLE.

The tree comes back as a diff only against a tree this cell already showed you, listing the elements that were removed, added or changed; unchanged rows are omitted and their indices stay valid. The first tree after binding, and the first after getScreenshot or elements(), are always complete. Pass { disableDiffing: true } for a full tree at any point. If a standalone observation reports no change, do not immediately repeat it without an intervening action.

A large tree is trimmed by priority (ancestors kept), the header says so, and indices then skip numbers. app.elements() returns every element with its index, trimmed ones included — filter it in JS, never guess an index.

A capture is scoped to one window: without a window_id the main/key window is re-resolved every observation, so a modal that just opened becomes the captured window. When an action fails or the tree reads like another part of the app, check the observation's window — a dialog shows up there. Read it, then act on it, or dismiss it (Escape or its own cancel) and observe again. A missing element is not proof the action worked. Pin one with getApp({pid, window_id}); list_windows has the id.

Output

Observations display themselves: getAXState, getScreenshot, getAXStateAndScreenshot, getState and listApps emit their own result. Never pass their return value to nodeRepl.write(...) or nodeRepl.emitImage(...): a second raster in one result breaks the one-raster rule and the frame is removed entirely, so you end up with no picture at all. Pass { emit: false } to suppress the display and still receive the value.

Action methods display nothing.

Show full SKILL.md (682 more words)Show less

Coordinates

Choose x and y only by looking at the current returned raster, with 0 <= x < width and 0 <= y < height for that raster. Submit those integers unchanged; CUA binds the current raster internally and owns every transform from the returned raster to native dispatch. Element and window bounds are diagnostic global screen points and must never be copied into a coordinate. When visual fallback begins, discard coordinate-like numbers from earlier text or accessibility results.

Act on the current raster when the target is clear; if it is too small or ambiguous, re-observe rather than guessing at geometry.

A pointer action accepted with no change usually means the app acted where the real pointer sits: repeating it will not help — use an element index or the keyboard.

A coordinate refusal may name an unexpected owner, or say the frame is stale — the window moved, resized or was replaced. Observe again for a current raster, or act on an element index, which does not depend on window geometry. Never move, resize or close a window to make a coordinate land.

Keyboard

pressKey takes a key or a +-separated chord and accepts both short names and X keysym style: "a", "Return", "Tab", "Control_L+a", "super+c", "Up". macOS uses cmd; Linux and Windows use ctrl. Use { holdSeconds } to hold a key or chord for a duration rather than simulating repeated presses.

Bind keyboard input to an app or element; never send it unbound. paste is for rich text, or a target setValue cannot set — not for plain text a settable element accepts.

performSecondaryAction accepts only an action the element advertises in the current tree. Do not guess an action name.

At most one element is focused — the one holding keyboard focus; absent means undetermined.

selectText locates text inside an editable element. Use prefix/suffix to disambiguate repeated matches and selectionType to place the cursor instead of selecting. An ambiguous match is refused rather than resolved to the first hit.

Waiting

Observations wait for the UI to settle before capturing. Do not pause or delay before reading state — no setTimeout, no polling loop.

Errors and stopping

Actions resolve to undefined on success and throw ComputerUseError otherwise:

  • code — PERMISSION_DENIED, NOT_AUTHORIZED, APP_NOT_FOUND, AMBIGUOUS_APP, LAUNCH_FAILED, INVALID_APP, ELEMENT_UNAVAILABLE, STALE_STATE, NOT_SETTABLE, NOT_SELECTABLE, ACTION_UNAVAILABLE, FOREGROUND_REQUIRED, CONTROLLER_BUSY, CONTROL_STOPPED, SCREEN_LOCKED, HELPER_UNAVAILABLE, VERSION_MISMATCH, TIMEOUT, STRUCTURED_STATE_UNAVAILABLE, INTERNAL.
  • actionSent — whether the action may already have reached the app. Retry a non-idempotent action only when this is false; otherwise observe first and decide from what you see.
  • retry — "reobserve", "retry" or "never".

CONTROLLER_BUSY means another live ZCode Computer Use session owns input. It is never retryable: report the owner from the error and ask the user to close that session.

Stop immediately after agent.computerUse.stop(), a kill switch, a permission refusal, or a non-retryable error. Do not switch to a different UI-automation technology after an access refusal.

Persist until the request is actually complete. Attempting an action is not completion: verify the returned state visibly shows the result. If it is unchanged or only intermediate, try another approach. Respond only when the requested state is visibly present, or explain a concrete blocker you cannot resolve.

Tool surface

agent.computerUse.computer.<tool>(args) is the low-level surface. Prefer the bound-object API above; reach for a tool only for what the API does not express — window enumeration, key repeat, or reading state back in the same call. Arguments are strict: an undeclared key is refused. Each tool takes one arguments object — the names below are its keys, not positional parameters: get_app_state({ app_ref: { bundle_id: "com.apple.Notes" }, include_screenshot: true }).

Name an app as the OS lists it (Windows: the Start-menu name, not a window title). Nothing takes the user's focus, except on Windows: launching a non-packaged app does, and include_screenshot=true un-minimizes — a minimized tree is fully usable, so keep the default.

list_apps({})
list_windows({app_ref})
get_app_state({app_ref, include_screenshot?=false, disable_diffing?=false})

left_click({target, mouse_button?="left", click_count?=1, modifiers?="",
           strategy?, app_ref?, return_state?})
left_click_drag({from_target, to, modifiers?="", app_ref?, return_state?})
scroll({target, scroll_direction, scroll_amount, strategy?, app_ref?,
       return_state?})

type({text, target?, app_ref?, strategy?, return_state?})
set_value({target, value, strategy?, app_ref?, return_state?})
select_text({target, text_range?, app_ref?, return_state?})
key({text, repeat?, hold_seconds?, app_ref?, strategy?, return_state?})
paste({text, format?="text", app_ref?, return_state?})
perform_action({target, action, app_ref?, return_state?})

request_access({capabilities?})
stop_computer_control({reason?})

Argument shapes: app_ref / app is an AppRef, but a bare string here is read as a bundle id, so pass {name: "Notes"} for a display name. scroll_direction is up|down|left|right, scroll_amount is pages, strategy is auto|a11y|event, and return_state is compact|full|none to return the app state in the same call. For target, text_range, modifiers and the response shapes, see nodeRepl.write(await agent.documentation.get("computer-use")).

© zai-org, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in apps/zcode-cli/packages/zcode-cua-plugin/skills/computer-use of zai-org/ZCode.

Open the folder on GitHubat commit aac4755

Compare with similar skills

Computer Use next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Computer Use compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Computer Use this skillzai-org/ZCode7.7k—~3.6kAutomated safety check: PassApache-2.0
Computer Use Action Pickermrmps/classifier-dev424—~1.5kAutomated safety check: PassMIT
Computer Useautonomous-ai/Physical-AI-Operating-System407—~2kAutomated safety check: PassApache-2.0
Codewhale Computer Use Controllercodewhale-hq/Codewhale41k—~1.6kAutomated safety check: PassMIT
Electron App Automationvercel-labs/agent-browser44k5 repos~1.7kAutomated safety check: PassApache-2.0
Isolated Linux Agent Workspaceagent-sh/agent-workspace-linux186—~2.1kAutomated safety check: PassMIT

Similar skills

  • Computer Use Action Picker

    mrmps/classifier-dev

    Pick a browser or desktop agent's next action by choosing among the actions actually on screen instead of inventing one.

    424 GitHub stars~1.5k tokensUpdated 4 days ago
    Productivity & AutomationAuto-check passed
  • Computer Use

    autonomous-ai/Physical-AI-Operating-System

    Operate apps/websites on the paired Mac via Buddy: Calendar, Notes, forms, screenshots, files.

    407 GitHub stars~2k tokensUpdated yesterday
    Productivity & AutomationAuto-check passed
  • Codewhale Computer Use Controller

    codewhale-hq/Codewhale

    Operates desktop apps and isolated or disposable browsers on registered computers through an observe, act, verify loop, separate from driving the user's own signed-in Chrome.

    41k GitHub stars~1.6k tokensUpdated today
    Productivity & AutomationAuto-check passed
  • Electron App Automation

    vercel-labs/agent-browser

    Official

    Automates Electron desktop apps such as VS Code, Slack or Discord by connecting agent-browser to their Chrome DevTools Protocol port.

    44k GitHub starsUsed in 5 repos~1.7k tokens
    Productivity & AutomationAuto-check passed
  • Isolated Linux Agent Workspace

    agent-sh/agent-workspace-linux

    Drives a hidden, agent-owned Linux desktop and browser over MCP for GUI testing and web automation without touching the user's real desktop.

    186 GitHub stars~2.1k tokensUpdated today
    Productivity & AutomationAuto-check passed
  • Altic Studio

    altic-dev/altic-mcp

    macOS automation skill for AppleScript actions and Chrome browser control via MCP CDP tools.

    174 GitHub stars~3.3k tokensUpdated 3 mo ago
    Productivity & AutomationAuto-check passed

More from zai-org/ZCode

All 25 skills in this repo
  • DOCX

    zai-org/ZCode

    Complete DOCX document creation, editing, and analysis capabilities with support for revisions, comments, formatting preservation, and text extraction.

    7.7k GitHub stars~4.9k tokensUpdated yesterday
    Auto-check: notes
  • Visualize

    zai-org/ZCode

    Create visualizations and interactive tools directly in conversation.

    7.7k GitHub stars~8.7k tokensUpdated yesterday
    Auto-check passed
  • PDF

    zai-org/ZCode

    Professional PDF toolkit covering four production workflows: reports, creative visuals, academic LaTeX, and existing PDF processing.

    7.7k GitHub stars~18k tokensUpdated yesterday
    Auto-check: notes
  • A skill your agent uses when ZCode needs to inspect, plan, or execute restoration of old ACP-era ZCode sessions from ~/.zcode/v2/sessions into the new ZCode task/session stores.

    7.7k GitHub stars~1.2k tokensUpdated yesterday
    Auto-check passed
  • Generate, verify, or remove large synthetic ZCode task fixtures in local ~/.zcode persistence for UI/session performance testing.

    7.7k GitHub stars~698 tokensUpdated yesterday
    Auto-check passed
  • Check ZCode module and layer boundaries for code changes. An agent skill from zai-org/ZCode.

    7.7k GitHub stars~1.3k tokensUpdated yesterday
    Auto-check passed

Questions about Computer Use

What does Computer Use do?

A skill your agent uses when a task needs a native desktop app's own UI or the OS. Computer Use is an agent skill from zai-org/ZCode. Use when a task needs a native desktop app's own UI or the OS.

When should I use Computer Use?

Computer Use fits situations like: A task needs a native desktop apps own UI; tasks that involve Desktop control; tasks that involve Browser automation.

How do I install Computer Use in Claude Code?

Run `npx skills add zai-org/ZCode --skill computer-use -a claude-code`. Or copy the skill folder (apps/zcode-cli/packages/zcode-cua-plugin/skills/computer-use in zai-org/ZCode) into .claude/skills/computer-use in your project. Claude Code loads it when a task matches its description.

How do I install Computer Use in Codex?

Run `npx skills add zai-org/ZCode --skill computer-use -a codex`. Or copy the skill folder (apps/zcode-cli/packages/zcode-cua-plugin/skills/computer-use in zai-org/ZCode) into .agents/skills/computer-use in your project. Codex loads it when a task matches its description.

Can I use Computer Use in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add zai-org/ZCode --skill computer-use -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/computer-use, .gemini/skills/computer-use, .github/skills/computer-use and .opencode/skills/computer-use in your project.

What does Computer Use need to run?

SKILL.md names no scripts, command-line tools or credentials: Computer Use is instructions for the agent only.

Does Computer Use access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Computer Use safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Computer Use use?

Computer Use is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Computer Use use?

About 3.6k tokens (SKILL.md is roughly 14k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Computer Use?

Skills that share tags, products or a category with Computer Use: Computer Use Action Picker (mrmps/classifier-dev, 424 stars), Computer Use (autonomous-ai/Physical-AI-Operating-System, 407 stars), Codewhale Computer Use Controller (codewhale-hq/Codewhale, 41k stars) and Electron App Automation (vercel-labs/agent-browser, 44k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Computer Use?

zai-org (a GitHub organization) maintains it in zai-org/ZCode, which has 7,659 GitHub stars. The repository holds 5 skills in this directory. The repository was last updated on October 10, 2026.

Source: zai-org/ZCode on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.