Agent skill

Computer Use

by QwenLM in QwenLM/qwen-code

Control local desktop applications through Computer Use for tasks that require reading or operating app UI.

Apache-2.0Auto-check passedProductivity & Automation

Install Computer Use

skills CLI
$ npx skills add QwenLM/qwen-code --skill computer-use -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install QwenLM/qwen-code computer-use --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/QwenLM/qwen-code.git skills-src && mkdir -p .claude/skills && cp -r skills-src/packages/core/src/skills/bundled/computer-use .claude/skills/computer-use && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
computer-use
GitHub stars
28k
Token cost
~3.6k tokens
SKILL.md length
1,584 words
Files
3
Skills in repo
41
Repo updated
First seen
Licence
Apache-2.0

At a glance

Control local desktop applications through Computer Use for tasks that require reading or operating app UI.

  • Works in 2 steps: Initialize → Actions using app
  • Tasks that require reading
  • SKILL.md covers node_repl + @qwen-code/cua-sdk…, Forwarding results in Codex…, Bootstrap and API surface, plus 3 more sections
  • Runs TypeScript scripts from its folder; calls npm

What it does

Computer Use is an agent skill from QwenLM/qwen-code. Control local desktop applications through Computer Use for tasks that require reading or operating app UI. Prefer purpose-built connectors, APIs, or CLIs when available.

Its SKILL.md is about 3.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files (for example `SKILL.test.ts` and `agents/openai.yaml`).

It sits in Productivity & Automation, covering Desktop control. It works with Qwen. The repository describes itself as: An open-source AI coding agent that lives in your terminal. The licence is Apache-2.0.

When your agent uses it

  • Tasks that require reading
  • Operating app UI

Example prompts

  • “/computer-use”

Requirements

  • Node.js

Workflow steps

2 steps, taken from the step headings in SKILL.md.

  1. Initialize
  2. Actions using app

What it can do on your machine

Read from SKILL.md and the folder at commit 4970bfa. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (TypeScript), which the agent can run.

    Shell commands in SKILL.md call:

    • npm

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use npm, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Computer Use loads about 3.6k tokens when it runs. Until then it costs about 46 tokens; SKILL.md has 1,584 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~46
When it runs · the whole SKILL.md, loaded when a task matches
~3.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from QwenLM/qwen-code at commit 4970bfa, republished under its Apache-2.0 licence (© QwenLM). 1,584 words, ~3,595 tokens.

Download SKILL.mdSave it as .claude/skills/computer-use/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
computer-use
description
Control local desktop applications through Computer Use for tasks that require reading or operating app UI. Prefer purpose-built connectors, APIs, or CLIs when available.

node_repl + @qwen-code/cua-sdk (Computer Use)

  • Use node_repl (JavaScript) for all Computer Use actions.
  • If both servers are available, use node_repl from the desktop-node-repl MCP server; the regular server runs on the Qwen Code host.
  • Do not use other technologies besides node_repl for computer interactions, unless specifically requested by the user (e.g. AppleScript, osascript, JXA, System Events, synthesized input).
  • Prefer a dedicated plugin or skill when it can complete the task; use Computer Use for app interactions that are not exposed through a more specific interface.
  • node_repl state is persistent across calls.
  • For text output, use nodeRepl.write(...). It takes a string; use JSON.stringify(...) only for textual metadata. For an observation, write its .text and emit its screenshots as images; do not stringify an observation or driver result containing image bytes.
  • Omit yield_time_ms for ordinary UI calls to use the default 10-second wait. A shorter yield does not speed up the action and can add a node_repl_wait round. Use a shorter yield only when you need control back before completion; if a cell is still running, collect its result with node_repl_wait before issuing dependent actions.

Forwarding results in Codex code mode

When calling node_repl through tools.* inside Codex's outer functions.exec, forward each returned content block by its type. nodeRepl.emitImage(...) produces an MCP image block; the outer script must pass that block to image() for the model to receive an image. Prefer the desktop relay tool when it is present; otherwise use the regular node-repl server:

js
// A code-mode `tools` object throws on an unknown key, so probe with `in`:
// reading an unbound tool would abort the script before any fallback ran.
const DESKTOP_NODE_REPL = 'mcp__desktop_node_repl__node_repl';
const nodeReplTool =
  DESKTOP_NODE_REPL in tools
    ? tools[DESKTOP_NODE_REPL]
    : tools.mcp__node_repl__node_repl;
const result = await nodeReplTool({ code });
for (const block of result.content ?? []) {
  if (block.type === 'text') {
    text(block.text);
  } else if (block.type === 'image') {
    image(block);
  }
}

Here code is the JavaScript to run in the persistent Node REPL. Keep the forwarding loop in the outer code-mode script, outside that code string. Apply the same loop to node_repl_wait results: images may arrive only when a running cell completes. Forward text blocks too, including running-cell IDs and errors. Direct MCP tool calls do not need this outer forwarding loop.

Do not use text(result), text(block), or JSON.stringify(result) to forward an MCP result containing images: this turns image base64 into text, consuming context without showing the image. image() accepts one image block, not the whole result, so do not use image(result) either.

Bootstrap

If desktop-node-repl is connected, skip the installation commands below: it already provides node_repl and the SDK on the connected computer. Continue with the computer initialization below. Otherwise, if node_repl is unavailable, run:

bash
qwen mcp add --scope user node-repl npx -y @qwen-code/node-repl-mcp@0.1.7
npm install --no-save --package-lock=false @qwen-code/cua-sdk@0.20.11

Tell the user to restart Qwen Code, then stop. If only the SDK import is missing, run the second command and retry.

Reuse an existing computer connected to the intended desktop. Otherwise import once per fresh node_repl session. The same App workflow below applies to macOS, Linux and Windows; no additional platform resource is needed:

js
globalThis.computer = await (
  await import('@qwen-code/cua-sdk/computer-use')
).ComputerUse.create();
var platform = await computer.getPlatform();
nodeRepl.write(`Connected platform: ${platform}`);

Use the connected platform for shortcuts, not the CLI or Node host operating system. When the task identifies an app, combine initialization with computer.getApp() and its first getState() in the same call. Read that state before editing or input.

API surface

ts
type Point = number | { x: number; y: number };
type ComputerUse = {
  getPlatform: () => Promise<'macos' | 'linux' | 'windows'>;
  getApp: (nameOrIdentifierOrPath: string) => Promise<App>;
  listApps: () => Promise<
    Array<{ id: string; displayName: string; isRunning: boolean }>
  >;
  close: () => Promise<void>;
};
type App = {
  getState: (options?: {
    disableDiff?: boolean;
    includeScreenshot?: boolean;
    maxTextChars?: number;
  }) => Promise<State>;
  click: (
    point: Point,
    options?: { button?: 'left' | 'right' | 'middle'; count?: number },
  ) => Promise<object>;
  doubleClick: (point: Point) => Promise<object>;
  rightClick: (
    point: Point,
    options?: { modifier?: string[] },
  ) => Promise<object>;
  setValue: (element: number, value: string) => Promise<object>;
  performSecondaryAction: (element: number, action: string) => Promise<object>;
  typeText: (text: string) => Promise<object>;
  paste: (
    text: string,
    options?: { format?: 'text' | 'md' | 'html'; signal?: AbortSignal },
  ) => Promise<object>;
  selectText: (
    element: number,
    text: string,
    options?: {
      prefix?: string;
      suffix?: string;
      selection?: 'text' | 'cursor_before' | 'cursor_after';
      signal?: AbortSignal;
    },
  ) => Promise<object>;
  pressKey: (
    key: string,
    options?: { modifiers?: string[] },
  ) => Promise<object>;
  hotkey: (keys: string[]) => Promise<object>;
  scroll: (
    point: Point,
    options: { direction: 'up' | 'down' | 'left' | 'right'; amount?: number },
  ) => Promise<object>;
  drag: (options: {
    fromX: number;
    fromY: number;
    toX: number;
    toY: number;
  }) => Promise<object>;
};
type State = {
  app: string;
  window: string;
  mode: 'full' | 'diff' | 'no_change';
  text: string;
  screenshot?: { images: Array<{ mimeType: string; dataBase64: string }> };
};

Workflow

paste and selectText currently require macOS. On Linux/Windows, use the shared typeText, setValue, pressKey and observed actions. Unsupported text methods fail with unsupported_platform; do not retry them as window failures.

1. Initialize

If initialization already bound the task's app and returned its state, reuse that app and observation. Otherwise, bind the app named by the task, then read its state. getApp() binds identity; getState() can open a discovered stopped app. Combine these steps in one Node REPL call:

js
var app = await computer.getApp('Microsoft Excel');
nodeRepl.write((await app.getState()).text);

The app handle tracks its current window and dialog. Read the returned window title to confirm the intended document. If the app is unknown or its name is ambiguous, discover applications with computer.listApps() and use a matching application id when it distinguishes the app. Two running instances may have the same ID; retrying that ID cannot resolve the ambiguity. Ask the user to keep only the intended instance open rather than guessing a target or retrying it.

AX text uses short numeric IDs, such as [37] TextField "Name". Use IDs from the current observation for element actions. IDs can change when the app's window or session changes. Disabled and static-text rows are observation-only.

For token efficiency, the accessibility tree will be returned as a diff when appropriate. Prefer this default diff output. A full state replaces the previous state; a diff updates it; no-change preserves it. If you need a full replacement, use disableDiff: true only when the previous state is unavailable or no longer useful. Do not disregard the text and then assume that a subsequent diff will reproduce the information you skipped.

Returned text defaults to at most 12,000 characters. Set maxTextChars (minimum 512) to adjust the limit. A truncation notice means some captured rows were omitted; request app.getState({ disableDiff: true, maxTextChars: 24000 }) when you need more full text. An omitted row does not prove an element is absent. Traversal-limited captures cover only the captured nodes: identical captures can return no-change, while changes return full captured state. Use current captured IDs; after a read failure, use only IDs from the latest observation.

Show full SKILL.md (768 more words)Show less
2. Actions using app

After performing one or more UI actions, call app.getState() before deciding what to do next. Batch actions whose target remains the same, then print only the state needed for the next decision:

js
await app.click(37);
await app.hotkey([platform === 'macos' ? 'super' : 'ctrl', 'a']);
await app.typeText('hello');
await app.pressKey('Return');
nodeRepl.write((await app.getState()).text);

Use the actual ID from your observation; 37 is only an example.

An observation is a decision boundary. When the current state already identifies the controls and the next actions are known, combine those actions and saving in the same call. Read state after the batch. End the batch at a new dialog, menu, changed target or uncertain result; use that state before choosing the next action. Do not split a known sequence merely to put each action in its own call.

  • Prefer element IDs to coordinates. setValue(id, value) changes a writable control, and performSecondaryAction(id, action) invokes a secondary action listed for that element. Use an observed action name rather than guessing.
  • When an action opens or closes a dialog, sheet or menu, end the batch and call app.getState() to read the new window and IDs before continuing.
  • No open application window. means the app is still running without a document window. If closing it completed the task, finish instead of retrying actions; otherwise open the intended file or window first.
  • An action error can occur after the UI already changed. Read state before deciding whether to retry. Partial, unconfirmed or cancelled actions must not be blindly repeated.
  • Coordinate actions use pixels in this app's current screenshot, with (0, 0) at its top-left. Every App observation refreshes that frame internally. Request includeScreenshot: true when you need to inspect the image, especially after a window change. Do not infer coordinates from another window or desktop screenshot.
  • pressKey sends one key, optionally with modifiers. hotkey sends a combination such as ['super', 's']. Use the connected platform's appropriate shortcut: macOS generally uses super; Linux/Windows generally use ctrl.
  • App input manages any required activation internally and restores the previous focus, unless the user has moved it elsewhere. If the platform cannot confirm the target or restore focus, the action reports an error. There is no delivery-mode choice and no automatic replay after an uncertain result.
  • On Linux, some compositor sessions cannot confirm an exact App target. An app_window_unavailable result can mean the session does not support this workflow; ask the user to use a supported desktop session instead of retrying input.
  • Literal \n or \r in typeText sends Return. In a composer or form this may submit rather than insert a newline.
  • If AX is incomplete or does not explain the interface, request a screenshot and inspect it. Only currently captured actionable IDs can be used for element actions.

Paste and select text

These App methods are available on macOS only. Read current state first and use an observed element ID for selection.

js
await app.selectText(37, 'draft', { prefix: 'Status: ' });
nodeRepl.write((await app.getState()).text);

Use a real ID and text from your observation; the example assumes that the selected element contains exactly one matching draft immediately after Status: .

After confirming the selection, replace it with plain text:

js
await app.paste('ready');
nodeRepl.write((await app.getState()).text);
  • selectText(element, text) selects one exact, case-sensitive match. Optional prefix and suffix must be immediately adjacent to that match. Missing or ambiguous matches fail; provide enough observed context to identify one.
  • selection defaults to 'text'. Use 'cursor_before' or 'cursor_after' to place the insertion point at a match boundary instead of selecting the text. The element must support a writable text selection; static labels may not.
  • paste(text) defaults to plain text. Use { format: 'md' } for Markdown or { format: 'html' } for HTML; the receiving app determines which supplied format it accepts. Use plain text when markup should remain literal.
  • Paste uses the clipboard temporarily. It restores the previous clipboard only while it still owns that transaction; a newer external clipboard change is preserved.
  • Both methods accept an optional signal. A completed dispatch, error or cancellation does not by itself prove what changed. Observe state before deciding whether to retry; do not blindly repeat unconfirmed text operations.

Reading screenshots

Every App observation captures the current screenshot internally, independent of whether AX returns full state, a diff or no-change. The default return omits the image. includeScreenshot: true exposes it when visual inspection is needed. Prefer the text-only default when it identifies the controls and confirms the requested change. Request an image to resolve missing or ambiguous information, choose coordinates or verify an appearance that AX does not describe.

js
var state = await app.getState({ includeScreenshot: true });
nodeRepl.write(state.text);
for (const image of state.screenshot?.images ?? []) {
  await nodeRepl.emitImage(`data:${image.mimeType};base64,${image.dataBase64}`);
}

Include connection cleanup at the end of the call that emits the final verification. Inspect that result before reporting success; reconnect and continue if it reveals unfinished work. A separate cleanup-only model turn is unnecessary:

js
await computer.close();
globalThis.computer = undefined;

Reset the Node REPL only when no other persistent state is needed.

© QwenLM, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files in packages/core/src/skills/bundled/computer-use of QwenLM/qwen-code.

  • SKILL.md
  • SKILL.test.ts
  • agents/openai.yaml

Open the folder on GitHubat commit 4970bfa

Compare with similar skills

Computer Use next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Computer Use compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Computer Use this skillQwenLM/qwen-code28k—~3.6kAutomated safety check: PassApache-2.0
Computer UseQwenLM/qwen-code-examples143—~632Automated safety check: PassNone
Computer Use HybridQwenLM/qwen-code-examples143—~623Automated safety check: PassNone
Vision SkillsAnionex/agent-vision-toolkit1.2k—~3.9kAutomated safety check: PassMIT
Mac Computer UseTo3akaRin/mac-computer-use1.1k—~495Automated safety check: PassMIT
Verify Reelrselbach/reel118—~2kAutomated safety check: PassUnlicense

Similar skills

  • Computer Use

    QwenLM/qwen-code-examples

    Control the local desktop using the computer MCP tool from computer-use-mcp.

    143 GitHub stars~632 tokensUpdated 4 mo ago
    Productivity & AutomationAuto-check passed
  • Computer Use Hybrid

    QwenLM/qwen-code-examples

    Control native macOS, Windows, and Linux desktop apps through the open-computer-use MCP server.

    143 GitHub stars~623 tokensUpdated 4 mo ago
    Productivity & AutomationAuto-check passed
  • Vision Skills

    Anionex/agent-vision-toolkit

    Local vision CLIs: glance (describe/ask/OCR an image), ground (locate a target, pixel box), detect (element inventory), trace (image to SVG geometry), crop (cut a pixel box to a file), and…

    1.2k GitHub stars~3.9k tokensUpdated 4 days ago
    Productivity & AutomationAuto-check passed
  • Mac Computer Use

    To3akaRin/mac-computer-use

    操作 macOS 桌面应用,探测窗口和自动化接口、截图、读取或修改辅助功能元素、执行鼠标键盘动作,以及通过 CDP 操作内嵌 Chromium 页面。适用于桌面应用自动化与界面验收;普通网页任务优先使用已有浏览器工具。

    1.1k GitHub stars~495 tokensUpdated 13 days ago
    Productivity & AutomationAuto-check passed
  • Verify Reel

    rselbach/reel

    Verify Reel, the macOS menu-bar screen recorder, by launching a disposable app bundle and driving its real UI with Computer Use.

    118 GitHub stars~2k tokensUpdated 1 mo ago
    Productivity & AutomationAuto-check passed
  • Cloud Computer Use

    davidondrej/cloudroom-core

    See and control desktop apps on this Cloud sandbox’s virtual Linux screen with cloudroom computer-use: launch GUI apps you build or install, read their UI, click, type, and take screenshots.

    263 GitHub stars~881 tokensUpdated today
    Productivity & AutomationAuto-check: notes

More from QwenLM/qwen-code

All 41 skills in this repo
  • Reproduces a feature from Codex or Claude Code in Qwen Code by running the reference agent under capture, reading the traces, then implementing matching behavior.

    28k GitHub stars~1.5k tokensUpdated today
    Auto-check passed
  • Qwen Code E2E Testing

    QwenLM/qwen-code

    Guides end-to-end testing of the Qwen Code CLI in headless mode with real model calls, MCP test servers and inspection of raw API traffic.

    28k GitHub stars~2.1k tokensUpdated today
    Auto-check passed
  • Scheduled CI skill that scans a repository for small, certain docs, test and code hygiene issues and fixes them on one branch with a commit per finding.

    28k GitHub stars~1.7k tokensUpdated today
    Auto-check passed
  • Builds a rebranded Qwen Code desktop package from the Tauri shell using only a brand id and a logo, with sensible derived defaults.

    28k GitHub stars~2.1k tokensUpdated today
    Auto-check passed
  • Walks through capturing and comparing V8 heap snapshots to find memory leaks in the Qwen Code Node.js CLI, using tmux and the chrome-devtools CLI.

    28k GitHub stars~1.3k tokensUpdated today
    Auto-check passed
  • tmux Real User Testing

    QwenLM/qwen-code

    Drives Qwen Code in a real tmux session the way a user would and saves a readable step-by-step transcript of each screen for maintainers to review.

    28k GitHub stars~2.3k tokensUpdated today
    Auto-check passed

Works with

Questions about Computer Use

What does Computer Use do?

Control local desktop applications through Computer Use for tasks that require reading or operating app UI. Computer Use is an agent skill from QwenLM/qwen-code. Control local desktop applications through Computer Use for tasks that require reading or operating app UI.

When should I use Computer Use?

Computer Use fits situations like: tasks that require reading; operating app UI.

How do I install Computer Use in Claude Code?

Run `npx skills add QwenLM/qwen-code --skill computer-use -a claude-code`. Or copy the skill folder (packages/core/src/skills/bundled/computer-use in QwenLM/qwen-code) into .claude/skills/computer-use in your project. Claude Code loads it when a task matches its description.

How do I install Computer Use in Codex?

Run `npx skills add QwenLM/qwen-code --skill computer-use -a codex`. Or copy the skill folder (packages/core/src/skills/bundled/computer-use in QwenLM/qwen-code) into .agents/skills/computer-use in your project. Codex loads it when a task matches its description.

Can I use Computer Use in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add QwenLM/qwen-code --skill computer-use -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/computer-use, .gemini/skills/computer-use, .github/skills/computer-use and .opencode/skills/computer-use in your project.

What does Computer Use need to run?

Going by SKILL.md and its folder, Computer Use needs TypeScript for the scripts in its folder and the command-line tools its instructions call (npm). Our summary lists: Node.js.

Does Computer Use access the network?

SKILL.md contains no URLs. Its commands use npm, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Computer Use safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Computer Use use?

Computer Use is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Computer Use use?

About 3.6k tokens (SKILL.md is roughly 14k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Computer Use?

Skills that share tags, products or a category with Computer Use: Computer Use (QwenLM/qwen-code-examples, 143 stars), Computer Use Hybrid (QwenLM/qwen-code-examples, 143 stars), Vision Skills (Anionex/agent-vision-toolkit, 1.2k stars) and Mac Computer Use (To3akaRin/mac-computer-use, 1.1k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Computer Use?

QwenLM (a GitHub organization) maintains it in QwenLM/qwen-code, which has 28,337 GitHub stars. The repository holds 41 skills in this directory. The repository was last updated on October 7, 2026.

Source: QwenLM/qwen-code on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.