Agent skill

Computer Use

by kunpengtalk in kunpengtalk/PeakCode

Drive a macOS app through the computer tool — read its accessibility tree, act on elements, and use real input only when the tree cannot reach the target.

MITAuto-check passedProductivity & Automation

Install Computer Use

skills CLI
$ npx skills add kunpengtalk/PeakCode --skill computer-use -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install kunpengtalk/PeakCode computer-use --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/kunpengtalk/PeakCode.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/computer-use/skills/computer-use .claude/skills/computer-use && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
computer-use
GitHub stars
166
Token cost
~2.5k tokens
SKILL.md length
1,533 words
Files
1
Skills in repo
3
Repo updated
First seen
Licence
MIT

At a glance

Drive a macOS app through the computer tool — read its accessibility tree, act on elements, and use real input only when the tree cannot reach the target.

  • Works in 4 steps: Files and shell commands. Copying a file… → The app's own CLI or config file. Most… → The browser tool, if the work is in a… → …
  • Tasks that involve Desktop control
  • SKILL.md covers Reach for it last, The loop, Acting and Permissions, plus 2 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Computer Use is an agent skill from kunpengtalk/PeakCode. Drive a macOS app through the computer tool — read its accessibility tree, act on elements, and use real input only when the tree cannot reach the target.

Its SKILL.md is about 2.5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Productivity & Automation, covering Desktop control and Accessibility. It works with macOS. The repository describes itself as: 一个基于PiAgent的简约AI 编程工具。支持任务看板,AI自驱开发、Goal目标、实时流式传输、内置 Git、本地优先、保护隐私。 The licence is MIT.

When your agent uses it

  • Tasks that involve Desktop control
  • Tasks that involve Accessibility

Example prompts

  • “/computer-use”

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Files and shell commands. Copying a file is cp, not a Finder drag.
  2. The app's own CLI or config file. Most macOS apps have one, and it is scriptable.
  3. The browser tool, if the work is in a web page. It has element refs and a visible
  4. Only then computer.

What it can do on your machine

Read from SKILL.md and the folder at commit d827a8f. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Computer Use loads about 2.5k tokens when it runs. Until then it costs about 42 tokens; SKILL.md has 1,533 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~42
When it runs · the whole SKILL.md, loaded when a task matches
~2.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from kunpengtalk/PeakCode at commit d827a8f, republished under its MIT licence (© kunpengtalk). 1,533 words, ~2,457 tokens.

Download SKILL.mdSave it as .claude/skills/computer-use/SKILL.md (or your agent's skills folder).
name
computer-use
description
Drive a macOS app through the `computer` tool — read its accessibility tree, act on elements, and use real input only when the tree cannot reach the target.

Computer Use

The computer tool drives apps on this Mac through a native helper. It is how you do work that lives in a GUI and nowhere else: a dialog with no CLI, a form in a native app, a preference toggle, confirming what an app actually shows.

Reach for it last

In order:

  1. Files and shell commands. Copying a file is cp, not a Finder drag.
  2. The app's own CLI or config file. Most macOS apps have one, and it is scriptable.
  3. The browser tool, if the work is in a web page. It has element refs and a visible pane the user can watch.
  4. Only then computer.

Skipping to step 4 for something 1–3 could do is slower, more brittle, and it moves the user's cursor around while they are trying to work.

The loop

Observe once, act once, observe again. Never chain GUI actions blind.

  1. action: "list_apps" — what is running. Apps are matched by the name the user would say or by bundle id, so ask for "Finder" even on a system that calls it 访达.
  2. action: "get_state" with that app — returns the accessibility tree of the window in front, with a numbered ref for each element, its position and size, and the actions it actually supports.
  3. action: "act" with a ref from that tree.
  4. Observe again before concluding anything.

Two other reads, for when the tree is not the right lens:

  • action: "list_windows" — every on-screen window with its id, app, title and bounds. Use it to find a window that is not the front one (a sheet, a popover, a background window) and to capture it with screenshot + window_id. Its bounds are already in screen points.
  • action: "displays" — the displays and their scale factors. This is the coordinate space itself: on a second monitor the same pixel is a different point, and this is where that is written down.

get_state also reports a state_id. Pass it back with the ref — a ref is only meaningful for the observation it came from, and a ref from an older one is refused on purpose. The element behind it may be a different control by now. Do not retry a refused ref; take a fresh get_state.

Reading the tree
state_id: s3
访达 (pid 2390) — 180 elements
[0] Window "Downloads" @(120,60) 1280x800
  [1] Button "名称" @(213,140) 1313x28  [press]
  [4] TextField "" @(20,80) 400x24  value=""  [set_value]
  [9] Checkbox "保持文件夹在顶部" @(220,300) 16x16  [press]

[press] and [set_value] are the actions that element supports — use act with element_action: "press" or "set_value" accordingly. [disabled] means it cannot be acted on right now.

The default scope is the focused window. scope: "all" reads every window of the app, which is much larger — use it only when you need to find something you were not told the location of. A tree can also come back marked partial: it stopped early because the app stopped answering accessibility, or because it hit the element ceiling. Do not conclude that a control is missing just because it is not listed.

Acting

An element action is almost always the right move. It is precise, it survives the window moving, and it does not disturb the user: no pointer movement, no focus change.

  • press — buttons, checkboxes, menu items, anything with [press].
  • set_value — text fields and other value-carrying controls. This writes the value directly, so the field does not need to be focused first. Prefer it over type for a field you can address.
  • focus / raise — give an element keyboard focus, or bring its window forward.
Moving through a window

action: "scroll" with a direction (down reveals what is further down) and an optional amount, aimed with a point or an app name. Scroll before deciding something is not there: a short list, a missing button and an empty panel are usually just content that has not been brought into view, and a tree only ever describes what is visible.

A scroll at a point moves the pointer there first, because that is how the system routes wheel events. Aiming at an app's window (no point) is the gentler option; either way the user's pointer ends up somewhere they did not put it.

When the tree cannot reach it

click, drag, scroll, type and key send real input. That means they move the user's actual pointer and type into whatever currently holds their keyboard focus — anything they were doing is interrupted. Use them only when the tree genuinely cannot express the target: a canvas, a custom-drawn control, a slider or a reorder that has no accessibility equivalent, a game.

Before using them, activate the app (open_app) so the input lands where you intend, and say what you are doing. Keep it brief.

  • click — count: 2 double-clicks, button: "right" opens a context menu.
  • drag — from one point to another in one gesture. The path matters: sliders, reorders, selections and canvases all follow it, so do not fake a drag with two clicks. Add modifiers: "cmd" for a modifier-held drag.
  • type — strategy: "keys" (default) types each character; strategy: "paste" puts the text on the clipboard and presses ⌘V. Use paste for long text and for Chinese, Japanese, Korean or emoji: it is one gesture instead of hundreds of key events, and input methods cannot reinterpret it. The clipboard is put back afterwards, so it does not cost the user anything — but say that you are using it if the text is sensitive.
  • key — one key or a chord, for example Enter, Escape, cmd+s, cmd+shift+n.

Coordinates for click, drag and scroll are pixels read off your most recent screenshot — pass the point you actually see, unchanged. That capture knows whether it was a window or the whole display, where it started on screen, and what the display's scale is, so the conversion to screen coordinates is already handled. A window capture's coordinates are therefore relative to that window's image: after the window moves, take a new capture.

Show full SKILL.md (575 more words)Show less
The clipboard

The clipboard is part of the desktop, so it is part of the tool:

  • action: "read_clipboard" — what the user (or the app they were just in) copied. This is often the cheapest way to get text a GUI produced: a rendered table, a link, a path, an error message. macOS may show its own paste notification the first time; that is the system's gate, not something to work around.
  • action: "write_clipboard" — put text there for an app that wants a paste. It replaces what was there: say so, and prefer type when a paste is not what the app expects.

Permissions

The tool runs through a separate helper app, Peak Code Computer Use, and macOS files its grants against that helper rather than against Peak Code. It needs:

WhatWhich permissionWhere the user grants it
Reading and acting on controlsAccessibilitySystem Settings → Privacy & Security → Accessibility
screenshotScreen RecordingSystem Settings → Privacy & Security → Screen Recording

action: "status" reports both. If something is missing, action: "request_access" asks macOS to show its own prompt — but the prompt is only a shortcut, and the user still has to switch it on themselves. Granting it once is enough and it survives later rebuilds of Peak Code, because the grant is pinned to the helper's own identity rather than to the app.

When a call comes back saying a permission is missing: tell the user which one, and stop. Do not retry it, do not look for another way to do the same thing, and never try to grant it yourself — no tccutil, no editing a database, no sudo. That is the user's decision. The same goes for a macOS permission dialog that appears mid-task: report it, do not click through it for them.

Discipline

  • Ask before anything irreversible. Sending a message, confirming a purchase, deleting something, quitting an app with unsaved work, accepting a dialog on the user's behalf. Describe what you are about to do and wait for the user.
  • Look before you act, and look after. Screenshots and trees are the only way you know what actually happened. An act result is a receipt, not a conclusion.
  • Coordinates come from the screenshot you just took. Do not reuse a point from earlier in the conversation; windows move. If you are aiming at something you saw in a tree, its @(x,y) WxH is already screen points — click the middle of it rather than converting anything.
  • Scroll before saying it is not there. Then look again.
  • Give the app a moment. A capture taken immediately after an action often still shows the previous state. Re-read before concluding it failed.
  • Do not fight a permission prompt. If a modal appears, report it.
  • The user's desktop is not your scratch space. Close what you opened, and leave their windows where you found them.

When to say no

Say so plainly, and offer the alternative, when:

  • the work is in a web page — use the browser tool;
  • the work is in files — use the shell;
  • the user asked for something this cannot do reliably (multi-touch gestures, a drag that depends on sub-pixel timing, anything that has to happen while they are using the machine);
  • permissions are not granted and the user has not chosen to grant them.

Computer use is genuinely useful and genuinely limited. Being straight about which one you are looking at is more helpful than a confident attempt that leaves the desktop in a strange state.

© kunpengtalk, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in plugins/computer-use/skills/computer-use of kunpengtalk/PeakCode.

Open the folder on GitHubat commit d827a8f

Compare with similar skills

Computer Use next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Computer Use compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Computer Use this skillkunpengtalk/PeakCode166—~2.5kAutomated safety check: PassMIT
Agent Desktopsuperagent-ai/grok-cli3.5k—~720Automated safety check: PassMIT
Cmux Cuamanaflow-ai/cmux28k—~4.5kAutomated safety check: PassCustom licence
Mac Computer UseTo3akaRin/mac-computer-use1.1k—~495Automated safety check: PassMIT
Verify Reelrselbach/reel118—~2kAutomated safety check: PassUnlicense
Steerdisler/mac-mini-agent272—~2.1kAutomated safety check: PassNone

Similar skills

  • Agent Desktop

    superagent-ai/grok-cli

    Use the built-in Computer sub-agent with agent-desktop for macOS desktop automation.

    3.5k GitHub stars~720 tokensUpdated 3 mo ago
    Productivity & AutomationAuto-check passed
  • Cmux Cua

    manaflow-ai/cmux

    Use only after the user explicitly asks for cmux Computer Use through the cmux-cua skill: drive real macOS apps from a cmux agent session via the bundled engine (accessibility tree + screenshots…

    28k GitHub stars~4.5k tokensUpdated today
    Productivity & AutomationAuto-check passed
  • Mac Computer Use

    To3akaRin/mac-computer-use

    操作 macOS 桌面应用,探测窗口和自动化接口、截图、读取或修改辅助功能元素、执行鼠标键盘动作,以及通过 CDP 操作内嵌 Chromium 页面。适用于桌面应用自动化与界面验收;普通网页任务优先使用已有浏览器工具。

    1.1k GitHub stars~495 tokensUpdated 13 days ago
    Productivity & AutomationAuto-check passed
  • Verify Reel

    rselbach/reel

    Verify Reel, the macOS menu-bar screen recorder, by launching a disposable app bundle and driving its real UI with Computer Use.

    118 GitHub stars~2k tokensUpdated 1 mo ago
    Productivity & AutomationAuto-check passed
  • Steer

    disler/mac-mini-agent

    macOS GUI automation CLI. An agent skill from disler/mac-mini-agent.

    272 GitHub stars~2.1k tokensUpdated 7 mo ago
    Productivity & AutomationAuto-check passed
  • Openbridge Debug

    AFK-surf/OpenBridge

    Debug and validate OpenBridge on a real macOS desktop through mini-machine plus full VNC computer-use.

    430 GitHub stars~2.2k tokensUpdated 3 mo ago
    Productivity & AutomationAuto-check passed

More from kunpengtalk/PeakCode

  • Jev Ultrafast

    kunpengtalk/PeakCode

    The browser decision policy: one numbered element table, one operation-plus-target decision per cycle, and the rules that keep long runs short.

    166 GitHub stars~1.5k tokensUpdated 15 days ago
    Auto-check passed
  • Browser Use

    kunpengtalk/PeakCode

    Drive the in-app browser — open a page, read it as an accessibility tree, click, type, fill forms, screenshot, and verify what the page really shows.

    166 GitHub stars~1.1k tokensUpdated 15 days ago
    Auto-check passed

Works with

Questions about Computer Use

What does Computer Use do?

Drive a macOS app through the computer tool — read its accessibility tree, act on elements, and use real input only when the tree cannot reach the target. Computer Use is an agent skill from kunpengtalk/PeakCode. Drive a macOS app through the computer tool — read its accessibility tree, act on elements, and use real input only when the tree cannot reach the target.

When should I use Computer Use?

Computer Use fits situations like: tasks that involve Desktop control; tasks that involve Accessibility.

How do I install Computer Use in Claude Code?

Run `npx skills add kunpengtalk/PeakCode --skill computer-use -a claude-code`. Or copy the skill folder (plugins/computer-use/skills/computer-use in kunpengtalk/PeakCode) into .claude/skills/computer-use in your project. Claude Code loads it when a task matches its description.

How do I install Computer Use in Codex?

Run `npx skills add kunpengtalk/PeakCode --skill computer-use -a codex`. Or copy the skill folder (plugins/computer-use/skills/computer-use in kunpengtalk/PeakCode) into .agents/skills/computer-use in your project. Codex loads it when a task matches its description.

Can I use Computer Use in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add kunpengtalk/PeakCode --skill computer-use -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/computer-use, .gemini/skills/computer-use, .github/skills/computer-use and .opencode/skills/computer-use in your project.

What does Computer Use need to run?

SKILL.md names no scripts, command-line tools or credentials: Computer Use is instructions for the agent only.

Does Computer Use access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Computer Use safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Computer Use use?

Computer Use is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Computer Use use?

About 2.5k tokens (SKILL.md is roughly 9.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Computer Use?

Skills that share tags, products or a category with Computer Use: Agent Desktop (superagent-ai/grok-cli, 3.5k stars), Cmux Cua (manaflow-ai/cmux, 28k stars), Mac Computer Use (To3akaRin/mac-computer-use, 1.1k stars) and Verify Reel (rselbach/reel, 118 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Computer Use?

kunpengtalk (a GitHub organization) maintains it in kunpengtalk/PeakCode, which has 166 GitHub stars. The repository holds 3 skills in this directory. The repository was last updated on September 22, 2026.

Source: kunpengtalk/PeakCode on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.