Agent skill

Computer Use

by katipally in katipally/openlive

Use before operating an app on the user's computer, such as clicking, filling a field, reading a window or moving through a website in their browser.

Apache-2.0Auto-check passedProductivity & Automation

Install Computer Use

skills CLI
$ npx skills add katipally/openlive --skill computer-use -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install katipally/openlive computer-use --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/katipally/openlive.git skills-src && mkdir -p .claude/skills && cp -r skills-src/services/agent/skills/computer-use .claude/skills/computer-use && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
computer-use
GitHub stars
336
Token cost
~1.4k tokens
SKILL.md length
893 words
Files
1
Skills in repo
4
Repo updated
First seen
Licence
Apache-2.0

At a glance

Use before operating an app on the user's computer, such as clicking, filling a field, reading a window or moving through a website in their browser.

  • Works in 5 steps: Look first → Act on meaning, not pixels → Verify by reading back → …
  • Tasks that involve Desktop control
  • SKILL.md covers The rules that never bend, 1. Look first, 2. Act on meaning, not pixels and 3. Verify by reading back, plus 5 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Computer Use is an agent skill from katipally/openlive. Use before operating an app on the user's computer, such as clicking, filling a field, reading a window or moving through a website in their browser.

Its SKILL.md is about 1.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Productivity & Automation, covering Desktop control and Text to speech and voice. The repository describes itself as: Opensource, on-device voice + vision layer for AI agents. Bring any model or coding agent; the whole speech loop (VAD, STT, TTS, barge-in) runs locally. An open alternative to… The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve Desktop control
  • Tasks that involve Text to speech and voice

Example prompts

  • “/computer-use”

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Look first
  2. Act on meaning, not pixels
  3. Verify by reading back
  4. Stale element numbers
  5. When to use read_screen_text

What it can do on your machine

Read from SKILL.md and the folder at commit c9121eb. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Computer Use loads about 1.4k tokens when it runs. Until then it costs about 41 tokens; SKILL.md has 893 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~41
When it runs · the whole SKILL.md, loaded when a task matches
~1.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from katipally/openlive at commit c9121eb, republished under its Apache-2.0 licence (© katipally). 893 words, ~1,395 tokens.

Download SKILL.mdSave it as .claude/skills/computer-use/SKILL.md (or your agent's skills folder).
name
computer-use
description
Use before operating an app on the user's computer, such as clicking, filling a field, reading a window or moving through a website in their browser.
license
Apache-2.0
metadata.source
Adapted in part from Orca (https://github.com/stablyai/orca), MIT. See THIRD_PARTY_NOTICES.

Working in apps on the user's computer

You see an app through its accessibility tree: every control numbered, with a picture of the window after it. You act on those numbers. Work in short loops: look, act once, read what came back, then decide the next step.

The rules that never bend

  • Do not send, submit, buy, delete, or change account settings unless the user asked for exactly that. When a step would do one of these and they did not ask, stop and ask.
  • Password managers are off limits. Do not open, read or fill from them.
  • Never tell the user something was sent, saved, bought or deleted unless the window's state shows it.

1. Look first

Call get_app_state with the app before any action. Name the app by its name or its id from list_apps. Without one you get the window in front that is not OpenLive's own.

  • More than one window? list_windows, then pass window_id.
  • Only need the controls? Pass screenshot: false. The tree alone is faster.
  • App not running? open_app. A web page? open_url is one call and cannot miss; prefer it over typing into the address bar.
  • Still loading? wait (a second or two), which looks again. Never guess what a page probably says by now.

2. Act on meaning, not pixels

In order of preference:

  1. set_value with an element number, for a text field, search box, slider or checkbox. It replaces the whole value and is the surest way to fill a field.
  2. click with an element number. It presses through accessibility when it can, which works even when the window is behind another.
  3. perform_action with an action the element lists under "Secondary Actions", such as opening a menu or expanding a row.
  4. type after clicking into a field, only when set_value cannot reach it. Pass paste: true for long text; the clipboard is put back after.
  5. keypress for shortcuts, as ["cmdorctrl", "s"]. cmdorctrl is Cmd on a Mac and Ctrl elsewhere.
  6. x and y from the latest picture, only when no element fits (a canvas, a map, a game). Use the coordinates of the picture you were given.

Window positions from list_windows are desktop coordinates for the window tools. They are never a place to click.

For putting the user's own words into their document, use insert_text, not type.

3. Verify by reading back

Every action answers with the window's new state. Read it before the next step. The result also says how sure it is:

  • "Done (...), and read back" means the change was checked.
  • "The element's own action ran" or "Input was posted" means nothing confirmed it. The state below is the only evidence. If it does not show the change, the step did not work: look again and try another way.

After a form or a multi-step task, read the final state and say what it shows, not what you meant to do.

4. Stale element numbers

Element numbers belong to the state they came from. Any action, a page load or a menu opening renumbers them. So:

  • Only ever use numbers from the latest state.
  • If an action fails with an unknown or wrong element, or the state looks different from what you expected, call get_app_state again and find the element by its label, not its old number.
  • If the same step fails twice, change approach (another element, a keyboard shortcut, a menu) rather than repeating it.
Show full SKILL.md (320 more words)Show less

5. When to use read_screen_text

The tree carries most text. Use read_screen_text (OCR) only for text the tree does not have: a canvas, an image, a PDF rendered as pictures, a remote desktop or a video. Its positions are in the picture's space, which is what click takes as x and y.

Browsers

  • To go to an address: open_url first. Already in the browser and need its own tab? set_value on the address field, then keypress ["Return"].
  • Web pages show up in the tree like any app. Find fields by their labels.
  • After a click that loads a page, wait, then read the new state.
  • Chromium browsers and Electron apps on Linux show their controls only when started with --force-renderer-accessibility. If the tree is empty, say so and use the picture.

Per system

macOS. The helper needs Accessibility and Screen Recording for "OpenLive Computer Use". Without Screen Recording there is no picture, but the tree still works.

Windows. An app running as administrator cannot be driven from a normal app (UIPI). The tool says so. Tell the user, and suggest they reopen the app without admin rights or do that step themselves. Windows may refuse keystrokes to a window it keeps in the background; activate the window first.

Linux. Controls come from the accessibility bus (AT-SPI). Apps open before accessibility was switched on need a restart. On Wayland the picture is often the whole screen, not one window, and on sway, Hyprland and other wlroots desktops controls can be pressed but clicks and keys cannot be posted. Prefer element actions there.

When to stop and ask

  • A login, a captcha, a payment or a permission dialog appears.
  • The next step would send, submit, buy, delete or change a setting the user did not ask for.
  • You tried two different ways and the state still does not show the change.

Say plainly what you see and what you need from them.

© katipally, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in services/agent/skills/computer-use of katipally/openlive.

Open the folder on GitHubat commit c9121eb

Compare with similar skills

Computer Use next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Computer Use compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Computer Use this skillkatipally/openlive336—~1.4kAutomated safety check: PassApache-2.0
Control APIPersonalJarvis/PersonalJarvis141—~1.3kAutomated safety check: PassApache-2.0
Vision SkillsAnionex/agent-vision-toolkit1.2k1 repos~4kAutomated safety check: PassMIT
Mac Computer UseTo3akaRin/mac-computer-use1.1k1 repos~495Automated safety check: PassMIT
Crabbox Appsopenclaw/openclaw392k—~1.5kAutomated safety check: PassMIT
Agent Managementautonomous-ai/Physical-AI-Operating-System381—~1.7kAutomated safety check: PassApache-2.0

Similar skills

  • Control API

    PersonalJarvis/PersonalJarvis

    Change Jarvis's own configuration through the local Control API instead of clicking the desktop UI or driving Computer-Use.

    141 GitHub stars~1.3k tokensUpdated today
    Productivity & AutomationAuto-check passed
  • Vision Skills

    Anionex/agent-vision-toolkit

    Local vision CLIs: glance (describe/ask/OCR an image), ground (locate a target, pixel box), detect (element inventory), trace (image to SVG geometry), crop (cut a pixel box to a file), and…

    1.2k GitHub starsUsed in 1 repo~4k tokens
    Productivity & AutomationAuto-check passed
  • Mac Computer Use

    To3akaRin/mac-computer-use

    操作 macOS 桌面应用,探测窗口和自动化接口、截图、读取或修改辅助功能元素、执行鼠标键盘动作,以及通过 CDP 操作内嵌 Chromium 页面。适用于桌面应用自动化与界面验收;普通网页任务优先使用已有浏览器工具。

    1.1k GitHub starsUsed in 1 repo~495 tokens
    Productivity & AutomationAuto-check passed
  • Crabbox Apps

    openclaw/openclaw

    A skill your agent uses when asked to open, run, test, or show an app in Crabbox, including native desktops, web previews, and computer use on a temporary machine.

    392k GitHub stars~1.5k tokensUpdated today
    Productivity & AutomationAuto-check passed
  • Agent Management

    autonomous-ai/Physical-AI-Operating-System

    Legacy Autonomous Buddy control for explicitly requested Buddy coding sessions.

    381 GitHub stars~1.7k tokensUpdated today
    Productivity & AutomationAuto-check passed
  • Computer Use

    bam-bam-2/solo-skills

    Use Orca's computer-use CLI to inspect and operate local desktop app windows through accessibility trees, screenshots, and safe UI actions.

    367 GitHub starsUsed in 1 repo~915 tokens
    Productivity & AutomationAuto-check passed

More from katipally/openlive

  • Connector Setup

    katipally/openlive

    A skill your agent uses when the user wants to add, connect, sign in to or fix an MCP connector (a server that gives you tools for an app such as GitHub, Linear or Notion).

    336 GitHub stars~860 tokensUpdated today
    Auto-check passed
  • Research

    katipally/openlive

    A skill your agent uses when the user asks you to research, compare, fact-check or find current information, and the answer should come with sources.

    336 GitHub stars~827 tokensUpdated today
    Auto-check passed
  • Skill Creator

    katipally/openlive

    A skill your agent uses when the user wants to teach you a workflow, save how they like something done, or make a new skill.

    336 GitHub stars~657 tokensUpdated today
    Auto-check passed

Questions about Computer Use

What does Computer Use do?

Use before operating an app on the user's computer, such as clicking, filling a field, reading a window or moving through a website in their browser. Computer Use is an agent skill from katipally/openlive. Use before operating an app on the user's computer, such as clicking, filling a field, reading a window or moving through a website in their browser.

When should I use Computer Use?

Computer Use fits situations like: tasks that involve Desktop control; tasks that involve Text to speech and voice.

How do I install Computer Use in Claude Code?

Run `npx skills add katipally/openlive --skill computer-use -a claude-code`. Or copy the skill folder (services/agent/skills/computer-use in katipally/openlive) into .claude/skills/computer-use in your project. Claude Code loads it when a task matches its description.

How do I install Computer Use in Codex?

Run `npx skills add katipally/openlive --skill computer-use -a codex`. Or copy the skill folder (services/agent/skills/computer-use in katipally/openlive) into .agents/skills/computer-use in your project. Codex loads it when a task matches its description.

Can I use Computer Use in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add katipally/openlive --skill computer-use -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/computer-use, .gemini/skills/computer-use, .github/skills/computer-use and .opencode/skills/computer-use in your project.

What does Computer Use need to run?

SKILL.md names no scripts, command-line tools or credentials: Computer Use is instructions for the agent only.

Does Computer Use access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Computer Use safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Computer Use use?

Computer Use is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Computer Use use?

About 1.4k tokens (SKILL.md is roughly 5.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Computer Use?

Skills that share tags, products or a category with Computer Use: Control API (PersonalJarvis/PersonalJarvis, 141 stars), Vision Skills (Anionex/agent-vision-toolkit, 1.2k stars), Mac Computer Use (To3akaRin/mac-computer-use, 1.1k stars) and Crabbox Apps (openclaw/openclaw, 392k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Computer Use?

katipally (a GitHub user) maintains it in katipally/openlive, which has 336 GitHub stars. The repository holds 4 skills in this directory. The repository was last updated on October 7, 2026.

Source: katipally/openlive on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.