Agent skill

Computer Use Automation

by borghei in borghei/Claude-Skills

This skill should be used when the user asks to "build a computer-use agent", "automate a GUI with an AI agent", "when to use computer use vs an API", "make browser automation reliable", or "design…

MITAuto-check passedProductivity & Automation

Install Computer Use Automation

skills CLI
$ npx skills add borghei/Claude-Skills --skill computer-use-automation -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install borghei/Claude-Skills computer-use-automation --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/borghei/Claude-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/engineering/computer-use-automation .claude/skills/computer-use-automation && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
computer-use-automation
GitHub stars
881
Token cost
~1.5k tokens
SKILL.md length
685 words
Files
4 (incl. scripts, references)
Skills in repo
349
Repo updated
First seen
Licence
MIT

At a glance

This skill should be used when the user asks to "build a computer-use agent", "automate a GUI with an AI agent", "when to use computer use vs an API", "make browser automation reliable", or "design…

  • Works in 4 steps: Run tool_choice_advisor.py with the… → If computer-use is justified, draft the… → Add a verification observation after… → …
  • Asks to build a computer-use agent
  • SKILL.md covers Overview, Clarify First, Quick Start and Tools Overview, plus 3 more sections
  • Runs Python scripts from its folder; calls python

What it does

Computer Use Automation is an agent skill from borghei/Claude-Skills. This skill should be used when the user asks to "build a computer-use agent", "automate a GUI with an AI agent", "when to use computer use vs an API", "make browser automation reliable", or "design screenshot-driven agent actions".

Its SKILL.md is about 1.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files, including scripts and reference files (for example `references/computer-use-patterns.md`, `scripts/action_safety_linter.py` and `scripts/tool_choice_advisor.py`).

It sits in Productivity & Automation, covering Desktop control. The repository describes itself as: 385 AI skills, 77 expert agents, and 900 stdlib Python tools for every team: engineering, PM, marketing, C-level, compliance, business ops, research, and a LinkedIn toolkit… The licence is MIT.

When your agent uses it

  • Asks to build a computer-use agent
  • Automate a GUI with an AI agent
  • To use computer use vs an API
  • Make browser automation reliable

Example prompts

  • “build a computer-use agent”
  • “automate a GUI with an AI agent”
  • “when to use computer use vs an API”
  • “/computer-use-automation”

Requirements

  • Python 3

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Run tool_choice_advisor.py with the target's API/MCP availability, GUI stability, and volume — if it says "use API/MCP," stop and build…
  2. If computer-use is justified, draft the action plan as the screenshot→reason→action loop: each step re-grounds on a fresh screenshot…
  3. Add a verification observation after every state-changing action (read back the resulting screen, not the intent).
  4. Insert confirmation gates before any destructive/irreversible step and choose a sandbox (throwaway profile, test account, isolated…

What it can do on your machine

Read from SKILL.md and the folder at commit 4a698e8. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 2 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Computer Use Automation loads about 1.5k tokens when it runs, and up to ~4k if it reads all its reference files. Until then it costs about 64 tokens; SKILL.md has 685 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~64
When it runs · the whole SKILL.md, loaded when a task matches
~1.5k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from borghei/Claude-Skills at commit 4a698e8, republished under its MIT licence (© borghei). 685 words, ~1,510 tokens.

Download SKILL.mdSave it as .claude/skills/computer-use-automation/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
computer-use-automation
description
This skill should be used when the user asks to "build a computer-use agent", "automate a GUI with an AI agent", "when to use computer use vs an API", "make browser automation reliable", or "design screenshot-driven agent actions".
license
MIT + Commons Clause
metadata.version
1.0.0
metadata.author
borghei
metadata.category
engineering
metadata.domain
ai-agents
metadata.updated
2026-06-29
metadata.tags
computer-use, browser-automation, agents, gui, reliability

Computer Use Automation

Category: Engineering Domain: AI Agents

Overview

The Computer Use Automation skill helps you design AI agents that operate a graphical interface the way a person does — take a screenshot, reason about what is on screen, then click, type, scroll, or navigate, and repeat. It covers the core perception→reason→action loop, the decision of when computer-use is the right tool versus a structured API/MCP tool (prefer a real API whenever one exists; reach for computer-use only for GUIs with no programmatic surface), reliability patterns (grounding every action in the current screenshot, verifying after each step, recovering from misclicks), safety guardrails (confirmation gates for destructive actions, sandboxing, avoiding blocking dialogs), and how to evaluate a computer-use agent. It is model-agnostic — the patterns apply to any computer-use-capable model and any GUI tool surface.

Clarify First

Before designing or auditing a computer-use agent, confirm these inputs. If any is unknown or vague, ASK — do not assume:

  • Does a real API/MCP tool exist? — whether the target exposes an API, SDK, CLI, or MCP server, or is GUI-only (the single biggest factor; if a real API exists, prefer it and skip computer-use)
  • Task & risk — what the agent must accomplish and whether any step is destructive or irreversible (delete, send, pay, submit), which sets the confirmation gates and sandboxing
  • Which tool — advise on tool choice for a target, or lint a planned action sequence for safety (selects tool_choice_advisor.py vs action_safety_linter.py)

Stop rule: ask only the 2-3 that most change the output. If the user says "just draft it," proceed and list your assumptions at the top of the artifact.

Quick Start

bash
# Decide computer-use vs API/MCP for a target
python scripts/tool_choice_advisor.py --api-exists no --gui-stability high --volume low --json

# Lint a planned action sequence for safety/reliability gaps
python scripts/action_safety_linter.py --file planned_actions.json

# Read actions from stdin and emit a markdown risk report
echo '[{"type":"click","target":"Delete"},{"type":"submit","target":"Confirm"}]' \
  | python scripts/action_safety_linter.py --format markdown

Tools Overview

ToolPurposeKey Flags
tool_choice_advisor.pyRecommend computer-use vs structured API/MCP for a target, with rationale--api-exists, --gui-stability, --volume, --reversible, --json
action_safety_linter.pyScan a planned action list for destructive verbs, missing verification, missing confirmation gates, and dialog-triggering patterns--file, --format, --json

All scripts: Python 3 standard library only, argparse CLI, --json and human-readable output. Run --help for full usage.

Workflows

Decide and Design a Computer-Use Agent
  1. Run tool_choice_advisor.py with the target's API/MCP availability, GUI stability, and volume — if it says "use API/MCP," stop and build against the real interface instead.
  2. If computer-use is justified, draft the action plan as the screenshot→reason→action loop: each step re-grounds on a fresh screenshot before acting.
  3. Add a verification observation after every state-changing action (read back the resulting screen, not the intent).
  4. Insert confirmation gates before any destructive/irreversible step and choose a sandbox (throwaway profile, test account, isolated VM/container).
Show full SKILL.md (270 more words)Show less
Audit a Planned Action Sequence
  1. Express the plan as a JSON/text list of actions (type, target, optional verified/confirmed).
  2. Run action_safety_linter.py --file plan.json to flag risky verbs, unverified state changes, ungated destructive actions, and dialog-triggering patterns.
  3. Resolve each finding — add verification steps, add confirmation gates, replace blocking-dialog flows.
  4. Re-run until clean, then dry-run in the sandbox before any real target.

Reference Documentation

  • Computer Use Patterns - The action loop; computer-use vs structured-tool decision matrix; reliability patterns (grounding, verification, recovery); safety guardrails (confirmation gates, sandboxing, blocking dialogs); evaluation approach; and common failure modes.

Common Patterns

Ground Every Action in the Current Screenshot
  • Never act on a stale screenshot or a remembered layout — re-capture before each action.
  • Reference elements by what is visible now (label, position) rather than a cached coordinate from a prior turn.
  • After acting, take a fresh screenshot and confirm the expected change actually happened before continuing.
Gate Destructive Actions and Sandbox by Default
  • Require an explicit confirmation step before delete, send, pay, submit, or any irreversible action.
  • Run in a sandbox first: throwaway browser profile, test account, or isolated VM/container.
  • Avoid flows that spawn blocking modal/native dialogs (file pickers, OS print dialogs) that the agent cannot see or dismiss; prefer paths that keep state on the page.
Prefer the Real Interface When It Exists
  • A documented API, SDK, CLI, or MCP tool is more reliable, cheaper, and more verifiable than pixels — use it.
  • Reserve computer-use for genuinely GUI-only targets, one-off tasks, or bridging gaps an API does not cover.
  • For high-volume or business-critical flows, the cost of computer-use flakiness usually justifies building or requesting an API.

© borghei, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files (scripts, references) in engineering/computer-use-automation of borghei/Claude-Skills.

  • SKILL.md
  • references/computer-use-patterns.md
  • scripts/action_safety_linter.py
  • scripts/tool_choice_advisor.py

Open the folder on GitHubat commit 4a698e8

Compare with similar skills

Computer Use Automation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Computer Use Automation compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Computer Use Automation this skillborghei/Claude-Skills881—~1.5kAutomated safety check: PassMIT
Vision SkillsAnionex/agent-vision-toolkit1.2k1 repos~4kAutomated safety check: PassMIT
Mac Computer UseTo3akaRin/mac-computer-use1.1k1 repos~495Automated safety check: PassMIT
Crabbox Appsopenclaw/openclaw392k—~1.5kAutomated safety check: PassMIT
Agent Managementautonomous-ai/Physical-AI-Operating-System381—~1.7kAutomated safety check: PassApache-2.0
Computer Usebam-bam-2/solo-skills3671 repos~915Automated safety check: PassMIT

Similar skills

  • Vision Skills

    Anionex/agent-vision-toolkit

    Local vision CLIs: glance (describe/ask/OCR an image), ground (locate a target, pixel box), detect (element inventory), trace (image to SVG geometry), crop (cut a pixel box to a file), and…

    1.2k GitHub starsUsed in 1 repo~4k tokens
    Productivity & AutomationAuto-check passed
  • Mac Computer Use

    To3akaRin/mac-computer-use

    操作 macOS 桌面应用,探测窗口和自动化接口、截图、读取或修改辅助功能元素、执行鼠标键盘动作,以及通过 CDP 操作内嵌 Chromium 页面。适用于桌面应用自动化与界面验收;普通网页任务优先使用已有浏览器工具。

    1.1k GitHub starsUsed in 1 repo~495 tokens
    Productivity & AutomationAuto-check passed
  • Crabbox Apps

    openclaw/openclaw

    A skill your agent uses when asked to open, run, test, or show an app in Crabbox, including native desktops, web previews, and computer use on a temporary machine.

    392k GitHub stars~1.5k tokensUpdated today
    Productivity & AutomationAuto-check passed
  • Agent Management

    autonomous-ai/Physical-AI-Operating-System

    Legacy Autonomous Buddy control for explicitly requested Buddy coding sessions.

    381 GitHub stars~1.7k tokensUpdated today
    Productivity & AutomationAuto-check passed
  • Computer Use

    bam-bam-2/solo-skills

    Use Orca's computer-use CLI to inspect and operate local desktop app windows through accessibility trees, screenshots, and safe UI actions.

    367 GitHub starsUsed in 1 repo~915 tokens
    Productivity & AutomationAuto-check passed
  • Cloud Computer Use

    davidondrej/cloudroom-core

    See and control desktop apps on this Cloud sandbox’s virtual Linux screen with cloudroom computer-use: launch GUI apps you build or install, read their UI, click, type, and take screenshots.

    272 GitHub stars~881 tokensUpdated today
    Productivity & AutomationAuto-check: notes

More from borghei/Claude-Skills

All 349 skills in this repo
  • Agents In The Team

    borghei/Claude-Skills

    Run delivery when AI coding and ops agents take tickets. An agent skill from borghei/Claude-Skills.

    881 GitHub stars~4.2k tokensUpdated today
    Auto-check passed
  • AI Content Disclosure

    borghei/Claude-Skills

    Check AI-generated marketing content and reviews for required disclosures under the EU AI Act, FTC rules and platform AI-label policies.

    881 GitHub stars~3.4k tokensUpdated today
    Auto-check passed
  • AI Prototyping

    borghei/Claude-Skills

    Idea to AI-generated prototype to customer validation to engineering handoff.

    881 GitHub stars~3.6k tokensUpdated today
    Auto-check passed
  • Analytics Engineer

    borghei/Claude-Skills

    Analytics engineering across data modeling, dbt, transformation, and semantic layers.

    881 GitHub stars~3.4k tokensUpdated today
    Auto-check passed
  • Ansoff Matrix

    borghei/Claude-Skills

    Ansoff Matrix — 4-quadrant framework for growth options: market penetration, market/product development, and diversification.

    881 GitHub stars~2.2k tokensUpdated today
    Auto-check passed
  • Brainstorm Okrs

    borghei/Claude-Skills

    OKR brainstorming and validation using the Radical Focus framework — outcome objectives, measurable key results, counter-metrics.

    881 GitHub stars~1.4k tokensUpdated today
    Auto-check passed

Questions about Computer Use Automation

What does Computer Use Automation do?

This skill should be used when the user asks to "build a computer-use agent", "automate a GUI with an AI agent", "when to use computer use vs an API", "make browser automation reliable", or "design…. Computer Use Automation is an agent skill from borghei/Claude-Skills. This skill should be used when the user asks to "build a computer-use agent", "automate a GUI with an AI agent", "when to use computer use vs an API", "make browser automation reliable", or "design screenshot-driven agent actions".

When should I use Computer Use Automation?

Computer Use Automation fits situations like: asks to build a computer-use agent; automate a GUI with an AI agent; to use computer use vs an API; make browser automation reliable.

How do I install Computer Use Automation in Claude Code?

Run `npx skills add borghei/Claude-Skills --skill computer-use-automation -a claude-code`. Or copy the skill folder (engineering/computer-use-automation in borghei/Claude-Skills) into .claude/skills/computer-use-automation in your project. Claude Code loads it when a task matches its description.

How do I install Computer Use Automation in Codex?

Run `npx skills add borghei/Claude-Skills --skill computer-use-automation -a codex`. Or copy the skill folder (engineering/computer-use-automation in borghei/Claude-Skills) into .agents/skills/computer-use-automation in your project. Codex loads it when a task matches its description.

Can I use Computer Use Automation in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add borghei/Claude-Skills --skill computer-use-automation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/computer-use-automation, .gemini/skills/computer-use-automation, .github/skills/computer-use-automation and .opencode/skills/computer-use-automation in your project.

What does Computer Use Automation need to run?

Going by SKILL.md and its folder, Computer Use Automation needs Python for the scripts in its folder and the command-line tools its instructions call (python). Our summary lists: Python 3.

Does Computer Use Automation access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Computer Use Automation safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Computer Use Automation use?

Computer Use Automation is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Computer Use Automation use?

About 1.5k tokens (SKILL.md is roughly 6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.5k tokens, read only when the agent opens those files.

What are the alternatives to Computer Use Automation?

Skills that share tags, products or a category with Computer Use Automation: Vision Skills (Anionex/agent-vision-toolkit, 1.2k stars), Mac Computer Use (To3akaRin/mac-computer-use, 1.1k stars), Crabbox Apps (openclaw/openclaw, 392k stars) and Agent Management (autonomous-ai/Physical-AI-Operating-System, 381 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Computer Use Automation?

borghei (a GitHub user) maintains it in borghei/Claude-Skills, which has 881 GitHub stars. The repository holds 349 skills in this directory. The repository was last updated on October 7, 2026.

Source: borghei/Claude-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.