Open Computer Use
iFurySt/open-codex-computer-use
Platform-neutral guidance for using Open Computer Use, the open-source Computer Use MCP server and CLI for macOS, Linux, and Windows.
Execute cross-platform desktop tasks through a packaged desktop automation skill that guides the main agent to observe the screen, focus apps and windows, call helper scripts for screenshots and…
$ npx skills add LeoYeAI/openclaw-master-skills --skill desktop-agent-ops -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install LeoYeAI/openclaw-master-skills desktop-agent-ops --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/LeoYeAI/openclaw-master-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/desktop-agent-ops .claude/skills/desktop-agent-ops && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "desktop-agent-ops" agent skill from https://github.com/LeoYeAI/openclaw-master-skills/tree/main/skills/desktop-agent-ops into .claude/skills/desktop-agent-ops/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "desktop-agent-ops", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/LeoYeAI/openclaw-master-skills/tree/main/skills/desktop-agent-opsType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add LeoYeAI/openclaw-master-skills --skill desktop-agent-ops -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install LeoYeAI/openclaw-master-skills desktop-agent-ops --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/LeoYeAI/openclaw-master-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/desktop-agent-ops .agents/skills/desktop-agent-ops && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "desktop-agent-ops" agent skill from https://github.com/LeoYeAI/openclaw-master-skills/tree/main/skills/desktop-agent-ops into .agents/skills/desktop-agent-ops/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "desktop-agent-ops", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add LeoYeAI/openclaw-master-skills --skill desktop-agent-ops -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install LeoYeAI/openclaw-master-skills desktop-agent-ops --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/LeoYeAI/openclaw-master-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/desktop-agent-ops .cursor/skills/desktop-agent-ops && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "desktop-agent-ops" agent skill from https://github.com/LeoYeAI/openclaw-master-skills/tree/main/skills/desktop-agent-ops into .cursor/skills/desktop-agent-ops/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "desktop-agent-ops", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/LeoYeAI/openclaw-master-skills.git --path skills/desktop-agent-ops--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add LeoYeAI/openclaw-master-skills --skill desktop-agent-ops -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install LeoYeAI/openclaw-master-skills desktop-agent-ops --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/LeoYeAI/openclaw-master-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/desktop-agent-ops .gemini/skills/desktop-agent-ops && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "desktop-agent-ops" agent skill from https://github.com/LeoYeAI/openclaw-master-skills/tree/main/skills/desktop-agent-ops into .gemini/skills/desktop-agent-ops/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "desktop-agent-ops", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install LeoYeAI/openclaw-master-skills desktop-agent-opsInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add LeoYeAI/openclaw-master-skills --skill desktop-agent-ops -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/LeoYeAI/openclaw-master-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/desktop-agent-ops .github/skills/desktop-agent-ops && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "desktop-agent-ops" agent skill from https://github.com/LeoYeAI/openclaw-master-skills/tree/main/skills/desktop-agent-ops into .github/skills/desktop-agent-ops/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "desktop-agent-ops", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add LeoYeAI/openclaw-master-skills --skill desktop-agent-ops -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install LeoYeAI/openclaw-master-skills desktop-agent-ops --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/LeoYeAI/openclaw-master-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/desktop-agent-ops .opencode/skills/desktop-agent-ops && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "desktop-agent-ops" agent skill from https://github.com/LeoYeAI/openclaw-master-skills/tree/main/skills/desktop-agent-ops into .opencode/skills/desktop-agent-ops/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "desktop-agent-ops", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
desktop-agent-opsExecute cross-platform desktop tasks through a packaged desktop automation skill that guides the main agent to observe the screen, focus apps and windows, call helper scripts for screenshots and…
Desktop Agent Ops is an agent skill from LeoYeAI/openclaw-master-skills. Execute cross-platform desktop tasks through a packaged desktop automation skill that guides the main agent to observe the screen, focus apps and windows, call helper scripts for screenshots and input actions, verify each step, clean up task context, and only escalate to multi-agent collaboration when tasks become clearly multi-window or multi-app. Use when the user wants desktop GUI control, native app operation, window focus, screenshots, click and type flows, or cross-platform desktop workflows on macOS…
Its SKILL.md is about 4.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 42 other files, including scripts and reference files (for example `_meta.json`, `agents/openai.yaml` and `references/app-wechat-desktop.md`).
It sits in Productivity & Automation, covering Desktop control. It works with Linux and macOS. The repository describes itself as: 🧠 Curated collection of 1209+ best OpenClaw skills — weekly updated by MyClaw.ai. The licence is MIT.
6 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit e5199b5. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 1 file in scripts/, which the agent can run.
Shell commands in SKILL.md call:
python3From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Desktop Agent Ops loads about 4.5k tokens when it runs, and up to ~24k if it reads all its reference files. Until then it costs about 137 tokens; SKILL.md has 1,204 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from LeoYeAI/openclaw-master-skills at commit e5199b5, republished under its MIT licence (© LeoYeAI). 1,204 words, ~4,513 tokens.
.claude/skills/desktop-agent-ops/SKILL.md (or your agent's skills folder). This skill also uses 40 other files; get the full folder from GitHub.Use this skill as a main-agent operating manual for desktop GUI tasks.
python3 <SKILL_DIR>/scripts/first_run_setup.py --checkIf "ready": false, run setup (installs EVERYTHING automatically):
python3 <SKILL_DIR>/scripts/first_run_setup.pyAuto-installs on first run:
cliclick + tesseract (macOS via brew; Linux guide printed)After setup, set $PY for ALL subsequent calls:
PY=<output.env.DESKTOP_AGENT_OPS_PYTHON>Do NOT proceed if setup is not ready.
Every desktop task follows this loop. No exceptions.
1. auto-setup gate ← run once per session
2. init task context ← create isolated temp directory
3. FOCUS the target app ← bring app to front, confirm frontmost
4. GET window bounds ← know exact position and size
5. CAPTURE that window ← screenshot ONLY the target window
6. ANALYZE the capture ← read screenshot or run OCR
7. LOCATE target via OCR ← find text/button within window bounds
8. VERIFY before acting ← move cursor, screenshot with cursor, confirm
9. EXECUTE one action ← click, type, scroll, press key
10. CAPTURE again ← screenshot to see result
11. VERIFY the result ← did the UI change as expected?
12. → if more steps, go to 5
13. CLEANUP ← remove task temp directoryKey principles:
NEVER do OCR or clicking on a full-screen screenshot. Always scope to the target app window.
┌─────────────────────────────────────────────────────────┐
│ Step 1: FOCUS the target app │
│ $PY desktop_ops.py focus-app --name "AppName" │
│ → brings app to front │
├─────────────────────────────────────────────────────────┤
│ Step 2: GET window bounds │
│ $PY desktop_ops.py front-window-bounds --app "AppName"│
│ → {x, y, width, height} in logical coordinates │
├─────────────────────────────────────────────────────────┤
│ Step 3: CAPTURE only that window │
│ $PY desktop_ops.py capture-region --x X --y Y │
│ --width W --height H --output /tmp/window.png │
├─────────────────────────────────────────────────────────┤
│ Step 4: OCR within the window │
│ $PY ocr_text.py --app "AppName" --python $PY │
│ → abs_box coordinates are INSIDE the window │
├─────────────────────────────────────────────────────────┤
│ Step 5: VERIFY before clicking │
│ $PY desktop_ops.py move --x TX --y TY │
│ $PY desktop_ops.py screenshot --with-cursor │
│ → confirm cursor is on the right element │
├─────────────────────────────────────────────────────────┤
│ Step 6: CLICK only if verified │
│ $PY desktop_ops.py click --x TX --y TY │
│ $PY desktop_ops.py screenshot → verify result │
└─────────────────────────────────────────────────────────┘$PY scripts/target_resolver.py --app "AppName" --text "按钮文字" --python $PYThis single command: focuses app → gets bounds → OCR within window → returns best_candidate with {x, y, within_window}.
| Approach | Risk |
|---|---|
| ❌ Full-screen OCR | "搜索" in WeChat AND Chrome → clicks wrong app |
| ✅ Window-scoped | "搜索" ONLY in WeChat window → correct click |
When something fails, follow these rules:
focus-app --name "AppName"front-window-bounds --app "AppName" (window may have moved/resized)content_area instead of bottom_input)--min-conf 30The pipeline works for any desktop application. Here is how to reason about new apps:
target_resolver.py --app "AppName" --text "target text"| Task | How to do it |
|---|---|
| Click a button | OCR find text → verify → click |
| Type in a field | OCR find field label → click field → type --text |
| Search for something | OCR find search box → click → type query → press return |
| Scroll a list | Get window bounds → scroll at window center with --x --y |
| Switch between apps | focus-app --name "OtherApp" → re-get bounds |
| Handle a dialog | Screenshot → OCR for dialog buttons → click appropriate one |
| Navigate menus | Click menu item → wait → screenshot → OCR new menu → click |
| Select from dropdown | Click dropdown → wait → OCR options → click selection |
| Read screen content | OCR the window → extract all text boxes |
| Verify an action | Screenshot before and after → compare or OCR for expected text |
| App type | Special considerations |
|---|---|
| Chat apps (WeChat, Slack, etc.) | Verify conversation title before typing; use insert-newline for multi-line; verify send mechanism |
| Browsers (Chrome, Safari, etc.) | Address bar at top; content area varies; may need to handle tabs |
| System Settings | Deep navigation; panels change; re-get bounds after each navigation |
| File managers (Finder, Explorer) | Sidebar + content area; double-click to open; path bar for navigation |
| Editors (VS Code, TextEdit, etc.) | Tab bar + editor area; use hotkeys for save/undo; type in editor area |
$PY scripts/desktop_ops.py type --text "your message"set the clipboard to + Cmd+V (single osascript call)Set-Clipboard + Ctrl+V (falls back to clip.exe)xclip + Ctrl+V$PY scripts/desktop_ops.py type --text "first line"
$PY scripts/desktop_ops.py insert-newline
$PY scripts/desktop_ops.py type --text "second line"insert-newline for literal line breaks\n in type --text — it may trigger send in some apps发送) via OCR, then click itpress --key return ONLY when the app is verified to use Enter-to-send| Operation | Primary | Fallback |
|---|---|---|
type | Clipboard paste | cliclick (ASCII only) |
press | AppleScript key code | cliclick kp: |
hotkey | cliclick kd:/t:/ku: | pyautogui |
click | cliclick | pyautogui |
Important: cliclick
kp:returnis NOT recognized by WeChat — always use AppleScript for key press. Important: cliclickt:silently drops CJK characters — always use clipboard paste for text input.
Handled automatically. No manual DPI work needed.
| Platform | Common scales | Detection method |
|---|---|---|
| macOS Retina | 2.0x | screenshot pixels ÷ logical screen bounds |
| Windows HiDPI | 1.25x, 1.5x, 2.0x | screenshot pixels ÷ pyautogui.size() |
| Linux X11 | 1.0x, 1.5x, 2.0x | screenshot pixels ÷ pyautogui.size() |
OCR output: box = logical (use for mouse), pixel_box = raw pixels, dpi_scale = factor.
CRITICAL: Use EXACTLY these names. Do NOT guess.
$PY scripts/desktop_ops.py screenshot [--output PATH] [--x X --y Y --width W --height H] [--with-cursor]
$PY scripts/desktop_ops.py capture-region --x X --y Y --width W --height H [--output PATH] [--with-cursor]
$PY scripts/desktop_ops.py frontmost
$PY scripts/desktop_ops.py list-apps
$PY scripts/desktop_ops.py front-window-bounds [--app NAME]
$PY scripts/desktop_ops.py focus-app --name "App Name"
$PY scripts/desktop_ops.py move --x X --y Y [--duration SECONDS]
$PY scripts/desktop_ops.py click [--x X --y Y] [--button left|right|middle]
$PY scripts/desktop_ops.py double-click [--x X --y Y] [--button left|right|middle]
$PY scripts/desktop_ops.py drag --x1 X1 --y1 Y1 --x2 X2 --y2 Y2 [--duration SEC] [--button left]
$PY scripts/desktop_ops.py scroll --amount N [--x X --y Y] [--direction vertical|horizontal]
$PY scripts/desktop_ops.py mouse-position
$PY scripts/desktop_ops.py press --key KEY
$PY scripts/desktop_ops.py type --text "text to type"
$PY scripts/desktop_ops.py insert-newline [--count N]
$PY scripts/desktop_ops.py hotkey --keys cmd c
$PY scripts/desktop_ops.py screen-size
$PY scripts/desktop_ops.py pixel-color --x X --y Y$PY scripts/ocr_text.py --app "AppName" --python $PY [--region-label LABEL] [--lang auto]
$PY scripts/ocr_text.py --image /path/to/capture.png --python $PY [--lang auto]$PY scripts/target_resolver.py --app "AppName" --text "text" --python $PY
$PY scripts/target_resolver.py --app "AppName" --template /path/icon.png --python $PY
$PY scripts/target_resolver.py --app "AppName" --text "text" --region-label LABEL --python $PY$PY scripts/task_context.py init --task-id "my-task" # aliases: create, --name
$PY scripts/task_context.py show --task-id "my-task"
$PY scripts/cleanup_task.py --task-id "my-task"$PY scripts/window_regions.py --window-x X --window-y Y --window-width W --window-height H [--label LABEL]Labels: top_search, left_sidebar, left_sidebar_top, title_header, content_area, toolbar_row, bottom_input, primary_action
1. $PY first_run_setup.py --check → ready: true
2. $PY task_context.py init --task-id "click-button"
3. $PY desktop_ops.py focus-app --name "AppName"
4. $PY desktop_ops.py front-window-bounds --app "AppName" → {x, y, w, h}
5. $PY target_resolver.py --app "AppName" --text "OK" --python $PY
→ best_candidate: {x:450, y:520, within_window:true}
6. $PY desktop_ops.py move --x 450 --y 520
7. $PY desktop_ops.py screenshot --with-cursor → verify cursor on "OK"
8. $PY desktop_ops.py click --x 450 --y 520
9. $PY desktop_ops.py screenshot → verify result
10. $PY cleanup_task.py --task-id "click-button"1. $PY desktop_ops.py focus-app --name "Safari"
2. $PY target_resolver.py --app "Safari" --text "Search" --region-label top_search --python $PY
→ {x:300, y:80, within_window:true}
3. $PY desktop_ops.py click --x 300 --y 80
4. $PY desktop_ops.py type --text "hello world"
5. $PY desktop_ops.py press --key return
6. $PY desktop_ops.py screenshot → verify search results1. $PY desktop_ops.py focus-app --name "WeChat"
2. $PY desktop_ops.py front-window-bounds --app "WeChat"
3. # Navigate to the right conversation (OCR sidebar or search)
4. $PY target_resolver.py --app "WeChat" --text "ContactName" --region-label left_sidebar --python $PY
5. $PY desktop_ops.py click --x <found_x> --y <found_y>
6. # Verify conversation is open
7. $PY desktop_ops.py screenshot → confirm conversation title
8. # Click the input field
9. $PY target_resolver.py --app "WeChat" --text "" --region-label bottom_input --python $PY
OR: click at the bottom center of the window
10. $PY desktop_ops.py type --text "Hello!"
11. # Send: prefer visible send button; if not available, use press --key return
12. $PY target_resolver.py --app "WeChat" --text "发送" --python $PY
IF found: $PY desktop_ops.py click --x <x> --y <y>
ELSE: $PY desktop_ops.py press --key return
13. $PY desktop_ops.py screenshot → verify message sent1. $PY desktop_ops.py focus-app --name "AppName"
2. $PY desktop_ops.py front-window-bounds --app "AppName" → {x:100, y:50, w:800, h:600}
3. # Scroll down in the window center
$PY desktop_ops.py scroll --amount -5 --x 500 --y 350
4. $PY desktop_ops.py screenshot → check if target visible
5. $PY target_resolver.py --app "AppName" --text "target item" --python $PY
6. IF not found: scroll more and retry (max 5 scrolls)
7. IF found: click it1. # During any operation, if the expected UI doesn't match:
2. $PY desktop_ops.py screenshot → examine what's on screen
3. # If a dialog is visible, OCR it:
$PY ocr_text.py --app "AppName" --python $PY
4. # Find and click the appropriate button (OK, Cancel, Allow, etc.)
$PY target_resolver.py --app "AppName" --text "OK" --python $PY
5. $PY desktop_ops.py click --x <x> --y <y>
6. # After dialog is dismissed, re-get window bounds and continue
$PY desktop_ops.py front-window-bounds --app "AppName"Load as needed:
| Document | When to read |
|---|---|
references/workflow.md | Core 8-step closed loop |
references/platform-macos.md | macOS-specific tools and permissions |
references/platform-windows.md | Windows setup |
references/platform-linux.md | Linux X11/Wayland setup |
references/operation-patterns.md | Reusable task templates |
references/validation-patterns.md | Two-stage validation |
references/precise-targeting.md | 5-layer precision targeting |
references/target-providers.md | Provider ordering and fallback contract |
references/coordinate-reconstruction.md | Rebuild click coordinates from screenshot evidence |
references/chat-app-macos.md | Chat app workflow |
references/app-wechat-desktop.md | Cross-platform WeChat guidance |
references/cleanup-rules.md | Cleanup timing and scope |
references/collaboration-rules.md | When multi-agent collaboration is justified |
references/example-cases.md | Repeatable task examples |
references/reproducible-setup.md | Host bring-up checklist |
Use this skill for: chat apps, browsers, file managers, editors, office apps, system settings, any closed desktop software with no usable API.
--python $PY to ocr_text.py and target_resolver.pyinsert-newline for line breaks; never use \n in type --textpress --key return only when verified© LeoYeAI, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 40 other files (scripts, references) in skills/desktop-agent-ops of LeoYeAI/openclaw-master-skills.
Open the folder on GitHubat commit e5199b5
Desktop Agent Ops next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Desktop Agent Ops this skillLeoYeAI/openclaw-master-skills | 2.2k | — | ~4.5k | Automated safety check: Pass | MIT | |
| Open Computer UseiFurySt/open-codex-computer-use | 2.4k | — | ~1.5k | Automated safety check: Pass | MIT | |
| Waku Computer Useegoist/waku | 1.6k | — | ~3.6k | Automated safety check: Pass | GPL-3.0 | |
| Drive Screencoleam00/skills | 674 | — | ~5.4k | Automated safety check: Pass | MIT | |
| Browser MCP Agentantibrow/anti-detect-browser-skills | 17 | 1 repos | ~4.2k | Automated safety check: Warn | MIT | |
| Gooeypi Computer Useam-will/gooey-pi | 941 | — | ~370 | Automated safety check: Pass | MIT |
iFurySt/open-codex-computer-use
Platform-neutral guidance for using Open Computer Use, the open-source Computer Use MCP server and CLI for macOS, Linux, and Windows.
egoist/waku
Control local macOS, Windows, and Linux apps through Waku Computer Use.
coleam00/skills
Take real control of the desktop - list and focus windows, type, paste, click, scroll, and screenshot - on Windows, macOS or Linux, and drive other coding-agent sessions running in terminals.
antibrow/anti-detect-browser-skills
Give an AI agent its own real browser over MCP tool calls - launch, navigate, click, fill, screenshot, extract text, run JS - with a kernel-level real-device fingerprint and a persistent profile, so…
am-will/gooey-pi
Drive native applications on macOS, Windows, or Linux through the separately installed TryCUA driver CLI.
johnson7788/MultiUserClaw
Drive the user's desktop in the background — clicking, typing, scrolling, dragging — without stealing the cursor, keyboard focus, or switching virtual desktops / Spaces.
LeoYeAI/openclaw-master-skills
Manages pipelines on a DevOps quality and efficiency platform through its OpenAPI: list workspaces and templates, create, update, run and cancel pipelines, and read run records.
LeoYeAI/openclaw-master-skills
Patches OpenClaw's Feishu extension so an edited document triggers an isolated agent session that reads the doc and replies inline, turning it into a live chat space.
LeoYeAI/openclaw-master-skills
Multi-context memory management system for OpenClaw agents with group-isolated storage, global shared memory, workspace organization, and group-specific skills isolation.
LeoYeAI/openclaw-master-skills
Runs a brand's AI-search visibility work end to end: diagnosing how AI platforms represent it, repositioning it, producing AI-optimized content and monitoring ongoing mentions.
LeoYeAI/openclaw-master-skills
Installs and authenticates the gws CLI, then automates Gmail, Drive, Sheets, Calendar, Docs, Chat and Tasks with ready-made recipes, persona bundles and security audits.
LeoYeAI/openclaw-master-skills
Runs four advisor roles, a fitness coach, nutritionist, data analyst and TCM practitioner, to build a health profile and track workouts, diet and wellness over time.
Categories
Execute cross-platform desktop tasks through a packaged desktop automation skill that guides the main agent to observe the screen, focus apps and windows, call helper scripts for screenshots and…. Desktop Agent Ops is an agent skill from LeoYeAI/openclaw-master-skills. Execute cross-platform desktop tasks through a packaged desktop automation skill that guides the main agent to observe the screen, focus apps and windows, call helper scripts for screenshots and input actions, verify each step, clean up task context, and only escalate to multi-agent collaboration when tasks become clearly multi-window or multi-app.
Desktop Agent Ops fits situations like: the user wants desktop GUI control; native app operation; click and type flows; cross-platform desktop workflows on macOS.
Run `npx skills add LeoYeAI/openclaw-master-skills --skill desktop-agent-ops -a claude-code`. Or copy the skill folder (skills/desktop-agent-ops in LeoYeAI/openclaw-master-skills) into .claude/skills/desktop-agent-ops in your project. Claude Code loads it when a task matches its description.
Run `npx skills add LeoYeAI/openclaw-master-skills --skill desktop-agent-ops -a codex`. Or copy the skill folder (skills/desktop-agent-ops in LeoYeAI/openclaw-master-skills) into .agents/skills/desktop-agent-ops in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add LeoYeAI/openclaw-master-skills --skill desktop-agent-ops -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/desktop-agent-ops, .gemini/skills/desktop-agent-ops, .github/skills/desktop-agent-ops and .opencode/skills/desktop-agent-ops in your project.
Going by SKILL.md and its folder, Desktop Agent Ops needs the command-line tools its instructions call (python3). Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Desktop Agent Ops is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 4.5k tokens (SKILL.md is roughly 18k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 20k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Desktop Agent Ops: Open Computer Use (iFurySt/open-codex-computer-use, 2.4k stars), Waku Computer Use (egoist/waku, 1.6k stars), Drive Screen (coleam00/skills, 674 stars) and Browser MCP Agent (antibrow/anti-detect-browser-skills, 17 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
LeoYeAI (a GitHub user) maintains it in LeoYeAI/openclaw-master-skills, which has 2,160 GitHub stars. The repository holds 1,235 skills in this directory. The repository was last updated on July 20, 2026.
Source: LeoYeAI/openclaw-master-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.