Agent skill

Desktop Agent Ops

by LeoYeAI in LeoYeAI/openclaw-master-skills

Execute cross-platform desktop tasks through a packaged desktop automation skill that guides the main agent to observe the screen, focus apps and windows, call helper scripts for screenshots and…

MITAuto-check passedProductivity & Automation

Install Desktop Agent Ops

skills CLI
$ npx skills add LeoYeAI/openclaw-master-skills --skill desktop-agent-ops -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install LeoYeAI/openclaw-master-skills desktop-agent-ops --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/LeoYeAI/openclaw-master-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/desktop-agent-ops .claude/skills/desktop-agent-ops && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
desktop-agent-ops
GitHub stars
2.2k
Token cost
~4.5k tokens
SKILL.md length
1,204 words
Files
41 (incl. scripts, references)
Skills in repo
1,235
Repo updated
First seen
Licence
MIT

At a glance

Execute cross-platform desktop tasks through a packaged desktop automation skill that guides the main agent to observe the screen, focus apps and windows, call helper scripts for screenshots and…

  • Works in 6 steps: Platform detection (macOS / Windows /… → cliclick + tesseract (macOS via brew;… → OCR language packs auto-detected from… → …
  • The user wants desktop GUI control
  • SKILL.md covers MANDATORY: Auto-setup gate…, Core Execution Loop, Window-Scoped Targeting (THE… and Failure Recovery (CRITICAL), plus 5 more sections
  • Calls python3

What it does

Desktop Agent Ops is an agent skill from LeoYeAI/openclaw-master-skills. Execute cross-platform desktop tasks through a packaged desktop automation skill that guides the main agent to observe the screen, focus apps and windows, call helper scripts for screenshots and input actions, verify each step, clean up task context, and only escalate to multi-agent collaboration when tasks become clearly multi-window or multi-app. Use when the user wants desktop GUI control, native app operation, window focus, screenshots, click and type flows, or cross-platform desktop workflows on macOS…

Its SKILL.md is about 4.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 42 other files, including scripts and reference files (for example `_meta.json`, `agents/openai.yaml` and `references/app-wechat-desktop.md`).

It sits in Productivity & Automation, covering Desktop control. It works with Linux and macOS. The repository describes itself as: 🧠 Curated collection of 1209+ best OpenClaw skills — weekly updated by MyClaw.ai. The licence is MIT.

When your agent uses it

  • The user wants desktop GUI control
  • Native app operation
  • Click and type flows
  • Cross-platform desktop workflows on macOS

Example prompts

  • “/desktop-agent-ops”

Requirements

  • Python 3

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Platform detection (macOS / Windows / Linux)
  2. cliclick + tesseract (macOS via brew; Linux guide printed)
  3. OCR language packs auto-detected from system locale (中文→chi_sim, 日本語→jpn, etc.)
  4. Python venv + pillow, pyautogui, pytesseract, opencv-python, numpy (via uv or pip)
  5. OS permissions (Screen Recording, Accessibility, Automation) with auto-open System Settings
  6. Smoke test (screenshot + mouse move verification)

What it can do on your machine

Read from SKILL.md and the folder at commit e5199b5. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/, which the agent can run.

    Shell commands in SKILL.md call:

    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Desktop Agent Ops loads about 4.5k tokens when it runs, and up to ~24k if it reads all its reference files. Until then it costs about 137 tokens; SKILL.md has 1,204 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~137
When it runs · the whole SKILL.md, loaded when a task matches
~4.5k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~24k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from LeoYeAI/openclaw-master-skills at commit e5199b5, republished under its MIT licence (© LeoYeAI). 1,204 words, ~4,513 tokens.

Download SKILL.mdSave it as .claude/skills/desktop-agent-ops/SKILL.md (or your agent's skills folder). This skill also uses 40 other files; get the full folder from GitHub.
name
desktop-agent-ops
description
Execute cross-platform desktop tasks through a packaged desktop automation skill that guides the main agent to observe the screen, focus apps and windows, call helper scripts for screenshots and input actions, verify each step, clean up task context, and only escalate to multi-agent collaboration when tasks become clearly multi-window or multi-app. Use when the user wants desktop GUI control, native app operation, window focus, screenshots, click and type flows, or cross-platform desktop workflows on macOS, Windows, or Linux.
version
1.0.2

Desktop Agent Ops

Use this skill as a main-agent operating manual for desktop GUI tasks.


MANDATORY: Auto-setup gate (FIRST ACTION, every time)

bash
python3 <SKILL_DIR>/scripts/first_run_setup.py --check

If "ready": false, run setup (installs EVERYTHING automatically):

bash
python3 <SKILL_DIR>/scripts/first_run_setup.py

Auto-installs on first run:

  1. Platform detection (macOS / Windows / Linux)
  2. cliclick + tesseract (macOS via brew; Linux guide printed)
  3. OCR language packs auto-detected from system locale (中文→chi_sim, 日本語→jpn, etc.)
  4. Python venv + pillow, pyautogui, pytesseract, opencv-python, numpy (via uv or pip)
  5. OS permissions (Screen Recording, Accessibility, Automation) with auto-open System Settings
  6. Smoke test (screenshot + mouse move verification)

After setup, set $PY for ALL subsequent calls:

PY=<output.env.DESKTOP_AGENT_OPS_PYTHON>

Do NOT proceed if setup is not ready.


Core Execution Loop

Every desktop task follows this loop. No exceptions.

 1. auto-setup gate           ← run once per session
 2. init task context          ← create isolated temp directory
 3. FOCUS the target app       ← bring app to front, confirm frontmost
 4. GET window bounds          ← know exact position and size
 5. CAPTURE that window        ← screenshot ONLY the target window
 6. ANALYZE the capture        ← read screenshot or run OCR
 7. LOCATE target via OCR      ← find text/button within window bounds
 8. VERIFY before acting       ← move cursor, screenshot with cursor, confirm
 9. EXECUTE one action         ← click, type, scroll, press key
10. CAPTURE again              ← screenshot to see result
11. VERIFY the result          ← did the UI change as expected?
12. → if more steps, go to 5
13. CLEANUP                    ← remove task temp directory

Key principles:

  • One action at a time. Never chain blind actions.
  • Always verify after each action. If verification fails, recapture and retry — do NOT guess.
  • Always work within a specific window. Never click based on full-screen assumptions.

Window-Scoped Targeting (THE CORRECT WAY)

NEVER do OCR or clicking on a full-screen screenshot. Always scope to the target app window.

The 6-Step Pipeline
┌─────────────────────────────────────────────────────────┐
│ Step 1: FOCUS the target app                            │
│   $PY desktop_ops.py focus-app --name "AppName"         │
│   → brings app to front                                 │
├─────────────────────────────────────────────────────────┤
│ Step 2: GET window bounds                               │
│   $PY desktop_ops.py front-window-bounds --app "AppName"│
│   → {x, y, width, height} in logical coordinates        │
├─────────────────────────────────────────────────────────┤
│ Step 3: CAPTURE only that window                        │
│   $PY desktop_ops.py capture-region --x X --y Y         │
│     --width W --height H --output /tmp/window.png       │
├─────────────────────────────────────────────────────────┤
│ Step 4: OCR within the window                           │
│   $PY ocr_text.py --app "AppName" --python $PY          │
│   → abs_box coordinates are INSIDE the window           │
├─────────────────────────────────────────────────────────┤
│ Step 5: VERIFY before clicking                          │
│   $PY desktop_ops.py move --x TX --y TY                 │
│   $PY desktop_ops.py screenshot --with-cursor            │
│   → confirm cursor is on the right element              │
├─────────────────────────────────────────────────────────┤
│ Step 6: CLICK only if verified                          │
│   $PY desktop_ops.py click --x TX --y TY                │
│   $PY desktop_ops.py screenshot → verify result          │
└─────────────────────────────────────────────────────────┘
bash
$PY scripts/target_resolver.py --app "AppName" --text "按钮文字" --python $PY

This single command: focuses app → gets bounds → OCR within window → returns best_candidate with {x, y, within_window}.

Why window-scoped matters:
ApproachRisk
❌ Full-screen OCR"搜索" in WeChat AND Chrome → clicks wrong app
✅ Window-scoped"搜索" ONLY in WeChat window → correct click

Failure Recovery (CRITICAL)

When something fails, follow these rules:

OCR finds nothing
  1. Re-focus the app: focus-app --name "AppName"
  2. Re-get bounds: front-window-bounds --app "AppName" (window may have moved/resized)
  3. Take a fresh screenshot and read it visually
  4. Try a different region label (e.g. content_area instead of bottom_input)
  5. Try lowering OCR confidence: --min-conf 30
Click doesn't work
  1. Screenshot with cursor to check cursor position
  2. The window may have moved — re-get bounds
  3. Try clicking a few pixels offset from the OCR center
  4. Check if a dialog/popup is blocking the target
App state changed (login screen, dialog, etc.)
  1. ALWAYS re-get window bounds after any major UI change
  2. ALWAYS re-run OCR after navigation or state change
  3. Never reuse old coordinates — they may be stale
General retry rule
  • Maximum 3 retries per action
  • Each retry must recapture fresh state
  • If 3 retries fail, report the failure with screenshots and stop

Generalization: How to Apply This to ANY App

The pipeline works for any desktop application. Here is how to reason about new apps:

Step-by-step for ANY new app:
  1. Identify the app name exactly as it appears in the system (e.g. "Google Chrome", "微信", "System Settings")
  2. Focus and get bounds — this tells you the window's exact position
  3. Screenshot the window — look at what's on screen
  4. Identify the target — what text, button, or area do you need to interact with?
  5. Use OCR to find it — target_resolver.py --app "AppName" --text "target text"
  6. Verify and click
Common patterns across apps:
TaskHow to do it
Click a buttonOCR find text → verify → click
Type in a fieldOCR find field label → click field → type --text
Search for somethingOCR find search box → click → type query → press return
Scroll a listGet window bounds → scroll at window center with --x --y
Switch between appsfocus-app --name "OtherApp" → re-get bounds
Handle a dialogScreenshot → OCR for dialog buttons → click appropriate one
Navigate menusClick menu item → wait → screenshot → OCR new menu → click
Select from dropdownClick dropdown → wait → OCR options → click selection
Read screen contentOCR the window → extract all text boxes
Verify an actionScreenshot before and after → compare or OCR for expected text
App-specific adaptations:
App typeSpecial considerations
Chat apps (WeChat, Slack, etc.)Verify conversation title before typing; use insert-newline for multi-line; verify send mechanism
Browsers (Chrome, Safari, etc.)Address bar at top; content area varies; may need to handle tabs
System SettingsDeep navigation; panels change; re-get bounds after each navigation
File managers (Finder, Explorer)Sidebar + content area; double-click to open; path bar for navigation
Editors (VS Code, TextEdit, etc.)Tab bar + editor area; use hotkeys for save/undo; type in editor area

Text Input and Send Rules

Typing text
bash
$PY scripts/desktop_ops.py type --text "your message"
  • Uses clipboard paste as primary method on all platforms — reliable for all languages including CJK
  • macOS: set the clipboard to + Cmd+V (single osascript call)
  • Windows: PowerShell Set-Clipboard + Ctrl+V (falls back to clip.exe)
  • Linux: xclip + Ctrl+V
  • First click on the input field to focus it before typing
Show full SKILL.md (478 more words)Show less
Multi-line messages
bash
$PY scripts/desktop_ops.py type --text "first line"
$PY scripts/desktop_ops.py insert-newline
$PY scripts/desktop_ops.py type --text "second line"
  • Use insert-newline for literal line breaks
  • Do NOT use \n in type --text — it may trigger send in some apps
Sending a message
  1. Preferred: Look for a visible send button (e.g., 发送) via OCR, then click it
  2. Alternative: Use press --key return ONLY when the app is verified to use Enter-to-send
  3. Never guess which send method to use — verify first
Backend priority (macOS)
OperationPrimaryFallback
typeClipboard pastecliclick (ASCII only)
pressAppleScript key codecliclick kp:
hotkeycliclick kd:/t:/ku:pyautogui
clickcliclickpyautogui

Important: cliclick kp:return is NOT recognized by WeChat — always use AppleScript for key press. Important: cliclick t: silently drops CJK characters — always use clipboard paste for text input.


DPI / HiDPI / Retina (All Platforms)

Handled automatically. No manual DPI work needed.

PlatformCommon scalesDetection method
macOS Retina2.0xscreenshot pixels ÷ logical screen bounds
Windows HiDPI1.25x, 1.5x, 2.0xscreenshot pixels ÷ pyautogui.size()
Linux X111.0x, 1.5x, 2.0xscreenshot pixels ÷ pyautogui.size()

OCR output: box = logical (use for mouse), pixel_box = raw pixels, dpi_scale = factor.


CLI Quick Reference (EXACT parameter names)

CRITICAL: Use EXACTLY these names. Do NOT guess.

desktop_ops.py
bash
$PY scripts/desktop_ops.py screenshot [--output PATH] [--x X --y Y --width W --height H] [--with-cursor]
$PY scripts/desktop_ops.py capture-region --x X --y Y --width W --height H [--output PATH] [--with-cursor]
$PY scripts/desktop_ops.py frontmost
$PY scripts/desktop_ops.py list-apps
$PY scripts/desktop_ops.py front-window-bounds [--app NAME]
$PY scripts/desktop_ops.py focus-app --name "App Name"
$PY scripts/desktop_ops.py move --x X --y Y [--duration SECONDS]
$PY scripts/desktop_ops.py click [--x X --y Y] [--button left|right|middle]
$PY scripts/desktop_ops.py double-click [--x X --y Y] [--button left|right|middle]
$PY scripts/desktop_ops.py drag --x1 X1 --y1 Y1 --x2 X2 --y2 Y2 [--duration SEC] [--button left]
$PY scripts/desktop_ops.py scroll --amount N [--x X --y Y] [--direction vertical|horizontal]
$PY scripts/desktop_ops.py mouse-position
$PY scripts/desktop_ops.py press --key KEY
$PY scripts/desktop_ops.py type --text "text to type"
$PY scripts/desktop_ops.py insert-newline [--count N]
$PY scripts/desktop_ops.py hotkey --keys cmd c
$PY scripts/desktop_ops.py screen-size
$PY scripts/desktop_ops.py pixel-color --x X --y Y
ocr_text.py
bash
$PY scripts/ocr_text.py --app "AppName" --python $PY [--region-label LABEL] [--lang auto]
$PY scripts/ocr_text.py --image /path/to/capture.png --python $PY [--lang auto]
target_resolver.py
bash
$PY scripts/target_resolver.py --app "AppName" --text "text" --python $PY
$PY scripts/target_resolver.py --app "AppName" --template /path/icon.png --python $PY
$PY scripts/target_resolver.py --app "AppName" --text "text" --region-label LABEL --python $PY
task_context.py / cleanup_task.py
bash
$PY scripts/task_context.py init --task-id "my-task"   # aliases: create, --name
$PY scripts/task_context.py show --task-id "my-task"
$PY scripts/cleanup_task.py --task-id "my-task"
window_regions.py
bash
$PY scripts/window_regions.py --window-x X --window-y Y --window-width W --window-height H [--label LABEL]

Labels: top_search, left_sidebar, left_sidebar_top, title_header, content_area, toolbar_row, bottom_input, primary_action


Workflow Examples

Example 1: Click a button by text (any app)
1. $PY first_run_setup.py --check                           → ready: true
2. $PY task_context.py init --task-id "click-button"
3. $PY desktop_ops.py focus-app --name "AppName"
4. $PY desktop_ops.py front-window-bounds --app "AppName"    → {x, y, w, h}
5. $PY target_resolver.py --app "AppName" --text "OK" --python $PY
   → best_candidate: {x:450, y:520, within_window:true}
6. $PY desktop_ops.py move --x 450 --y 520
7. $PY desktop_ops.py screenshot --with-cursor               → verify cursor on "OK"
8. $PY desktop_ops.py click --x 450 --y 520
9. $PY desktop_ops.py screenshot                             → verify result
10. $PY cleanup_task.py --task-id "click-button"
1. $PY desktop_ops.py focus-app --name "Safari"
2. $PY target_resolver.py --app "Safari" --text "Search" --region-label top_search --python $PY
   → {x:300, y:80, within_window:true}
3. $PY desktop_ops.py click --x 300 --y 80
4. $PY desktop_ops.py type --text "hello world"
5. $PY desktop_ops.py press --key return
6. $PY desktop_ops.py screenshot                             → verify search results
Example 3: Send a chat message (WeChat, Slack, etc.)
1. $PY desktop_ops.py focus-app --name "WeChat"
2. $PY desktop_ops.py front-window-bounds --app "WeChat"
3. # Navigate to the right conversation (OCR sidebar or search)
4. $PY target_resolver.py --app "WeChat" --text "ContactName" --region-label left_sidebar --python $PY
5. $PY desktop_ops.py click --x <found_x> --y <found_y>
6. # Verify conversation is open
7. $PY desktop_ops.py screenshot → confirm conversation title
8. # Click the input field
9. $PY target_resolver.py --app "WeChat" --text "" --region-label bottom_input --python $PY
   OR: click at the bottom center of the window
10. $PY desktop_ops.py type --text "Hello!"
11. # Send: prefer visible send button; if not available, use press --key return
12. $PY target_resolver.py --app "WeChat" --text "发送" --python $PY
    IF found: $PY desktop_ops.py click --x <x> --y <y>
    ELSE: $PY desktop_ops.py press --key return
13. $PY desktop_ops.py screenshot → verify message sent
Example 4: Scroll a list and find an item
1. $PY desktop_ops.py focus-app --name "AppName"
2. $PY desktop_ops.py front-window-bounds --app "AppName"   → {x:100, y:50, w:800, h:600}
3. # Scroll down in the window center
   $PY desktop_ops.py scroll --amount -5 --x 500 --y 350
4. $PY desktop_ops.py screenshot                             → check if target visible
5. $PY target_resolver.py --app "AppName" --text "target item" --python $PY
6. IF not found: scroll more and retry (max 5 scrolls)
7. IF found: click it
Example 5: Handle an unexpected dialog
1. # During any operation, if the expected UI doesn't match:
2. $PY desktop_ops.py screenshot → examine what's on screen
3. # If a dialog is visible, OCR it:
   $PY ocr_text.py --app "AppName" --python $PY
4. # Find and click the appropriate button (OK, Cancel, Allow, etc.)
   $PY target_resolver.py --app "AppName" --text "OK" --python $PY
5. $PY desktop_ops.py click --x <x> --y <y>
6. # After dialog is dismissed, re-get window bounds and continue
   $PY desktop_ops.py front-window-bounds --app "AppName"

Reference Documents

Load as needed:

DocumentWhen to read
references/workflow.mdCore 8-step closed loop
references/platform-macos.mdmacOS-specific tools and permissions
references/platform-windows.mdWindows setup
references/platform-linux.mdLinux X11/Wayland setup
references/operation-patterns.mdReusable task templates
references/validation-patterns.mdTwo-stage validation
references/precise-targeting.md5-layer precision targeting
references/target-providers.mdProvider ordering and fallback contract
references/coordinate-reconstruction.mdRebuild click coordinates from screenshot evidence
references/chat-app-macos.mdChat app workflow
references/app-wechat-desktop.mdCross-platform WeChat guidance
references/cleanup-rules.mdCleanup timing and scope
references/collaboration-rules.mdWhen multi-agent collaboration is justified
references/example-cases.mdRepeatable task examples
references/reproducible-setup.mdHost bring-up checklist

Scope

Use this skill for: chat apps, browsers, file managers, editors, office apps, system settings, any closed desktop software with no usable API.

Hard Rules

  1. Always run auto-setup gate first
  2. Always use EXACT parameter names from CLI reference — never guess
  3. Always scope OCR to the target app window — NEVER full-screen OCR
  4. Always: focus-app → front-window-bounds → OCR within window → verify → act
  5. Always pass --python $PY to ocr_text.py and target_resolver.py
  6. Always verify coordinates are within window bounds before clicking
  7. Always re-get window bounds after any UI state change (login, dialog, navigation)
  8. Use insert-newline for line breaks; never use \n in type --text
  9. For send actions: prefer visible send button; use press --key return only when verified
  10. One action at a time; verify after each
  11. Maximum 3 retries per action; each retry must recapture fresh state
  12. Cleanup is mandatory at task end
  13. If verification fails, recapture and rebuild — do not retry blindly

© LeoYeAI, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 40 other files (scripts, references) in skills/desktop-agent-ops of LeoYeAI/openclaw-master-skills.

  • SKILL.md
  • _meta.json
  • agents/openai.yaml
  • references/app-wechat-desktop.md
  • references/app-wechat-macos.md
  • references/app-wechat-windows.md
  • references/chat-app-macos.md
  • references/cleanup-rules.md
  • references/collaboration-rules.md
  • references/coordinate-reconstruction.md
  • references/eval-scenarios.md
  • references/example-cases.md
  • references/market-precision-targeting-gap-analysis.md
  • references/operation-patterns.md
  • references/platform-linux.md
  • references/platform-macos.md
  • references/platform-windows.md
  • references/precise-targeting.md
  • references/reproducible-setup.md
  • … and 22 more

Open the folder on GitHubat commit e5199b5

Compare with similar skills

Desktop Agent Ops next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Desktop Agent Ops compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Desktop Agent Ops this skillLeoYeAI/openclaw-master-skills2.2k—~4.5kAutomated safety check: PassMIT
Open Computer UseiFurySt/open-codex-computer-use2.4k—~1.5kAutomated safety check: PassMIT
Waku Computer Useegoist/waku1.6k—~3.6kAutomated safety check: PassGPL-3.0
Drive Screencoleam00/skills674—~5.4kAutomated safety check: PassMIT
Browser MCP Agentantibrow/anti-detect-browser-skills171 repos~4.2kAutomated safety check: WarnMIT
Gooeypi Computer Useam-will/gooey-pi941—~370Automated safety check: PassMIT

Similar skills

  • Open Computer Use

    iFurySt/open-codex-computer-use

    Platform-neutral guidance for using Open Computer Use, the open-source Computer Use MCP server and CLI for macOS, Linux, and Windows.

    2.4k GitHub stars~1.5k tokensUpdated 2 days ago
    Productivity & AutomationAuto-check passed
  • Control local macOS, Windows, and Linux apps through Waku Computer Use.

    1.6k GitHub stars~3.6k tokensUpdated 7 days ago
    Productivity & AutomationAuto-check passed
  • Drive Screen

    coleam00/skills

    Take real control of the desktop - list and focus windows, type, paste, click, scroll, and screenshot - on Windows, macOS or Linux, and drive other coding-agent sessions running in terminals.

    674 GitHub stars~5.4k tokensUpdated 2 days ago
    Productivity & AutomationAuto-check passed
  • Browser MCP Agent

    antibrow/anti-detect-browser-skills

    Give an AI agent its own real browser over MCP tool calls - launch, navigate, click, fill, screenshot, extract text, run JS - with a kernel-level real-device fingerprint and a persistent profile, so…

    17 GitHub starsUsed in 1 repo~4.2k tokens
    Productivity & AutomationAuto-check: warnings
  • Gooeypi Computer Use

    am-will/gooey-pi

    Drive native applications on macOS, Windows, or Linux through the separately installed TryCUA driver CLI.

    941 GitHub stars~370 tokensUpdated 29 days ago
    Productivity & AutomationAuto-check passed
  • Computer Use

    johnson7788/MultiUserClaw

    Drive the user's desktop in the background — clicking, typing, scrolling, dragging — without stealing the cursor, keyboard focus, or switching virtual desktops / Spaces.

    327 GitHub starsUsed in 1 repo~2.8k tokens
    Productivity & AutomationAuto-check: warnings

More from LeoYeAI/openclaw-master-skills

All 1,235 skills in this repo
  • DevOps Pipeline Management

    LeoYeAI/openclaw-master-skills

    Manages pipelines on a DevOps quality and efficiency platform through its OpenAPI: list workspaces and templates, create, update, run and cancel pipelines, and read run records.

    2.2k GitHub stars~4.2k tokensUpdated 2 mo ago
    Auto-check: notes
  • Feishu Document Collaboration

    LeoYeAI/openclaw-master-skills

    Patches OpenClaw's Feishu extension so an edited document triggers an isolated agent session that reads the doc and replies inline, turning it into a live chat space.

    2.2k GitHub stars~2k tokensUpdated 2 mo ago
    Auto-check passed
  • Files Memory System

    LeoYeAI/openclaw-master-skills

    Multi-context memory management system for OpenClaw agents with group-isolated storage, global shared memory, workspace organization, and group-specific skills isolation.

    2.2k GitHub stars~3.8k tokensUpdated 2 mo ago
    Auto-check passed
  • GEO-Claw AI Visibility Agent

    LeoYeAI/openclaw-master-skills

    Runs a brand's AI-search visibility work end to end: diagnosing how AI platforms represent it, repositioning it, producing AI-optimized content and monitoring ongoing mentions.

    2.2k GitHub stars~4.7k tokensUpdated 2 mo ago
    Auto-check passed
  • Google Workspace CLI

    LeoYeAI/openclaw-master-skills

    Installs and authenticates the gws CLI, then automates Gmail, Drive, Sheets, Calendar, Docs, Chat and Tasks with ready-made recipes, persona bundles and security audits.

    2.2k GitHub stars~2.6k tokensUpdated 2 mo ago
    Auto-check: notes
  • HealthFit Health Advisors

    LeoYeAI/openclaw-master-skills

    Runs four advisor roles, a fitness coach, nutritionist, data analyst and TCM practitioner, to build a health profile and track workouts, diet and wellness over time.

    2.2k GitHub stars~4.4k tokensUpdated 2 mo ago
    Auto-check passed

Works with

Questions about Desktop Agent Ops

What does Desktop Agent Ops do?

Execute cross-platform desktop tasks through a packaged desktop automation skill that guides the main agent to observe the screen, focus apps and windows, call helper scripts for screenshots and…. Desktop Agent Ops is an agent skill from LeoYeAI/openclaw-master-skills. Execute cross-platform desktop tasks through a packaged desktop automation skill that guides the main agent to observe the screen, focus apps and windows, call helper scripts for screenshots and input actions, verify each step, clean up task context, and only escalate to multi-agent collaboration when tasks become clearly multi-window or multi-app.

When should I use Desktop Agent Ops?

Desktop Agent Ops fits situations like: the user wants desktop GUI control; native app operation; click and type flows; cross-platform desktop workflows on macOS.

How do I install Desktop Agent Ops in Claude Code?

Run `npx skills add LeoYeAI/openclaw-master-skills --skill desktop-agent-ops -a claude-code`. Or copy the skill folder (skills/desktop-agent-ops in LeoYeAI/openclaw-master-skills) into .claude/skills/desktop-agent-ops in your project. Claude Code loads it when a task matches its description.

How do I install Desktop Agent Ops in Codex?

Run `npx skills add LeoYeAI/openclaw-master-skills --skill desktop-agent-ops -a codex`. Or copy the skill folder (skills/desktop-agent-ops in LeoYeAI/openclaw-master-skills) into .agents/skills/desktop-agent-ops in your project. Codex loads it when a task matches its description.

Can I use Desktop Agent Ops in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add LeoYeAI/openclaw-master-skills --skill desktop-agent-ops -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/desktop-agent-ops, .gemini/skills/desktop-agent-ops, .github/skills/desktop-agent-ops and .opencode/skills/desktop-agent-ops in your project.

What does Desktop Agent Ops need to run?

Going by SKILL.md and its folder, Desktop Agent Ops needs the command-line tools its instructions call (python3). Our summary lists: Python 3.

Does Desktop Agent Ops access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Desktop Agent Ops safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Desktop Agent Ops use?

Desktop Agent Ops is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Desktop Agent Ops use?

About 4.5k tokens (SKILL.md is roughly 18k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 20k tokens, read only when the agent opens those files.

What are the alternatives to Desktop Agent Ops?

Skills that share tags, products or a category with Desktop Agent Ops: Open Computer Use (iFurySt/open-codex-computer-use, 2.4k stars), Waku Computer Use (egoist/waku, 1.6k stars), Drive Screen (coleam00/skills, 674 stars) and Browser MCP Agent (antibrow/anti-detect-browser-skills, 17 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Desktop Agent Ops?

LeoYeAI (a GitHub user) maintains it in LeoYeAI/openclaw-master-skills, which has 2,160 GitHub stars. The repository holds 1,235 skills in this directory. The repository was last updated on July 20, 2026.

Source: LeoYeAI/openclaw-master-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.