Agent skill

Virtual Desktop

by LeoYeAI in LeoYeAI/openclaw-master-skills

Full Computer Use for OpenClaw via kasmweb/chrome Docker sidecar.

MITAuto-check: notesProductivity & Automation

Install Virtual Desktop

skills CLI
$ npx skills add LeoYeAI/openclaw-master-skills --skill virtual-desktop -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install LeoYeAI/openclaw-master-skills virtual-desktop --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/LeoYeAI/openclaw-master-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/virtual-desktop .claude/skills/virtual-desktop && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
virtual-desktop
GitHub stars
2.2k
Token cost
~4.5k tokens
SKILL.md length
308 words
Files
5
Skills in repo
972
Repo updated
First seen
Licence
MIT

At a glance

Full Computer Use for OpenClaw via kasmweb/chrome Docker sidecar.

  • OpenClaw via kasmweb/chrome Docker sidecar
  • SKILL.md covers What this skill does, Required Workspace Structure, Setup — Run Once and Initial Login — Once Per…, plus 11 more sections
  • Runs Python scripts from its folder; calls docker, curl and python3; reaches github.com and x.com; needs CAPSOLVER_API_KEY and BROWSERBASE_API_KEY
  • Tasks that involve Desktop control

What it does

Virtual Desktop is an agent skill from LeoYeAI/openclaw-master-skills. Full Computer Use for OpenClaw via kasmweb/chrome Docker sidecar. Navigate any website, click, type, fill forms, extract data, upload files, screenshot on any platform including private authenticated accounts. Principal logs in once via noVNC. Sessions saved permanently in Docker volume. After one-time manual login via noVNC, agent can access authenticated platforms. CapSolver solves CAPTCHAs automatically. Browserbase profile available for residential proxy and stealth. Claude vision analyses screenshots and…

Its SKILL.md is about 4.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files (for example `CONFIGURATION.md`, `README.md` and `_meta.json`).

It sits in Productivity & Automation, covering Desktop control, Forms and invoices and Containers. It works with Docker and Browserbase. The repository describes itself as: 🧠 Curated collection of 1209+ best OpenClaw skills — weekly updated by MyClaw.ai. The licence is MIT.

When your agent uses it

  • OpenClaw via kasmweb/chrome Docker sidecar
  • Tasks that involve Desktop control
  • Tasks that involve Forms and invoices

Example prompts

  • “/virtual-desktop”

Requirements

  • Python 3
  • Docker
  • A credential in CAPSOLVER_API_KEY
  • A credential in BROWSERBASE_API_KEY

What it can do on your machine

Read from SKILL.md and the folder at commit e5199b5. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • docker
    • curl
    • python3
    • apt-get
    • ssh

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • github.com
    • x.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • CAPSOLVER_API_KEY
    • BROWSERBASE_API_KEY
    • CAPSOLVER_KEY
    • ANTHROPIC_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Virtual Desktop loads about 4.5k tokens when it runs. Until then it costs about 159 tokens; SKILL.md has 308 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~159
When it runs · the whole SKILL.md, loaded when a task matches
~4.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteMentions a .env fileSKILL.md:133
    # 2. Update .env
  • NoteMentions a .env fileSKILL.md:134
    grep -q "VNC_PW"              .env || echo "VNC_PW=CHANGE_ME_NOW"                  >> .env
  • NoteMentions a .env fileSKILL.md:135
    grep -q "BROWSER_CDP_URL"     .env || echo "BROWSER_CDP_URL=http://browser:9222"   >> .env
  • NoteMentions a .env fileSKILL.md:136
    grep -q "CAPSOLVER_API_KEY"   .env || echo "CAPSOLVER_API_KEY="                    >> .env
  • NoteMentions a .env fileSKILL.md:137
    grep -q "BROWSERBASE_API_KEY" .env || echo "BROWSERBASE_API_KEY="                  >> .env
  • NoteMentions a .env fileSKILL.md:180
    CAPSOLVER_KEY=$(grep CAPSOLVER_API_KEY .env | cut -d= -f2)
  • NoteMentions a .env fileSKILL.md:359
    resolution (if CAPSOLVER_API_KEY set in .env)
  • NoteMentions a .env fileSKILL.md:368
    To enable: add CAPSOLVER_API_KEY=xxx to .env (~$0.001 per CAPTCHA)
  • NoteMentions a .env fileSKILL.md:378
    # Add to .env: BROWSERBASE_API_KEY=xxx
  • NoteMentions a .env fileSKILL.md:382
    # Add to .env: PROXY_URL=http://user:pass@proxy:port

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from LeoYeAI/openclaw-master-skills at commit e5199b5, republished under its MIT licence (© LeoYeAI). 308 words, ~4,466 tokens.

Download SKILL.mdSave it as .claude/skills/virtual-desktop/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
virtual-desktop
description
Full Computer Use for OpenClaw via kasmweb/chrome Docker sidecar. Navigate any website, click, type, fill forms, extract data, upload files, screenshot on any platform including private authenticated accounts. Principal logs in once via noVNC. Sessions saved permanently in Docker volume. After one-time manual login via noVNC, agent can access authenticated platforms. CapSolver solves CAPTCHAs automatically. Browserbase profile available for residential proxy and stealth. Claude vision analyses screenshots and AI-generated images natively. Every action logged. Every discovery improves performance via .learnings/.
version
3.0.0
author
Georges Andronescu (Wesley Armando)
license
MIT

Virtual Desktop — Universal Execution Layer

What this skill does

Gives the agent a persistent authenticated browser (kasmweb/chrome) running as a Docker sidecar. Principal logs in once via noVNC. Sessions saved permanently. After one-time manual login via noVNC, agent can access authenticated platforms — no credentials needed after setup.

CapabilityWhat it means
ANALYZERead any page, extract structured data, monitor changes over time
PLANMap the UI, identify selectors, prepare multi-step action sequences
EXECUTEClick, type, fill forms, submit, upload, download, navigate any flow
SELF-CORRECTScreenshot error state, identify root cause, retry with alternate approach
IMPROVEWrite UI patterns and selector maps to .learnings/ after every session

Use cases: Google Workspace · social platforms · admin dashboards · e-commerce · forms · market research · data extraction · any platform with or without an API


Required Workspace Structure

/workspace/
├── screenshots/          ← visual proof of every action (auto-created)
├── logs/browser/         ← full tracebacks (auto-created)
├── AUDIT.md              ← append-only action log
├── memory/YYYY-MM-DD.md  ← daily session summary
└── .learnings/
    ├── ERRORS.md         ← errors, broken selectors, ref maps
    └── LEARNINGS.md      ← patterns, timing, navigation per platform

Setup — Run Once

Agent executes all steps automatically:

bash
OPENCLAW_DIR="${OPENCLAW_DIR:-$(pwd)}"
cd "$OPENCLAW_DIR"
CONTAINER="${OPENCLAW_CONTAINER:-$(docker ps --format '{{.Names}}' | grep openclaw | head -1)}"

# 1. Add kasmweb/chrome to docker-compose.yml
python3 -c "
import yaml, os
VNC_PW = os.environ.get('VNC_PW', 'CHANGE_ME_NOW')
with open('docker-compose.yml') as f:
    data = yaml.safe_load(f)
data.setdefault('services', {})['browser'] = {
    'image': 'kasmweb/chrome:1.15.0',
    'container_name': 'browser',
    'restart': 'unless-stopped',
    'shm_size': '1gb',
    'ports': ['6901:6901', '9222:9222'],
    'environment': [
        'VNC_PW=' + VNC_PW,
        'RESOLUTION=1920x1080',
        'CHROME_ARGS=--remote-debugging-port=9222 --remote-debugging-address=0.0.0.0 --no-sandbox --disable-blink-features=AutomationControlled --disable-infobars'
    ],
    'volumes': ['browser-profile:/home/kasm-user/chrome-profile'],
    'networks': list(data.get('networks', {'default': None}).keys())
}
data.setdefault('volumes', {})['browser-profile'] = None
with open('docker-compose.yml', 'w') as f:
    yaml.dump(data, f, default_flow_style=False, allow_unicode=True)
print('docker-compose.yml updated')
"

# 2. Update .env
grep -q "VNC_PW"              .env || echo "VNC_PW=CHANGE_ME_NOW"                  >> .env
grep -q "BROWSER_CDP_URL"     .env || echo "BROWSER_CDP_URL=http://browser:9222"   >> .env
grep -q "CAPSOLVER_API_KEY"   .env || echo "CAPSOLVER_API_KEY="                    >> .env
grep -q "BROWSERBASE_API_KEY" .env || echo "BROWSERBASE_API_KEY="                  >> .env

# 3. Update openclaw.json (hot reload — no restart needed)
# OpenClaw watches openclaw.json and applies changes automatically
python3 -c "
import json, os
f = 'data/.openclaw/openclaw.json'
with open(f) as fp: cfg = json.load(fp)
cfg.setdefault('browser', {}).update({
    'enabled': True, 'headless': False,
    'noSandbox': True, 'defaultProfile': 'chrome-sidecar'
})
profiles = cfg['browser'].setdefault('profiles', {})
profiles['chrome-sidecar'] = {'cdpUrl': 'http://browser:9222', 'color': '#4285F4'}
bb_key = os.environ.get('BROWSERBASE_API_KEY', '')
if bb_key:
    profiles['browserbase'] = {'cdpUrl': f'wss://connect.browserbase.com?apiKey={bb_key}', 'color': '#F97316'}
    print('Browserbase profile enabled')
with open(f, 'w') as fp: json.dump(cfg, fp, indent=2)
print('openclaw.json updated — hot reload applied automatically')
"

# 4. Start ONLY the new browser container (no need to restart OpenClaw)
# docker compose up -d --no-deps starts only the specified service
# OpenClaw keeps running without interruption
docker compose up -d --no-deps browser
echo "Chrome Desktop container started"
sleep 12

# 5. Install Python dependencies inside the OpenClaw container
docker exec "$CONTAINER" pip install requests playwright --break-system-packages -q
echo "✅ Python dependencies installed (requests, playwright)"

# 6. Install Playwright Chromium browser binaries
docker exec "$CONTAINER"   node /app/node_modules/playwright-core/cli.js install chromium
echo "✅ Playwright Chromium binaries installed"

# 7. Download CapSolver extension for autonomous CAPTCHA solving
docker exec "$CONTAINER" bash -c "
apt-get install -y unzip curl -qq
curl -sL https://github.com/capsolver/capsolver-browser-extension/releases/latest/download/chrome.zip -o /tmp/capsolver.zip
unzip -q /tmp/capsolver.zip -d /data/.openclaw/capsolver-extension
"
CAPSOLVER_KEY=$(grep CAPSOLVER_API_KEY .env | cut -d= -f2)
if [ -n "$CAPSOLVER_KEY" ]; then
  docker exec "$CONTAINER" bash -c "
  sed -i \"s/apiKey: \"\"/apiKey: \"$CAPSOLVER_KEY\"/\" /data/.openclaw/capsolver-extension/assets/config.js 2>/dev/null
  "
fi

# 8. Create workspace directories
docker exec "$CONTAINER" bash -c "
mkdir -p /data/.openclaw/workspace/skills/virtual-desktop
mkdir -p /workspace/screenshots /workspace/logs/browser /workspace/.learnings
touch /workspace/AUDIT.md /workspace/.learnings/ERRORS.md /workspace/.learnings/LEARNINGS.md
"

# 9. Deploy browser_control.py from skill directory
docker cp {baseDir}/browser_control.py   "$CONTAINER":/data/.openclaw/workspace/skills/virtual-desktop/browser_control.py
echo "✅ browser_control.py deployed"

# 10. Verify
docker ps | grep -E "openclaw|browser"
curl -s http://localhost:9222/json > /dev/null && echo "✅ Chrome CDP active" || echo "⏳ Chrome starting..."
docker exec "$CONTAINER"   python3 /data/.openclaw/workspace/skills/virtual-desktop/browser_control.py status

# 11. Notify principal
VPS_IP=$(curl -s ifconfig.me 2>/dev/null || echo "YOUR_VPS_IP")
echo ""
echo "Virtual Desktop ready — https://${VPS_IP}:6901"
echo "Log in to all your platforms then reply DONE."

Initial Login — Once Per Platform

https://YOUR_VPS_IP:6901  —  login: kasm_user  /  password: your VNC_PW value

Open Chrome via noVNC and log in to all platforms. Sessions saved in Docker volume browser-profile — survive restarts — valid forever. Session expired → agent notifies via Telegram → principal reconnects in 2 min.


Native OpenClaw Browser Commands — Quick Reference

These commands are native to OpenClaw. The agent already knows them. This reference is here for quick lookup during missions.

bash
# Navigation & tabs
openclaw browser open <url>
openclaw browser navigate <url>
openclaw browser go-back
openclaw browser reload
openclaw browser tab new | select 2 | close 2 | tabs
openclaw browser resize 1920 1080

# Inspection
openclaw browser snapshot                          # numeric refs
openclaw browser snapshot --interactive            # role refs e12 — best for actions
openclaw browser snapshot --efficient              # token-efficient mode
openclaw browser snapshot --selector "#main"       # scoped to element
openclaw browser snapshot --labels                 # screenshot with ref labels
openclaw browser screenshot
openclaw browser screenshot --full-page
openclaw browser screenshot --ref e12              # capture specific element
openclaw browser pdf

# Actions
openclaw browser click e12
openclaw browser click e12 --double
openclaw browser hover e12
openclaw browser type e12 "text"
openclaw browser type e12 "text" --submit
openclaw browser press Enter | Tab | Escape | "Control+a" | "Control+c" | "Control+v"
openclaw browser select e9 "option"
openclaw browser drag e10 e11
openclaw browser scrollintoview e12
openclaw browser fill --fields '[{"ref":"e1","type":"text","value":"text"}]'
openclaw browser dialog --accept | --dismiss
openclaw browser evaluate --fn '(el) => el.textContent' --ref e7
openclaw browser highlight e12

# Wait — critical for dynamic pages
openclaw browser wait "#selector"
openclaw browser wait --text "expected text"
openclaw browser wait --url "**/dashboard"
openclaw browser wait --load networkidle
openclaw browser wait --load domcontentloaded
openclaw browser wait "#el" --load networkidle --fn "window.ready===true" --timeout-ms 15000

# Files
openclaw browser upload /tmp/openclaw/uploads/file.pdf
openclaw browser download e12 file.pdf
openclaw browser waitfordownload file.pdf

# Cookies & storage
openclaw browser cookies | cookies set k v --url "https://x.com" | cookies clear
openclaw browser storage local get | set k v | clear
openclaw browser storage session clear

# Browser configuration
openclaw browser set offline on | off
openclaw browser set headers --headers-json '{"X-Custom":"val"}'
openclaw browser set geo 48.8566 2.3522 --origin "https://example.com"
openclaw browser set media dark
openclaw browser set timezone Europe/Paris
openclaw browser set locale fr-FR
openclaw browser set device "iPhone 14"

# Debug & monitoring
openclaw browser console --level error
openclaw browser errors
openclaw browser requests --filter api
openclaw browser responsebody "**/api" --max-chars 5000
openclaw browser trace start | stop
openclaw browser status | start | stop

# Stealth — if VPS is blocked
openclaw browser --browser-profile browserbase open <url>

browser_control.py — Commands (auto-logging + CAPTCHA + vision)

bash
BC="python3 /data/.openclaw/workspace/skills/virtual-desktop/browser_control.py"

$BC screenshot  <url> [label]
$BC navigate    <url> [selector]
$BC click       <url> <selector>
$BC click_xy    <url> <x> <y>
$BC fill        <url> <selector> <value>
$BC select      <url> <selector> <value>
$BC hover       <url> <selector>
$BC scroll      <url> <direction> [pixels]
$BC keyboard    <url> <selector> <key>
$BC extract     <url> <selector> [output_file]
$BC wait_for    <url> <selector> [timeout_ms]
$BC upload      <url> <file_selector> <file_path>
$BC analyze     <url_or_image> [question]        ← CLAUDE VISION
$BC captcha     <url>                            ← AUTONOMOUS CAPTCHA
$BC workflow    <json_steps_file>                ← MULTI-STEP WORKFLOW
$BC status

Workflow JSON Format

json
[
  { "action": "goto",       "target": "https://TARGET_URL" },
  { "action": "captcha" },
  { "action": "analyze",    "value": "Identify the key elements on this page" },
  { "action": "wait_for",   "target": ".loaded", "timeout_ms": 5000 },
  { "action": "fill",       "target": "#field", "value": "text" },
  { "action": "click",      "target": "#btn" },
  { "action": "click_xy",   "x": 960, "y": 540 },
  { "action": "scroll",     "direction": "down" },
  { "action": "hover",      "target": "#menu" },
  { "action": "select",     "target": "#list", "value": "option" },
  { "action": "keyboard",   "target": "#input", "value": "Enter" },
  { "action": "extract",    "target": ".data", "value": "/workspace/tasks/out.json" },
  { "action": "screenshot" },
  { "action": "wait",       "value": "2" }
]

CAPTCHA — Autonomous Strategy

1. Auto-detection on every page load
   → reCAPTCHA v2/v3, hCaptcha, Cloudflare Turnstile

2. CapSolver API resolution (if CAPSOLVER_API_KEY set in .env)
   → Extracts sitekey → sends to API → receives token → injects → continues

3. Cloudflare Turnstile
   → CapSolver Chrome extension handles it in background → wait 60s → continues

4. Fallback
   → Screenshot → Telegram → principal opens noVNC → solves → agent continues

To enable: add CAPSOLVER_API_KEY=xxx to .env (~$0.001 per CAPTCHA)

Residential Proxy — If Site Blocks the VPS

bash
# Option 1 — Browserbase (CAPTCHA + stealth + residential proxy built-in)
# Free tier: 1 concurrent session, 1h/month — browserbase.com
# Add to .env: BROWSERBASE_API_KEY=xxx
# Use: openclaw browser --browser-profile browserbase open <url>

# Option 2 — Custom proxy in browser_control.py
# Add to .env: PROXY_URL=http://user:pass@proxy:port
# In get_browser(): ctx = browser.new_context(proxy={"server": os.environ["PROXY_URL"]}, ...)

Claude Vision — Analyze Images and Pages

bash
# Web page → auto screenshot + analysis
$BC analyze https://example.com "What does this page sell?"

# AI-generated image
$BC analyze https://site.com/image.png "Describe the visual elements"

# Existing screenshot
$BC analyze /workspace/screenshots/capture.png "Is there a form on this page?"

# Inside a JSON workflow
{ "action": "analyze", "value": "Identify all form fields on this page" }

Execution Protocol

BEFORE EVERY BROWSER ACTION:
  1. Log to AUDIT.md: "BEFORE [action] on [url]"
  2. Detect CAPTCHA → resolve automatically if present
  3. Execute
  4. Screenshot as proof
  5. Log to AUDIT.md: "OK/FAILED [action]"
  6. Telegram report if real-world consequences

NEVER:
  → Access platforms not authorized by the principal
  → Execute payments without explicit approval
  → Fail silently — always log
  → Retry more than 3 times without alerting principal

Error Recovery

CAPTCHA          → CapSolver auto → fallback noVNC
CLOUDFLARE       → switch to --browser-profile browserbase
SESSION EXPIRED  → Telegram → principal opens noVNC → reconnects
ELEMENT MISSING  → use analyze to understand the new layout
                 → log to .learnings/ERRORS.md with ref map
TIMEOUT          → check /workspace/logs/browser/YYYY-MM-DD.log

Security

This skill opens port 6901 (noVNC) on your VPS and stores authenticated browser sessions permanently. Before installing, understand what this means:

REQUIRED before running:
  1. Set a strong VNC_PW in .env — never use the default
  2. Firewall port 6901 to your IP only:
     → Hostinger: Panel → VPS → Firewall → restrict port 6901 to your IP
     → Or use SSH tunnel instead of opening the port publicly:
        ssh -L 6901:localhost:6901 user@YOUR_VPS_IP
        Then access via http://localhost:6901

  3. The agent will have autonomous access to whatever accounts
     you log into via noVNC — only log into accounts you trust
     the agent to access

  4. CAPSOLVER_API_KEY, BROWSERBASE_API_KEY, ANTHROPIC_API_KEY
     are optional — only add them if you trust those services
     and understand their costs

Files Written By This Skill

FileWhenContent
/workspace/AUDIT.mdEvery actionBefore + after log, append-only
/workspace/screenshots/YYYY-MM-DD_*.pngEvery actionVisual proof
/workspace/screenshots/YYYY-MM-DD_*_analysis.txtAfter analyzeVision result
/workspace/logs/browser/YYYY-MM-DD.logOn exceptionFull traceback
/workspace/.learnings/ERRORS.mdOn failureErrors + ref maps
/workspace/.learnings/LEARNINGS.mdOn discoveryPatterns + timing
/workspace/tasks/lessons.mdDuring missionImmediate task capture
/workspace/memory/YYYY-MM-DD.mdDailySession summary

Self-Improvement

After every browser session, write immediately:

# ERRORS.md
## [YYYY-MM-DD] [Platform] — [Title]
**Priority**: low|medium|high — **Status**: pending|resolved
**What happened**: ... **Root cause**: ... **Fix**: ... **Ref map**: {"e12":"e15"}

# LEARNINGS.md
## [YYYY-MM-DD] [Platform] — [Pattern]
**Category**: navigation|interaction|timing|auth_flow|captcha|vision
**Discovery**: ... **Usage**: ...

© LeoYeAI, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files in skills/virtual-desktop of LeoYeAI/openclaw-master-skills.

  • SKILL.md
  • CONFIGURATION.md
  • README.md
  • _meta.json
  • browser_control.py

Open the folder on GitHubat commit e5199b5

Compare with similar skills

Virtual Desktop next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Virtual Desktop compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Virtual Desktop this skillLeoYeAI/openclaw-master-skills2.2k—~4.5kAutomated safety check: NotesMIT
Gui Automationam-will/gooey-pi941—~1.4kAutomated safety check: PassMIT
Cloudroutermanaflow-ai/manaflow1.1k—~4.7kAutomated safety check: PassMIT
Cua Driverdavidondrej/skills4.1k—~821Automated safety check: PassMIT
Local Vscuse ValidationOfficeDev/microsoft-365-agents-toolkit781—~10kAutomated safety check: WarnCustom licence
Copaw Opschujianyun/skills740—~1.3kAutomated safety check: PassCustom licence

Similar skills

  • Gui Automation

    am-will/gooey-pi

    A skill your agent uses when you need to visually interact with a GUI — test buttons, fill forms, verify visual layouts, fuzz web pages, automate user flows, take screenshots, or perform end-to-end…

    941 GitHub stars~1.4k tokensUpdated 28 days ago
    Productivity & AutomationAuto-check passed
  • Cloudrouter

    manaflow-ai/manaflow

    Manage cloud development sandboxes with cloudrouter. An agent skill from manaflow-ai/manaflow.

    1.1k GitHub stars~4.7k tokensUpdated 2 days ago
    Productivity & AutomationAuto-check passed
  • Cua Driver

    davidondrej/skills

    Use Cua Driver for desktop or browser tasks that are awkward or unavailable through Bash/APIs, or when the user explicitly wants GUI interaction: app testing, visual bug reproduction, form filling…

    4.1k GitHub stars~821 tokensUpdated yesterday
    Productivity & AutomationAuto-check passed
  • Local Vscuse Validation

    OfficeDev/microsoft-365-agents-toolkit

    A skill your agent uses when: setting up shared local vscuse prerequisites, credentials, runner, local or pinned published Docker images, local VSIX, env variables, and common rules for Microsoft…

    781 GitHub stars~10k tokensUpdated today
    Documents & OfficeAuto-check: warnings
  • Copaw Ops

    chujianyun/skills

    CoPaw 运维助手。用于用户提到 copaw 运维、服务无响应、渠道断连、MCP 失败、模型调用失败、cron 不执行、Docker 部署、重载、重启或重置恢复时使用。优先执行状态检查与故障分流;涉及重启、重载、重置、配置修改等高影响动作时,先向用户说明再执行。

    740 GitHub stars~1.3k tokensUpdated 13 days ago
    DevOps & CloudAuto-check passed
  • GreptimeDB Dev Docker Image

    GreptimeTeam/greptimedb

    Packages a locally built GreptimeDB debug binary into a development-only Docker image for local-cluster testing, with an optional push to a dev registry.

    6.7k GitHub stars~4k tokensUpdated today
    DevOps & CloudAuto-check: notes

More from LeoYeAI/openclaw-master-skills

All 972 skills in this repo
  • DevOps Pipeline Management

    LeoYeAI/openclaw-master-skills

    Manages pipelines on a DevOps quality and efficiency platform through its OpenAPI: list workspaces and templates, create, update, run and cancel pipelines, and read run records.

    2.2k GitHub stars~4.2k tokensUpdated 2 mo ago
    Auto-check: notes
  • Feishu Document Collaboration

    LeoYeAI/openclaw-master-skills

    Patches OpenClaw's Feishu extension so an edited document triggers an isolated agent session that reads the doc and replies inline, turning it into a live chat space.

    2.2k GitHub stars~2k tokensUpdated 2 mo ago
    Auto-check passed
  • Files Memory System

    LeoYeAI/openclaw-master-skills

    Multi-context memory management system for OpenClaw agents with group-isolated storage, global shared memory, workspace organization, and group-specific skills isolation.

    2.2k GitHub stars~3.8k tokensUpdated 2 mo ago
    Auto-check passed
  • GEO-Claw AI Visibility Agent

    LeoYeAI/openclaw-master-skills

    Runs a brand's AI-search visibility work end to end: diagnosing how AI platforms represent it, repositioning it, producing AI-optimized content and monitoring ongoing mentions.

    2.2k GitHub stars~4.7k tokensUpdated 2 mo ago
    Auto-check passed
  • Google Workspace CLI

    LeoYeAI/openclaw-master-skills

    Installs and authenticates the gws CLI, then automates Gmail, Drive, Sheets, Calendar, Docs, Chat and Tasks with ready-made recipes, persona bundles and security audits.

    2.2k GitHub stars~2.6k tokensUpdated 2 mo ago
    Auto-check: notes
  • HealthFit Health Advisors

    LeoYeAI/openclaw-master-skills

    Runs four advisor roles, a fitness coach, nutritionist, data analyst and TCM practitioner, to build a health profile and track workouts, diet and wellness over time.

    2.2k GitHub stars~4.4k tokensUpdated 2 mo ago
    Auto-check passed

Questions about Virtual Desktop

What does Virtual Desktop do?

Full Computer Use for OpenClaw via kasmweb/chrome Docker sidecar. Virtual Desktop is an agent skill from LeoYeAI/openclaw-master-skills. Full Computer Use for OpenClaw via kasmweb/chrome Docker sidecar.

When should I use Virtual Desktop?

Virtual Desktop fits situations like: openClaw via kasmweb/chrome Docker sidecar; tasks that involve Desktop control; tasks that involve Forms and invoices.

How do I install Virtual Desktop in Claude Code?

Run `npx skills add LeoYeAI/openclaw-master-skills --skill virtual-desktop -a claude-code`. Or copy the skill folder (skills/virtual-desktop in LeoYeAI/openclaw-master-skills) into .claude/skills/virtual-desktop in your project. Claude Code loads it when a task matches its description.

How do I install Virtual Desktop in Codex?

Run `npx skills add LeoYeAI/openclaw-master-skills --skill virtual-desktop -a codex`. Or copy the skill folder (skills/virtual-desktop in LeoYeAI/openclaw-master-skills) into .agents/skills/virtual-desktop in your project. Codex loads it when a task matches its description.

Can I use Virtual Desktop in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add LeoYeAI/openclaw-master-skills --skill virtual-desktop -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/virtual-desktop, .gemini/skills/virtual-desktop, .github/skills/virtual-desktop and .opencode/skills/virtual-desktop in your project.

What does Virtual Desktop need to run?

Going by SKILL.md and its folder, Virtual Desktop needs Python for the scripts in its folder, the command-line tools its instructions call (docker, curl, python3, apt-get and ssh) and credentials named CAPSOLVER_API_KEY, BROWSERBASE_API_KEY, CAPSOLVER_KEY and ANTHROPIC_API_KEY. Our summary lists: Python 3; Docker; A credential in CAPSOLVER_API_KEY; A credential in BROWSERBASE_API_KEY.

Does Virtual Desktop access the network?

SKILL.md names 2 domains. In commands or code: github.com and x.com; the agent is likely to contact these when it follows the instructions. This is read from the text; nothing was executed.

Is Virtual Desktop safe to install?

Our automated static check of SKILL.md found notes only (mentions a .env file), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Virtual Desktop use?

Virtual Desktop is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Virtual Desktop use?

About 4.5k tokens (SKILL.md is roughly 18k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Virtual Desktop?

Skills that share tags, products or a category with Virtual Desktop: Gui Automation (am-will/gooey-pi, 941 stars), Cloudrouter (manaflow-ai/manaflow, 1.1k stars), Cua Driver (davidondrej/skills, 4.1k stars) and Local Vscuse Validation (OfficeDev/microsoft-365-agents-toolkit, 781 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Virtual Desktop?

LeoYeAI (a GitHub user) maintains it in LeoYeAI/openclaw-master-skills, which has 2,159 GitHub stars. The repository holds 972 skills in this directory. The repository was last updated on July 20, 2026.

Source: LeoYeAI/openclaw-master-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.