Gui Automation
am-will/gooey-pi
A skill your agent uses when you need to visually interact with a GUI — test buttons, fill forms, verify visual layouts, fuzz web pages, automate user flows, take screenshots, or perform end-to-end…
Full Computer Use for OpenClaw via kasmweb/chrome Docker sidecar.
$ npx skills add LeoYeAI/openclaw-master-skills --skill virtual-desktop -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install LeoYeAI/openclaw-master-skills virtual-desktop --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/LeoYeAI/openclaw-master-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/virtual-desktop .claude/skills/virtual-desktop && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "virtual-desktop" agent skill from https://github.com/LeoYeAI/openclaw-master-skills/tree/main/skills/virtual-desktop into .claude/skills/virtual-desktop/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "virtual-desktop", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/LeoYeAI/openclaw-master-skills/tree/main/skills/virtual-desktopType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add LeoYeAI/openclaw-master-skills --skill virtual-desktop -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install LeoYeAI/openclaw-master-skills virtual-desktop --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/LeoYeAI/openclaw-master-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/virtual-desktop .agents/skills/virtual-desktop && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "virtual-desktop" agent skill from https://github.com/LeoYeAI/openclaw-master-skills/tree/main/skills/virtual-desktop into .agents/skills/virtual-desktop/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "virtual-desktop", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add LeoYeAI/openclaw-master-skills --skill virtual-desktop -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install LeoYeAI/openclaw-master-skills virtual-desktop --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/LeoYeAI/openclaw-master-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/virtual-desktop .cursor/skills/virtual-desktop && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "virtual-desktop" agent skill from https://github.com/LeoYeAI/openclaw-master-skills/tree/main/skills/virtual-desktop into .cursor/skills/virtual-desktop/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "virtual-desktop", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/LeoYeAI/openclaw-master-skills.git --path skills/virtual-desktop--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add LeoYeAI/openclaw-master-skills --skill virtual-desktop -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install LeoYeAI/openclaw-master-skills virtual-desktop --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/LeoYeAI/openclaw-master-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/virtual-desktop .gemini/skills/virtual-desktop && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "virtual-desktop" agent skill from https://github.com/LeoYeAI/openclaw-master-skills/tree/main/skills/virtual-desktop into .gemini/skills/virtual-desktop/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "virtual-desktop", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install LeoYeAI/openclaw-master-skills virtual-desktopInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add LeoYeAI/openclaw-master-skills --skill virtual-desktop -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/LeoYeAI/openclaw-master-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/virtual-desktop .github/skills/virtual-desktop && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "virtual-desktop" agent skill from https://github.com/LeoYeAI/openclaw-master-skills/tree/main/skills/virtual-desktop into .github/skills/virtual-desktop/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "virtual-desktop", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add LeoYeAI/openclaw-master-skills --skill virtual-desktop -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install LeoYeAI/openclaw-master-skills virtual-desktop --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/LeoYeAI/openclaw-master-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/virtual-desktop .opencode/skills/virtual-desktop && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "virtual-desktop" agent skill from https://github.com/LeoYeAI/openclaw-master-skills/tree/main/skills/virtual-desktop into .opencode/skills/virtual-desktop/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "virtual-desktop", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
virtual-desktopFull Computer Use for OpenClaw via kasmweb/chrome Docker sidecar.
Virtual Desktop is an agent skill from LeoYeAI/openclaw-master-skills. Full Computer Use for OpenClaw via kasmweb/chrome Docker sidecar. Navigate any website, click, type, fill forms, extract data, upload files, screenshot on any platform including private authenticated accounts. Principal logs in once via noVNC. Sessions saved permanently in Docker volume. After one-time manual login via noVNC, agent can access authenticated platforms. CapSolver solves CAPTCHAs automatically. Browserbase profile available for residential proxy and stealth. Claude vision analyses screenshots and…
Its SKILL.md is about 4.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files (for example `CONFIGURATION.md`, `README.md` and `_meta.json`).
It sits in Productivity & Automation, covering Desktop control, Forms and invoices and Containers. It works with Docker and Browserbase. The repository describes itself as: 🧠 Curated collection of 1209+ best OpenClaw skills — weekly updated by MyClaw.ai. The licence is MIT.
Read from SKILL.md and the folder at commit e5199b5. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships script files (Python), which the agent can run.
Shell commands in SKILL.md call:
dockercurlpython3apt-getsshFrom the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
github.comx.comFrom URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
CAPSOLVER_API_KEYBROWSERBASE_API_KEYCAPSOLVER_KEYANTHROPIC_API_KEYFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Virtual Desktop loads about 4.5k tokens when it runs. Until then it costs about 159 tokens; SKILL.md has 308 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check noted patterns worth knowing about, such as sudo or a known installer.
# 2. Update .envgrep -q "VNC_PW" .env || echo "VNC_PW=CHANGE_ME_NOW" >> .envgrep -q "BROWSER_CDP_URL" .env || echo "BROWSER_CDP_URL=http://browser:9222" >> .envgrep -q "CAPSOLVER_API_KEY" .env || echo "CAPSOLVER_API_KEY=" >> .envgrep -q "BROWSERBASE_API_KEY" .env || echo "BROWSERBASE_API_KEY=" >> .envCAPSOLVER_KEY=$(grep CAPSOLVER_API_KEY .env | cut -d= -f2)resolution (if CAPSOLVER_API_KEY set in .env)To enable: add CAPSOLVER_API_KEY=xxx to .env (~$0.001 per CAPTCHA)# Add to .env: BROWSERBASE_API_KEY=xxx# Add to .env: PROXY_URL=http://user:pass@proxy:portAutomated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from LeoYeAI/openclaw-master-skills at commit e5199b5, republished under its MIT licence (© LeoYeAI). 308 words, ~4,466 tokens.
.claude/skills/virtual-desktop/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.Gives the agent a persistent authenticated browser (kasmweb/chrome) running as a Docker sidecar. Principal logs in once via noVNC. Sessions saved permanently. After one-time manual login via noVNC, agent can access authenticated platforms — no credentials needed after setup.
| Capability | What it means |
|---|---|
| ANALYZE | Read any page, extract structured data, monitor changes over time |
| PLAN | Map the UI, identify selectors, prepare multi-step action sequences |
| EXECUTE | Click, type, fill forms, submit, upload, download, navigate any flow |
| SELF-CORRECT | Screenshot error state, identify root cause, retry with alternate approach |
| IMPROVE | Write UI patterns and selector maps to .learnings/ after every session |
Use cases: Google Workspace · social platforms · admin dashboards · e-commerce · forms · market research · data extraction · any platform with or without an API
/workspace/
├── screenshots/ ← visual proof of every action (auto-created)
├── logs/browser/ ← full tracebacks (auto-created)
├── AUDIT.md ← append-only action log
├── memory/YYYY-MM-DD.md ← daily session summary
└── .learnings/
├── ERRORS.md ← errors, broken selectors, ref maps
└── LEARNINGS.md ← patterns, timing, navigation per platformAgent executes all steps automatically:
OPENCLAW_DIR="${OPENCLAW_DIR:-$(pwd)}"
cd "$OPENCLAW_DIR"
CONTAINER="${OPENCLAW_CONTAINER:-$(docker ps --format '{{.Names}}' | grep openclaw | head -1)}"
# 1. Add kasmweb/chrome to docker-compose.yml
python3 -c "
import yaml, os
VNC_PW = os.environ.get('VNC_PW', 'CHANGE_ME_NOW')
with open('docker-compose.yml') as f:
data = yaml.safe_load(f)
data.setdefault('services', {})['browser'] = {
'image': 'kasmweb/chrome:1.15.0',
'container_name': 'browser',
'restart': 'unless-stopped',
'shm_size': '1gb',
'ports': ['6901:6901', '9222:9222'],
'environment': [
'VNC_PW=' + VNC_PW,
'RESOLUTION=1920x1080',
'CHROME_ARGS=--remote-debugging-port=9222 --remote-debugging-address=0.0.0.0 --no-sandbox --disable-blink-features=AutomationControlled --disable-infobars'
],
'volumes': ['browser-profile:/home/kasm-user/chrome-profile'],
'networks': list(data.get('networks', {'default': None}).keys())
}
data.setdefault('volumes', {})['browser-profile'] = None
with open('docker-compose.yml', 'w') as f:
yaml.dump(data, f, default_flow_style=False, allow_unicode=True)
print('docker-compose.yml updated')
"
# 2. Update .env
grep -q "VNC_PW" .env || echo "VNC_PW=CHANGE_ME_NOW" >> .env
grep -q "BROWSER_CDP_URL" .env || echo "BROWSER_CDP_URL=http://browser:9222" >> .env
grep -q "CAPSOLVER_API_KEY" .env || echo "CAPSOLVER_API_KEY=" >> .env
grep -q "BROWSERBASE_API_KEY" .env || echo "BROWSERBASE_API_KEY=" >> .env
# 3. Update openclaw.json (hot reload — no restart needed)
# OpenClaw watches openclaw.json and applies changes automatically
python3 -c "
import json, os
f = 'data/.openclaw/openclaw.json'
with open(f) as fp: cfg = json.load(fp)
cfg.setdefault('browser', {}).update({
'enabled': True, 'headless': False,
'noSandbox': True, 'defaultProfile': 'chrome-sidecar'
})
profiles = cfg['browser'].setdefault('profiles', {})
profiles['chrome-sidecar'] = {'cdpUrl': 'http://browser:9222', 'color': '#4285F4'}
bb_key = os.environ.get('BROWSERBASE_API_KEY', '')
if bb_key:
profiles['browserbase'] = {'cdpUrl': f'wss://connect.browserbase.com?apiKey={bb_key}', 'color': '#F97316'}
print('Browserbase profile enabled')
with open(f, 'w') as fp: json.dump(cfg, fp, indent=2)
print('openclaw.json updated — hot reload applied automatically')
"
# 4. Start ONLY the new browser container (no need to restart OpenClaw)
# docker compose up -d --no-deps starts only the specified service
# OpenClaw keeps running without interruption
docker compose up -d --no-deps browser
echo "Chrome Desktop container started"
sleep 12
# 5. Install Python dependencies inside the OpenClaw container
docker exec "$CONTAINER" pip install requests playwright --break-system-packages -q
echo "✅ Python dependencies installed (requests, playwright)"
# 6. Install Playwright Chromium browser binaries
docker exec "$CONTAINER" node /app/node_modules/playwright-core/cli.js install chromium
echo "✅ Playwright Chromium binaries installed"
# 7. Download CapSolver extension for autonomous CAPTCHA solving
docker exec "$CONTAINER" bash -c "
apt-get install -y unzip curl -qq
curl -sL https://github.com/capsolver/capsolver-browser-extension/releases/latest/download/chrome.zip -o /tmp/capsolver.zip
unzip -q /tmp/capsolver.zip -d /data/.openclaw/capsolver-extension
"
CAPSOLVER_KEY=$(grep CAPSOLVER_API_KEY .env | cut -d= -f2)
if [ -n "$CAPSOLVER_KEY" ]; then
docker exec "$CONTAINER" bash -c "
sed -i \"s/apiKey: \"\"/apiKey: \"$CAPSOLVER_KEY\"/\" /data/.openclaw/capsolver-extension/assets/config.js 2>/dev/null
"
fi
# 8. Create workspace directories
docker exec "$CONTAINER" bash -c "
mkdir -p /data/.openclaw/workspace/skills/virtual-desktop
mkdir -p /workspace/screenshots /workspace/logs/browser /workspace/.learnings
touch /workspace/AUDIT.md /workspace/.learnings/ERRORS.md /workspace/.learnings/LEARNINGS.md
"
# 9. Deploy browser_control.py from skill directory
docker cp {baseDir}/browser_control.py "$CONTAINER":/data/.openclaw/workspace/skills/virtual-desktop/browser_control.py
echo "✅ browser_control.py deployed"
# 10. Verify
docker ps | grep -E "openclaw|browser"
curl -s http://localhost:9222/json > /dev/null && echo "✅ Chrome CDP active" || echo "⏳ Chrome starting..."
docker exec "$CONTAINER" python3 /data/.openclaw/workspace/skills/virtual-desktop/browser_control.py status
# 11. Notify principal
VPS_IP=$(curl -s ifconfig.me 2>/dev/null || echo "YOUR_VPS_IP")
echo ""
echo "Virtual Desktop ready — https://${VPS_IP}:6901"
echo "Log in to all your platforms then reply DONE."https://YOUR_VPS_IP:6901 — login: kasm_user / password: your VNC_PW valueOpen Chrome via noVNC and log in to all platforms.
Sessions saved in Docker volume browser-profile — survive restarts — valid forever.
Session expired → agent notifies via Telegram → principal reconnects in 2 min.
These commands are native to OpenClaw. The agent already knows them. This reference is here for quick lookup during missions.
# Navigation & tabs
openclaw browser open <url>
openclaw browser navigate <url>
openclaw browser go-back
openclaw browser reload
openclaw browser tab new | select 2 | close 2 | tabs
openclaw browser resize 1920 1080
# Inspection
openclaw browser snapshot # numeric refs
openclaw browser snapshot --interactive # role refs e12 — best for actions
openclaw browser snapshot --efficient # token-efficient mode
openclaw browser snapshot --selector "#main" # scoped to element
openclaw browser snapshot --labels # screenshot with ref labels
openclaw browser screenshot
openclaw browser screenshot --full-page
openclaw browser screenshot --ref e12 # capture specific element
openclaw browser pdf
# Actions
openclaw browser click e12
openclaw browser click e12 --double
openclaw browser hover e12
openclaw browser type e12 "text"
openclaw browser type e12 "text" --submit
openclaw browser press Enter | Tab | Escape | "Control+a" | "Control+c" | "Control+v"
openclaw browser select e9 "option"
openclaw browser drag e10 e11
openclaw browser scrollintoview e12
openclaw browser fill --fields '[{"ref":"e1","type":"text","value":"text"}]'
openclaw browser dialog --accept | --dismiss
openclaw browser evaluate --fn '(el) => el.textContent' --ref e7
openclaw browser highlight e12
# Wait — critical for dynamic pages
openclaw browser wait "#selector"
openclaw browser wait --text "expected text"
openclaw browser wait --url "**/dashboard"
openclaw browser wait --load networkidle
openclaw browser wait --load domcontentloaded
openclaw browser wait "#el" --load networkidle --fn "window.ready===true" --timeout-ms 15000
# Files
openclaw browser upload /tmp/openclaw/uploads/file.pdf
openclaw browser download e12 file.pdf
openclaw browser waitfordownload file.pdf
# Cookies & storage
openclaw browser cookies | cookies set k v --url "https://x.com" | cookies clear
openclaw browser storage local get | set k v | clear
openclaw browser storage session clear
# Browser configuration
openclaw browser set offline on | off
openclaw browser set headers --headers-json '{"X-Custom":"val"}'
openclaw browser set geo 48.8566 2.3522 --origin "https://example.com"
openclaw browser set media dark
openclaw browser set timezone Europe/Paris
openclaw browser set locale fr-FR
openclaw browser set device "iPhone 14"
# Debug & monitoring
openclaw browser console --level error
openclaw browser errors
openclaw browser requests --filter api
openclaw browser responsebody "**/api" --max-chars 5000
openclaw browser trace start | stop
openclaw browser status | start | stop
# Stealth — if VPS is blocked
openclaw browser --browser-profile browserbase open <url>BC="python3 /data/.openclaw/workspace/skills/virtual-desktop/browser_control.py"
$BC screenshot <url> [label]
$BC navigate <url> [selector]
$BC click <url> <selector>
$BC click_xy <url> <x> <y>
$BC fill <url> <selector> <value>
$BC select <url> <selector> <value>
$BC hover <url> <selector>
$BC scroll <url> <direction> [pixels]
$BC keyboard <url> <selector> <key>
$BC extract <url> <selector> [output_file]
$BC wait_for <url> <selector> [timeout_ms]
$BC upload <url> <file_selector> <file_path>
$BC analyze <url_or_image> [question] ← CLAUDE VISION
$BC captcha <url> ← AUTONOMOUS CAPTCHA
$BC workflow <json_steps_file> ← MULTI-STEP WORKFLOW
$BC status[
{ "action": "goto", "target": "https://TARGET_URL" },
{ "action": "captcha" },
{ "action": "analyze", "value": "Identify the key elements on this page" },
{ "action": "wait_for", "target": ".loaded", "timeout_ms": 5000 },
{ "action": "fill", "target": "#field", "value": "text" },
{ "action": "click", "target": "#btn" },
{ "action": "click_xy", "x": 960, "y": 540 },
{ "action": "scroll", "direction": "down" },
{ "action": "hover", "target": "#menu" },
{ "action": "select", "target": "#list", "value": "option" },
{ "action": "keyboard", "target": "#input", "value": "Enter" },
{ "action": "extract", "target": ".data", "value": "/workspace/tasks/out.json" },
{ "action": "screenshot" },
{ "action": "wait", "value": "2" }
]1. Auto-detection on every page load
→ reCAPTCHA v2/v3, hCaptcha, Cloudflare Turnstile
2. CapSolver API resolution (if CAPSOLVER_API_KEY set in .env)
→ Extracts sitekey → sends to API → receives token → injects → continues
3. Cloudflare Turnstile
→ CapSolver Chrome extension handles it in background → wait 60s → continues
4. Fallback
→ Screenshot → Telegram → principal opens noVNC → solves → agent continues
To enable: add CAPSOLVER_API_KEY=xxx to .env (~$0.001 per CAPTCHA)# Option 1 — Browserbase (CAPTCHA + stealth + residential proxy built-in)
# Free tier: 1 concurrent session, 1h/month — browserbase.com
# Add to .env: BROWSERBASE_API_KEY=xxx
# Use: openclaw browser --browser-profile browserbase open <url>
# Option 2 — Custom proxy in browser_control.py
# Add to .env: PROXY_URL=http://user:pass@proxy:port
# In get_browser(): ctx = browser.new_context(proxy={"server": os.environ["PROXY_URL"]}, ...)# Web page → auto screenshot + analysis
$BC analyze https://example.com "What does this page sell?"
# AI-generated image
$BC analyze https://site.com/image.png "Describe the visual elements"
# Existing screenshot
$BC analyze /workspace/screenshots/capture.png "Is there a form on this page?"
# Inside a JSON workflow
{ "action": "analyze", "value": "Identify all form fields on this page" }BEFORE EVERY BROWSER ACTION:
1. Log to AUDIT.md: "BEFORE [action] on [url]"
2. Detect CAPTCHA → resolve automatically if present
3. Execute
4. Screenshot as proof
5. Log to AUDIT.md: "OK/FAILED [action]"
6. Telegram report if real-world consequences
NEVER:
→ Access platforms not authorized by the principal
→ Execute payments without explicit approval
→ Fail silently — always log
→ Retry more than 3 times without alerting principalCAPTCHA → CapSolver auto → fallback noVNC
CLOUDFLARE → switch to --browser-profile browserbase
SESSION EXPIRED → Telegram → principal opens noVNC → reconnects
ELEMENT MISSING → use analyze to understand the new layout
→ log to .learnings/ERRORS.md with ref map
TIMEOUT → check /workspace/logs/browser/YYYY-MM-DD.logThis skill opens port 6901 (noVNC) on your VPS and stores authenticated browser sessions permanently. Before installing, understand what this means:
REQUIRED before running:
1. Set a strong VNC_PW in .env — never use the default
2. Firewall port 6901 to your IP only:
→ Hostinger: Panel → VPS → Firewall → restrict port 6901 to your IP
→ Or use SSH tunnel instead of opening the port publicly:
ssh -L 6901:localhost:6901 user@YOUR_VPS_IP
Then access via http://localhost:6901
3. The agent will have autonomous access to whatever accounts
you log into via noVNC — only log into accounts you trust
the agent to access
4. CAPSOLVER_API_KEY, BROWSERBASE_API_KEY, ANTHROPIC_API_KEY
are optional — only add them if you trust those services
and understand their costs| File | When | Content |
|---|---|---|
/workspace/AUDIT.md | Every action | Before + after log, append-only |
/workspace/screenshots/YYYY-MM-DD_*.png | Every action | Visual proof |
/workspace/screenshots/YYYY-MM-DD_*_analysis.txt | After analyze | Vision result |
/workspace/logs/browser/YYYY-MM-DD.log | On exception | Full traceback |
/workspace/.learnings/ERRORS.md | On failure | Errors + ref maps |
/workspace/.learnings/LEARNINGS.md | On discovery | Patterns + timing |
/workspace/tasks/lessons.md | During mission | Immediate task capture |
/workspace/memory/YYYY-MM-DD.md | Daily | Session summary |
After every browser session, write immediately:
# ERRORS.md
## [YYYY-MM-DD] [Platform] — [Title]
**Priority**: low|medium|high — **Status**: pending|resolved
**What happened**: ... **Root cause**: ... **Fix**: ... **Ref map**: {"e12":"e15"}
# LEARNINGS.md
## [YYYY-MM-DD] [Platform] — [Pattern]
**Category**: navigation|interaction|timing|auth_flow|captcha|vision
**Discovery**: ... **Usage**: ...© LeoYeAI, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 4 other files in skills/virtual-desktop of LeoYeAI/openclaw-master-skills.
Open the folder on GitHubat commit e5199b5
Virtual Desktop next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Virtual Desktop this skillLeoYeAI/openclaw-master-skills | 2.2k | — | ~4.5k | Automated safety check: Notes | MIT | |
| Gui Automationam-will/gooey-pi | 941 | — | ~1.4k | Automated safety check: Pass | MIT | |
| Cloudroutermanaflow-ai/manaflow | 1.1k | — | ~4.7k | Automated safety check: Pass | MIT | |
| Cua Driverdavidondrej/skills | 4.1k | — | ~821 | Automated safety check: Pass | MIT | |
| Local Vscuse ValidationOfficeDev/microsoft-365-agents-toolkit | 781 | — | ~10k | Automated safety check: Warn | Custom licence | |
| Copaw Opschujianyun/skills | 740 | — | ~1.3k | Automated safety check: Pass | Custom licence |
am-will/gooey-pi
A skill your agent uses when you need to visually interact with a GUI — test buttons, fill forms, verify visual layouts, fuzz web pages, automate user flows, take screenshots, or perform end-to-end…
manaflow-ai/manaflow
Manage cloud development sandboxes with cloudrouter. An agent skill from manaflow-ai/manaflow.
davidondrej/skills
Use Cua Driver for desktop or browser tasks that are awkward or unavailable through Bash/APIs, or when the user explicitly wants GUI interaction: app testing, visual bug reproduction, form filling…
OfficeDev/microsoft-365-agents-toolkit
A skill your agent uses when: setting up shared local vscuse prerequisites, credentials, runner, local or pinned published Docker images, local VSIX, env variables, and common rules for Microsoft…
chujianyun/skills
CoPaw 运维助手。用于用户提到 copaw 运维、服务无响应、渠道断连、MCP 失败、模型调用失败、cron 不执行、Docker 部署、重载、重启或重置恢复时使用。优先执行状态检查与故障分流;涉及重启、重载、重置、配置修改等高影响动作时,先向用户说明再执行。
GreptimeTeam/greptimedb
Packages a locally built GreptimeDB debug binary into a development-only Docker image for local-cluster testing, with an optional push to a dev registry.
LeoYeAI/openclaw-master-skills
Manages pipelines on a DevOps quality and efficiency platform through its OpenAPI: list workspaces and templates, create, update, run and cancel pipelines, and read run records.
LeoYeAI/openclaw-master-skills
Patches OpenClaw's Feishu extension so an edited document triggers an isolated agent session that reads the doc and replies inline, turning it into a live chat space.
LeoYeAI/openclaw-master-skills
Multi-context memory management system for OpenClaw agents with group-isolated storage, global shared memory, workspace organization, and group-specific skills isolation.
LeoYeAI/openclaw-master-skills
Runs a brand's AI-search visibility work end to end: diagnosing how AI platforms represent it, repositioning it, producing AI-optimized content and monitoring ongoing mentions.
LeoYeAI/openclaw-master-skills
Installs and authenticates the gws CLI, then automates Gmail, Drive, Sheets, Calendar, Docs, Chat and Tasks with ready-made recipes, persona bundles and security audits.
LeoYeAI/openclaw-master-skills
Runs four advisor roles, a fitness coach, nutritionist, data analyst and TCM practitioner, to build a health profile and track workouts, diet and wellness over time.
Works with
Full Computer Use for OpenClaw via kasmweb/chrome Docker sidecar. Virtual Desktop is an agent skill from LeoYeAI/openclaw-master-skills. Full Computer Use for OpenClaw via kasmweb/chrome Docker sidecar.
Virtual Desktop fits situations like: openClaw via kasmweb/chrome Docker sidecar; tasks that involve Desktop control; tasks that involve Forms and invoices.
Run `npx skills add LeoYeAI/openclaw-master-skills --skill virtual-desktop -a claude-code`. Or copy the skill folder (skills/virtual-desktop in LeoYeAI/openclaw-master-skills) into .claude/skills/virtual-desktop in your project. Claude Code loads it when a task matches its description.
Run `npx skills add LeoYeAI/openclaw-master-skills --skill virtual-desktop -a codex`. Or copy the skill folder (skills/virtual-desktop in LeoYeAI/openclaw-master-skills) into .agents/skills/virtual-desktop in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add LeoYeAI/openclaw-master-skills --skill virtual-desktop -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/virtual-desktop, .gemini/skills/virtual-desktop, .github/skills/virtual-desktop and .opencode/skills/virtual-desktop in your project.
Going by SKILL.md and its folder, Virtual Desktop needs Python for the scripts in its folder, the command-line tools its instructions call (docker, curl, python3, apt-get and ssh) and credentials named CAPSOLVER_API_KEY, BROWSERBASE_API_KEY, CAPSOLVER_KEY and ANTHROPIC_API_KEY. Our summary lists: Python 3; Docker; A credential in CAPSOLVER_API_KEY; A credential in BROWSERBASE_API_KEY.
SKILL.md names 2 domains. In commands or code: github.com and x.com; the agent is likely to contact these when it follows the instructions. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found notes only (mentions a .env file), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.
Virtual Desktop is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 4.5k tokens (SKILL.md is roughly 18k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Virtual Desktop: Gui Automation (am-will/gooey-pi, 941 stars), Cloudrouter (manaflow-ai/manaflow, 1.1k stars), Cua Driver (davidondrej/skills, 4.1k stars) and Local Vscuse Validation (OfficeDev/microsoft-365-agents-toolkit, 781 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
LeoYeAI (a GitHub user) maintains it in LeoYeAI/openclaw-master-skills, which has 2,159 GitHub stars. The repository holds 972 skills in this directory. The repository was last updated on July 20, 2026.
Source: LeoYeAI/openclaw-master-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.