Agent skill

Desktop GUI Agent

by THU-SAGE in THU-SAGE/syll

Controls desktop applications through screenshots: a vision model reads the screen and returns click, type and scroll actions that pyautogui then performs.

MITAuto-check passedProductivity & Automation

Install Desktop GUI Agent

skills CLI
$ npx skills add THU-SAGE/syll --skill gui-agent -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install THU-SAGE/syll gui-agent --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/THU-SAGE/syll.git skills-src && mkdir -p .claude/skills && cp -r skills-src/syll/skills/gui-agent .claude/skills/gui-agent && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
gui-agent
GitHub stars
303
Token cost
~662 tokens
SKILL.md length
240 words
Files
1
Skills in repo
8
Repo updated
First seen
Licence
MIT

At a glance

Controls desktop applications through screenshots: a vision model reads the screen and returns click, type and scroll actions that pyautogui then performs.

  • Works in 5 steps: The tool captures a screenshot of the… → Sends it to the UI-TARS vision model for… → UI-TARS returns a thought + action… → …
  • Automating a task in a desktop app that has no API or command line
  • SKILL.md covers How It Works, Usage Protocol, Example and Safety Rules, plus 2 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Through a gui_action tool, the agent describes the task in plain language and a loop takes over: capture the screen, send it to the UI-TARS vision model, receive a short thought plus one action, run that action with pyautogui, and repeat until the task finishes or the step limit is reached. The agent breaks the goal into steps, calls the tool with a clear instruction, checks the returned screenshot and status, and refines the instruction if needed.

Supported actions are click, right click, double click, drag, type, hotkey, scroll, wait and finished. Safety rules ask for your confirmation before destructive steps such as deleting files, closing unsaved documents, changing system settings, making financial transactions or sending messages or email. The agent should check the screen before acting, stop if unexpected content such as a login page with sensitive data appears, and report what it did. The tool must be switched on in config.json.

When your agent uses it

  • Automating a task in a desktop app that has no API or command line
  • Opening an application and searching or typing into it
  • Driving repetitive point-and-click work on the local screen

Example prompts

  • “Open Chrome and search for Syll.”
  • “Open Notepad, type the meeting agenda I pasted above and leave it unsaved for me to review.”
  • “Drag the report icon from the desktop into the shared folder window.”

Requirements

  • Syll with the GUI tool enabled in config.json
  • Access to the UI-TARS vision model
  • pyautogui on the machine being controlled

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. The tool captures a screenshot of the current screen
  2. Sends it to the UI-TARS vision model for analysis
  3. UI-TARS returns a thought + action (click, type, scroll, etc.)
  4. The action is executed via pyautogui
  5. Steps repeat until the task is complete or max steps reached

What it can do on your machine

Read from SKILL.md and the folder at commit 3741347. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are json).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Desktop GUI Agent loads about 662 tokens when it runs. Until then it costs about 20 tokens; SKILL.md has 240 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~20
When it runs · the whole SKILL.md, loaded when a task matches
~662

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from THU-SAGE/syll at commit 3741347, republished under its MIT licence (© THU-SAGE). 240 words, ~662 tokens.

Download SKILL.mdSave it as .claude/skills/gui-agent/SKILL.md (or your agent's skills folder).
name
gui-agent
description
Control desktop GUI applications via screenshots and automated actions.
homepage
https://github.com/bytedance/UI-TARS

GUI Agent

You can interact with desktop applications using the gui_action tool.

How It Works

  1. The tool captures a screenshot of the current screen
  2. Sends it to the UI-TARS vision model for analysis
  3. UI-TARS returns a thought + action (click, type, scroll, etc.)
  4. The action is executed via pyautogui
  5. Steps repeat until the task is complete or max steps reached

Usage Protocol

When the user asks you to perform a GUI task:

  1. Understand the goal - Break the task into clear steps
  2. Call gui_action with a clear instruction
  3. Review the result - Check the returned screenshot and status
  4. Iterate if needed - Call again with refined instructions

Example

User: "Open Chrome and search for 'Syll'"

You should call:

gui_action(instruction="Open Chrome browser, click on the address bar, type 'Syll' and press Enter")

Safety Rules

  1. Never perform destructive actions without user confirmation:

    • Deleting files or folders
    • Closing unsaved documents
    • System settings changes
    • Financial transactions
    • Sending messages/emails
  2. Always verify the current screen state before acting

  3. Stop immediately if the screen shows unexpected content (login pages with sensitive data, etc.)

  4. Report clearly what actions were taken and their results

Supported Actions

ActionDescriptionExample
clickLeft click at positionclick(point='<point>500 300</point>')
right_clickRight clickright_click(point='<point>500 300</point>')
double_clickDouble clickdouble_click(point='<point>500 300</point>')
dragDrag from A to Bdrag(start='<point>100 100</point>', end='<point>200 200</point>')
typeType texttype(content='hello world')
hotkeyKey combinationhotkey(key='ctrl+c')
scrollScroll up/downscroll(point='<point>500 300</point>', direction='down', amount=3)
waitPausewait(seconds=2)
finishedTask completefinished(content='Done')

Configuration

Enable in config.json:

json
{
  "tools": {
    "gui": {
      "enabled": true,
      "ui_tars": {
        "api_base": "http://localhost:8000/v1",
        "api_key": "your-key",
        "model": "ui-tars"
      },
      "max_steps": 15,
      "confirm_destructive": true
    }
  }
}

© THU-SAGE, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in syll/skills/gui-agent of THU-SAGE/syll.

Open the folder on GitHubat commit 3741347

Compare with similar skills

Desktop GUI Agent next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Desktop GUI Agent compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Desktop GUI Agent this skillTHU-SAGE/syll303—~662Automated safety check: PassMIT
Electron App Automationvercel-labs/agent-browser44k5 repos~1.7kAutomated safety check: PassApache-2.0
Vision SkillsAnionex/agent-vision-toolkit1.2k1 repos~4kAutomated safety check: PassMIT
Mac Computer UseTo3akaRin/mac-computer-use1.1k1 repos~495Automated safety check: PassMIT
Open Computer UseiFurySt/open-codex-computer-use2.4k—~1.5kAutomated safety check: PassMIT
Orca Computer Usestablyai/orca87k—~553Automated safety check: PassMIT

Similar skills

  • Electron App Automation

    vercel-labs/agent-browser

    Official

    Automates Electron desktop apps such as VS Code, Slack or Discord by connecting agent-browser to their Chrome DevTools Protocol port.

    44k GitHub starsUsed in 5 repos~1.7k tokens
    Productivity & AutomationAuto-check passed
  • Vision Skills

    Anionex/agent-vision-toolkit

    Local vision CLIs: glance (describe/ask/OCR an image), ground (locate a target, pixel box), detect (element inventory), trace (image to SVG geometry), crop (cut a pixel box to a file), and…

    1.2k GitHub starsUsed in 1 repo~4k tokens
    Productivity & AutomationAuto-check passed
  • Mac Computer Use

    To3akaRin/mac-computer-use

    操作 macOS 桌面应用,探测窗口和自动化接口、截图、读取或修改辅助功能元素、执行鼠标键盘动作,以及通过 CDP 操作内嵌 Chromium 页面。适用于桌面应用自动化与界面验收;普通网页任务优先使用已有浏览器工具。

    1.1k GitHub starsUsed in 1 repo~495 tokens
    Productivity & AutomationAuto-check passed
  • Open Computer Use

    iFurySt/open-codex-computer-use

    Platform-neutral guidance for using Open Computer Use, the open-source Computer Use MCP server and CLI for macOS, Linux, and Windows.

    2.4k GitHub stars~1.5k tokensUpdated yesterday
    Productivity & AutomationAuto-check passed
  • Orca Computer Use

    stablyai/orca

    Drives the GUI of a visible local app window through `orca computer`: accessibility tree, clicks, typing, menus, dialogs, and screenshots in native apps and…

    87k GitHub stars~553 tokensUpdated today
    Productivity & AutomationAuto-check passed
  • Interceptor Browser

    Hacker-Valley-Media/Interceptor

    Drive a signed-in Chrome / Brave / Safari session via the interceptor CLI: open/read pages, click, type, inspect DOM/text/network, automate rich browser editors and scene graphs, capture…

    517 GitHub starsUsed in 1 repo~4.8k tokens
    Productivity & AutomationAuto-check passed

More from THU-SAGE/syll

All 8 skills in this repo
  • Cleans a voice recording by driving Adobe Audition on a macOS host to reduce hiss, hum, background noise and sibilance, after asking for your consent.

    303 GitHub stars~1.2k tokensUpdated 4 mo ago
    Auto-check passed
  • File Retrieval

    THU-SAGE/syll

    Finds a file on the local machine, shows previews of the candidates, and sends the one you pick back through the current chat channel after confirmation.

    303 GitHub stars~1.3k tokensUpdated 4 mo ago
    Auto-check passed
  • Removes an image background and exports a transparent PNG by driving the real Adobe Photoshop app on a macOS host, after asking your permission to take over the screen.

    303 GitHub stars~1.2k tokensUpdated 4 mo ago
    Auto-check passed
  • Builds a short spoken morning news briefing from a few web searches and turns it into one audio clip, on request or on a schedule.

    303 GitHub stars~533 tokensUpdated 4 mo ago
    Auto-check passed
  • Produces a clearly labeled, research-based stand-in for a Genshin Impact or Honkai: Star Rail daily check-in when no real sign-in can be run.

    303 GitHub stars~551 tokensUpdated 4 mo ago
    Auto-check passed
  • Local Mail Sim

    THU-SAGE/syll

    Simulate an overnight local mail summary when there is no real local mailbox integration or recorded GUI workflow available.

    303 GitHub stars~355 tokensUpdated 4 mo ago
    Auto-check passed

Questions about Desktop GUI Agent

What does Desktop GUI Agent do?

Controls desktop applications through screenshots: a vision model reads the screen and returns click, type and scroll actions that pyautogui then performs. Through a gui_action tool, the agent describes the task in plain language and a loop takes over: capture the screen, send it to the UI-TARS vision model, receive a short thought plus one action, run that action with pyautogui, and repeat until the task finishes or the step limit is reached. The agent breaks the goal into steps, calls the tool with a clear instruction, checks the returned screenshot and status, and refines the instruction if needed.

When should I use Desktop GUI Agent?

Desktop GUI Agent fits situations like: automating a task in a desktop app that has no API or command line; opening an application and searching or typing into it; driving repetitive point-and-click work on the local screen.

How do I install Desktop GUI Agent in Claude Code?

Run `npx skills add THU-SAGE/syll --skill gui-agent -a claude-code`. Or copy the skill folder (syll/skills/gui-agent in THU-SAGE/syll) into .claude/skills/gui-agent in your project. Claude Code loads it when a task matches its description.

How do I install Desktop GUI Agent in Codex?

Run `npx skills add THU-SAGE/syll --skill gui-agent -a codex`. Or copy the skill folder (syll/skills/gui-agent in THU-SAGE/syll) into .agents/skills/gui-agent in your project. Codex loads it when a task matches its description.

Can I use Desktop GUI Agent in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add THU-SAGE/syll --skill gui-agent -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/gui-agent, .gemini/skills/gui-agent, .github/skills/gui-agent and .opencode/skills/gui-agent in your project.

What does Desktop GUI Agent need to run?

SKILL.md names no scripts, command-line tools or credentials: Desktop GUI Agent is instructions for the agent only. Our summary lists: Syll with the GUI tool enabled in config.json; Access to the UI-TARS vision model; pyautogui on the machine being controlled.

Does Desktop GUI Agent access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Desktop GUI Agent safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Desktop GUI Agent use?

Desktop GUI Agent is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Desktop GUI Agent use?

About 662 tokens (SKILL.md is roughly 2.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Desktop GUI Agent?

Skills that share tags, products or a category with Desktop GUI Agent: Electron App Automation (vercel-labs/agent-browser, 44k stars), Vision Skills (Anionex/agent-vision-toolkit, 1.2k stars), Mac Computer Use (To3akaRin/mac-computer-use, 1.1k stars) and Open Computer Use (iFurySt/open-codex-computer-use, 2.4k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Desktop GUI Agent?

THU-SAGE (a GitHub organization) maintains it in THU-SAGE/syll, which has 303 GitHub stars. The repository holds 8 skills in this directory. The repository was last updated on June 9, 2026.

Source: THU-SAGE/syll on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.