Agent skill

Computer Use Action Picker

by mrmps in mrmps/classifier-dev

Pick a browser or desktop agent's next action by choosing among the actions actually on screen instead of inventing one.

MITAuto-check passedProductivity & Automation

Install Computer Use Action Picker

skills CLI
$ npx skills add mrmps/classifier-dev --skill computer-use-action-picker -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install mrmps/classifier-dev computer-use-action-picker --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/mrmps/classifier-dev.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/computer-use-action-picker .claude/skills/computer-use-action-picker && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
computer-use-action-picker
GitHub stars
424
Token cost
~1.5k tokens
SKILL.md length
536 words
Files
1
Skills in repo
21
Repo updated
First seen
Licence
MIT

At a glance

Pick a browser or desktop agent's next action by choosing among the actions actually on screen instead of inventing one.

  • Works in 4 steps: Candidates from the accessibility tree → One call per step → The loop → …
  • Debugging computer use
  • SKILL.md covers When not to use it, 1. Candidates from the…, 2. One call per step and 3. The loop, plus 2 more sections
  • Calls curl, jq and go; reaches classifier.dev

What it does

Computer Use Action Picker is an agent skill from mrmps/classifier-dev. Pick a browser or desktop agent's next action by choosing among the actions actually on screen instead of inventing one. Candidates come from the accessibility tree, the task goes in instructions, and a calibrated confidence decides whether to click or hand the step back to your own model. Use when building or debugging computer use, browser automation or a web agent, or on "it clicked the wrong thing" and "how do I stop it looping".

Its SKILL.md is about 1.5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Productivity & Automation, covering Desktop control, Accessibility and Browser automation. The repository describes itself as: Zero-shot text classification over plain HTTP — no API key, no account. One Cloudflare Worker, a CLI, and an MCP server. https://classifier.dev. The licence is MIT.

When your agent uses it

  • Debugging computer use
  • Browser automation
  • On it clicked the wrong thing and how do I stop it looping

Example prompts

  • “it clicked the wrong thing”
  • “how do I stop it looping”
  • “/computer-use-action-picker”

Requirements

  • Python 3

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Candidates from the accessibility tree
  2. One call per step
  3. The loop
  4. The gate

What it can do on your machine

Read from SKILL.md and the folder at commit b9211dd. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • curl
    • jq
    • go
    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • classifier.dev

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Computer Use Action Picker loads about 1.5k tokens when it runs. Until then it costs about 116 tokens; SKILL.md has 536 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~116
When it runs · the whole SKILL.md, loaded when a task matches
~1.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from mrmps/classifier-dev at commit b9211dd, republished under its MIT licence (© mrmps). 536 words, ~1,498 tokens.

Download SKILL.mdSave it as .claude/skills/computer-use-action-picker/SKILL.md (or your agent's skills folder).
name
computer-use-action-picker
description
Pick a browser or desktop agent's next action by choosing among the actions actually on screen instead of inventing one. Candidates come from the accessibility tree, the task goes in instructions, and a calibrated confidence decides whether to click or hand the step back to your own model. Use when building or debugging computer use, browser automation or a web agent, or on "it clicked the wrong thing" and "how do I stop it looping".
license
MIT

Choose the next action from what is on screen

An agent that writes its own action can name a button that is not there. Make the step a choice over the page's real elements and that failure goes away: the answer indexes a list you built, and carries a calibrated probability.

When not to use it

  • Steps needing reasoning the page does not show, like comparing prices across tabs. Classify the action, not the plan.
  • Anything irreversible: payments, deletions, sending mail. Gate those on a person, never on a number.
  • Fewer than about five obviously different candidates. Your own model knows already; the call is a round trip you do not need.

1. Candidates from the accessibility tree

Snapshot it — Playwright's page.accessibility.snapshot(), or Accessibility.getFullAXTree over the DevTools Protocol — and keep the nodes a person could act on:

button    "Add to cart"
combobox  "Size"              value "Choose a size"
link      "Continue shopping"
link      "Sign in"
textbox   "Search products"

Turn each node into one label in the words a person would use — click the Add to cart button, type into the Search products field — plus the moves that are not elements: scroll down, go back, the task is done, stop.

  • Prune by visibility. Off-screen, aria-hidden and zero-size nodes go; those get chosen and then fail to click.
  • Cap at 100, the label limit. Past that, keep the viewport and the nav landmarks and let scroll down reach the rest.
  • Keep names distinct. Three labels reading click the Edit button split the score three ways and none clears the gate. Say click Edit on the billing address row.
  • Always include the task is done, stop. Every call returns one of your labels, so with no stop candidate a finished task keeps clicking.

2. One call per step

The task goes in instructions, the page state is the input, the candidates are the labels — step.json:

{
  "labels": ["click the Add to cart button", "click the Size dropdown", "go back",
    "click the Continue shopping link", "type into the Search products field",
    "click the Sign in link", "scroll down", "the task is done, stop"],
  "instructions": "Choose the single next action for a browser agent. The task is: buy one medium blue t-shirt. Pick what makes progress and is possible on this page.",
  "inputs": ["Page: Blue cotton t-shirt. Size dropdown reads 'Choose a size'. Add to cart button present. Cart is empty."]
}
curl -s https://classifier.dev/v1/classify -H 'content-type: application/json' --data @step.json \
  | jq -r '.results[0] | "\(.confidence)  \(.label)"'
1  click the Size dropdown

Put one line of history in the input, or the picker re-picks the action it just took — where most loops start.

Show full SKILL.md (224 more words)Show less

3. The loop

pick.py, with a step budget and a stop condition. replay.json holds [page state, candidates] pairs recorded from a real run, so you can dry-run a change without driving a browser:

python
import json, urllib.request
API, ACT_AT, BUDGET = "https://classifier.dev/v1/classify", 0.8, 12
TASK = "buy one medium blue t-shirt"
INSTR = (f"Choose the single next action for a browser agent. The task is: {TASK}. "
         "The input is the page the agent sees and what it has already done. "
         "Pick the action that makes progress and is possible on this page.")

def pick(state, candidates):
    body = json.dumps({"labels": candidates[:100], "instructions": INSTR,
                       "inputs": [state]}).encode()
    req = urllib.request.Request(API, data=body, headers={
        "content-type": "application/json", "user-agent": "action-picker/1.0"})
    r = json.load(urllib.request.urlopen(req))["results"][0]
    return r["label"], r["confidence"]

for step, (state, cands) in enumerate(json.load(open("replay.json"))[:BUDGET], 1):
    label, conf = pick(state, cands)
    print(f"step {step}  {conf:.2f}  {label}" + ("" if conf >= ACT_AT else "  -> ask your own model"))
    if conf >= ACT_AT and label.startswith("the task is done"):
        print(f"stopped after {step} steps"); break

python3 pick.py over a five-state replay:

step 1  0.49  type into the Search products field  -> ask your own model
step 2  1.00  click the Size dropdown
step 3  1.00  click the Add to cart button
step 4  1.00  click the Proceed to checkout button
step 5  0.99  the task is done, stop
stopped after 5 steps

Set the user-agent header: Python's urllib default is blocked at the edge and returns 403 before the call is classified.

4. The gate

  • 0.8 and above — act. Navigation inside a task is the easy case: four of the five steps came back at 0.99 or 1.00.
  • 0.5 to 0.8 — hand the step to your own model with the same candidate list, to choose or to ask the user. Step 1 is a search page whose results do not hold the product; searching again and scrolling are both defensible, and the picker says so at 0.45 to 0.53 across runs.
  • Below 0.5 — stop and ask. Repeated low confidence means the task is not reachable from this screen; say so instead of clicking.

Two more stops whatever the confidence: the step budget, and the same action twice on an unchanged page. Both mean it is stuck.

What done looks like

Each step logs the candidate count, chosen label and confidence. Every acted step is at or above 0.8, the run ends on the stop label or the budget, and the replay file reproduces it without a browser.

© mrmps, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/computer-use-action-picker of mrmps/classifier-dev.

Open the folder on GitHubat commit b9211dd

Compare with similar skills

Computer Use Action Picker next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Computer Use Action Picker compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Computer Use Action Picker this skillmrmps/classifier-dev424—~1.5kAutomated safety check: PassMIT
Computer Usebam-bam-2/solo-skills3671 repos~915Automated safety check: PassMIT
Computer Usekunpengtalk/PeakCode166—~2.5kAutomated safety check: PassMIT
Computer Useautonomous-ai/Physical-AI-Operating-System381—~2kAutomated safety check: PassApache-2.0
Agentic Browser Testingpetrkindlmann/qa-skills163—~4.5kAutomated safety check: PassMIT
Codewhale Computer Use Controllercodewhale-hq/Codewhale41k—~1.4kAutomated safety check: PassMIT

Similar skills

  • Computer Use

    bam-bam-2/solo-skills

    Use Orca's computer-use CLI to inspect and operate local desktop app windows through accessibility trees, screenshots, and safe UI actions.

    367 GitHub starsUsed in 1 repo~915 tokens
    Productivity & AutomationAuto-check passed
  • Computer Use

    kunpengtalk/PeakCode

    Drive a macOS app through the computer tool — read its accessibility tree, act on elements, and use real input only when the tree cannot reach the target.

    166 GitHub stars~2.5k tokensUpdated 15 days ago
    Productivity & AutomationAuto-check passed
  • Computer Use

    autonomous-ai/Physical-AI-Operating-System

    Operate apps/websites on the paired Mac via Buddy: Calendar, Notes, forms, screenshots, files.

    381 GitHub stars~2k tokensUpdated today
    Productivity & AutomationAuto-check passed
  • Agentic Browser Testing

    petrkindlmann/qa-skills

    Goal-driven E2E testing where a browser agent (Playwright MCP / computer-use) reads a natural-language goal and explores the app via the accessibility tree to assert outcomes — no pre-written script.

    163 GitHub stars~4.5k tokensUpdated 3 mo ago
    Testing & QAAuto-check passed
  • Codewhale Computer Use Controller

    codewhale-hq/Codewhale

    Operates desktop apps and isolated or disposable browsers on registered computers through an observe, act, verify loop, separate from driving the user's own signed-in Chrome.

    41k GitHub stars~1.4k tokensUpdated today
    Productivity & AutomationAuto-check passed
  • Covers 原生 Playwright Browser Tool、Chromium 进程池、Session Context、页面与 ARIA snapshot/ref 权威、同源安全、诊断/截图和 Web Browser 面板投影。

    180 GitHub stars~2k tokensUpdated 10 days ago
    Productivity & AutomationAuto-check passed

More from mrmps/classifier-dev

All 21 skills in this repo
  • Bulk Classify

    mrmps/classifier-dev

    Sort many texts into your own categories without reading them, using a keyless HTTP API that returns a calibrated confidence per answer.

    424 GitHub stars~3.1k tokensUpdated yesterday
    Auto-check passed
  • Content Moderation Gate

    mrmps/classifier-dev

    Check user-generated text against a written policy before it is published.

    424 GitHub stars~1.5k tokensUpdated yesterday
    Auto-check passed
  • Label each context chunk keep, drop or replace-with-a-pointer and pass the survivors through byte for byte instead of summarising, with key-shaped chunks decided locally and never sent, and a…

    424 GitHub stars~1.5k tokensUpdated yesterday
    Auto-check passed
  • Document Intake Routing

    mrmps/classifier-dev

    Label each page of an intake packet with a document type and a page role before extraction runs, so only confident pages reach an extractor and the rest reach a person.

    424 GitHub stars~1.5k tokensUpdated yesterday
    Auto-check passed
  • Headline Filter Map Reduce

    mrmps/classifier-dev

    Filter hundreds or thousands of headlines, search results or feed items against a written brief before opening any of them, using a two-stage cascade that spends a fast model on everything and a…

    424 GitHub stars~1.5k tokensUpdated yesterday
    Auto-check passed
  • Type candidate (subject, sentence, object) triples against a fixed relation schema and flag triples that contradict each other, batched, with a calibrated confidence per edge so only confident edges…

    424 GitHub stars~1.5k tokensUpdated yesterday
    Auto-check passed

Questions about Computer Use Action Picker

What does Computer Use Action Picker do?

Pick a browser or desktop agent's next action by choosing among the actions actually on screen instead of inventing one. Computer Use Action Picker is an agent skill from mrmps/classifier-dev. Pick a browser or desktop agent's next action by choosing among the actions actually on screen instead of inventing one.

When should I use Computer Use Action Picker?

Computer Use Action Picker fits situations like: debugging computer use; browser automation; on it clicked the wrong thing and how do I stop it looping.

How do I install Computer Use Action Picker in Claude Code?

Run `npx skills add mrmps/classifier-dev --skill computer-use-action-picker -a claude-code`. Or copy the skill folder (skills/computer-use-action-picker in mrmps/classifier-dev) into .claude/skills/computer-use-action-picker in your project. Claude Code loads it when a task matches its description.

How do I install Computer Use Action Picker in Codex?

Run `npx skills add mrmps/classifier-dev --skill computer-use-action-picker -a codex`. Or copy the skill folder (skills/computer-use-action-picker in mrmps/classifier-dev) into .agents/skills/computer-use-action-picker in your project. Codex loads it when a task matches its description.

Can I use Computer Use Action Picker in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add mrmps/classifier-dev --skill computer-use-action-picker -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/computer-use-action-picker, .gemini/skills/computer-use-action-picker, .github/skills/computer-use-action-picker and .opencode/skills/computer-use-action-picker in your project.

What does Computer Use Action Picker need to run?

Going by SKILL.md and its folder, Computer Use Action Picker needs the command-line tools its instructions call (curl, jq, go and python3). Our summary lists: Python 3.

Does Computer Use Action Picker access the network?

SKILL.md names 1 domain. In commands or code: classifier.dev; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is Computer Use Action Picker safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Computer Use Action Picker use?

Computer Use Action Picker is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Computer Use Action Picker use?

About 1.5k tokens (SKILL.md is roughly 6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Computer Use Action Picker?

Skills that share tags, products or a category with Computer Use Action Picker: Computer Use (bam-bam-2/solo-skills, 367 stars), Computer Use (kunpengtalk/PeakCode, 166 stars), Computer Use (autonomous-ai/Physical-AI-Operating-System, 381 stars) and Agentic Browser Testing (petrkindlmann/qa-skills, 163 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Computer Use Action Picker?

mrmps (a GitHub user) maintains it in mrmps/classifier-dev, which has 424 GitHub stars. The repository holds 21 skills in this directory. The repository was last updated on October 6, 2026.

Source: mrmps/classifier-dev on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.