Agent skill

Screen Annotation

by vellum-ai in vellum-ai/vellum-assistant

Point at things on the screen the user is showing you. An agent skill from vellum-ai/vellum-assistant.

MITAuto-check passed

Install Screen Annotation

skills CLI
$ npx skills add vellum-ai/vellum-assistant --skill screen-annotation -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install vellum-ai/vellum-assistant screen-annotation --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/vellum-ai/vellum-assistant.git skills-src && mkdir -p .claude/skills && cp -r skills-src/assistant/src/config/bundled-skills/screen-annotation .claude/skills/screen-annotation && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
screen-annotation
GitHub stars
1.4k
Token cost
~2k tokens
SKILL.md length
1,366 words
Files
5
Skills in repo
108
Repo updated
First seen
Licence
MIT

At a glance

Point at things on the screen the user is showing you. An agent skill from vellum-ai/vellum-assistant.

  • SKILL.md covers Requires a screen share, Say what to point at, Coordinates, for a ring and How to point, plus 2 more sections
  • Runs TypeScript scripts from its folder

What it does

Screen Annotation is an agent skill from vellum-ai/vellum-assistant. Point at things on the screen the user is showing you

Its SKILL.md is about 2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files (for example `TOOLS.json`, `tools/screen-annotation.test.ts` and `tools/screen-clear-marks.ts`). Compatibility notes: Designed for Vellum personal assistants

The repository describes itself as: An AI Assistant that’s easy to setup, does your work 24/7, knows your preferences and gets better over time. The licence is MIT.

Example prompts

  • “/screen-annotation”

Requirements

  • Node.js
  • Compatibility (from SKILL.md): Designed for Vellum personal assistants

What it can do on your machine

Read from SKILL.md and the folder at commit 844117a. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (TypeScript), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Designed for Vellum personal assistants

    From compatibility in the SKILL.md frontmatter.

Context cost

Screen Annotation loads about 2k tokens when it runs. Until then it costs about 18 tokens; SKILL.md has 1,366 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~18
When it runs · the whole SKILL.md, loaded when a task matches
~2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from vellum-ai/vellum-assistant at commit 844117a, republished under its MIT licence (© vellum-ai). 1,366 words, ~2,048 tokens.

Download SKILL.mdSave it as .claude/skills/screen-annotation/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
screen-annotation
description
Point at things on the screen the user is showing you
compatibility
Designed for Vellum personal assistants
metadata.emoji
🫵

Drawing on the screen the user is showing you, so they can go and do the thing themselves.

This is the opposite errand from computer use. Nothing here clicks, types or drives anything: the marks are a way of pointing while you talk, for someone who wants to learn where a control is rather than have it operated for them. A mark is drawn clear of what it indicates and never takes the mouse, so what you point at stays visible and clickable the whole time.

Requires a screen share

Marks are drawn on the frame around the surface the user is sharing with the call. With nothing shared there is nowhere to draw, and screen_point_at fails saying so. Ask them to share their screen from the call, then point.

Say what to point at

Prefer naming the thing for an arrow. {"target": "color balance", "caption": "Click this"}. The name is looked up on the surface itself, which knows where its controls actually are, and an arrow is drawn at it.

The label, not a description of it. What is matched is the control's own name. Casing, spacing and punctuation are forgiven, so Color Balance finds color balance; nothing beyond that is, so "the stabilization button" finds nothing, because no control is called that. Give the label on its own: stabilization, Send, Search.

The arrow points at the middle of the control and stops just short, so what you are sending someone to stays visible the whole time.

You are answered with what was drawn and the accessibility name it resolved to. Some apps expose internal names such as ToolbarExportControl instead of the visible label. When a returned name identifies the intended control in the shared image, use that exact name as target; use the visible label in the caption and speech. Do not invent an internal name or choose an unclear candidate.

A failed name lookup draws nothing. It means the accessibility name could not be resolved, not that the control is absent from the picture. Read the returned names and retry with an exact name when it identifies the intended control. Do not cycle through invented internal names.

Use the picture when names cannot identify the control. If the intended control is clearly visible in a fresh image of the shared surface, retry in the same turn with tight x/y/width/height bounds measured from that image. This also applies when accessibility information is missing or a name matches multiple controls. A failed name lookup does not require the user to ask again before you can draw a ring.

If you cannot confidently locate the intended control, get a fresh view or ask the user to clarify. Describe its location verbally only when you can identify it. Do not invent a location or claim a mark was drawn after a failed tool call.

Coordinates, for a ring

Use bounds for a region, a visual fallback after a failed name lookup, or an explicit request to circle a clearly visible control. When the user asks for a circle or ring, use bounds directly; a named lookup draws an arrow. The same requirement to identify the target in a fresh image applies in every case.

Fractions of the shared surface, 0 to 1, measured against the picture of that surface you were last shown. x and y are the top-left corner, width and height the size. Give the bounds of the thing itself: the ring is drawn around them, so a box tight on a button reads as a ring around that button, and a box drawn where you think the ring should go puts the ring outside that instead.

Bounds are measured off a picture that has been scaled on its way to you. If the user has scrolled, moved a window, or changed the layout since that frame, get a fresh view before measuring. Say what you are pointing at as well as drawing it, so the user can check the mark against the intended control.

Moving the share is the one kind of drift that is caught for you. A mark measured against the surface before the move is refused rather than drawn, because those fractions land somewhere arbitrary on the surface that replaced it. Wait for a frame of the new one and point again.

How to point

One thing at a time. A mark is where to look next. A screen with four marks on it is not four times as helpful; it is a diagram, and nobody knows which one to start with. Point at the current step, talk, then point at the next one.

Captions are imperatives, not explanations. "Click Share", "Type the name here", "This is the tempo". Whatever else needs saying, say out loud: the caption is drawn over the user's own work in a window they cannot scroll or dismiss, and it is capped at 80 characters for that reason.

Say it as well as draw it. The marks are a gesture that accompanies speech, the way a person points while explaining. A mark with no words is a riddle.

Take them down when they stop being true. Call screen_clear_marks when the step is done, when the user has moved on, or when the conversation has left the screen behind. Marks come down on their own if the share ends or moves, but a mark left standing over a finished step is one the user has to work out is stale.

Show full SKILL.md (457 more words)Show less

Walking someone through several steps

Sometimes the answer is one pointer. Sometimes it is a route: four places to click, in order, before the thing they asked about happens. Decide which it is before you draw anything, because the two are paced differently.

Say the route before you start it. "There are three steps. First the Share menu, then the format, then Export." If they asked how and you inferred they want to be walked through it rather than told, this is where they wave it off and just want the answer. Keep it to the count and the landmarks; the detail belongs to each step as you reach it.

One step, then stop. Point at it, say what to do, and then wait. The temptation is to narrate the next step while the mark for this one is still up, and that leaves them doing step one with instructions for step two in their ear. Silence is the cue that it is their turn.

Advance on evidence, not on time. Move to the next step when a fresh picture shows this one done, or when they tell you it is. Do not move on because a plausible amount of time has passed.

The most direct telling is automatic. When the user clicks the control an arrow is pointing at, a message arrives as their turn saying they clicked it, by name, and that mark comes down on its own. Treat it as the step done: say what comes next and point at it. Nothing arrives for a click anywhere else, and nothing arrives for a ring, so for those the picture and their words are still the evidence. If the next picture shows the step not done, or done to the wrong thing, point at the same place again with a shorter caption and say what you saw. Pointing at the next step while the previous one is still open is how someone ends up two steps behind a mark.

Going back is just pointing again. There is no undo. If they went past something, or want to see step two again, point at step two. Say which step it is, so the words and the mark agree about where you both are.

Close it out. When the last step is done, clear the marks and say so, in a word. A mark left on the final button is one the user has to work out is stale, and a walkthrough that ends without an ending leaves them waiting for step five of four.

Shapes

A named control gives you an arrow. Bounds give you a ring for an extent, a control located visually after a failed name lookup, or the circle the user explicitly requested.

© vellum-ai, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files in assistant/src/config/bundled-skills/screen-annotation of vellum-ai/vellum-assistant.

  • SKILL.md
  • TOOLS.json
  • tools/screen-annotation.test.ts
  • tools/screen-clear-marks.ts
  • tools/screen-point-at.ts

Open the folder on GitHubat commit 844117a

Compare with similar skills

Screen Annotation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Screen Annotation compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Screen Annotation this skillvellum-ai/vellum-assistant1.4k—~2kAutomated safety check: PassMIT
Wireframe Annotatednexu-io/open-design100k—~960Automated safety check: PassApache-2.0
Screen Recordinggithub/awesome-copilot40k—~2kAutomated safety check: PassMIT
Arize Annotationgithub/awesome-copilot40k1 repos~2.7kAutomated safety check: NotesMIT
Bio Crispr Screens Screen QcFreedomIntelligence/OpenClaw-Medical-Skills3.1k—~2.1kAutomated safety check: PassNone
Image Annotationsgithub/awesome-copilot40k—~6kAutomated safety check: PassMIT

Similar skills

  • Wireframe Annotated

    nexu-io/open-design

    An annotated / redline lo-fi wireframe — a desktop landing/marketing page drawn as flat greyboxes inside a browser chrome frame, overlaid with numbered annotation pins (①②③④⑤) in a single accent…

    100k GitHub stars~960 tokensUpdated today
    Frontend & DesignAuto-check passed
  • Screen Recording

    github/awesome-copilot

    Official

    Create annotated animated GIF demos and screen recordings for pull requests and documentation.

    40k GitHub stars~2k tokensUpdated today
    DevelopmentAuto-check passed
  • Arize Annotation

    github/awesome-copilot

    Official

    Creates and manages annotation configs (categorical, continuous, freeform label schemas) and annotation queues (human review workflows) on Arize.

    40k GitHub starsUsed in 1 repo~2.7k tokens
    Auto-check: notes
  • Bio Crispr Screens Screen Qc

    FreedomIntelligence/OpenClaw-Medical-Skills

    Quality control for pooled CRISPR screens. An agent skill from FreedomIntelligence/OpenClaw-Medical-Skills.

    3.1k GitHub stars~2.1k tokensUpdated 2 mo ago
    Research & ScienceAuto-check passed
  • Image Annotations

    github/awesome-copilot

    Official

    Annotate screenshots, diagrams, and images with callout rectangles, arrows, labels, and color-coded highlights using PIL.

    40k GitHub stars~6k tokensUpdated today
    DevelopmentAuto-check passed
  • macOS Screen Recorder

    sickn33/agentic-awesome-skills

    macOS screen recorder that captures the main display PLUS system audio via ScreenCaptureKit — no BlackHole/loopback driver, no sudo, just the standard Screen Recording permission.

    47k GitHub starsUsed in 1 repo~727 tokens
    Auto-check passed

More from vellum-ai/vellum-assistant

All 108 skills in this repo
  • Vellum GitHub App Setup

    vellum-ai/vellum-assistant

    Create and configure a GitHub App so the assistant can push commits, open PRs, and comment under its own bot identity.

    1.4k GitHub stars~3.1k tokensUpdated today
    Auto-check passed
  • Discord App Setup

    vellum-ai/vellum-assistant

    Connect a Discord bot to the assistant via the Discord Gateway with guided application creation and intent configuration

    1.4k GitHub stars~4.2k tokensUpdated today
    Auto-check passed
  • Sentry App Setup

    vellum-ai/vellum-assistant

    Create and configure a Sentry internal integration so the assistant can manage issues, alerts, and releases under its own identity

    1.4k GitHub stars~1.3k tokensUpdated today
    Auto-check passed
  • Memory Corpus Ingest

    vellum-ai/vellum-assistant

    Ingest a large dataset into memory as a skimmed map. An agent skill from vellum-ai/vellum-assistant.

    1.4k GitHub stars~3k tokensUpdated today
    Auto-check: notes
  • Plugin Builder

    vellum-ai/vellum-assistant

    A skill your agent uses when the user wants to build, scaffold, ship, or edit a Vellum plugin that bundles multiple surfaces (hooks, tools, skills, and more) into one installable package.

    1.4k GitHub stars~3.1k tokensUpdated today
    Auto-check passed
  • Slack App Setup

    vellum-ai/vellum-assistant

    Connect a Slack app to the Vellum Assistant via Socket Mode.

    1.4k GitHub stars~2.5k tokensUpdated today
    Auto-check: warnings

Questions about Screen Annotation

What does Screen Annotation do?

Point at things on the screen the user is showing you. An agent skill from vellum-ai/vellum-assistant. Screen Annotation is an agent skill from vellum-ai/vellum-assistant.

How do I install Screen Annotation in Claude Code?

Run `npx skills add vellum-ai/vellum-assistant --skill screen-annotation -a claude-code`. Or copy the skill folder (assistant/src/config/bundled-skills/screen-annotation in vellum-ai/vellum-assistant) into .claude/skills/screen-annotation in your project. Claude Code loads it when a task matches its description.

How do I install Screen Annotation in Codex?

Run `npx skills add vellum-ai/vellum-assistant --skill screen-annotation -a codex`. Or copy the skill folder (assistant/src/config/bundled-skills/screen-annotation in vellum-ai/vellum-assistant) into .agents/skills/screen-annotation in your project. Codex loads it when a task matches its description.

Can I use Screen Annotation in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add vellum-ai/vellum-assistant --skill screen-annotation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/screen-annotation, .gemini/skills/screen-annotation, .github/skills/screen-annotation and .opencode/skills/screen-annotation in your project.

What does Screen Annotation need to run?

Going by SKILL.md and its folder, Screen Annotation needs TypeScript for the scripts in its folder. Our summary lists: Node.js. Compatibility (from SKILL.md): Designed for Vellum personal assistants.

Does Screen Annotation access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Screen Annotation safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Screen Annotation use?

Screen Annotation is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Screen Annotation use?

About 2k tokens (SKILL.md is roughly 8.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Screen Annotation?

Skills that share tags, products or a category with Screen Annotation: Wireframe Annotated (nexu-io/open-design, 100k stars), Screen Recording (github/awesome-copilot, 40k stars), Arize Annotation (github/awesome-copilot, 40k stars) and Bio Crispr Screens Screen Qc (FreedomIntelligence/OpenClaw-Medical-Skills, 3.1k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Screen Annotation?

vellum-ai (a GitHub organization) maintains it in vellum-ai/vellum-assistant, which has 1,400 GitHub stars. The repository holds 108 skills in this directory. The repository was last updated on October 9, 2026.

Source: vellum-ai/vellum-assistant on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.