Agent skill

Watch

by oxbshw in oxbshw/watch-skill

Watch any video (URL, stream, or local path) via Watch Skill.

MITAuto-check: notesAI & LLM Engineering

Install Watch

skills CLI
$ npx skills add oxbshw/watch-skill --skill watch -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install oxbshw/watch-skill watch --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/oxbshw/watch-skill.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/watch .claude/skills/watch && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
watch
GitHub stars
469
Token cost
~1.3k tokens
SKILL.md length
632 words
Files
1
Skills in repo
10
Repo updated
First seen
Licence
MIT

At a glance

Watch any video (URL, stream, or local path) via Watch Skill.

  • Works in 4 steps: Preflight (first invocation per session) → Watch → Read the frames → …
  • Tasks that involve Transcription
  • SKILL.md covers Step 0 — Preflight (first…, Step 1 — Watch, Step 2 — Read the frames and Step 3 — Answer, plus 3 more sections
  • Calls pip and uv

What it does

Watch is an agent skill from oxbshw/watch-skill. Watch any video (URL, stream, or local path) via Watch Skill. Downloads, extracts scene-aware deduped frames, OCRs them, transcribes (captions first, then local Whisper — offline by default), indexes everything, and hands the result to the agent. Follow-up questions are answered from the persistent index without re-processing.

Its SKILL.md is about 1.3k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering Transcription and Speech recognition and synthesis. It works with DeepSeek. The repository describes itself as: Give AI agents eyes, ears, and verifiable results. Watch Skill turns video, audio and screen activity into searchable, timestamped evidence and proves work with deterministic… The licence is MIT.

When your agent uses it

  • Tasks that involve Transcription
  • Tasks that involve Speech recognition and synthesis

Example prompts

  • “/watch”

Requirements

  • Python 3
  • Pre-approved tools (allowed-tools): Bash, Read, AskUserQuestion

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Preflight (first invocation per session)
  2. Watch
  3. Read the frames
  4. Answer

What it can do on your machine

Read from SKILL.md and the folder at commit f1317c8. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Bash
    • Read
    • AskUserQuestion

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pip
    • uv

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip and uv, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Watch loads about 1.3k tokens when it runs. Until then it costs about 84 tokens; SKILL.md has 632 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~84
When it runs · the whole SKILL.md, loaded when a task matches
~1.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteMentions a .env fileSKILL.md:123
    - API keys live in env vars / `.env`; they are never logged or echoed.
  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Bash, Read, AskUserQuestion

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from oxbshw/watch-skill at commit f1317c8, republished under its MIT licence (© oxbshw). 632 words, ~1,334 tokens.

Download SKILL.mdSave it as .claude/skills/watch/SKILL.md (or your agent's skills folder).
name
watch
description
Watch any video (URL, stream, or local path) via Watch Skill. Downloads, extracts scene-aware deduped frames, OCRs them, transcribes (captions first, then local Whisper — offline by default), indexes everything, and hands the result to the agent. Follow-up questions are answered from the persistent index without re-processing.
allowed-tools
Bash, Read, AskUserQuestion
version
1.4.3
argument-hint
<video-url-or-path> [question]
license
MIT
user-invocable
true

/watch (Watch Skill)

You don't have a video input; this skill gives you one. It is a thin wrapper around the watch-skill CLI — all logic lives in the engine, so this skill works identically on every harness (Claude Code, Codex, Cursor, ...).

This is a drop-in upgrade of the classic claude-video /watch skill: same invocation shape, plus a persistent index (ask answers follow-ups without re-processing), OCR on frames, scene-aware sampling with perceptual dedup, local Whisper (offline by default, no API key needed), and THE LOOP (capture -> critique -> fix -> re-capture) for iterating on your own output.

Step 0 — Preflight (first invocation per session)

bash
watch-skill doctor --json
  • Exit 0 → proceed silently. Do NOT announce that setup is fine.
  • Non-zero → the JSON lists each failing check with a fix. doctor auto-bootstraps ffmpeg and yt-dlp into a managed bin dir on Windows/macOS/ Linux; re-run once after it reports fixes. Only involve the user when a check still fails after remediation.
  • If watch-skill itself is not on PATH: pip install watch-skill (or uv tool install watch-skill), then re-run the doctor.

No API key is required for acquisition, transcription, OCR, indexing, or search: transcription falls back to local faster-whisper. Visual synthesis and verification can use the user's existing Anthropic, OpenAI, Gemini, or OpenRouter key, or an optional local Ollama model. The agent and provider are independent; see the configuring-vision skill. Cloud STT is opt-in (--cloud-stt) and only ever uploads extracted mono audio — the video file never leaves the machine.

Step 1 — Watch

Parse the user input into source + optional question, then:

bash
watch-skill watch "<source>" [--start T --end T] [--max-frames N] [--transcript-only]
  • Any yt-dlp-supported site (1800+), direct media URLs, HLS/DASH manifests (--duration 60 bounds live streams), and local files all work.
  • --start / --end (SS, MM:SS, HH:MM:SS) switch to dense focused sampling of that window — use for "what happens at 2:30?" questions and for any video over ~10 minutes when the user cares about one section.
  • --timestamps T1,T2,... pins frames at transcript-flagged moments ("look here", "as you can see") that visual selection may miss.
  • --transcript-only skips frames entirely (fastest; no video download when captions exist).
  • --max-frames N tightens the token budget (default: duration-tiered, hard cap 100, max 2 fps).

The report prints an Indexed: video_id ... line, frames with t=MM:SS timestamps, OCR text, and the transcript.

Step 2 — Read the frames

Read every frame path the report lists, in a single message (parallel Read calls), so you see them together in chronological order.

Show full SKILL.md (242 more words)Show less

Step 3 — Answer

Answer from frames + OCR + transcript, citing timestamps. No question → summarize structure, key moments, notable visuals, spoken content.

Follow-ups — use the index, not re-processing

The video is already indexed. For any follow-up question in this or a LATER session:

bash
watch-skill ask <video_id> "<question>"      # self-healing answer + evidence
watch-skill search "<phrase>"                 # across every video ever watched

ask (v0.6) answers text-first with timestamped evidence, a confidence score, and a ~N tokens saved line. It escalates on its own when unsure (dense re-sampling, zoom-crop re-OCR) and states plainly when the video does not clearly show the answer — trust that refusal; do NOT invent an answer past it. Frame paths are listed only when the engine wants you to look yourself (or pass --frames); Read them then. Never re-run watch for a follow-up on an already-indexed video.

If the user corrects one of your video answers, report it so the system learns (locally):

bash
watch-skill lessons add <video_id> "<question>" "<your wrong answer>" "<the correction>"

THE LOOP — iterate on your own output

When the user asks you to fix UI/visual output and verify the fix:

bash
watch-skill loop start "<url-or-screen:-or-file>" "<pass criteria>" [--script '<json steps>']
# ... you apply the suggested fixes ...
watch-skill loop iterate <loop_id>

The critique returns structured issues with timestamps and suggested fixes. YOU change the code; the loop only observes. On pass it renders a before/after MP4+GIF proof. watch-skill capture "<target>" records without critiquing.

Security posture

  • The video file itself NEVER leaves the machine. Only extracted mono-16kHz audio may go to a cloud STT API, and only with explicit --cloud-stt.
  • No cookies, no logins — only public data is requested.
  • API keys live in env vars / .env; they are never logged or echoed.
  • Downloads are cached under ~/.watch-skill/cache (LRU, size-capped).

© oxbshw, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/watch of oxbshw/watch-skill.

Open the folder on GitHubat commit f1317c8

Compare with similar skills

Watch next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Watch compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Watch this skilloxbshw/watch-skill469—~1.3kAutomated safety check: NotesMIT
Yichen Asrmcncarl/yichen-skills4.4k—~780Automated safety check: PassCustom licence
Youtube FetcherJimmySadek/youtube-fetcher-to-markdown485—~3.1kAutomated safety check: PassMIT
Ax Audiodosco/aithy107—~2.5kAutomated safety check: PassApache-2.0
Stt Integrationrapidaai/voice-ai745—~778Automated safety check: PassCustom licence
Whisper Speech RecognitionOrchestra-Research/AI-Research-SKILLs13k7 repos~1.9kAutomated safety check: NotesMIT

Similar skills

  • Yichen Asr

    mcncarl/yichen-skills

    逸尘自用的统一音视频转写入口,在 StepFun Step ASR 与火山引擎豆包 ASR 之间按输出需求、安全边界和可用状态路由。用于本地音频或视频的纯文本转写、时间戳、SRT 字幕、口播粗剪,以及转写前体检;用户明确指定服务商时不得静默切换。Use when a local audio or video file needs transcription and the correct…

    4.4k GitHub stars~780 tokensUpdated 5 days ago
    AI & LLM EngineeringAuto-check passed
  • Youtube Fetcher

    JimmySadek/youtube-fetcher-to-markdown

    Retrieve transcripts from YouTube, Instagram, TikTok, X, Vimeo and other video sites, summarize or analyze what was said (and shown on screen), or save an Obsidian-ready Markdown knowledge-base note…

    485 GitHub stars~3.1k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • Ax Audio

    dosco/aithy

    This skill helps an LLM generate correct audio code with @ax-llm/ax.

    107 GitHub stars~2.5k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Stt Integration

    rapidaai/voice-ai

    Add or modify speech-to-text providers in assistant-api with transport-aware ingestion (WS/SDK/HTTP), transcript packet correctness, and UI/provider wiring.

    745 GitHub stars~778 tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • Whisper Speech Recognition

    Orchestra-Research/AI-Research-SKILLs

    Transcribes audio with OpenAI's Whisper: 99 languages, translation to English, language detection, six model sizes and word-level timestamps, from Python or the CLI.

    13k GitHub starsUsed in 7 repos~1.9k tokens
    AI & LLM EngineeringAuto-check: notes
  • AI SDK Development

    trypostit/trypost

    TRIGGER when working with ai-sdk which is Laravel official first-party AI SDK.

    685 GitHub starsUsed in 1 repo~3.5k tokens
    AI & LLM EngineeringAuto-check passed

More from oxbshw/watch-skill

All 10 skills in this repo
  • Asking With Evidence

    oxbshw/watch-skill

    The user asks a question about a video that was already watched or indexed — "what did they say about X", "what error code appears", "what happens at 2:30", "does the video show Y".

    469 GitHub stars~550 tokensUpdated 24 days ago
    Auto-check: notes
  • Configuring Vision

    oxbshw/watch-skill

    The user wants to connect an LLM or vision provider, already has an API key, asks "can I use OpenAI/Anthropic/Gemini/OpenRouter", wants local Ollama, or needs different cheap and strong models.

    469 GitHub stars~509 tokensUpdated 24 days ago
    Auto-check: notes
  • The Loop

    oxbshw/watch-skill

    The user built or changed something visual — a UI, an animation, a game, a generated video — and wants it verified, or asks "why does my UI look wrong", "check that the fix actually worked", "does…

    469 GitHub stars~622 tokensUpdated 24 days ago
    Auto-check: notes
  • Video Memory

    oxbshw/watch-skill

    The user asks about videos watched in the past or across sessions — "have we watched anything about X", "which video showed that error", "what did that meeting decide", "search my videos", or a…

    469 GitHub stars~569 tokensUpdated 24 days ago
    Auto-check: notes
  • Watching Videos

    oxbshw/watch-skill

    The user shared a video URL, a YouTube/TikTok/stream link, a local video file, a screen recording, a meeting recording, or a playlist/folder of videos — "watch this", "summarize this video", "what's…

    469 GitHub stars~599 tokensUpdated 24 days ago
    Auto-check: notes
  • Extracting Structure

    oxbshw/watch-skill

    The user wants structure pulled out of a watched video — "make chapters for this video", "where does the bug appear in this recording", "turn this screen recording into a bug report", "how strong is…

    469 GitHub stars~490 tokensUpdated 24 days ago
    Auto-check: notes

Works with

Questions about Watch

What does Watch do?

Watch any video (URL, stream, or local path) via Watch Skill. Watch is an agent skill from oxbshw/watch-skill. Watch any video (URL, stream, or local path) via Watch Skill.

When should I use Watch?

Watch fits situations like: tasks that involve Transcription; tasks that involve Speech recognition and synthesis.

How do I install Watch in Claude Code?

Run `npx skills add oxbshw/watch-skill --skill watch -a claude-code`. Or copy the skill folder (skills/watch in oxbshw/watch-skill) into .claude/skills/watch in your project. Claude Code loads it when a task matches its description.

How do I install Watch in Codex?

Run `npx skills add oxbshw/watch-skill --skill watch -a codex`. Or copy the skill folder (skills/watch in oxbshw/watch-skill) into .agents/skills/watch in your project. Codex loads it when a task matches its description.

Can I use Watch in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add oxbshw/watch-skill --skill watch -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/watch, .gemini/skills/watch, .github/skills/watch and .opencode/skills/watch in your project.

What does Watch need to run?

Going by SKILL.md and its folder, Watch needs the command-line tools its instructions call (pip and uv). Our summary lists: Python 3. Its frontmatter pre-approves these tools: Bash, Read, AskUserQuestion.

Does Watch access the network?

SKILL.md contains no URLs. Its commands use pip and uv, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Watch safe to install?

Our automated static check of SKILL.md found notes only (mentions a .env file; pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Watch use?

Watch is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Watch use?

About 1.3k tokens (SKILL.md is roughly 5.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Watch?

Skills that share tags, products or a category with Watch: Yichen Asr (mcncarl/yichen-skills, 4.4k stars), Youtube Fetcher (JimmySadek/youtube-fetcher-to-markdown, 485 stars), Ax Audio (dosco/aithy, 107 stars) and Stt Integration (rapidaai/voice-ai, 745 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Watch?

oxbshw (a GitHub user) maintains it in oxbshw/watch-skill, which has 469 GitHub stars. The repository holds 10 skills in this directory. The repository was last updated on September 14, 2026.

Source: oxbshw/watch-skill on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.