Agent skill

Watch

by mathiaschu in mathiaschu/watch

Watch a video from YouTube, Instagram, X/Twitter, Vimeo, TikTok or any of ~1800 yt-dlp sites (or a local path).

MITAuto-check: warningsMedia & Creative

Install Watch

The automated check flagged lines worth reading first. See the safety section below.

skills CLI
$ npx skills add mathiaschu/watch --skill watch -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install mathiaschu/watch watch --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
watch
GitHub stars
141
Token cost
~4k tokens
SKILL.md length
2,174 words
Files
19 (incl. scripts)
Skills in repo
1
Repo updated
First seen
Licence
MIT

At a glance

Watch a video from YouTube, Instagram, X/Twitter, Vimeo, TikTok or any of ~1800 yt-dlp sites (or a local path).

  • Works in 4 steps: Install a cookies-export extension —… → Open and log into the site (e.g.… → Click the extension → Export → save the… → …
  • Tasks that involve Speech recognition and synthesis
  • SKILL.md covers Step 0 — Setup preflight (runs…, When to use, Recommended limits and How to invoke, plus 4 more sections
  • Runs Shell and Python scripts from its folder; calls python3, pip3 and whisper; reaches youtu.be and instagram.com

What it does

Watch is an agent skill from mathiaschu/watch. Watch a video from YouTube, Instagram, X/Twitter, Vimeo, TikTok or any of ~1800 yt-dlp sites (or a local path). Downloads with yt-dlp, extracts auto-scaled frames with ffmpeg, pulls the transcript from captions (or local mlx-whisper fallback, no API key), and hands the result to Claude so it can answer questions about what's in the video.

Its SKILL.md is about 4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 25 other files, including scripts (for example `.claude-plugin/marketplace.json`, `.claude-plugin/plugin.json` and `.github/workflows/release.yml`).

It sits in Media & Creative, covering Speech recognition and synthesis, Transcription and Video production. It works with FFmpeg, Instagram, TikTok and X (Twitter). The repository describes itself as: Give Claude a video input. /watch downloads from YouTube/Instagram/X/Vimeo/any yt-dlp site, extracts frames, and transcribes locally with mlx-whisper — no API key. Fork of… The licence is MIT.

When your agent uses it

  • Tasks that involve Speech recognition and synthesis
  • Tasks that involve Transcription
  • Tasks that involve Video production

Example prompts

  • “/watch”

Requirements

  • Python 3
  • A Bash shell
  • Pre-approved tools (allowed-tools): Bash, Read

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Install a cookies-export extension — "Get cookies.txt LOCALLY" (open-source, exports Netscape format) for Chrome/Edge, or "cookies.txt"…
  2. Open and log into the site (e.g. instagram.com).
  3. Click the extension → Export → save the file (e.g. ~/Downloads/cookies.txt).
  4. Re-run: python3 "${CLAUDE_SKILL_DIR}/scripts/watch.py" "" --cookies ~/Downloads/cookies.txt

What it can do on your machine

Read from SKILL.md and the folder at commit 14c780e. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Bash
    • Read

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 2 files in scripts/ (Shell and Python, from the files we listed), which the agent can run.

    Shell commands in SKILL.md call:

    • python3
    • pip3
    • whisper
    • ffmpeg

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • youtu.be
    • instagram.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Watch loads about 4k tokens when it runs. Until then it costs about 87 tokens; SKILL.md has 2,174 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~87
When it runs · the whole SKILL.md, loaded when a task matches
~4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: warnings

The automated check found patterns that need a careful read before installing.

  • NoteMentions a .env fileSKILL.md:43
    esent. **No API key, no config file, no `.env` — transcription runs entirely on-device.**
  • WarningMentions a credentials file (SSH keys, cloud or package-manager tokens)SKILL.md:103
    - **Chrome on macOS** locks its cookie DB while open *and* its cookies are encrypted. Two things may happen: (1) extract
  • WarningMentions a credentials file (SSH keys, cloud or package-manager tokens)SKILL.md:109
    en-source, exports Netscape format) for Chrome/Edge, or **"cookies.txt"** for Firefox.
  • NoteMentions a .env fileSKILL.md:195
    e, or require any API key — there is no `.env`, no config file, no secrets
  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Bash, Read

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from mathiaschu/watch at commit 14c780e, republished under its MIT licence (© mathiaschu). 2,174 words, ~4,047 tokens.

Download SKILL.mdSave it as .claude/skills/watch/SKILL.md (or your agent's skills folder). This skill also uses 18 other files; get the full folder from GitHub.
name
watch
description
Watch a video from YouTube, Instagram, X/Twitter, Vimeo, TikTok or any of ~1800 yt-dlp sites (or a local path). Downloads with yt-dlp, extracts auto-scaled frames with ffmpeg, pulls the transcript from captions (or local mlx-whisper fallback, no API key), and hands the result to Claude so it can answer questions about what's in the video.
allowed-tools
Bash, Read
argument-hint
<video-url-or-path> [question]
homepage
https://github.com/mathiaschu/claude-video
repository
https://github.com/mathiaschu/claude-video
author
Mathias Schusterman (fork of bradautomates/claude-video)
license
MIT
user-invocable
true

/watch — Claude watches a video

You don't have a video input; this skill gives you one. A Python script downloads the video, extracts frames as JPEGs, gets a timestamped transcript (native captions first, then local mlx-whisper as fallback — runs on-device, no API and no key), and prints frame paths. You then Read each frame path to see the images and combine them with the transcript to answer the user.

Step 0 — Setup preflight (runs every /watch invocation, silent on success)

Python interpreter: every python3 ... command in this skill is for macOS/Linux. On Windows, substitute python — the python3 command on Windows is the Microsoft Store stub and will not run the script.

Before every /watch run, verify that dependencies are in place:

bash
python3 "${CLAUDE_SKILL_DIR}/scripts/setup.py" --check

This is a <100ms lookup. On exit 0, the script emits nothing — proceed to Step 1 without comment. Do NOT announce "setup is complete" to the user — they don't need a status message on every turn. The only acceptable user-visible output from Step 0 is when remediation is required.

On non-zero exit, follow the table:

ExitMeaningAction
2Missing binaries (ffmpeg / ffprobe / yt-dlp)Run installer
3No local whisper engine (mlx-whisper / openai-whisper)Run installer, then tell user the pip3 command it prints
4Both missingRun installer

The installer is idempotent — safe to re-run:

bash
python3 "${CLAUDE_SKILL_DIR}/scripts/setup.py"

On macOS with Homebrew, it auto-installs ffmpeg and yt-dlp. On Linux/Windows, it prints the exact install commands for the user to run. For transcription it checks for a local whisper engine (mlx-whisper preferred on Apple Silicon, openai-whisper as a CPU fallback) and prints the pip3 install command if neither is present. No API key, no config file, no .env — transcription runs entirely on-device.

If no whisper engine is installed: run the installer and relay the exact pip3 install … command it prints (mlx-whisper on Apple Silicon, openai-whisper on Windows/Linux/Intel Macs — do not assume mlx, it only installs on Apple Silicon). If they don't want to install it, proceed with --no-whisper and tell them videos without native captions will come back frames-only.

Structured mode (optional): python3 "${CLAUDE_SKILL_DIR}/scripts/setup.py" --json emits {status, missing_binaries, whisper_backend, has_whisper, platform} where status is one of ready | needs_install | needs_whisper | needs_install_and_whisper.

Within a single session, you can skip Step 0 on follow-up /watch calls — once --check returned 0, nothing about the environment changes between turns.

When to use

  • User pastes a video URL (YouTube, Vimeo, X, TikTok, Twitch clip, most yt-dlp-supported sites) and asks about it.
  • User points at a local video file (.mp4, .mov, .mkv, .webm, etc.) and asks about it.
  • User types /watch <url-or-path> [question].
  • Best accuracy: videos under 10 minutes. Frame coverage scales inversely with duration.
  • Hard caps: 100 frames total and 2 fps. Token cost grows with frame count, so the script targets a frame budget by duration (and never exceeds 2 fps even when the budget would imply more):
    • ≤30s → ~1-2 fps (up to 30 frames)
    • 30s-1min → ~40 frames
    • 1-3min → ~60 frames
    • 3-10min → ~80 frames
    • >10min → 100 frames, sparsely spaced (warning printed)
  • If the user hands you a long video, consider asking whether they want a specific section before burning tokens on a sparse scan.

How to invoke

Step 1 — parse the user input. Separate the video source (URL or path) from any question the user asked. Example: /watch https://youtu.be/abc what language is this in? → source = https://youtu.be/abc, question = what language is this in?.

Step 2 — run the watch script. Pass the source verbatim. Do not shell-escape it yourself beyond normal quoting:

bash
python3 "${CLAUDE_SKILL_DIR}/scripts/watch.py" "<source>"

Optional flags:

  • --start T / --end T — focus on a section. Accepts SS, MM:SS, or HH:MM:SS. When either is set, fps auto-scales denser (see "Focusing on a section" below).
  • --max-frames N — lower the cap for tighter token budget (e.g. --max-frames 40)
  • --resolution W — change frame width in px (default 512; bump to 1024 only if the user needs to read on-screen text)
  • --fps F — override auto-fps (clamped to 2 fps max)
  • --out-dir DIR — keep working files somewhere specific (default: an auto-generated tmp dir)
  • --cookies-from-browser B — read cookies from a local browser (chrome, firefox, safari, edge, brave, …) for login-gated sources
  • --cookies FILE — path to a Netscape-format cookies.txt (alternative to --cookies-from-browser)
  • --whisper mlx|openai-whisper — force a specific local Whisper engine (default: prefer mlx-whisper, fall back to openai-whisper)
  • --no-whisper — disable the local Whisper fallback entirely (frames-only if no captions)
Login-gated sources (Instagram, X, private/age-restricted videos)

Public videos (most of YouTube, Vimeo, TikTok, Loom, etc.) download with no auth. But some sources gate the download behind a login: Instagram, X/Twitter, age-restricted or private/unlisted YouTube, members-only content. Those need the user's own cookies.

Do NOT pass cookies pre-emptively. Always try the plain download first. Only reach for cookies when it fails with a login / private / 403 / "login required" / "rate-limit" error. The user never types the flag themselves — you add it and re-run. When that happens, walk the user through it (these are sub-steps of the main Step 2, not the main flow):

(a) Ask which browser they're logged into. "To grab this Instagram video I need to borrow the cookies from a browser where you're logged into Instagram. Which one are you logged in on — Chrome, Safari, Firefox, Edge, or Brave?" Supported values: chrome, firefox, safari, edge, brave, chromium, opera, vivaldi.

(b) Re-run with that browser (on Windows use python, not python3 — see Step 0):

bash
python3 "${CLAUDE_SKILL_DIR}/scripts/watch.py" "https://www.instagram.com/reel/XXXX/" --cookies-from-browser chrome

(c) Handle the common per-browser snags (tell the user the specific fix, don't just retry):

  • Chrome on macOS locks its cookie DB while open and its cookies are encrypted. Two things may happen: (1) extraction fails with "could not copy/open the cookie database" → tell the user to fully quit Chrome (Cmd-Q, not just close the window) and retry; (2) a macOS Keychain prompt pops up ("… wants to use your confidential information stored in Chrome Safe Storage") → tell the user to click Always Allow. If Chrome keeps fighting it, suggest they switch to Safari or Firefox.
  • Chrome on Windows also locks the DB while running → tell the user to fully close it (check the system tray) and retry.
  • Safari on macOS needs the app running this (the terminal / Claude Code) to have Full Disk Access (System Settings → Privacy & Security → Full Disk Access). If Safari extraction fails, point them there, or fall back to another browser.
  • Firefox usually works without closing it — good fallback on any OS when Chrome is stubborn.

(d) Manual fallback if browser extraction just won't cooperate (most reliable, works on macOS / Windows / Linux): guide the user to export a cookies.txt and pass it with --cookies:

  1. Install a cookies-export extension — "Get cookies.txt LOCALLY" (open-source, exports Netscape format) for Chrome/Edge, or "cookies.txt" for Firefox.
  2. Open and log into the site (e.g. instagram.com).
  3. Click the extension → Export → save the file (e.g. ~/Downloads/cookies.txt).
  4. Re-run: python3 "${CLAUDE_SKILL_DIR}/scripts/watch.py" "<url>" --cookies ~/Downloads/cookies.txt

Privacy note to reassure the user: cookies are read live from their own machine and piped straight into the yt-dlp subprocess. The skill never copies, stores, logs, or transmits them anywhere. The cookies.txt file (if they used the manual fallback) stays on their disk — they can delete it after.

Show full SKILL.md (1,012 more words)Show less
Focusing on a section (higher frame rate)

When the user asks about a specific moment — "what happens at the 2 minute mark?", "zoom into 0:45 to 1:00", "the first 10 seconds" — pass --start and/or --end. The script switches to focused-mode budgets, which are denser than full-video budgets (still capped at 2 fps):

  • ≤5s → 2 fps (up to 10 frames)
  • 5-15s → 2 fps (up to 30 frames)
  • 15-30s → ~2 fps (up to 60 frames)
  • 30-60s → ~1.3 fps (up to 80 frames)
  • 60-180s → ~0.6 fps (100 frames, capped)

Focused mode is the right call for:

  • Any moment/range the user names explicitly ("around 2:30", "the intro", "the last 30 seconds").
  • Any video longer than ~10 minutes where the user's question is about a specific part — running focused on the relevant section is far more useful than a sparse scan of the whole thing.
  • Re-runs after a full scan didn't have enough detail in some region.

Transcript is auto-filtered to the same range. Frame timestamps are absolute (real video timeline, not offset-from-start).

Examples:

bash
# Last 10 seconds of a 1 minute video
python3 "${CLAUDE_SKILL_DIR}/scripts/watch.py" video.mp4 --start 50 --end 60

# Zoom into 2:15 → 2:45 at 3 fps (90 frames)
python3 "${CLAUDE_SKILL_DIR}/scripts/watch.py" "$URL" --start 2:15 --end 2:45 --fps 3

# From 1h12m to the end of the video
python3 "${CLAUDE_SKILL_DIR}/scripts/watch.py" "$URL" --start 1:12:00

Step 3 — Read every frame path the script lists. The Read tool renders JPEGs directly as images for you. Read all frames in a single message (parallel tool calls) so you see them together. The frames are in chronological order with a t=MM:SS timestamp so you can align them to the transcript.

Step 4 — answer the user. You now have two streams of evidence:

  • Frames — what's on screen at each timestamp
  • Transcript — what's said at each timestamp. The report's header shows the source (captions = yt-dlp pulled native subs; whisper (mlx) or whisper (openai-whisper) = transcribed locally on-device).

If the user asked a specific question, answer it directly citing timestamps. If they didn't ask anything, summarize what happens in the video — structure, key moments, notable visuals, spoken content.

Step 5 — clean up. The script prints a working directory at the end. If the user isn't going to ask follow-ups about this video, delete it with rm -rf <dir>. If they might, leave it in place.

Transcription

The script gets a timestamped transcript in one of two ways:

  1. Native captions (free, preferred). yt-dlp pulls manual or auto-generated subtitles from the source platform if available.
  2. Local Whisper fallback (on-device, no API, no key). If no captions came back (or the source is a local file), the script extracts audio (ffmpeg -vn -ac 1 -ar 16000 -b:a 64k, ~0.5 MB/min) and transcribes it locally:
    • mlx-whisper — mlx-community/whisper-large-v3-turbo. Preferred on Apple Silicon: fast, runs on the GPU/Neural Engine. Same engine ig-scraper uses. Install: pip3 install mlx-whisper.
    • openai-whisper — base model on CPU. Cross-platform fallback when mlx isn't available. Install: pip3 install openai-whisper.

The audio never leaves the machine. The script prefers mlx-whisper; override with --whisper openai-whisper. Language is auto-detected. Use --no-whisper to skip the fallback entirely.

Failure modes and handling

  • Setup preflight failed → run python3 "${CLAUDE_SKILL_DIR}/scripts/setup.py" (auto-installs ffmpeg/yt-dlp via brew on macOS; prints exact commands on Windows/Linux). If it reports no whisper engine, relay the exact pip3 install … command it printed (mlx-whisper only on Apple Silicon, openai-whisper elsewhere).
  • No transcript available → captions missing AND (no local whisper engine OR --no-whisper set OR transcription failed). Script prints a hint pointing to setup. Proceed frames-only and tell the user.
  • Long video warning printed → acknowledge it in your answer. Offer to re-run focused on a specific section via --start/--end rather than a sparse full-video scan.
  • Download fails (login/private/403) → the source needs auth (common on Instagram, X, age-restricted or private videos). Re-run with --cookies-from-browser <browser> using a browser the user is logged into (see "Login-gated sources" above). If it's region-locked or genuinely unavailable, tell the user plainly; do not keep retrying.
  • Whisper fails → the error is printed to stderr (likely: engine not installed, or a video with no audio track). The report will say "none available" for transcript. The first mlx run also downloads the model (~1.5 GB) once, then caches it.

Token efficiency

This skill burns tokens primarily on frames. Order of magnitude:

  • 80 frames at 512px wide is roughly 50-80k image tokens depending on aspect ratio.
  • The transcript is cheap (a few thousand tokens at most for a 10-minute video).
  • Bumping --resolution to 1024 roughly quadruples the image tokens per frame. Only do it when necessary.

If you already watched a video this session and the user asks a follow-up, do not re-run the script — you already have the frames and transcript in context. Just answer from what you have.

Security & Permissions

What this skill does:

  • Runs yt-dlp locally to download the video and pull native captions when the source supports them (public data; the request goes directly to whatever host the URL points at)
  • Runs ffmpeg / ffprobe locally to extract frames as JPEGs and, when Whisper is needed, a mono 16 kHz audio clip
  • Transcribes the audio clip locally on-device with mlx-whisper (or openai-whisper as a CPU fallback) — no network call, no API, no key
  • Writes the downloaded video, frames, audio, and an intermediate transcript to a working directory under the system temp dir (or --out-dir if specified) so Claude can Read them
  • On first mlx run, downloads the whisper model (~1.5 GB) from Hugging Face once and caches it under ~/.cache/huggingface

What this skill does NOT do:

  • Does not upload the video OR the audio to any API — transcription is fully local. The only outbound traffic is yt-dlp fetching the video/captions from the source URL (and the one-time model download)
  • Does not log into or post to any account. It only reads browser cookies when the user explicitly passes --cookies-from-browser / --cookies for a login-gated source, and only to authenticate the yt-dlp download. Those cookies are read live and never copied, stored, logged, or transmitted by the skill
  • Does not use, store, or require any API key — there is no .env, no config file, no secrets
  • Does not persist anything outside the working directory and the Hugging Face model cache — clean up the working directory when you're done (Step 5)

Bundled scripts: scripts/watch.py (entry point), scripts/download.py (yt-dlp wrapper), scripts/frames.py (ffmpeg frame extraction), scripts/transcribe.py (caption parsing), scripts/whisper.py (local mlx/openai-whisper transcription), scripts/setup.py (preflight + installer)

Review scripts before first use to verify behavior.

© mathiaschu, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 18 other files (scripts) in the repository root of mathiaschu/watch.

  • SKILL.md
  • .claude-plugin/marketplace.json
  • .claude-plugin/plugin.json
  • .gitattributes
  • .github/workflows/release.yml
  • .gitignore
  • CHANGELOG.md
  • LICENSE
  • README.md
  • commands/watch.md
  • hooks/hooks.json
  • hooks/scripts/check-setup.sh
  • scripts/build-skill.sh
  • scripts/download.py
  • … and 5 more

Open the folder on GitHubat commit 14c780e

Compare with similar skills

Watch next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Watch compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Watch this skillmathiaschu/watch141—~4kAutomated safety check: WarnMIT
Ffmpeg Skillkajisho5/ffmpeg-skill1.9k—~7.4kAutomated safety check: PassMIT
WatchTheCraigHewitt/skills157—~1.6kAutomated safety check: NotesMIT
AutoshortsUpload-Post/skill-autoshorts151—~5.3kAutomated safety check: NotesMIT
Douyin DownloaderOpenMinis/MinisSkills444—~1.4kAutomated safety check: PassMIT
Claude Real VideoHUANGCHIHHUNGLeo/claude-real-video2.2k—~639Automated safety check: PassMIT

Similar skills

  • Ffmpeg Skill

    kajisho5/ffmpeg-skill

    Edit video and audio with local FFmpeg from natural-language requests: cut, trim, join, resize/reframe (9:16, 1:1), speed change, captions and subtitles (SRT/ASS, animated, karaoke), logos and text…

    1.9k GitHub stars~7.4k tokensUpdated 3 days ago
    Media & CreativeAuto-check passed
  • Watch

    TheCraigHewitt/skills

    When the user wants to read, transcribe, summarize, or research a video — YouTube link, podcast clip, Loom, TikTok, X/Twitter video, local file, or any URL yt-dlp supports.

    157 GitHub stars~1.6k tokensUpdated 4 mo ago
    Media & CreativeAuto-check: notes
  • Autoshorts

    Upload-Post/skill-autoshorts

    Daily pipeline that picks one long video from a folder, transcribes it with Whisper, uses Gemini 3 Flash multimodal to find every viral short-form moment, cuts each candidate with FFmpeg, adds a…

    151 GitHub stars~5.3k tokensUpdated 5 mo ago
    Media & CreativeAuto-check: notes
  • Douyin Downloader

    OpenMinis/MinisSkills

    Download Douyin (TikTok) videos from share links. An agent skill from OpenMinis/MinisSkills.

    444 GitHub stars~1.4k tokensUpdated today
    Media & CreativeAuto-check passed
  • Claude Real Video

    HUANGCHIHHUNGLeo/claude-real-video

    Watch a video for the user. An agent skill from HUANGCHIHHUNGLeo/claude-real-video.

    2.2k GitHub stars~639 tokensUpdated 6 days ago
    Media & CreativeAuto-check passed
  • Shorts

    AgriciDaniel/claude-shorts

    Interactive longform-to-shortform video creator. An agent skill from AgriciDaniel/claude-shorts.

    218 GitHub stars~3.2k tokensUpdated 2 mo ago
    Media & CreativeAuto-check: notes

Questions about Watch

What does Watch do?

Watch a video from YouTube, Instagram, X/Twitter, Vimeo, TikTok or any of ~1800 yt-dlp sites (or a local path). Watch is an agent skill from mathiaschu/watch. Watch a video from YouTube, Instagram, X/Twitter, Vimeo, TikTok or any of ~1800 yt-dlp sites (or a local path).

When should I use Watch?

Watch fits situations like: tasks that involve Speech recognition and synthesis; tasks that involve Transcription; tasks that involve Video production.

How do I install Watch in Claude Code?

Run `npx skills add mathiaschu/watch --skill watch -a claude-code`. Or copy the skill folder (the mathiaschu/watch repository) into .claude/skills/watch in your project. Claude Code loads it when a task matches its description.

How do I install Watch in Codex?

Run `npx skills add mathiaschu/watch --skill watch -a codex`. Or copy the skill folder (the mathiaschu/watch repository) into .agents/skills/watch in your project. Codex loads it when a task matches its description.

Can I use Watch in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add mathiaschu/watch --skill watch -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/watch, .gemini/skills/watch, .github/skills/watch and .opencode/skills/watch in your project.

What does Watch need to run?

Going by SKILL.md and its folder, Watch needs a shell and Python for the scripts in its folder and the command-line tools its instructions call (python3, pip3, whisper and ffmpeg). Our summary lists: Python 3; A Bash shell. Its frontmatter pre-approves these tools: Bash, Read.

Does Watch access the network?

SKILL.md names 2 domains. In commands or code: youtu.be and instagram.com; the agent is likely to contact these when it follows the instructions. This is read from the text; nothing was executed.

Is Watch safe to install?

Our automated static check of SKILL.md flagged 2 warning(s): mentions a credentials file (ssh keys, cloud or package-manager tokens). Read the flagged lines before installing; the check is not a guarantee either way. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Watch use?

Watch is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Watch use?

About 4k tokens (SKILL.md is roughly 16k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Watch?

Skills that share tags, products or a category with Watch: Ffmpeg Skill (kajisho5/ffmpeg-skill, 1.9k stars), Watch (TheCraigHewitt/skills, 157 stars), Autoshorts (Upload-Post/skill-autoshorts, 151 stars) and Douyin Downloader (OpenMinis/MinisSkills, 444 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Watch?

mathiaschu (a GitHub user) maintains it in mathiaschu/watch, which has 141 GitHub stars. The repository was last updated on May 29, 2026.

Source: mathiaschu/watch on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.