When the user wants to read, transcribe, summarize, or research a video — YouTube link, podcast clip, Loom, TikTok, X/Twitter video, local file, or any URL yt-dlp supports.

MITAuto-check: notesMedia & Creative

Install Watch

skills CLI
$ npx skills add TheCraigHewitt/skills --skill watch -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install TheCraigHewitt/skills watch --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/TheCraigHewitt/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/general/watch .claude/skills/watch && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
watch
GitHub stars
157
Token cost
~1.6k tokens
SKILL.md length
769 words
Files
2
Skills in repo
65
Repo updated
First seen
Licence
MIT

At a glance

When the user wants to read, transcribe, summarize, or research a video — YouTube link, podcast clip, Loom, TikTok, X/Twitter video, local file, or any URL yt-dlp supports.

  • Works in 3 steps: Download audio only (transcript-only… → Upload it to Groq's whisper-large-v3… → Return the same [MM:SS] timestamped format
  • Research a video — YouTube link
  • SKILL.md covers Requirements, How to invoke, When the script needs Whisper and Output flow, plus 4 more sections
  • Runs Python scripts from its folder; calls python3, brew and winget; reaches youtu.be; needs GROQ_API_KEY and OPENAI_API_KEY

What it does

Watch is an agent skill from TheCraigHewitt/skills. When the user wants to read, transcribe, summarize, or research a video — YouTube link, podcast clip, Loom, TikTok, X/Twitter video, local file, or any URL yt-dlp supports. Use when they paste a video URL, say 'transcribe this,' 'summarize this video,' 'what does this video say about X,' 'pull the transcript,' 'analyze this YouTube video,' or hand you a video for content research. Default is fast transcript-only (no video download). Pass --with-frames when the visual layer matters (UI bugs, on-screen text, visual…

Its SKILL.md is about 1.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 1 other file (for example `watch.py`).

It sits in Media & Creative, covering Transcription and Speech recognition and synthesis. It works with YouTube, TikTok, X (Twitter) and FFmpeg. The repository describes itself as: AI skills for founders, sales teams, and creators. 47 skills across CEO, Sales, YouTube, and General categories. Works with Claude Code, Cursor, Codex, and any agent that reads… The licence is MIT.

When your agent uses it

  • Research a video — YouTube link
  • X/Twitter video
  • Any URL yt-dlp supports
  • They paste a video URL

Example prompts

  • “transcribe this,”
  • “summarize this video,”
  • “what does this video say about X,”
  • “/watch”

Requirements

  • Python 3
  • A credential in GROQ_API_KEY
  • A credential in OPENAI_API_KEY

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. Download audio only (transcript-only mode) or extract from the downloaded video (--with-frames)
  2. Upload it to Groq's whisper-large-v3 (preferred — cheap, fast) or OpenAI's whisper-1
  3. Return the same [MM:SS] timestamped format

What it can do on your machine

Read from SKILL.md and the folder at commit fdbf39b. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python3
    • brew
    • winget
    • pipx
    • apt

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • youtu.be

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • GROQ_API_KEY
    • OPENAI_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Watch loads about 1.6k tokens when it runs. Until then it costs about 134 tokens; SKILL.md has 769 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~134
When it runs · the whole SKILL.md, loaded when a task matches
~1.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteRuns commands with sudoSKILL.md:20
    ` + `ffprobe` — `brew install ffmpeg` / `sudo apt install ffmpeg` / `winget install Gyan.FFmpeg`. Only required for `--w
  • NoteMentions a .env fileSKILL.md:83
    cript reads env vars or `~/.config/watch/.env`):
  • NoteMentions a .env fileSKILL.md:119
    to stdout. They live in `~/.config/watch/.env` (mode 0600).

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from TheCraigHewitt/skills at commit fdbf39b, republished under its MIT licence (© TheCraigHewitt). 769 words, ~1,614 tokens.

Download SKILL.mdSave it as .claude/skills/watch/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
watch
description
When the user wants to read, transcribe, summarize, or research a video — YouTube link, podcast clip, Loom, TikTok, X/Twitter video, local file, or any URL yt-dlp supports. Use when they paste a video URL, say 'transcribe this,' 'summarize this video,' 'what does this video say about X,' 'pull the transcript,' 'analyze this YouTube video,' or hand you a video for content research. Default is fast transcript-only (no video download). Pass --with-frames when the visual layer matters (UI bugs, on-screen text, visual breakdowns).
metadata.version
1.0.0

Watch — read a video like a PDF

You don't have a video input. This skill bundles a Python script that fetches the timestamped transcript (and optionally frames) so you can answer questions about video content.

Default mode is transcript-only. No video download, no frame extraction — just captions pulled via yt-dlp in a few seconds. That covers ~every YouTube video and is the right default for research, summarization, and content analysis.

Opt into --with-frames only when the visual layer matters: debugging a screen recording, breaking down a thumbnail/hook visually, reading on-screen text, analyzing UI behavior.

Requirements

Locally installed:

  • yt-dlp — brew install yt-dlp (macOS) / pipx install yt-dlp (Linux) / winget install yt-dlp.yt-dlp (Windows)
  • ffmpeg + ffprobe — brew install ffmpeg / sudo apt install ffmpeg / winget install Gyan.FFmpeg. Only required for --with-frames and the Whisper audio fallback. Transcript-only on a captioned YouTube video does not need ffmpeg.

If the user's missing one, tell them the install command and stop. Do not auto-install.

How to invoke

The script lives next to this SKILL.md. Invoke it with the absolute path or ${CLAUDE_SKILL_DIR}/watch.py:

bash
python3 "${CLAUDE_SKILL_DIR}/watch.py" "<url-or-path>"

If ${CLAUDE_SKILL_DIR} isn't set in this agent, resolve the path yourself — it's the directory holding this SKILL.md file.

Default (transcript-only)
bash
python3 "${CLAUDE_SKILL_DIR}/watch.py" "https://youtu.be/abc123"

Returns a markdown report with title, channel, duration, and a timestamped transcript like:

[00:01] All right, so here we are, in front of the elephants
[00:05] the cool thing about these guys is that they have really...
With frames (visual layer)
bash
python3 "${CLAUDE_SKILL_DIR}/watch.py" "<url>" --with-frames

Returns frame paths + transcript. Read each frame path in parallel with the Read tool so you see the full chronological visual flow alongside the transcript.

Focus on a section
bash
python3 "${CLAUDE_SKILL_DIR}/watch.py" "<url>" --start 1:30 --end 2:00

Works in both modes. Times accept SS, MM:SS, or HH:MM:SS. In transcript-only mode it filters the transcript. In --with-frames mode it also limits frame extraction and switches to a denser frame budget.

Other flags
FlagPurpose
--no-whisperDon't fall back to Whisper if captions are missing. Fail with a clear error instead.
--whisper groq|openaiForce a Whisper backend. Default: prefer Groq, fall back to OpenAI.
--max-frames N(frames mode) Cap on frame count. Default 80, hard max 100.
--resolution W(frames mode) Frame width in px. Default 512. Bump to 1024 only if reading on-screen text matters.
--fps F(frames mode) Override auto-fps (clamped to 2 fps).
--out-dir DIRKeep working files somewhere specific (default: tmp).
--json(transcript mode) Emit machine-readable JSON instead of markdown.

When the script needs Whisper

If a video has no captions (rare for YouTube, common for Loom / Instagram / local files) and you didn't pass --no-whisper, the script will:

  1. Download audio only (transcript-only mode) or extract from the downloaded video (--with-frames)
  2. Upload it to Groq's whisper-large-v3 (preferred — cheap, fast) or OpenAI's whisper-1
  3. Return the same [MM:SS] timestamped format

To enable Whisper, set one of these (script reads env vars or ~/.config/watch/.env):

GROQ_API_KEY=...        # console.groq.com/keys
OPENAI_API_KEY=...      # platform.openai.com/api-keys

If neither key is set and captions are missing, the script exits 2 with a clear message — tell the user how to enable Whisper or that this particular video can't be transcribed.

Show full SKILL.md (299 more words)Show less

Output flow

Transcript-only (default). The script prints a self-contained markdown report. Quote timestamps when citing claims. Don't re-print the full transcript back to the user — summarize, extract, or answer their question.

--with-frames. Same report plus a list of frame paths. After running:

  1. Read every frame path in a single parallel batch — you need them all to follow the visual flow.
  2. Combine frames with the transcript when answering. If the user asked about the hook, look at frames 1-3 + transcript 0:00–0:15.

Cleanup

The script prints a working directory. If the user isn't going to ask follow-ups about this video, rm -rf <dir> when you're done. Otherwise leave it in place — re-runs can reuse it via --out-dir.

Common research workflows

Hook deconstruction. python3 watch.py "<url>" --start 0 --end 15 --with-frames → analyze cold open: words said, what's on screen, pattern interrupt timing.

Channel research. Run transcript-only on 5-10 competitor videos. Compare structure, topic mix, CTA placement, transcript length-to-duration ratio.

Quote extraction. Transcript-only + grep for a topic in the report (use Bash on the printed work dir).

UI bug repro. --with-frames on a screen recording, then ask which frame the bug appears. Frames mode shines here.

Don't

  • Don't re-run the script on a video you already watched this session — you have the transcript in context. Just answer.
  • Don't pass --with-frames for pure-text research. It downloads the full video and burns image tokens for no benefit.
  • Don't write API keys into the repo or to stdout. They live in ~/.config/watch/.env (mode 0600).

Security

The script runs yt-dlp and ffmpeg locally — no third-party server in the middle. Audio only leaves the machine when Whisper is needed AND a key is configured (Groq or OpenAI endpoint, whichever matches). The video itself never gets uploaded.

Inspect watch.py before first run if you want to verify.

© TheCraigHewitt, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in general/watch of TheCraigHewitt/skills.

  • SKILL.md
  • watch.py

Open the folder on GitHubat commit fdbf39b

Compare with similar skills

Watch next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Watch compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Watch this skillTheCraigHewitt/skills157—~1.6kAutomated safety check: NotesMIT
Watchmathiaschu/watch142—~4kAutomated safety check: WarnMIT
Lecture To Notesysyecust/lecture-to-notes273—~14kAutomated safety check: NotesCustom licence
Video Transcribewendy7756/AI-Video-Transcriber3.3k—~937Automated safety check: NotesApache-2.0
Ffmpeg Skillkajisho5/ffmpeg-skill1.9k—~7.4kAutomated safety check: PassMIT
Claude Real VideoHUANGCHIHHUNGLeo/claude-real-video2.2k—~639Automated safety check: PassMIT

Similar skills

  • Watch

    mathiaschu/watch

    Watch a video from YouTube, Instagram, X/Twitter, Vimeo, TikTok or any of ~1800 yt-dlp sites (or a local path).

    142 GitHub stars~4k tokensUpdated 4 mo ago
    Media & CreativeAuto-check: warnings
  • Lecture To Notes

    ysyecust/lecture-to-notes

    A skill your agent uses when users provide YouTube, Bilibili, or X/Twitter lecture URLs and want reader-first Chinese LaTeX/PDF notes with source-faithful claims, fluent authored prose, and verified…

    273 GitHub stars~14k tokensUpdated 6 days ago
    Documents & OfficeAuto-check: notes
  • Video Transcribe

    wendy7756/AI-Video-Transcriber

    Transcribe and summarize a video or podcast from a URL (YouTube, TikTok, Bilibili, Apple Podcasts, SoundCloud, 30+ platforms) or from a local media/.txt file.

    3.3k GitHub stars~937 tokensUpdated 24 days ago
    Media & CreativeAuto-check: notes
  • Ffmpeg Skill

    kajisho5/ffmpeg-skill

    Edit video and audio with local FFmpeg from natural-language requests: cut, trim, join, resize/reframe (9:16, 1:1), speed change, captions and subtitles (SRT/ASS, animated, karaoke), logos and text…

    1.9k GitHub stars~7.4k tokensUpdated 5 days ago
    Media & CreativeAuto-check passed
  • Claude Real Video

    HUANGCHIHHUNGLeo/claude-real-video

    Watch a video for the user. An agent skill from HUANGCHIHHUNGLeo/claude-real-video.

    2.2k GitHub stars~639 tokensUpdated yesterday
    Media & CreativeAuto-check passed
  • Video Clip Extractor

    linzzzzzz/openclip

    Processes videos to identify engaging moments, generate transcripts, and create highlight clips with artistic titles and custom cover images.

    569 GitHub stars~2.8k tokensUpdated 1 mo ago
    Media & CreativeAuto-check: warnings

More from TheCraigHewitt/skills

All 65 skills in this repo
  • Carousels

    TheCraigHewitt/skills

    Turns a piece of Craig's content (an email, a YouTube script, an essay, or pasted text) into a polished image carousel publishable to both LinkedIn and Instagram from one set of 1080x1350 slides.

    157 GitHub stars~2.3k tokensUpdated 4 mo ago
    Auto-check passed
  • Audience Research

    TheCraigHewitt/skills

    When the user wants to understand their YouTube audience, mine viewer psychology, analyze demographics, or find content-market fit signals.

    157 GitHub stars~2.4k tokensUpdated 4 mo ago
    Auto-check passed
  • Brain Dump

    TheCraigHewitt/skills

    Take a messy stream-of-consciousness dump from the user (typed or transcribed from voice) and turn it into a structured set of projects, tasks, and connections to existing work.

    157 GitHub stars~1.3k tokensUpdated 4 mo ago
    Auto-check passed
  • Channel Audit

    TheCraigHewitt/skills

    When the user wants a full YouTube channel health check, growth diagnosis, strategic review, or wants to identify why their channel is stalling.

    157 GitHub stars~2.4k tokensUpdated 4 mo ago
    Auto-check passed
  • Channel Strategy

    TheCraigHewitt/skills

    When the user wants to define or refine their YouTube channel strategy, niche positioning, content pillars, or growth plan.

    157 GitHub stars~2.2k tokensUpdated 4 mo ago
    Auto-check passed
  • Cold Call

    TheCraigHewitt/skills

    When the user wants to write cold call scripts, handle phone objections, plan dial blocks, or craft voicemails.

    157 GitHub stars~5k tokensUpdated 4 mo ago
    Auto-check passed

Questions about Watch

What does Watch do?

When the user wants to read, transcribe, summarize, or research a video — YouTube link, podcast clip, Loom, TikTok, X/Twitter video, local file, or any URL yt-dlp supports. Watch is an agent skill from TheCraigHewitt/skills. When the user wants to read, transcribe, summarize, or research a video — YouTube link, podcast clip, Loom, TikTok, X/Twitter video, local file, or any URL yt-dlp supports.

When should I use Watch?

Watch fits situations like: research a video — YouTube link; X/Twitter video; any URL yt-dlp supports; they paste a video URL.

How do I install Watch in Claude Code?

Run `npx skills add TheCraigHewitt/skills --skill watch -a claude-code`. Or copy the skill folder (general/watch in TheCraigHewitt/skills) into .claude/skills/watch in your project. Claude Code loads it when a task matches its description.

How do I install Watch in Codex?

Run `npx skills add TheCraigHewitt/skills --skill watch -a codex`. Or copy the skill folder (general/watch in TheCraigHewitt/skills) into .agents/skills/watch in your project. Codex loads it when a task matches its description.

Can I use Watch in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add TheCraigHewitt/skills --skill watch -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/watch, .gemini/skills/watch, .github/skills/watch and .opencode/skills/watch in your project.

What does Watch need to run?

Going by SKILL.md and its folder, Watch needs Python for the scripts in its folder, the command-line tools its instructions call (python3, brew, winget, pipx and apt) and credentials named GROQ_API_KEY and OPENAI_API_KEY. Our summary lists: Python 3; A credential in GROQ_API_KEY; A credential in OPENAI_API_KEY.

Does Watch access the network?

SKILL.md names 1 domain. In commands or code: youtu.be; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is Watch safe to install?

Our automated static check of SKILL.md found notes only (runs commands with sudo; mentions a .env file), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Watch use?

Watch is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Watch use?

About 1.6k tokens (SKILL.md is roughly 6.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Watch?

Skills that share tags, products or a category with Watch: Watch (mathiaschu/watch, 142 stars), Lecture To Notes (ysyecust/lecture-to-notes, 273 stars), Video Transcribe (wendy7756/AI-Video-Transcriber, 3.3k stars) and Ffmpeg Skill (kajisho5/ffmpeg-skill, 1.9k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Watch?

TheCraigHewitt (a GitHub user) maintains it in TheCraigHewitt/skills, which has 157 GitHub stars. The repository holds 65 skills in this directory. The repository was last updated on May 22, 2026.

Source: TheCraigHewitt/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.