Agent skill

Watch Video

by coreyhaines31 in coreyhaines31/makerskills

When you want to extract content from a video — YouTube, Loom, Vimeo, Riverside, Zoom recording, local MP4, X/IG video, anything yt-dlp supports.

MITAuto-check passedMedia & Creative

Install Watch Video

skills CLI
$ npx skills add coreyhaines31/makerskills --skill watch-video -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install coreyhaines31/makerskills watch-video --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/coreyhaines31/makerskills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/watch-video .claude/skills/watch-video && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
watch-video
GitHub stars
851
Token cost
~3.8k tokens
SKILL.md length
1,359 words
Files
1
Skills in repo
21
Repo updated
First seen
Licence
MIT

At a glance

When you want to extract content from a video — YouTube, Loom, Vimeo, Riverside, Zoom recording, local MP4, X/IG video, anything yt-dlp supports.

  • Works in 10 steps: Parse input → Parse depth mode → Pull metadata → …
  • /watch-video <url
  • SKILL.md covers Step 1 — Parse input, Step 2 — Parse depth mode, Step 3 — Pull metadata and Step 4 — Build workdir, plus 10 more sections
  • Calls yt-dlp, ffmpeg and curl; reaches generativelanguage.googleapis.com and loom.com; needs GEMINI_API_KEY and LOOM_API_KEY

What it does

Watch Video is an agent skill from coreyhaines31/makerskills. When you want to extract content from a video — YouTube, Loom, Vimeo, Riverside, Zoom recording, local MP4, X/IG video, anything yt-dlp supports. Three depth modes user picks per invocation — transcript (just words, fast/free), visual (transcript + ffmpeg frame extraction + Claude vision pass on key moments), multimodal (Gemini native video ingestion if $GEMINIAPIKEY set, else dense Claude vision). Uses local Whisper for transcription (MLX-Whisper on Apple Silicon, faster-whisper elsewhere), falls back to…

Its SKILL.md is about 3.8k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Media & Creative, covering Transcription, Video production and Speech recognition and synthesis. It works with YouTube, FFmpeg, Google Gemini and Whisper. The repository describes itself as: AI agent skills for the personal operator's craft — decisions, research, second-brain, content rotation, scenario modeling, and meta-skills to author more. Works with Claude… The licence is MIT.

When your agent uses it

  • /watch-video <url
  • Watch this video
  • Transcribe this loom
  • Analyze this video

Example prompts

  • “/watch-video <url,”
  • “watch this video,”
  • “transcribe this loom,”
  • “/watch-video”

Requirements

  • Python 3
  • A credential in GEMINI_API_KEY
  • A credential in LOOM_API_KEY

Workflow steps

10 steps, taken from the step headings in SKILL.md.

  1. Parse input
  2. Parse depth mode
  3. Pull metadata
  4. Build workdir
  5. Get the transcript
  6. If transcript mode: stop here
  7. If visual mode: extract frames + vision pass
  8. If multimodal mode
  9. Optional: capture to second-brain
  10. Report

What it can do on your machine

Read from SKILL.md and the folder at commit cc31579. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • yt-dlp
    • ffmpeg
    • curl
    • pip
    • jq
    • brew
    • ffprobe
    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • generativelanguage.googleapis.com
    • loom.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • GEMINI_API_KEY
    • LOOM_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Watch Video loads about 3.8k tokens when it runs. Until then it costs about 248 tokens; SKILL.md has 1,359 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~248
When it runs · the whole SKILL.md, loaded when a task matches
~3.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from coreyhaines31/makerskills at commit cc31579, republished under its MIT licence (© coreyhaines31). 1,359 words, ~3,760 tokens.

Download SKILL.mdSave it as .claude/skills/watch-video/SKILL.md (or your agent's skills folder).
name
watch-video
description
When you want to extract content from a video — YouTube, Loom, Vimeo, Riverside, Zoom recording, local MP4, X/IG video, anything yt-dlp supports. Three depth modes user picks per invocation — transcript (just words, fast/free), visual (transcript + ffmpeg frame extraction + Claude vision pass on key moments), multimodal (Gemini native video ingestion if $GEMINI_API_KEY set, else dense Claude vision). Uses local Whisper for transcription (MLX-Whisper on Apple Silicon, faster-whisper elsewhere), falls back to platform-provided transcripts when available (Loom, Riverside, YouTube auto-subs). Saves to ~/Documents/videos/<source>-<slug>-<date>/ and optionally captures summary to second-brain raw/ as call-/meeting-/note-. Triggers on "/watch-video <url>," "watch this video," "transcribe this loom," "analyze this video," "summarize this recording," "key moments from this," "what happened in this video." This skill replaces and broadens the prior youtube-transcript skill.
metadata.version
0.2.3

/watch-video — Transcribe and analyze any video at the depth you choose

Replaces and broadens the prior youtube-transcript skill. YouTube is now one of many sources; depth is user-controlled.

Step 1 — Parse input

Accept:

  • YouTube: full URL, youtu.be/<id>, youtube.com/shorts/<id>, raw 11-char ID
  • Loom: loom.com/share/<id> or loom.com/embed/<id>
  • Vimeo: vimeo.com/<id>
  • Riverside: download URL or local file
  • Zoom: local .mp4 from a downloaded recording
  • X / IG / TikTok video: URL — defers to social-fetch for metadata, uses yt-dlp for the file
  • Local file: any path to an .mp4 / .mov / .webm / .mkv

Detect source from URL pattern or file extension. If ambiguous, ask.

Step 2 — Parse depth mode

InvocationModeWhat you get
/watch-video <url>transcript (default)Clean text, metadata, optional chapters
/watch-video <url> transcripttranscriptSame as default
/watch-video <url> visualvisualTranscript + frames at intervals + Claude vision pass identifying key moments
/watch-video <url> multimodalmultimodalNative video to Gemini (if $GEMINI_API_KEY), else dense Claude vision frame-by-frame

If the depth isn't specified and the video is >10 minutes, ask before defaulting (visual/multimodal cost real money on long videos).

Step 3 — Pull metadata

For URL sources, use yt-dlp:

bash
yt-dlp --print "%(title)s|%(uploader)s|%(duration_string)s|%(upload_date>%Y-%m-%d)s|%(description)s" \
  --print "%(chapters)j" --skip-download "<url>"

Capture: title, uploader/channel, duration, upload date, description (first paragraph), chapters (JSON or null).

For local files, use ffprobe:

bash
ffprobe -v error -show_entries format=duration -of default=noprint_wrappers=1:nokey=1 "<file>"

Step 4 — Build workdir

~/Documents/videos/<source>-<slug>-<date>/

Where:

  • source: youtube / loom / vimeo / riverside / zoom / local
  • slug: kebab-case of title (first 4–6 words, max 50 chars)
  • date: YYYY-MM-DD

Step 5 — Get the transcript

Backend selection (in order):

  1. Platform-provided transcript if it exists and looks complete:

    • YouTube: yt-dlp --write-sub --write-auto-sub --skip-download --sub-lang en --sub-format vtt
    • Loom: fetch via https://www.loom.com/share/<id> page metadata or Loom API if $LOOM_API_KEY set
    • Riverside: built-in transcripts available on the recording's share page
    • If platform transcript exists and has timestamps, use it. Skip Whisper.
  2. MLX-Whisper local (default on Apple Silicon Macs):

    bash
    # Install once: pip install mlx-whisper
    python3 -c "import mlx_whisper; mlx_whisper.transcribe('<file>', path_or_hf_repo='mlx-community/whisper-large-v3-turbo')" \
      > "<workdir>/transcript-raw.json"

    Or via the CLI: mlx_whisper <file> --model mlx-community/whisper-large-v3-turbo --output-dir <workdir>

  3. Not on Apple Silicon (Linux, Intel Mac, Windows, cloud agents): faster-whisper (pip install faster-whisper, CPU or CUDA) or whisper.cpp. Same model size, same output handling.

Download the video file first if it's a URL (use yt-dlp; Loom/Vimeo/YT all supported):

bash
yt-dlp -f "bv*[height<=720]+ba/b[height<=720]" -o "<workdir>/video.%(ext)s" "<url>"

720p is plenty for transcription and frame analysis (smaller download, faster processing).

Clean the transcript (only needed for YouTube auto-subs which have rolling captions; Whisper output is already clean):

bash
# YouTube VTT cleanup — de-dup rolling captions, strip tags, paragraph-break on cue gaps >2s
awk '
  /^WEBVTT/ || /^Kind:/ || /^Language:/ || /^NOTE/ { next }
  /-->/ { in_cue = 1; last = ""; next }
  /^$/ { if (last) print last; in_cue = 0; last = ""; next }
  in_cue { gsub(/<[^>]+>/, "", $0); last = $0 }
  END { if (last) print last }
' "<workdir>/transcript.en.vtt" | awk '!seen[$0]++' > "<workdir>/transcript.txt"

Save final to <workdir>/transcript.txt.

Step 6 — If transcript mode: stop here

Output:

  • transcript.txt
  • metadata.json
  • One-line summary in chat: title, source, duration, word count
  • Path to workdir
  • (Optional) Step 9 — offer to capture to second-brain

Step 7 — If visual mode: extract frames + vision pass

Frame extraction (ffmpeg)

Cadence by source heuristic:

Source typeFrame cadence
Screen-share / Loom / demo1 frame per 5s (UI changes fast)
Talking head / podcast1 frame per 30s (slow change)
Slide presentation1 frame per 10s + force a frame on each detected scene change
Default if unsure1 frame per 15s
bash
mkdir -p "<workdir>/frames"
ffmpeg -i "<workdir>/video.mp4" -vf "fps=1/15" "<workdir>/frames/frame-%04d.png" -y

For scene-change detection (slide decks especially):

bash
ffmpeg -i "<workdir>/video.mp4" -vf "select='gt(scene,0.3)',showinfo" -vsync vfr "<workdir>/frames/scene-%04d.png" 2> "<workdir>/scene-detection.log"
Vision pass

Pair each frame with the transcript chunk for the same timestamp window. Then batch-send to Claude vision for synthesis.

Per-frame batch prompt (up to ~10 frames per call):

Here are N frames from a video at timestamps T1..TN. For each frame, describe what's on screen in 1–2 sentences. Flag: (a) UI changes from previous frame, (b) text visible on screen, (c) any moment that looks like a decision, action, or notable event. Also note the transcript text spoken during this window.

Save the output as <workdir>/moments.md:

markdown
# Key moments — <title>

## 00:00:15 (frame-001.png)
**On screen**: Login form, email field focused
**Transcript**: "So you just open it up and..."
**Note**: Beginning of UI demo

## 00:00:45 (frame-002.png)
**On screen**: Dashboard with 4 cards
**Transcript**: "And here's where you see all your projects."
**Note**: Major view change — first time the dashboard appears
Generate summary

After moments are identified, synthesize the whole video into <workdir>/summary.md:

markdown
# Summary — <title>

**Source:** <source URL / file>
**Duration:** <hh:mm:ss>
**Watched at:** <date>
**Mode:** visual

## TL;DR
<2–4 sentences>

## Key moments
- 00:00:15 — <one-line>
- 00:00:45 — <one-line>

## Action items flagged
- <item> [timestamp]

## Decisions flagged
- <decision> [timestamp] — consider routing to /decide

## Quotes worth keeping
- "..." [timestamp]

## Open questions
- <question raised but not answered>

Step 8 — If multimodal mode

Backend selection
  1. Gemini native if $GEMINI_API_KEY is set (much cheaper + faster than per-frame for long videos):

    Default model: gemini-3.5-flash (released May 2026, ~$1.50 input / $9 output per 1M tokens; ~$0.15/sec of video; beats 3.1 Pro on coding/agentic benchmarks at 4× the speed). Override to gemini-3.1-pro for brand audits / high-stakes analysis where details matter; gemini-2.5-flash-lite for bulk cheap processing.

    bash
    # Step 1: Upload video via Files API
    FILE_URI=$(curl -s -X POST "https://generativelanguage.googleapis.com/upload/v1beta/files?key=$GEMINI_API_KEY" \
      -H "X-Goog-Upload-Command: start, upload, finalize" \
      -H "Content-Type: video/mp4" \
      --data-binary "@<workdir>/video.mp4" | jq -r '.file.uri')
    
    # Wait until file is ACTIVE (Gemini processes the video first)
    while true; do
      STATE=$(curl -s "$FILE_URI?key=$GEMINI_API_KEY" | jq -r '.state')
      [ "$STATE" = "ACTIVE" ] && break
      sleep 3
    done
    
    # Step 2: Generate content with the file + multimodal-analysis prompt
    curl -s -X POST "https://generativelanguage.googleapis.com/v1beta/models/gemini-3.5-flash:generateContent?key=$GEMINI_API_KEY" \
      -H "Content-Type: application/json" \
      -d "{
        \"contents\":[{
          \"parts\":[
            {\"file_data\":{\"mime_type\":\"video/mp4\",\"file_uri\":\"$FILE_URI\"}},
            {\"text\":\"<multimodal analysis prompt — see Step 7's summary template + use-case extensions>\"}
          ]
        }]
      }"

    Files persist in Gemini Files API for ~48 hours — useful for re-querying the same video with different prompts.

  2. Dense Claude vision fallback if no Gemini key:

    • Frame cadence: 1 frame per 3s (much denser than visual mode)
    • Batch through Claude vision with the multimodal-analysis prompt
    • Slower and more expensive than Gemini for long videos — warn the user before running on >10min content
Multimodal output

Same summary.md template as Step 7 + an extended section:

markdown
## Multimodal observations
- **Body language / delivery**: <observations on talking-head video>
- **Pacing**: <fast/slow/uneven>
- **Visual style**: <brand audit, ad review, design observations>
- **Audio quality / atmosphere**: <music, silence, background>

Exact extra sections depend on the use case (brand audit, ad review, talk delivery review, client-call read). Use case is inferred from the source + the user's verbal framing when invoking.

Step 9 — Optional: capture to second-brain

After any mode completes, offer:

"Want to capture this to second-brain? I'll write a call-<slug>.md (or meeting- / note- / resource-) to ${SECOND_BRAIN_VAULT:-$HOME/Documents/SecondBrain}/raw/ with the summary, source URL, and transcript link."

Type prefix by source:

SourcePrefix
Loom / Zoom / Riverside / Otter / call recordingcall-
Meeting (own notes, not a transcript)meeting-
Talk / keynote / conferencenote-
Ad / landing-page video / marketing reference / competitor videoresource-

File body: 1-line source, the summary, link to full workdir.

Show full SKILL.md (545 more words)Show less

Step 10 — Report

In chat:

  • One-line headline: <source> · <title> · <duration> · <mode> · <word count> words
  • Workdir path
  • For visual / multimodal: brief list of top 3 key moments
  • For all modes: any action items / decisions flagged for triage
  • If captured to second-brain: that path too

Sources reference

SourceDownloadBuilt-in transcriptNotes
YouTubeyt-dlpAuto-subs (--write-auto-sub)Same as the prior youtube-transcript skill
Loomyt-dlp (Loom supported)Yes — fetch via embed metadata or Loom APIAsync screenshare focus — prime use case
Vimeoyt-dlpSometimesMarketing/embed videos
RiversideDirect URL from export, or local fileYes — Riverside generates themPodcast episodes
ZoomLocal .mp4 (downloaded recordings)Sometimes (Zoom audio transcript file)Client calls
X / IG / TikTokDefer to social-fetch for metadata, yt-dlp for fileNoShort-form
Local filen/an/aDrop a path

Composes with

  • social-fetch — for X/IG/TikTok URL metadata (engagement, author, replies) before video processing
  • second-brain — capture summary as raw/call-<slug>.md, meeting-, note-, or resource- per source type
  • decide — when a video contains a flagged decision, route to /decide for structured capture
  • pm — action items flagged in summary can be triaged to project boards
  • slide-deck — talk recordings → outline extraction → deck draft (loop)
  • jab-hook — quotes + clip-worthy moments from podcast/talk videos feed BIP/promo posts
  • skillify from-video — primary use case for visual mode on process recordings. the user records themselves doing a workflow (Loom/screen-share), this skill extracts transcript + key visual moments, then skillify synthesizes the workflow into a SKILL.md. "Record once, AI converts to skill."

Error handling

FailureResponse
Video unavailable / private / region-lockedReport and stop
No subtitles + Whisper not installedTell the user: pip install mlx-whisper (Apple Silicon) or pip install faster-whisper (anything else)
ffmpeg missing (for visual/multimodal)Tell the user: brew install ffmpeg
Vision pass returns empty / unclearLower the frame count, retry, or fall back to transcript-only with a note
Multimodal requested but no $GEMINI_API_KEY and >30min videoWarn cost, offer to fall back to visual mode
yt-dlp binary missingbrew install yt-dlp

Notes on quality

  • User picks depth, not the skill. Transcript / visual / multimodal are 3 different cost + latency profiles. Long videos (>10 min) always confirm before spending on visual/multimodal.
  • Platform transcript first, Whisper second. YouTube auto-subs, Loom transcripts, Riverside built-in transcripts — all free + instant when they exist. Fall back to local Whisper only when nothing platform-provided works.
  • Local Whisper is the fast path. MLX on M-series Macs transcribes faster than real-time; faster-whisper is the equivalent elsewhere. Cloud Whisper is a distant second choice — costs money, network dependency, worse latency on typical durations.
  • Frame cadence by source type. Screen-share / demos need 1 frame per 5s (UI changes fast); talking-head podcasts need 1 per 30s (slow change). Default 15s if unsure. Wrong cadence = missed key moments OR wasted vision-pass cost.
  • 720p is plenty. Downloading 1080p / 4K for transcription + frame analysis wastes bandwidth + storage. yt-dlp -f "bv*[height<=720]+ba/b[height<=720]" is the default.
  • Scene-change detection catches slide transitions. When the video is a slide presentation, add ffmpeg -vf "select='gt(scene,0.3)'" to force a frame on each detected slide change — more reliable than pure time-based sampling.
  • Multimodal cost warning is non-optional. Gemini multimodal on a 60-min video is meaningfully expensive. Warn before running; offer transcript-only as fallback if the user isn't sure.
  • Summary format includes routing hints. ## Decisions flagged + ## Action items flagged sections signal /decide and /pm follow-ups. Downstream composability lives in the summary structure.

© coreyhaines31, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/watch-video of coreyhaines31/makerskills.

Open the folder on GitHubat commit cc31579

Compare with similar skills

Watch Video next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Watch Video compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Watch Video this skillcoreyhaines31/makerskills851—~3.8kAutomated safety check: PassMIT
Watchmathiaschu/watch142—~4kAutomated safety check: WarnMIT
Video Clip Extractorlinzzzzzz/openclip569—~2.8kAutomated safety check: WarnMIT
Bggg Tiktok Readvideobinggandata/bggg-skills605—~1.6kAutomated safety check: PassMIT
Video Transcript Downloadersundial-org/awesome-openclaw-skills6632 repos~574Automated safety check: PassNone
9Router Speech-to-Textdecolua/9router31k—~914Automated safety check: PassMIT

Similar skills

  • Watch

    mathiaschu/watch

    Watch a video from YouTube, Instagram, X/Twitter, Vimeo, TikTok or any of ~1800 yt-dlp sites (or a local path).

    142 GitHub stars~4k tokensUpdated 4 mo ago
    Media & CreativeAuto-check: warnings
  • Video Clip Extractor

    linzzzzzz/openclip

    Processes videos to identify engaging moments, generate transcripts, and create highlight clips with artistic titles and custom cover images.

    569 GitHub stars~2.8k tokensUpdated 1 mo ago
    Media & CreativeAuto-check: warnings
  • Bggg Tiktok Readvideo

    binggandata/bggg-skills

    把 TikTok、Reels、YouTube Shorts、UGC 广告、本地 MP4/MOV/WebM 等视频拆成 Codex 可读的视频上下文。

    605 GitHub stars~1.6k tokensUpdated 1 mo ago
    Media & CreativeAuto-check passed
  • Video Transcript Downloader

    sundial-org/awesome-openclaw-skills

    Download videos, audio, subtitles, and clean paragraph-style transcripts from YouTube and any other yt-dlp supported site.

    663 GitHub starsUsed in 2 repos~574 tokens
    Media & CreativeAuto-check passed
  • 9Router Speech-to-Text

    decolua/9router

    Transcribes audio files into text or subtitles through 9Router's Whisper-compatible endpoint, using models from OpenAI, Groq, Gemini, Deepgram and others.

    31k GitHub stars~914 tokensUpdated 2 days ago
    Media & CreativeAuto-check passed
  • Video To Notes

    KIRVO-REPORTING/video-to-notes

    Use immediately for any bare YouTube or YouTube Shorts URL, youtu.be link, Bilibili or b23.tv link, or other video URL; do not ask what the user wants.

    105 GitHub stars~1.5k tokensUpdated 1 mo ago
    Media & CreativeAuto-check passed

More from coreyhaines31/makerskills

All 21 skills in this repo
  • Business Brainstorm

    coreyhaines31/makerskills

    When you want to pressure-test a potential new business, product, or side project against the serial-founder filter.

    851 GitHub stars~1.7k tokensUpdated 2 days ago
    Auto-check passed
  • Company Brain

    coreyhaines31/makerskills

    Your team's shared, AI-ready knowledge base — people, companies, meetings, SOPs, and decisions structured so an agent can answer on your team's behalf.

    851 GitHub stars~4.9k tokensUpdated 2 days ago
    Auto-check passed
  • Company Cfo

    coreyhaines31/makerskills

    Monthly CFO workflow for a company or agency — pull raw data from bank + payment processor + payroll + expense management, categorize and reconcile, compute end-of-month cash via transaction-sum…

    851 GitHub stars~4k tokensUpdated 2 days ago
    Auto-check: notes
  • Decide

    coreyhaines31/makerskills

    When you have a decision to make and want a structured workflow that picks the load-bearing questions, walks through them, reaches a call (or "wait"), and archives the rationale for future reference.

    851 GitHub stars~1.9k tokensUpdated 2 days ago
    Auto-check passed
  • Ingest

    coreyhaines31/makerskills

    When you paste raw human input — a call transcript (Grain, Zoom, Granola, Fathom), a text or email from a client/partner/friend, a voice-memo dump, or meeting notes — and want it converted into…

    851 GitHub stars~1.8k tokensUpdated 2 days ago
    Auto-check passed
  • Maker Council

    coreyhaines31/makerskills

    When you want multiple expert perspectives on a founder/operator question — a simulated board of advisors (Jason Fried, Elon Musk, Jeff Bezos, Jensen Huang, Bob Iger, Paul Graham, Naval Ravikant…

    851 GitHub stars~2.9k tokensUpdated 2 days ago
    Auto-check passed

Questions about Watch Video

What does Watch Video do?

When you want to extract content from a video — YouTube, Loom, Vimeo, Riverside, Zoom recording, local MP4, X/IG video, anything yt-dlp supports. Watch Video is an agent skill from coreyhaines31/makerskills. When you want to extract content from a video — YouTube, Loom, Vimeo, Riverside, Zoom recording, local MP4, X/IG video, anything yt-dlp supports.

When should I use Watch Video?

Watch Video fits situations like: /watch-video <url; watch this video; transcribe this loom; analyze this video.

How do I install Watch Video in Claude Code?

Run `npx skills add coreyhaines31/makerskills --skill watch-video -a claude-code`. Or copy the skill folder (skills/watch-video in coreyhaines31/makerskills) into .claude/skills/watch-video in your project. Claude Code loads it when a task matches its description.

How do I install Watch Video in Codex?

Run `npx skills add coreyhaines31/makerskills --skill watch-video -a codex`. Or copy the skill folder (skills/watch-video in coreyhaines31/makerskills) into .agents/skills/watch-video in your project. Codex loads it when a task matches its description.

Can I use Watch Video in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add coreyhaines31/makerskills --skill watch-video -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/watch-video, .gemini/skills/watch-video, .github/skills/watch-video and .opencode/skills/watch-video in your project.

What does Watch Video need to run?

Going by SKILL.md and its folder, Watch Video needs the command-line tools its instructions call (yt-dlp, ffmpeg, curl, pip, jq and brew) and credentials named GEMINI_API_KEY and LOOM_API_KEY. Our summary lists: Python 3; A credential in GEMINI_API_KEY; A credential in LOOM_API_KEY.

Does Watch Video access the network?

SKILL.md names 2 domains. In commands or code: generativelanguage.googleapis.com and loom.com; the agent is likely to contact these when it follows the instructions. This is read from the text; nothing was executed.

Is Watch Video safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Watch Video use?

Watch Video is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Watch Video use?

About 3.8k tokens (SKILL.md is roughly 15k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Watch Video?

Skills that share tags, products or a category with Watch Video: Watch (mathiaschu/watch, 142 stars), Video Clip Extractor (linzzzzzz/openclip, 569 stars), Bggg Tiktok Readvideo (binggandata/bggg-skills, 605 stars) and Video Transcript Downloader (sundial-org/awesome-openclaw-skills, 663 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Watch Video?

coreyhaines31 (a GitHub user) maintains it in coreyhaines31/makerskills, which has 851 GitHub stars. The repository holds 21 skills in this directory. The repository was last updated on October 8, 2026.

Source: coreyhaines31/makerskills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.