Agent skill

Watchless

by chenzixin1 in chenzixin1/watchless

A skill your agent uses when turning a YouTube URL or local presentation, explainer, interview, podcast, or product-demo video into complete screenshot-led notes, faithful light-polished text, HTML…

MITAuto-check: warningsMedia & Creative

Install Watchless

The automated check flagged lines worth reading first. See the safety section below.

skills CLI
$ npx skills add chenzixin1/watchless --skill watchless -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install chenzixin1/watchless watchless --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
watchless
GitHub stars
144
Token cost
~3.7k tokens
SKILL.md length
1,529 words
Files
60 (incl. scripts, references, assets)
Skills in repo
1
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when turning a YouTube URL or local presentation, explainer, interview, podcast, or product-demo video into complete screenshot-led notes, faithful light-polished text, HTML…

  • Works in 6 steps: Acquire, transcribe, and inspect the… → Confirm route and visual strategies → Confirm selected frames → …
  • Turning a YouTube URL
  • SKILL.md covers Hard Requirements, Private-Use Compliance Gate, Route Model and Workflow, plus 2 more sections
  • Local presentation

What it does

Watchless is an agent skill from chenzixin1/watchless. Use when turning a YouTube URL or local presentation, explainer, interview, podcast, or product-demo video into complete screenshot-led notes, faithful light-polished text, HTML, PDF, or a shareable ZIP.

Its SKILL.md is about 3.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 61 other files, including scripts, reference files and assets (for example `LEGAL.en.md`, `LEGAL.md` and `README.en.md`).

It sits in Media & Creative, covering Video production, Transcription and PDF. It works with YouTube. The repository describes itself as: Codex Skill that turns videos into complete keyframe-led visual documents. The licence is MIT.

When your agent uses it

  • Turning a YouTube URL
  • Local presentation
  • Product-demo video into complete screenshot-led notes
  • Faithful light-polished text

Example prompts

  • “/watchless”

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Acquire, transcribe, and inspect the whole video
  2. Confirm route and visual strategies
  3. Confirm selected frames
  4. Conversation speaker identity pass
  5. Write complete scene notes
  6. Build and verify outputs

What it can do on your machine

Read from SKILL.md and the folder at commit 34e2fa8. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/, which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Watchless loads about 3.7k tokens when it runs, and up to ~4.2k if it reads all its reference files. Until then it costs about 53 tokens; SKILL.md has 1,529 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~53
When it runs · the whole SKILL.md, loaded when a task matches
~3.7k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~4.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: warnings

The automated check found patterns that need a careful read before installing.

  • WarningMentions a credentials file (SSH keys, cloud or package-manager tokens)SKILL.md:30
    - Chrome browser cookies may be read only from the user's local browser profile. Never export, print, copy, upload, log,
  • WarningMentions a credentials file (SSH keys, cloud or package-manager tokens)SKILL.md:214
    ied, try anonymous `yt-dlp`, then local Chrome browser cookies without exporting them. Do not use the cookie fallback to

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from chenzixin1/watchless at commit 34e2fa8, republished under its MIT licence (© chenzixin1). 1,529 words, ~3,664 tokens.

Download SKILL.mdSave it as .claude/skills/watchless/SKILL.md (or your agent's skills folder). This skill also uses 59 other files; get the full folder from GitHub.
name
watchless
description
Use when turning a YouTube URL or local presentation, explainer, interview, podcast, or product-demo video into complete screenshot-led notes, faithful light-polished text, HTML, PDF, or a shareable ZIP.

Watchless

English | 简体中文

Turn one video into a complete visual article. The transcript is the factual source; screenshots preserve the visual evidence. The reader should not need to watch the original video.

Hard Requirements

  • Run locally. Do not call PodSum, its website, MCP, Cloudflare, APIFY, D1, or R2.
  • Keep the source immutable and reuse cached downloads and transcripts.
  • Use Tencent Cloud word-level ASR with speaker diarization by default (--provider tencent; auto also selects Tencent). Never automatically fall back to Volcengine; select other providers explicitly. Do not silently replace it with YouTube automatic captions or local Whisper.
  • Keep the output chronological and complete. Light-plus is not a summary: preserve reasoning, examples, figures, caveats, disagreement, questions, answers, and repeated emphasis.
  • Show process images during the run: whole-video route overview, boundary evidence where applicable, candidate frames, selected frames, and PDF page overview.
  • Use no OpenCV anywhere in this Skill.
  • Visual algorithms are route-specific. slides uses ffmpeg keyframe recall plus local SSIM refinement. explainer, conversation, and demo do not use SSIM, perceptual hash, histogram scoring, face scoring, or other visual ranking; Codex directly reads their candidate images.
  • Resolve speaker identities for every conversation video before writing notes. Keep Speaker N when evidence remains weak or conflicting.
  • Record model token usage and cost in work/token-usage.json whenever the runtime exposes real counts. Never invent unavailable usage or silently apply stale prices.
  • Do not claim completion until HTML image references, rendered PDF pages, and ZIP contents pass verification.

Private-Use Compliance Gate

Before acquiring a remote source, confirm that the user owns the content, has permission to process it, or has independently established another lawful basis. If authorization is unclear, stop and request an authorized local file instead of downloading.

  • This repository and its outputs are private by default. Do not publish or distribute generated notes without a separate rights, privacy, confidentiality, and attribution review.
  • Chrome browser cookies may be read only from the user's local browser profile. Never export, print, copy, upload, log, or commit cookie values. Cookie access does not establish a right to download or reuse content.
  • Never use this Skill to bypass DRM, paywalls, members-only access, private-video access, geographic restrictions, CAPTCHAs, account enforcement, or other access controls.
  • Before sending non-public or sensitive audio to the selected cloud ASR provider, confirm that the user is authorized to make that third-party transfer. Use explicitly requested local Whisper or stop when authorization is unavailable.
  • Resolve speaker names only from explicit public metadata, self-introduction, lower thirds, or similarly reliable evidence. Do not perform face or voice biometric identification. Keep Speaker N when uncertain and require human review before external use.
  • Use the minimum screenshots and quotations needed for any approved external use. Do not publish a complete transcript or visual reconstruction that substitutes for the source without a separate legal review.

Read LEGAL.en.md before running this Skill on third-party, confidential, commercial, or sensitive material. These controls reduce risk but do not constitute legal advice.

Route Model

Choose one content route after directly inspecting verify/mode-overview.jpg and the transcript. Then add visual strategies. Do not route from the title or channel alone.

RouteObservable patternSegmentationFrame extraction
slidesStable PPT, slide deck, document pages, or a fixed presentation regionEvery meaningful slide/build stateHybrid H.264 keyframe recall, local SSIM refinement, accurate scan fallback
explainerScripted visual argument: presenter, charts, animation, maps, documents, B-rollComplete argument/example/visual functionDense time-distributed candidates inside each semantic scene; Codex selects
conversationInterview, panel, podcast, or question-and-answer exchangeComplete Q&A or topic unit, not every speaker turnActive speaker/two-shot candidates; evidence or B-roll can override faces
demoUI walkthrough, product demonstration, tutorial, or physical procedureExecutable step and observable resultCandidates biased toward before/after and completed states; Codex selects

Visual strategies are composable:

  • slide-state: stable, complete slide/build state.
  • speaker: active speaker or useful group shot.
  • evidence: chart, quote, interface, object, or source that carries the claim.
  • broll: relevant external footage in an edited interview or documentary.
  • dense-visual: increase candidate density for XiaoLin-style edited explainers.
  • document-evidence: prefer papers, article excerpts, tables, and diagrams.
  • screen-state: prefer a legible UI result rather than cursor motion or transitions.

For conversation, also choose one extraction profile without creating a new main route:

ProfileUse whenCandidate behavior
studioStable studio/remote podcast with little visual evidenceThree representative speaker/group candidates per semantic scene
editedInterview with B-roll, locations, documents, or archival footageNine distributed candidates plus local evidence-trigger candidates
newsBroadcast interview with lower thirds, tickers, charts, and product graphicsEight distributed candidates plus local evidence-trigger candidates
chapteredLong interview whose official chapters are useful topic hintsFive candidates; chapters seed review but never force a boundary through an answer
generalEvidence is insufficient to choose a more specific profileBackward-compatible five-candidate behavior

Typical routing examples:

  • XiaoLin-style scripted finance/technology video: explainer + dense-visual + evidence.
  • Best Partners paper/article explanation: explainer + document-evidence.
  • a16z or YC Lightcone studio podcast: conversation + speaker.
  • Silicon Valley 101 edited field interview: conversation + evidence + broll.
  • YC Design Review screen walkthrough: demo + screen-state + evidence.
  • Low-information entertainment chat: conversation + speaker, with fewer, larger topic units.

Workflow

Set SKILL_DIR to this Skill directory. Run commands from the workspace where outputs/video-notes/ should be created.

1. Acquire, transcribe, and inspect the whole video
bash
"$SKILL_DIR/.venv/bin/python" "$SKILL_DIR/scripts/00_build_video_notes.py" "$SOURCE" --stage prepare

Read PROJECT_DIR from output. Display verify/mode-overview.jpg, inspect the transcript and source metadata, and decide the route from observable structure.

2. Confirm route and visual strategies
bash
"$SKILL_DIR/.venv/bin/python" "$SKILL_DIR/scripts/00_build_video_notes.py" \
  --stage route --project-dir "$PROJECT_DIR" \
  --mode explainer \
  --visual-strategy dense-visual \
  --visual-strategy evidence \
  --mode-reason "scripted host alternates with explanatory charts and B-roll"

Later stages may use --mode auto; they read the confirmed work/route-decision.json. An explicit --mode remains a manual override. Compatibility aliases: presentation -> slides, editorial -> explainer.

For a Bloomberg-style news interview:

bash
"$SKILL_DIR/.venv/bin/python" "$SKILL_DIR/scripts/00_build_video_notes.py" \
  --stage route --project-dir "$PROJECT_DIR" \
  --mode conversation --conversation-profile news \
  --visual-strategy speaker --visual-strategy evidence \
  --mode-reason "broadcast interview with lower thirds and sparse data graphics"
Show full SKILL.md (617 more words)Show less
3A. Slides route

Slides do not use semantic timeline boundaries. If the slide occupies only part of the frame, directly inspect the overview and pass a relative crop rectangle.

bash
"$SKILL_DIR/.venv/bin/python" "$SKILL_DIR/scripts/00_build_video_notes.py" \
  --stage scenes --project-dir "$PROJECT_DIR" --mode auto \
  --slide-rect "0.04,0.08,0.72,0.92"

The default backend is hybrid-keyframe: scan H.264 keyframes at low resolution, compare the fixed slide region with SSIM, locally refine clustered changes on a 2-second grid, and fall back to accurate full scanning when keyframe recall is unavailable. Use --slide-backend accurate, --slide-interval, or --slide-threshold only when the default misses or over-splits states.

Display verify/selected-keyframes.jpg. Confirm that frames are complete slides, not black frames, fades, or partial transitions.

3B. Explainer, conversation, and demo routes

Prepare an audio-first semantic review:

bash
"$SKILL_DIR/.venv/bin/python" "$SKILL_DIR/scripts/00_build_video_notes.py" \
  --stage timeline --project-dir "$PROJECT_DIR" --mode auto

Read work/video-use/takes_packed.md and work/timeline/boundary-review.md. Propose boundaries only at complete phrases and semantic transitions. For each real decision point T, inspect approximately T-4s to T+4s with video-use/helpers/timeline_view.py; do not scan the full video at fixed intervals.

When official chapters exist, the timeline stage writes work/timeline/chapter-hints.json. Treat chapters only as coarse topic proposals: merge or split them to preserve complete questions, answers, examples, and qualifications.

Write work/timeline/scene-boundaries.json:

json
{"mode":"explainer","scenes":[{"end_sec":42.35,"reason":"hook and first claim complete"},{"end_sec":113.8,"reason":"chart explanation completes"}]}

The final boundary must cover the final transcript phrase. Then build scenes:

bash
"$SKILL_DIR/.venv/bin/python" "$SKILL_DIR/scripts/00_build_video_notes.py" \
  --stage scenes --project-dir "$PROJECT_DIR" --mode auto \
  --boundaries "$PROJECT_DIR/work/timeline/scene-boundaries.json"

Display candidate contact sheets and verify/selected-keyframes.jpg. Select frames by direct inspection. For explainers, prefer visual evidence over a face; for conversations, prefer the active speaker unless evidence/B-roll adds information; for demos, prefer the completed UI or physical result.

For edited and news conversations, the scene builder also looks for spoken references such as charts, data, reports, papers, products, chips, and screens. It adds nearby candidates labeled evidence to the contact sheet. These are recall hints, not automatic visual scores; Codex must still inspect and choose the frame.

4. Confirm selected frames

If default candidates are suitable:

bash
"$SKILL_DIR/.venv/bin/python" "$SKILL_DIR/scripts/00_build_video_notes.py" \
  --stage select --project-dir "$PROJECT_DIR"

Otherwise write selections such as {"1":2,"7":5} and pass --selections. Do not proceed until keyframe_review.status=complete.

5. Conversation speaker identity pass

For conversation, complete work/speaker-map.json before batches. Use evidence in this order:

  1. YouTube metadata and official description: title, credits, chapters, linked show notes.
  2. Self-introductions and reliable subtitles.
  3. Lower thirds, on-screen names, camera handoffs, and visible turn agreement.
  4. Official episode pages.
  5. Earlier episodes from the same channel for recurring-host corroboration.

Record display_name, role, confidence, and evidence. high requires explicit naming plus turn agreement; medium is a documented inference; low remains Speaker N. Use timestamped turn_overrides when ASR merges people or changes labels for one person. Set status=complete only after checking handoffs throughout the video.

6. Write complete scene notes
bash
"$SKILL_DIR/.venv/bin/python" "$SKILL_DIR/scripts/00_build_video_notes.py" \
  --stage batches --project-dir "$PROJECT_DIR"

Inspect every referenced image and transcript. Write each requested work/codex-notes/scene_NNN.md. Preserve the original order and voice; do not turn dialogue into third-person editorial narration. Visual descriptions must remain factual and useful.

If the runtime reports actual model usage, append it after each model-heavy phase. Pass current explicit rates only when cost should be calculated:

bash
"$SKILL_DIR/.venv/bin/python" "$SKILL_DIR/scripts/00_build_video_notes.py" \
  --stage usage --project-dir "$PROJECT_DIR" \
  --usage-stage notes --model-name "$MODEL" \
  --input-tokens "$INPUT_TOKENS" --output-tokens "$OUTPUT_TOKENS" \
  --cached-input-tokens "$CACHED_INPUT_TOKENS" \
  --input-rate-per-million "$INPUT_RATE" \
  --cached-input-rate-per-million "$CACHED_RATE" \
  --output-rate-per-million "$OUTPUT_RATE"
7. Build and verify outputs
bash
"$SKILL_DIR/.venv/bin/python" "$SKILL_DIR/scripts/00_build_video_notes.py" \
  --stage finalize --project-dir "$PROJECT_DIR"

Finalize automatically writes verify/quality-audit.json and verify/quality-audit.md. Display verify/pdf-pages-contact-sheet.jpg, read the audit, and resolve warnings about transcript coverage, note/frame counts, speaker identity, duplicate exact frames, or missing usage records. Report Markdown, HTML, PDF, ZIP, scene count, image count, page count, recorded tokens, and known cost. Unknown usage remains unknown.

Fallbacks

After the compliance gate is satisfied, try anonymous yt-dlp, then local Chrome browser cookies without exporting them. Do not use the cookie fallback to defeat access controls. Volcengine credentials may be discovered from the environment, scripts/config.py, or WATCHLESS_VOLCENGINE_CONFIG; never print or copy credential values. Use source/manual subtitles only with --use-source-subtitles. Use local Whisper only when explicitly requested with --provider whisper. If acquisition still fails, report the classified error and request an authorized local video.

Stages are idempotent and resumable. --target-seconds is an explicit legacy fallback for non-slide modes, never the normal segmentation method.

Tencent ASR / 腾讯云通道

See references/tencent-asr.md for credentials, engine selection, resumable jobs and cross-chunk speaker-label limitations.

© chenzixin1, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 59 other files (scripts, references, assets) in the repository root of chenzixin1/watchless.

  • SKILL.md
  • .gitignore
  • LEGAL.en.md
  • LEGAL.md
  • LICENSE
  • README.en.md
  • README.md
  • SKILL.zh-CN.md
  • THIRD_PARTY_NOTICES.md
  • assets/examples/conversation.en.png
  • assets/examples/conversation.jpg
  • assets/examples/demo.en.png
  • assets/examples/demo.jpg
  • assets/examples/explainer.en.png
  • assets/examples/explainer.jpg
  • assets/examples/news-interview.en.png
  • assets/examples/news-interview.jpg
  • assets/watchless-hero.png
  • assets/watchless-workflow.en.svg
  • … and 41 more

Open the folder on GitHubat commit 34e2fa8

Compare with similar skills

Watchless next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Watchless compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Watchless this skillchenzixin1/watchless144—~3.7kAutomated safety check: WarnMIT
Yt Dlp DownloaderMapleShaw/yt-dlp-downloader-skill2091 repos~1.5kAutomated safety check: PassNone
Clean Cuthassancs91/claude-youtube-editor322—~4.3kAutomated safety check: NotesMIT
Video Downloadcalesthio/OpenMontage65k—~885Automated safety check: PassAGPL-3.0
Videohub Youtubecacity/VideoHub167—~550Automated safety check: PassMIT
AI MultimodalMicrock/ordinary-claude-skills4031 repos~2.7kAutomated safety check: NotesMIT

Similar skills

  • Yt Dlp Downloader

    MapleShaw/yt-dlp-downloader-skill

    Download videos from YouTube, Bilibili, Twitter, and thousands of other sites using yt-dlp.

    209 GitHub starsUsed in 1 repo~1.5k tokens
    Media & CreativeAuto-check passed
  • Clean Cut

    hassancs91/claude-youtube-editor

    Step 1 of the AI Video Editor pipeline — turn raw talking-head footage into a clean master cut.

    322 GitHub stars~4.3k tokensUpdated 1 mo ago
    Media & CreativeAuto-check: notes
  • Video Download

    calesthio/OpenMontage

    Download video and audio from YouTube and 1000+ sites using yt-dlp.

    65k GitHub stars~885 tokensUpdated 5 days ago
    Media & CreativeAuto-check passed
  • Videohub Youtube

    cacity/VideoHub

    处理 YouTube、Twitter(X)、Bilibili 和本地音视频/文本的转写、字幕、翻译与总结。优先复用 src/youtubetranscriber.py 现有 CLI。

    167 GitHub stars~550 tokensUpdated 6 days ago
    Media & CreativeAuto-check passed
  • AI Multimodal

    Microck/ordinary-claude-skills

    Process and generate multimedia content using Google Gemini API.

    403 GitHub starsUsed in 1 repo~2.7k tokens
    Media & CreativeAuto-check: notes
  • Video Transcript Downloader

    sundial-org/awesome-openclaw-skills

    Download videos, audio, subtitles, and clean paragraph-style transcripts from YouTube and any other yt-dlp supported site.

    663 GitHub starsUsed in 2 repos~574 tokens
    Media & CreativeAuto-check passed

Works with

Questions about Watchless

What does Watchless do?

A skill your agent uses when turning a YouTube URL or local presentation, explainer, interview, podcast, or product-demo video into complete screenshot-led notes, faithful light-polished text, HTML…. Watchless is an agent skill from chenzixin1/watchless. Use when turning a YouTube URL or local presentation, explainer, interview, podcast, or product-demo video into complete screenshot-led notes, faithful light-polished text, HTML, PDF, or a shareable ZIP.

When should I use Watchless?

Watchless fits situations like: turning a YouTube URL; local presentation; product-demo video into complete screenshot-led notes; faithful light-polished text.

How do I install Watchless in Claude Code?

Run `npx skills add chenzixin1/watchless --skill watchless -a claude-code`. Or copy the skill folder (the chenzixin1/watchless repository) into .claude/skills/watchless in your project. Claude Code loads it when a task matches its description.

How do I install Watchless in Codex?

Run `npx skills add chenzixin1/watchless --skill watchless -a codex`. Or copy the skill folder (the chenzixin1/watchless repository) into .agents/skills/watchless in your project. Codex loads it when a task matches its description.

Can I use Watchless in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add chenzixin1/watchless --skill watchless -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/watchless, .gemini/skills/watchless, .github/skills/watchless and .opencode/skills/watchless in your project.

What does Watchless need to run?

SKILL.md names no scripts, command-line tools or credentials: Watchless is instructions for the agent only.

Does Watchless access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Watchless safe to install?

Our automated static check of SKILL.md flagged 2 warning(s): mentions a credentials file (ssh keys, cloud or package-manager tokens). Read the flagged lines before installing; the check is not a guarantee either way. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Watchless use?

Watchless is published under the MIT licence (from the LICENSE file in the skill folder). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Watchless use?

About 3.7k tokens (SKILL.md is roughly 15k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 495 tokens, read only when the agent opens those files.

What are the alternatives to Watchless?

Skills that share tags, products or a category with Watchless: Yt Dlp Downloader (MapleShaw/yt-dlp-downloader-skill, 209 stars), Clean Cut (hassancs91/claude-youtube-editor, 322 stars), Video Download (calesthio/OpenMontage, 65k stars) and Videohub Youtube (cacity/VideoHub, 167 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Watchless?

chenzixin1 (a GitHub user) maintains it in chenzixin1/watchless, which has 144 GitHub stars. The repository was last updated on September 16, 2026.

Source: chenzixin1/watchless on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.