Agent skill

Video Transcribe

by wendy7756 in wendy7756/AI-Video-Transcriber

Transcribe and summarize a video or podcast from a URL (YouTube, TikTok, Bilibili, Apple Podcasts, SoundCloud, 30+ platforms) or from a local media/.txt file.

Apache-2.0Auto-check: notesMedia & Creative

Install Video Transcribe

skills CLI
$ npx skills add wendy7756/AI-Video-Transcriber --skill video-transcribe -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install wendy7756/AI-Video-Transcriber video-transcribe --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/wendy7756/AI-Video-Transcriber.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/video-transcribe .claude/skills/video-transcribe && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
video-transcribe
GitHub stars
3.3k
Token cost
~937 tokens
SKILL.md length
397 words
Files
1
Skills in repo
1
Repo updated
First seen
Licence
Apache-2.0

At a glance

Transcribe and summarize a video or podcast from a URL (YouTube, TikTok, Bilibili, Apple Podcasts, SoundCloud, 30+ platforms) or from a local media/.txt file.

  • The user shares a video/podcast link
  • SKILL.md covers Before running, Run it, Reading the result and Critical: no_speech, plus 1 more section
  • Calls python, python3 and pip; reaches openrouter.ai; needs OPENAI_API_KEY
  • Media file and wants a transcript

What it does

Video Transcribe is an agent skill from wendy7756/AI-Video-Transcriber. Transcribe and summarize a video or podcast from a URL (YouTube, TikTok, Bilibili, Apple Podcasts, SoundCloud, 30+ platforms) or from a local media/.txt file. Use when the user shares a video/podcast link or media file and wants a transcript, summary, translation, or the original video downloaded. Triggers on "transcribe", "summarize this video", "what does this video say", "get the transcript", "转录", "视频摘要".

Its SKILL.md is about 940 tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Media & Creative, covering Transcription and Podcasting. It works with TikTok, YouTube, Bilibili and OpenAI. The repository describes itself as: Transcribe and summarize videos and podcasts using AI. Open-source, multi-platform, and supports multiple languages. The licence is Apache-2.0.

When your agent uses it

  • The user shares a video/podcast link
  • Media file and wants a transcript
  • The original video downloaded
  • Summarize this video

Example prompts

  • “transcribe”
  • “summarize this video”
  • “what does this video say”
  • “/video-transcribe”

Requirements

  • Python 3
  • A credential in OPENAI_API_KEY

What it can do on your machine

Read from SKILL.md and the folder at commit 634e6a2. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python
    • python3
    • pip
    • brew
    • apt

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • openrouter.ai

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • OPENAI_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Video Transcribe loads about 937 tokens when it runs. Until then it costs about 107 tokens; SKILL.md has 397 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~107
When it runs · the whole SKILL.md, loaded when a task matches
~937

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteRuns commands with sudoSKILL.md:21
    with `brew install ffmpeg` on macOS or `sudo apt install ffmpeg` on Debian/Ubuntu. Do not proceed without it; audio ext

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from wendy7756/AI-Video-Transcriber at commit 634e6a2, republished under its Apache-2.0 licence (© wendy7756). 397 words, ~937 tokens.

Download SKILL.mdSave it as .claude/skills/video-transcribe/SKILL.md (or your agent's skills folder).
name
video-transcribe
description
Transcribe and summarize a video or podcast from a URL (YouTube, TikTok, Bilibili, Apple Podcasts, SoundCloud, 30+ platforms) or from a local media/.txt file. Use when the user shares a video/podcast link or media file and wants a transcript, summary, translation, or the original video downloaded. Triggers on "transcribe", "summarize this video", "what does this video say", "get the transcript", "转录", "视频摘要".

Video Transcribe

Runs this repo's pipeline headlessly: platform subtitles when available (seconds), local Whisper as fallback, then optimize -> translate -> summarize.

Before running

The command must be run from the AI Video Transcriber repository root and needs the project virtualenv plus ffmpeg. Verify once per session:

bash
ls venv/bin/python && command -v ffmpeg
  • No venv: run ./install.sh, or python3 -m venv venv && venv/bin/pip install -r requirements.txt.
  • No ffmpeg: install it first with brew install ffmpeg on macOS or sudo apt install ffmpeg on Debian/Ubuntu. Do not proceed without it; audio extraction will fail.

Run it

bash
venv/bin/python transcribe.py "<URL or file path>" --json

Always pass --json: it puts machine-readable output on stdout and keeps progress chatter on stderr. Parse the JSON rather than scraping log lines.

Useful flags:

FlagWhen to use
-l, --summary-language <code>Summary language: en, zh, es, fr, de, it, pt, ru, ja, ko, ar. Default en
--no-llmNo API key is available, or the user only wants the raw transcript. Skips optimize/translate/summarize
--no-videoThe user does not want the source video kept
--whisper-model smallAccuracy matters more than speed. tiny to large, default base
-o <dir>Write the Markdown somewhere other than ./temp

The LLM steps need an OpenAI-compatible key via OPENAI_API_KEY and optionally OPENAI_BASE_URL. Without one the pipeline still transcribes but falls back to basic formatting; prefer --no-llm when the user only needs the transcript.

Provider settings:

bash
export OPENAI_API_KEY="your_api_key_here"
export OPENAI_BASE_URL="https://openrouter.ai/api/v1"
export OPENAI_TRANSLATION_MODEL="gpt-4o"  # optional

For a one-off run, transcribe.py also accepts --api-key, --base-url, and --model. Prefer environment variables when possible so API keys are not written to shell history.

Show full SKILL.md (156 more words)Show less

Reading the result

json
{
  "title": "...",
  "no_speech": false,
  "detected_language": "en",
  "files": { "raw": "...", "transcript": "...", "summary": "...", "translation": "..." },
  "media": { "path": "...", "kind": "video", "size_bytes": 2243691 }
}

Read the transcript and summary paths to get the content. translation is present only when the source language differs from the summary language.

Critical: no_speech

When "no_speech": true the source contains no speech. There is no transcript, summary, or translation. The pipeline deliberately skips the LLM so that nothing gets invented.

If you see this, tell the user the video has no speech. Do not guess at the content, infer it from the title, or describe what the video probably says.

Failure modes

  • Exit 2: bad input, such as a missing file, unsupported extension, or empty text file. Report the message and fix the argument.
  • Exit 1: download/transcode failure, usually an unreachable URL, region-blocked/private video, or missing ffmpeg. Report the actual error.
  • First run downloads Whisper weights, so it can look stalled for a minute.
  • Long videos are slow in Whisper mode. Warn the user before starting if no platform subtitles are available.

© wendy7756, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/video-transcribe of wendy7756/AI-Video-Transcriber.

Open the folder on GitHubat commit 634e6a2

Compare with similar skills

Video Transcribe next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Video Transcribe compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Video Transcribe this skillwendy7756/AI-Video-Transcriber3.3k—~937Automated safety check: NotesApache-2.0
BibiJimmyLv/BibiGPT-v16.2k—~885Automated safety check: PassGPL-3.0
Ffmpeg Skillkajisho5/ffmpeg-skill1.9k—~7.4kAutomated safety check: PassMIT
Video Clip Extractorlinzzzzzz/openclip569—~2.8kAutomated safety check: WarnMIT
Watchmathiaschu/watch142—~4kAutomated safety check: WarnMIT
VideoiBigQiang/feedgrab614—~1.6kAutomated safety check: PassMIT

Similar skills

  • Bibi

    JimmyLv/BibiGPT-v1

    BibiGPT CLI for summarizing videos, audio, and podcasts directly in the terminal.

    6.2k GitHub stars~885 tokensUpdated 5 mo ago
    Media & CreativeAuto-check passed
  • Ffmpeg Skill

    kajisho5/ffmpeg-skill

    Edit video and audio with local FFmpeg from natural-language requests: cut, trim, join, resize/reframe (9:16, 1:1), speed change, captions and subtitles (SRT/ASS, animated, karaoke), logos and text…

    1.9k GitHub stars~7.4k tokensUpdated yesterday
    Media & CreativeAuto-check passed
  • Video Clip Extractor

    linzzzzzz/openclip

    Processes videos to identify engaging moments, generate transcripts, and create highlight clips with artistic titles and custom cover images.

    569 GitHub stars~2.8k tokensUpdated 1 mo ago
    Media & CreativeAuto-check: warnings
  • Watch

    mathiaschu/watch

    Watch a video from YouTube, Instagram, X/Twitter, Vimeo, TikTok or any of ~1800 yt-dlp sites (or a local path).

    142 GitHub stars~4k tokensUpdated 4 mo ago
    Media & CreativeAuto-check: warnings
  • Video

    iBigQiang/feedgrab

    Video & Podcast Digest — send a video/podcast link, get full transcript + structured summary.

    614 GitHub stars~1.6k tokensUpdated 1 mo ago
    Media & CreativeAuto-check passed
  • Video Clipper

    gooseworks-ai/goose-skills

    Repurposes long-form video (podcasts, interviews, talks) into short-form vertical clips for Instagram Reels, TikTok, and YouTube Shorts.

    1.2k GitHub starsUsed in 1 repo~3.1k tokens
    Media & CreativeAuto-check: notes

Questions about Video Transcribe

What does Video Transcribe do?

Transcribe and summarize a video or podcast from a URL (YouTube, TikTok, Bilibili, Apple Podcasts, SoundCloud, 30+ platforms) or from a local media/.txt file. Video Transcribe is an agent skill from wendy7756/AI-Video-Transcriber.txt file.

When should I use Video Transcribe?

Video Transcribe fits situations like: the user shares a video/podcast link; media file and wants a transcript; the original video downloaded; summarize this video.

How do I install Video Transcribe in Claude Code?

Run `npx skills add wendy7756/AI-Video-Transcriber --skill video-transcribe -a claude-code`. Or copy the skill folder (skills/video-transcribe in wendy7756/AI-Video-Transcriber) into .claude/skills/video-transcribe in your project. Claude Code loads it when a task matches its description.

How do I install Video Transcribe in Codex?

Run `npx skills add wendy7756/AI-Video-Transcriber --skill video-transcribe -a codex`. Or copy the skill folder (skills/video-transcribe in wendy7756/AI-Video-Transcriber) into .agents/skills/video-transcribe in your project. Codex loads it when a task matches its description.

Can I use Video Transcribe in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add wendy7756/AI-Video-Transcriber --skill video-transcribe -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/video-transcribe, .gemini/skills/video-transcribe, .github/skills/video-transcribe and .opencode/skills/video-transcribe in your project.

What does Video Transcribe need to run?

Going by SKILL.md and its folder, Video Transcribe needs the command-line tools its instructions call (python, python3, pip, brew and apt) and credentials named OPENAI_API_KEY. Our summary lists: Python 3; A credential in OPENAI_API_KEY.

Does Video Transcribe access the network?

SKILL.md names 1 domain. In commands or code: openrouter.ai; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is Video Transcribe safe to install?

Our automated static check of SKILL.md found notes only (runs commands with sudo), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Video Transcribe use?

Video Transcribe is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Video Transcribe use?

About 937 tokens (SKILL.md is roughly 3.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Video Transcribe?

Skills that share tags, products or a category with Video Transcribe: Bibi (JimmyLv/BibiGPT-v1, 6.2k stars), Ffmpeg Skill (kajisho5/ffmpeg-skill, 1.9k stars), Video Clip Extractor (linzzzzzz/openclip, 569 stars) and Watch (mathiaschu/watch, 142 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Video Transcribe?

wendy7756 (a GitHub user) maintains it in wendy7756/AI-Video-Transcriber, which has 3,334 GitHub stars. The repository was last updated on September 15, 2026.

Source: wendy7756/AI-Video-Transcriber on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.