Video To Notes
KIRVO-REPORTING/video-to-notes
Use immediately for any bare YouTube or YouTube Shorts URL, youtu.be link, Bilibili or b23.tv link, or other video URL; do not ask what the user wants.
Pulls text out of videos and podcasts by sending the media to a Quark cloud-drive skill for transcription, then proofreads the result lightly.
SKILL.md written in Chinese; this summary is our English description.
$ npx skills add Backtthefuture/video-transcript --skill video-transcript -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install Backtthefuture/video-transcript video-transcript --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
Claude Code skills documentation · loads skills from .claude/skills/
Install the "video-transcript" agent skill from https://github.com/Backtthefuture/video-transcript/tree/main into .claude/skills/video-transcript/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "video-transcript", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Backtthefuture/video-transcript --skill video-transcript -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install Backtthefuture/video-transcript video-transcript --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "video-transcript" agent skill from https://github.com/Backtthefuture/video-transcript/tree/main into .agents/skills/video-transcript/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "video-transcript", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Backtthefuture/video-transcript --skill video-transcript -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install Backtthefuture/video-transcript video-transcript --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "video-transcript" agent skill from https://github.com/Backtthefuture/video-transcript/tree/main into .cursor/skills/video-transcript/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "video-transcript", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Backtthefuture/video-transcript --skill video-transcript -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install Backtthefuture/video-transcript video-transcript --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "video-transcript" agent skill from https://github.com/Backtthefuture/video-transcript/tree/main into .gemini/skills/video-transcript/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "video-transcript", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install Backtthefuture/video-transcript video-transcriptInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add Backtthefuture/video-transcript --skill video-transcript -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "video-transcript" agent skill from https://github.com/Backtthefuture/video-transcript/tree/main into .github/skills/video-transcript/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "video-transcript", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Backtthefuture/video-transcript --skill video-transcript -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install Backtthefuture/video-transcript video-transcript --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "video-transcript" agent skill from https://github.com/Backtthefuture/video-transcript/tree/main into .opencode/skills/video-transcript/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "video-transcript", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
video-transcriptPulls text out of videos and podcasts by sending the media to a Quark cloud-drive skill for transcription, then proofreads the result lightly.
The SKILL.md is written in Chinese. It takes a video or podcast link, a local media file or a file already on the Quark cloud drive and returns text. Each item is registered, downloaded or matched, uploaded and checked, then the Quark skill is asked for the text of one file at a time. After quality checks, the model corrects the text by writing a small patch file instead of rewriting the transcript.
Single items and batches are supported. Every file keeps a fixed asset ID, so a batch can be resumed after an interruption with `transcript.py --resume`, and statuses such as text_received, waiting_analysis and auth_required show where each item stands. Returned text counts as reviewable, not as verbatim or verified, and suspicious results such as summaries, truncated or very short text are flagged for a manual read of the raw output.
No local speech recognition runs by default, and the skill makes no promise of speaker separation, sentence-level timestamps or SRT files. Waiting items are rechecked after 60 seconds at first and then at longer intervals; nothing wakes the task automatically after it ends unless you ask for background monitoring.
Read from SKILL.md and the folder at commit 49f2dca. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 3 files in scripts/ (Python and Shell, from the files we listed), which the agent can run.
Shell commands in SKILL.md call:
python3From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Video Transcript Extractor loads about 970 tokens when it runs, and up to ~2.1k if it reads all its reference files. Until then it costs about 38 tokens; SKILL.md has 164 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check noted patterns worth knowing about, such as sudo or a known installer.
安装步骤。需要安装下载依赖才运行本 Skill 的 install.sh;已有 .env、outputs 和模型缓存保留。Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from Backtthefuture/video-transcript at commit 49f2dca, republished under its MIT licence (© Backtthefuture). 164 words, ~970 tokens.
.claude/skills/video-transcript/SKILL.md (or your agent's skills folder). This skill also uses 29 other files; get the full folder from GitHub.链接 / 本地媒体 / 已指定的网盘文件 → 登记 → 下载或匹配 → 上传并核对 → 单文件问答取正文 → 质量检查 → 模型 patch 轻校对。
使用本次读取的这份 SKILL.md 所在目录作为 VT_HOME。不要按 WorkBuddy/Codex 等优先级另找一份 video-transcript;不能把其他副本当成这份的运行目录。
VT_PY 使用可运行脚本的 Python 3.9+,不必寻找装有 FunASR 的解释器。
首次执行先读已安装 quarkclouddrive/SKILL.md 及本次涉及的 references/file-upload.md、references/file-search.md、references/assistant.md。
下载沿用 video-download;平台解析所需登录态仍按该 Skill 处理。夸克 CLI 是外部依赖,不复制其源码或凭据。本脚本会在每次调用夸克命令前运行官方环境检查,并使用同一批次会话 ID。
python3 "$VT_HOME/scripts/transcript.py" --doctor该体检只检查本地工具与依赖位置,不代表授权或云端链路成功。缺少夸克 Skill 时按其安装说明安装;不要运行旧版 ASR 安装步骤。需要安装下载依赖才运行本 Skill 的 install.sh;已有 .env、outputs 和模型缓存保留。
用户明确选择此 Skill 的云端工作流,即按任务范围上传指定媒体。用户要求只下载时转交 video-download。
保存用户本轮原始请求为临时 UTF-8 文件,通过 --session-input-file 传入;这是夸克要求的服务质量追踪参数,不应塞入额外对话或凭据。
python3 "$VT_HOME/scripts/transcript.py" "<链接或本地路径>" --session-input-file "<原始请求文件>" --wait 1800
python3 "$VT_HOME/scripts/transcript.py" --batch "<inputs.json>" --session-input-file "<原始请求文件>" --wait 1800等待值是本次进程的复查预算,不是预计完成时间;单次网络请求可能在预算结束后才返回。
批次 JSON 是输入字符串数组,或包含 input、可选 title 的对象数组。
截图 PDF 保留 --keep-video 兼容参数;新版媒体默认持久保存,不自动删本地文件或云端文件。
已上传网盘文件的选定、参数、缓存与错误恢复见 references/quark-workflow.md。
每条文件有固定 asset_id;所有批次复用同一 output-dir 内的条目记录。通过来源链接或本地内容 SHA256 识别任务,不靠标题和完成顺序关联。 上传原始 ID 与搜索返回 FID 分开存。读取完整 Artifact,并按唯一文件名、大小及上传时间(时长返回时另核对)绑定;匹配不唯一就停该条。 一条问答只传一个文件,原始响应与每次请求单独留档。上传结果未知不盲目重传。
text_received:取回了可供审阅的正文,不是逐字准确或全文已验证。waiting_analysis:保持 FID,按 next_poll_at 复查,初期 60 秒,随后 120/300 秒。response_needs_review:疑似摘要、缺来源、过短或不完整。读 raw.md,不直接润色成“完整稿”。mapping_needs_review / upload_outcome_unknown:核对对应关系,不猜、不重新上传。auth_required:按夸克 Skill 引导 login,完成后用原批次明确重试。qa_failed / failed:说明原因,修复后恢复,不自动无限重试。python3 "$VT_HOME/scripts/transcript.py" --resume "<batch_manifest>" --status
python3 "$VT_HOME/scripts/transcript.py" --resume "<batch_manifest>" --session-input-file "<原始请求文件>" --wait 1800退出码 0=本次各条已取回正文;2=仍有等待;1=需要处理错误或人工核对。退出码不代表质量验收。 任务终止后没有自动唤醒;只在用户请求后台监测时另接调度。不能用“已保存等待状态”冒充会自动继续。
输出包含 batch_manifest、各条状态,以及可用的 transcript_path、preorganized_path、polish_brief_path。 先读完整原始返回与预整理稿,核对来源是否对应、是否摘要/拒绝/明显截断,条件允许时抽查首中尾字幕或听校。字数和首尾覆盖只能作为抽查,不能证明逐句完整。
普通视频:读 polish_brief,写一份 patch.json,再调用:
python3 "$VT_HOME/scripts/make_optimized.py" --from-md "<preorganized_path>" --patch "<patch.json>" --filename "<asset_id>_整理稿" --output-dir "<该条产物目录>"patch 使用 title、headings、fixes(from/to/confidence/basis)和必要的 paragraph_edits。保留顺序、数字、观点;高确信且有依据才替换,低确信只列待核。不要补写缺段、推测说话人或编造时间轴。渲染器支持无时间标记的章节,并保留补丁记录。 修改后回读,核对没有整段丢失。保留夸克原始返回,润色稿独立交付。单条用户要求全文时在对话输出全文;批次优先交付汇总表与每条稿件链接,避免四份长稿互相混杂。
微信视频号:读 skills/weixin-layout.md,保留默认口语稿、文字 PDF、截图 PDF 的分流。无时间轴时截图与正文需实际核对,不得按虚构时间自动配图。 播客:同样先取纯文字,再轻校对。旧 --speakers/--host/--guest 明确报不支持,不转本地模型、不伪造人物归属。
分别报告:文件对应、正文取回、抽查或听校程度、润色完成。历史专名、古文和无字幕材料要保留待核点。 --force 只重新请求正文,复用原上传;--reformat 只从已保存原文重建预整理。不自动读取旧 FunASR 缓存,旧文件不删除。
供 deep-search、视频博主拆解使用的适配器为 scripts/caller_client.py,不复制夸克运行时。先通过 --batch ... --status 本地登记,再持久保存 asset_id/manifest 后开始云端请求;适配器按完整 stdout JSON 和条目 state.json 读取结果,schema_version=1。
同一调用方工作目录保存 session.json;--session-id 可在新批次登记时指定 {timestamp}-{6位随机字符},已有批次继续使用原值。所有实际云端调用仍传原始请求文件。
适配器保留 waiting_analysis、response_needs_review、mapping_needs_review、upload_outcome_unknown、auth_required 等状态。无音轨标 no_audio,媒体保留且不上传/问答;缺衍生稿可本地重建。取帧只接受记录中且 SHA256 相符的媒体,未核对的目录文件不能作为替代。
调用方显式使用 --repair-media 时,适配器仅补下载缺失媒体,要求与原哈希一致;不重新上传或问答。字节不同保留候选并停下核对。没有可核对的原链接或哈希时不能自动修复。
© Backtthefuture, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 29 other files (scripts, references) in the repository root of Backtthefuture/video-transcript.
Open the folder on GitHubat commit 49f2dca
Video Transcript Extractor next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Video Transcript Extractor this skillBacktthefuture/video-transcript | 115 | — | ~970 | Automated safety check: Notes | MIT | |
| Video To NotesKIRVO-REPORTING/video-to-notes | 105 | — | ~1.5k | Automated safety check: Pass | MIT | |
| Baoyu Youtube TranscriptJimLiu/baoyu-skills | 26k | 1 repos | ~2.4k | Automated safety check: Pass | MIT | |
| Video Link Transcript ExtractorSpaceZephyr/creator-buddy | 1.6k | — | ~498 | Automated safety check: Pass | None | |
| Watch Videocoreyhaines31/makerskills | 844 | — | ~3.7k | Automated safety check: Pass | MIT | |
| Video Transcriptnotque/vexjoy-agent | 435 | — | ~576 | Automated safety check: Notes | MIT |
KIRVO-REPORTING/video-to-notes
Use immediately for any bare YouTube or YouTube Shorts URL, youtu.be link, Bilibili or b23.tv link, or other video URL; do not ask what the user wants.
JimLiu/baoyu-skills
Downloads YouTube video transcripts/subtitles and cover images by URL or video ID.
SpaceZephyr/creator-buddy
Extracts subtitles or a full transcript from YouTube, Xiaoyuzhou, Bilibili, Douyin and Xiaohongshu links, using speech recognition when a video has no subtitles.
coreyhaines31/makerskills
When you want to extract content from a video — YouTube, Loom, Vimeo, Riverside, Zoom recording, local MP4, X/IG video, anything yt-dlp supports.
notque/vexjoy-agent
Extract video transcripts: yt-dlp subtitles to clean paragraphs.
bozhouDev/video-skills-toolkit
Deprecated compatibility package. An agent skill from bozhouDev/video-skills-toolkit.
Categories
Pulls text out of videos and podcasts by sending the media to a Quark cloud-drive skill for transcription, then proofreads the result lightly. md is written in Chinese. It takes a video or podcast link, a local media file or a file already on the Quark cloud drive and returns text.
Video Transcript Extractor fits situations like: turning a video or podcast link into a text transcript; extracting copy from a batch of recordings and resuming after an interruption; getting text from media files already uploaded to the Quark cloud drive.
Run `npx skills add Backtthefuture/video-transcript --skill video-transcript -a claude-code`. Or copy the skill folder (the Backtthefuture/video-transcript repository) into .claude/skills/video-transcript in your project. Claude Code loads it when a task matches its description.
Run `npx skills add Backtthefuture/video-transcript --skill video-transcript -a codex`. Or copy the skill folder (the Backtthefuture/video-transcript repository) into .agents/skills/video-transcript in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Backtthefuture/video-transcript --skill video-transcript -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/video-transcript, .gemini/skills/video-transcript, .github/skills/video-transcript and .opencode/skills/video-transcript in your project.
Going by SKILL.md and its folder, Video Transcript Extractor needs Python and a shell for the scripts in its folder and the command-line tools its instructions call (python3). Our summary lists: Python 3.9 or newer; The Quark cloud drive skill (quarkclouddrive), installed and logged in; The video-download skill for fetching media from links.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found notes only (mentions a .env file), nothing it rates as a warning. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Video Transcript Extractor is published under the MIT licence (from the LICENSE file in the skill folder). It allows redistribution, so the full SKILL.md is shown on this page.
About 970 tokens (SKILL.md is roughly 3.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.1k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Video Transcript Extractor: Video To Notes (KIRVO-REPORTING/video-to-notes, 105 stars), Baoyu Youtube Transcript (JimLiu/baoyu-skills, 26k stars), Video Link Transcript Extractor (SpaceZephyr/creator-buddy, 1.6k stars) and Watch Video (coreyhaines31/makerskills, 844 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
Backtthefuture (a GitHub user) maintains it in Backtthefuture/video-transcript, which has 115 GitHub stars. The repository was last updated on September 29, 2026.
Source: Backtthefuture/video-transcript on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.