Agent skill

Video Reference

by zenstory-ai in zenstory-ai/video-recap-skills

按需把一部成片拆成可复用的制作参考:测镜头节奏与响度,标注段落与音轨分工,把原片事实与可迁移方法分开, 导出不含原片人名台词的 productionreference.json 供下次制作参考。不在默认生产路径上。

MITAuto-check passedMedia & Creative

Install Video Reference

skills CLI
$ npx skills add zenstory-ai/video-recap-skills --skill video-reference -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install zenstory-ai/video-recap-skills video-reference --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/zenstory-ai/video-recap-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/video-reference .claude/skills/video-reference && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
video-reference
GitHub stars
561
Token cost
~1.3k tokens
SKILL.md length
352 words
Files
8 (incl. scripts, references)
Skills in repo
7
Repo updated
First seen
Licence
MIT

At a glance

按需把一部成片拆成可复用的制作参考:测镜头节奏与响度,标注段落与音轨分工,把原片事实与可迁移方法分开, 导出不含原片人名台词的 productionreference.json 供下次制作参考。不在默认生产路径上。

  • Works in 6 steps: 定位 → 输入 → 流程 → …
  • Tasks that involve Text to speech and voice
  • SKILL.md covers 1. 定位, 2. 输入, 3. 流程 and 4. 分离规则, plus 2 more sections
  • Runs Python scripts from its folder; calls python3

What it does

Video Reference is an agent skill from zenstory-ai/video-recap-skills. 按需把一部成片拆成可复用的制作参考:测镜头节奏与响度,标注段落与音轨分工,把原片事实与可迁移方法分开, 导出不含原片人名台词的 productionreference.json 供下次制作参考。不在默认生产路径上。 触发词:拆片、拆解成片、制作参考、参考模板、production reference。

Its SKILL.md is about 1.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 9 other files, including scripts and reference files (for example `references/reference-schema.md`, `scripts/lib.py` and `scripts/reference.py`).

It sits in Media & Creative, covering Text to speech and voice and Speech recognition and synthesis. The repository describes itself as: Claude Code / Codex skills that turn a video into a Chinese narration recap (视频解说): scene detection, ASR, VLM, script, TTS, ffmpeg assembly, optional editable JianYing / CapCut… The licence is MIT.

When your agent uses it

  • Tasks that involve Text to speech and voice
  • Tasks that involve Speech recognition and synthesis

Example prompts

  • “/video-reference”

Requirements

  • Python 3

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. 定位
  2. 输入
  3. 流程
  4. 分离规则
  5. 交付与使用
  6. 能力边界

What it can do on your machine

Read from SKILL.md and the folder at commit 5391686. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 6 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Video Reference loads about 1.3k tokens when it runs, and up to ~3.6k if it reads all its reference files. Until then it costs about 42 tokens; SKILL.md has 352 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~42
When it runs · the whole SKILL.md, loaded when a task matches
~1.3k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~3.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from zenstory-ai/video-recap-skills at commit 5391686, republished under its MIT licence (© zenstory-ai). 352 words, ~1,325 tokens.

Download SKILL.mdSave it as .claude/skills/video-reference/SKILL.md (or your agent's skills folder). This skill also uses 7 other files; get the full folder from GitHub.
name
video-reference
description
按需把一部成片拆成可复用的制作参考:测镜头节奏与响度,标注段落与音轨分工,把原片事实与可迁移方法分开, 导出不含原片人名台词的 production_reference.json 供下次制作参考。不在默认生产路径上。 触发词:拆片、拆解成片、制作参考、参考模板、production reference。
user-invocable
true

1. 定位

本技能把一部已完成的成片拆成下一次制作能直接参考的方法与数值。它是按需的参考分析,不是生产阶段,也不是质检: 不调用 MiMo,不改其他产物,不给任何运行打分或拦截。

Agent 的角色是拆片编辑:先如实记录"这部片子怎么做的"(事实),再提炼"换一部素材还能怎么用"(方法)。 两者必须分开:事实只留在本地工作目录,导出物只含方法和测得的数值。

2. 输入

  • 成片文件(.mp4 / .mov / .mkv / .webm)。
  • 对这部成片跑出的视频理解产物目录 U(建议 ASR_SEGMENT_SECONDS=5,窗口越短,旁白语速越准)。 本技能读取其中可选的 asr_result.json、asr_timing_evidence.json、background_research.json、understanding_index.json (其 characters 的名字、别名与 ASR 提及都进入人名泄漏扫描); 标注时 Agent 还应看故事板 / contact sheet 与 vlm_analysis.json。

没有理解产物也能跑,但旁白语速为空,泄漏扫描只剩 Agent 自己写的 entities,check 会给出警告; 看画面改用 frames --span 0,<时长> --step 2 的接触表。

下面的 scripts/... 均相对于本技能目录。

3. 流程

bash
python3 scripts/reference.py measure <成片> --work-dir U        # 一次 ffmpeg:切点 + 响度,按文件身份缓存
python3 scripts/reference.py frames  <成片> --work-dir U --review          # 逐帧看被压下的疑似切点
python3 scripts/reference.py frames  <成片> --work-dir U --longest 5       # 看最长镜头里有没有漏切
python3 scripts/reference.py check   --work-dir U [--json]      # 校验并打印派生值
python3 scripts/reference.py export  --work-dir U --out <下次运行的 work_dir>/production_reference.json
  1. measure 写 U/reference_measurements.json(只由脚本写):镜头切点、镜长分布、每分钟切点数、10 秒切点曲线、 整体响度 / LRA / 真峰值、逐秒短期响度。它要解码整片:5 分钟 720p 约 15–70 秒,4K 约 7 分钟;成片不变时复用, 只改 --soft-score/--hard-score 不重新解码。切点规则:scdet 分数 ≥10 必算;≥4 且是前后 0.3 秒内其他帧(相邻帧除外) 两倍以上的孤立峰才算。运动镜头、急推拉、闪光会被压下,并列进 shots.review_windows;scdet 分数会减去前一帧的帧差, 从快速运动切进静止镜头的硬切可能只有 1 分,所以帧差本身的单侧峰(mafd_peaks)也进待复核窗口,但不会自动算切点。
  2. 复核切点:固定阈值在真实成片上两头都错(暗场硬切只有 5–7 分;快速运动的单个镜头每 0.2 秒一个 7–8 分的峰), 所以每次都要看图。frames --review 逐帧拼出每个待复核窗口,frames --longest 5 在最长的 5 个镜头里均匀取 12 帧; 页面在 U/reference_frames/,命令打印每格对应的秒数。看完写 labels.cut_fixes:add 漏掉的切点秒数, remove 误报的切点秒数(±0.1 秒内对上测得的切点);看过无需改动就写 {}。重新 measure 后待复核窗口数变了,旧的 cut_fixes 不再算数,按新窗口重看一遍。之后所有镜头数值、段内切点密度和导出都用复核后的切点。
  3. 标注:Agent 写 U/reference_breakdown.json 的 labels——音轨归属 audio_spans、叙事段落 sections、 字幕形态 subtitles、标注依据 basis。audio_spans 的边界放在声音实际起止处(听得到的人声起点与止点), 不放在字幕或旁白块的开始处(按字幕出现帧定的旁白结束点实测晚 0.24–0.42 秒)。U/speech_boundary_anchors.json 存在时, 用它的 acoustic_pauses(start 是人声停下处,end 是下一句开口处)对齐边界;switch_on_cut_share 只容差 ±0.25 秒, 边界放错它就只反映标注习惯。原片自带的画外音、内心独白也是原片音轨,记 original_dialogue;要区分时写进 fact 和方法。 只有 labels 时就可以跑 check,它会打印 derived(各音轨占比、段内切点密度、 旁白语速、声音切换与画面切点的对齐比例、分数位置的结构、第一次原声出现位置)。
  4. 事实与方法:在同一文件写 source_facts(带时间或测量锚点、显式 entities)和 methods (rule、applies_when、avoid_when、applies_to、evidence、targets)。targets 只写 {"from": "<测量路径>"}, 数值由 export 从当前测量填入,Agent 永远不手写数字。
  5. check → export:零 error 才导出;导出后对产物再做一次泄漏扫描。

字段、枚举和好/坏方法示例见 references/reference-schema.md。

Show full SKILL.md (190 more words)Show less

4. 分离规则

check 的 error(退出码 1):

规则要求
R1顶层、labels、fact、method、target、cut_fixes 都是封闭键集与封闭枚举;fact 不能带方法字段,method 不能带事实字段;id 为 f1… / m1… 且唯一;subtitles 值类型固定;skipped_dimensions 的值都是非空字符串
R2audio_spans、sections 按时间排序、不重叠、间隙 ≤0.5s、覆盖整片;字幕证据时间在时长内;cut_fixes.remove 对得上测得的切点,add 不与测得的切点重复
R3每条 fact 二选一锚定:t:[a,b] 在时长内,或 measure:[路径] 解析到 shots / loudness / derived 下的非字符串值;必须显式写 entities
R4每条 method 至少一条证据:已有 fact id 或同样只认这三个根的 measure:<路径>
R5target 只写 from:shots / loudness / derived 下的数值叶子,或白名单派生对象(见 schema);不得带列表下标
R6rule / applies_when / avoid_when 和 skipped_dimensions 的原因不得含:原片实体名(忽略空白)、与台词、背景资料或事实共有的连续 8 个汉字(标点隔开、跨相邻 ASR 窗口也算)或 5 个英文词、绝对时间码、"第 N 秒"或"N 分 M 秒"(含中文数字)、绝对路径
R7五个维度各至少一条 method,或在 skipped_dimensions 写明原因
R8导出物的每个键和字符串再扫一遍 R6,且不得出现 source_facts、labels、entities、evidence、statement、from、path 键

实体名来自 fact entities、understanding_index.json 的 characters / entities / research_glossary 名字与别名、background_research.json 的角色与 character_details 别名、≤5 字的 cultural_notes 条目,以及资料里《》「」引号中的短词。与事实共有 8 字时错误会注明来自 source_facts:通用剪辑措辞改写任一侧即可。

警告不阻断:method 缺 applies_when、rule 正文写了数字、ASR 不是 AVAILABLE_COARSE、有中文对白但理解产物与背景资料都没给出名字(此时只扫 fact entities)、basis 为空、 理解产物记录的成片身份与测量的成片不一致、有待复核窗口却没写 cut_fixes。

R6 只能拦住字面泄漏,拦不住改写过的剧情;写方法时要写"什么情况下怎么做",不要复述"这部片里发生了什么"。 规则也不判断语义:方法的方向必须与 check 打印的派生值一致(例如先看各音轨的 cuts_per_min 再写"哪类段落切得密")。

5. 交付与使用

production_reference.json 只含:时长与画布、cut_detection(切点规则的参数与 Agent 增删的切点数)、 profile(节奏数值,provenance 为 measured / reviewed(写了 cut_fixes,{} 也算)/ labeled)、 分数位置的 structure、字幕形态、methods、skipped_dimensions。

使用方式:把它复制进下一次制作的 work_dir。写稿阶段只在文件存在时阅读它, 并可在 recap_story_plan.json 的可选字段 reference_methods 记录每条方法的 adopt / adapt / skip。 它是参考,不是配额:与新素材的证据冲突时以素材为准。没有任何脚本读取它,也不会因为它的存在改变任何阶段的通过与否。

U/ 里的 breakdown 与 reference_frames/ 含原片事实和画面,只留在本地,不要复制或提交。

6. 能力边界

  • 不调用 MiMo 或任何 LLM,不改动理解产物,不给成片打分,不做"新片与参考的差距"判定。
  • 切点来自 ffmpeg scdet 分数:慢叠化测不到;相隔 0.3 秒内分数相近的两个真切点、刚好压在 soft 线上的切点(不同 ffmpeg 构建 4.00 与 3.989 分)、运动后的暗场硬切(0.9 分)都不算切点,只进待复核窗口。在一部 5 分钟剧集解说上测得 82 个切点, 11 个待复核窗口里 4 个含漏掉的 5 个真切点,其余是急推拉、翻页和运动镜头;补上后,三段逐帧真值(开场、暗场、中段共 27 个)全部对上、无误报。
  • 旁白占比与语速是 labeled 精度:来自 Agent 的音轨标注与粗 ASR 窗口,不是逐词对齐。
  • 泄漏扫描基于字面匹配,Agent 仍需自查方法是否在复述原片剧情。

© zenstory-ai, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 7 other files (scripts, references) in skills/video-reference of zenstory-ai/video-recap-skills.

  • SKILL.md
  • references/reference-schema.md
  • scripts/lib.py
  • scripts/reference.py
  • scripts/reference_check.py
  • scripts/reference_frames.py
  • scripts/reference_measure.py
  • scripts/reference_profile.py

Open the folder on GitHubat commit 5391686

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders. This page covers the copy in zenstory-ai/video-recap-skills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Video Reference next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Video Reference compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Video Reference this skillzenstory-ai/video-recap-skills561—~1.3kAutomated safety check: PassMIT
Dotty Av TestBrettKinny/dotty-stackchan113—~866Automated safety check: NotesMIT
Arkcli Deployvolcengine/ark-cli140—~2.6kAutomated safety check: PassApache-2.0
Arkcli Infer Endpointvolcengine/ark-cli140—~1.8kAutomated safety check: PassApache-2.0
Tts Integrationrapidaai/voice-ai745—~780Automated safety check: PassCustom licence
Speech Buildcnemri/google-genai-skills127—~430Automated safety check: PassMIT

Similar skills

  • Dotty Av Test

    BrettKinny/dotty-stackchan

    Run local black-box voice tests against the physical Dotty robot by playing a TTS prompt through the workstation speakers while the C920 records video and room audio.

    113 GitHub stars~866 tokensUpdated 5 days ago
    Media & CreativeAuto-check: notes
  • Arkcli Deploy

    volcengine/ark-cli

    arkcli +deploy:普通创建推理接入点(Endpoint)的统一首选入口。用户说『创建/新建/create 一个 endpoint/接入点』或『部署/上线/deploy 某模型』时优先走这里;但脚本化 / CI / 无护栏 / 原始 raw CRUD 创建是唯一例外,必须改走 arkcli-infer-endpoint,不能由本 skill…

    140 GitHub stars~2.6k tokensUpdated 2 days ago
    Media & CreativeAuto-check passed
  • Arkcli Infer Endpoint

    volcengine/ark-cli

    arkcli 推理接入点管理与显式 raw CRUD 创建能力。正向触发:用户明确要求脚本化 / CI / 无护栏 / 原始 raw CRUD 创建 Endpoint 时,必须使用本 skill 的 arkcli infer endpoint create 路径,不能切到 arkcli-deploy;模型缺失或只有品牌/家族名时仍留在本 skill,先执行实时有界候选与 0/1/N…

    140 GitHub stars~1.8k tokensUpdated 2 days ago
    Media & CreativeAuto-check passed
  • Tts Integration

    rapidaai/voice-ai

    Add or modify text-to-speech providers in assistant-api with transport-aware behavior (WS/SSE/SDK/HTTP), packet lifecycle correctness, and UI/provider wiring.

    745 GitHub stars~780 tokensUpdated 4 days ago
    Media & CreativeAuto-check passed
  • Speech Build

    cnemri/google-genai-skills

    Generate and transcribe speech using Google's Gemini-TTS and Chirp 3 models.

    127 GitHub stars~430 tokensUpdated 8 mo ago
    Media & CreativeAuto-check passed
  • Tts Integration

    rapidaai/voice-ai

    Add or modify text-to-speech providers in assistant-api with transport-aware behavior (WS/SSE/SDK/HTTP), correct packet lifecycle, and UI/provider wiring.

    745 GitHub stars~895 tokensUpdated 4 days ago
    Media & CreativeAuto-check passed

More from zenstory-ai/video-recap-skills

  • Video Recap

    zenstory-ai/video-recap-skills

    从输入视频生成中文解说成片或原声剧情短片。用户提供 .mp4 / .mov / .mkv / .webm,并要求剪辑、添加旁白、 配音、总结、短剧/电视剧/电影/纪录片/科普解说时使用。负责编排 video- 技能链:视频理解 → Agent 制定故事与视听方案 → 剪辑 → 配音 → 合成。触发词:视频解说、视频旁白、生成解说、 视频 recap、video…

    561 GitHub stars~2.4k tokensUpdated 7 days ago
    Auto-check passed
  • Video Assemble

    zenstory-ai/video-recap-skills

    合成视频解说最终成片:把旁白音频铺到源视频上,按旁白窗口压低原声,生成 SRT / ASS 字幕并可烧录, 最后做响度标准化。作为最终合成阶段使用。输入源视频、ttsmeta.json 与旁白位置; 输出 recap 成片和字幕。触发词:视频合成、混音、字幕、压字幕、assemble video、mux、ducking、subtitles、成片。

    561 GitHub stars~1.7k tokensUpdated 7 days ago
    Auto-check passed
  • Video Cut

    zenstory-ai/video-recap-skills

    把长视频按 Agent 选择的原片区间剪成短片。作为两阶段创作流程中的剪辑环节,读取 clipplan.json 与源视频, 输出 editedsource.mp4;随后 Agent 按输出时间线写 narration.json。支持单视频与多视频(sources manifest)拼剪, 本工具不读取、不映射旁白。

    561 GitHub stars~1.6k tokensUpdated 7 days ago
    Auto-check passed
  • Video Script

    zenstory-ai/video-recap-skills

    对已完成分析的视频进行导演与剪辑策划,再写带时间戳的中文解说并校验;也处理已有短片的 宣发标题、花字修订和外部文案回填。普通策划输入 workdir 的 agentnarrationbrief.md 与 vlmanalysis.json;文案返修输入当前成片的工程与内容证据。策划输出 recapstoryplan.json、visualaudioboard.json、 可选…

    561 GitHub stars~2.4k tokensUpdated 7 days ago
    Auto-check passed
  • Video Understanding

    zenstory-ai/video-recap-skills

    把视频分析为结构化理解索引:场景检测、ASR 转写、逐场景 VLM 观察、静音窗口、融合时间线和写作 brief. An agent skill from zenstory-ai/video-recap-skills.

    561 GitHub stars~1.1k tokensUpdated 7 days ago
    Auto-check passed
  • Video Voiceover

    zenstory-ai/video-recap-skills

    把带时间戳的 narration.json 合成为中文解说音频。使用 MiMo TTS(mimo-v2.5-tts)或 Fish Audio(s2.1-pro-free)或显式配置的通用 IndexTTS HTTP 服务逐段生成语音, 按时间窗动态适配语速并处理响度;输入输出时间线上的旁白,产出 ttssegments 与 ttsmeta.json。

    561 GitHub stars~1.6k tokensUpdated 7 days ago
    Auto-check passed

Questions about Video Reference

What does Video Reference do?

按需把一部成片拆成可复用的制作参考:测镜头节奏与响度,标注段落与音轨分工,把原片事实与可迁移方法分开, 导出不含原片人名台词的 productionreference.json 供下次制作参考。不在默认生产路径上。. Video Reference is an agent skill from zenstory-ai/video-recap-skills.

When should I use Video Reference?

Video Reference fits situations like: tasks that involve Text to speech and voice; tasks that involve Speech recognition and synthesis.

How do I install Video Reference in Claude Code?

Run `npx skills add zenstory-ai/video-recap-skills --skill video-reference -a claude-code`. Or copy the skill folder (skills/video-reference in zenstory-ai/video-recap-skills) into .claude/skills/video-reference in your project. Claude Code loads it when a task matches its description.

How do I install Video Reference in Codex?

Run `npx skills add zenstory-ai/video-recap-skills --skill video-reference -a codex`. Or copy the skill folder (skills/video-reference in zenstory-ai/video-recap-skills) into .agents/skills/video-reference in your project. Codex loads it when a task matches its description.

Can I use Video Reference in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add zenstory-ai/video-recap-skills --skill video-reference -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/video-reference, .gemini/skills/video-reference, .github/skills/video-reference and .opencode/skills/video-reference in your project.

What does Video Reference need to run?

Going by SKILL.md and its folder, Video Reference needs Python for the scripts in its folder and the command-line tools its instructions call (python3). Our summary lists: Python 3.

Does Video Reference access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Video Reference safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Video Reference use?

Video Reference is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Video Reference use?

About 1.3k tokens (SKILL.md is roughly 5.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.3k tokens, read only when the agent opens those files.

What are the alternatives to Video Reference?

Skills that share tags, products or a category with Video Reference: Dotty Av Test (BrettKinny/dotty-stackchan, 113 stars), Arkcli Deploy (volcengine/ark-cli, 140 stars), Arkcli Infer Endpoint (volcengine/ark-cli, 140 stars) and Tts Integration (rapidaai/voice-ai, 745 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Video Reference?

zenstory-ai (a GitHub organization) maintains it in zenstory-ai/video-recap-skills, which has 561 GitHub stars. The repository holds 7 skills in this directory. The repository was last updated on October 4, 2026.

Source: zenstory-ai/video-recap-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.