Agent skill

Video Understanding

by zenstory-ai in zenstory-ai/video-recap-skills

把视频分析为结构化理解索引:场景检测、ASR 转写、逐场景 VLM 观察、静音窗口、融合时间线和写作 brief. An agent skill from zenstory-ai/video-recap-skills.

MITAuto-check passedAI & LLM Engineering

Install Video Understanding

skills CLI
$ npx skills add zenstory-ai/video-recap-skills --skill video-understanding -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install zenstory-ai/video-recap-skills video-understanding --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/zenstory-ai/video-recap-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/video-understanding .claude/skills/video-understanding && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
video-understanding
GitHub stars
561
Token cost
~1.1k tokens
SKILL.md length
281 words
Files
25 (incl. scripts, references)
Skills in repo
7
Repo updated
First seen
Licence
MIT

At a glance

把视频分析为结构化理解索引:场景检测、ASR 转写、逐场景 VLM 观察、静音窗口、融合时间线和写作 brief. An agent skill from zenstory-ai/video-recap-skills.

  • Works in 7 steps: 定位 → 处理阶段 → 环境要求 → …
  • Tasks that involve Computer vision
  • SKILL.md covers 1. 定位, 2. 处理阶段, 3. 环境要求 and 4. 运行命令, plus 3 more sections
  • Runs Python scripts from its folder; calls python3 and rsync; needs MIMO_API_KEY

What it does

Video Understanding is an agent skill from zenstory-ai/video-recap-skills. 把视频分析为结构化理解索引:场景检测、ASR 转写、逐场景 VLM 观察、静音窗口、融合时间线和写作 brief。 用于理解、索引或总结视频,也作为后续创作前的分析阶段。输入视频文件;输出 scenes.json、 asrresult.json、vlmanalysis.json、silenceperiods.json、timelinefusion.json、agentnarrationbrief.md。 触发词:视频理解、视频分析、视频索引、video understanding、analyze video、看懂视频。

Its SKILL.md is about 1.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 27 other files, including scripts and reference files (for example `references/data-schema.md`, `references/prompt-templates.md` and `references/research-guide.md`).

It sits in AI & LLM Engineering, covering Computer vision, Speech recognition and synthesis and Text to speech and voice. The repository describes itself as: Claude Code / Codex skills that turn a video into a Chinese narration recap (视频解说): scene detection, ASR, VLM, script, TTS, ffmpeg assembly, optional editable JianYing / CapCut… The licence is MIT.

When your agent uses it

  • Tasks that involve Computer vision
  • Tasks that involve Speech recognition and synthesis
  • Tasks that involve Text to speech and voice

Example prompts

  • “/video-understanding”

Requirements

  • Python 3
  • A credential in MIMO_API_KEY

Workflow steps

7 steps, taken from the step headings in SKILL.md.

  1. 定位
  2. 处理阶段
  3. 环境要求
  4. 运行命令
  5. 输出契约
  6. 参考资料
  7. 能力边界

What it can do on your machine

Read from SKILL.md and the folder at commit 5391686. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 14 files in scripts/ (Python, from the files we listed), which the agent can run.

    Shell commands in SKILL.md call:

    • python3
    • rsync

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use rsync, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • MIMO_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Video Understanding loads about 1.1k tokens when it runs, and up to ~3.8k if it reads all its reference files. Until then it costs about 72 tokens; SKILL.md has 281 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~72
When it runs · the whole SKILL.md, loaded when a task matches
~1.1k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~3.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from zenstory-ai/video-recap-skills at commit 5391686, republished under its MIT licence (© zenstory-ai). 281 words, ~1,115 tokens.

Download SKILL.mdSave it as .claude/skills/video-understanding/SKILL.md (or your agent's skills folder). This skill also uses 24 other files; get the full folder from GitHub.
name
video-understanding
description
把视频分析为结构化理解索引:场景检测、ASR 转写、逐场景 VLM 观察、静音窗口、融合时间线和写作 brief。 用于理解、索引或总结视频,也作为后续创作前的分析阶段。输入视频文件;输出 scenes.json、 asr_result.json、vlm_analysis.json、silence_periods.json、timeline_fusion.json、agent_narration_brief.md。 触发词:视频理解、视频分析、视频索引、video understanding、analyze video、看懂视频。
user-invocable
false

1. 定位

本技能把源视频转成 Agent 与下游阶段可读取的理解索引。它的创作角色是素材观察员 / 场记,不是导演:

  • 先观察,再解释;事实与推断分开。
  • 除了“发生了什么”,还要让下游看见知识、权力、目标、关系或情绪在哪一刻变化。
  • 标出由谁的 POV 承载变化、哪个反应或表演不可替代,以及哪里存在完整台词/动作的自然剪辑边界。
  • 证据不足时保留不确定性,不制造戏剧结论。

2. 处理阶段

  1. 场景检测:写 scenes.json,包含切点、时长和废片段过滤结果。
  2. 抽帧:为视觉分析提取代表帧。
  3. ASR:通过 mimo-v2.5-asr 写粗分段对白 asr_result.json,并写 asr_timing_evidence.json 说明可用性、有限时间精度与文本修正来源。
  4. 静音检测:写 silence_periods.json,标注安静窗口与 has_speech。
  5. VLM 观察:写 vlm_analysis.json,包含场景描述、深层分析和 frame_facts。
  6. 时间线融合与创作 brief:写 timeline_fusion.json、asr_writing_chunks.json 和 agent_narration_brief.md。

各阶段只有在输出产物与 provenance sidecar 同时匹配当前视频及影响结果的设置时才会复用;--force 强制重算。

3. 环境要求

bash
# ffmpeg: brew install ffmpeg | apt install ffmpeg | choco install ffmpeg
export MIMO_API_KEY=***

ASR 使用 mimo-v2.5-asr;VLM 使用 mimo-v2.5。--skip-asr 可跳过对白转写,但完整理解仍需要 MIMO_API_KEY 运行 VLM。--mimo-video-overview 可开启按场景块的视频概览。未设置 key 时重跑会复用已缓存的转写、画面分析、概览与故事索引(key 决定的默认 endpoint 不参与比对);需要请求模型的 consolidation 记为 skipped_no_key,不发请求。缓存对不上(例如复制 work_dir 时没保留文件时间)而已有转写时,运行停下并保留转写:用 cp -p / cp -Rp / rsync -t 保留时间重新复制,或设置 key 后重跑(会重新转写)。不要用 --skip-asr 绕过,它会把现有转写替换成 []。

若 work_dir/background_research.json 存在,本技能会把剧情梗概和角色名折入 VLM 上下文;--context 可补充一条简短提示。

下面的 scripts/... 均相对于本技能目录。若执行器从仓库根目录启动,请给脚本路径加上本技能的绝对目录。

4. 运行命令

bash
python3 scripts/understand.py <video> --work-dir <work_dir> [选项]
选项默认作用
<video>必填源视频
--work-dir必填产物目录;不存在时创建
--context "..."空补充给 VLM 的简短上下文(节目名、角色名),与 background_research.json 合并
--scene-threshold0.1场景检测阈值
--style纪录片写进创作简报的解说风格
--edit-mode full|cut不设写进简报的 recap 模式;cut 时按剪后时长估算旁白预算,已有 edited_source.mp4 时句末锚点改用剪后时间
--target-duration不设写进简报的 cut 目标时长;尚无 clip_plan_validated.json 时用它估算旁白预算
--skip-asr关不转写对白,把 asr_result.json 写成 [](已有转写会被覆盖),ASR 证据标为显式跳过
--mimo-video-overview关按场景块运行 MiMo 视频概览,并作为逐场景主描述
--force关忽略缓存,全部重算
--brief-only关只用现有产物重建 agent_narration_brief.md,不抽帧、不调 API
--edited-storyboard-only关只按 clip_plan_validated.json 写剪后时间线 storyboard/edited_storyboard.*,并在已有的 agent_narration_brief.md 顶部加 storyboard 指引;多源计划(带 sources)从各来源 source_work_dir 的 frames/ 按其 frames_manifest.json 的 fps 取帧,tile 标 S1/S2…;不抽帧、不调 API,故事板生成失败只记日志。与 --brief-only 互斥
--consolidate / --no-consolidate开生成全局故事索引 understanding_index.*
--consolidate-asr关另外清洗 ASR 文本,写 asr_clean.json

5. 输出契约

默认运行写出下表产物;各阶段产物旁的 *.meta.json 是缓存 provenance sidecar。

文件内容
frames/frame_*.jpg、frames_manifest.json按 fps 抽出的帧及其清单
scenes.json场景切点、起止时间与时长
audio.wav16 kHz 单声道音频,供 ASR、静音检测与句末锚点使用
asr_result.json[{start, end, text}] 时间戳对白
asr_timing_evidence.jsonASR 可用性状态、粗窗口精度、glossary 前后文本,以及它所描述的源视频/音频/结果文件(路径存在性 + size/mtime)
silence_periods.json[{start, end, duration, has_speech}] 安静窗口
speech_boundary_anchors.jsonASR 句末标点对齐到短停顿的句末锚点;缺音频或 ASR 时 status: unavailable
vlm_analysis.json逐场景描述、深层分析与 frame_facts
mimo_video_overview.status.json视频概览状态(未启用时为 disabled);启用成功另写 mimo_video_overview.json
understanding_index.json、understanding_index.md全局故事索引(--no-consolidate 时不写)
consolidation.status.json故事索引与 ASR 清洗的运行状态
storyboard/source_storyboard.{json,jpg}原片时间线联系表(超页时续写 _001.jpg 等);已有 clip_plan_validated.json 时另写 edited_storyboard.*;STORYBOARD=0 关闭
timeline_fusion.jsonVLM、ASR 与静音信息的统一时间线
asr_writing_chunks.json按句界和场景切分的 ASR 写作块
agent_narration_brief.mdAgent 首先阅读的创作简报

后续写作阶段根据创作简报与索引制定方案并写 narration.json。

6. 参考资料

  • 背景调研:references/research-guide.md,产出 background_research.json。
  • JSON 结构:references/data-schema.md。

7. 能力边界

  • 不写解说词,也不做解说评分;只负责生成理解索引与创作简报。
  • 不编造信号无法支持的剧情;当 ASR / VLM 过薄时输出素材警告。
  • MiMo ASR 的 start/end 是固定分片形成的粗窗口,不是词级对齐;空文本只表示原因未知, 不能当作已证实静音。asr_timing_evidence.json 的状态字段见 references/data-schema.md。

© zenstory-ai, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 24 other files (scripts, references) in skills/video-understanding of zenstory-ai/video-recap-skills.

  • SKILL.md
  • references/data-schema.md
  • references/prompt-templates.md
  • references/research-guide.md
  • scripts/agent_text.py
  • scripts/asr.py
  • scripts/asr_timing_evidence.py
  • scripts/briefing/__init__.py
  • scripts/briefing/builder.py
  • scripts/briefing/context.py
  • scripts/briefing/inputs.py
  • scripts/briefing/timeline.py
  • scripts/consolidate.py
  • scripts/detect.py
  • scripts/extract.py
  • scripts/index_normalize.py
  • scripts/lib.py
  • scripts/storyboard.py
  • … and 7 more

Open the folder on GitHubat commit 5391686

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders. This page covers the copy in zenstory-ai/video-recap-skills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Video Understanding next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Video Understanding compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Video Understanding this skillzenstory-ai/video-recap-skills561—~1.1kAutomated safety check: PassMIT
Hriterrense/ros2-multimodal-robot-collab111—~268Automated safety check: PassMIT
Agentstadaspetra/loop2961 repos~2.5kAutomated safety check: PassMIT
Piper Tts Trainingsammcj/agentic-coding162—~1.4kAutomated safety check: PassApache-2.0
Agentselevenlabs/skills482—~6.5kAutomated safety check: PassMIT
Voice Agentsdavila7/claude-code-templates33k3 repos~565Automated safety check: PassMIT

Similar skills

  • Hri

    terrense/ros2-multimodal-robot-collab

    A skill your agent uses when an Agent needs to speak to the operator through TTS, interpret ASR text, request clarification, or confirm a robot delivery action.

    111 GitHub stars~268 tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Agents

    tadaspetra/loop

    Build voice AI agents with ElevenLabs. An agent skill from tadaspetra/loop.

    296 GitHub starsUsed in 1 repo~2.5k tokens
    AI & LLM EngineeringAuto-check passed
  • Piper Tts Training

    sammcj/agentic-coding

    Train custom TTS voices for Piper (ONNX format) using fine-tuning or from-scratch approaches.

    162 GitHub stars~1.4k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • Agents

    elevenlabs/skills

    Build voice AI agents with ElevenLabs. An agent skill from elevenlabs/skills.

    482 GitHub stars~6.5k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • Voice Agents

    davila7/claude-code-templates

    Voice agents represent the frontier of AI interaction - humans speaking naturally with AI systems.

    33k GitHub starsUsed in 3 repos~565 tokens
    AI & LLM EngineeringAuto-check passed
  • Official

    Stage 1 of Clinical ASR Flywheel. An agent skill from NVIDIA/skills.

    3.6k GitHub stars~4.5k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check: notes

More from zenstory-ai/video-recap-skills

  • Video Recap

    zenstory-ai/video-recap-skills

    从输入视频生成中文解说成片或原声剧情短片。用户提供 .mp4 / .mov / .mkv / .webm,并要求剪辑、添加旁白、 配音、总结、短剧/电视剧/电影/纪录片/科普解说时使用。负责编排 video- 技能链:视频理解 → Agent 制定故事与视听方案 → 剪辑 → 配音 → 合成。触发词:视频解说、视频旁白、生成解说、 视频 recap、video…

    561 GitHub stars~2.4k tokensUpdated 7 days ago
    Auto-check passed
  • Video Assemble

    zenstory-ai/video-recap-skills

    合成视频解说最终成片:把旁白音频铺到源视频上,按旁白窗口压低原声,生成 SRT / ASS 字幕并可烧录, 最后做响度标准化。作为最终合成阶段使用。输入源视频、ttsmeta.json 与旁白位置; 输出 recap 成片和字幕。触发词:视频合成、混音、字幕、压字幕、assemble video、mux、ducking、subtitles、成片。

    561 GitHub stars~1.7k tokensUpdated 7 days ago
    Auto-check passed
  • Video Cut

    zenstory-ai/video-recap-skills

    把长视频按 Agent 选择的原片区间剪成短片。作为两阶段创作流程中的剪辑环节,读取 clipplan.json 与源视频, 输出 editedsource.mp4;随后 Agent 按输出时间线写 narration.json。支持单视频与多视频(sources manifest)拼剪, 本工具不读取、不映射旁白。

    561 GitHub stars~1.6k tokensUpdated 7 days ago
    Auto-check passed
  • Video Reference

    zenstory-ai/video-recap-skills

    按需把一部成片拆成可复用的制作参考:测镜头节奏与响度,标注段落与音轨分工,把原片事实与可迁移方法分开, 导出不含原片人名台词的 productionreference.json 供下次制作参考。不在默认生产路径上。

    561 GitHub stars~1.3k tokensUpdated 7 days ago
    Auto-check passed
  • Video Script

    zenstory-ai/video-recap-skills

    对已完成分析的视频进行导演与剪辑策划,再写带时间戳的中文解说并校验;也处理已有短片的 宣发标题、花字修订和外部文案回填。普通策划输入 workdir 的 agentnarrationbrief.md 与 vlmanalysis.json;文案返修输入当前成片的工程与内容证据。策划输出 recapstoryplan.json、visualaudioboard.json、 可选…

    561 GitHub stars~2.4k tokensUpdated 7 days ago
    Auto-check passed
  • Video Voiceover

    zenstory-ai/video-recap-skills

    把带时间戳的 narration.json 合成为中文解说音频。使用 MiMo TTS(mimo-v2.5-tts)或 Fish Audio(s2.1-pro-free)或显式配置的通用 IndexTTS HTTP 服务逐段生成语音, 按时间窗动态适配语速并处理响度;输入输出时间线上的旁白,产出 ttssegments 与 ttsmeta.json。

    561 GitHub stars~1.6k tokensUpdated 7 days ago
    Auto-check passed

Questions about Video Understanding

What does Video Understanding do?

把视频分析为结构化理解索引:场景检测、ASR 转写、逐场景 VLM 观察、静音窗口、融合时间线和写作 brief. An agent skill from zenstory-ai/video-recap-skills. Video Understanding is an agent skill from zenstory-ai/video-recap-skills.

When should I use Video Understanding?

Video Understanding fits situations like: tasks that involve Computer vision; tasks that involve Speech recognition and synthesis; tasks that involve Text to speech and voice.

How do I install Video Understanding in Claude Code?

Run `npx skills add zenstory-ai/video-recap-skills --skill video-understanding -a claude-code`. Or copy the skill folder (skills/video-understanding in zenstory-ai/video-recap-skills) into .claude/skills/video-understanding in your project. Claude Code loads it when a task matches its description.

How do I install Video Understanding in Codex?

Run `npx skills add zenstory-ai/video-recap-skills --skill video-understanding -a codex`. Or copy the skill folder (skills/video-understanding in zenstory-ai/video-recap-skills) into .agents/skills/video-understanding in your project. Codex loads it when a task matches its description.

Can I use Video Understanding in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add zenstory-ai/video-recap-skills --skill video-understanding -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/video-understanding, .gemini/skills/video-understanding, .github/skills/video-understanding and .opencode/skills/video-understanding in your project.

What does Video Understanding need to run?

Going by SKILL.md and its folder, Video Understanding needs Python for the scripts in its folder, the command-line tools its instructions call (python3 and rsync) and credentials named MIMO_API_KEY. Our summary lists: Python 3; A credential in MIMO_API_KEY.

Does Video Understanding access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Video Understanding safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Video Understanding use?

Video Understanding is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Video Understanding use?

About 1.1k tokens (SKILL.md is roughly 4.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.7k tokens, read only when the agent opens those files.

What are the alternatives to Video Understanding?

Skills that share tags, products or a category with Video Understanding: Hri (terrense/ros2-multimodal-robot-collab, 111 stars), Agents (tadaspetra/loop, 296 stars), Piper Tts Training (sammcj/agentic-coding, 162 stars) and Agents (elevenlabs/skills, 482 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Video Understanding?

zenstory-ai (a GitHub organization) maintains it in zenstory-ai/video-recap-skills, which has 561 GitHub stars. The repository holds 7 skills in this directory. The repository was last updated on October 4, 2026.

Source: zenstory-ai/video-recap-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.