Hri
terrense/ros2-multimodal-robot-collab
A skill your agent uses when an Agent needs to speak to the operator through TTS, interpret ASR text, request clarification, or confirm a robot delivery action.
把视频分析为结构化理解索引:场景检测、ASR 转写、逐场景 VLM 观察、静音窗口、融合时间线和写作 brief. An agent skill from zenstory-ai/video-recap-skills.
$ npx skills add zenstory-ai/video-recap-skills --skill video-understanding -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install zenstory-ai/video-recap-skills video-understanding --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/zenstory-ai/video-recap-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/video-understanding .claude/skills/video-understanding && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "video-understanding" agent skill from https://github.com/zenstory-ai/video-recap-skills/tree/main/skills/video-understanding into .claude/skills/video-understanding/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "video-understanding", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/zenstory-ai/video-recap-skills/tree/main/skills/video-understandingType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add zenstory-ai/video-recap-skills --skill video-understanding -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install zenstory-ai/video-recap-skills video-understanding --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/zenstory-ai/video-recap-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/video-understanding .agents/skills/video-understanding && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "video-understanding" agent skill from https://github.com/zenstory-ai/video-recap-skills/tree/main/skills/video-understanding into .agents/skills/video-understanding/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "video-understanding", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add zenstory-ai/video-recap-skills --skill video-understanding -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install zenstory-ai/video-recap-skills video-understanding --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/zenstory-ai/video-recap-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/video-understanding .cursor/skills/video-understanding && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "video-understanding" agent skill from https://github.com/zenstory-ai/video-recap-skills/tree/main/skills/video-understanding into .cursor/skills/video-understanding/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "video-understanding", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/zenstory-ai/video-recap-skills.git --path skills/video-understanding--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add zenstory-ai/video-recap-skills --skill video-understanding -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install zenstory-ai/video-recap-skills video-understanding --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/zenstory-ai/video-recap-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/video-understanding .gemini/skills/video-understanding && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "video-understanding" agent skill from https://github.com/zenstory-ai/video-recap-skills/tree/main/skills/video-understanding into .gemini/skills/video-understanding/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "video-understanding", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install zenstory-ai/video-recap-skills video-understandingInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add zenstory-ai/video-recap-skills --skill video-understanding -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/zenstory-ai/video-recap-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/video-understanding .github/skills/video-understanding && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "video-understanding" agent skill from https://github.com/zenstory-ai/video-recap-skills/tree/main/skills/video-understanding into .github/skills/video-understanding/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "video-understanding", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add zenstory-ai/video-recap-skills --skill video-understanding -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install zenstory-ai/video-recap-skills video-understanding --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/zenstory-ai/video-recap-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/video-understanding .opencode/skills/video-understanding && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "video-understanding" agent skill from https://github.com/zenstory-ai/video-recap-skills/tree/main/skills/video-understanding into .opencode/skills/video-understanding/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "video-understanding", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
video-understanding把视频分析为结构化理解索引:场景检测、ASR 转写、逐场景 VLM 观察、静音窗口、融合时间线和写作 brief. An agent skill from zenstory-ai/video-recap-skills.
Video Understanding is an agent skill from zenstory-ai/video-recap-skills. 把视频分析为结构化理解索引:场景检测、ASR 转写、逐场景 VLM 观察、静音窗口、融合时间线和写作 brief。 用于理解、索引或总结视频,也作为后续创作前的分析阶段。输入视频文件;输出 scenes.json、 asrresult.json、vlmanalysis.json、silenceperiods.json、timelinefusion.json、agentnarrationbrief.md。 触发词:视频理解、视频分析、视频索引、video understanding、analyze video、看懂视频。
Its SKILL.md is about 1.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 27 other files, including scripts and reference files (for example `references/data-schema.md`, `references/prompt-templates.md` and `references/research-guide.md`).
It sits in AI & LLM Engineering, covering Computer vision, Speech recognition and synthesis and Text to speech and voice. The repository describes itself as: Claude Code / Codex skills that turn a video into a Chinese narration recap (视频解说): scene detection, ASR, VLM, script, TTS, ffmpeg assembly, optional editable JianYing / CapCut… The licence is MIT.
7 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 5391686. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 14 files in scripts/ (Python, from the files we listed), which the agent can run.
Shell commands in SKILL.md call:
python3rsyncFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use rsync, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
MIMO_API_KEYFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Video Understanding loads about 1.1k tokens when it runs, and up to ~3.8k if it reads all its reference files. Until then it costs about 72 tokens; SKILL.md has 281 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from zenstory-ai/video-recap-skills at commit 5391686, republished under its MIT licence (© zenstory-ai). 281 words, ~1,115 tokens.
.claude/skills/video-understanding/SKILL.md (or your agent's skills folder). This skill also uses 24 other files; get the full folder from GitHub.本技能把源视频转成 Agent 与下游阶段可读取的理解索引。它的创作角色是素材观察员 / 场记,不是导演:
scenes.json,包含切点、时长和废片段过滤结果。mimo-v2.5-asr 写粗分段对白 asr_result.json,并写
asr_timing_evidence.json 说明可用性、有限时间精度与文本修正来源。silence_periods.json,标注安静窗口与 has_speech。vlm_analysis.json,包含场景描述、深层分析和 frame_facts。timeline_fusion.json、asr_writing_chunks.json 和 agent_narration_brief.md。各阶段只有在输出产物与 provenance sidecar 同时匹配当前视频及影响结果的设置时才会复用;--force 强制重算。
# ffmpeg: brew install ffmpeg | apt install ffmpeg | choco install ffmpeg
export MIMO_API_KEY=***ASR 使用 mimo-v2.5-asr;VLM 使用 mimo-v2.5。--skip-asr 可跳过对白转写,但完整理解仍需要 MIMO_API_KEY 运行 VLM。--mimo-video-overview 可开启按场景块的视频概览。未设置 key 时重跑会复用已缓存的转写、画面分析、概览与故事索引(key 决定的默认 endpoint 不参与比对);需要请求模型的 consolidation 记为 skipped_no_key,不发请求。缓存对不上(例如复制 work_dir 时没保留文件时间)而已有转写时,运行停下并保留转写:用 cp -p / cp -Rp / rsync -t 保留时间重新复制,或设置 key 后重跑(会重新转写)。不要用 --skip-asr 绕过,它会把现有转写替换成 []。
若 work_dir/background_research.json 存在,本技能会把剧情梗概和角色名折入 VLM 上下文;--context 可补充一条简短提示。
下面的 scripts/... 均相对于本技能目录。若执行器从仓库根目录启动,请给脚本路径加上本技能的绝对目录。
python3 scripts/understand.py <video> --work-dir <work_dir> [选项]| 选项 | 默认 | 作用 |
|---|---|---|
<video> | 必填 | 源视频 |
--work-dir | 必填 | 产物目录;不存在时创建 |
--context "..." | 空 | 补充给 VLM 的简短上下文(节目名、角色名),与 background_research.json 合并 |
--scene-threshold | 0.1 | 场景检测阈值 |
--style | 纪录片 | 写进创作简报的解说风格 |
--edit-mode full|cut | 不设 | 写进简报的 recap 模式;cut 时按剪后时长估算旁白预算,已有 edited_source.mp4 时句末锚点改用剪后时间 |
--target-duration | 不设 | 写进简报的 cut 目标时长;尚无 clip_plan_validated.json 时用它估算旁白预算 |
--skip-asr | 关 | 不转写对白,把 asr_result.json 写成 [](已有转写会被覆盖),ASR 证据标为显式跳过 |
--mimo-video-overview | 关 | 按场景块运行 MiMo 视频概览,并作为逐场景主描述 |
--force | 关 | 忽略缓存,全部重算 |
--brief-only | 关 | 只用现有产物重建 agent_narration_brief.md,不抽帧、不调 API |
--edited-storyboard-only | 关 | 只按 clip_plan_validated.json 写剪后时间线 storyboard/edited_storyboard.*,并在已有的 agent_narration_brief.md 顶部加 storyboard 指引;多源计划(带 sources)从各来源 source_work_dir 的 frames/ 按其 frames_manifest.json 的 fps 取帧,tile 标 S1/S2…;不抽帧、不调 API,故事板生成失败只记日志。与 --brief-only 互斥 |
--consolidate / --no-consolidate | 开 | 生成全局故事索引 understanding_index.* |
--consolidate-asr | 关 | 另外清洗 ASR 文本,写 asr_clean.json |
默认运行写出下表产物;各阶段产物旁的 *.meta.json 是缓存 provenance sidecar。
| 文件 | 内容 |
|---|---|
frames/frame_*.jpg、frames_manifest.json | 按 fps 抽出的帧及其清单 |
scenes.json | 场景切点、起止时间与时长 |
audio.wav | 16 kHz 单声道音频,供 ASR、静音检测与句末锚点使用 |
asr_result.json | [{start, end, text}] 时间戳对白 |
asr_timing_evidence.json | ASR 可用性状态、粗窗口精度、glossary 前后文本,以及它所描述的源视频/音频/结果文件(路径存在性 + size/mtime) |
silence_periods.json | [{start, end, duration, has_speech}] 安静窗口 |
speech_boundary_anchors.json | ASR 句末标点对齐到短停顿的句末锚点;缺音频或 ASR 时 status: unavailable |
vlm_analysis.json | 逐场景描述、深层分析与 frame_facts |
mimo_video_overview.status.json | 视频概览状态(未启用时为 disabled);启用成功另写 mimo_video_overview.json |
understanding_index.json、understanding_index.md | 全局故事索引(--no-consolidate 时不写) |
consolidation.status.json | 故事索引与 ASR 清洗的运行状态 |
storyboard/source_storyboard.{json,jpg} | 原片时间线联系表(超页时续写 _001.jpg 等);已有 clip_plan_validated.json 时另写 edited_storyboard.*;STORYBOARD=0 关闭 |
timeline_fusion.json | VLM、ASR 与静音信息的统一时间线 |
asr_writing_chunks.json | 按句界和场景切分的 ASR 写作块 |
agent_narration_brief.md | Agent 首先阅读的创作简报 |
后续写作阶段根据创作简报与索引制定方案并写 narration.json。
references/research-guide.md,产出 background_research.json。references/data-schema.md。start/end 是固定分片形成的粗窗口,不是词级对齐;空文本只表示原因未知,
不能当作已证实静音。asr_timing_evidence.json 的状态字段见 references/data-schema.md。© zenstory-ai, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 24 other files (scripts, references) in skills/video-understanding of zenstory-ai/video-recap-skills.
Open the folder on GitHubat commit 5391686
We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders. This page covers the copy in zenstory-ai/video-recap-skills, which our catalogue first saw on October 7, 2026.
Video Understanding next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Video Understanding this skillzenstory-ai/video-recap-skills | 561 | — | ~1.1k | Automated safety check: Pass | MIT | |
| Hriterrense/ros2-multimodal-robot-collab | 111 | — | ~268 | Automated safety check: Pass | MIT | |
| Agentstadaspetra/loop | 296 | 1 repos | ~2.5k | Automated safety check: Pass | MIT | |
| Piper Tts Trainingsammcj/agentic-coding | 162 | — | ~1.4k | Automated safety check: Pass | Apache-2.0 | |
| Agentselevenlabs/skills | 482 | — | ~6.5k | Automated safety check: Pass | MIT | |
| Voice Agentsdavila7/claude-code-templates | 33k | 3 repos | ~565 | Automated safety check: Pass | MIT |
terrense/ros2-multimodal-robot-collab
A skill your agent uses when an Agent needs to speak to the operator through TTS, interpret ASR text, request clarification, or confirm a robot delivery action.
tadaspetra/loop
Build voice AI agents with ElevenLabs. An agent skill from tadaspetra/loop.
sammcj/agentic-coding
Train custom TTS voices for Piper (ONNX format) using fine-tuning or from-scratch approaches.
elevenlabs/skills
Build voice AI agents with ElevenLabs. An agent skill from elevenlabs/skills.
davila7/claude-code-templates
Voice agents represent the frontier of AI interaction - humans speaking naturally with AI systems.
NVIDIA/skills
Stage 1 of Clinical ASR Flywheel. An agent skill from NVIDIA/skills.
zenstory-ai/video-recap-skills
从输入视频生成中文解说成片或原声剧情短片。用户提供 .mp4 / .mov / .mkv / .webm,并要求剪辑、添加旁白、 配音、总结、短剧/电视剧/电影/纪录片/科普解说时使用。负责编排 video- 技能链:视频理解 → Agent 制定故事与视听方案 → 剪辑 → 配音 → 合成。触发词:视频解说、视频旁白、生成解说、 视频 recap、video…
zenstory-ai/video-recap-skills
合成视频解说最终成片:把旁白音频铺到源视频上,按旁白窗口压低原声,生成 SRT / ASS 字幕并可烧录, 最后做响度标准化。作为最终合成阶段使用。输入源视频、ttsmeta.json 与旁白位置; 输出 recap 成片和字幕。触发词:视频合成、混音、字幕、压字幕、assemble video、mux、ducking、subtitles、成片。
zenstory-ai/video-recap-skills
把长视频按 Agent 选择的原片区间剪成短片。作为两阶段创作流程中的剪辑环节,读取 clipplan.json 与源视频, 输出 editedsource.mp4;随后 Agent 按输出时间线写 narration.json。支持单视频与多视频(sources manifest)拼剪, 本工具不读取、不映射旁白。
zenstory-ai/video-recap-skills
按需把一部成片拆成可复用的制作参考:测镜头节奏与响度,标注段落与音轨分工,把原片事实与可迁移方法分开, 导出不含原片人名台词的 productionreference.json 供下次制作参考。不在默认生产路径上。
zenstory-ai/video-recap-skills
对已完成分析的视频进行导演与剪辑策划,再写带时间戳的中文解说并校验;也处理已有短片的 宣发标题、花字修订和外部文案回填。普通策划输入 workdir 的 agentnarrationbrief.md 与 vlmanalysis.json;文案返修输入当前成片的工程与内容证据。策划输出 recapstoryplan.json、visualaudioboard.json、 可选…
zenstory-ai/video-recap-skills
把带时间戳的 narration.json 合成为中文解说音频。使用 MiMo TTS(mimo-v2.5-tts)或 Fish Audio(s2.1-pro-free)或显式配置的通用 IndexTTS HTTP 服务逐段生成语音, 按时间窗动态适配语速并处理响度;输入输出时间线上的旁白,产出 ttssegments 与 ttsmeta.json。
Categories
把视频分析为结构化理解索引:场景检测、ASR 转写、逐场景 VLM 观察、静音窗口、融合时间线和写作 brief. An agent skill from zenstory-ai/video-recap-skills. Video Understanding is an agent skill from zenstory-ai/video-recap-skills.
Video Understanding fits situations like: tasks that involve Computer vision; tasks that involve Speech recognition and synthesis; tasks that involve Text to speech and voice.
Run `npx skills add zenstory-ai/video-recap-skills --skill video-understanding -a claude-code`. Or copy the skill folder (skills/video-understanding in zenstory-ai/video-recap-skills) into .claude/skills/video-understanding in your project. Claude Code loads it when a task matches its description.
Run `npx skills add zenstory-ai/video-recap-skills --skill video-understanding -a codex`. Or copy the skill folder (skills/video-understanding in zenstory-ai/video-recap-skills) into .agents/skills/video-understanding in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add zenstory-ai/video-recap-skills --skill video-understanding -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/video-understanding, .gemini/skills/video-understanding, .github/skills/video-understanding and .opencode/skills/video-understanding in your project.
Going by SKILL.md and its folder, Video Understanding needs Python for the scripts in its folder, the command-line tools its instructions call (python3 and rsync) and credentials named MIMO_API_KEY. Our summary lists: Python 3; A credential in MIMO_API_KEY.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Video Understanding is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.1k tokens (SKILL.md is roughly 4.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.7k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Video Understanding: Hri (terrense/ros2-multimodal-robot-collab, 111 stars), Agents (tadaspetra/loop, 296 stars), Piper Tts Training (sammcj/agentic-coding, 162 stars) and Agents (elevenlabs/skills, 482 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
zenstory-ai (a GitHub organization) maintains it in zenstory-ai/video-recap-skills, which has 561 GitHub stars. The repository holds 7 skills in this directory. The repository was last updated on October 4, 2026.
Source: zenstory-ai/video-recap-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.