Agent skill

Talking-Head Video Pipeline

by naive-kun in naive-kun/naive-video-skill

Turns raw or rough-cut talking-head footage into a captioned, animated final video through a staged, resumable production pipeline.

MITAuto-check passedMedia & Creative

Install Talking-Head Video Pipeline

skills CLI
$ npx skills add naive-kun/naive-video-skill --skill talking-head-video-pipeline -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install naive-kun/naive-video-skill talking-head-video-pipeline --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
talking-head-video-pipeline
GitHub stars
132
Token cost
~3.3k tokens
SKILL.md length
1,485 words
Files
77 (incl. scripts, references)
Skills in repo
1
Repo updated
First seen
Licence
MIT

At a glance

Turns raw or rough-cut talking-head footage into a captioned, animated final video through a staged, resumable production pipeline.

  • Works in 12 steps: Source is immutable. Never overwrite or… → Main audio is the clock. Do not retime,… → Evidence stays readable. Never cover… → …
  • Turning rough talking-head footage into a captioned final video
  • SKILL.md covers Beginner Promise, Non-Negotiable Rules, Route Table and Mode Detection, plus 6 more sections
  • Runs Shell scripts from its folder; calls python3 and bash

What it does

This skill runs talking-head video production as a staged pipeline for someone who may never have edited video with an agent before: it asks one question at a time only when it can't infer the answer, defaults to a working choice over a menu of technical options, and explains the next visible result rather than renderer internals.

It treats the main audio track as the immovable clock and the original source footage as untouchable, optionally helps with a rough cut, transcription and caption revision, groups content for the viewer, places screenshots and demos, builds a staged keyframe and HyperFrames or GSAP preview, and only produces a synchronized final export after that preview is approved. Work can be interrupted and resumed, diagnosed, and revised within a limited scope rather than restarted from scratch, and it can optionally pull in raw footage through a separate video-use integration.

It keeps project paths, media names, brand rules, screenshots and API keys inside the project rather than inside the skill, and it only learns a lasting style preference from explicit feedback, never from silence or a single unconfirmed draft.

When your agent uses it

  • Turning rough talking-head footage into a captioned final video
  • Resuming a talking-head edit that was interrupted partway through
  • Reviewing and revising captions or keyframes before a final export

Example prompts

  • “Turn this raw talking-head recording into a captioned final video.”
  • “Resume the video edit we started yesterday from where it stopped.”
  • “Revise the captions in the third section and re-render the preview.”

Workflow steps

12 steps, taken from the first numbered list in SKILL.md.

  1. Source is immutable. Never overwrite or destructively edit the supplied source media.
  2. Main audio is the clock. Do not retime, reorder, resample, or replace it unless the user explicitly approves a rough-cut strategy. After…
  3. Evidence stays readable. Never cover screenshot body text, product UI, faces, or user-marked safe zones. Do not blur evidence assets.
  4. Preview before final. Build an official preview and wait for approval unless the user explicitly requests direct export.
  5. Original-quality final. Do not present an upscaled low-resolution preview as a true high-resolution final. Use the original source as the…
  6. Do not interrupt progress. If a render is advancing, wait and report progress. Change pipelines only after a real error or explicit…
  7. No private data in the skill. Project-local paths, media names, brand rules, screenshots, customer details, and API keys stay in the…
  8. Learn only from explicit feedback. Do not infer permanent style rules from silence or one unconfirmed draft.
  9. Typography is part of correctness. Declare and verify font family, real weight, caption line policy, baseline alignment, and longest-label…
  10. Public defaults stay brand-neutral. Store creator palettes, recurring card systems, and private visual rules only in the local profile.
  11. Public install exposes one Skill. Keep stage modules under references/workflows/; do not add nested SKILL.md files that package managers…
  12. Install operations are recoverable. Back up an existing installation before replacement and archive uninstall targets instead of…

What it can do on your machine

Read from SKILL.md and the folder at commit 3e9c7c5. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Shell, from the files we listed), which the agent can run.

    Shell commands in SKILL.md call:

    • python3
    • bash

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Talking-Head Video Pipeline loads about 3.3k tokens when it runs, and up to ~28k if it reads all its reference files. Until then it costs about 161 tokens; SKILL.md has 1,485 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~161
When it runs · the whole SKILL.md, loaded when a task matches
~3.3k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~28k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from naive-kun/naive-video-skill at commit 3e9c7c5, republished under its MIT licence (© naive-kun). 1,485 words, ~3,320 tokens.

Download SKILL.mdSave it as .claude/skills/talking-head-video-pipeline/SKILL.md (or your agent's skills folder). This skill also uses 76 other files; get the full folder from GitHub.
name
talking-head-video-pipeline
description
End-to-end, stateful talking-head video production for complete beginners and experienced editors. Use when a user wants optional rough-cut guidance, transcription or caption revision, viewer-facing content logic groups, screenshot and demo placement, visual design, staged keyframe plus HyperFrames/GSAP preview, a synchronized final export, interrupted-work recovery, diagnosis, scoped revision, or reusable style learning from explicit feedback. Routes through internal workflow modules, optionally integrates video-use for raw footage, and preserves originals, evidence readability, privacy, and preview approval.

Talking Head Video Pipeline

Turn raw or rough-cut talking-head footage into an approved working cut, editable captions, a motion-packaged preview, and a verified final export. Treat this file as the router and shared contract. Read only the routed internal workflow and references needed for the current stage.

Beginner Promise

Assume the user has never edited with Codex.

  • Ask one short question at a time only when an answer cannot be inferred safely.
  • Prefer a working default over presenting many technical choices.
  • Explain the next visible result, not the renderer internals.
  • Introduce rough cutting, transcription providers, or API setup only when raw footage actually needs that branch.
  • Explain paid and local choices plainly, recommend one when useful, and never force it.
  • End every stage with exact output paths and the next phrase the user can say.
  • Resume existing work instead of restarting it.

Non-Negotiable Rules

  1. Source is immutable. Never overwrite or destructively edit the supplied source media.
  2. Main audio is the clock. Do not retime, reorder, resample, or replace it unless the user explicitly approves a rough-cut strategy. After approval, the derived working_video becomes the downstream clock while the original remains immutable.
  3. Evidence stays readable. Never cover screenshot body text, product UI, faces, or user-marked safe zones. Do not blur evidence assets.
  4. Preview before final. Build an official preview and wait for approval unless the user explicitly requests direct export.
  5. Original-quality final. Do not present an upscaled low-resolution preview as a true high-resolution final. Use the original source as the final base whenever possible.
  6. Do not interrupt progress. If a render is advancing, wait and report progress. Change pipelines only after a real error or explicit approval.
  7. No private data in the skill. Project-local paths, media names, brand rules, screenshots, customer details, and API keys stay in the user's project.
  8. Learn only from explicit feedback. Do not infer permanent style rules from silence or one unconfirmed draft.
  9. Typography is part of correctness. Declare and verify font family, real weight, caption line policy, baseline alignment, and longest-label fit before approval.
  10. Public defaults stay brand-neutral. Store creator palettes, recurring card systems, and private visual rules only in the local profile.
  11. Public install exposes one Skill. Keep stage modules under references/workflows/; do not add nested SKILL.md files that package managers may register separately.
  12. Install operations are recoverable. Back up an existing installation before replacement and archive uninstall targets instead of performing recursive force deletion.
  13. Optional providers stay optional. The curated default shot pack is native HyperFrames + GSAP. Never clone, install, update, or require ShotCraft, Remotion, or another provider without explicit user consent. Native recipes must remain a complete fallback.

Read references/quality-gates.md before preview or export work.

Route Table

User intent or phraseRouteTypical output
初始化视频项目, 第一次用, start a video projectreferences/workflows/init.mdState, edit plan, design profile
原片还没剪, 删口误和停顿, 帮我粗剪, raw takes, rough cutreferences/workflows/rough-cut.mdApproved non-destructive rough cut and timing handoff
抽字幕轴, 只做字幕, transcribe, 改字幕references/workflows/captions.mdSRT, CSV, transcript JSON
帮我设计, 换颜色风格, 规划弹窗, designreferences/workflows/design.mdDESIGN and complete edit plan
参考这张截图设计, 按这张图的视觉语言, style referencereferences/workflows/design.md reference branchProject-local STYLE_REFERENCE
按语义匹配动效, GSAP 动效丰富一点, semantic motionreferences/workflows/design.md, then previewValidated MOTION_PLAN and preview
参考 ShotCraft, 推荐电影感镜头, cinematic shot referencereferences/workflows/design.md ShotCraft branchOptional mapped shot references in MOTION_PLAN
做预览, 加动效, build previewreferences/workflows/preview.mdOfficial preview URL
出成片, 导出 4K, export finalreferences/workflows/export.mdVerified final video
改这个, 恢复上一版, revisereferences/workflows/revise.mdScoped revision and new preview/final
进度, 到哪了, statusreferences/workflows/status.mdCurrent stage and next action
体检, 为什么失败, doctorreferences/workflows/doctor.mdRead-only diagnosis
以后都这样, 记住这个风格, learn thisreferences/workflows/learn.mdConfirmed project-local lessons
复盘成片, 下次怎么改进, retroreferences/workflows/retro.mdStructured delivery retrospective
升级状态, 迁移, migratereferences/workflows/migrate.mdState schema upgrade

Mode Detection

Before routing:

  1. Identify the user's video project directory, not this skill repository.
  2. Look for .naive-video-state.json.
  3. If absent and the task is more than a one-off caption text edit, load the init workflow.
  4. If present, read it plus EDIT_PLAN.md, DESIGN.md, and VIDEO_LESSONS.md only when relevant. Use working_video when present; otherwise use main_video.
  5. Run python3 tools/video_doctor.py --project <project_dir> when state and files disagree.

Use these stages:

text
initialized -> captions_ready -> design_ready -> preview_ready
-> approved -> rendering -> final_ready

Never mark a later stage until its quality gate passes.

Default Pipeline

When the user asks for the complete workflow:

  1. Inspect source media with ffprobe.
  2. Initialize project state and safe output directories.
  3. Ask whether the footage is already rough-cut only when that status is unclear. If it is raw, load the optional rough-cut workflow and require strategy approval.
  4. Produce or validate caption files from working_video when present, otherwise main_video.
  5. Ask whether screenshots, recordings, product images, charts, or demos must appear. Let the user choose semantic placement, exact seconds/spoken-sentence anchors, or a hybrid; explain that exact anchors are more precise.
  6. Normalize word-level timing when available and split the speech into viewer-facing logic groups. Do not fake word precision from whole-caption starts.
  7. Collect or infer a style profile and safe zones. Offer optional screenshot-reference matching to beginners.
  8. Write an edit plan and semantic motion plan with every requested insert, logic-group link, and time range.
  9. For new projects, automatically map a few safe semantic nodes to the curated ShotCraft-inspired native pack unless the user chooses skip. Preserve existing approved projects, the native recipe, and offline fallback.
  10. For first-use, new-style, or reasoning-heavy work, render representative static keyframes and resolve composition problems before a full dynamic preview.
  11. Build a HyperFrames + GSAP preview using a browser-friendly proxy if needed.
  12. Run preview quality gates and provide the official preview URL.
  13. Wait for approval unless explicitly told to skip.
  14. Render overlays or final composition using the approved working video and its audio clock.
  15. Verify file existence, duration, resolution, frame rate, and audio.
  16. Record only confirmed feedback as project-local lessons.
  17. After delivery, run a factual retrospective and promote only explicit, privacy-safe rules.
Show full SKILL.md (515 more words)Show less

Project Files

The init workflow creates this minimal project memory without touching source media:

text
<project_dir>/
├── .naive-video-state.json
├── EDIT_PLAN.md
├── DESIGN.md
├── CONTENT_LOGIC.json       # viewer-facing reasoning groups and timed beats
├── STYLE_REFERENCE.md       # created only when a reference image is used
├── MOTION_PLAN.json         # created when semantic motion is planned
├── VIDEO_LESSONS.md
├── VIDEO_RETRO.md
├── edit/
│   ├── rough-cut.mp4            # optional; original media is never overwritten
│   ├── rough-cut-edl.json       # optional decision record
│   ├── script-aligned.srt
│   ├── caption-table.csv
│   └── transcripts/
├── preview/
├── final/
└── qa/
    └── KEYFRAME_REVIEW.md   # static-composition review before dynamic preview

State is operational metadata, not a media database. Keep it small and never store transcript bodies, private screenshots, or secrets in it. See references/state-management.md.

Style Defaults

If the user wants speed and gives no reference:

  • Preserve the source aspect ratio.
  • Use a clean, readable card system with one configurable accent color.
  • Use bold captions with restrained keyword emphasis.
  • Use GSAP transforms and opacity for motion.
  • Place cards in verified empty space.
  • Reduce motion density while screenshots or demos are visible.
  • Keep readable text level and baseline-aligned; animate the containing component instead of skewing labels.
  • Choose an explicit caption maximum line count and wrap policy.

Ask for an accent color only if brand consistency matters. Otherwise use the neutral preset and make it easy to change later. Beginners may optionally provide a screenshot and choose low, medium, or high reference strength; explain that the skill copies visual language, never the source brand or content. Motion density is restrained, balanced, or energetic. See references/style-onboarding.md.

Detailed GSAP recipes and semantic mappings live in references/motion-recipes.md. Runtime and plugin selection live in references/gsap-runtime.md. Screenshot extraction and anti-copy rules live in references/style-reference-workflow.md. Asset timing choices live in references/asset-onboarding.md. Load them only for the matching design branch.

ShotCraft discovery and adaptation rules live in references/shotcraft-integration.md; the native beginner pack lives in references/shotcraft-default-pack.md. Load them for new automatic projects, named cards, or cinematic shot requests. Never install a provider silently.

Word-level timing, viewer-facing logic groups, and accumulation/exit behavior live in references/content-logic-workflow.md. Load it after captions and before semantic design.

Typography, component geometry, glass notifications, and seek-safe focus/type/split adaptation live in references/visual-quality-rules.md. Load it for every design or preview task.

Self-Iteration Contract

Self-iteration means improving the current user's workflow without leaking it into public defaults.

Classify feedback into one scope:

  • project: applies only to this video.
  • profile: applies to this user's future videos; write it only after explicit confirmation.
  • product: a privacy-safe, general reliability improvement; propose it to the skill maintainer separately.

Record profile feedback in VIDEO_LESSONS.md with the user's words, the confirmed rule, and the affected stage. Never copy media paths or private evidence into this repository. See references/self-iteration.md.

Use the retro workflow after delivery to separate failures, environment issues, and taste feedback before promoting any rule. New projects import active private-profile rules into VIDEO_LESSONS.md so the editor can actually apply them.

Recovery Rules

  • If preview stops responding, check the process and port before rebuilding.
  • If a browser cannot decode HEVC/Main10 smoothly, make an H.264 proxy for preview only.
  • If a render session is running, poll it; do not terminate it merely because it is slow.
  • If partial frames exist, compare expected and actual frame counts before resuming or restarting.
  • Move damaged partial outputs to a project-local quarantine only after confirming they are not the sole valid result.
  • Keep logs outside command pipes that can terminate long jobs.

Privacy Before Publishing

Run:

bash
bash scripts/doctor.sh --privacy-scan .
python3 tools/validate_skill.py .

Block publication if either command reports personal absolute paths, secrets, private media names, invalid frontmatter, multiple skill manifests, missing internal workflows, unsafe install commands, or broken templates.

© naive-kun, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 76 other files (scripts, references) in the repository root of naive-kun/naive-video-skill.

  • SKILL.md
  • .github/workflows/validate.yml
  • .gitignore
  • CHANGELOG.md
  • CONTRIBUTING.md
  • LICENSE
  • README.md
  • VERSION
  • agents/openai.yaml
  • examples/prompts.md
  • install.sh
  • migrations/0-to-1.0.md
  • migrations/registry.md
  • references/asr-adapters.md
  • references/asset-onboarding.md
  • … and 62 more

Open the folder on GitHubat commit 3e9c7c5

Compare with similar skills

Talking-Head Video Pipeline next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Talking-Head Video Pipeline compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Talking-Head Video Pipeline this skillnaive-kun/naive-video-skill132—~3.3kAutomated safety check: PassMIT
Embedded Video Captionsheygen-com/hyperframes60k3 repos~8.6kAutomated safety check: PassApache-2.0
Yuv Viral Videohoodini/ai-agents-skills282—~7.5kAutomated safety check: NotesNone
Noti Tiktok Full Textnotivn/AIEV127—~4.8kAutomated safety check: PassMIT
Noti Tiktok Vnnotivn/AIEV127—~5.4kAutomated safety check: PassMIT
HyperFrames Animationheygen-com/hyperframes60k3 repos~2.1kAutomated safety check: PassApache-2.0

Similar skills

  • Embedded Video Captions

    heygen-com/hyperframes

    Adds captions to a single-subject talking-head video without editing the footage, from plain subtitles to cinematic text placed behind the speaker.

    60k GitHub starsUsed in 3 repos~8.6k tokens
    Media & CreativeAuto-check passed
  • Yuv Viral Video

    hoodini/ai-agents-skills

    Edit any selfie or screen-share footage into a viral short-form video in YUV.AI's signature style — Apple-style liquid-glass cards (real CSS backdrop-filter), dark-mode polish, MrBeast-paced cuts…

    282 GitHub stars~7.5k tokensUpdated 3 mo ago
    Media & CreativeAuto-check: notes
  • Build a Vietnamese vertical TikTok explainer in the "MỔ XẺ PAPER AI" (AI paper dissection) format with HyperFrames (HTML/CSS/GSAP → MP4), Noti.vn style.

    127 GitHub stars~4.8k tokensUpdated yesterday
    Media & CreativeAuto-check passed
  • Noti Tiktok Vn

    notivn/AIEV

    Edit a Vietnamese vertical TikTok video (9:16) with HyperFrames following the Noti.vn/GĐT standard - talking-head + kinetic typography + karaoke captions + zoom/punch-in camera + timestamp-synced…

    127 GitHub stars~5.4k tokensUpdated yesterday
    Media & CreativeAuto-check passed
  • HyperFrames Animation

    heygen-com/hyperframes

    Collects motion rules, scene blueprints, transitions and runtime adapters for HyperFrames video compositions, with GSAP as the default animation runtime.

    60k GitHub starsUsed in 3 repos~2.1k tokens
    Media & CreativeAuto-check passed
  • Figma to HyperFrames

    heygen-com/hyperframes

    Imports Figma assets, brand tokens, components and motion into a HyperFrames video composition, using the Figma REST API with a connector or native export for shaders.

    60k GitHub starsUsed in 3 repos~4.5k tokens
    Media & CreativeAuto-check: notes

Works with

Questions about Talking-Head Video Pipeline

What does Talking-Head Video Pipeline do?

Turns raw or rough-cut talking-head footage into a captioned, animated final video through a staged, resumable production pipeline. This skill runs talking-head video production as a staged pipeline for someone who may never have edited video with an agent before: it asks one question at a time only when it can't infer the answer, defaults to a working choice over a menu of technical options, and explains the next visible result rather than renderer internals.

When should I use Talking-Head Video Pipeline?

Talking-Head Video Pipeline fits situations like: turning rough talking-head footage into a captioned final video; resuming a talking-head edit that was interrupted partway through; reviewing and revising captions or keyframes before a final export.

How do I install Talking-Head Video Pipeline in Claude Code?

Run `npx skills add naive-kun/naive-video-skill --skill talking-head-video-pipeline -a claude-code`. Or copy the skill folder (the naive-kun/naive-video-skill repository) into .claude/skills/talking-head-video-pipeline in your project. Claude Code loads it when a task matches its description.

How do I install Talking-Head Video Pipeline in Codex?

Run `npx skills add naive-kun/naive-video-skill --skill talking-head-video-pipeline -a codex`. Or copy the skill folder (the naive-kun/naive-video-skill repository) into .agents/skills/talking-head-video-pipeline in your project. Codex loads it when a task matches its description.

Can I use Talking-Head Video Pipeline in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add naive-kun/naive-video-skill --skill talking-head-video-pipeline -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/talking-head-video-pipeline, .gemini/skills/talking-head-video-pipeline, .github/skills/talking-head-video-pipeline and .opencode/skills/talking-head-video-pipeline in your project.

What does Talking-Head Video Pipeline need to run?

Going by SKILL.md and its folder, Talking-Head Video Pipeline needs a shell for the scripts in its folder and the command-line tools its instructions call (python3 and bash).

Does Talking-Head Video Pipeline access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Talking-Head Video Pipeline safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Talking-Head Video Pipeline use?

Talking-Head Video Pipeline is published under the MIT licence (from the LICENSE file in the skill folder). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Talking-Head Video Pipeline use?

About 3.3k tokens (SKILL.md is roughly 13k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 25k tokens, read only when the agent opens those files.

What are the alternatives to Talking-Head Video Pipeline?

Skills that share tags, products or a category with Talking-Head Video Pipeline: Embedded Video Captions (heygen-com/hyperframes, 60k stars), Yuv Viral Video (hoodini/ai-agents-skills, 282 stars), Noti Tiktok Full Text (notivn/AIEV, 127 stars) and Noti Tiktok Vn (notivn/AIEV, 127 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Talking-Head Video Pipeline?

naive-kun (a GitHub user) maintains it in naive-kun/naive-video-skill, which has 132 GitHub stars. The repository was last updated on August 11, 2026.

Source: naive-kun/naive-video-skill on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.