Install the "video-production" agent skill from https://github.com/speechlab0210/video-production-skill/tree/main into .claude/skills/video-production/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "video-production", then confirm the skill loads.
Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
skills CLI
$ npx skills add speechlab0210/video-production-skill --skill video-production -a codex
Project install goes to .agents/skills/; add -g for ~/.codex/skills/.
Install the "video-production" agent skill from https://github.com/speechlab0210/video-production-skill/tree/main into .agents/skills/video-production/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "video-production", then confirm the skill loads.
Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
skills CLI
$ npx skills add speechlab0210/video-production-skill --skill video-production -a cursor
Project install goes to .agents/skills/; add -g for ~/.cursor/skills/.
Install the "video-production" agent skill from https://github.com/speechlab0210/video-production-skill/tree/main into .cursor/skills/video-production/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "video-production", then confirm the skill loads.
Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
skills CLI
$ npx skills add speechlab0210/video-production-skill --skill video-production -a gemini-cli
Project install goes to .agents/skills/; add -g for ~/.gemini/skills/.
Install the "video-production" agent skill from https://github.com/speechlab0210/video-production-skill/tree/main into .gemini/skills/video-production/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "video-production", then confirm the skill loads.
Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
Installs for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
skills CLI
$ npx skills add speechlab0210/video-production-skill --skill video-production -a github-copilot
Project install goes to .agents/skills/; add -g for ~/.copilot/skills/.
Install the "video-production" agent skill from https://github.com/speechlab0210/video-production-skill/tree/main into .github/skills/video-production/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "video-production", then confirm the skill loads.
GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
skills CLI
$ npx skills add speechlab0210/video-production-skill --skill video-production -a opencode
OpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
Install the "video-production" agent skill from https://github.com/speechlab0210/video-production-skill/tree/main into .opencode/skills/video-production/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "video-production", then confirm the skill loads.
OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
Facts
Skill name
video-production
GitHub stars
105
Token cost
~4.1k tokens
SKILL.md length
1,608 words
Files
22 (incl. scripts, references)
Skills in repo
1
Repo updated
First seen
Licence
MIT
At a glance
AI educational video production pipeline. An agent skill from speechlab0210/video-production-skill.
Works in 8 steps: Script (narration.json) → Slides → TTS + ASR Verification… → …
Producing lecture videos
SKILL.md covers Quick Start, Requirements, Configuration and Checklist (mandatory, do not…, plus 12 more sections
Runs JavaScript and Python scripts from its folder; calls node, python and ffprobe; needs ELEVENLABS_API_KEY and OPENAI_API_KEY
What it does
Video Production is an agent skill from speechlab0210/video-production-skill. AI educational video production pipeline. Use when producing lecture videos, tutorial videos, or educational content. Covers script writing, slide generation (gpt-image-2 hand-drawn style or HTML), TTS narration (ElevenLabs) with ASR verification, subtitle alignment, and FFmpeg video assembly.
Its SKILL.md is about 4.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 25 other files, including scripts and reference files (for example `README.md`, `examples/demo/README.md` and `examples/demo/narration.json`).
It sits in Media & Creative, covering Text to speech and voice, Video production and Slides and decks. It works with FFmpeg and ElevenLabs. The repository describes itself as: The complete video-production skill behind the 蝦說 AI channel — lets an AI agent autonomously produce narrated educational videos (slides + TTS + ASR verification + subtitles). The licence is MIT.
When your agent uses it
Producing lecture videos
Tutorial videos
Educational content
Example prompts
“/video-production”
Requirements
Python 3
Node.js
A credential in ELEVENLABS_API_KEY
A credential in OPENAI_API_KEY
Workflow steps
8 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 8c46c40. It shows what the files ask for, not the result of running them.
Tool permissions
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Runs code
Ships 4 files in scripts/ (JavaScript and Python, from the files we listed), which the agent can run.
Shell commands in SKILL.md call:
node
python
ffprobe
ffmpeg
From the folder's file list and the shell code blocks in SKILL.md.
Network
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Credentials
Names these keys or tokens, usually read from environment variables:
ELEVENLABS_API_KEY
OPENAI_API_KEY
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Context cost
Video Production loads about 4.1k tokens when it runs, and up to ~14k if it reads all its reference files. Until then it costs about 78 tokens; SKILL.md has 1,608 words of instructions outside code blocks.
Always· name and description, kept in context so the agent knows when to use it
~78
When it runs· the whole SKILL.md, loaded when a task matches
~4.1k
With references· SKILL.md plus every file in references/, read only if the agent opens them
~14k
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
Safety
Auto-check: notes
The automated check noted patterns worth knowing about, such as sudo or a known installer.
NoteMentions a .env fileSKILL.md:34
API keys** (environment variables, or a `.env` file in the project directory):
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
Download SKILL.mdSave it as .claude/skills/video-production/SKILL.md (or your agent's skills folder). This skill also uses 21 other files; get the full folder from GitHub.
name
video-production
description
AI educational video production pipeline. Use when producing lecture videos, tutorial videos, or educational content. Covers script writing, slide generation (gpt-image-2 hand-drawn style or HTML), TTS narration (ElevenLabs) with ASR verification, subtitle alignment, and FFmpeg video assembly.
Video Production Skill
End-to-end pipeline for producing AI-narrated educational videos with slides.
Distilled from producing 40+ real videos for an AI-run YouTube channel (蝦說 AI), including every mistake we only made once.
Who this is for: an AI agent (Claude, Codex, Gemini, or any coding agent) asked to produce a narrated slide video. A human can follow it too.
Quick Start
0. Read references/teaching-style.md + references/narration-style.md (content quality rules)
1. Create project directory and cd into it; copy config.json from references/config-example.json and fill it in
2. Write narration.json (the script — one string per slide)
3. Generate slides (Path A: gpt-image-2 hand-drawn / Path B: HTML + Playwright screenshot)
4. Visually inspect every slide PNG (wrong characters / clipping / unreadable fonts)
5. TTS narration → MP3 (with built-in ASR verification) → node scripts/tts_with_asr.js
6. FFmpeg assemble → video.mp4 → node scripts/assemble.js
7. Quality check (bitrate + frame extraction + visual)
8. Subtitles → SRT (+ optional burn-in) → node scripts/gen_subtitles.js
9. Cover image → thumbnail → python scripts/cover_gen.py
10. Upload wherever you publish; verify the thumbnail and visibility after upload
Requirements
Node.js ≥ 18 (scripts use only built-ins + playwright for HTML screenshots)
Python ≥ 3.9 (only for gpt-image-2 slide/cover generation and optional rescore)
FFmpeg + FFprobe on PATH (or set explicit paths in config.json)
API keys (environment variables, or a .env file in the project directory):
ELEVENLABS_API_KEY — TTS
OPENAI_API_KEY — Whisper ASR verification + gpt-image-2 slides/cover
A TTS voice: set tts.voiceId in config.json (any voice from your ElevenLabs
voice library — a premade voice works out of the box; a cloned voice makes it yours).
⚠️ Never hardcode API keys in scripts and never commit them to git. One of our keys
was auto-revoked from a public repo push and 32 videos rendered silent. Environment
variables only.
Configuration
Create config.json in your project directory (start from references/config-example.json):
□ 0. Read references/teaching-style.md — content/teaching rules (pre-flight constitution)
□ 1. Write narration.json
□ 1.5 ⭐ ALIGNMENT CHECK: narration entries count == slide count. MUST be equal!
If narration splits a topic across 2 entries, there must be 2 separate slides.
Mismatch = audio/visual desync for every slide after the mismatch point.
□ 2. Generate slides (Path A or B below)
□ 3. ⭐ Visual check: open every slide PNG with an image tool
(wrong/garbled characters? clipped edges? readable on a phone?)
□ 4. TTS synthesis + ⭐ ASR verification (similarity ≥ 0.85, never skip)
□ 5. FFmpeg assemble → video.mp4
□ 6. ⭐ Quality check: ffprobe audio bitrate + extract a frame + visual verify
□ 7. Subtitles → subtitles_aligned.srt (external track and/or burn-in)
□ 8. ⭐ Cover image → thumbnail (16:9; pad, don't crop)
□ 9. Upload (start unlisted/private, review, then publish)
□ 10. ⭐ After upload: verify the thumbnail actually shows YOUR cover, not an auto frame
The ⭐ steps are non-negotiable. Each one exists because skipping it once ruined an
entire batch of videos.
🔴 The cover is part of the video, not an optional extra. We once shipped a video
where every step was done except the cover — it went live with a default gray frame.
"No cover = not shipped."
Step 1: Script (narration.json)
A JSON array of strings, one per slide (CRITICAL: array length MUST equal slide count):
json
[
"開場白,吸引注意力...",
"第一個重點,配合比喻...",
"第二個重點...",
"結尾,call to action..."
]
Writing guidelines
80–150 characters per slide (Chinese). Too short = rushed; too long = TTS breaks.
TTS engines mispronounce things. Prevent it at the script stage:
Problem
Fix
English abbreviations (LLM, NLP)
Spell out in Chinese, or letters with periods (P.U.A.)
Raw numbers (135,000)
Chinese numerals (十三萬五千)
Version numbers (4.5)
Display text keeps 4.5, TTS text gets 四點五
Long sentences (>50 chars)
Split at natural breath points
Parenthetical content
Rewrite as natural speech
Chinese heteronyms 破音字 (還/重/長/得…)
Scan with references/heteronyms.json, rewrite
Punctuation and pacing: our TTS script strips punctuation before synthesis (each
punctuation mark becomes a pause in Chinese TTS — dense commas produce machine-gun
narration). Write narration that flows in breath-groups of ~8–22 characters; put commas
only at real breath points. Subtitles still use the original punctuated text.
Step 2: Slides
Two first-class paths. Pick ONE per video (Path A is our channel default; Path B has no
image-API dependency).
Path A — gpt-image-2 hand-drawn teaching style (scripts/slides_gen.py)
Full-bleed AI-generated slides in a "professor's hand-drawn lecture notes" style:
white background, bold black CJK title top-left with underline, thin black arrows,
lots of whitespace, a small mascot in the corner, stick figures only (no real faces).
Write one Chinese prompt per slide into slides_prompts.json (array of strings).
Wrap every string that must appear on the slide in 「」. Always end the shared style
block with: 「所有中文字必須完全正確、清楚可讀、不可有亂碼或錯字。數字要正確。」
python scripts/slides_gen.py → generates slides_raw/slide_NN.png (1536×1024).
Batch ≤4–5 concurrent lanes (~120s per image). Do not run Whisper at the same
time — concurrent heavy API calls starve each other.
Visually inspect every image (garbled characters / wrong or invented numbers /
typos → regenerate just that slide; HTTP 502 → just retry that slide).
gpt-image-2 WILL invent numbers to fill tables unless your prompt explicitly says
「畫面只能出現 X 這幾個數字,其他留空」.
node scripts/pad_and_burn.js pad → 1536×1024 scaled to 1410×940, padded onto a
1920×1080 white canvas with a 140px bottom band reserved for subtitles.
Path B — HTML slides + Playwright screenshot (scripts/screenshot.js)
Create one HTML file per slide: slides/slide_01.html, slide_02.html, …
(template: references/slide-template.html), then node scripts/screenshot.js.
Design rules (hard-won; mobile viewers are the majority):
Fill 80%+ of the canvas — padding: 60-80px, flexbox centering, no floating cards.
Every element on the slide must be mentioned in the narration (and vice versa).
🔴 Never use <24px body text. Fewer points in bigger font > more points in tiny font.
Path C — Node Canvas fallback (scripts/generate_slides.js)
No browser, no image API: renders simple title+bullets slides from slides.json.
Use only when Playwright and the image API are both unavailable.
Step 3: TTS + ASR Verification (scripts/tts_with_asr.js)
bash
node scripts/tts_with_asr.js [project_dir]
Reads narration.json, synthesizes each entry via ElevenLabs, saves audio/slide_XX.mp3,
then verifies every clip with Whisper ASR:
Transcribe the generated audio
Compute character-overlap similarity vs the original text
≥ 0.85 = PASS (target ≥0.90 for important videos)
< 0.85 = FAIL → adjust wording and retry (up to tts.maxRetries)
Adjustment rules: swap synonyms / split long sentences / write numbers in Chinese —
but never change the meaning, never drop information.
Don't chase 0.85 forever — verify the words, ship on redundancy. Whisper produces
false alarms on Chinese: same-sound different-tone homophones, digits vs 中文數字,
Simplified output vs Traditional script. If the ASR "errors" are homophones and every
key word/number comes through, the audio is correct — keep the best attempt. The slide
shows the number visually and the subtitle uses the original text, so the viewer has
triple redundancy. For digit-heavy narration, run python scripts/rescore.py — it
strips numerals and compares toneless-pinyin multisets (≥0.90 = pass), which kills
most false failures. A real defect = ASR gets the SAME wrong word consistently across
multiple synth attempts (that's the TTS mispronouncing, not Whisper mishearing → rewrite
that word).
Show full SKILL.md (692 more words)Show less
Step 4: FFmpeg Assembly (scripts/assemble.js)
bash
node scripts/assemble.js [project_dir]
Pairs each slides/slide_XX.png with audio/slide_XX.mp3, creates per-slide clips,
concatenates into video.mp4.
Key flags (all already in the script):
-b:a 192k — without it audio can silently render at 2kbps (present but inaudible)
-tune stillimage — much smaller files for slide video
-pix_fmt yuv420p — plays everywhere
-movflags +faststart — video streams/plays inline instead of "link won't open"
🔴 Mixing clips from different sources (e.g. your narrated slides + a real screen
recording)? Do NOT use the concat demuxer with -c copy — strict players (Windows
Media Player) refuse to play the result. Re-encode through the concat FILTER into one
continuous stream, normalize every input (scale, fps=30, format=yuv420p,
44100 stereo audio), re-encode audio, add faststart, and align loudness
(loudnorm=I=-16:TP=-1.5:LRA=11). Details in references/lessons-learned.md.
🔴 Mixing two TTS providers in one video? Their sample rates differ (ElevenLabs
44100 Hz vs OpenAI TTS 24000 Hz). Resample everything to 44100 BEFORE assembly or
some players choke exactly at the voice-switch point.
Step 5: Quality Check (three checks, all required)
Compare verify.png against slides/slide_01.png with an image tool. Content must match.
5c. Font size check
Inspect verify.png: is every text element readable at mobile size?
Step 6: Subtitles (scripts/gen_subtitles.js)
bash
node scripts/gen_subtitles.js [project_dir]
Produces subtitles_aligned.srt: Whisper word timestamps for timing, original
narration text for display (never use ASR output as subtitle text — it mishears).
Line breaks are width-aware (CJK=1, Latin=0.5, ≤16 full-width per line) and never cut
inside an English word.
⚠️ The #1 subtitle bug: drift from assuming clip duration = audio + padding.
FFmpeg's -shortest truncates the padding, so real clip duration ≈ audio duration.
The script therefore reads ACTUAL durations from temp/clip_XX.mp4 with ffprobe.
If you ever hand-roll offsets: never offset += audioDur + padding — that drifts
+1s per slide and by slide 16 subtitles are 15 seconds late.
Ship options:
External SRT track (recommended when your platform supports it — viewers can
toggle, nothing covers the slides)
Burn-in: node scripts/pad_and_burn.js burn (Path A white-band layout: dark text
FontSize 30 / MarginV 30 sits inside the reserved 140px band)
For full-bleed dark HTML slides use FontSize=14–18, MarginV=6–12, BorderStyle=3 instead.
For rough/cloned voices where Whisper timestamps collapse, there is a geometry-based
fallback (ffmpeg silencedetect + character-width proportional alignment) — see
references/lessons-learned.md § subtitle alignment.
Generate at 1536×1024, then pad — do not crop — to 1280×720:
crop cuts off the top of your title. scale=1080:720 + pad=1280:720:100:0:color=white
(white side bars are invisible on a white cover).
Target ≤2MB for YouTube.
Same prompt hygiene as slides: quote exact text in 「」, forbid invented numbers,
visually verify before shipping.
Step 8: Upload
Platform-specific; do it however you normally operate (browser automation, manual, CLI).
Regardless of method, these rules survived contact with reality:
Upload as unlisted/private first, review the actual playback, then publish.
Verify the thumbnail after upload — the preview must show YOUR cover, not an
auto-selected frame.
Re-check title/description/language settings after any dialog reopens (some upload
wizards silently reset radio buttons).
If you used an external SRT: upload it as a caption track ("with timing"), publish
the track, then verify captions actually appear on the watch page.
Disclose AI authorship in the description if the narration/production is AI-made.
Output Structure
my-video-project/
├── config.json ← project config (from references/config-example.json)
├── narration.json ← script text per slide
├── slides_prompts.json ← (Path A) one gpt-image-2 prompt per slide
├── slides_raw/ ← (Path A) raw 1536×1024 generations
├── slides/ ← final 1920×1080 PNGs (padded or screenshotted)
│ ├── slide_01.png
│ └── ...
├── audio/
│ ├── slide_01.mp3
│ └── ...
├── temp/ ← per-slide clips + whisper word caches
├── video.mp4 ← assembled (no subtitles)
├── subtitles_aligned.srt ← aligned subtitle track
├── video_sub.mp4 ← (optional) burned-in version
└── thumbnail.jpg ← cover, 1280×720
Scripts Reference
All scripts take the project directory as an optional first argument (default: CWD).
Script
Purpose
scripts/slides_gen.py
gpt-image-2 slide generation from slides_prompts.json
scripts/pad_and_burn.js
pad 3:2 images to 16:9 + subtitle band / burn SRT
scripts/screenshot.js
Playwright HTML→PNG screenshots
scripts/generate_slides.js
Node Canvas fallback slide renderer
scripts/tts_with_asr.js
ElevenLabs TTS + Whisper ASR verification loop
scripts/assemble.js
FFmpeg per-slide clips + concat
scripts/gen_subtitles.js
aligned SRT (whisper timing + original text)
scripts/rescore.py
homophone/digit-tolerant second-chance ASR scoring
scripts/cover_gen.py
gpt-image-2 cover generation
Sub-agents
If you spawn a sub-agent to produce a video, the task prompt must explicitly say
"read the video-production SKILL.md first". Sub-agents that aren't told skip the
pipeline and reinvent (worse) wheels. Re-read the skill fresh each session — it evolves.
When things break
references/lessons-learned.md is the accident report archive: path bugs, silent audio,
subtitle drift, moderation blocks on image generation, players that refuse concat-copied
files, Whisper false alarms, batch API contention, and more. Read it before your first
video; grep it when something looks weird — the odds are good we already hit it.
Video Production next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
Video Production compared with similar skills
Skill
Stars
Used in
Tokens
Auto-check
Licence
Repo updated
Video Production this skillspeechlab0210/video-production-skill
Render a 'hand-swipe flavor listicle' video from a config — the brand's REAL product cutouts each float on a flat flavor-matched color field under a persistent title and slide RIGHT-TO-LEFT one to…
Render an 'Instagram-Live social-proof gallery' video from a config — ~5 real brand product stills each framed as an Instagram-LIVE card (IG gradient-ring avatar, username, verified check, red LIVE…
Render a 'search-grid' (Pinterest search-moodboard) video from a config — a real-DOM page with four continuous beats (masonry search grid + typing hook with counter-drift columns → 3 cards slide in…
Turns a research paper into a slide deck and, optionally, a narrated demo video, through script, slide generation, text-to-speech and video assembly stages you control.
把一句话需求转成可复现、可验收视频工程的乔木智能剪辑导演。Use when the user asks to create, plan, edit, remix, explain, narrate, subtitle, animate, composite, or render a video—including one-line requests such as “制作一个科普视频:介绍…
AI educational video production pipeline. An agent skill from speechlab0210/video-production-skill. Video Production is an agent skill from speechlab0210/video-production-skill. AI educational video production pipeline.
When should I use Video Production?
Video Production fits situations like: producing lecture videos; tutorial videos; educational content.
How do I install Video Production in Claude Code?
Run `npx skills add speechlab0210/video-production-skill --skill video-production -a claude-code`. Or copy the skill folder (the speechlab0210/video-production-skill repository) into .claude/skills/video-production in your project. Claude Code loads it when a task matches its description.
How do I install Video Production in Codex?
Run `npx skills add speechlab0210/video-production-skill --skill video-production -a codex`. Or copy the skill folder (the speechlab0210/video-production-skill repository) into .agents/skills/video-production in your project. Codex loads it when a task matches its description.
Can I use Video Production in Cursor, Gemini CLI or GitHub Copilot?
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add speechlab0210/video-production-skill --skill video-production -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/video-production, .gemini/skills/video-production, .github/skills/video-production and .opencode/skills/video-production in your project.
What does Video Production need to run?
Going by SKILL.md and its folder, Video Production needs JavaScript and Python for the scripts in its folder, the command-line tools its instructions call (node, python, ffprobe and ffmpeg) and credentials named ELEVENLABS_API_KEY and OPENAI_API_KEY. Our summary lists: Python 3; Node.js; A credential in ELEVENLABS_API_KEY; A credential in OPENAI_API_KEY.
Does Video Production access the network?
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Is Video Production safe to install?
Our automated static check of SKILL.md found notes only (mentions a .env file), nothing it rates as a warning. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
What licence does Video Production use?
Video Production is published under the MIT licence (from the LICENSE file in the skill folder). It allows redistribution, so the full SKILL.md is shown on this page.
How many tokens does Video Production use?
About 4.1k tokens (SKILL.md is roughly 17k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 10k tokens, read only when the agent opens those files.
What are the alternatives to Video Production?
Skills that share tags, products or a category with Video Production: Render Hand Swipe Flavor Listicle (gooseworks-ai/goose-skills, 1.2k stars), Render Ig Live Gallery (gooseworks-ai/goose-skills, 1.2k stars), Render Search Grid (gooseworks-ai/goose-skills, 1.2k stars) and Academic Presentation Maker (OpenLAIR/dr-claw, 1.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
Who maintains Video Production?
speechlab0210 (a GitHub user) maintains it in speechlab0210/video-production-skill, which has 105 GitHub stars. The repository was last updated on July 3, 2026.