Embedded Video Captions
heygen-com/hyperframes
Adds captions to a single-subject talking-head video without editing the footage, from plain subtitles to cinematic text placed behind the speaker.
Edits video through conversation: transcribes speech, proposes cuts, color grades, adds overlay animations and burns in subtitles, confirming the plan before each edit.
The automated check flagged lines worth reading first. See the safety section below.
$ npx skills add browser-use/video-use --skill video-use -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install browser-use/video-use video-use --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
Claude Code skills documentation · loads skills from .claude/skills/
Install the "video-use" agent skill from https://github.com/browser-use/video-use/tree/main into .claude/skills/video-use/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "video-use", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add browser-use/video-use --skill video-use -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install browser-use/video-use video-use --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "video-use" agent skill from https://github.com/browser-use/video-use/tree/main into .agents/skills/video-use/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "video-use", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add browser-use/video-use --skill video-use -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install browser-use/video-use video-use --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "video-use" agent skill from https://github.com/browser-use/video-use/tree/main into .cursor/skills/video-use/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "video-use", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add browser-use/video-use --skill video-use -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install browser-use/video-use video-use --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "video-use" agent skill from https://github.com/browser-use/video-use/tree/main into .gemini/skills/video-use/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "video-use", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install browser-use/video-use video-useInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add browser-use/video-use --skill video-use -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "video-use" agent skill from https://github.com/browser-use/video-use/tree/main into .github/skills/video-use/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "video-use", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add browser-use/video-use --skill video-use -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install browser-use/video-use video-use --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "video-use" agent skill from https://github.com/browser-use/video-use/tree/main into .opencode/skills/video-use/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "video-use", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
video-useEdits video through conversation: transcribes speech, proposes cuts, color grades, adds overlay animations and burns in subtitles, confirming the plan before each edit.
The agent works as a video editor that reasons from a transcript. It builds a packed, phrase-level transcript file, takes_packed.md, and picks cut candidates from speech boundaries and silences, looking at visuals only at decision points. It does not assume the type of video, so it examines the material and asks you before editing.
The working loop is ask, confirm, execute, iterate and persist, and the cut is not touched until you have agreed to the strategy in plain language. Most creative choices, such as fonts, colors, durations and techniques, are treated as worked examples to adapt, with ffmpeg and PIL as the underlying helpers, and the agent checks its own output before showing it to you.
A short list of hard rules guards against silent failures: subtitles are applied last in the filter chain so overlays do not hide them, segments are extracted one at a time and joined with lossless stream copy, 30 ms audio fades are applied at each segment boundary to avoid pops, and overlays are time-shifted so an animation starts at its first frame. Helper scripts cover transcription, batch transcription, transcript packing, color grading, timeline view and rendering.
7 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit b877063. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships script files (Python, from the files we listed), which the agent can run.
Shell commands in SKILL.md call:
npxuvpipffmpegFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use npx, uv and pip, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
ELEVENLABS_API_KEYFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Conversational Video Editing loads about 6.4k tokens when it runs. Until then it costs about 84 tokens; SKILL.md has 3,046 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found patterns that need a careful read before installing.
ey can do anything the format supports. Do not wait for permission.olves — either in the environment or in `.env` at the video-use repo root. If missing, ask the user to paste one and wriAutomated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from browser-use/video-use at commit b877063, republished under its MIT licence (© browser-use). 3,046 words, ~6,444 tokens.
.claude/skills/video-use/SKILL.md (or your agent's skills folder). This skill also uses 19 other files; get the full folder from GitHub.takes_packed.md). Everything else — filler tagging, retake detection, shot classification, emphasis scoring — you derive at decision time.These are the things where deviation produces silent failures or broken output. They are not taste, they are correctness. Memorize them.
-c copy concat, not single-pass filtergraph. Otherwise you double-encode every segment when overlays are added.afade=t=in:st=0:d=0.03,afade=t=out:st={dur-0.03}:d=0.03). Otherwise audible pops at every cut.setpts=PTS-STARTPTS+T/TB to shift the overlay's frame 0 to its window start. Otherwise you see the middle of the animation during the overlay window.output_time = word.start - segment_start + segment_offset. Otherwise captions misalign after segment concat.Agent tool; total wall time ≈ slowest one.<videos_dir>/edit/. Never write inside the video-use/ project directory.Everything else in this document is a worked example. Deviate whenever the material calls for it.
The skill lives in video-use/. User footage lives wherever they put it. All session outputs go into <videos_dir>/edit/.
<videos_dir>/
├── <source files, untouched>
└── edit/
├── project.md ← memory; appended every session
├── takes_packed.md ← phrase-level transcripts, the LLM's primary reading view
├── edl.json ← cut decisions
├── transcripts/<name>.json ← cached raw Scribe JSON
├── animations/slot_<id>/ ← per-animation source + render + reasoning
├── clips_graded/ ← per-segment extracts with grade + fades
├── master.srt ← output-timeline subtitles
├── downloads/ ← yt-dlp outputs
├── verify/ ← debug frames / timeline PNGs
├── preview.mp4
└── final.mp4First-time install lives in install.md (clone, deps, ffmpeg, skill registration, API key). Don't re-run it every session; on cold start just verify:
ELEVENLABS_API_KEY resolves — either in the environment or in .env at the video-use repo root. If missing, ask the user to paste one and write it to .env (never to the user's <videos_dir>).ffmpeg + ffprobe on PATH.uv sync or pip install -e . inside the repo).yt-dlp, HyperFrames, Remotion, Manim installed only on first use.npx --yes hyperframes ...; Remotion can be scaffolded with npx create-video@latest or installed as a project-local dependency before using its remotion render command.skills/manim-video/. Read its SKILL.md when building a Manim slot.Helpers (helpers/transcribe.py, helpers/render.py, etc.) live alongside this SKILL.md. Resolve their paths relative to the directory containing this file — the skill is typically symlinked at ~/.claude/skills/video-use/ or ~/.codex/skills/video-use/.
transcribe.py <video> — single-file Scribe call. --num-speakers N optional. Cached.transcribe_batch.py <videos_dir> — 4-worker parallel transcription. Use for multi-take.pack_transcripts.py --edit-dir <dir> — transcripts/*.json → takes_packed.md (phrase-level, break on silence ≥ 0.5s).timeline_view.py <video> <start> <end> — filmstrip + waveform PNG. On-demand visual drill-down. Not a scan tool — use it at decision points, not constantly.render.py <edl.json> -o <out> — per-segment extract → concat → overlays (PTS-shifted) → subtitles LAST. --preview for 720p fast. --build-subtitles to generate master.srt inline.grade.py <in> -o <out> — ffmpeg filter chain grade. Presets + --filter '<raw>' for custom.For animations, create <edit>/animations/slot_<id>/ with Bash and spawn a sub-agent via the Agent tool.
Inventory. ffprobe every source. transcribe_batch.py on the directory. pack_transcripts.py to produce takes_packed.md. Sample one or two timeline_views for a visual first impression.
Pre-scan for problems. One pass over takes_packed.md to note verbal slips, obvious mis-speaks, or phrasings to avoid. Plain list, feed into the editor brief.
Converse. Describe what you see in plain English. Ask questions shaped by the material. Collect: content type, target length/aspect, aesthetic/brand direction, pacing feel, must-preserve moments, must-cut moments, animation and grade preferences, subtitle needs. Do not use a fixed checklist — the right questions are different every time.
Propose strategy. 4–8 sentences: shape, take choices, cut direction, animation plan, grade direction, subtitle style, length estimate. Wait for confirmation.
Execute. Produce edl.json via the editor sub-agent brief. Drill into timeline_view at ambiguous moments. Build animations in parallel sub-agents. Apply grade per-segment. Compose via render.py.
Preview. render.py --preview.
Self-eval (before showing the user). Run timeline_view on the rendered output (not the sources) at every cut boundary (±1.5s window). Check each image for:
Also sample: first 2s, last 2s, and 2–3 mid-points — check grade consistency, subtitle readability, overall coherence. Run ffprobe on the output to verify duration matches the EDL expectation.
Measure the audio, don't assume it: ffmpeg -i out.mp4 -af ebur128=peak=true -f null - for integrated loudness and true peak, plus RMS per section (dialogue, music-only, end card). An end card 15 dB under the dialogue, or effects louder than speech, is a bug. You cannot listen: say so, and report the numbers.
For anything the user will publish (launch, promo, ad), also spawn one critic sub-agent with the rendered file, the EDL, and any reference videos the user gave. Brief it to roast, not to praise: a verdict, ranked problems with timecodes and evidence (frames, levels), and the 5 fixes to do first. Fresh eyes catch what the author stopped seeing — cut-off payoff lines, 0.5s memes, unreadable 28px text at phone size.
If anything fails: fix → re-render → re-eval. Cap at 3 self-eval passes — if issues remain after 3, flag them to the user rather than looping forever. Only present the preview once the self-eval passes.
Iterate + persist. Natural-language feedback, re-plan, re-render. Never re-transcribe. Final render on confirmation. Append to project.md.
(laughs), (sighs), (applause) mark beats. Extend past them.pack_transcripts.py reads all transcripts/*.json and produces one markdown file where each take is a list of phrase-level lines, each prefixed with its [start-end] time range. Phrases break on any silence ≥ 0.5s OR speaker change. This is the artifact the editor sub-agent reads to pick cuts — it gives word-boundary precision from text alone at 1/10 the tokens of raw JSON.
Example line:
## C0103 (duration: 43.0s, 8 phrases)
[002.52-005.36] S0 Ninety percent of what a web agent does is completely wasted.
[006.08-006.74] S0 We fixed this.When the task is "pick the best take of each beat across many clips," spawn a dedicated sub-agent with a brief shaped like this. The structure is load-bearing; the pitch-shape example is not.
You are editing a <type> video. Pick the best take of each beat and
assemble them chronologically by beat, not by source clip order.
INPUTS:
- takes_packed.md (time-annotated phrase-level transcripts of all takes)
- Product/narrative context: <2 sentences from the user>
- Speaker(s): <name, role, delivery style note>
- Expected structure: <pick an archetype or invent one>
- Verbal slips to avoid: <list from the pre-scan pass>
- Target runtime: <seconds>
Common structural archetypes (pick, adapt, or invent):
- Tech launch / demo: HOOK → PROBLEM → SOLUTION → BENEFIT → EXAMPLE → CTA
- Tutorial: INTRO → SETUP → STEPS → GOTCHAS → RECAP
- Interview: (QUESTION → ANSWER → FOLLOWUP) repeat
- Travel / event: ARRIVAL → HIGHLIGHTS → QUIET MOMENTS → DEPARTURE
- Documentary: THESIS → EVIDENCE → COUNTERPOINT → CONCLUSION
- Music / performance: INTRO → VERSE → CHORUS → BRIDGE → OUTRO
- Or invent your own.
RULES:
- Start/end times must fall on word boundaries from the transcript.
- Pad cut boundaries (working window 30–200ms).
- Prefer silences ≥ 400ms as cut targets.
- Unavoidable slips are kept if no better take exists. Note them in "reason".
- If over budget, revise: drop a beat or trim tails. Report total and self-correct.
OUTPUT (JSON array, no prose):
[{"source": "C0103", "start": 2.42, "end": 6.85, "beat": "HOOK",
"quote": "...", "reason": "..."}, ...]
Return the final EDL and a one-line total runtime check.Your job is to reason about the image, not apply a preset. Look at a frame (via timeline_view), decide what's wrong, adjust one thing, look again.
Mental model is ASC CDL. Per channel: out = (in * slope + offset) ** power, then global saturation. slope → highlights, offset → shadows, power → midtones.
Example filter chains (grade.py has --list-presets; use them as starting points or mix your own):
warm_cinematic — retro/technical, subtle teal/orange split, desaturated. Shipped in a real launch video. Safe for talking heads.neutral_punch — minimal corrective: contrast bump + gentle S-curve. No hue shifts.none — straight copy. Default when the user hasn't asked.For anything else — portraiture, nature, product, music video, documentary — invent your own chain. grade.py --filter '<raw ffmpeg>' accepts any filter string.
Hard rules: apply per-segment during extraction (not post-concat, which re-encodes twice). Never go aggressive without testing skin tones.
Subtitles have three dimensions worth reasoning about: chunking (1/2/3/sentence per line), case (UPPER/Title/Natural), and placement (margin from bottom). The right combo depends on content.
Worked styles — pick, adapt, or invent:
bold-overlay — short-form tech launch, fast-paced social. ~2-word chunks, UPPERCASE, break on punctuation and pauses ≥ 0.3s, grow to 3 words rather than flash a cue < 0.35s (chunk_words in render.py), Helvetica 18 Bold, white-on-outline, MarginV=35. render.py ships with this as SUB_FORCE_STYLE.
FontName=Helvetica,FontSize=18,Bold=1,
PrimaryColour=&H00FFFFFF,OutlineColour=&H00000000,BackColour=&H00000000,
BorderStyle=1,Outline=2,Shadow=0,
Alignment=2,MarginV=35natural-sentence (if you invent this mode) — narrative, documentary, education. 4–7 word chunks, sentence case, break on natural pauses, MarginV=60–80, larger font for readability, slightly wider max-width. No shipped force_style — design one if you need it.
Invent a third style if neither fits. Hard rules: subtitles LAST (Rule 1), output-timeline offsets (Rule 5).
Animations match the content and the brand. Get the palette, font, and visual language from the conversation — never assume a default. If the user hasn't told you, propose a palette in the strategy phase and wait for confirmation before building anything.
Tool options:
Pick the engine per animation slot. Do not default to Remotion just because the animation is web-adjacent.
skills/manim-video/SKILL.md and its references for depth.For HyperFrames slots, scaffold the slot inside edit/animations/slot_<id>/ with npx --yes hyperframes init . --example blank --non-interactive --skip-skills, build the HTML composition there, run the HyperFrames checks that fit the slot (lint, validate, and a draft render when practical), then produce the final overlay video with npx --yes hyperframes render . -o render.mp4 or --format webm -o render.webm when alpha is required. Point the EDL overlay file at the actual rendered path.
For Remotion slots, keep the Remotion project isolated inside the same slot directory, scaffold with npx create-video@latest or install Remotion locally there, render the composition to render.mp4 with the project-local remotion render command, and verify duration and dimensions with ffprobe.
None is mandatory. Invent hybrids if useful (e.g., PIL background with a HyperFrames or Remotion layer on top).
Duration rules of thumb, context-dependent:
narration_length + 1s (universal).Animation payoff timing (rule for sync-to-narration): get the payoff word's timestamp. Start the overlay reveal_duration seconds earlier so the landing frame coincides with the spoken payoff word. Without this sync the animation feels disconnected.
Easing (universal — never linear, it looks robotic):
def ease_out_cubic(t): return 1 - (1 - t) ** 3
def ease_in_out_cubic(t):
if t < 0.5: return 4 * t ** 3
return 1 - (-2 * t + 2) ** 3 / 2ease_out_cubic for single reveals (slow landing). ease_in_out_cubic for continuous draws.
Typing text anchor trick: center on the FULL string's width, not the partial-string width — otherwise text slides left during reveal.
Example palette (the launch video — one aesthetic among infinite):
(10, 10, 10) near-black#FF5A00 / (255, 90, 0) orange(110, 110, 110) dim gray/System/Library/Fonts/Menlo.ttc (index 1)This is one style. If the brand is warm and serif, use that. If it's colorful and playful, use that. If the user handed you a style guide, follow it. If they didn't, propose one and confirm.
Fonts fail silently. A web font that didn't load renders in a fallback face with no error — the video ships in "almost Arial". In HyperFrames/Remotion, await the font load and then assert it: if (!document.fonts.check('700 76px "Inter"')) throw new Error(...). In PIL, pass an explicit font path; never rely on the default.
Worked example — "show the edit" hook (a launch video for this tool). Instead of a title card, the first 3s visualize the editing itself: each transcript word pops in on its Scribe timestamp, with bars under it drawn from the real audio envelope; a filler ("ummm") grows letter by letter while it is spoken, turns orange and is cut out on screen at the same frame the audio cuts, and the next line lands immediately. It works because the picture is driven by the same data as the sound — one composition (Remotion) reads frame-exact word/envelope JSON produced in Python, so nothing can drift. Use the idea whenever the story is "we removed something": make the removal visible.
Parallel sub-agent brief — each animation is one sub-agent spawned via the Agent tool. Each prompt is self-contained (sub-agents have no parent context). Include:
<edit>/animations/slot_<id>/render.mp4)One sub-agent = one file (unique filenames, parallel agents don't overwrite each other).
Sound is where generated videos sound cheap. Worked rules from launch edits:
attack seconds before the visible contact frame.Match the source unless the user asked for something specific. Common targets: 1920×1080@24 cinematic, 1920×1080@30 screen content, 1080×1920@30 vertical social, 3840×2160@24 4K cinema, 1080×1080@30 square. render.py defaults the scale to 1080p from any source; pass --filter or edit the extract command for other targets. Worth asking the user which delivery format matters.
{
"version": 1,
"sources": {"C0103": "/abs/path/C0103.MP4", "C0108": "/abs/path/C0108.MP4"},
"ranges": [
{"source": "C0103", "start": 2.42, "end": 6.85,
"beat": "HOOK", "quote": "...", "reason": "Cleanest delivery, stops before slip at 38.46."},
{"source": "C0108", "start": 14.30, "end": 28.90,
"beat": "SOLUTION", "quote": "...", "reason": "Only take without the false start."}
],
"grade": "warm_cinematic",
"overlays": [
{"file": "edit/animations/slot_1/render.mp4", "start_in_output": 0.0, "duration": 5.0}
],
"subtitles": "edit/master.srt",
"total_duration_s": 87.4
}grade is a preset name or raw ffmpeg filter. overlays are rendered animation clips. subtitles is optional and applied LAST.
project.mdAppend one section per session at <edit>/project.md:
## Session N — YYYY-MM-DD
**Strategy:** one paragraph describing the approach
**Decisions:** take choices, cuts, grades, animations + why
**Reasoning log:** one-line rationale for non-obvious decisions
**Outstanding:** deferred itemsOn startup, read project.md if it exists and summarize the last session in one sentence before asking whether to continue.
Things that consistently fail regardless of style:
© browser-use, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 19 other files in the repository root of browser-use/video-use.
Open the folder on GitHubat commit b877063
We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 2 other GitHub owners. This page covers the copy in browser-use/video-use, which our catalogue first saw on October 7, 2026.
Conversational Video Editing next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Conversational Video Editing this skillbrowser-use/video-use | 28k | 2 repos | ~6.4k | Automated safety check: Warn | MIT | |
| Embedded Video Captionsheygen-com/hyperframes | 59k | 3 repos | ~8.6k | Automated safety check: Pass | Apache-2.0 | |
| Video Understandcalesthio/OpenMontage | 65k | — | ~841 | Automated safety check: Pass | AGPL-3.0 | |
| VideoCaptioner SubtitlesWEIFENG2333/VideoCaptioner | 16k | — | ~1.7k | Automated safety check: Pass | GPL-3.0 | |
| Karaoke CaptionsAI-Builder-Club/skills | 1.3k | 1 repos | ~850 | Automated safety check: Pass | None | |
| Ffmpeg Skillkajisho5/ffmpeg-skill | 1.9k | — | ~7.4k | Automated safety check: Pass | MIT |
heygen-com/hyperframes
Adds captions to a single-subject talking-head video without editing the footage, from plain subtitles to cinematic text placed behind the speaker.
calesthio/OpenMontage
Understand video content locally using ffmpeg frame extraction and Whisper transcription.
WEIFENG2333/VideoCaptioner
Adds subtitles to video with the videocaptioner CLI: transcribes speech, tidies and translates the text, and burns styled subtitles into the video or exports SRT and ASS files.
AI-Builder-Club/skills
Generate TikTok/Shorts-style karaoke captions using MLX Whisper, ASS subtitles, and FFmpeg libass.
kajisho5/ffmpeg-skill
Edit video and audio with local FFmpeg from natural-language requests: cut, trim, join, resize/reframe (9:16, 1:1), speed change, captions and subtitles (SRT/ASS, animated, karaoke), logos and text…
AgriciDaniel/claude-shorts
Interactive longform-to-shortform video creator. An agent skill from AgriciDaniel/claude-shorts.
browser-use/video-use
Produces math and technical explainer videos with Manim Community Edition: concept animations, equation derivations, algorithm walkthroughs and data stories.
Works with
Categories
Edits video through conversation: transcribes speech, proposes cuts, color grades, adds overlay animations and burns in subtitles, confirming the plan before each edit. The agent works as a video editor that reasons from a transcript.md, and picks cut candidates from speech boundaries and silences, looking at visuals only at decision points.
Conversational Video Editing fits situations like: cutting filler and retakes out of a talking-head recording; assembling a travel or tutorial video from raw clips; burning subtitles into a finished edit; adding overlay animations or color grading to an interview.
Run `npx skills add browser-use/video-use --skill video-use -a claude-code`. Or copy the skill folder (the browser-use/video-use repository) into .claude/skills/video-use in your project. Claude Code loads it when a task matches its description.
Run `npx skills add browser-use/video-use --skill video-use -a codex`. Or copy the skill folder (the browser-use/video-use repository) into .agents/skills/video-use in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add browser-use/video-use --skill video-use -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/video-use, .gemini/skills/video-use, .github/skills/video-use and .opencode/skills/video-use in your project.
Going by SKILL.md and its folder, Conversational Video Editing needs Python for the scripts in its folder, the command-line tools its instructions call (npx, uv, pip and ffmpeg) and credentials named ELEVENLABS_API_KEY. Our summary lists: ffmpeg; Python with PIL.
SKILL.md contains no URLs. Its commands use npx, uv and pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md flagged 1 warning(s): tells the agent its actions are pre-authorized / not to stop for confirmation. Read the flagged lines before installing; the check is not a guarantee either way.
Conversational Video Editing is published under the MIT licence (from the LICENSE file in the skill folder). It allows redistribution, so the full SKILL.md is shown on this page.
About 6.4k tokens (SKILL.md is roughly 26k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Conversational Video Editing: Embedded Video Captions (heygen-com/hyperframes, 59k stars), Video Understand (calesthio/OpenMontage, 65k stars), VideoCaptioner Subtitles (WEIFENG2333/VideoCaptioner, 16k stars) and Karaoke Captions (AI-Builder-Club/skills, 1.3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
browser-use (a GitHub organization) maintains it in browser-use/video-use, which has 28,387 GitHub stars. The repository holds 2 skills in this directory. The repository was last updated on October 2, 2026.
Source: browser-use/video-use on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.