Agent skill

Video Toolkit

by calesthio in calesthio/OpenMontage

Create professional videos autonomously using claude-code-video-toolkit — AI voiceovers, image generation, music, talking heads, and Remotion rendering.

AGPL-3.0Auto-check: notesMedia & Creative

Install Video Toolkit

skills CLI
$ npx skills add calesthio/OpenMontage --skill video-toolkit -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install calesthio/OpenMontage video-toolkit --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/calesthio/OpenMontage.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/video-toolkit .claude/skills/video-toolkit && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
video-toolkit
GitHub stars
66k
Used in
1 other repo
Token cost
~3.8k tokens
SKILL.md length
869 words
Files
1
Skills in repo
41
Repo updated
First seen
Licence
AGPL-3.0

At a glance

Create professional videos autonomously using claude-code-video-toolkit — AI voiceovers, image generation, music, talking heads, and Remotion rendering.

  • Works in 11 steps: Check Current State → Install Python Dependencies → Configure Cloud GPU Endpoints → …
  • Tasks that involve Video production
  • SKILL.md covers CRITICAL: Toolkit Path, Setup, Creating a Video and Composition Patterns, plus 2 more sections
  • Calls python3, modal and npm; needs HF_TOKEN and R2_ACCESS_KEY_ID

What it does

Video Toolkit is an agent skill from calesthio/OpenMontage. Create professional videos autonomously using claude-code-video-toolkit — AI voiceovers, image generation, music, talking heads, and Remotion rendering.

Its SKILL.md is about 3.8k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Media & Creative, covering Video production, Text to speech and voice and Image generation. It works with Remotion and Python. The repository describes itself as: World's first open-source, agentic video production system. 12 production pipelines, 100+ tools, 700+ agent skill and production-knowledge files. Turn your AI coding assistant… The licence is AGPL-3.0.

When your agent uses it

  • Tasks that involve Video production
  • Tasks that involve Text to speech and voice
  • Tasks that involve Image generation

Example prompts

  • “/video-toolkit”

Requirements

  • Python 3
  • Node.js
  • Docker
  • A credential in R2_SECRET_ACCESS_KEY

Workflow steps

11 steps, taken from the step headings in SKILL.md.

  1. Check Current State
  2. Install Python Dependencies
  3. Configure Cloud GPU Endpoints
  4. Verify and Quick Test
  5. Create Project
  6. Write Config
  7. Write Voiceover Script
  8. Generate Assets
  9. Sync Timing
  10. Review Still Frames
  11. Render

What it can do on your machine

Read from SKILL.md and the folder at commit 9327439. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python3
    • modal
    • npm
    • pip3
    • npx
    • ffprobe

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • huggingface.co

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • HF_TOKEN
    • R2_ACCESS_KEY_ID
    • R2_SECRET_ACCESS_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Video Toolkit loads about 3.8k tokens when it runs. Until then it costs about 42 tokens; SKILL.md has 869 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~42
When it runs · the whole SKILL.md, loaded when a task matches
~3.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteMentions a .env fileSKILL.md:50
    oolkit needs cloud GPU endpoint URLs in `.env`. Check if `.env` exists and has Modal endpoints:
  • NoteMentions a .env fileSKILL.md:53
    claw/workspace/claude-code-video-toolkit/.env | grep MODAL
  • NoteMentions a .env fileSKILL.md:79
    Add each URL to `.env`:
  • NoteMentions a .env fileSKILL.md:114
    - `.env` has the correct `MODAL_QWEN3_TTS_ENDPOINT_URL`
  • NoteMentions a .env fileSKILL.md:433
    _*_ENDPOINT_URL not configured" | Check `.env` has the endpoint URL. Run `python3 tools/verify_setup.py` |

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from calesthio/OpenMontage at commit 9327439, republished under its AGPL-3.0 licence (© calesthio). 869 words, ~3,782 tokens.

Download SKILL.mdSave it as .claude/skills/video-toolkit/SKILL.md (or your agent's skills folder).
name
video-toolkit
description
Create professional videos autonomously using claude-code-video-toolkit — AI voiceovers, image generation, music, talking heads, and Remotion rendering.

Video Toolkit

Create professional explainer videos from a text brief. The toolkit uses open-source AI models on cloud GPUs (Modal or RunPod) for voiceover, image generation, music, and talking head animation. Remotion (React) handles composition and rendering.

CRITICAL: Toolkit Path

The toolkit lives at a fixed path. ALWAYS cd here before running any tool command.

bash
TOOLKIT=~/.openclaw/workspace/claude-code-video-toolkit
cd $TOOLKIT

NEVER run tool commands from inside a project directory. Tools resolve paths relative to the toolkit root.

Setup

Step 1: Check Current State
bash
cd ~/.openclaw/workspace/claude-code-video-toolkit
python3 tools/verify_setup.py

If everything shows [x], skip to "Quick Test" below. Otherwise continue setup.

Step 2: Install Python Dependencies
bash
cd ~/.openclaw/workspace/claude-code-video-toolkit
pip3 install --break-system-packages -r tools/requirements.txt

Note: --break-system-packages is needed on Debian/Ubuntu with managed Python (PEP 668). Safe inside containers.

Step 3: Configure Cloud GPU Endpoints

The toolkit needs cloud GPU endpoint URLs in .env. Check if .env exists and has Modal endpoints:

bash
cat ~/.openclaw/workspace/claude-code-video-toolkit/.env | grep MODAL

If Modal endpoints are configured, you're ready. If not, ask the user to provide Modal endpoint URLs or set up Modal:

bash
pip3 install --break-system-packages modal
python3 -m modal setup   # Opens browser for authentication

# Deploy each tool — capture the endpoint URL from output
cd ~/.openclaw/workspace/claude-code-video-toolkit
modal deploy docker/modal-qwen3-tts/app.py
modal deploy docker/modal-flux2/app.py
modal deploy docker/modal-music-gen/app.py
modal deploy docker/modal-sadtalker/app.py
modal deploy docker/modal-image-edit/app.py
modal deploy docker/modal-upscale/app.py
modal deploy docker/modal-propainter/app.py
modal deploy docker/modal-ltx2/app.py      # Requires: modal secret create huggingface-token HF_TOKEN=hf_...

LTX-2 prerequisite: Before deploying LTX-2, create a HuggingFace secret and accept the Gemma 3 license:

bash
modal secret create huggingface-token HF_TOKEN=hf_your_read_access_token

Add each URL to .env:

MODAL_QWEN3_TTS_ENDPOINT_URL=https://...modal.run
MODAL_FLUX2_ENDPOINT_URL=https://...modal.run
MODAL_MUSIC_GEN_ENDPOINT_URL=https://...modal.run
MODAL_SADTALKER_ENDPOINT_URL=https://...modal.run
MODAL_IMAGE_EDIT_ENDPOINT_URL=https://...modal.run
MODAL_UPSCALE_ENDPOINT_URL=https://...modal.run
MODAL_DEWATERMARK_ENDPOINT_URL=https://...modal.run
MODAL_LTX2_ENDPOINT_URL=https://...modal.run

Optional but recommended — Cloudflare R2 for reliable file transfer:

R2_ACCOUNT_ID=...
R2_ACCESS_KEY_ID=...
R2_SECRET_ACCESS_KEY=...
R2_BUCKET_NAME=video-toolkit
Step 4: Verify and Quick Test
bash
cd ~/.openclaw/workspace/claude-code-video-toolkit
python3 tools/verify_setup.py

All tools should show [x]. Then run a quick test to confirm the GPU pipeline works:

bash
cd ~/.openclaw/workspace/claude-code-video-toolkit
python3 tools/qwen3_tts.py --text "Hello, this is a test." --speaker Ryan --tone warm --output /tmp/video-toolkit-test.mp3 --cloud modal

If you get a valid .mp3 file, setup is complete. If it fails, check:

  • .env has the correct MODAL_QWEN3_TTS_ENDPOINT_URL
  • Run python3 tools/verify_setup.py --json and check modal_tools for which endpoints are missing

Cost: Modal includes $30/month free compute. A typical 60s video costs $1-3.


Creating a Video

Step 1: Create Project
bash
cd ~/.openclaw/workspace/claude-code-video-toolkit
cp -r templates/product-demo projects/PROJECT_NAME
cd projects/PROJECT_NAME
npm install

Templates: product-demo (marketing/explainer), sprint-review, sprint-review-v2 (composable scenes).

Step 2: Write Config

Edit projects/PROJECT_NAME/src/config/demo-config.ts:

typescript
export const demoConfig: ProductDemoConfig = {
  product: {
    name: 'My Product',
    tagline: 'What it does in one line',
    website: 'example.com',
  },
  scenes: [
    { type: 'title', durationSeconds: 9, content: { headline: '...', subheadline: '...' } },
    { type: 'problem', durationSeconds: 14, content: { headline: '...', problems: ['...', '...'] } },
    { type: 'solution', durationSeconds: 13, content: { headline: '...', highlights: ['...', '...'] } },
    { type: 'stats', durationSeconds: 12, content: { stats: [{value: '99%', label: '...'}, ...] } },
    { type: 'cta', durationSeconds: 10, content: { headline: '...', links: ['...'] } },
  ],
  audio: {
    backgroundMusicFile: 'audio/bg-music.mp3',
    backgroundMusicVolume: 0.12,
  },
};

Scene types: title, problem, solution, demo, feature, stats, cta.

Duration rule: Estimate durationSeconds as ceil(word_count / 2.5) + 2. You will adjust this after generating audio in Step 4.

Step 3: Write Voiceover Script

Create projects/PROJECT_NAME/VOICEOVER-SCRIPT.md:

markdown
## Scene 1: Title (9s, ~17 words)
Build videos with AI. The product name toolkit makes it easy.

## Scene 2: Problem (14s, ~30 words)
The problem statement goes here. Keep it punchy and relatable.

Word budget per scene: (durationSeconds - 2) * 2.5 words. The -2 accounts for 1s audio delay + 1s padding.

Step 4: Generate Assets

CRITICAL: All commands below MUST be run from the toolkit root, not the project directory.

bash
cd ~/.openclaw/workspace/claude-code-video-toolkit
4a. Background Music
bash
cd ~/.openclaw/workspace/claude-code-video-toolkit
python3 tools/music_gen.py \
  --preset corporate-bg \
  --duration 90 \
  --output projects/PROJECT_NAME/public/audio/bg-music.mp3 \
  --cloud modal

Presets: corporate-bg, upbeat-tech, ambient, dramatic, tension, hopeful, cta, lofi.

4b. Voiceover (per-scene)

Generate ONE .mp3 file PER SCENE. Do NOT generate a single voiceover file.

bash
cd ~/.openclaw/workspace/claude-code-video-toolkit

# Scene 01
python3 tools/qwen3_tts.py \
  --text "The voiceover text for scene one." \
  --speaker Ryan --tone warm \
  --output projects/PROJECT_NAME/public/audio/scenes/01.mp3 \
  --cloud modal

# Scene 02
python3 tools/qwen3_tts.py \
  --text "The voiceover text for scene two." \
  --speaker Ryan --tone warm \
  --output projects/PROJECT_NAME/public/audio/scenes/02.mp3 \
  --cloud modal

# ... repeat for each scene

Speakers: Ryan, Aiden, Vivian, Serena, Uncle_Fu, Dylan, Eric, Ono_Anna, Sohee Tones: neutral, warm, professional, excited, calm, serious, storyteller, tutorial

For voice cloning (needs a reference recording):

bash
cd ~/.openclaw/workspace/claude-code-video-toolkit
python3 tools/qwen3_tts.py \
  --text "Text to speak" \
  --ref-audio assets/voices/reference.m4a \
  --ref-text "Exact transcript of the reference audio" \
  --output projects/PROJECT_NAME/public/audio/scenes/01.mp3 \
  --cloud modal
4c. Scene Images
bash
cd ~/.openclaw/workspace/claude-code-video-toolkit
python3 tools/flux2.py \
  --prompt "Dark tech background with blue geometric grid, cinematic lighting" \
  --width 1920 --height 1080 \
  --output projects/PROJECT_NAME/public/images/title-bg.png \
  --cloud modal

Image presets (use --preset instead of --prompt --width --height): title-bg, problem, solution, demo-bg, stats-bg, cta, thumbnail, portrait-bg

bash
cd ~/.openclaw/workspace/claude-code-video-toolkit
python3 tools/flux2.py \
  --preset title-bg \
  --output projects/PROJECT_NAME/public/images/title-bg.png \
  --cloud modal
4d. Video Clips — B-Roll & Animated Backgrounds (optional)

Generate AI video clips for b-roll cutaways, animated slide backgrounds, or intro/outro sequences:

bash
cd ~/.openclaw/workspace/claude-code-video-toolkit

# B-roll clip from text
python3 tools/ltx2.py \
  --prompt "Aerial drone shot over a European city at golden hour, cinematic wide angle" \
  --output projects/PROJECT_NAME/public/videos/broll-europe.mp4 \
  --cloud modal

# Animate a slide/screenshot (image-to-video)
python3 tools/ltx2.py \
  --prompt "Gentle particle effects, soft ambient light shifts, very slight camera drift" \
  --input projects/PROJECT_NAME/public/images/title-bg.png \
  --output projects/PROJECT_NAME/public/videos/animated-title.mp4 \
  --cloud modal

# Abstract intro/outro background
python3 tools/ltx2.py \
  --prompt "Dark moody abstract background with flowing blue light streaks, bokeh particles, cinematic" \
  --output projects/PROJECT_NAME/public/videos/intro-bg.mp4 \
  --cloud modal

Use in Remotion compositions with <OffthreadVideo>:

tsx
<OffthreadVideo src={staticFile('videos/broll-europe.mp4')} />

LTX-2 rules:

  • Max ~8 seconds per clip (193 frames at 24fps). Default is ~5s (121 frames).
  • Width/height must be divisible by 64. Default: 768x512.
  • ~$0.20-0.25 per clip, ~2.5 min generation time.
  • Cold start ~60-90s. Subsequent clips on warm GPU are faster.
  • Generated audio is ambient only — use voiceover/music tools for speech and music.
  • ~30% of generations may have training data artifacts (logos/text). Re-run with --seed to vary.
Show full SKILL.md (365 more words)Show less
4e. Talking Head Narrator (optional)

Generate a presenter portrait, then animate per-scene clips:

bash
cd ~/.openclaw/workspace/claude-code-video-toolkit

# 1. Generate portrait
python3 tools/flux2.py \
  --prompt "Professional presenter portrait, clean style, dark background, facing camera, upper body" \
  --width 1024 --height 576 \
  --output projects/PROJECT_NAME/public/images/presenter.png \
  --cloud modal

# 2. Generate per-scene narrator clips (one per scene, NOT one long video)
python3 tools/sadtalker.py \
  --image projects/PROJECT_NAME/public/images/presenter.png \
  --audio projects/PROJECT_NAME/public/audio/scenes/01.mp3 \
  --preprocess full --still --expression-scale 0.8 \
  --output projects/PROJECT_NAME/public/narrator-01.mp4 \
  --cloud modal

# Repeat for each scene that needs a narrator

SadTalker rules — follow these exactly:

  • ALWAYS use --preprocess full (default crop outputs a square, wrong aspect ratio)
  • ALWAYS use --still (reduces head movement, looks professional)
  • ALWAYS generate per-scene clips (6-15s each), NEVER one long video
  • Processing: ~3-4 min per 10s of audio on Modal A10G
  • --expression-scale 0.8 keeps expressions subtle (range 0.0-1.5)
4e. Image Editing (optional)

Create scene variants from existing images:

bash
cd ~/.openclaw/workspace/claude-code-video-toolkit
python3 tools/image_edit.py \
  --input projects/PROJECT_NAME/public/images/title-bg.png \
  --prompt "Make it darker with red tones, more ominous" \
  --output projects/PROJECT_NAME/public/images/problem-bg.png \
  --cloud modal
4f. Upscaling (optional)
bash
cd ~/.openclaw/workspace/claude-code-video-toolkit
python3 tools/upscale.py \
  --input projects/PROJECT_NAME/public/images/some-image.png \
  --output projects/PROJECT_NAME/public/images/some-image-4x.png \
  --scale 4 --cloud modal
Step 5: Sync Timing

ALWAYS do this after generating voiceover. Audio duration differs from estimates.

bash
cd ~/.openclaw/workspace/claude-code-video-toolkit
for f in projects/PROJECT_NAME/public/audio/scenes/*.mp3; do
  echo "$(basename $f): $(ffprobe -v error -show_entries format=duration -of csv=p=0 "$f")s"
done

Update each scene's durationSeconds in demo-config.ts to: ceil(actual_audio_duration + 2).

Example: if 01.mp3 is 6.8s, set scene 1 durationSeconds to 9 (ceil(6.8 + 2) = 9).

Step 6: Review Still Frames
bash
cd ~/.openclaw/workspace/claude-code-video-toolkit/projects/PROJECT_NAME
npx remotion still src/index.ts ProductDemo --frame=100 --output=/tmp/review-scene1.png
npx remotion still src/index.ts ProductDemo --frame=400 --output=/tmp/review-scene2.png

Check: text truncation, animation timing, narrator PiP positioning, background contrast.

Step 7: Render
bash
cd ~/.openclaw/workspace/claude-code-video-toolkit/projects/PROJECT_NAME
npm run render

Output: out/ProductDemo.mp4


Composition Patterns

Per-Scene Audio

Use per-scene audio with a 1-second delay (from={30} = 30 frames = 1s at 30fps):

tsx
<Sequence from={30}>
  <Audio src={staticFile('audio/scenes/01.mp3')} volume={1} />
</Sequence>
Per-Scene Narrator PiP
tsx
<Sequence from={30}>
  <OffthreadVideo
    src={staticFile('narrator-01.mp4')}
    style={{ width: 320, height: 180, objectFit: 'cover' }}
    muted
  />
</Sequence>

ALWAYS use <OffthreadVideo>, NEVER <video>. Remotion requires its own component for frame-accurate rendering.

Transitions
tsx
import { TransitionSeries, linearTiming } from '@remotion/transitions';
import { fade } from '@remotion/transitions/fade';
import { glitch } from '../../../lib/transitions/presentations/glitch';
import { lightLeak } from '../../../lib/transitions/presentations/light-leak';

NEVER import from lib/transitions barrel — import custom transitions from lib/transitions/presentations/ directly.


Error Recovery

ProblemSolution
Tool command fails with "No module named..."Run pip3 install --break-system-packages -r tools/requirements.txt from toolkit root
"MODAL_*_ENDPOINT_URL not configured"Check .env has the endpoint URL. Run python3 tools/verify_setup.py
SadTalker output is square/croppedYou forgot --preprocess full. Re-run with that flag
Audio too short/long for sceneRe-run Step 5 (sync timing) and update config
npm run render failsMake sure you're in the project dir, not toolkit root. Run npm install first
"Cannot find module" in RemotionCheck import paths. Custom components use ../../../lib/ relative paths
Cold start timeout on ModalFirst call after idle takes 30-120s. Retry once — second call uses warm GPU

Cost Estimates (Modal)

ToolTypical CostNotes
Qwen3-TTS~$0.01/scene~20s per scene on warm GPU
FLUX.2~$0.01/image~3s warm, ~30s cold
ACE-Step~$0.02-0.05Depends on duration
SadTalker~$0.05-0.20/scene~3-4 min per 10s audio
Qwen-Edit~$0.03-0.15~8 min cold start (25GB model)
RealESRGAN~$0.005/imageVery fast
LTX-2.3~$0.20-0.25/clip~2.5 min per 5s clip, A100-80GB

Total for a 60s video: ~$1-3 depending on scenes and narrator clips.

Modal Starter plan: $30/month free compute. Apps scale to zero when idle.

© calesthio, AGPL-3.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/video-toolkit of calesthio/OpenMontage.

Open the folder on GitHubat commit 9327439

Used in 1 other repository

We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in calesthio/OpenMontage, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Video Toolkit next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Video Toolkit compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Video Toolkit this skillcalesthio/OpenMontage66k1 repos~3.8kAutomated safety check: NotesAGPL-3.0
Whiteboard Videognipbao/codex-whiteboard-video-skill327—~7.2kAutomated safety check: NotesMIT
Remotionzhuzhaoyun/Molio433—~4kAutomated safety check: PassCustom licence
AI Mediaericrisco/rsc-harness180—~3.3kAutomated safety check: PassMIT
Video Podcast Makerdtsola/xiaoyaosearch1k—~3.4kAutomated safety check: PassMIT
Making Demo Videosnukeop/nuclear19k—~892Automated safety check: PassAGPL-3.0

Similar skills

  • Whiteboard Video

    gnipbao/codex-whiteboard-video-skill

    Generate smooth hand-drawn whiteboard and story videos directly inside Codex from text scripts, GPT Image 2 color storyboards, scene plans, SVGs, line art, or local images.

    327 GitHub stars~7.2k tokensUpdated 1 mo ago
    Media & CreativeAuto-check: notes
  • Remotion

    zhuzhaoyun/Molio

    Molio's builtin skill for MAKING a video from any source — wiki notes, articles, scripts, product info, or a brief — and rendering it to MP4.

    433 GitHub stars~4k tokensUpdated today
    Media & CreativeAuto-check passed
  • AI Media

    ericrisco/rsc-harness

    A skill your agent uses when a creative goal must become a finished media file: pick and order generative-media models per modality — AI voiceover, image-to-video clips, score — then glue them with…

    180 GitHub stars~3.3k tokensUpdated today
    Media & CreativeAuto-check passed
  • Video Podcast Maker

    dtsola/xiaoyaosearch

    Turns a topic into a 4K horizontal video podcast through research, scripting, text-to-speech, Remotion rendering and background music, and can learn styles from references.

    1k GitHub stars~3.4k tokensUpdated 8 days ago
    Media & CreativeAuto-check passed
  • Making Demo Videos

    nukeop/nuclear

    A skill your agent uses when making a demo, tutorial, or feature video of Nuclear.

    19k GitHub stars~892 tokensUpdated 3 days ago
    Media & CreativeAuto-check passed
  • Vox Director

    Alisa0808/vox-director

    Turn ONE topic into a finished Vox-style paper-collage explainer / ad video, end to end on the Atlas Cloud API + local ffmpeg — script, collage keyframes, motion, voice-over, music, captions, all…

    2.2k GitHub stars~5.6k tokensUpdated 4 days ago
    Media & CreativeAuto-check passed

More from calesthio/OpenMontage

All 41 skills in this repo
  • Video Understand

    calesthio/OpenMontage

    Understand video content locally using ffmpeg frame extraction and Whisper transcription.

    66k GitHub stars~841 tokensUpdated 6 days ago
    Auto-check passed
  • Avatar Video

    calesthio/OpenMontage

    Create AI avatar videos with precise control over avatars, voices, scripts, scenes, and backgrounds using HeyGen's v2 API.

    66k GitHub stars~1.6k tokensUpdated 6 days ago
    Auto-check passed
  • D3 Viz

    calesthio/OpenMontage

    Creating interactive data visualisations using d3.js. An agent skill from calesthio/OpenMontage.

    66k GitHub starsUsed in 3 repos~5.4k tokens
    Auto-check passed
  • Create Video

    calesthio/OpenMontage

    Create videos from a text prompt using HeyGen's Video Agent.

    66k GitHub stars~1.3k tokensUpdated 6 days ago
    Auto-check passed
  • Threejs World Generation

    calesthio/OpenMontage

    Build deterministic, editable, free-viewpoint Three.js worlds from text or structured briefs.

    66k GitHub stars~2k tokensUpdated 6 days ago
    Auto-check passed
  • Video Edit

    calesthio/OpenMontage

    Edit videos locally using ffmpeg. An agent skill from calesthio/OpenMontage.

    66k GitHub stars~855 tokensUpdated 6 days ago
    Auto-check: notes

Works with

Questions about Video Toolkit

What does Video Toolkit do?

Create professional videos autonomously using claude-code-video-toolkit — AI voiceovers, image generation, music, talking heads, and Remotion rendering. Video Toolkit is an agent skill from calesthio/OpenMontage. Create professional videos autonomously using claude-code-video-toolkit — AI voiceovers, image generation, music, talking heads, and Remotion rendering.

When should I use Video Toolkit?

Video Toolkit fits situations like: tasks that involve Video production; tasks that involve Text to speech and voice; tasks that involve Image generation.

How do I install Video Toolkit in Claude Code?

Run `npx skills add calesthio/OpenMontage --skill video-toolkit -a claude-code`. Or copy the skill folder (.agents/skills/video-toolkit in calesthio/OpenMontage) into .claude/skills/video-toolkit in your project. Claude Code loads it when a task matches its description.

How do I install Video Toolkit in Codex?

Run `npx skills add calesthio/OpenMontage --skill video-toolkit -a codex`. Or copy the skill folder (.agents/skills/video-toolkit in calesthio/OpenMontage) into .agents/skills/video-toolkit in your project. Codex loads it when a task matches its description.

Can I use Video Toolkit in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add calesthio/OpenMontage --skill video-toolkit -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/video-toolkit, .gemini/skills/video-toolkit, .github/skills/video-toolkit and .opencode/skills/video-toolkit in your project.

What does Video Toolkit need to run?

Going by SKILL.md and its folder, Video Toolkit needs the command-line tools its instructions call (python3, modal, npm, pip3, npx and ffprobe) and credentials named HF_TOKEN, R2_ACCESS_KEY_ID and R2_SECRET_ACCESS_KEY. Our summary lists: Python 3; Node.js; Docker; A credential in R2_SECRET_ACCESS_KEY.

Does Video Toolkit access the network?

SKILL.md names 1 domain. As links in the text: huggingface.co. This is read from the text; nothing was executed.

Is Video Toolkit safe to install?

Our automated static check of SKILL.md found notes only (mentions a .env file), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Video Toolkit use?

Video Toolkit is published under the AGPL-3.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Video Toolkit use?

About 3.8k tokens (SKILL.md is roughly 15k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Video Toolkit?

Skills that share tags, products or a category with Video Toolkit: Whiteboard Video (gnipbao/codex-whiteboard-video-skill, 327 stars), Remotion (zhuzhaoyun/Molio, 433 stars), AI Media (ericrisco/rsc-harness, 180 stars) and Video Podcast Maker (dtsola/xiaoyaosearch, 1k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Video Toolkit?

calesthio (a GitHub user) maintains it in calesthio/OpenMontage, which has 65,930 GitHub stars. The repository holds 41 skills in this directory. The repository was last updated on October 3, 2026.

Source: calesthio/OpenMontage on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.