Agent skill

Creating Video

by oaustegard in oaustegard/claude-skills

Create video from prompts by overseeing multi-clip AI generation end to end: write a shot list, generate each scene with Gemini Omni Flash (via the Cloudflare AI Gateway), review the results, and…

MITAuto-check passedMedia & Creative

Install Creating Video

skills CLI
$ npx skills add oaustegard/claude-skills --skill creating-video -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install oaustegard/claude-skills creating-video --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/oaustegard/claude-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/creating-video .claude/skills/creating-video && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
creating-video
GitHub stars
150
Token cost
~2k tokens
SKILL.md length
877 words
Files
5 (incl. scripts)
Skills in repo
66
Repo updated
First seen
Licence
MIT

At a glance

Create video from prompts by overseeing multi-clip AI generation end to end: write a shot list, generate each scene with Gemini Omni Flash (via the Cloudflare AI Gateway), review the results, and…

  • Works in 4 steps: Generation (Omni via Cloudflare AI… → Review (use the parsing-video skill) → Assembly (scripts/assemble.py) → …
  • The user asks to make/generate a video
  • SKILL.md covers The five stages, Stage 2 — Generation (Omni via…, Stage 3 — Review (use the… and Stage 4 — Assembly…, plus 3 more sections
  • Runs Python scripts from its folder; calls python3; needs CF_API_TOKEN

What it does

Creating Video is an agent skill from oaustegard/claude-skills. Create video from prompts by overseeing multi-clip AI generation end to end: write a shot list, generate each scene with Gemini Omni Flash (via the Cloudflare AI Gateway), review the results, and assemble them into a finished cut. Use when the user asks to make/generate a video, a short film, an animatic, or a multi-scene clip from a script or idea; when they mention Omni, Veo, text-to-video, or image-to-video; or when acting as the editing/director agent over generated footage. Triggers on 'make a video'…

Its SKILL.md is about 2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files, including scripts (for example `CHANGELOG.md`, `scripts/assemble.py` and `scripts/omni_generate.py`).

It sits in Media & Creative, covering AI video generation. It works with Workers AI. The repository describes itself as: My collection of Claude skills. The licence is MIT.

When your agent uses it

  • The user asks to make/generate a video
  • A multi-scene clip from a script
  • They mention Omni
  • Acting as the editing/director agent over generated footage

Example prompts

  • “make a video”
  • “generate a clip”
  • “short film”
  • “/creating-video”

Requirements

  • Python 3
  • A credential in CF_API_TOKEN

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Generation (Omni via Cloudflare AI Gateway)
  2. Review (use the parsing-video skill)
  3. Assembly (scripts/assemble.py)
  4. Iterate by editing, not regenerating

What it can do on your machine

Read from SKILL.md and the folder at commit cf49d47. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 3 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • CF_API_TOKEN

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Creating Video loads about 2k tokens when it runs. Until then it costs about 201 tokens; SKILL.md has 877 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~201
When it runs · the whole SKILL.md, loaded when a task matches
~2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from oaustegard/claude-skills at commit cf49d47, republished under its MIT licence (© oaustegard). 877 words, ~1,963 tokens.

Download SKILL.mdSave it as .claude/skills/creating-video/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
creating-video
description
Create video from prompts by overseeing multi-clip AI generation end to end: write a shot list, generate each scene with Gemini Omni Flash (via the Cloudflare AI Gateway), review the results, and assemble them into a finished cut. Use when the user asks to make/generate a video, a short film, an animatic, or a multi-scene clip from a script or idea; when they mention Omni, Veo, text-to-video, or image-to-video; or when acting as the editing/director agent over generated footage. Triggers on 'make a video', 'generate a clip', 'short film', 'video from this script', 'turn this into a video', 'omni', 'veo', 'text to video', 'storyboard to video'. For transcoding/trimming/merging/GIF/subtitles use processing-video; for reading or summarizing existing video content use parsing-video.
metadata.version
0.2.1

Creating Video

Claude cannot render video itself, but it can direct a generator. This skill drives a multi-clip pipeline — script → per-scene generation → review → assembly — with Claude as the editing agent overseeing continuity and cut.

Requires ffmpeg/ffprobe and Cloudflare AI Gateway creds (/mnt/project/proxy.env).

Default model: Gemini Omni Flash (gemini-omni-flash-preview, Interactions API) — Google's own guidance as of 2026-07. It beats Veo on character consistency, reference-image control, text rendering, and conversational editing, and generates faster (~30–90 s/clip vs 1–3 min). Fall back to Veo 3.1 (scripts/veo_generate.py, kept intact) only for scene extension, first+last-frame interpolation, or legacy pipelines — Omni supports neither.

The five stages

  1. Shot list — break the story into beats, up to 10 s each, one prompt per shot.
  2. Generate — scripts/omni_generate.py, run detached.
  3. Review — parsing-video on each clip and on the assembled cut.
  4. Assemble — scripts/assemble.py, run detached.
  5. Iterate — re-review, then edit failing scenes conversationally (cheaper and more stable than regenerating from scratch).

Stage 2 — Generation (Omni via Cloudflare AI Gateway)

Auth is CF AI Gateway BYOK: proxy.env supplies CF_ACCOUNT_ID, CF_GATEWAY_ID, CF_API_TOKEN; the Google key lives inside the gateway. Both the interactions call and the files/{id}:download retrieval route through it (verified 2026-07-20).

Run detached. The Interactions API is synchronous — the HTTP call holds open for the whole generation (~30–90 s/clip). bash_tool caps at ~50 s. Launch and adaptive-wait on the DONE sentinel:

bash
# prompts.json = {"1": "...", "2": {"text": "...", "images": ["ref0.png"]}, ...}
set -a; . /mnt/project/proxy.env; set +a
(setsid python3 scripts/omni_generate.py prompts.json --out omni/ \
    --negative "on-screen watermark" &)
# then, in a separate call:
timeout 45 sh -c 'while [ ! -f omni/DONE ]; do sleep 3; done'; cat omni/generate.log

Gotchas (all verified 2026-07-20 — the script already handles them):

  • No dry run exists. Even input:"" returns 200 and bills a full generation (~58k video output tokens). Every request costs money; never "validate" with a throwaway call.
  • No negativePrompt parameter (Veo had one; Omni 400s on unsupported params). --negative folds into the prompt as "Do not include: X."
  • Videos >4MB need delivery:"uri" → poll files/{id} to ACTIVE → download via the gateway. The script always uses uri delivery, inline base64 as fallback.
  • Output is up to 10 s 1280×720 mp4 with synchronized audio + SynthID watermark. No video extension, no first↔last-frame interpolation, no multi-video referencing — those are Veo territory.
  • The egress proxy 503s on cold start → retry with backoff.

Omni defaults to multi-shot. Left alone it invents its own cuts inside a clip, which fights a shot-list pipeline. Every per-scene prompt should say "single continuous shot" / "no scene cuts" unless the beat wants internal cuts.

Stage 3 — Review (use the parsing-video skill)

Contact-sheet each clip and the assembled cut. Per-clip sheets miss cross-clip continuity; the full-cut sheet is where character drift, prop jumps, and logic breaks show up in one read (Oskar, 2026-07-18: review the whole assembly, not just scenes). Scan every sheet against the continuity checklist:

  • Character — same face/hair/wardrobe across every shot they appear in.
  • Prop — same identity, size, color, and attachment point shot to shot.
  • Physical logic — is the world coherent (a window must be open before a bird lands on the sill)?
  • Action completeness — is the key action actually shown, not cut around?

Stage 4 — Assembly (scripts/assemble.py)

Trims each clip to its beat, crossfades, drops audio by default, optionally burns an overlay word, and holds the final frame so the ending lands.

bash
(setsid python3 scripts/assemble.py omni/scene_1.mp4 ... omni/scene_6.mp4 \
    --out film.mp4 --tail-hold 1.3 &)

Run detached too — six trims plus a 30 s stitch exceed the bash ceiling on the single-core container. Diagnosed: abrupt ending → --tail-hold; jarring cuts between independently generated ambiences → audio dropped by default (--keep-audio to acrossfade instead).

Show full SKILL.md (337 more words)Show less

Stage 5 — Iterate by editing, not regenerating

results.json stores each scene's interaction_id. A failed detail (wrong color, unwanted object, lighting) is a one-line conversational edit that preserves everything else:

bash
python3 scripts/omni_generate.py --edit <interaction_id> \
    "Make the scarf red. Keep everything else the same." --out omni/scene_3_v2.mp4

Simple edit prompts work best; append "Keep everything else the same." Reserve full regeneration for scenes whose composition or action is wrong. (Editing requires store=true, the script's default.)

Prompt craft — continuity is the hard part

The failure modes below came from real crits on 2026-07-18 (Veo era). Omni narrows several of them but the disciplines still pay on the first pass.

  • Pin characters and props with reference images — the first-class fix for the drift that text locks never fully solved under Veo. Generate one canonical still per character/prop, put it in the scene's images list, and bind it in the prompt: the woman <IMAGE_REF_0> is holding <IMAGE_REF_1> (refs start at 0; <FIRST_FRAME> pins a starting frame instead).
  • Still lock a character sheet in text. One verbatim appearance/wardrobe block, reused in every shot with that character — references constrain, text directs.
  • Lock the setting. One room/location description across co-located scenes.
  • Name the actor in every shot ("the same woman's hands") — never leave it to the model to infer who is on screen.
  • Spell out scene logic — physical coherence, the correct order of beats, and show the payoff action rather than cutting away from it.
  • Timing is promptable. [0-3s] ... [3-6s] ... timecode syntax or natural language ("after 3 seconds, ...") controls beats inside a clip.
  • On-screen text now renders reliably — spell out exactly what any visible text should say (signs, labels, title cards). Burning text in assembly (assemble.py --overlay-last) remains an option for cross-clip title cards.
  • Audio is promptable. Describe the track ("no dialogue", "calm background music") or Omni invents one per clip — which still won't match across clips; assembly drops audio by default for that reason.

Scope

  • Transcode / trim / merge / GIF / subtitles → processing-video.
  • Reading, summarizing, or QA of existing video content → parsing-video (also the review step here).
  • Scene extension / first+last-frame control → the retained Veo path (scripts/veo_generate.py --list to enumerate models; strings drift).

© oaustegard, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files (scripts) in creating-video of oaustegard/claude-skills.

  • SKILL.md
  • CHANGELOG.md
  • scripts/assemble.py
  • scripts/omni_generate.py
  • scripts/veo_generate.py

Open the folder on GitHubat commit cf49d47

Compare with similar skills

Creating Video next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Creating Video compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Creating Video this skilloaustegard/claude-skills150—~2kAutomated safety check: PassMIT
Video Generationbytedance/deer-flow84k3 repos~1.4kAutomated safety check: PassMIT
Video Cover Imageitwanger/toBeBetterJavaer18k—~3.3kAutomated safety check: PassNone
Seedancesongguoxs/seedance-prompt-skill2.9k1 repos~2.5kAutomated safety check: PassNone
HyperFrames Video Entry Pointheygen-com/hyperframes59k3 repos~5.2kAutomated safety check: PassApache-2.0
Lanshu Create AI Presenter Videocclank/lanshu-create-ai-presenter-video2.6k—~3.6kAutomated safety check: PassMIT

Similar skills

  • Video Generation

    bytedance/deer-flow

    Generates short videos from a structured JSON prompt, optionally guided by a reference image used as the first or last frame.

    84k GitHub starsUsed in 3 repos~1.4k tokens
    Media & CreativeAuto-check passed
  • Video Cover Image

    itwanger/toBeBetterJavaer

    Generate matched 3:4, 16:9, and 4:3 short-video cover images from toBeBetterJavaer video scripts or AI/Java technical topics.

    18k GitHub stars~3.3k tokensUpdated yesterday
    Media & CreativeAuto-check passed
  • Seedance

    songguoxs/seedance-prompt-skill

    This skill should be used when the user asks to "generate video prompts", "create Seedance prompts", "write video descriptions", mentions "Seedance", "seedance", "即梦", "即梦平台", "视频提示词", "视频生成"…

    2.9k GitHub starsUsed in 1 repo~2.5k tokens
    Media & CreativeAuto-check passed
  • HyperFrames Video Entry Point

    heygen-com/hyperframes

    Entry point for making, editing and rendering videos from HTML compositions with HyperFrames, routing each request to the right workflow.

    59k GitHub starsUsed in 3 repos~5.2k tokens
    Media & CreativeAuto-check passed
  • Lanshu Create AI Presenter Video

    cclank/lanshu-create-ai-presenter-video

    Turn a topic or finished script into a complete, publish-ready explainer video — led by an AI presenter from an authorized adult presenter image, or performed in one of nine visual explainer styles…

    2.6k GitHub stars~3.6k tokensUpdated 3 days ago
    Media & CreativeAuto-check passed
  • Video Shots

    eternityspring/reelbench-skills

    拉片:把一条成片拆成逐镜头的分析表——每个镜头的时长、景别、类别、运镜、画面. An agent skill from eternityspring/reelbench-skills.

    872 GitHub starsUsed in 1 repo~1.8k tokens
    Media & CreativeAuto-check: notes

More from oaustegard/claude-skills

All 66 skills in this repo
  • Vega-Lite Interactive Charts

    oaustegard/claude-skills

    Builds interactive Vega-Lite charts from uploaded data: analyzes the fields, picks five to ten fitting chart types, and produces a React artifact with the data embedded inline.

    150 GitHub stars~2.1k tokensUpdated yesterday
    Auto-check passed
  • Single-File HTML Composer

    oaustegard/claude-skills

    Builds self-contained single-file HTML pages such as reports, decks, postmortems, flowcharts and prototypes from a small spec using a bundled Python composer and templates.

    150 GitHub stars~3.2k tokensUpdated yesterday
    Auto-check passed
  • Declauding

    oaustegard/claude-skills

    Rewrites model-sounding prose into plain technical writing and checks that every claim survives, for PR text, docs, commit messages and similar drafts.

    150 GitHub stars~5.2k tokensUpdated yesterday
    Auto-check passed
  • Preact Developer

    oaustegard/claude-skills

    Guides building standards-based Preact apps with native-first choices, HTM syntax, import maps and vendored ESM, from single-file demos to larger builds.

    150 GitHub stars~4.6k tokensUpdated yesterday
    Auto-check passed
  • Bluesky Zeitgeist Sampler

    oaustegard/claude-skills

    Deprecated sampler that captures short windows of the Bluesky firehose, clusters trending terms and builds an HTML report; replaced by the browsing-bluesky skill.

    150 GitHub stars~1.4k tokensUpdated yesterday
    Auto-check passed
  • Adversarial Review Before Shipping

    oaustegard/claude-skills

    Has a fresh-context adversary attack a blog post, recommendation, analysis brief or piece of code before you ship it, using a profile suited to that kind of artifact.

    150 GitHub stars~3.4k tokensUpdated yesterday
    Auto-check passed

Works with

Questions about Creating Video

What does Creating Video do?

Create video from prompts by overseeing multi-clip AI generation end to end: write a shot list, generate each scene with Gemini Omni Flash (via the Cloudflare AI Gateway), review the results, and…. Creating Video is an agent skill from oaustegard/claude-skills. Create video from prompts by overseeing multi-clip AI generation end to end: write a shot list, generate each scene with Gemini Omni Flash (via the Cloudflare AI Gateway), review the results, and assemble them into a finished cut.

When should I use Creating Video?

Creating Video fits situations like: the user asks to make/generate a video; A multi-scene clip from a script; they mention Omni; acting as the editing/director agent over generated footage.

How do I install Creating Video in Claude Code?

Run `npx skills add oaustegard/claude-skills --skill creating-video -a claude-code`. Or copy the skill folder (creating-video in oaustegard/claude-skills) into .claude/skills/creating-video in your project. Claude Code loads it when a task matches its description.

How do I install Creating Video in Codex?

Run `npx skills add oaustegard/claude-skills --skill creating-video -a codex`. Or copy the skill folder (creating-video in oaustegard/claude-skills) into .agents/skills/creating-video in your project. Codex loads it when a task matches its description.

Can I use Creating Video in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add oaustegard/claude-skills --skill creating-video -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/creating-video, .gemini/skills/creating-video, .github/skills/creating-video and .opencode/skills/creating-video in your project.

What does Creating Video need to run?

Going by SKILL.md and its folder, Creating Video needs Python for the scripts in its folder, the command-line tools its instructions call (python3) and credentials named CF_API_TOKEN. Our summary lists: Python 3; A credential in CF_API_TOKEN.

Does Creating Video access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Creating Video safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Creating Video use?

Creating Video is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Creating Video use?

About 2k tokens (SKILL.md is roughly 7.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Creating Video?

Skills that share tags, products or a category with Creating Video: Video Generation (bytedance/deer-flow, 84k stars), Video Cover Image (itwanger/toBeBetterJavaer, 18k stars), Seedance (songguoxs/seedance-prompt-skill, 2.9k stars) and HyperFrames Video Entry Point (heygen-com/hyperframes, 59k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Creating Video?

oaustegard (a GitHub user) maintains it in oaustegard/claude-skills, which has 150 GitHub stars. The repository holds 66 skills in this directory. The repository was last updated on October 8, 2026.

Source: oaustegard/claude-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.