Agent skill

Video Generation

by bytedance in bytedance/deer-flow

Generates short videos from a structured JSON prompt, optionally guided by a reference image used as the first or last frame.

MITAuto-check passedMedia & Creative

Install Video Generation

skills CLI
$ npx skills add bytedance/deer-flow --skill video-generation -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install bytedance/deer-flow video-generation --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/bytedance/deer-flow.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/public/video-generation .claude/skills/video-generation && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
video-generation
GitHub stars
84k
Used in
3 other repos
Token cost
~1.4k tokens
SKILL.md length
460 words
Files
2 (incl. scripts)
Skills in repo
23
Repo updated
First seen
Licence
MIT

At a glance

Generates short videos from a structured JSON prompt, optionally guided by a reference image used as the first or last frame.

  • Works in 4 steps: Understand Requirements → Create Structured Prompt → Create Reference Image (Optional when… → …
  • Asked to create or imagine a short video clip
  • SKILL.md covers Overview, Core Capabilities, Workflow and Video Generation Example, plus 3 more sections
  • Runs Python scripts from its folder; calls python; reaches api.minimaxi.com; needs GEMINI_API_KEY and MINIMAX_API_KEY

What it does

The agent starts by pinning down what the video should show: the subject, the style and mood, technical specs such as aspect ratio and lighting, and any reference image. It then writes a structured JSON prompt file into `/mnt/user-data/workspace/` and, when the image-generation skill is available, can create a reference image first. A single supplied image is used as the guiding frame of the video.

Generation runs through `scripts/generate.py`, called with a prompt file, an output file, optional reference images and an aspect ratio that defaults to 16:9. The skill tells the agent to call the script with its parameters rather than read its source. The worked example builds a short clip of a film's opening scene after searching for that scene online. The excerpt is cut off before the remaining example steps.

When your agent uses it

  • Asked to create or imagine a short video clip
  • Animating from a reference image used as the first or last frame
  • Writing a structured JSON prompt for a video model

Example prompts

  • “Generate a short video of a fox running through snowy woods at dawn.”
  • “Make a clip that starts from this product photo: ./assets/lamp.jpg.”
  • “Create a 16:9 video of a paper boat drifting down a rain-filled street.”

Requirements

  • Python to run `scripts/generate.py`

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Understand Requirements
  2. Create Structured Prompt
  3. Create Reference Image (Optional when image-generation skill is available)
  4. Execute Generation

What it can do on your machine

Read from SKILL.md and the folder at commit ab9920f. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • api.minimaxi.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • GEMINI_API_KEY
    • MINIMAX_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Video Generation loads about 1.4k tokens when it runs. Until then it costs about 41 tokens; SKILL.md has 460 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~41
When it runs · the whole SKILL.md, loaded when a task matches
~1.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from bytedance/deer-flow at commit ab9920f, republished under its MIT licence (© bytedance). 460 words, ~1,371 tokens.

Download SKILL.mdSave it as .claude/skills/video-generation/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
video-generation
description
Use this skill when the user requests to generate, create, or imagine videos. Supports structured prompts and reference image for guided generation.

Video Generation Skill

Overview

This skill generates high-quality videos using structured prompts and a Python script. The workflow includes creating JSON-formatted prompts and executing video generation with optional reference image.

Core Capabilities

  • Create structured JSON prompts for AIGC video generation
  • Support reference image as guidance or the first/last frame of the video
  • Generate videos through automated Python script execution

Workflow

Step 1: Understand Requirements

When a user requests video generation, identify:

  • Subject/content: What should be in the image
  • Style preferences: Art style, mood, color palette
  • Technical specs: Aspect ratio, composition, lighting
  • Reference image: Any image to guide generation
  • You don't need to check the folder under /mnt/user-data
Step 2: Create Structured Prompt

Generate a structured JSON file in /mnt/user-data/workspace/ with naming pattern: {descriptive-name}.json

Step 3: Create Reference Image (Optional when image-generation skill is available)

Generate reference image for the video generation.

  • If only 1 image is provided, use it as the guided frame of the video
Step 3: Execute Generation

Call the Python script:

bash
python /mnt/skills/public/video-generation/scripts/generate.py \
  --prompt-file /mnt/user-data/workspace/prompt-file.json \
  --reference-images /path/to/ref1.jpg \
  --output-file /mnt/user-data/outputs/generated-video.mp4 \
  --aspect-ratio 16:9

Parameters:

  • --prompt-file: Absolute path to JSON prompt file (required)
  • --reference-images: Absolute paths to reference image (optional)
  • --output-file: Absolute path to output image file (required)
  • --aspect-ratio: Aspect ratio of the generated image (optional, default: 16:9)

[!NOTE] Do NOT read the python file, instead just call it with the parameters.

Video Generation Example

User request: "Generate a short video clip depicting the opening scene from "The Chronicles of Narnia: The Lion, the Witch and the Wardrobe"

Step 1: Search for the opening scene of "The Chronicles of Narnia: The Lion, the Witch and the Wardrobe" online

Step 2: Create a JSON prompt file with the following content:

json
{
  "title": "The Chronicles of Narnia - Train Station Farewell",
  "background": {
    "description": "World War II evacuation scene at a crowded London train station. Steam and smoke fill the air as children are being sent to the countryside to escape the Blitz.",
    "era": "1940s wartime Britain",
    "location": "London railway station platform"
  },
  "characters": ["Mrs. Pevensie", "Lucy Pevensie"],
  "camera": {
    "type": "Close-up two-shot",
    "movement": "Static with subtle handheld movement",
    "angle": "Profile view, intimate framing",
    "focus": "Both faces in focus, background soft bokeh"
  },
  "dialogue": [
    {
      "character": "Mrs. Pevensie",
      "text": "You must be brave for me, darling. I'll come for you... I promise."
    },
    {
      "character": "Lucy Pevensie",
      "text": "I will be, mother. I promise."
    }
  ],
  "audio": [
    {
      "type": "Train whistle blows (signaling departure)",
      "volume": 1
    },
    {
      "type": "Strings swell emotionally, then fade",
      "volume": 0.5
    },
    {
      "type": "Ambient sound of the train station",
      "volume": 0.5
    }
  ]
}

Step 3: Use the image-generation skill to generate the reference image

Load the image-generation skill and generate a single reference image narnia-farewell-scene-01.jpg according to the skill.

Step 4: Use the generate.py script to generate the video

bash
python /mnt/skills/public/video-generation/scripts/generate.py \
  --prompt-file /mnt/user-data/workspace/narnia-farewell-scene.json \
  --reference-images /mnt/user-data/outputs/narnia-farewell-scene-01.jpg \
  --output-file /mnt/user-data/outputs/narnia-farewell-scene-01.mp4 \
  --aspect-ratio 16:9

Do NOT read the python file, just call it with the parameters.

Show full SKILL.md (143 more words)Show less

Output Handling

After generation:

  • Videos are typically saved in /mnt/user-data/outputs/
  • Share generated videos (come first) with user as well as generated image if applicable, using present_files tool
  • Provide brief description of the generation result
  • Offer to iterate if adjustments needed

Notes

  • Always use English for prompts regardless of user's language
  • JSON format ensures structured, parsable prompts
  • Reference image enhance generation quality significantly
  • Iterative refinement is normal for optimal results

Providers (Gemini / MiniMax)

Provider credentials are read from the runtime environment, not embedded in the script. Do not put their values in prompt files or command-line arguments.

Auto-selected by environment variables (CLI unchanged):

  • GEMINI_API_KEY set → Gemini Veo (default, unchanged).
  • Only MINIMAX_API_KEY set → MiniMax video (/v1/video_generation, async 3-step poll/download).
  • Force with VIDEO_GENERATION_PROVIDER=gemini|minimax.

MiniMax overrides: MINIMAX_API_HOST (default https://api.minimaxi.com), MINIMAX_VIDEO_MODEL (default MiniMax-Hailuo-2.3). The first reference image is used as MiniMax first_frame_image. MiniMax ignores --aspect-ratio (it uses resolution/duration).

© bytedance, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (scripts) in skills/public/video-generation of bytedance/deer-flow.

  • SKILL.md
  • scripts/generate.py

Open the folder on GitHubat commit ab9920f

Used in 3 other repositories

We found 3 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 3 other GitHub owners. This page covers the copy in bytedance/deer-flow, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Video Generation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Video Generation compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Video Generation this skillbytedance/deer-flow84k3 repos~1.4kAutomated safety check: PassMIT
MoneyPrinterTurbo Video Generatorharry0703/MoneyPrinterTurbo129k—~2.1kAutomated safety check: WarnMIT
Vox DirectorAlisa0808/vox-director2.2k—~5.6kAutomated safety check: PassMIT
Gbro Collage Brollpyang5166/gbro-collage-broll1.3k—~2.6kAutomated safety check: NotesMIT
VRGDG H3 Short Film Pipelinevrgamegirl19/comfyui-vrgamedevgirl765—~4.2kAutomated safety check: PassCustom licence
Forge Media Route Layer0x0funky/agent-sprite-forge4.4k—~2.2kAutomated safety check: PassMIT

Similar skills

  • MoneyPrinterTurbo Video Generator

    harry0703/MoneyPrinterTurbo

    Installs and runs MoneyPrinterTurbo to turn a topic or script into a finished short video with voice-over, subtitles, stock footage and music.

    129k GitHub stars~2.1k tokensUpdated today
    Media & CreativeAuto-check: warnings
  • Vox Director

    Alisa0808/vox-director

    Turn ONE topic into a finished Vox-style paper-collage explainer / ad video, end to end on the Atlas Cloud API + local ffmpeg — script, collage keyframes, motion, voice-over, music, captions, all…

    2.2k GitHub stars~5.6k tokensUpdated 4 days ago
    Media & CreativeAuto-check passed
  • Gbro Collage Broll

    pyang5166/gbro-collage-broll

    将约 5 秒口播文稿、观点句或抽象概念做成高级 editorial halftone paper-collage / 半调纸拼贴 B-roll。用户说“collage b-roll”“纸拼贴 b-roll”“半调拼贴”“拼贴风格配画面”“用这段文稿做拼贴动画”“gbro-collage-broll”,或希望把一句文稿转成拼贴视觉隐喻时,必须使用此…

    1.3k GitHub stars~2.6k tokensUpdated 2 mo ago
    Media & CreativeAuto-check: notes
  • VRGDG H3 Short Film Pipeline

    vrgamegirl19/comfyui-vrgamedevgirl

    Builds an AI short film in a local ComfyUI with the VRGDG Video Builder, MiniMax H3 scenes, reference images, a music score, QA and a final edit.

    765 GitHub stars~4.2k tokensUpdated today
    Media & CreativeAuto-check passed
  • Forge Media Route Layer

    0x0funky/agent-sprite-forge

    Generates an image or an image-to-video clip through a configured provider API or a signed-in Codex or Grok CLI, and reports the route, file, hash and cost estimate.

    4.4k GitHub stars~2.2k tokensUpdated 3 days ago
    Media & CreativeAuto-check passed
  • Cassette Video Edit

    Cassette-Editor/oh-my-cassette

    Edit, trim, cut, caption, subtitle, reframe, combine, add background music to, or export video, audio, and image files through Cassette.

    119 GitHub starsUsed in 1 repo~3.4k tokens
    Media & CreativeAuto-check passed

More from bytedance/deer-flow

All 23 skills in this repo
  • Vercel Deploy

    bytedance/deer-flow

    Deploys a project to Vercel with one script and no login, then returns a live preview URL and a claim link for moving the deployment into your own Vercel account.

    84k GitHub starsUsed in 10 repos~797 tokens
    Auto-check passed
  • GitHub Deep Research

    bytedance/deer-flow

    Researches a GitHub repository over four rounds using the GitHub API and web search, then writes a structured markdown report with timeline, metrics and Mermaid diagrams.

    84k GitHub starsUsed in 4 repos~1.3k tokens
    Auto-check passed
  • Chart Visualization

    bytedance/deer-flow

    Picks a suitable chart type from 26 options for your data, maps the data to that chart's parameters and generates a chart image through a JavaScript script.

    84k GitHub starsUsed in 1 repo~840 tokens
    Auto-check passed
  • Excel and CSV Data Analysis

    bytedance/deer-flow

    Analyzes uploaded Excel and CSV files with SQL through DuckDB, producing schema inspections, statistical summaries and exports to CSV, JSON or Markdown.

    84k GitHub starsUsed in 4 repos~2.2k tokens
    Auto-check passed
  • Structured Image Generation

    bytedance/deer-flow

    Turns an image request into a structured JSON prompt and runs a bundled Python script to generate the picture, optionally guided by reference images.

    84k GitHub starsUsed in 4 repos~2.9k tokens
    Auto-check passed
  • DeerFlow Smoke Test

    bytedance/deer-flow

    Walks through an end-to-end smoke test of a DeerFlow deployment: pull the latest code, deploy with Docker or locally, verify services, run health checks and write a report.

    84k GitHub stars~2.5k tokensUpdated today
    Auto-check: notes

Works with

Questions about Video Generation

What does Video Generation do?

Generates short videos from a structured JSON prompt, optionally guided by a reference image used as the first or last frame. The agent starts by pinning down what the video should show: the subject, the style and mood, technical specs such as aspect ratio and lighting, and any reference image. It then writes a structured JSON prompt file into `/mnt/user-data/workspace/` and, when the image-generation skill is available, can create a reference image first.

When should I use Video Generation?

Video Generation fits situations like: asked to create or imagine a short video clip; animating from a reference image used as the first or last frame; writing a structured JSON prompt for a video model.

How do I install Video Generation in Claude Code?

Run `npx skills add bytedance/deer-flow --skill video-generation -a claude-code`. Or copy the skill folder (skills/public/video-generation in bytedance/deer-flow) into .claude/skills/video-generation in your project. Claude Code loads it when a task matches its description.

How do I install Video Generation in Codex?

Run `npx skills add bytedance/deer-flow --skill video-generation -a codex`. Or copy the skill folder (skills/public/video-generation in bytedance/deer-flow) into .agents/skills/video-generation in your project. Codex loads it when a task matches its description.

Can I use Video Generation in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add bytedance/deer-flow --skill video-generation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/video-generation, .gemini/skills/video-generation, .github/skills/video-generation and .opencode/skills/video-generation in your project.

What does Video Generation need to run?

Going by SKILL.md and its folder, Video Generation needs Python for the scripts in its folder, the command-line tools its instructions call (python) and credentials named GEMINI_API_KEY and MINIMAX_API_KEY. Our summary lists: Python to run `scripts/generate.py`.

Does Video Generation access the network?

SKILL.md names 1 domain. In commands or code: api.minimaxi.com; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is Video Generation safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Video Generation use?

Video Generation is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Video Generation use?

About 1.4k tokens (SKILL.md is roughly 5.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Video Generation?

Skills that share tags, products or a category with Video Generation: MoneyPrinterTurbo Video Generator (harry0703/MoneyPrinterTurbo, 129k stars), Vox Director (Alisa0808/vox-director, 2.2k stars), Gbro Collage Broll (pyang5166/gbro-collage-broll, 1.3k stars) and VRGDG H3 Short Film Pipeline (vrgamegirl19/comfyui-vrgamedevgirl, 765 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Video Generation?

bytedance (a GitHub organization) maintains it in bytedance/deer-flow, which has 83,611 GitHub stars. The repository holds 23 skills in this directory. The repository was last updated on October 10, 2026.

Source: bytedance/deer-flow on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.