Agent skill

Capy Video Gen Skill

by happycapy-ai in happycapy-ai/Happycapy-skills

Multi-shot AI video generation pipeline with face identity consistency.

MITAuto-check: notesMedia & Creative

Install Capy Video Gen Skill

skills CLI
$ npx skills add happycapy-ai/Happycapy-skills --skill capy-video-gen-skill -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install happycapy-ai/Happycapy-skills capy-video-gen-skill --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/happycapy-ai/Happycapy-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/capy-video-gen-skill .claude/skills/capy-video-gen-skill && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
capy-video-gen-skill
GitHub stars
137
Token cost
~2.4k tokens
SKILL.md length
673 words
Files
75 (incl. assets)
Skills in repo
28
Repo updated
First seen
Licence
MIT

At a glance

Multi-shot AI video generation pipeline with face identity consistency.

  • Works in 8 steps: Extract characters from script with… → Generate front/side/back character… → Design shot-by-shot storyboard → …
  • The user asks to create a video from a script
  • SKILL.md covers Overview, Installation Location, Prerequisites and Quick Start, plus 8 more sections
  • Runs Python scripts from its folder; calls python; reaches ai-gateway.happycapy.ai; needs AI_GATEWAY_API_KEY

What it does

Capy Video Gen Skill is an agent skill from happycapy-ai/Happycapy-skills. Multi-shot AI video generation pipeline with face identity consistency. Converts scripts or ideas into complete videos using character extraction, storyboarding, frame generation, and video assembly. 300 experiments validated, 70% face distance improvement. Use when the user asks to create a video from a script, story, idea, or wants multi-shot video with consistent characters.

Its SKILL.md is about 2.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 75 other files, including assets (for example `Communication.md`, `FACE_IDENTITY_GUIDE.md` and `README.md`).

It sits in Media & Creative, covering AI video generation. The repository describes itself as: A curated collection of high-quality Claude Code skills to enhance your development workflow. The licence is MIT.

When your agent uses it

  • The user asks to create a video from a script
  • Wants multi-shot video with consistent characters

Example prompts

  • “/capy-video-gen-skill”

Requirements

  • Python 3
  • A credential in AI_GATEWAY_API_KEY
  • Pre-approved tools (allowed-tools): Bash, Read, Write, Edit

Workflow steps

8 steps, taken from the first numbered list in SKILL.md.

  1. Extract characters from script with detailed physical features
  2. Generate front/side/back character portraits
  3. Design shot-by-shot storyboard
  4. Decompose each shot into first_frame, last_frame, and motion descriptions
  5. Build camera tree for shot relationships
  6. Generate frames with reference image selection (face identity as top priority)
  7. Generate video clips from frames
  8. Concatenate into final video

What it can do on your machine

Read from SKILL.md and the folder at commit 9ff72fe. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Bash
    • Read
    • Write
    • Edit

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python, from the files we listed), which the agent can run.

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • ai-gateway.happycapy.ai

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • AI_GATEWAY_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Capy Video Gen Skill loads about 2.4k tokens when it runs. Until then it costs about 100 tokens; SKILL.md has 673 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~100
When it runs · the whole SKILL.md, loaded when a task matches
~2.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Bash, Read, Write, Edit

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from happycapy-ai/Happycapy-skills at commit 9ff72fe, republished under its MIT licence (© happycapy-ai). 673 words, ~2,426 tokens.

Download SKILL.mdSave it as .claude/skills/capy-video-gen-skill/SKILL.md (or your agent's skills folder). This skill also uses 74 other files; get the full folder from GitHub.
name
capy-video-gen-skill
description
Multi-shot AI video generation pipeline with face identity consistency. Converts scripts or ideas into complete videos using character extraction, storyboarding, frame generation, and video assembly. 300 experiments validated, 70% face distance improvement. Use when the user asks to create a video from a script, story, idea, or wants multi-shot video with consistent characters.
allowed-tools
Bash, Read, Write, Edit

Capy Video Gen Skill - Script-to-Video Pipeline

Generate complete multi-shot videos from scripts or ideas with consistent character faces across all scenes. Built for HappyCapy AI Gateway. 300 experiments validated, 70% face distance improvement.

Overview

ViMax converts text scripts into full videos through an automated pipeline:

  1. Extract characters from script with detailed physical features
  2. Generate front/side/back character portraits
  3. Design shot-by-shot storyboard
  4. Decompose each shot into first_frame, last_frame, and motion descriptions
  5. Build camera tree for shot relationships
  6. Generate frames with reference image selection (face identity as top priority)
  7. Generate video clips from frames
  8. Concatenate into final video

Installation Location

The ViMax pipeline code is at: /home/node/a0/workspace/527fb591-1439-4b5b-ad5d-90f972773f95/workspace/tmp/ViMax/

All commands must be run from this directory using the venv:

bash
cd /home/node/a0/workspace/527fb591-1439-4b5b-ad5d-90f972773f95/workspace/tmp/ViMax

Prerequisites

  • AI_GATEWAY_API_KEY environment variable (auto-configured in HappyCapy)
  • Python venv at .venv/ (already set up)

Quick Start

Script-to-Video

Edit the script, requirements, and style in the entry script, then run:

bash
cd /home/node/a0/workspace/527fb591-1439-4b5b-ad5d-90f972773f95/workspace/tmp/ViMax
.venv/bin/python main_happycapy_script2video.py
Idea-to-Video

For generating from a brief idea (auto-generates script first):

bash
cd /home/node/a0/workspace/527fb591-1439-4b5b-ad5d-90f972773f95/workspace/tmp/ViMax
.venv/bin/python main_happycapy_idea2video.py

Programmatic Usage

python
import asyncio
from langchain.chat_models import init_chat_model
from tools.render_backend import RenderBackend
from utils.config_loader import load_config
from pipelines.script2video_pipeline import Script2VideoPipeline

config = load_config("configs/happycapy_script2video.yaml")
chat_model = init_chat_model(**config["chat_model"]["init_args"])
backend = RenderBackend.from_config(config)

pipeline = Script2VideoPipeline(
    chat_model=chat_model,
    image_generator=backend.image_generator,
    video_generator=backend.video_generator,
    working_dir=config["working_dir"],
)

# Run the pipeline
asyncio.run(pipeline(
    script="Your script here...",
    user_requirement="No more than 8 shots total.",
    style="Cinematic, warm lighting"
))

Pipelines

Script2VideoPipeline
  • Input: A formatted screenplay/script with character dialogue and scene descriptions
  • Output: Concatenated video at {working_dir}/final_video.mp4
  • Config: configs/happycapy_script2video.yaml
Idea2VideoPipeline
  • Input: A brief idea/concept (1-3 paragraphs)
  • Output: Auto-generates a script, then produces video
  • Config: configs/happycapy_idea2video.yaml

Configuration

HappyCapy configs at configs/happycapy_script2video.yaml:

yaml
chat_model:
  init_args:
    model: gpt-4.1
    model_provider: openai
    api_key: ${AI_GATEWAY_API_KEY}
    base_url: https://ai-gateway.happycapy.ai/api/v1/openai/v1

image_generator:
  class_path: tools.ImageGeneratorHappyCapyAPI
  init_args:
    api_key: ${AI_GATEWAY_API_KEY}
    model: google/gemini-3.1-flash-image-preview

video_generator:
  class_path: tools.VideoGeneratorHappyCapyAPI
  init_args:
    api_key: ${AI_GATEWAY_API_KEY}
    model: google/veo-3.1-generate-preview

working_dir: .working_dir/script2video

Key Components

Agents (AI Processing)
AgentFilePurpose
CharacterExtractoragents/character_extractor.pyExtract characters with static/dynamic features from script
CharacterPortraitsGeneratoragents/character_portraits_generator.pyGenerate front/side/back portraits for each character
StoryboardArtistagents/storyboard_artist.pyDesign shot-by-shot storyboard with first/last frames and motion
ReferenceImageSelectoragents/reference_image_selector.pySelect best reference images for each frame (face identity #1 priority)
CameraImageGeneratoragents/camera_image_generator.pyBuild camera trees and generate transition videos
BestImageSelectoragents/best_image_selector.pySelect best generated image from candidates
Screenwriteragents/screenwriter.pyGenerate scripts from ideas
Tools (Generation Backends)
ToolFilePurpose
ImageGeneratorHappyCapyAPItools/image_generator_happycapy_api.pyImage generation via HappyCapy Gateway (Gemini)
VideoGeneratorHappyCapyAPItools/video_generator_happycapy_api.pyVideo generation via HappyCapy Gateway (Veo)
RenderBackendtools/render_backend.pyFactory for instantiating generators from config
Interfaces (Data Models)
  • CharacterInScene - Character with identifier, static_features, dynamic_features
  • ShotDescription - Shot with ff_desc, lf_desc, motion_desc, variation_type
  • Camera - Camera with parent-child relationships
  • Frame - Frame with shot_idx, frame_type, visible characters
  • ImageOutput / VideoOutput - Generation outputs with save methods

Face Identity Consistency (CRITICAL)

This pipeline includes face identity improvements validated through 257 experiments (70% improvement in face distance, from 0.74 to 0.22):

Built-In Protections
  1. Reference Image Selector: Face identity is the #1 priority when selecting reference images. The front-view portrait is always included when a character's face is visible.

  2. Character Portraits: Enhanced prompts generate identity-critical details (exact nose shape, eye spacing, jawline, distinguishing marks) for cross-scene recognition.

  3. Video Prompt Face Lock: Every video generation prompt is prepended with a face identity instruction requiring the character's face to remain identical to the starting frame throughout the clip.

Show full SKILL.md (233 more words)Show less
Best Practices When Using ViMax
  • Hyper-detailed character descriptions: Include ethnicity, age, hair texture/style/color, eye shape, facial hair, glasses, skin tone, build, and distinguishing marks in your script's character introductions
  • Extreme close-up shots: Include at least one extreme close-up per character to anchor identity
  • Consistent lighting: Specify similar lighting across scenes to prevent face drift
  • User-provided reference photos: Place photos in the working directory and pass them as character_portraits_registry to skip AI portrait generation
What Does NOT Work
  • Complex prompt engineering (viseme morphing, phoneme anchoring) does not improve face identity
  • Simple, direct prompts with detailed physical descriptions outperform clever prompts
  • Lip-sync to external audio is NOT possible (Veo generates its own internal audio)

See FACE_IDENTITY_GUIDE.md in the ViMax directory for full details.

Output Structure

After a run, the working directory contains:

.working_dir/script2video/
  characters.json                      # Extracted characters
  character_portraits_registry.json    # Portrait paths registry
  character_portraits/                 # Generated portraits
    0_CharacterName/
      front.png
      side.png
      back.png
  storyboard.json                     # Shot descriptions
  camera_tree.json                    # Camera relationships
  shots/
    0/
      shot_description.json
      first_frame.png
      last_frame.png (if medium/large variation)
      video.mp4
    1/
      ...
  final_video.mp4                     # Final concatenated output

Customization

Using Your Own Reference Photos

To use real photos instead of AI-generated portraits:

python
# Build a portrait registry pointing to your photos
character_portraits_registry = {
    "Alice": {
        "front": {"path": "/path/to/alice_front.png", "description": "Front view of Alice"},
        "side": {"path": "/path/to/alice_side.png", "description": "Side view of Alice"},
        "back": {"path": "/path/to/alice_back.png", "description": "Back view of Alice"},
    }
}

# Pass to pipeline (skips portrait generation)
await pipeline(
    script=script,
    user_requirement=user_requirement,
    style=style,
    character_portraits_registry=character_portraits_registry,
)
Changing Models

Edit the YAML config to use different models:

  • Image: google/gemini-3.1-flash-image-preview (recommended for face identity)
  • Video: google/veo-3.1-generate-preview (recommended) or openai/sora-2
  • Chat: gpt-4.1 (recommended) or any OpenAI-compatible model

Troubleshooting

"No module named 'tools'" or similar import errors

Run from the ViMax root directory:

bash
cd /home/node/a0/workspace/527fb591-1439-4b5b-ad5d-90f972773f95/workspace/tmp/ViMax
.venv/bin/python main_happycapy_script2video.py
API rate limit errors

Reduce max_requests_per_minute in the YAML config.

Face identity drift in generated videos
  • Add more physical detail to character descriptions in your script
  • Use user-provided reference photos instead of AI-generated portraits
  • Include extreme close-up shots for important characters
  • Keep lighting consistent across scenes

© happycapy-ai, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 74 other files (assets) in skills/capy-video-gen-skill of happycapy-ai/Happycapy-skills.

  • SKILL.md
  • .gitignore
  • Communication.md
  • FACE_IDENTITY_GUIDE.md
  • LICENSE
  • README.md
  • README_ZH.md
  • TODO
  • agents/__init__.py
  • agents/best_image_selector.py
  • agents/camera_image_generator.py
  • agents/character_extractor.py
  • agents/character_portraits_generator.py
  • agents/event_extractor.py
  • agents/global_information_planner.py
  • agents/novel_compressor.py
  • agents/reference_image_selector.py
  • agents/scene_extractor.py
  • agents/screenwriter.py
  • agents/script_enhancer.py
  • … and 55 more

Open the folder on GitHubat commit 9ff72fe

Compare with similar skills

Capy Video Gen Skill next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Capy Video Gen Skill compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Capy Video Gen Skill this skillhappycapy-ai/Happycapy-skills137—~2.4kAutomated safety check: NotesMIT
Video Generationbytedance/deer-flow84k3 repos~1.4kAutomated safety check: PassMIT
Video Cover Imageitwanger/toBeBetterJavaer18k—~3.3kAutomated safety check: PassNone
Seedancesongguoxs/seedance-prompt-skill2.9k1 repos~2.5kAutomated safety check: PassNone
HyperFrames Video Entry Pointheygen-com/hyperframes60k3 repos~5.2kAutomated safety check: PassApache-2.0
Lanshu Create AI Presenter Videocclank/lanshu-create-ai-presenter-video2.6k—~3.6kAutomated safety check: PassMIT

Similar skills

  • Video Generation

    bytedance/deer-flow

    Generates short videos from a structured JSON prompt, optionally guided by a reference image used as the first or last frame.

    84k GitHub starsUsed in 3 repos~1.4k tokens
    Media & CreativeAuto-check passed
  • Video Cover Image

    itwanger/toBeBetterJavaer

    Generate matched 3:4, 16:9, and 4:3 short-video cover images from toBeBetterJavaer video scripts or AI/Java technical topics.

    18k GitHub stars~3.3k tokensUpdated yesterday
    Media & CreativeAuto-check passed
  • Seedance

    songguoxs/seedance-prompt-skill

    This skill should be used when the user asks to "generate video prompts", "create Seedance prompts", "write video descriptions", mentions "Seedance", "seedance", "即梦", "即梦平台", "视频提示词", "视频生成"…

    2.9k GitHub starsUsed in 1 repo~2.5k tokens
    Media & CreativeAuto-check passed
  • HyperFrames Video Entry Point

    heygen-com/hyperframes

    Entry point for making, editing and rendering videos from HTML compositions with HyperFrames, routing each request to the right workflow.

    60k GitHub starsUsed in 3 repos~5.2k tokens
    Media & CreativeAuto-check passed
  • Lanshu Create AI Presenter Video

    cclank/lanshu-create-ai-presenter-video

    Turn a topic or finished script into a complete, publish-ready explainer video — led by an AI presenter from an authorized adult presenter image, or performed in one of nine visual explainer styles…

    2.6k GitHub stars~3.6k tokensUpdated today
    Media & CreativeAuto-check passed
  • Video Shots

    eternityspring/reelbench-skills

    拉片:把一条成片拆成逐镜头的分析表——每个镜头的时长、景别、类别、运镜、画面. An agent skill from eternityspring/reelbench-skills.

    878 GitHub starsUsed in 1 repo~1.8k tokens
    Media & CreativeAuto-check: notes

More from happycapy-ai/Happycapy-skills

All 28 skills in this repo
  • Happycapy Skill Creator

    happycapy-ai/Happycapy-skills

    Automate HappyCapy skill creation by finding and adapting existing skills from anthropics/skills repository.

    137 GitHub stars~627 tokensUpdated 1 mo ago
    Auto-check passed
  • 360 Panorama Viewer

    happycapy-ai/Happycapy-skills

    Build a fully self-contained 360° equirectangular panorama viewer as a single HTML file.

    137 GitHub stars~2.1k tokensUpdated 1 mo ago
    Auto-check passed
  • Generate Image

    happycapy-ai/Happycapy-skills

    Generate and transform images using AI Gateway API. An agent skill from happycapy-ai/Happycapy-skills.

    137 GitHub stars~2k tokensUpdated 1 mo ago
    Auto-check: notes
  • LLM Council

    happycapy-ai/Happycapy-skills

    Multi-model LLM Council with live dashboard. An agent skill from happycapy-ai/Happycapy-skills.

    137 GitHub stars~581 tokensUpdated 1 mo ago
    Auto-check passed
  • PPTX

    happycapy-ai/Happycapy-skills

    Create polished PowerPoint (.pptx) presentations directly from a topic or content description.

    137 GitHub stars~1.5k tokensUpdated 1 mo ago
    Auto-check passed
  • World Class Carousel

    happycapy-ai/Happycapy-skills

    Generate world-class Instagram carousel content on any topic.

    137 GitHub stars~16k tokensUpdated 1 mo ago
    Auto-check: notes

Questions about Capy Video Gen Skill

What does Capy Video Gen Skill do?

Multi-shot AI video generation pipeline with face identity consistency. Capy Video Gen Skill is an agent skill from happycapy-ai/Happycapy-skills. Multi-shot AI video generation pipeline with face identity consistency.

When should I use Capy Video Gen Skill?

Capy Video Gen Skill fits situations like: the user asks to create a video from a script; wants multi-shot video with consistent characters.

How do I install Capy Video Gen Skill in Claude Code?

Run `npx skills add happycapy-ai/Happycapy-skills --skill capy-video-gen-skill -a claude-code`. Or copy the skill folder (skills/capy-video-gen-skill in happycapy-ai/Happycapy-skills) into .claude/skills/capy-video-gen-skill in your project. Claude Code loads it when a task matches its description.

How do I install Capy Video Gen Skill in Codex?

Run `npx skills add happycapy-ai/Happycapy-skills --skill capy-video-gen-skill -a codex`. Or copy the skill folder (skills/capy-video-gen-skill in happycapy-ai/Happycapy-skills) into .agents/skills/capy-video-gen-skill in your project. Codex loads it when a task matches its description.

Can I use Capy Video Gen Skill in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add happycapy-ai/Happycapy-skills --skill capy-video-gen-skill -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/capy-video-gen-skill, .gemini/skills/capy-video-gen-skill, .github/skills/capy-video-gen-skill and .opencode/skills/capy-video-gen-skill in your project.

What does Capy Video Gen Skill need to run?

Going by SKILL.md and its folder, Capy Video Gen Skill needs Python for the scripts in its folder, the command-line tools its instructions call (python) and credentials named AI_GATEWAY_API_KEY. Our summary lists: Python 3; A credential in AI_GATEWAY_API_KEY. Its frontmatter pre-approves these tools: Bash, Read, Write, Edit.

Does Capy Video Gen Skill access the network?

SKILL.md names 1 domain. In commands or code: ai-gateway.happycapy.ai; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is Capy Video Gen Skill safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Capy Video Gen Skill use?

Capy Video Gen Skill is published under the MIT licence (from the LICENSE file in the skill folder). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Capy Video Gen Skill use?

About 2.4k tokens (SKILL.md is roughly 9.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Capy Video Gen Skill?

Skills that share tags, products or a category with Capy Video Gen Skill: Video Generation (bytedance/deer-flow, 84k stars), Video Cover Image (itwanger/toBeBetterJavaer, 18k stars), Seedance (songguoxs/seedance-prompt-skill, 2.9k stars) and HyperFrames Video Entry Point (heygen-com/hyperframes, 60k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Capy Video Gen Skill?

happycapy-ai (a GitHub user) maintains it in happycapy-ai/Happycapy-skills, which has 137 GitHub stars. The repository holds 28 skills in this directory. The repository was last updated on September 3, 2026.

Source: happycapy-ai/Happycapy-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.