Agent skill

LTX-2.3 Video Generation

by digitalsamba in digitalsamba/claude-code-video-toolkit

Generates roughly five-second video clips from a text prompt or a still image with the LTX-2.3 22B model, run through a Modal endpoint by `tools/ltx2.py`.

MITAuto-check: notesMedia & Creative

Install LTX-2.3 Video Generation

skills CLI
$ npx skills add digitalsamba/claude-code-video-toolkit --skill ltx2 -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install digitalsamba/claude-code-video-toolkit ltx2 --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/digitalsamba/claude-code-video-toolkit.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/ltx2 .claude/skills/ltx2 && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
ltx2
GitHub stars
2.2k
Used in
2 other repos
Token cost
~2.4k tokens
SKILL.md length
711 words
Files
1
Skills in repo
11
Repo updated
First seen
Licence
MIT

At a glance

Generates roughly five-second video clips from a text prompt or a still image with the LTX-2.3 22B model, run through a Modal endpoint by `tools/ltx2.py`.

  • Generating a short b-roll clip from a text description
  • SKILL.md covers Quick Reference, Parameters, Style LoRAs and Valid Frame Counts, plus 5 more sections
  • Calls uv and ffmpeg; needs HF_TOKEN
  • Animating a still image into a few seconds of motion

What it does

The agent calls `tools/ltx2.py` with `uv run` to produce a clip of about five seconds, either from a text prompt or from an input image for image-to-video. The model runs on Modal on an A100-80GB GPU, so a `MODAL_LTX2_ENDPOINT_URL` must be set in `.env`.

Parameters include width and height (defaults 768 and 512, both divisible by 64), frame count (default 121, which must satisfy (n-1) % 8 == 0), frames per second (24), a `standard` quality mode of 30 steps or a `fast` one of 15, plus seed, output path and negative prompt. A style LoRA option currently offers `crt-terminal` for CRT and pixel-art terminal looks. It adds a trigger word, defaults to 1024 by 1024 at 121 frames and loosens the negative prompt so on-screen text survives. Switching LoRAs rebuilds the pipeline, about 60 seconds. On-screen text should stay to one to three words.

When your agent uses it

  • Generating a short b-roll clip from a text description
  • Animating a still image into a few seconds of motion
  • Making animated backgrounds or motion content for a video project
  • Creating a CRT-terminal style animation of typed text

Example prompts

  • “Generate a five second clip of a sunset over the ocean with golden light on the waves.”
  • “Animate this product photo into a slow push-in video.”
  • “Make a CRT terminal clip of a command being typed in green pixel font.”

Requirements

  • A Modal endpoint with `MODAL_LTX2_ENDPOINT_URL` set in `.env`
  • `uv` to run `tools/ltx2.py`

What it can do on your machine

Read from SKILL.md and the folder at commit 2c99460. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • uv
    • ffmpeg

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • huggingface.co

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • HF_TOKEN

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

LTX-2.3 Video Generation loads about 2.4k tokens when it runs. Until then it costs about 80 tokens; SKILL.md has 711 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~80
When it runs · the whole SKILL.md, loaded when a task matches
~2.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteMentions a .env fileSKILL.md:9
    . Requires `MODAL_LTX2_ENDPOINT_URL` in `.env`.
  • NoteMentions a .env fileSKILL.md:215
    # 3. Save endpoint URL to .env
  • NoteMentions a .env fileSKILL.md:216
    toolkit-ltx2-ltx2-generate.modal.run" >> .env

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from digitalsamba/claude-code-video-toolkit at commit 2c99460, republished under its MIT licence (© digitalsamba). 711 words, ~2,435 tokens.

Download SKILL.mdSave it as .claude/skills/ltx2/SKILL.md (or your agent's skills folder).
name
ltx2
description
AI video generation with LTX-2.3 22B — text-to-video, image-to-video clips for video production. Use when generating video clips, animating images, creating b-roll, animated backgrounds, or motion content. Triggers include video generation, animate image, b-roll, motion, video clip, text-to-video, image-to-video.

LTX-2.3 Video Generation

Generate ~5 second video clips from text prompts or images using the LTX-2.3 22B DiT model. Runs on Modal (A100-80GB). Requires MODAL_LTX2_ENDPOINT_URL in .env.

Quick Reference

bash
# Text-to-video
uv run tools/ltx2.py --prompt "A sunset over the ocean, golden light on waves, cinematic" --output sunset.mp4

# Image-to-video (animate a still image)
uv run tools/ltx2.py --prompt "Gentle camera drift, soft ambient motion" --input photo.jpg --output animated.mp4

# Custom resolution and duration
uv run tools/ltx2.py --prompt "..." --width 1024 --height 576 --num-frames 161 --output wide.mp4

# Fast mode (fewer steps, quicker)
uv run tools/ltx2.py --prompt "..." --quality fast --output quick.mp4

# Reproducible output
uv run tools/ltx2.py --prompt "..." --seed 42 --output reproducible.mp4

Parameters

ParameterDefaultDescription
--prompt(required)Text description of the video
--input-Input image for image-to-video
--width768Video width (divisible by 64)
--height512Video height (divisible by 64)
--num-frames121Frame count, must satisfy (n-1) % 8 == 0
--fps24Frames per second
--qualitystandardstandard (30 steps) or fast (15 steps)
--steps30Override inference steps directly
--seedrandomSeed for reproducibility
--outputautoOutput file path
--negative-promptsensible defaultWhat to avoid
--loranoneStyle LoRA preset. Currently: crt-terminal.

Style LoRAs

Style LoRAs bias the output toward a specific visual aesthetic. They're baked into the Modal image and selected per-request; switching LoRAs forces a pipeline rebuild (~60s one-time cost per container lifetime per switch).

crt-terminal — CRT / pixel-art terminals

Base: LTX-2.3 22B, trained by @lovis93 (Apache 2.0).

bash
# Trigger word is auto-prepended — write the prompt normally
uv run tools/ltx2.py --lora crt-terminal \
  --prompt "a terminal typing out \"\\$ claude --continue\" character by character in glowing green pixel font, scanlines, phosphor glow, low choppy frame rate, hacker mood" \
  --output crt_claude.mp4

What the preset changes:

  • Prepends crtanim, to the prompt (the LoRA's trigger word)
  • Defaults to 1024×1024, 121 frames (the ratio it was trained on)
  • Relaxes the default negative prompt so on-screen text isn't filtered out

Prompt pattern: <CRT aesthetic> → <color palette> → <animation style> → <subject> → <literal text in quotes> → <mood>. Keep on-screen text to 1–3 words — the model can't render long strings reliably. The LoRA prefers static framing; ask for camera moves explicitly if you want them.

Valid Frame Counts

(n - 1) % 8 == 0: 25 (~1s), 49 (~2s), 73 (~3s), 97 (~4s), 121 (~5s default), 161 (~6.7s), 193 (~8s max practical).

Common Resolutions

ResolutionRatioNotes
768x5123:2Default, good balance
512x5121:1Square, fastest
1024x57616:9Widescreen
576x10249:16Portrait/vertical

Prompting Guide

LTX-2 responds well to cinematographic descriptions. Layer these dimensions:

  • Camera: "Slow dolly forward", "Aerial drone shot", "Tracking shot", "Static wide angle"
  • Lighting: "Golden hour", "Cinematic lighting", "Neon-lit", "Soft diffused light"
  • Motion: "Timelapse of...", "Slow motion", "Gentle camera drift", "Gradually transitions"
  • Style: "Shot on 35mm film", "Documentary style", "Clean minimal aesthetic"
  • Negative: Always implicitly avoids "worst quality, blurry, jittery, watermark, text, logo"

Keep prompts under 200 words. Be specific about the scene.

Good Prompts
# Atmospheric b-roll
"Aerial drone shot slowly flying over turquoise ocean waves breaking on white sand, golden hour sunlight, cinematic"

# Product/tech scene
"Close-up of hands typing on a mechanical keyboard, shallow depth of field, soft desk lamp lighting, cozy atmosphere"

# Abstract background
"Dark moody abstract background with flowing blue light streaks, subtle geometric grid, bokeh particles floating, cinematic tech atmosphere"

# Animate a portrait
"Professional headshot, subtle natural head movement, confident warm expression, studio lighting, shallow depth of field"

# Animate a slide/screenshot
"Gentle subtle particle effects floating across a presentation slide, soft ambient light shifts, very slight camera drift"
Bad Prompts
# Too vague
"A cool video"

# Too many competing ideas
"A cat riding a skateboard while juggling fire on the moon during a thunderstorm"

# Describing text/UI (model can't render text reliably)
"A website showing the text 'Welcome to our platform'"

Video Production Use Cases

B-Roll Clips

Generate atmospheric 5s shots for cutaways between narrated scenes:

bash
uv run tools/ltx2.py --prompt "Futuristic holographic interface, glowing data visualizations, clean workspace, cinematic" --output broll_tech.mp4
uv run tools/ltx2.py --prompt "Aerial view of European city at golden hour, modern architecture" --output broll_europe.mp4
Animated Slide Backgrounds

Feed a slide screenshot and add subtle motion:

bash
uv run tools/ltx2.py --prompt "Gentle particle effects, soft ambient light shifts, very slight camera drift" --input slide.png --output animated_slide.mp4
Animated Portraits

Bring still headshots to life:

bash
uv run tools/ltx2.py --prompt "Subtle natural head movement, warm expression, professional lighting" --input headshot.png --output animated_portrait.mp4
Show full SKILL.md (323 more words)Show less
Stylized Character Cameo (SadTalker Alternative)

For non-realistic faces — fantasy characters, masked figures, heavy beards, helmets, illustrations — SadTalker often produces uncanny or broken lip sync because it's trained on photoreal humans. LTX-2 image-to-video is frequently a better choice when lip-sync precision isn't critical (the viewer's brain fills in the gap as long as something is moving). Prompt for motion + atmosphere, not phonemes:

bash
uv run tools/ltx2.py \
  --input character_portrait.png \
  --prompt "Ancient warrior speaks slowly with gravitas, beard shifts subtly, glowing aura pulses, embers drift past, slow head movement, cinematic close-up, mystical atmosphere" \
  --width 768 --height 768 \
  --output character_speaking.mp4

When LTX-2 wins over SadTalker:

  • Stylized / illustrated / fantasy characters
  • Heavy facial hair or accessories obscuring the mouth
  • Masked or helmeted figures
  • Short cameo lines where atmosphere matters more than precision
  • Dramatic VO rather than dialogue

When SadTalker still wins:

  • Photoreal human presenters
  • Full sentences where mouth shape needs to match phonemes
  • Tutorials / talking-head explainers where the viewer is effectively reading lips
Branded Intro/Outro

Generate abstract motion backgrounds for title cards:

bash
uv run tools/ltx2.py --prompt "Dark moody background with flowing blue and coral light streaks, bokeh particles, cinematic tech atmosphere, no text" --output intro_bg.mp4
Combining with Other Tools

LTX-2 generates raw clips. Combine with the rest of the toolkit:

WorkflowTools
Generate clip → upscaleltx2.py → upscale.py
Generate clip → add to Remotionltx2.py → use as <OffthreadVideo> in composition
Generate image → animateflux2.py → ltx2.py --input
Generate clip → extract audioltx2.py → ffmpeg -i clip.mp4 -vn audio.wav
Generate clip → add voiceoverltx2.py → mix with qwen3_tts.py output

Technical Details

  • Model: LTX-2.3 22B DiT (Lightricks), bf16
  • GPU: A100-80GB on Modal (~$4.68/hr)
  • Inference: ~2.5 min per clip (768x512, 121 frames, 30 steps)
  • Cost: ~$0.20-0.25 per 5s clip
  • Cold start: ~60-90s (loading ~55GB weights)
  • Output: H.264 MP4 with synchronized ambient audio (24fps)
  • Max duration: ~8s (193 frames) per clip
Known Limitations
  • Training data artifacts: ~30% of generations may have unwanted logos/text from training data. Re-run with different --seed.
  • Text rendering: Cannot reliably generate readable text in video. Use Remotion overlays instead.
  • Max duration: ~8s per clip. Longer content needs stitching.
  • Audio: Generated audio is ambient/environmental only. Use voiceover/music tools for speech and music.
  • License: Community License — free under $10M revenue, commercial license needed above that.

Setup

bash
# 1. Create Modal secret for HuggingFace (one-time)
uv run modal secret create huggingface-token HF_TOKEN=hf_your_token

# 2. Deploy (downloads ~55GB of weights, takes ~10 min)
uv run modal deploy docker/modal-ltx2/app.py

# 3. Save endpoint URL to .env
echo "MODAL_LTX2_ENDPOINT_URL=https://yourname--video-toolkit-ltx2-ltx2-generate.modal.run" >> .env

# 4. Test
uv run tools/ltx2.py --prompt "A candle flickering on a dark table, cinematic" --output test.mp4

Important: HuggingFace token needs read-access scope. Accept the Gemma 3 license before deploying. Unauthenticated downloads are severely rate-limited.

© digitalsamba, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/ltx2 of digitalsamba/claude-code-video-toolkit.

Open the folder on GitHubat commit 2c99460

Used in 2 other repositories

We found 5 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 2 other GitHub owners. This page covers the copy in digitalsamba/claude-code-video-toolkit, which our catalogue first saw on October 7, 2026.

Compare with similar skills

LTX-2.3 Video Generation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

LTX-2.3 Video Generation compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
LTX-2.3 Video Generation this skilldigitalsamba/claude-code-video-toolkit2.2k2 repos~2.4kAutomated safety check: NotesMIT
Video Shotseternityspring/reelbench-skills8682 repos~1.8kAutomated safety check: NotesApache-2.0
HyperFrames Video Entry Pointheygen-com/hyperframes59k3 repos~5.2kAutomated safety check: PassApache-2.0
Lanshu Create AI Presenter Videocclank/lanshu-create-ai-presenter-video2.5k—~3.6kAutomated safety check: PassMIT
Video Scrubeternityspring/reelbench-skills8681 repos~1.3kAutomated safety check: NotesApache-2.0
Video ComposeUtopai-Research/pai-code3541 repos~4.1kAutomated safety check: PassCustom licence

Similar skills

  • Video Shots

    eternityspring/reelbench-skills

    拉片:把一条成片拆成逐镜头的分析表——每个镜头的时长、景别、类别、运镜、画面. An agent skill from eternityspring/reelbench-skills.

    868 GitHub starsUsed in 2 repos~1.8k tokens
    Media & CreativeAuto-check: notes
  • HyperFrames Video Entry Point

    heygen-com/hyperframes

    Entry point for making, editing and rendering videos from HTML compositions with HyperFrames, routing each request to the right workflow.

    59k GitHub starsUsed in 3 repos~5.2k tokens
    Media & CreativeAuto-check passed
  • Lanshu Create AI Presenter Video

    cclank/lanshu-create-ai-presenter-video

    Turn a topic or finished script into a complete, publish-ready explainer video — led by an AI presenter from an authorized adult presenter image, or performed in one of nine visual explainer styles…

    2.5k GitHub stars~3.6k tokensUpdated 2 days ago
    Media & CreativeAuto-check passed
  • Video Scrub

    eternityspring/reelbench-skills

    把一条视频重建成「只有画面和声音」的干净文件——源片的元数据一概不搬: GPS、设备型号、账号 ID、创建时间、章节、GoPro 的遥测轨,全部留在原地。

    868 GitHub starsUsed in 1 repo~1.3k tokens
    Media & CreativeAuto-check: notes
  • Video Compose

    Utopai-Research/pai-code

    Generates and prompts video clips on the filmmaking canvas. An agent skill from Utopai-Research/pai-code.

    354 GitHub starsUsed in 1 repo~4.1k tokens
    Media & CreativeAuto-check passed
  • Ergo Remotion Video

    itwanger/toBeBetterJavaer

    把口播稿做成二哥风格的 Remotion 视频,包括整理视频用稿、火山 TTS 配音、音画对齐、逐章动画预览和导出带配音的 MP4。用户说“做视频”“口播稿转视频”“Remotion”“继续做下一章”“出片”“渲染”“改读音”“配音读错了”,或给出 docs/src/ai/video/ 下的稿子要做成视频时使用。共享工具、配置和素材在…

    18k GitHub stars~1.1k tokensUpdated today
    Media & CreativeAuto-check passed

More from digitalsamba/claude-code-video-toolkit

All 11 skills in this repo
  • FFmpeg for Video Production

    digitalsamba/claude-code-video-toolkit

    Command recipes for converting, resizing, compressing, trimming and extracting audio from video with FFmpeg, including settings for Remotion projects.

    2.2k GitHub starsUsed in 3 repos~3.3k tokens
    Auto-check passed
  • ElevenLabs Voiceover Generator

    digitalsamba/claude-code-video-toolkit

    Generates narration, sound effects and cloned voices through the ElevenLabs API, with model and setting choices tuned to the content's style.

    2.2k GitHub starsUsed in 1 repo~2.7k tokens
    Auto-check: notes
  • Ideogram 4 Prompt Builder

    digitalsamba/claude-code-video-toolkit

    Turns a casual image request into the structured JSON caption Ideogram 4 needs for legible on-image text, exact brand colors and controlled layout.

    2.2k GitHub starsUsed in 1 repo~1.3k tokens
    Auto-check: notes
  • Playwright Browser Demo Recording

    digitalsamba/claude-code-video-toolkit

    Records browser interactions as video with Playwright, covering viewport sizing, cursor highlighting, and converting output for Remotion.

    2.2k GitHub starsUsed in 1 repo~3.2k tokens
    Auto-check passed
  • Qwen Edit

    digitalsamba/claude-code-video-toolkit

    AI image editing prompting patterns for Qwen-Image-Edit. An agent skill from digitalsamba/claude-code-video-toolkit.

    2.2k GitHub starsUsed in 1 repo~711 tokens
    Auto-check passed
  • ACE-Step Music Generation

    digitalsamba/claude-code-video-toolkit

    Generates background music, vocal tracks, covers and stems with ACE-Step 1.5 through a bundled music_gen.py tool, using cloud or self-hosted providers.

    2.2k GitHub stars~3.3k tokensUpdated 3 days ago
    Auto-check: notes

Questions about LTX-2.3 Video Generation

What does LTX-2.3 Video Generation do?

Generates roughly five-second video clips from a text prompt or a still image with the LTX-2.3 22B model, run through a Modal endpoint by `tools/ltx2.py`. py` with `uv run` to produce a clip of about five seconds, either from a text prompt or from an input image for image-to-video.env`.

When should I use LTX-2.3 Video Generation?

LTX-2.3 Video Generation fits situations like: generating a short b-roll clip from a text description; animating a still image into a few seconds of motion; making animated backgrounds or motion content for a video project; creating a CRT-terminal style animation of typed text.

How do I install LTX-2.3 Video Generation in Claude Code?

Run `npx skills add digitalsamba/claude-code-video-toolkit --skill ltx2 -a claude-code`. Or copy the skill folder (.claude/skills/ltx2 in digitalsamba/claude-code-video-toolkit) into .claude/skills/ltx2 in your project. Claude Code loads it when a task matches its description.

How do I install LTX-2.3 Video Generation in Codex?

Run `npx skills add digitalsamba/claude-code-video-toolkit --skill ltx2 -a codex`. Or copy the skill folder (.claude/skills/ltx2 in digitalsamba/claude-code-video-toolkit) into .agents/skills/ltx2 in your project. Codex loads it when a task matches its description.

Can I use LTX-2.3 Video Generation in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add digitalsamba/claude-code-video-toolkit --skill ltx2 -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ltx2, .gemini/skills/ltx2, .github/skills/ltx2 and .opencode/skills/ltx2 in your project.

What does LTX-2.3 Video Generation need to run?

Going by SKILL.md and its folder, LTX-2.3 Video Generation needs the command-line tools its instructions call (uv and ffmpeg) and credentials named HF_TOKEN. Our summary lists: A Modal endpoint with `MODAL_LTX2_ENDPOINT_URL` set in `.env`; `uv` to run `tools/ltx2.py`.

Does LTX-2.3 Video Generation access the network?

SKILL.md names 1 domain. As links in the text: huggingface.co. This is read from the text; nothing was executed.

Is LTX-2.3 Video Generation safe to install?

Our automated static check of SKILL.md found notes only (mentions a .env file), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does LTX-2.3 Video Generation use?

LTX-2.3 Video Generation is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does LTX-2.3 Video Generation use?

About 2.4k tokens (SKILL.md is roughly 9.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to LTX-2.3 Video Generation?

Skills that share tags, products or a category with LTX-2.3 Video Generation: Video Shots (eternityspring/reelbench-skills, 868 stars), HyperFrames Video Entry Point (heygen-com/hyperframes, 59k stars), Lanshu Create AI Presenter Video (cclank/lanshu-create-ai-presenter-video, 2.5k stars) and Video Scrub (eternityspring/reelbench-skills, 868 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains LTX-2.3 Video Generation?

digitalsamba (a GitHub organization) maintains it in digitalsamba/claude-code-video-toolkit, which has 2,174 GitHub stars. The repository holds 11 skills in this directory. The repository was last updated on October 5, 2026.

Source: digitalsamba/claude-code-video-toolkit on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.