Agent skill

Ltx Video

by LeoYeAI in LeoYeAI/openclaw-master-skills

Generate videos via LTX-2.3 API (ltx.video). An agent skill from LeoYeAI/openclaw-master-skills.

MITAuto-check: notesMedia & Creative

Install Ltx Video

skills CLI
$ npx skills add LeoYeAI/openclaw-master-skills --skill ltx-video -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install LeoYeAI/openclaw-master-skills ltx-video --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/LeoYeAI/openclaw-master-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/ltx-video .claude/skills/ltx-video && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
ltx-video
GitHub stars
2.2k
Token cost
~1.6k tokens
SKILL.md length
458 words
Files
2
Skills in repo
1,235
Repo updated
First seen
Licence
MIT

At a glance

Generate videos via LTX-2.3 API (ltx.video). An agent skill from LeoYeAI/openclaw-master-skills.

  • Works in 7 steps: Use Verbs, Not Nouns → Block the Scene Like a Director → Describe Audio Explicitly (for… → …
  • : generating AI video from text/image/audio
  • SKILL.md covers API Reference, Quick Examples, ⚠️ Critical Rules (learned… and Prompting Guide (LTX-2.3), plus 2 more sections
  • Calls curl, python3 and ffmpeg; reaches api.ltx.video and uguu.se; needs LTX_API_KEY and API_KEY

What it does

Ltx Video is an agent skill from LeoYeAI/openclaw-master-skills. Generate videos via LTX-2.3 API (ltx.video). Supports text-to-video, image-to-video, audio-to-video (lip-sync from audio + image), extend, and retake. Use when: generating AI video from text/image/audio, animating a portrait, creating lip-sync video from an existing image + audio recording.

Its SKILL.md is about 1.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 1 other file.

It sits in Media & Creative, covering AI video generation. The repository describes itself as: 🧠 Curated collection of 1209+ best OpenClaw skills — weekly updated by MyClaw.ai. The licence is MIT.

When your agent uses it

  • : generating AI video from text/image/audio
  • Animating a portrait
  • Creating lip-sync video from an existing image + audio recording

Example prompts

  • “/ltx-video”

Requirements

  • Python 3
  • A credential in API_KEY
  • A credential in LTX_API_KEY

Workflow steps

7 steps, taken from the step headings in SKILL.md.

  1. Use Verbs, Not Nouns
  2. Block the Scene Like a Director
  3. Describe Audio Explicitly (for text-to-video)
  4. Avoid Static Photo-Like Prompts
  5. Describe Texture and Material
  6. Portrait (9:16) Native
  7. Complex Shots Work Now

What it can do on your machine

Read from SKILL.md and the folder at commit e5199b5. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • curl
    • python3
    • ffmpeg
    • git

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • api.ltx.video
    • uguu.se
    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • LTX_API_KEY
    • API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Ltx Video loads about 1.6k tokens when it runs. Until then it costs about 75 tokens; SKILL.md has 458 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~75
When it runs · the whole SKILL.md, loaded when a task matches
~1.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteMentions a .env fileSKILL.md:194
    Paul's key: stored in `~/clawd/.env` as `LTX_API_KEY`

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from LeoYeAI/openclaw-master-skills at commit e5199b5, republished under its MIT licence (© LeoYeAI). 458 words, ~1,648 tokens.

Download SKILL.mdSave it as .claude/skills/ltx-video/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
ltx-video
description
Generate videos via LTX-2.3 API (ltx.video). Supports text-to-video, image-to-video, audio-to-video (lip-sync from audio + image), extend, and retake. Use when: generating AI video from text/image/audio, animating a portrait, creating lip-sync video from an existing image + audio recording.

LTX-2.3 Video API

API Reference

Base URL: https://api.ltx.video/v1
Auth: Authorization: Bearer <API_KEY>
Response: MP4 binary (direct download, no polling)

Endpoints
EndpointInputUse
/v1/text-to-videopromptGenerate video from text
/v1/image-to-videoimage_uri + promptAnimate a still image
/v1/audio-to-videoaudio_uri + image_uri + promptLip-sync video from audio + image
/v1/extendvideo_uri + promptExtend a video at start or end
/v1/retakevideo_uri + time rangeRegenerate a section of a video
Models
ModelSpeedQuality
ltx-2-3-fast~17sGood (use for tests)
ltx-2-3-pro~30-60sBest (use for final)
Supported Resolutions
  • 1920x1080 (landscape 16:9)
  • 1080x1920 (portrait 9:16 — native vertical, trained on vertical data)
  • 1440x1080, 4096x2160 (text-to-video only)

audio-to-video only supports: 1920x1080 or 1080x1920


Quick Examples

Text to Video
bash
curl -X POST "https://api.ltx.video/v1/text-to-video" \
  -H "Authorization: Bearer $LTX_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "prompt": "A man in a navy blue suit sits at a luxury restaurant table...",
    "model": "ltx-2-3-pro",
    "duration": 8,
    "resolution": "1920x1080"
  }' -o output.mp4
Audio to Video (Lip-sync)
bash
curl -X POST "https://api.ltx.video/v1/audio-to-video" \
  -H "Authorization: Bearer $LTX_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "audio_uri": "https://example.com/voice.mp3",
    "image_uri": "https://example.com/portrait.jpg",
    "prompt": "A man speaks directly to camera...",
    "model": "ltx-2-3-pro",
    "resolution": "1920x1080"
  }' -o output.mp4
Python Wrapper
python
import requests

def ltx_audio_to_video(audio_url, image_url, prompt, api_key,
                        model="ltx-2-3-pro", resolution="1920x1080",
                        output_path="output.mp4"):
    r = requests.post(
        "https://api.ltx.video/v1/audio-to-video",
        headers={"Authorization": f"Bearer {api_key}",
                 "Content-Type": "application/json"},
        json={"audio_uri": audio_url, "image_uri": image_url,
              "prompt": prompt, "model": model, "resolution": resolution},
        timeout=300, stream=True
    )
    if r.status_code != 200:
        raise RuntimeError(f"LTX error {r.status_code}: {r.text}")
    with open(output_path, "wb") as f:
        for chunk in r.iter_content(8192): f.write(chunk)
    return output_path

⚠️ Critical Rules (learned from experience)

File Hosting
  • URLs must be HTTPS — HTTP is rejected
  • Files must return correct MIME type (not application/octet-stream)
  • uguu.se works: upload with curl -F "files[]=@file.mp3" https://uguu.se/upload
  • Audio: upload as MP3 (not WAV) → uguu returns audio/mpeg ✅
  • 4K images fail → resize to 1920x1080 before uploading
bash
# Upload MP3 to uguu.se
AUDIO_URL=$(curl -s -F "files[]=@audio.mp3" "https://uguu.se/upload" | \
  python3 -c "import sys,json; print(json.load(sys.stdin)['files'][0]['url'])")

# Upload image
IMAGE_URL=$(curl -s -F "files[]=@portrait.jpg" "https://uguu.se/upload" | \
  python3 -c "import sys,json; print(json.load(sys.stdin)['files'][0]['url'])")
Image Size Limit
bash
# Resize large images before upload
ffmpeg -y -i input_4k.png -vf "scale=1920:1080" output_1080.jpg
Face Consistency
  • Avoid prompts where the character looks down — breaks face consistency
  • Keep head level and gaze forward throughout
  • Place objects already in frame instead of having character reach below frame
Last Frame
  • LTX does not support first+last frame natively
  • Workaround: generate clip A, generate clip B, then use /v1/extend to chain them

Prompting Guide (LTX-2.3)

LTX-2.3 has a much stronger text connector. Specificity wins.

1. Use Verbs, Not Nouns

❌ "A dramatic portrait of a man standing"
✅ "A man stands on a rooftop. His coat flaps in the wind. He adjusts his collar and steps forward as the camera tracks right."

Show full SKILL.md (191 more words)Show less
2. Block the Scene Like a Director
  • Specify left vs right, foreground vs background
  • Describe who moves, what moves, how they move, what the camera does
  • Spatial relationships are now respected
3. Describe Audio Explicitly (for text-to-video)
  • Name the type of sound: dialogue, ambient, music
  • Specify tone and intensity
  • Example: "His voice is clear and warm. Restaurant ambient sound softly in the background."
4. Avoid Static Photo-Like Prompts
  • If the prompt reads like a still image → the output behaves like one
  • Add wind, motion, breathing, gestures, camera movement
5. Describe Texture and Material
  • Hair, fabric, surface finish, lighting fall-off
  • "Individual hair strands visible in the backlight" → now renders correctly
6. Portrait (9:16) Native
  • resolution: "1080x1920" → trained on vertical data
  • Frame for vertical intentionally, don't treat as cropped landscape
7. Complex Shots Work Now
  • Layer multiple actions: "He picks up the banana, raises it to his ear, and smirks"
  • Combine character performance + environment + camera motion
Lip-Sync Prompt Template
A [description of person] sits/stands [location]. He/she speaks directly 
to camera, lips moving in perfect sync with his/her voice. [Gesture details]. 
Head stays level and gaze remains locked on camera throughout. 
[Environment description softly blurred in background]. 
[Lighting]. [Camera: holds steady at eye level, front-on].

ComfyUI Node

Custom nodes for ComfyUI (no manual API calls):

bash
cd ComfyUI/custom_nodes
git clone https://github.com/PauldeLavallaz/comfyui-ltx-node

Nodes: LTX Text to Video, LTX Image to Video, LTX Extend Video
Category: LTX Video


API Key

Paul's key: stored in ~/clawd/.env as LTX_API_KEY

ltxv_RfSU5hdKJb_g5dwbECZWnilE1P8dJzbavz6niP_0LQJ942ARHIVhrBCfebcytEL1efLVx_63S_PJyWTzicrBcWEkOXfCbGTl8JSzlJJk329MwRViEgOoE2KnE9LIA5t6QSFeBy7DLnTIcX0AZNbV9Jv0TuC7qcq2gV33G6ROhUVUDCuN

© LeoYeAI, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in skills/ltx-video of LeoYeAI/openclaw-master-skills.

  • SKILL.md
  • LICENSE.txt

Open the folder on GitHubat commit e5199b5

Compare with similar skills

Ltx Video next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Ltx Video compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Ltx Video this skillLeoYeAI/openclaw-master-skills2.2k—~1.6kAutomated safety check: NotesMIT
Video Generationbytedance/deer-flow84k3 repos~1.4kAutomated safety check: PassMIT
Video Cover Imageitwanger/toBeBetterJavaer18k—~3.3kAutomated safety check: PassNone
Seedancesongguoxs/seedance-prompt-skill2.9k1 repos~2.5kAutomated safety check: PassNone
HyperFrames Video Entry Pointheygen-com/hyperframes60k3 repos~5.2kAutomated safety check: PassApache-2.0
Lanshu Create AI Presenter Videocclank/lanshu-create-ai-presenter-video2.6k—~3.6kAutomated safety check: PassMIT

Similar skills

  • Video Generation

    bytedance/deer-flow

    Generates short videos from a structured JSON prompt, optionally guided by a reference image used as the first or last frame.

    84k GitHub starsUsed in 3 repos~1.4k tokens
    Media & CreativeAuto-check passed
  • Video Cover Image

    itwanger/toBeBetterJavaer

    Generate matched 3:4, 16:9, and 4:3 short-video cover images from toBeBetterJavaer video scripts or AI/Java technical topics.

    18k GitHub stars~3.3k tokensUpdated today
    Media & CreativeAuto-check passed
  • Seedance

    songguoxs/seedance-prompt-skill

    This skill should be used when the user asks to "generate video prompts", "create Seedance prompts", "write video descriptions", mentions "Seedance", "seedance", "即梦", "即梦平台", "视频提示词", "视频生成"…

    2.9k GitHub starsUsed in 1 repo~2.5k tokens
    Media & CreativeAuto-check passed
  • HyperFrames Video Entry Point

    heygen-com/hyperframes

    Entry point for making, editing and rendering videos from HTML compositions with HyperFrames, routing each request to the right workflow.

    60k GitHub starsUsed in 3 repos~5.2k tokens
    Media & CreativeAuto-check passed
  • Lanshu Create AI Presenter Video

    cclank/lanshu-create-ai-presenter-video

    Turn a topic or finished script into a complete, publish-ready explainer video — led by an AI presenter from an authorized adult presenter image, or performed in one of nine visual explainer styles…

    2.6k GitHub stars~3.6k tokensUpdated 4 days ago
    Media & CreativeAuto-check passed
  • Video Shots

    eternityspring/reelbench-skills

    拉片:把一条成片拆成逐镜头的分析表——每个镜头的时长、景别、类别、运镜、画面. An agent skill from eternityspring/reelbench-skills.

    878 GitHub starsUsed in 1 repo~1.8k tokens
    Media & CreativeAuto-check: notes

More from LeoYeAI/openclaw-master-skills

All 1,235 skills in this repo
  • DevOps Pipeline Management

    LeoYeAI/openclaw-master-skills

    Manages pipelines on a DevOps quality and efficiency platform through its OpenAPI: list workspaces and templates, create, update, run and cancel pipelines, and read run records.

    2.2k GitHub stars~4.2k tokensUpdated 2 mo ago
    Auto-check: notes
  • Feishu Document Collaboration

    LeoYeAI/openclaw-master-skills

    Patches OpenClaw's Feishu extension so an edited document triggers an isolated agent session that reads the doc and replies inline, turning it into a live chat space.

    2.2k GitHub stars~2k tokensUpdated 2 mo ago
    Auto-check passed
  • Files Memory System

    LeoYeAI/openclaw-master-skills

    Multi-context memory management system for OpenClaw agents with group-isolated storage, global shared memory, workspace organization, and group-specific skills isolation.

    2.2k GitHub stars~3.8k tokensUpdated 2 mo ago
    Auto-check passed
  • GEO-Claw AI Visibility Agent

    LeoYeAI/openclaw-master-skills

    Runs a brand's AI-search visibility work end to end: diagnosing how AI platforms represent it, repositioning it, producing AI-optimized content and monitoring ongoing mentions.

    2.2k GitHub stars~4.7k tokensUpdated 2 mo ago
    Auto-check passed
  • Google Workspace CLI

    LeoYeAI/openclaw-master-skills

    Installs and authenticates the gws CLI, then automates Gmail, Drive, Sheets, Calendar, Docs, Chat and Tasks with ready-made recipes, persona bundles and security audits.

    2.2k GitHub stars~2.6k tokensUpdated 2 mo ago
    Auto-check: notes
  • HealthFit Health Advisors

    LeoYeAI/openclaw-master-skills

    Runs four advisor roles, a fitness coach, nutritionist, data analyst and TCM practitioner, to build a health profile and track workouts, diet and wellness over time.

    2.2k GitHub stars~4.4k tokensUpdated 2 mo ago
    Auto-check passed

Questions about Ltx Video

What does Ltx Video do?

Generate videos via LTX-2.3 API (ltx.video). An agent skill from LeoYeAI/openclaw-master-skills. Ltx Video is an agent skill from LeoYeAI/openclaw-master-skills.video).

When should I use Ltx Video?

Ltx Video fits situations like: : generating AI video from text/image/audio; animating a portrait; creating lip-sync video from an existing image + audio recording.

How do I install Ltx Video in Claude Code?

Run `npx skills add LeoYeAI/openclaw-master-skills --skill ltx-video -a claude-code`. Or copy the skill folder (skills/ltx-video in LeoYeAI/openclaw-master-skills) into .claude/skills/ltx-video in your project. Claude Code loads it when a task matches its description.

How do I install Ltx Video in Codex?

Run `npx skills add LeoYeAI/openclaw-master-skills --skill ltx-video -a codex`. Or copy the skill folder (skills/ltx-video in LeoYeAI/openclaw-master-skills) into .agents/skills/ltx-video in your project. Codex loads it when a task matches its description.

Can I use Ltx Video in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add LeoYeAI/openclaw-master-skills --skill ltx-video -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ltx-video, .gemini/skills/ltx-video, .github/skills/ltx-video and .opencode/skills/ltx-video in your project.

What does Ltx Video need to run?

Going by SKILL.md and its folder, Ltx Video needs the command-line tools its instructions call (curl, python3, ffmpeg and git) and credentials named LTX_API_KEY and API_KEY. Our summary lists: Python 3; A credential in API_KEY; A credential in LTX_API_KEY.

Does Ltx Video access the network?

SKILL.md names 3 domains. In commands or code: api.ltx.video, uguu.se and github.com; the agent is likely to contact these when it follows the instructions. This is read from the text; nothing was executed.

Is Ltx Video safe to install?

Our automated static check of SKILL.md found notes only (mentions a .env file), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Ltx Video use?

Ltx Video is published under the MIT licence (from the LICENSE file in the skill folder). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Ltx Video use?

About 1.6k tokens (SKILL.md is roughly 6.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Ltx Video?

Skills that share tags, products or a category with Ltx Video: Video Generation (bytedance/deer-flow, 84k stars), Video Cover Image (itwanger/toBeBetterJavaer, 18k stars), Seedance (songguoxs/seedance-prompt-skill, 2.9k stars) and HyperFrames Video Entry Point (heygen-com/hyperframes, 60k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Ltx Video?

LeoYeAI (a GitHub user) maintains it in LeoYeAI/openclaw-master-skills, which has 2,161 GitHub stars. The repository holds 1,235 skills in this directory. The repository was last updated on July 20, 2026.

Source: LeoYeAI/openclaw-master-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.