Agent skill

Ltxv2 Video

by artokun in artokun/comfyui-mcp

Build Lightricks LTX-2 / LTX-2.3 video workflows covering text-to-video, image-to-video, GGUF and bundled checkpoints, distilled model, camera control LoRAs, synchronized audio, two-stage upscaling…

MITAuto-check: warningsMedia & Creative

Install Ltxv2 Video

The automated check flagged lines worth reading first. See the safety section below.

skills CLI
$ npx skills add artokun/comfyui-mcp --skill ltxv2-video -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install artokun/comfyui-mcp ltxv2-video --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/artokun/comfyui-mcp.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugin/skills/ltxv2-video .claude/skills/ltxv2-video && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
ltxv2-video
GitHub stars
803
Token cost
~6.8k tokens
SKILL.md length
2,786 words
Files
2 (incl. references)
Skills in repo
42
Repo updated
First seen
Licence
MIT

At a glance

Build Lightricks LTX-2 / LTX-2.3 video workflows covering text-to-video, image-to-video, GGUF and bundled checkpoints, distilled model, camera control LoRAs, synchronized audio, two-stage upscaling…

  • Works in 6 steps: Use VAEDecodeTiled or… → Start at 768x512 resolution, upscale in… → Use FP4 Gemma text encoder (installed) → …
  • Tasks that involve AI video generation
  • SKILL.md covers Version naming (read this first), ⭐ Render-verified correct…, Overview and Models, plus 4 more sections
  • Calls python and ffmpeg

What it does

Ltxv2 Video is an agent skill from artokun/comfyui-mcp. Build Lightricks LTX-2 / LTX-2.3 video workflows covering text-to-video, image-to-video, GGUF and bundled checkpoints, distilled model, camera control LoRAs, synchronized audio, two-stage upscaling, and swapping alternate/GGUF base models

Its SKILL.md is about 6.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including reference files (for example `references/workflows.md`).

It sits in Media & Creative, covering AI video generation and Fine-tuning. It works with llama.cpp. The repository describes itself as: Local-first, agent-native control plane for ComfyUI — MCP server + sidebar agent that generates images, video & audio, authors and runs workflows, and edits your live graph in… The licence is MIT.

When your agent uses it

  • Tasks that involve AI video generation
  • Tasks that involve Fine-tuning

Example prompts

  • “/ltxv2-video”

Requirements

  • Python 3

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Use VAEDecodeTiled or LTXVSpatioTemporalTiledVAEDecode instead of standard VAEDecode
  2. Start at 768x512 resolution, upscale in Stage 2
  3. Use FP4 Gemma text encoder (installed)
  4. For LTX-2.3, pick the GGUF quant to match VRAM: Q4_K_S (<12GB), Q5_K_S (12 to 16GB), Q8_0 (24GB+). The dev GGUF needs ~20+ steps; the…
  5. Always clear_vram before switching to LTX-V2 from another model family
  6. Reduce frame count to 81 or 49 if OOM persists

What it can do on your machine

Read from SKILL.md and the folder at commit 6ad6fc0. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python
    • ffmpeg

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Ltxv2 Video loads about 6.8k tokens when it runs, and up to ~8.8k if it reads all its reference files. Until then it costs about 63 tokens; SKILL.md has 2,786 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~63
When it runs · the whole SKILL.md, loaded when a task matches
~6.8k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~8.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: warnings

The automated check found patterns that need a careful read before installing.

  • WarningContains zero-width charactersSKILL.md:204
    > the `filename_prefix` (e.g. `ltxv2_…⟨U+200B⟩.mp4`), check the mtime is fresh, then

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from artokun/comfyui-mcp at commit 6ad6fc0, republished under its MIT licence (© artokun). 2,786 words, ~6,754 tokens.

Download SKILL.mdSave it as .claude/skills/ltxv2-video/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
ltxv2-video
description
Build Lightricks LTX-2 / LTX-2.3 video workflows covering text-to-video, image-to-video, GGUF and bundled checkpoints, distilled model, camera control LoRAs, synchronized audio, two-stage upscaling, and swapping alternate/GGUF base models
globs
**/*.json

LTX-2 / LTX-2.3 Video Workflows

Version naming (read this first)

There is no "LTX 3.2" or "LTX2.3" as separate products. The user's shorthand refers to Lightricks LTX-2.3, a point release of the LTX-2 family. The lineage is:

  • LTX-Video (2024): first text-to-video model from Lightricks.
  • LTX-2 / LTX-V2 (Oct 2025): 19B-class DiT audio-video foundation model. Bundled checkpoint ltx-2-19b-distilled.safetensors, Gemma 3 12B text encoder.
  • LTX-2.3 (released ~March 2026): 22B-parameter DiT update. Rebuilt VAE (sharper textures/faces/hair/text), ~4x larger text connector (text projection) for prompt adherence, native 9:16 portrait, LoRA support, HiFi-GAN vocoder for cleaner synchronized audio, up to 4K@50fps / ~20s clips. Apache 2.0. Distributed primarily as GGUF UNets (community quants) plus separate VAE / text-encoder / text-projection files, NOT a single bundled checkpoint like LTX-2.

When the user says "LTX3.2" / "LTX2.3", treat it as LTX-2.3. This skill covers both LTX-2 (bundled checkpoint path) and LTX-2.3 (GGUF UNet path).


⭐ Render-verified correct setup (read this FIRST — 2026-06-19)

The GGUF-UNet + DualCLIPLoader + gemma_3_12B_it_fp4_mixed path documented later in this skill (the Aitrepreneur installer path) produces soft/mushy video with inaccurate faces and eyes. It runs, but it is NOT the quality path. The setup below is the official Comfy-Org template, render-proven sharp (1280×704, accurate faces, synchronized 48 kHz stereo audio).

Models (exact, render-verified)
ComponentFileSource repoFolderNotes
Checkpointltx-2.3-22b-dev.safetensors (46 GB, max quality) or ltx-2.3-22b-dev-fp8.safetensors (~23 GB, official VRAM-friendly)Lightricks/LTX-2.3 / Lightricks/LTX-2.3-fp8checkpoints/ (NOT unet/)The checkpoint carries the transformer and the audio VAE. Loaded by CheckpointLoaderSimple + reused by LTXVAudioVAELoader + LTXAVTextEncoderLoader.
Gemma text encodergemma_3_12B_it_fp8_scaled.safetensors (13 GB)Comfy-Org/ltx-2 → split_files/text_encoders/text_encoders/Use fp8_scaled (unpacked). The Aitrepreneur fp4_mixed mirror file is truncated (5.3 GB vs 9.4 GB) AND a packed-fp4 layout core can't reshape → shape [15360,1920] invalid for input 27582328.
Distilled speed LoRAltx_2.3_22b_distilled_1.1_lora_dynamic_fro09_avg_rank_111_bf16.safetensors @ 0.5Comfy-Org/ltx-2.3 → split_files/loras/loras/The newer dynamic rank-111 distilled LoRA — NOT the older ...384-1.1.
Gemma abliterated LoRA ⭐gemma-3-12b-it-abliterated_lora_rank64_bf16.safetensors @ 1.0Comfy-Org/ltx-2 → split_files/loras/loras/Applied to the text-encoder CLIP via a LoraLoader. This is the prompt-accuracy / correct-eyes fix. Missing this = subtly-wrong faces.
Spatial upscalerltx-2.3-spatial-upscaler-x2-1.1.safetensorsLightricks/LTX-2.3latent_upscale_models/Used by the stage-2 LTXVLatentUpsampler. Use x2-1.1, not x2-1.0.
Node stack (the right one)
  • LTXAVTextEncoderLoader (CORE, comfy_extras/nodes_lt_audio.py) loads gemma + the full checkpoint together via comfy.sd.load_clip([gemma, ckpt], type=LTXV). This is the audio-video encoder driving both video and audio/voice. Do NOT use DualCLIPLoader(type=ltxv) + a separate ltx-2.3_text_projection file. That is the legacy video-only path and yields mush.
  • Gemma abliterated LoRA via a LoraLoader (CLIP LoRA) on the encoder output → CLIPTextEncode.
  • Two-stage: base sample (~768×512) → LTXVLatentUpsampler (×2 spatial, uses the upscaler model + the checkpoint VAE) → refine sample → 1280×704 output. The upscale is the sharpness. A single-stage graph is visibly softer.
  • Guider: the Comfy-Org template uses plain CFGGuider cfg=1 (distilled); the LTXVideo repo example uses MultimodalGuider + GuiderParameters (separate AUDIO/VIDEO) + ClownSampler_Beta (RES4LYF). Both produce sharp output. The LoRAs + two-stage matter more than the guider.
  • ffmpeg is required for the final mux: <comfy-venv>/python -m pip install imageio-ffmpeg, then reboot. CreateVideo/SaveVideo/VHS_VideoCombine fail with ffmpeg ... could not be found otherwise.
Custom nodes

ComfyUI-LTXVideo (LTXV* nodes, MultimodalGuider, GuiderParameters, LTXVPreprocess, LTXVTiledVAEDecode, GemmaAPITextEncode, LTXFloatToInt) + RES4LYF (ClownSampler_Beta, only for the repo-example sampler). LTXAVTextEncoderLoader, ResizeImageMaskNode, CreateVideo, SaveVideo, ManualSigmas, LTXVScheduler, the Primitive* nodes are all CORE ComfyUI.

Quality troubleshooting (symptom → cause → fix)
  • Mushy/garbage, no clear subject → empty positive prompt, or DualCLIPLoader+projection text encoder. Fix: set a prompt; use LTXAVTextEncoderLoader.
  • Coherent but soft/blurry, faces & eyes slightly wrong → no two-stage upscale and/or missing the gemma abliterated LoRA and/or the old distilled LoRA. Fix: full two-stage template + both LoRAs above.
  • status: success but no video file / outputs only has a math or text node → the output node (SaveVideo/VHS) failed validation and was silently dropped; the graph short-circuited. Check the ComfyUI log for Failed to validate prompt for output N and fix that node (missing ffmpeg, a broken connection, a model-not-in-list).
  • DualCLIPLoader reshape [15360,1920] invalid for input 27582328 → wrong/truncated gemma → use gemma_3_12B_it_fp8_scaled.
  • LatentUpscaleModelLoader: ...x2-1.0 not in list → reference ...x2-1.1.
  • SaveVideo writes to a subfolder (video/<prefix>_NNNNN.mp4). Its history outputs entry isn't under images/videos/gifs, so a naive "find the video" check misses it. Look on disk under output/video/.
MCP UI→API converter gotchas (src/services/workflow-converter.ts)

The official template exercised several convertUiToApi gaps, all now fixed. Keep them in mind if a template still mis-converts:

  • V3 dynamic combos (COMFY_DYNAMICCOMBO_V3, e.g. ResizeImageMaskNode.resize_type): each selected option's nested input must be keyed <combo>.<nested> (e.g. resize_type.longer_size, resize_type.width), NOT flat. ComfyUI rebuilds the nested dict via dynamic_paths/finalize_prefix. A flat key is rejected required_input_missing.
  • Reroute is virtual. Its connections must be passed through (consumer resolves to the Reroute's input), else everything downstream dangles and the graph short-circuits.
  • VHS_VideoCombine stores widgets_values as a name→value object, not a positional array.
  • Typed Primitive* nodes (PrimitiveInt/Float/Boolean/StringMultiline) are real executable nodes. Keep them as link sources; don't bake their values into a consumer's widgets_values by index (mis-positions V3 nested inputs).
Pack

packs/ltx-2.3-txt2vid (and the i2v/flf/extender variants) should be built on this official two-stage template. For a no-input-file T2V pack, set the template's bypass_i2v / "Switch to Text to Video?" boolean true and feed the I2V image input a blank EmptyImage (discarded at runtime but still validates).


Source note: the install scripts below pull LTX-2.3 files from a third-party mirror repo huggingface.co/Aitrepreneur/FLX, not the official Lightricks/LTX-2.3 repo. The official weights live at huggingface.co/Lightricks/LTX-2.3. Filenames/quants match what those scripts download.

Overview

LTX-2 is a DiT-based video foundation model from Lightricks. It uses a Gemma 3 12B text encoder and supports both text-to-video (T2V) and image-to-video (I2V). Key features:

  • Distilled model for fast 8-step generation; dev model for higher quality (~20+ steps)
  • Two-stage pipeline: Generate at low res, then 2x spatial upscale in latent space
  • Camera control LoRAs for cinematic movements
  • Synchronized audio-video generation in a single pass (LTX-2.3 audio VAE + HiFi-GAN vocoder)
  • GGUF quantization (LTX-2.3) for low-VRAM local inference via ComfyUI-GGUF

Models

LTX-2 (bundled checkpoint path)
ComponentNodeModelNotes
CheckpointCheckpointLoaderSimpleltx-2-19b-distilled.safetensors41GB bf16, distilled variant; bundles VAE internally
Gemma 3CLIPLoader (type=ltxv)gemma_3_12B_it_fp4_mixed.safetensors9GB FP4, in text_encoders/

Loading note (LTX-2): The bundled checkpoint contains the VAE internally. The Gemma 3 text encoder loads separately via CLIPLoader with type: "ltxv" pointing at text_encoders/.

LTX-2.3 (GGUF UNet path — current install)

LTX-2.3 ships as a separate GGUF UNet + standalone VAE + text encoder + text projection, not a single bundled checkpoint. The install scripts (see below) place files like this:

ComponentNodeModel fileFolderNotes
UNet (GGUF)UnetLoaderGGUF ("Unet Loader (GGUF)", bootleg category, from ComfyUI-GGUF)ltx-2.3-22b-dev-Q4_K_S.gguf / -Q5_K_S.gguf / -Q8_0.ggufmodels/unet/22B dev model. Q4_K_S <12GB VRAM, Q5_K_S 12–16GB, Q8_0 24GB+
Video VAEVAELoaderLTX23_video_vae_bf16.safetensorsmodels/vae/rebuilt LTX-2.3 VAE
Audio VAEVAELoaderLTX23_audio_vae_bf16.safetensorsmodels/vae/only for audio-sync output
Gemma 3CLIPLoader (type=ltxv)gemma_3_12B_it_fp4_mixed.safetensorsmodels/text_encoders/same FP4 encoder as LTX-2
Text projectionloaded with the text encoderltx-2.3_text_projection_bf16.safetensorsmodels/text_encoders/the enlarged text connector new in 2.3
Spatial upscalerLatentUpscaleModelLoaderltx-2.3-spatial-upscaler-x2-1.1.safetensorsmodels/latent_upscale_models/replaces LTX-2's ...x2-1.0

Loading note (LTX-2.3): Because the UNet is a bare GGUF, the VAE no longer comes "for free" with a checkpoint. Load LTX23_video_vae_bf16.safetensors explicitly with VAELoader. Place GGUF UNets in models/unet/ and use the GGUF Unet loader. Some community 2.3 workflows pair gemma_3_12B_it.safetensors (full) instead of the FP4 mixed file; the installer uses the FP4 mixed one.

Install scripts

The exact download commands for both paths live in references/workflows.md.

LoRAs (Installed)
LoRAFilePurpose
Distilled LoRA (384, 2.3)loras/ltx-2.3-22b-distilled-lora-384-1.1.safetensorsApply to the 2.3 dev UNet for fast distilled behavior
IC-LoRA detailerloras/ltx-2-19b-ic-lora-detailer.safetensorsDetail/refinement IC-LoRA
Distilled LoRA (384, LTX-2)ltx2/ltx-2-19b-distilled-lora-384.safetensorsApply to LTX-2 base for distilled behavior
Camera Dolly Leftltx-2-19b-lora-camera-control-dolly-left.safetensorsCamera movement (see Camera Control section)
Concept/Style LoRAs (Installed)

Located in loras/LTXV2/:

  • style/PLORAV7_LTX_000010500.safetensors
  • concept/head_swap_v1_13500_first_frame.safetensors
  • concept/LTX-2 - Better Female Nudity.safetensors
  • action/LTX2-i2v-OralSuite.safetensors
  • action/LTX2-i2v-SexThrust.safetensors
  • And more in concept/ and action/ subfolders

Key Nodes

LTXVConditioning

Binds text conditioning with frame rate information:

json
{
  "class_type": "LTXVConditioning",
  "inputs": {
    "positive": ["<clip_text_encode>", 0],
    "negative": ["<clip_text_encode_neg>", 0],
    "frame_rate": 25
  }
}
EmptyLTXVLatentVideo

Creates the initial video latent (for T2V):

json
{
  "class_type": "EmptyLTXVLatentVideo",
  "inputs": {
    "width": 768,
    "height": 512,
    "length": 97,
    "batch_size": 1
  }
}

Frame count constraint: Must be 8n + 1 (9, 17, 25, 33, 41, 49, 57, 65, 73, 81, 89, 97, 105, 113, 121).

LTXVScheduler

Dedicated sigma schedule for LTX-V2 latent space:

json
{
  "class_type": "LTXVScheduler",
  "inputs": {
    "steps": 8,
    "max_shift": 2.05,
    "base_shift": 0.95,
    "stretch": true,
    "terminal": 0.1
  }
}

Connect the optional latent input for latent-aware shift scaling.

Feeding a prior stage's output into I2V (e.g. Krea2 image → LTX video). The LoadImage that feeds LTXVImgToVideo.image needs the source frame registered as a ComfyUI INPUT. When that frame is an OUTPUT from an earlier stage, call upload_image (action:"stage") with its { filename, subfolder?, type? } and drop the returned input filename into LoadImage. (For a file already on local disk, upload_image (action:"image").) NEVER copy the output file into, or guess, a filesystem input/ path. ComfyUI's input/output dirs may be CUSTOM (--input-directory / --output-directory), so a guessed path makes LoadImage reject the file (Invalid image file) and wastes the render. upload_image (action:"stage") goes through the server API (/view → /upload/image) and resolves the real dirs correctly.

VERIFY A VIDEO RENDER VIA THE FILESYSTEM, NOT /history. VHS_VideoCombine (and similar video nodes) write the .mp4 but frequently do NOT register the output in ComfyUI's /history. The prompt shows done with an empty outputs map and no error. Do NOT conclude the render "silently dropped" from get_history / queue (action:"status") alone. Confirm the file with get_image (action:"list_outputs") (it now lists videos too, with kind: "video"): match the filename_prefix (e.g. ltxv2_….mp4), check the mtime is fresh, then chain it into the next stage with upload_image (action:"stage").

LTXVImgToVideo (For I2V)

All-in-one node that encodes image, creates latent, and wraps conditioning:

json
{
  "class_type": "LTXVImgToVideo",
  "inputs": {
    "positive": ["<conditioning>", 0],
    "negative": ["<conditioning>", 0],
    "vae": ["<checkpoint>", 2],
    "image": ["<load_image>", 0],
    "width": 768,
    "height": 512,
    "length": 97,
    "batch_size": 1,
    "strength": 0.6
  }
}

Gotcha: strength controls motion; DON'T set it to 1.0. LTXVImgToVideo.strength is how strongly the output adheres to the start image: higher = more adherence = LESS motion. Setting it to 1.0 pins every frame to the start image → a FROZEN i2v with ZERO motion (the storyboard frames come out nearly identical). Keep the verified value ~0.6 (as in the example above) for proper motion. If a generated i2v clip shows little/no motion, the FIRST thing to check is that strength wasn't bumped toward 1.0.

LTXVLatentUpsampler (For Two-Stage Upscale)
json
{
  "class_type": "LTXVLatentUpsampler",
  "inputs": {
    "latent": ["<sampler_output>", 0],
    "upscale_model": ["<upscale_loader>", 0]
  }
}

Requires LatentUpscaleModelLoader. Use ltx-2.3-spatial-upscaler-x2-1.1.safetensors for LTX-2.3 (or ltx-2-spatial-upscaler-x2-1.0.safetensors for LTX-2).

Sampler Settings

Distilled Model (Installed)

Uses SamplerCustomAdvanced with manual sigmas, NOT standard KSampler:

ParameterStage 1 (Generate)Stage 2 (Upscale)
samplereulereuler
steps84
cfg1.01.0
schedulerLTXVSchedulerManual sigmas

Stage 1 sigmas (via LTXVScheduler): max_shift=2.05, base_shift=0.95, stretch=true, terminal=0.1

Stage 2 sigmas (manual, for upscale refinement): 0.909375, 0.725, 0.421875, 0.0

Base Model (If Using Distilled LoRA on Base)
ParameterValue
samplerres_2s
steps20
cfg4.0
schedulerLTXVScheduler
distilled_lora_strength0.6

Resolution and Frame Count

Show full SKILL.md (1,136 more words)Show less
Resolutions (Must be multiples of 32)
AspectStage 1After 2x UpscaleNotes
3:2 landscape768x5121536x1024Default
16:9 landscape960x5441920x1088Official example
1:1 square640x6401280x1280
4:3 landscape704x5121408x1024

Start at lower resolution for Stage 1 to manage VRAM, then upscale.

Frame Count (8n + 1)
FramesDuration @25fpsDuration @24fpsNotes
491.96s2.04sQuick test
813.24s3.38sShort clip
973.88s4.04sDefault
1214.84s5.04sOfficial example, recommended
1616.44s6.71sLonger clip
25710.28s10.71sMaximum
Frame Rate

Standard: 25 fps (conditioned via LTXVConditioning). 24 and 30 fps also supported.

Pipeline Flow: T2V Distilled

CheckpointLoaderSimple → MODEL + VAE
CLIPLoader (ltxv, gemma_3_12B_it_fp4_mixed) → CLIP
  ├─ CLIPTextEncode (positive) → CONDITIONING
  └─ CLIPTextEncode (negative) → CONDITIONING

LTXVConditioning (positive, negative, frame_rate=25) → pos/neg CONDITIONING
EmptyLTXVLatentVideo (768x512, 121 frames) → LATENT
LTXVScheduler (steps=8, max_shift=2.05, base_shift=0.95) → SIGMAS

SamplerCustomAdvanced (model, sigmas, positive, negative, latent)
  → Stage 1 LATENT

[Optional: LTXVLatentUpsampler → 2x LATENT → SamplerCustomAdvanced Stage 2]

VAEDecode (or LTXVSpatioTemporalTiledVAEDecode for VRAM savings) → IMAGE
VHS_VideoCombine (or CreateVideo + SaveVideo) → MP4

Complete workflows (API JSON)

Both end-to-end graphs, T2V Distilled (8-Step) and LTX-2.3 GGUF (dev, T2V), are in references/workflows.md.

Camera Control LoRAs

Seven official camera control LoRAs from Lightricks:

MovementLoRA File
Dolly Leftltx-2-19b-lora-camera-control-dolly-left.safetensors
Dolly Rightltx-2-19b-lora-camera-control-dolly-right.safetensors
Dolly Inltx-2-19b-lora-camera-control-dolly-in.safetensors
Dolly Outltx-2-19b-lora-camera-control-dolly-out.safetensors
Jib Upltx-2-19b-lora-camera-control-jib-up.safetensors
Jib Downltx-2-19b-lora-camera-control-jib-down.safetensors
Staticltx-2-19b-lora-camera-control-static.safetensors

Usage: Apply with LoraLoaderModelOnly at strength 1.0. Do NOT describe camera movement in your prompt. The LoRA handles it.

json
{
  "class_type": "LoraLoaderModelOnly",
  "inputs": {
    "model": ["<checkpoint>", 0],
    "lora_name": "ltx-2-19b-lora-camera-control-dolly-left.safetensors",
    "strength_model": 1.0
  }
}

Cannot combine camera control LoRA with IC-LoRA (canny/depth/pose) in the same generation.

Concept/Style LoRAs

Apply with LoraLoaderModelOnly. Typical strength: 0.5 to 1.0.

json
{
  "class_type": "LoraLoaderModelOnly",
  "inputs": {
    "model": ["<checkpoint_or_camera_lora>", 0],
    "lora_name": "LTXV2\\concept\\LTX-2 - Better Female Nudity.safetensors",
    "strength_model": 0.8
  }
}

Concept/style LoRAs CAN be stacked with camera control LoRAs.

VRAM Considerations

ConfigVRAMNotes
bf16 checkpoint + FP4 Gemma~24GB+Tight on RTX 4090, may OOM
FP8 checkpoint + FP4 Gemma~16-20GBRecommended for 24GB GPUs
bf16 + tiled VAE decode~22GBUse LTXVSpatioTemporalTiledVAEDecode

VRAM warnings from MEMORY.md: "LTXV2 can OOM on 24GB — suggest FP8 quantized models or --lowvram"

Tips for 24GB GPUs
  1. Use VAEDecodeTiled or LTXVSpatioTemporalTiledVAEDecode instead of standard VAEDecode
  2. Start at 768x512 resolution, upscale in Stage 2
  3. Use FP4 Gemma text encoder (installed)
  4. For LTX-2.3, pick the GGUF quant to match VRAM: Q4_K_S (<12GB), Q5_K_S (12 to 16GB), Q8_0 (24GB+). The dev GGUF needs ~20+ steps; the distilled LoRA path runs ~8 steps
  5. Always clear_vram before switching to LTX-V2 from another model family
  6. Reduce frame count to 81 or 49 if OOM persists

Prompt Style

Natural language descriptions. Be specific about motion, camera angles, and temporal progression:

Good: "A woman with flowing auburn hair walks through a sun-dappled forest, leaves falling gently around her, soft golden hour lighting, cinematic depth of field"
Bad: "woman, forest, walking"

Describe the entire scene progression, not a single moment. Include lighting, mood, and motion cues.

Two-Stage Upscale Pattern

For production quality, generate at low resolution then upscale:

  1. Stage 1: Generate at 768x512, 121 frames, 8 steps (distilled)
  2. Upscale: LTXVLatentUpsampler (2x spatial) → 1536x1024
  3. Stage 2: Resample the upscaled latent with 3-4 steps at CFG 1.0
  4. Decode: Use tiled VAE decode for the larger resolution

This requires the spatial upscaler model in models/latent_upscale_models/: ltx-2.3-spatial-upscaler-x2-1.1.safetensors (LTX-2.3) or ltx-2-spatial-upscaler-x2-1.0.safetensors (LTX-2).

Using alternate / GGUF base models (incl. the "sulphur" model)

You can swap the LTX UNet for any LTX-2.3-compatible base model. The most-asked-about one is Sulphur 2 (the user's "sulphur2Base_dev.safetensors"; see name note below).

What Sulphur 2 actually is (verified June 2026)
  • It exists and is real. Sulphur 2 is an uncensored, realism-leaning finetune/derivative of LTX-2.3 (22B DiT), marketed as a drop-in replacement inside existing LTX-2.3 ComfyUI graphs (T2V + I2V + the other 2.3 formats). It is NOT its own architecture and is not LTX-2 (19B) compatible. It targets the LTX-2.3 stack (2.3 VAE + Gemma 3 text encoder + 2.3 text projection).
  • Filename caveat: there is no file literally named sulphur2Base_dev.safetensors. The real base checkpoints are sulphur_dev_bf16.safetensors (~46 GB) and sulphur_dev_fp8mixed.safetensors (~29 GB). There is also a distilled variant (sulphur_distil_bf16.safetensors) and a LoRA (sulphur_lora_rank_768.safetensors). Treat "sulphur2Base_dev" as the user's shorthand for the Sulphur 2 base dev checkpoint.
  • GGUF version: confirmed. vantagewithai/Sulphur-2-Base-GGUF hosts sulphur_dev-<quant>.gguf for Q3_K_S/M, Q4_0/1/K_S/K_M, Q5_0/1/K_S/K_M, Q6_K, Q8_0 (~10 to 23 GB). There is also a Civitai/Sulphur-2-distilled-fp8 and Civitai listings ("Sulphur 2 Base", "Rebels Sulphur 2 GGUF").
  • Hosting: HF SulphurAI/Sulphur-2-base (safetensors + a bundled Qwen-based prompt-enhancer GGUF), HF vantagewithai/Sulphur-2-Base-GGUF (the GGUF quants), and Civitai mirrors. Uncensored open weights are in scope to document. Nothing here is fabricated, but verify the exact repo/license yourself before downloading.
How to load it (it slots straight into the LTX-2.3 GGUF workflow above)

The GGUF quant is a different UNet and nothing more. Load it with the same UnetLoaderGGUF node and keep the rest of the 2.3 graph identical:

  1. Put sulphur_dev-Q8_0.gguf (or your chosen quant) in models/unet/.
  2. In the LTX-2.3 GGUF workflow above, change node "1":
    json
    "1": { "class_type": "UnetLoaderGGUF", "inputs": { "unet_name": "sulphur_dev-Q8_0.gguf" }}
  3. Keep the same LTX-2.3 companions: VAELoader → LTX23_video_vae_bf16.safetensors, CLIPLoader (type=ltxv) → gemma_3_12B_it_fp4_mixed.safetensors, plus ltx-2.3_text_projection_bf16.safetensors. These must match the LTX-2.3 architecture. Do not pair it with LTX-2 (19B) VAE/encoder.
  4. For the bf16/fp8 safetensors (non-GGUF) variants, load with the LTX checkpoint/diffusion-model loader the workflow uses for the safetensors path (Lightricks recommends the native LTX Video nodes documented at docs.ltx.video, not the auto-generated Diffusers snippet) rather than UnetLoaderGGUF.
  5. Obey the same constraints as any LTX-2.3 gen: frame count 8n+1, resolution multiples of 32, LTXVConditioning frame_rate, dev model ~20+ steps / distilled ~8 steps.
General rule for ANY alternate LTX base model

To verify a third-party model is usable before wiring it up:

  • Confirm the architecture/version it was trained on (LTX-2 19B vs LTX-2.3 22B). Mixing a 2.3 UNet with a 2.0 VAE/encoder will fail or produce garbage.
  • For GGUF: requires the ComfyUI-GGUF custom node (installed by the scripts), file in models/unet/, loaded via UnetLoaderGGUF. Match the correct VAE + text encoder + text projection for that LTX version.
  • For safetensors finetunes: load like the matching official checkpoint, keep the official VAE/encoder of the same version.
  • If you only have a LoRA (e.g. sulphur_lora_rank_768.safetensors), apply it to the matching base UNet with LoraLoaderModelOnly instead of swapping the whole model.

Troubleshooting

LTXVideo "kornia" import error (pad ImportError)

Symptom: ComfyUI-LTXVideo fails to load with an ImportError from kornia.geometry.transform.pyramid because pad can no longer be imported. This happens with kornia 0.8.3+, which stopped exporting pad from that module.

What the fix does (FIX-LTXVIDEO-KORNIA.bat, run from the ComfyUI_windows_portable folder): it patches ComfyUI/custom_nodes/ComfyUI-LTXVideo/pyramid_blending.py:

  1. Backs the file up to pyramid_blending.py.bak_kornia_fix.
  2. Removes the broken pad, line from the from kornia.geometry.transform.pyramid import ( ... ) block.
  3. Inserts a compatibility shim right after import torch.nn.functional as F:
    python
    # Compatibility fix for Kornia 0.8.3+ where pad is no longer exported here
    pad = F.pad
  4. Verifies pad = F.pad is present and the broken import is gone.

Manual equivalent if you don't run the .bat: edit pyramid_blending.py to delete pad, from the kornia import list and add pad = F.pad after the import torch.nn.functional as F line, then restart ComfyUI. (Alternatively, pin kornia to a pre-0.8.3 release, but the patch is the lighter-touch fix and is what the install set ships.)

LTXVideo version / workflow mismatch

The RunPod installer pins ComfyUI-LTXVideo to commit cd5d371518afb07d6b3641be8012f644f25269fc for workflow compatibility. If 2.3 workflows error on the latest LTXVideo, check out that commit. Torch is pinned to 2.4.0 + cu121; do not let a node's requirements.txt upgrade torch (the installers sanitize requirements to prevent this).

Sources

  • Official: none found.
  • Empirical: sampler values, wiring, and prompt notes from working graphs in packs/ and observed renders; not a vendor prompting guide.

© artokun, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. 1 hidden character (zero-width or bidirectional) removed. Raw file

Files

SKILL.md and 1 other file (references) in plugin/skills/ltxv2-video of artokun/comfyui-mcp.

  • SKILL.md
  • references/workflows.md

Open the folder on GitHubat commit 6ad6fc0

Compare with similar skills

Ltxv2 Video next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Ltxv2 Video compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Ltxv2 Video this skillartokun/comfyui-mcp803—~6.8kAutomated safety check: WarnMIT
Nsfw VideoLeoYeAI/openclaw-master-skills2.2k—~4.6kAutomated safety check: PassMIT
Gemini Live APIgoogle/skills21k—~2.5kAutomated safety check: NotesApache-2.0
Tao Finetune Cosmos EmbedNVIDIA/skills3.6k—~3.5kAutomated safety check: NotesApache-2.0
Gemma Trainergoogle-gemma/gemma-skills1k—~1.9kAutomated safety check: PassApache-2.0
Hugging Face LLM Trainerhuggingface/skills11k1 repos~7.2kAutomated safety check: PassApache-2.0

Similar skills

  • Nsfw Video

    LeoYeAI/openclaw-master-skills

    Generate AI videos for mature creative projects using Wan 2.2 Spicy (LoRA-tuned for NSFW, top recommended), Wan 2.6, Seedance 1.5, Vidu Q3-Pro, and other models with relaxed content policies via…

    2.2k GitHub stars~4.6k tokensUpdated 2 mo ago
    Media & CreativeAuto-check passed
  • Gemini Live API

    google/skills

    Official

    Generates a Gemini LiveAPI client service class in the user's chosen programming language.

    21k GitHub stars~2.5k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check: notes
  • Official

    Cosmos-Embed1 video-text embedding for text-to-video retrieval, video-to-video search, semantic deduplication, and fine-tuning.

    3.6k GitHub stars~3.5k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check: notes
  • Gemma Trainer

    google-gemma/gemma-skills

    Trigger this skill when the user wants to train, fine-tune, or adapt Gemma models (e.g.

    1k GitHub stars~1.9k tokensUpdated 4 days ago
    AI & LLM EngineeringAuto-check passed
  • Hugging Face LLM Trainer

    huggingface/skills

    Official

    Trains or fine-tunes language and vision models with TRL or Unsloth on Hugging Face Jobs cloud GPUs, then converts the results to GGUF.

    11k GitHub starsUsed in 1 repo~7.2k tokens
    AI & LLM EngineeringAuto-check passed
  • Huggingface LLM Trainer

    waybarrios/opencode-power-pack

    Train or fine-tune language models with TRL or Unsloth on Hugging Face Jobs, including SFT, DPO, GRPO, reward models, and GGUF conversion.

    534 GitHub stars~3k tokensUpdated 5 days ago
    AI & LLM EngineeringAuto-check passed

More from artokun/comfyui-mcp

All 42 skills in this repo
  • AI Toolkit Trainer

    artokun/comfyui-mcp

    Train custom LoRAs with ostris AI-Toolkit. An agent skill from artokun/comfyui-mcp.

    803 GitHub stars~2.7k tokensUpdated 6 days ago
    Auto-check passed
  • Anima Base

    artokun/comfyui-mcp

    Anime/illustration text-to-image (ANIMA 1.0, ~2B Cosmos DiT).

    803 GitHub stars~4k tokensUpdated 6 days ago
    Auto-check passed
  • Civitai

    artokun/comfyui-mcp

    Discover Civitai models with the BUILT-IN downloadmodel action:"searchcivitai" and install/generate them locally.

    803 GitHub stars~1.1k tokensUpdated 6 days ago
    Auto-check passed
  • Color Correction

    artokun/comfyui-mcp

    Diagnose and fix video/image color OBJECTIVELY with the getimage (action:"analyzecolor") tool (scopes/stats such as black/white points, contrast, saturation, clipping, cast) instead of eyeballing a…

    803 GitHub stars~2.4k tokensUpdated 6 days ago
    Auto-check passed
  • Comfyui Frontend Extensions

    artokun/comfyui-mcp

    Authoring ComfyUI v2 frontend extensions with @comfyorg/extension-api, covering defineNode/defineExtension/defineWidget, shell UI (sidebar tabs, commands, hotkeys), typed events, and handles.

    803 GitHub stars~5.4k tokensUpdated 6 days ago
    Auto-check passed
  • Comfyui Launch Flags

    artokun/comfyui-mcp

    Pick the right ComfyUI startup flags for VRAM, attention, caching, and speed.

    803 GitHub stars~3.1k tokensUpdated 6 days ago
    Auto-check passed

Works with

Questions about Ltxv2 Video

What does Ltxv2 Video do?

Build Lightricks LTX-2 / LTX-2.3 video workflows covering text-to-video, image-to-video, GGUF and bundled checkpoints, distilled model, camera control LoRAs, synchronized audio, two-stage upscaling…. Ltxv2 Video is an agent skill from artokun/comfyui-mcp.

When should I use Ltxv2 Video?

Ltxv2 Video fits situations like: tasks that involve AI video generation; tasks that involve Fine-tuning.

How do I install Ltxv2 Video in Claude Code?

Run `npx skills add artokun/comfyui-mcp --skill ltxv2-video -a claude-code`. Or copy the skill folder (plugin/skills/ltxv2-video in artokun/comfyui-mcp) into .claude/skills/ltxv2-video in your project. Claude Code loads it when a task matches its description.

How do I install Ltxv2 Video in Codex?

Run `npx skills add artokun/comfyui-mcp --skill ltxv2-video -a codex`. Or copy the skill folder (plugin/skills/ltxv2-video in artokun/comfyui-mcp) into .agents/skills/ltxv2-video in your project. Codex loads it when a task matches its description.

Can I use Ltxv2 Video in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add artokun/comfyui-mcp --skill ltxv2-video -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ltxv2-video, .gemini/skills/ltxv2-video, .github/skills/ltxv2-video and .opencode/skills/ltxv2-video in your project.

What does Ltxv2 Video need to run?

Going by SKILL.md and its folder, Ltxv2 Video needs the command-line tools its instructions call (python and ffmpeg). Our summary lists: Python 3.

Does Ltxv2 Video access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Ltxv2 Video safe to install?

Our automated static check of SKILL.md flagged 1 warning(s): contains zero-width characters. Read the flagged lines before installing; the check is not a guarantee either way.

What licence does Ltxv2 Video use?

Ltxv2 Video is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Ltxv2 Video use?

About 6.8k tokens (SKILL.md is roughly 27k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.1k tokens, read only when the agent opens those files.

What are the alternatives to Ltxv2 Video?

Skills that share tags, products or a category with Ltxv2 Video: Nsfw Video (LeoYeAI/openclaw-master-skills, 2.2k stars), Gemini Live API (google/skills, 21k stars), Tao Finetune Cosmos Embed (NVIDIA/skills, 3.6k stars) and Gemma Trainer (google-gemma/gemma-skills, 1k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Ltxv2 Video?

artokun (a GitHub user) maintains it in artokun/comfyui-mcp, which has 803 GitHub stars. The repository holds 42 skills in this directory. The repository was last updated on October 5, 2026.

Source: artokun/comfyui-mcp on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.