Agent skill

Run Longcat Avatar Video

by drawthingsai in drawthingsai/draw-things-community

Generate, benchmark, validate, and troubleshoot LongCat-Video-Avatar 1.5 videos with draw-things-cli.

GPL-3.0Auto-check passedAI & LLM Engineering

Install Run Longcat Avatar Video

skills CLI
$ npx skills add drawthingsai/draw-things-community --skill run-longcat-avatar-video -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install drawthingsai/draw-things-community run-longcat-avatar-video --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/drawthingsai/draw-things-community.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/run-longcat-avatar-video .claude/skills/run-longcat-avatar-video && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
run-longcat-avatar-video
GitHub stars
584
Token cost
~2.8k tokens
SKILL.md length
1,106 words
Files
2
Skills in repo
8
Repo updated
First seen
Licence
GPL-3.0

At a glance

Generate, benchmark, validate, and troubleshoot LongCat-Video-Avatar 1.5 videos with draw-things-cli.

  • Codex needs to prepare LongCat q8p
  • SKILL.md covers Build And Inspect The CLI, Prepare Models, Audio API Boundary and Validate Inputs, plus 5 more sections
  • Calls bazel and xcrun
  • I8x checkpoints

What it does

Run Longcat Avatar Video is an agent skill from drawthingsai/draw-things-community. Generate, benchmark, validate, and troubleshoot LongCat-Video-Avatar 1.5 videos with draw-things-cli. Use when Codex needs to prepare LongCat q8p or i8x checkpoints, drive an avatar from a reference image and audio file, generate a fixed 4k+1 frame clip, run local segmented AVC for long audio with the canonical 93/13 configuration, compare i8x against q8p on Apple silicon, estimate full-video runtime, or verify LongCat MP4 frame count, duration, codec, and audio muxing.

Its SKILL.md is about 2.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files (for example `agents/openai.yaml`).

It sits in AI & LLM Engineering. The repository describes itself as: The community repository for the Draw Things app. The licence is GPL-3.0.

When your agent uses it

  • Codex needs to prepare LongCat q8p
  • I8x checkpoints
  • Drive an avatar from a reference image and audio file
  • Generate a fixed 4k+1 frame clip

Example prompts

  • “/run-longcat-avatar-video”

What it can do on your machine

Read from SKILL.md and the folder at commit 4357b8d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • bazel
    • xcrun

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Run Longcat Avatar Video loads about 2.8k tokens when it runs. Until then it costs about 125 tokens; SKILL.md has 1,106 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~125
When it runs · the whole SKILL.md, loaded when a task matches
~2.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from drawthingsai/draw-things-community at commit 4357b8d, republished under its GPL-3.0 licence (© drawthingsai). 1,106 words, ~2,839 tokens.

Download SKILL.mdSave it as .claude/skills/run-longcat-avatar-video/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
run-longcat-avatar-video
description
Generate, benchmark, validate, and troubleshoot LongCat-Video-Avatar 1.5 videos with draw-things-cli. Use when Codex needs to prepare LongCat q8p or i8x checkpoints, drive an avatar from a reference image and audio file, generate a fixed 4k+1 frame clip, run local segmented AVC for long audio with the canonical 93/13 configuration, compare i8x against q8p on Apple silicon, estimate full-video runtime, or verify LongCat MP4 frame count, duration, codec, and audio muxing.

Run LongCat Avatar Video

Use the local draw-things-cli generate command for LongCat-Video-Avatar 1.5. Keep i8x and q8p comparisons identical except for the model checkpoint and output path.

Build And Inspect The CLI

Work from the Draw Things repository root. Build once, then reuse the same optimized binary for all runs in one comparison:

sh
bazel build --compilation_mode=opt //Apps:DrawThingsCLI
CLI=bazel-bin/Apps/DrawThingsCLI
"$CLI" generate --help

Prefer the app model directory on macOS unless the user provides another one:

sh
MODELS_DIR="${DRAWTHINGS_MODELS_DIR:-$HOME/Library/Containers/com.liuliu.draw-things/Data/Documents/Models}"

Do not switch builds, power modes, resolutions, segment sizes, or preview settings during a benchmark.

Prepare Models

Choose one of these DiT checkpoints:

  • longcat_video_avatar_1.5_dmd_i8x.ckpt: 8-bit S model; prefer it for throughput on supported Apple silicon.
  • longcat_video_avatar_1.5_dmd_q8p.ckpt: q8p baseline and fallback; use it for matched quality and performance comparisons.

Ensure the selected model and its registered dependencies:

sh
"$CLI" models ensure \
  --models-dir "$MODELS_DIR" \
  --model longcat_video_avatar_1.5_dmd_i8x.ckpt

LongCat also needs these files in MODELS_DIR:

text
umt5_xxl_encoder_q8p.ckpt
wan_v2.1_video_vae_f16.ckpt
whisper_large_v3_f16.ckpt

The current model dependency list covers UMT5 and the Wan VAE. Verify Whisper separately because --audio-encoder-file defaults to whisper_large_v3_f16.ckpt, but it is not a registered LongCat model dependency:

sh
test -f "$MODELS_DIR/whisper_large_v3_f16.ckpt"

Pass --audio-encoder-file NAME.ckpt only when using a different compatible Whisper checkpoint.

Audio API Boundary

Keep model audio separate from ControlNet hints. The reusable input and model-specific encoders live in Libraries/AudioConverter; LocalImageGenerator only consumes finished conditioning. LongCat audio follows this internal path:

text
AudioInput -> LongCatAudioConditioningEncoder -> LongCatAudioFeatures
           -> LongCatAudioConditioning -> AudioConditioning.longCat

AudioInput owns decoded PCM and builds the waveform used for output audio muxing. Run the Whisper encoder once to produce LongCatAudioFeatures; derive one LongCatAudioConditioning for a normal generation or one per AVC segment. AudioConditioning is the model-dispatch boundary where a future LTX audio-conditioning case can be added. Do not represent model audio as ControlHintType.audio or route it through ControlModel.

Validate Inputs

Use both a reference image and driving audio. The image is aspect-scaled and center-cropped to the requested output size. Width and height must be multiples of 64.

Confirm that AVFoundation can decode the audio before starting a long run:

sh
afinfo "$AUDIO"

Reject an input that reports zero packets or zero duration even if its file size is nonzero. Use a valid CAF, MP3, M4A, or WAV container that afinfo and AVAudioFile can decode; do not fix a bad container by renaming its extension.

The DMD checkpoint is step-distilled. Use these baseline values unless the task explicitly changes them:

text
steps = 8
cfg = 1
shift = 7
fps = 25

LongCat temporal counts must be 4k + 1. Canonical values are 93 generated frames and 13 AVC condition frames. Values such as 77 and 109 are valid experiments, but do not mix them into a canonical 93/13 comparison.

Generate An AVC Video

Use AVC for audio longer than one generated clip. AVC is local-only and currently supports only LongCat-Video-Avatar 1.5. Its output duration follows the audio; do not pass --frames with --avc.

Use 93/13 explicitly even though they are the current defaults:

sh
IMAGE=/path/to/reference.png
AUDIO=/path/to/driving-audio.caf
OUTPUT=/path/to/longcat_avc_i8x.mp4
PROMPT='A person speaks naturally to the camera with stable posture and synchronized mouth motion.'

"$CLI" generate --avc \
  --models-dir "$MODELS_DIR" \
  --model longcat_video_avatar_1.5_dmd_i8x.ckpt \
  --image "$IMAGE" \
  --audio "$AUDIO" \
  --prompt "$PROMPT" \
  --steps 8 --cfg 1 \
  --segment-frames 93 --cond-frames 13 \
  --width 448 --height 320 \
  --seed 42 \
  --config-json '{"shift":7}' \
  --no-download-missing \
  --disable-preview \
  --video-format h264 \
  --output "$OUTPUT"

With 93/13, each later segment contributes 80 new frames. For an audio target of T frames:

text
stride = 93 - 13 = 80
segments = T <= 93 ? 1 : ceil((T - 93) / 80) + 1
sampling steps = segments * 8

The CLI computes Whisper features for the generated span, reuses the last 13 decoded frames as the next segment's clean condition, drops overlap frames, trims to the audio target, and muxes the input speech into the output container.

For the full approximately 82-second reference workload, keep 93/13 and change only the audio, resolution, and output:

sh
"$CLI" generate --avc \
  --models-dir "$MODELS_DIR" \
  --model longcat_video_avatar_1.5_dmd_i8x.ckpt \
  --image "$IMAGE" --audio /path/to/man.mp3 \
  --prompt "$PROMPT" \
  --steps 8 --cfg 1 \
  --segment-frames 93 --cond-frames 13 \
  --width 832 --height 512 \
  --seed 42 --config-json '{"shift":7}' \
  --no-download-missing --disable-preview \
  --video-format h264 \
  --output /path/to/man_832x512_i8x.mp4

Generate One 93-Frame Clip Without AVC

Omit --avc, --segment-frames, and --cond-frames. Set --frames 93 explicitly:

sh
"$CLI" generate \
  --models-dir "$MODELS_DIR" \
  --model longcat_video_avatar_1.5_dmd_i8x.ckpt \
  --image "$IMAGE" \
  --audio "$AUDIO" \
  --prompt "$PROMPT" \
  --steps 8 --cfg 1 \
  --frames 93 \
  --width 448 --height 320 \
  --seed 42 \
  --config-json '{"shift":7}' \
  --no-download-missing \
  --disable-preview \
  --video-format h264 \
  --output /path/to/longcat_93f_i8x.mp4

At 25 fps, 93 frames produce 3.72 seconds. The exported audio is trimmed or padded to the same duration.

Compare i8x And q8p

Run comparisons serially on the same machine. Keep all of these identical:

  • CLI binary and build mode
  • image, audio, and prompt
  • width, height, and frame or AVC settings
  • steps, CFG, shift, and seed
  • preview and video-format flags
  • system power mode and competing workload

Change only:

text
--model longcat_video_avatar_1.5_dmd_i8x.ckpt
--model longcat_video_avatar_1.5_dmd_q8p.ckpt

Wrap each command with /usr/bin/time -p. Record both the CLI's generation/sampling summary and the external real time:

sh
/usr/bin/time -p "$CLI" generate ...

Do not call model-loading, sampling-step, generation, and wall-clock speedups interchangeable. Report each metric by name.

Show full SKILL.md (484 more words)Show less
Measured M5 Max Baseline

The following results were measured on 2026-07-13 on an Apple M5 Max with 48 GB memory. Both models used 448x320, 8 steps, CFG 1, shift 7, seed 42, H.264, disabled preview, and the same image, prompt, and valid 7.988-second audio. Use these as reference data, not a universal performance guarantee.

WorkloadModelSampling stepsGenerationMedian / stepWall time
AVC 93/13, 200 output framesi8x24368.45 s10.69 s421.77 s
AVC 93/13, 200 output framesq8p24497.04 s14.92 s550.07 s
No AVC, 93 output framesi8x886.59 s7.98 s134.70 s
No AVC, 93 output framesq8p8136.40 s13.23 s184.26 s

Observed speedups from that matched run:

  • AVC wall time: i8x was 1.30x faster; median sampling step was 1.40x faster.
  • Non-AVC wall time: i8x was 1.37x faster; median sampling step was 1.66x faster.

An earlier 832x512, approximately 82-second, AVC 93/13 workload was roughly 9 hours with q8p and 3 hours or more with i8x. Treat that as a historical high-resolution observation, not a canonical 3x claim, because it was not captured with the same benchmark ledger as the table above. Re-run both checkpoints with matched commands before publishing a 3x result.

Verify Outputs

Check existence, size, codecs, dimensions, and duration:

sh
ls -lh "$OUTPUT"
mdls \
  -name kMDItemCodecs \
  -name kMDItemDurationSeconds \
  -name kMDItemPixelWidth \
  -name kMDItemPixelHeight \
  "$OUTPUT"

On macOS, count actual video samples with AVFoundation when ffprobe is unavailable:

sh
VIDEO="$OUTPUT" xcrun swift -e '
import AVFoundation
import Foundation
let asset = AVURLAsset(url: URL(fileURLWithPath: ProcessInfo.processInfo.environment["VIDEO"]!))
let track = asset.tracks(withMediaType: .video).first!
let reader = try AVAssetReader(asset: asset)
let output = AVAssetReaderTrackOutput(track: track, outputSettings: nil)
reader.add(output)
reader.startReading()
var samples = 0
while let buffer = output.copyNextSampleBuffer() {
  samples += CMSampleBufferGetNumSamples(buffer)
}
print("video_samples=\(samples) status=\(reader.status.rawValue)")
'

Expect 93 samples for --frames 93. For AVC, expect the audio-derived target; a valid 7.988-second input at 25 fps produces 200 samples and an 8.00-second MP4.

Inspect visual continuity around each AVC boundary. With 93/13, the first output boundary is near frame 93, and later boundaries advance by 80 frames. Compare i8x and q8p boundary frames before making a quality claim.

Troubleshoot

  • --frames must be 4k + 1: use 93 for the canonical single-shot run.
  • --segment-frames or --cond-frames validation fails: use 93/13; both values must be 4k + 1, and segment frames must exceed condition frames.
  • --frames cannot be used with --avc: remove --frames; AVC duration follows audio.
  • --avc currently supports only local generation: remove --remote and --cloud-compute.
  • Missing Whisper error: place the selected audio encoder in MODELS_DIR or pass --audio-encoder-file with its filename.
  • Audio reports zero duration or zero packets: replace or properly transcode the source file before generation.
  • High-resolution run exhausts memory: validate first at 448x320, close competing GPU workloads, then retry 832x512. Do not change segment size during an i8x/q8p comparison.
  • Unexpectedly weak speedup: compare sampling metrics separately from Whisper, model loading, VAE, and video encoding. Fixed costs dominate short low-resolution clips.
  • Segment boundary jump: confirm both runs use the same 93/13, seed, audio, and reference image; inspect frames around each 80-frame stride before changing the continuation policy.
  • --zero-audio-features is a hidden pipeline diagnostic. Never use it for a quality result.

Keep the current stable LongCat behavior unless debugging model internals. In particular, do not change masked-reference attention or continuation policy merely to improve a benchmark number.

© drawthingsai, GPL-3.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in .agents/skills/run-longcat-avatar-video of drawthingsai/draw-things-community.

  • SKILL.md
  • agents/openai.yaml

Open the folder on GitHubat commit 4357b8d

Compare with similar skills

Run Longcat Avatar Video next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Run Longcat Avatar Video compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Run Longcat Avatar Video this skilldrawthingsai/draw-things-community584—~2.8kAutomated safety check: PassGPL-3.0
Agent BuildershareAI-lab/learn-claude-code78k4 repos~1.2kAutomated safety check: PassMIT
Add Uint Supportpytorch/pytorch104k2 repos~2.3kAutomated safety check: PassCustom licence
LLM Benchmarking with lm-evaluation-harnessOrchestra-Research/AI-Research-SKILLs13k8 repos~3kAutomated safety check: PassMIT
Segment Anything Model GuideOrchestra-Research/AI-Research-SKILLs13k8 repos~3.3kAutomated safety check: PassMIT
1passwordtrpc-group/trpc-agent-go1.9k14 repos~656Automated safety check: PassApache-2.0

Similar skills

  • Agent Builder

    shareAI-lab/learn-claude-code

    Design and build AI agents for any domain. An agent skill from shareAI-lab/learn-claude-code.

    78k GitHub starsUsed in 4 repos~1.2k tokens
    AI & LLM EngineeringAuto-check passed
  • Add Uint Support

    pytorch/pytorch

    Add unsigned integer (uint) type support to PyTorch operators by updating ATDISPATCH macros.

    104k GitHub starsUsed in 2 repos~2.3k tokens
    AI & LLM EngineeringAuto-check passed
  • LLM Benchmarking with lm-evaluation-harness

    Orchestra-Research/AI-Research-SKILLs

    Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.

    13k GitHub starsUsed in 8 repos~3k tokens
    AI & LLM EngineeringAuto-check passed
  • Segment Anything Model Guide

    Orchestra-Research/AI-Research-SKILLs

    Guide to using Meta's Segment Anything Model for zero-shot image segmentation with point, box or mask prompts, or automatic mask generation.

    13k GitHub starsUsed in 8 repos~3.3k tokens
    AI & LLM EngineeringAuto-check passed
  • 1password

    trpc-group/trpc-agent-go

    Set up and use 1Password CLI (op). An agent skill from trpc-group/trpc-agent-go.

    1.9k GitHub starsUsed in 14 repos~656 tokens
    AI & LLM EngineeringAuto-check passed
  • Planning With Files

    jarrodwatts/claude-code-config

    Transforms workflow to use Manus-style persistent markdown files for planning, progress tracking, and knowledge storage.

    1.1k GitHub starsUsed in 5 repos~967 tokens
    AI & LLM EngineeringAuto-check passed

More from drawthingsai/draw-things-community

All 8 skills in this repo
  • Setup Cpu Proxy Server

    drawthingsai/draw-things-community

    Set up and verify a new Draw Things CPU proxy and Envoy server using the scripts in Scripts/ServerManagement/CPUScript.

    584 GitHub stars~3.8k tokensUpdated yesterday
    Auto-check passed
  • Use Drawthings CLI

    drawthingsai/draw-things-community

    Use and troubleshoot an already installed Homebrew draw-things-cli for model discovery, authentication, local, cloud, or remote image generation, basic image-to-image and video generation, output…

    584 GitHub stars~3k tokensUpdated yesterday
    Auto-check passed
  • Init GPU Server

    drawthingsai/draw-things-community

    Initialize a Draw Things GPU server with GPUScript, including script sync, Docker/CUDA/NVIDIA runtime setup, 7T data disk mounting, mergerfs, and end-to-end GPU verification.

    584 GitHub stars~2.2k tokensUpdated yesterday
    Auto-check passed
  • New Model Integration

    drawthingsai/draw-things-community

    Add a new image or video generative model to the Draw Things app / CLI with a compile-first, end-to-end workflow across SwiftDiffusion, tokenizer plumbing, text encoder, fixed encoder, UNet / DiT…

    584 GitHub stars~4.8k tokensUpdated yesterday
    Auto-check passed
  • Train Lora

    drawthingsai/draw-things-community

    Validate Draw Things LoRA training end to end with draw-things-cli, including tiny-dataset training, loss and scaler checks, checkpoint sanity, and base-versus-LoRA generation comparison.

    584 GitHub stars~2k tokensUpdated yesterday
    Auto-check passed
  • Trainer Integration

    drawthingsai/draw-things-community

    Add or tighten Draw Things LoRA trainer support for generative models available in the Draw Things app / CLI, covering LoRA builders, trainer dispatch, tokenizer and fixed-encoder wiring, checkpoint…

    584 GitHub stars~2.6k tokensUpdated yesterday
    Auto-check passed

Questions about Run Longcat Avatar Video

What does Run Longcat Avatar Video do?

Generate, benchmark, validate, and troubleshoot LongCat-Video-Avatar 1.5 videos with draw-things-cli. Run Longcat Avatar Video is an agent skill from drawthingsai/draw-things-community.5 videos with draw-things-cli.

When should I use Run Longcat Avatar Video?

Run Longcat Avatar Video fits situations like: Codex needs to prepare LongCat q8p; I8x checkpoints; drive an avatar from a reference image and audio file; generate a fixed 4k+1 frame clip.

How do I install Run Longcat Avatar Video in Claude Code?

Run `npx skills add drawthingsai/draw-things-community --skill run-longcat-avatar-video -a claude-code`. Or copy the skill folder (.agents/skills/run-longcat-avatar-video in drawthingsai/draw-things-community) into .claude/skills/run-longcat-avatar-video in your project. Claude Code loads it when a task matches its description.

How do I install Run Longcat Avatar Video in Codex?

Run `npx skills add drawthingsai/draw-things-community --skill run-longcat-avatar-video -a codex`. Or copy the skill folder (.agents/skills/run-longcat-avatar-video in drawthingsai/draw-things-community) into .agents/skills/run-longcat-avatar-video in your project. Codex loads it when a task matches its description.

Can I use Run Longcat Avatar Video in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add drawthingsai/draw-things-community --skill run-longcat-avatar-video -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/run-longcat-avatar-video, .gemini/skills/run-longcat-avatar-video, .github/skills/run-longcat-avatar-video and .opencode/skills/run-longcat-avatar-video in your project.

What does Run Longcat Avatar Video need to run?

Going by SKILL.md and its folder, Run Longcat Avatar Video needs the command-line tools its instructions call (bazel and xcrun).

Does Run Longcat Avatar Video access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Run Longcat Avatar Video safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Run Longcat Avatar Video use?

Run Longcat Avatar Video is published under the GPL-3.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Run Longcat Avatar Video use?

About 2.8k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Run Longcat Avatar Video?

Skills that share tags, products or a category with Run Longcat Avatar Video: Agent Builder (shareAI-lab/learn-claude-code, 78k stars), Add Uint Support (pytorch/pytorch, 104k stars), LLM Benchmarking with lm-evaluation-harness (Orchestra-Research/AI-Research-SKILLs, 13k stars) and Segment Anything Model Guide (Orchestra-Research/AI-Research-SKILLs, 13k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Run Longcat Avatar Video?

drawthingsai (a GitHub organization) maintains it in drawthingsai/draw-things-community, which has 584 GitHub stars. The repository holds 8 skills in this directory. The repository was last updated on October 9, 2026.

Source: drawthingsai/draw-things-community on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.