Agent skill

Wan Multitalk

by artokun in artokun/comfyui-mcp

Build WAN MultiTalk audio-driven talking-avatar / lip-sync video workflows.

MITAuto-check passedAI & LLM Engineering

Install Wan Multitalk

skills CLI
$ npx skills add artokun/comfyui-mcp --skill wan-multitalk -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install artokun/comfyui-mcp wan-multitalk --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/artokun/comfyui-mcp.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugin/skills/wan-multitalk .claude/skills/wan-multitalk && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
wan-multitalk
GitHub stars
795
Token cost
~1.3k tokens
SKILL.md length
465 words
Files
1
Skills in repo
42
Repo updated
First seen
Licence
MIT

At a glance

Build WAN MultiTalk audio-driven talking-avatar / lip-sync video workflows.

  • Tasks that involve Diffusion and image models
  • SKILL.md covers Overview, Pipeline (node graph), Models and Inputs & key parameters, plus 3 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Wan Multitalk is an agent skill from artokun/comfyui-mcp. Build WAN MultiTalk audio-driven talking-avatar / lip-sync video workflows. MeiGen-AI MultiTalk on WAN 2.1 14B I2V via kijai WanVideoWrapper (portrait + audio → lip-synced video)

Its SKILL.md is about 1.3k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering Diffusion and image models. The repository describes itself as: Local-first, agent-native control plane for ComfyUI — MCP server + sidebar agent that generates images, video & audio, authors and runs workflows, and edits your live graph in… The licence is MIT.

When your agent uses it

  • Tasks that involve Diffusion and image models

Example prompts

  • “/wan-multitalk”

What it can do on your machine

Read from SKILL.md and the folder at commit 6ad6fc0. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Wan Multitalk loads about 1.3k tokens when it runs. Until then it costs about 48 tokens; SKILL.md has 465 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~48
When it runs · the whole SKILL.md, loaded when a task matches
~1.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from artokun/comfyui-mcp at commit 6ad6fc0, republished under its MIT licence (© artokun). 465 words, ~1,275 tokens.

Download SKILL.mdSave it as .claude/skills/wan-multitalk/SKILL.md (or your agent's skills folder).
name
wan-multitalk
description
Build WAN MultiTalk audio-driven talking-avatar / lip-sync video workflows. MeiGen-AI MultiTalk on WAN 2.1 14B I2V via kijai WanVideoWrapper (portrait + audio → lip-synced video)
globs
**/*.json

WAN MultiTalk — Audio-Driven Talking Avatar

Overview

MultiTalk (MeiGen-AI) drives a still portrait's lip-sync and head motion from an audio track. It runs on WAN 2.1 14B Image-to-Video via kijai's ComfyUI-WanVideoWrapper. Wav2Vec speech embeddings condition the WAN sampler so the mouth and expression follow the speech, while the lightx2v step-distill LoRA keeps it to a few sampling steps.

Use it for talking heads, dubbing, and single-speaker avatar clips (~10s at 480p). It is distinct from wan-animate (pose/motion-driven character animation). This is audio → lip-sync, not reference-video motion transfer.

Pack: wan-multitalk (480p, ~10s). Higher-res/longer variants exist in the source bundle (720p, long-context) as VRAM/duration knobs on the same graph.

Pipeline (node graph)

LoadImage (portrait) ─┐
LoadAudio ─ AudioSeparation ─ AudioCrop ─ DownloadAndLoadWav2VecModel ─ MultiTalkWav2VecEmbeds ─┐
                                                                                                 ▼
WanVideoModelLoader (WAN 2.1 14B I2V GGUF) ─ MultiTalkModelLoader ─ WanVideoLoraSelect (lightx2v)
   + LoadWanVideoT5TextEncoder (umt5) + WanVideoTextEncode + WanVideoClipVisionEncode (clip_vision_h)
   + WanVideoVAELoader ──────────────────────────────────────────────────────────────────────────┘
                                                     ▼
                       WanVideoImageToVideoMultiTalk ─ WanVideoSampler ─ WanVideoDecode ─ VHS_VideoCombine

Key nodes (all kijai WanVideoWrapper unless noted):

  • DownloadAndLoadWav2VecModel. Auto-downloads the Wav2Vec speech model on first run (no manifest entry needed).
  • MultiTalkWav2VecEmbeds. Turns the (separated, cropped) speech into the embeddings that steer the mouth and expression.
  • MultiTalkModelLoader + WanVideoImageToVideoMultiTalk. The MultiTalk head on top of the WAN I2V model.
  • AudioSeparation and AudioCrop (audio-separation-nodes-comfyui). Isolate the voice from music/noise before embedding and trim the segment you want to animate.
  • ImageResizeKJv2 (KJNodes), VHS_VideoCombine (VideoHelperSuite). Resize and mux to mp4.

Models

FileLoaderFolder
Wan2.1_14b_Image_to_Video_480p_GGUF_Q8.ggufWanVideoModelLoaderdiffusion_models/
WanVideo_2_1_Multitalk_14B_fp8_e4m3fn.safetensorsMultiTalkModelLoaderdiffusion_models/
umt5_xxl_fp16.safetensorsLoadWanVideoT5TextEncodertext_encoders/
Wan2_1_VAE_bf16.safetensorsWanVideoVAELoadervae/
clip_vision_h.safetensorsCLIPVisionLoaderclip_vision/
Wan21_I2V_14B_lightx2v_cfg_step_distill_lora_rank64.safetensorsWanVideoLoraSelectloras/

Sources: kijai Kijai/WanVideo_comfy, MeiGen-AI MeiGen-AI/MeiGen-MultiTalk, GGUF city96/Wan2.1-I2V-14B-480P-gguf, and Comfy-Org's repackaged UMT5. See packs/wan-multitalk/manifest.yaml (some URLs are best-effort; verify per mirror). Wav2Vec auto-downloads. The bundled WanVideoWrapper loader rejects the scaled_fp8 UMT5 checkpoint; use the UMT5 fp16 file above, not generic t5xxl_fp16 weights.

Show full SKILL.md (218 more words)Show less

Inputs & key parameters

  • Portrait (LoadImage): front-facing, clear face, neutral-ish expression works best. Resized by ImageResizeKJv2 to the target (480p).
  • Audio (LoadAudio): the speech track. AudioSeparation isolates the voice; AudioCrop selects the segment (drives clip length).
  • Steps: low (the lightx2v distill LoRA is why; typically ~4 to 8). Raising steps rarely helps and costs time.
  • BlockSwap (WanVideoBlockSwap): trade VRAM for speed. Increase blocks swapped to CPU on lower-VRAM cards.

VRAM tiers (from the source bundle's variants)

TargetApprox VRAMLever
480p 10s~8–12 GBbase
480p low-VRAM~6–8.4 GBmore BlockSwap, GGUF quant, lower quality
720p 10s~11–16 GBhigher res

Pair with the VRAM launch-flags guidance (see troubleshooting): --use-sage-attention

  • appropriate --*vram mode; MultiTalk benefits from --reserve-vram headroom for the Wav2Vec + VAE round-trips.

Gotchas

  • Audio must be voice-isolated for good lip-sync. Skipping AudioSeparation on a music-heavy track makes the mouth chase the wrong signal.
  • One speaker. This graph is single-speaker; multi-speaker MultiTalk needs the multi-embed variant (not in this pack).
  • Wav2Vec first run downloads a model, so the first render is slower.
  • If lips look under-driven, check the MultiTalk embeds are actually wired into WanVideoImageToVideoMultiTalk (not bypassed), and that the audio isn't silent after AudioCrop.

Sources

  • Official: none found.
  • Empirical: sampler values, wiring, and prompt notes from working graphs in packs/ and observed renders; not a vendor prompting guide.

© artokun, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in plugin/skills/wan-multitalk of artokun/comfyui-mcp.

Open the folder on GitHubat commit 6ad6fc0

Compare with similar skills

Wan Multitalk next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Wan Multitalk compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Wan Multitalk this skillartokun/comfyui-mcp795—~1.3kAutomated safety check: PassMIT
Adapt New Diffusion Modelintel/auto-round1.6k—~2.8kAutomated safety check: PassApache-2.0
Add Pipelineverl-project/verl-omni1.2k—~1kAutomated safety check: PassApache-2.0
Comfyui AnimatoolShiroEirin/comfyui-good-anima478—~4.6kAutomated safety check: PassGPL-3.0
Stage1 Add VaeEnd2End-Diffusion/diffusion-bench105—~1.1kAutomated safety check: PassNone
Comfyui Agent Skill MieMieMieeeee/comfyui-agent-skill116—~3.9kAutomated safety check: PassApache-2.0

Similar skills

  • Official

    Adapt AutoRound to support a new diffusion model architecture (DiT, UNet, hybrid AR+DiT).

    1.6k GitHub stars~2.8k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Add Pipeline

    verl-project/verl-omni

    Router for adding a diffusion or omni pipeline to verl-omni.

    1.2k GitHub stars~1k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Comfyui Animatool

    ShiroEirin/comfyui-good-anima

    Route ALL Anima image generation: validate Danbooru hard anchors, form visual brief, assemble English prompts and args, then load comfyui-manager for workflow execution.

    478 GitHub stars~4.6k tokensUpdated 3 mo ago
    AI & LLM EngineeringAuto-check passed
  • Stage1 Add Vae

    End2End-Diffusion/diffusion-bench

    Add a new HuggingFace-supported VAE to the stage1 tokenizer pipeline.

    105 GitHub stars~1.1k tokensUpdated 3 mo ago
    AI & LLM EngineeringAuto-check passed
  • Comfyui Agent Skill Mie

    MieMieeeee/comfyui-agent-skill

    Agent skill for running registered ComfyUI workflows through a stable CLI, and for importing a user's own ComfyUI workflow into their private registry after review.

    116 GitHub stars~3.9k tokensUpdated 4 days ago
    AI & LLM EngineeringAuto-check passed
  • Importing Subgraphs

    Comfy-Org/workflow_templates

    Imports and registers subgraph blueprints into the ComfyUI workflowtemplates repository.

    1.3k GitHub stars~1.5k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed

More from artokun/comfyui-mcp

All 42 skills in this repo
  • AI Toolkit Trainer

    artokun/comfyui-mcp

    Train custom LoRAs with ostris AI-Toolkit. An agent skill from artokun/comfyui-mcp.

    795 GitHub stars~2.7k tokensUpdated 3 days ago
    Auto-check passed
  • Anima Base

    artokun/comfyui-mcp

    Anime/illustration text-to-image (ANIMA 1.0, ~2B Cosmos DiT).

    795 GitHub stars~4k tokensUpdated 3 days ago
    Auto-check passed
  • Civitai

    artokun/comfyui-mcp

    Discover Civitai models with the BUILT-IN downloadmodel action:"searchcivitai" and install/generate them locally.

    795 GitHub stars~1.1k tokensUpdated 3 days ago
    Auto-check passed
  • Color Correction

    artokun/comfyui-mcp

    Diagnose and fix video/image color OBJECTIVELY with the getimage (action:"analyzecolor") tool (scopes/stats such as black/white points, contrast, saturation, clipping, cast) instead of eyeballing a…

    795 GitHub stars~2.4k tokensUpdated 3 days ago
    Auto-check passed
  • Comfyui Frontend Extensions

    artokun/comfyui-mcp

    Authoring ComfyUI v2 frontend extensions with @comfyorg/extension-api, covering defineNode/defineExtension/defineWidget, shell UI (sidebar tabs, commands, hotkeys), typed events, and handles.

    795 GitHub stars~5.4k tokensUpdated 3 days ago
    Auto-check passed
  • Comfyui Launch Flags

    artokun/comfyui-mcp

    Pick the right ComfyUI startup flags for VRAM, attention, caching, and speed.

    795 GitHub stars~3.1k tokensUpdated 3 days ago
    Auto-check passed

Questions about Wan Multitalk

What does Wan Multitalk do?

Build WAN MultiTalk audio-driven talking-avatar / lip-sync video workflows. Wan Multitalk is an agent skill from artokun/comfyui-mcp. Build WAN MultiTalk audio-driven talking-avatar / lip-sync video workflows.

When should I use Wan Multitalk?

Wan Multitalk fits situations like: tasks that involve Diffusion and image models.

How do I install Wan Multitalk in Claude Code?

Run `npx skills add artokun/comfyui-mcp --skill wan-multitalk -a claude-code`. Or copy the skill folder (plugin/skills/wan-multitalk in artokun/comfyui-mcp) into .claude/skills/wan-multitalk in your project. Claude Code loads it when a task matches its description.

How do I install Wan Multitalk in Codex?

Run `npx skills add artokun/comfyui-mcp --skill wan-multitalk -a codex`. Or copy the skill folder (plugin/skills/wan-multitalk in artokun/comfyui-mcp) into .agents/skills/wan-multitalk in your project. Codex loads it when a task matches its description.

Can I use Wan Multitalk in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add artokun/comfyui-mcp --skill wan-multitalk -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/wan-multitalk, .gemini/skills/wan-multitalk, .github/skills/wan-multitalk and .opencode/skills/wan-multitalk in your project.

What does Wan Multitalk need to run?

SKILL.md names no scripts, command-line tools or credentials: Wan Multitalk is instructions for the agent only.

Does Wan Multitalk access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Wan Multitalk safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Wan Multitalk use?

Wan Multitalk is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Wan Multitalk use?

About 1.3k tokens (SKILL.md is roughly 5.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Wan Multitalk?

Skills that share tags, products or a category with Wan Multitalk: Adapt New Diffusion Model (intel/auto-round, 1.6k stars), Add Pipeline (verl-project/verl-omni, 1.2k stars), Comfyui Animatool (ShiroEirin/comfyui-good-anima, 478 stars) and Stage1 Add Vae (End2End-Diffusion/diffusion-bench, 105 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Wan Multitalk?

artokun (a GitHub user) maintains it in artokun/comfyui-mcp, which has 795 GitHub stars. The repository holds 42 skills in this directory. The repository was last updated on October 5, 2026.

Source: artokun/comfyui-mcp on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.