Agent skill

Wan Flf Video

by artokun in artokun/comfyui-mcp

Build WAN 2.2 First-Last-Frame video workflows. An agent skill from artokun/comfyui-mcp.

MITAuto-check passedAI & LLM Engineering

Install Wan Flf Video

skills CLI
$ npx skills add artokun/comfyui-mcp --skill wan-flf-video -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install artokun/comfyui-mcp wan-flf-video --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/artokun/comfyui-mcp.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugin/skills/wan-flf-video .claude/skills/wan-flf-video && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
wan-flf-video
GitHub stars
803
Token cost
~5.1k tokens
SKILL.md length
1,723 words
Files
2 (incl. references)
Skills in repo
42
Repo updated
First seen
Licence
MIT

At a glance

Build WAN 2.2 First-Last-Frame video workflows. An agent skill from artokun/comfyui-mcp.

  • Works in 2 steps: Native Dual Hi-Lo (Default):… → WanVideoWrapper:…
  • Tasks that involve Fine-tuning
  • SKILL.md covers Overview, CRITICAL: Dual Hi-Lo…, Models and ModelSamplingSD3 (REQUIRED), plus 13 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Wan Flf Video is an agent skill from artokun/comfyui-mcp. Build WAN 2.2 First-Last-Frame video workflows. Native dual hi-lo (required), and WanVideoWrapper VACE approaches

Its SKILL.md is about 5.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including reference files (for example `references/workflows.md`).

It sits in AI & LLM Engineering, covering Fine-tuning. It works with llama.cpp. The repository describes itself as: Local-first, agent-native control plane for ComfyUI — MCP server + sidebar agent that generates images, video & audio, authors and runs workflows, and edits your live graph in… The licence is MIT.

When your agent uses it

  • Tasks that involve Fine-tuning

Example prompts

  • “/wan-flf-video”

Workflow steps

2 steps, taken from the first numbered list in SKILL.md.

  1. Native Dual Hi-Lo (Default): WanFirstLastFrameToVideo + dual KSamplerAdvanced two-pass
  2. WanVideoWrapper: WanVideoVACEStartToEndFrame + WanVideoVACEEncode + WanVideoSampler (VACE, caching, context windows)

What it can do on your machine

Read from SKILL.md and the folder at commit 6ad6fc0. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are json).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • huggingface.co
    • civitai.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Wan Flf Video loads about 5.1k tokens when it runs, and up to ~6.6k if it reads all its reference files. Until then it costs about 32 tokens; SKILL.md has 1,723 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~32
When it runs · the whole SKILL.md, loaded when a task matches
~5.1k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~6.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from artokun/comfyui-mcp at commit 6ad6fc0, republished under its MIT licence (© artokun). 1,723 words, ~5,122 tokens.

Download SKILL.mdSave it as .claude/skills/wan-flf-video/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
wan-flf-video
description
Build WAN 2.2 First-Last-Frame video workflows. Native dual hi-lo (required), and WanVideoWrapper VACE approaches
globs
**/*.json

WAN 2.2 First-Last-Frame (FLF) Video Workflows

Overview

First-Last-Frame (FLF) video generation takes a start image and an end image and generates a smooth video transition between them. The WAN 2.2 I2V (Image-to-Video) 14B model is good at this.

CRITICAL: Dual Hi-Lo Architecture (REQUIRED)

WAN 2.2 I2V uses a split-noise architecture. Unlike WAN 2.1, the 2.2 model was trained with separate HighNoise and LowNoise components that handle different denoising ranges. You MUST use both models in a two-pass KSamplerAdvanced setup. Using a single model produces low-quality, broken output.

  • HighNoise model (pass 1, steps 0→N/2) establishes structure, motion, and composition
  • LowNoise model (pass 2, steps N/2→N) refines details and keeps fidelity to input frames
  • Both passes share the same conditioning from WanFirstLastFrameToVideo
  • Pass 1 returns noisy latent → Pass 2 continues from there

NEVER use a single KSampler with only one model for WAN 2.2 I2V.

Two native approaches are available:

  1. Native Dual Hi-Lo (Default): WanFirstLastFrameToVideo + dual KSamplerAdvanced two-pass
  2. WanVideoWrapper: WanVideoVACEStartToEndFrame + WanVideoVACEEncode + WanVideoSampler (VACE, caching, context windows)

Models

UNET Pairs (Always load BOTH Hi and Lo)

Remix NSFW (Recommended, built-in lightning, fp16):

ModelLoaderNotes
Wan2.2_Remix_NSFW_i2v_14b_high_lighting_fp16_v2.1.safetensorsUNETLoaderHighNoise, built-in lightning acceleration
Wan2.2_Remix_NSFW_i2v_14b_low_lighting_fp16_v2.1.safetensorsUNETLoaderLowNoise, built-in lightning acceleration

GGUF Q8 (Alternative, needs external lightning LoRAs):

ModelLoaderNotes
Wan2.2-I2V-A14B-HighNoise-Q8_0.ggufUnetLoaderGGUFHighNoise, quantized
Wan2.2-I2V-A14B-LowNoise-Q8_0.ggufUnetLoaderGGUFLowNoise, quantized

Official fp8:

ModelLoaderNotes
wan2.2_i2v_high_noise_14B_fp8_scaled.safetensorsUNETLoaderHighNoise, needs lightning LoRA
wan2.2_i2v_low_noise_14B_fp8_scaled.safetensorsUNETLoaderLowNoise, needs lightning LoRA
Text Encoder
ModelNodeNotes
nsfw_wan_umt5-xxl_bf16_fixed.safetensorsCLIPLoaderGGUF (type=wan)NSFW-tuned, pair with Remix models
umt5_xxl_fp8_e4m3fn_scaled.safetensorsCLIPLoader (type=wan)Standard UMT5-XXL fp8
CLIP Vision + VAE
ComponentNodeModel
CLIP VisionCLIPVisionLoaderclip_vision_h.safetensors
VAEVAELoaderwan_2.1_vae.safetensors

ModelSamplingSD3 (REQUIRED)

WAN 2.2 uses flow matching and requires ModelSamplingSD3 applied to each UNET:

json
{"class_type": "ModelSamplingSD3", "inputs": {"model": ["<unet>", 0], "shift": 5}}

shift=5 for lightning/Remix models. shift=8 for standard (non-lightning) models.

Lightning LoRAs

Remix NSFW models have lightning baked in. No external LoRA needed.

For GGUF/fp8 models, use paired hi/lo lightning LoRAs:

  • wan2.2_i2v_lightx2v_4steps_lora_v1_high_noise.safetensors → HighNoise UNET
  • wan2.2_i2v_lightx2v_4steps_lora_v1_low_noise.safetensors → LowNoise UNET

LoRA Stacks (rgthree)

Each model path has two stacked loaders (Common + Specific), each supporting 4 LoRA slots:

Hi path: UNETLoader(HN) → ModelSamplingSD3(shift=5) → Hi Common Stack → Hi Lora Stack → MODEL_HI
Lo path: UNETLoader(LN) → ModelSamplingSD3(shift=5) → Lo Common Stack → Lo Lora Stack → MODEL_LO

Common stacks hold shared LoRAs (quality/style). Specific stacks hold model-variant LoRAs. Set slots to "None" when unused. Even with no LoRAs, include the stacks. They pass CLIP through for text encoding.

Image Resizing (ImageResizeKJv2)

Input frames MUST be resized to the target video resolution before FLF and CLIPVisionEncode. The end frame inherits width/height from the start frame's resize so the dimensions match.

json
{"class_type": "ImageResizeKJv2", "inputs": {
  "image": ["<load_image>", 0], "width": 480, "height": 720,
  "upscale_method": "nearest-exact", "keep_proportion": "crop",
  "pad_color": "0, 0, 0", "crop_position": "center", "divisible_by": 2
}}

KSamplerAdvanced Two-Pass Settings

ParameterPass 1 (Hi)Pass 2 (Lo)
modelHi LoRA stack outputLo LoRA stack output
add_noiseenabledisable
steps44
cfg11
sampler_nameuni_pcuni_pc
schedulerbetabeta
start_at_step02
end_at_step24
return_with_leftover_noiseenabledisable
latent_imageWanFLF output[2]Pass 1 output[0]

Both passes share the same positive/negative conditioning from WanFirstLastFrameToVideo outputs [0] and [1].

For standard (non-lightning) models: steps=20, split at step 10, cfg=4, sampler=euler, scheduler=simple, shift=8.

Negative Prompt (REQUIRED)

Always include a quality negative prompt:

The tones are vibrant, overexposed, static, details are unclear, subtitles, style, work, painting, image, still, overall grayish, worst quality, low quality, JPEG compression artifacts, ugly, incomplete, extra fingers, poorly drawn hands, poorly drawn faces, deformed, disfigured, distorted limbs, merged fingers, motionless image, cluttered background, three legs, many people in the background, walking backwards

Node: WanFirstLastFrameToVideo

Required Inputs:
  - positive: CONDITIONING (from CLIPTextEncode)
  - negative: CONDITIONING (from CLIPTextEncode with negative prompt)
  - vae: VAE
  - width: INT (from ImageResizeKJv2 end frame output[1])
  - height: INT (from ImageResizeKJv2 end frame output[2])
  - length: INT (default 81, step 4) — number of frames
  - batch_size: INT (default 1)

Optional Inputs:
  - clip_vision_start_image: CLIP_VISION_OUTPUT (from CLIPVisionEncode)
  - clip_vision_end_image: CLIP_VISION_OUTPUT (from CLIPVisionEncode)
  - start_image: IMAGE (resized start frame)
  - end_image: IMAGE (resized end frame)

Outputs:
  - [0] positive: CONDITIONING → feed to BOTH Hi and Lo KSamplerAdvanced
  - [1] negative: CONDITIONING → feed to BOTH Hi and Lo KSamplerAdvanced
  - [2] latent: LATENT → feed to Hi Pass only (Lo Pass gets Hi Pass output)

Pipeline Flow

UNETLoader (HighNoise) → ModelSamplingSD3 (shift=5) → Hi Common Stack → Hi Lora Stack → MODEL_HI
UNETLoader (LowNoise) → ModelSamplingSD3 (shift=5) → Lo Common Stack → Lo Lora Stack → MODEL_LO
CLIPLoaderGGUF (wan) → CLIP
  ├─ CLIPTextEncode (positive) → CONDITIONING
  └─ CLIPTextEncode (negative) → CONDITIONING
CLIPVisionLoader → CLIPVisionEncode (start) + CLIPVisionEncode (end)
VAELoader → VAE
LoadImage (start) → ImageResizeKJv2 (480x720) → resized start
LoadImage (end) → ImageResizeKJv2 (match dims) → resized end

WanFirstLastFrameToVideo (positive, negative, vae, clip_vision_start, clip_vision_end,
  start_image, end_image, width/height from resize)
  → modified positive [0], modified negative [1], latent [2]

KSamplerAdvanced (Hi: MODEL_HI, steps 0→2, add_noise=enable, return_leftover=enable)
  → noisy LATENT
KSamplerAdvanced (Lo: MODEL_LO, steps 2→4, add_noise=disable, return_leftover=disable)
  → final LATENT

VAEDecode → IMAGE → VHS_VideoCombine (raw output)
                   → VRAM_Debug → SeedVR2VideoUpscaler (1080p) → VHS_VideoCombine (upscaled)

Complete workflow (API JSON)

The full Native FLF (Remix NSFW + Lightning) graph is in references/workflows.md.

Optional: Video Upscaling with SeedVR2

Add after VAEDecode for AI-powered video upscaling to 1080p. Use VRAM_Debug to free VRAM between generation and upscaling:

json
{
  "25": { "class_type": "VRAM_Debug", "inputs": {
    "image_pass": ["23", 0], "empty_cache": true, "gc_collect": true, "unload_all_models": true
  }},
  "26": { "class_type": "SeedVR2LoadDiTModel", "inputs": {
    "model": "seedvr2_ema_3b_fp8_e4m3fn.safetensors", "device": "cuda:0",
    "blocks_to_swap": 0, "swap_io_components": false, "cache_model": false, "attention_mode": "sdpa"
  }},
  "27": { "class_type": "SeedVR2LoadVAEModel", "inputs": {
    "model": "ema_vae_fp16.safetensors", "device": "cuda:0",
    "encode_tiled": false, "decode_tiled": false, "cache_model": false
  }},
  "28": { "class_type": "SeedVR2VideoUpscaler", "inputs": {
    "image": ["25", 1], "dit": ["26", 0], "vae": ["27", 0],
    "seed": 0, "resolution": 1080, "max_resolution": 0,
    "batch_size": 5, "uniform_batch_size": false, "color_correction": "lab"
  }},
  "29": { "class_type": "VHS_VideoCombine", "inputs": {
    "images": ["28", 0], "frame_rate": 16, "loop_count": 0,
    "filename_prefix": "wan_flf_upscaled", "format": "video/h264-mp4",
    "pingpong": false, "save_output": true,
    "pix_fmt": "yuv420p", "crf": 19, "save_metadata": true, "trim_to_audio": false
  }}
}

Alternative: GGUF Models with Lightning LoRAs

When using GGUF Q8 models instead of Remix, add paired lightning LoRAs:

Hi path: UnetLoaderGGUF(HN Q8) → ModelSamplingSD3(shift=5) → LoraLoaderModelOnly(hi_noise_lightning) → Hi Common Stack → Hi Lora Stack
Lo path: UnetLoaderGGUF(LN Q8) → ModelSamplingSD3(shift=5) → LoraLoaderModelOnly(lo_noise_lightning) → Lo Common Stack → Lo Lora Stack

LoRA files:

  • Unknown\no tags\wan2.2_i2v_lightx2v_4steps_lora_v1_high_noise.safetensors
  • Unknown\no tags\wan2.2_i2v_lightx2v_4steps_lora_v1_low_noise.safetensors

Approach 2: WanVideoWrapper (Advanced Control)

Uses the WanVideoWrapper custom node pack for more control over conditioning, caching, context windows, and advanced features.

Key Differences from Native
  • Uses WANVIDEOMODEL type instead of generic MODEL
  • Uses WANVIDIMAGE_EMBEDS for conditioning instead of CONDITIONING
  • Has own sampler (WanVideoSampler) with shift parameter and scheduler options
  • Supports TeaCache, MagCache, EasyCache for speed optimization
  • Supports context windows for longer videos
  • VACE module provides more flexible frame conditioning
VACE-Based FLF Pipeline
WanVideoModelLoader → WANVIDEOMODEL
WanVideoVAELoader → WANVAE
WanVideoTextEncode → WANVIDEOTEXTEMBEDS
WanVideoClipVisionEncode (start + end images) → WANVIDIMAGE_CLIPEMBEDS

WanVideoVACEStartToEndFrame (start_image, end_image, num_frames=81)
  → images batch, masks

WanVideoVACEEncode (vae, input_frames, input_masks, width, height, num_frames)
  → WANVIDIMAGE_EMBEDS (vace_embeds)

WanVideoSampler (model, image_embeds, text_embeds, steps, cfg, shift, scheduler)
  → LATENT

WanVideoDecode (vae, samples) → IMAGE → VHS_VideoCombine → MP4
WanVideoSampler Settings
ParameterStandardLightningNotes
steps304
cfg6.01.0
shift5.05.0Flow matching shift
schedulerunipceulerWanVideoWrapper has own schedulers
force_offloadtruetrueMove model to CPU after sampling
When to Use WanVideoWrapper vs Native
FeatureNativeWanVideoWrapper
SimplicitySimplerMore complex
Dual Hi-LoManual two-passMay handle internally
LoRA loadingLora Loader Stack (rgthree)WanVideoLoraSelect → WanVideoModelLoader lora (see merge_loras caveat)
Caching (TeaCache)Not availableBuilt-in
Context windowsNot availableWanVideoContextOptions
Block swap (VRAM)Not availableWanVideoBlockSwap
VACE conditioningNot availableFull VACE support
Long video (>81 frames)LimitedInfiniteTalk / context windows

Recommendation: use Native dual hi-lo for standard FLF transitions. Use WanVideoWrapper when you need caching, context windows, VRAM management, or advanced conditioning.

⚠️ CRITICAL: merge_loras=false with fp8-scaled models

When loading a LoRA through WanVideoLoraSelect → WanVideoModelLoader's lora input on an fp8-quantized model (quantization=fp8_e4m3fn_scaled, e.g. the official wan2.2_i2v_high/low_noise_14B_fp8_scaled weights), you MUST set the WanVideoLoraSelect widget merge_loras=false.

  • merge_loras=true (the node default) tries to bake the LoRA deltas into the already-quantized fp8 weights. That merge path hard-crashes ComfyUI during LoRA loading. The process dies with no Python traceback (so panel_get_errors / the frontend show nothing; only a process restart/OOM-style symptom). This is the #1 cause of a "crashed on lora loading" report with the wrapper.
  • merge_loras=false applies the LoRA as a runtime patch during the forward pass instead of merging. It is fp8-safe with negligible speed cost. This is the correct setting for the lightx2v 4-step lightning LoRAs (hi + lo) on the fp8 hi/lo I2V models.
  • It also pairs cleanly with block swap: WanVideoBlockSwap (e.g. 20 to 30 of 40 blocks → RAM) + merge_loras=false is the verified combo for fp8 14B I2V at 720p/81f on a 24GB card. (If you instead use a non-quantized bf16/fp16 model, merge_loras=true is fine.)

Separately, at 720p/81f enable enable_vae_tiling=true on WanVideoDecode. The full-frame decode is the other common uncaught-OOM crash point.

Resolution & Frame Count

Standard Resolutions
AspectResolutionMegapixels
Portrait 2:3480x7200.35MP (recommended default)
Landscape 16:9832x4800.4MP
Portrait 9:16480x8320.4MP
Square640x6400.4MP

Width and height must be divisible by 16. Use ImageResizeKJv2 with divisible_by: 2 and keep_proportion: crop.

Frame Count
  • 81 frames at 16fps = ~5 seconds (default, recommended)
  • 49 frames at 16fps = ~3 seconds (faster, less motion)
  • 121 frames at 16fps = ~7.5 seconds (longer, more VRAM)
  • Frame count should be 4n + 1 (1, 5, 9, ..., 49, 81, 121)
Frame Rate

Standard: 16 fps for WAN 2.2 output.

Video Output

VHS_VideoCombine
json
{
  "class_type": "VHS_VideoCombine",
  "inputs": {
    "images": ["<vae_decode>", 0],
    "frame_rate": 16,
    "loop_count": 0,
    "filename_prefix": "wan_flf",
    "format": "video/h264-mp4",
    "pingpong": false,
    "save_output": true,
    "pix_fmt": "yuv420p",
    "crf": 19,
    "save_metadata": true,
    "trim_to_audio": false
  }
}

VRAM Considerations

Show full SKILL.md (707 more words)Show less
Dual Hi-Lo with Remix fp16
  • Two UNETs loaded sequentially (ComfyUI offloads between passes): ~14GB each
  • NSFW UMT5-XXL bf16: ~8GB (offloaded after text encoding)
  • CLIP Vision H: ~1.5GB (offloaded after encoding)
  • VAE: ~200MB
  • Latent (81 frames at 480x720): ~1-2GB

ComfyUI manages VRAM by offloading models between passes. The Hi UNET is offloaded before the Lo UNET loads.

Tips
  1. Always clear_vram before switching to WAN from another model family
  2. Use VRAM_Debug node between generation and SeedVR2 upscaling to free all VRAM
  3. For 24GB GPUs, 81 frames at 480x720 is the practical maximum
  4. Remix NSFW models have lightning baked in. No separate LoRA needed, 4 steps total

Morph LoRAs (Smooth Metamorphosis)

By default, FLF produces a transition/dissolve between frames. For true morphing (one shape continuously reshaping into another), use a morph LoRA on both Hi and Lo paths.

VariantFileStrengthNotes
HighNoisewan2.2_i2v_magical_morph_highnoise.safetensors0.7-1.0Apply to Hi Common stack
LowNoisewan2.2_i2v_magical_morph_lownoise.safetensors0.7-1.0Apply to Lo Common stack
  • Source: NikolaSigmoid/wan2.2-i2v-loras-magical-morph
  • No trigger word needed. The LoRA modifies the denoising behavior
  • Strength 1.0 can add visual sparkle/particle effects. Reduce to 0.7-0.8 for cleaner morphs
  • Works with Remix NSFW models (no conflict with built-in lightning)
SkinMorph Redmond (Alternative — Face/Body Focus)

For person-to-person morphs (identity, gender transforms):

  • Trigger word: Skin morph
  • Strength: 0.8-1.0
  • Source: CivitAI

Prompt Tips

Describe the transition motion in addition to the start/end states:

Good: "A small cat sitting on the ground smoothly transforms and grows into a woman standing tall, seamless transformation, cinematic"
Bad: "A cat and a girl"

IMPORTANT: prompt language affects visuals.

  • AVOID words like "magical", "enchanted", "mystical". They cause literal sparkle/particle effects
  • USE clean motion language: "smoothly transforms", "gradually reshapes", "seamlessly morphs", "transitions into"
  • The morph LoRA handles the morphing effect. The prompt should describe motion and form change, not style
  • Include scale/position cues when subjects differ in size: "grows into", "expands upward", "shrinks down"

Settings Quick Reference

ConfigLightning (Remix)Standard
ModelsRemix NSFW Hi+Lo fp16Official Hi+Lo fp8
CLIPnsfw_wan_umt5-xxl_bf16_fixedumt5_xxl_fp8_e4m3fn_scaled
ModelSamplingSD3 shift58
Total steps420
Hi pass end_at_step210
CFG14
Sampleruni_pceuler
Schedulerbetasimple
External LoRA neededNo (built-in)Yes (paired hi/lo)

Multi-Step Pipeline Pattern

Anchor Frame Strategy (Proportions)

When the start and end frames have different subject sizes (e.g., small cat → tall person), generate the "anchor" frame first (the one with the most complex composition), then use Qwen Edit to create the other frame from it. This gives you:

  • Consistent background/scene between frames
  • Correct relative proportions (the edit inherits the scene scale)
  • Better FLF results since both frames share the same visual context

Example, cat-to-girl morph:

  1. Generate girl standing in front of barn with Z-Image (she fills the frame)
  2. Qwen Edit: "Replace the woman with a small cat sitting at the bottom of the image"
  3. FLF: cat (start) → girl (end). Proportions are correct because the barn establishes scale

Anti-pattern: generating cat and girl independently produces mismatched scale.

Full Pipeline
  1. Generate anchor frame with Z-Image/SDXL/Flux (portrait orientation for standing subjects)
  2. Qwen Edit to create second frame. The edit preserves scene context
  3. Clear VRAM between model families
  4. Stage both frames as inputs. When the frames are ComfyUI OUTPUTS from a prior stage (the generated/edited frames above), use upload_image (action:"stage") with each output's { filename, subfolder?, type? } and feed the returned input filename into each LoadImage. (For a frame already on local disk, use upload_image (action:"image").) NEVER copy the output file into, or guess, a filesystem input/ path. ComfyUI's input/output dirs may be CUSTOM (--input-directory / --output-directory), so a guessed path makes LoadImage reject the file (Invalid image file) and wastes the render. upload_image (action:"stage") routes through the server API (/view → /upload/image), which resolves the real dirs correctly.
  5. Run dual hi-lo FLF with morph LoRA if morphing is desired
  6. Optionally upscale with SeedVR2 to 1080p

Proven timing on RTX 4090: Z-Image (35s) → Qwen Edit (78s) → WAN FLF 81 frames (139s) = ~4 minutes total.

Working with Saved Workflows

Use get_workflow (action:"analyze") to understand any saved WAN FLF workflow before modifying or executing it. It returns a structured summary with sections, node IDs, key settings, and virtual wire connections. No raw JSON needed.

get_workflow(action="analyze", filename="Wan FirstLastFrame Advanced.json")                # summary view (default)
get_workflow(action="analyze", filename="Wan FirstLastFrame Advanced.json", view="flat")   # mermaid diagram

Only use get_workflow when you need the raw JSON for enqueue_workflow or create_workflow (action:"modify").

Sources

  • Official: none found.
  • Empirical: sampler values, wiring, and prompt notes from working graphs in packs/ and observed renders; not a vendor prompting guide.

© artokun, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (references) in plugin/skills/wan-flf-video of artokun/comfyui-mcp.

  • SKILL.md
  • references/workflows.md

Open the folder on GitHubat commit 6ad6fc0

Compare with similar skills

Wan Flf Video next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Wan Flf Video compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Wan Flf Video this skillartokun/comfyui-mcp803—~5.1kAutomated safety check: PassMIT
Gemma Trainergoogle-gemma/gemma-skills1k—~1.9kAutomated safety check: PassApache-2.0
Hugging Face LLM Trainerhuggingface/skills11k1 repos~7.2kAutomated safety check: PassApache-2.0
Unsloth Finetuningsickn33/agentic-awesome-skills47k1 repos~4.1kAutomated safety check: PassApache-2.0
Quantized Exportwshobson/agents40k—~2kAutomated safety check: PassMIT
ML Research LabAnastasiyaW/codex-claude-code-config154—~794Automated safety check: PassMIT

Similar skills

  • Gemma Trainer

    google-gemma/gemma-skills

    Trigger this skill when the user wants to train, fine-tune, or adapt Gemma models (e.g.

    1k GitHub stars~1.9k tokensUpdated 4 days ago
    AI & LLM EngineeringAuto-check passed
  • Hugging Face LLM Trainer

    huggingface/skills

    Official

    Trains or fine-tunes language and vision models with TRL or Unsloth on Hugging Face Jobs cloud GPUs, then converts the results to GGUF.

    11k GitHub starsUsed in 1 repo~7.2k tokens
    AI & LLM EngineeringAuto-check passed
  • Unsloth Finetuning

    sickn33/agentic-awesome-skills

    Fine-tune and post-train LLMs with Unsloth Core on a single consumer GPU: VRAM sizing, LoRA/QLoRA, GRPO/DPO, chat-template correctness, and GGUF export.

    47k GitHub starsUsed in 1 repo~4.1k tokens
    AI & LLM EngineeringAuto-check passed
  • Quantized Export

    wshobson/agents

    Export a promoted fine-tuned model in the right deployment format — merged safetensors, LoRA-only, GGUF with imatrix, or FP8.

    40k GitHub stars~2k tokensUpdated 6 days ago
    AI & LLM EngineeringAuto-check passed
  • ML Research Lab

    AnastasiyaW/codex-claude-code-config

    Machine-learning research loop for dataset curation, fine-tuning, evaluation, inference deployment, experiment tracking, and model explainability.

    154 GitHub stars~794 tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • Finetuning

    ericrisco/rsc-harness

    A skill your agent uses when adapting an open-weight model to a target form or behavior — tone, output format, reasoning pattern — via LoRA/QLoRA or full fine-tuning with TRL SFTTrainer, then…

    180 GitHub stars~3.8k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed

More from artokun/comfyui-mcp

All 42 skills in this repo
  • AI Toolkit Trainer

    artokun/comfyui-mcp

    Train custom LoRAs with ostris AI-Toolkit. An agent skill from artokun/comfyui-mcp.

    803 GitHub stars~2.7k tokensUpdated 6 days ago
    Auto-check passed
  • Anima Base

    artokun/comfyui-mcp

    Anime/illustration text-to-image (ANIMA 1.0, ~2B Cosmos DiT).

    803 GitHub stars~4k tokensUpdated 6 days ago
    Auto-check passed
  • Civitai

    artokun/comfyui-mcp

    Discover Civitai models with the BUILT-IN downloadmodel action:"searchcivitai" and install/generate them locally.

    803 GitHub stars~1.1k tokensUpdated 6 days ago
    Auto-check passed
  • Color Correction

    artokun/comfyui-mcp

    Diagnose and fix video/image color OBJECTIVELY with the getimage (action:"analyzecolor") tool (scopes/stats such as black/white points, contrast, saturation, clipping, cast) instead of eyeballing a…

    803 GitHub stars~2.4k tokensUpdated 6 days ago
    Auto-check passed
  • Comfyui Frontend Extensions

    artokun/comfyui-mcp

    Authoring ComfyUI v2 frontend extensions with @comfyorg/extension-api, covering defineNode/defineExtension/defineWidget, shell UI (sidebar tabs, commands, hotkeys), typed events, and handles.

    803 GitHub stars~5.4k tokensUpdated 6 days ago
    Auto-check passed
  • Comfyui Launch Flags

    artokun/comfyui-mcp

    Pick the right ComfyUI startup flags for VRAM, attention, caching, and speed.

    803 GitHub stars~3.1k tokensUpdated 6 days ago
    Auto-check passed

Works with

Questions about Wan Flf Video

What does Wan Flf Video do?

Build WAN 2.2 First-Last-Frame video workflows. An agent skill from artokun/comfyui-mcp. Wan Flf Video is an agent skill from artokun/comfyui-mcp.2 First-Last-Frame video workflows.

When should I use Wan Flf Video?

Wan Flf Video fits situations like: tasks that involve Fine-tuning.

How do I install Wan Flf Video in Claude Code?

Run `npx skills add artokun/comfyui-mcp --skill wan-flf-video -a claude-code`. Or copy the skill folder (plugin/skills/wan-flf-video in artokun/comfyui-mcp) into .claude/skills/wan-flf-video in your project. Claude Code loads it when a task matches its description.

How do I install Wan Flf Video in Codex?

Run `npx skills add artokun/comfyui-mcp --skill wan-flf-video -a codex`. Or copy the skill folder (plugin/skills/wan-flf-video in artokun/comfyui-mcp) into .agents/skills/wan-flf-video in your project. Codex loads it when a task matches its description.

Can I use Wan Flf Video in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add artokun/comfyui-mcp --skill wan-flf-video -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/wan-flf-video, .gemini/skills/wan-flf-video, .github/skills/wan-flf-video and .opencode/skills/wan-flf-video in your project.

What does Wan Flf Video need to run?

SKILL.md names no scripts, command-line tools or credentials: Wan Flf Video is instructions for the agent only.

Does Wan Flf Video access the network?

SKILL.md names 2 domains. As links in the text: huggingface.co and civitai.com. This is read from the text; nothing was executed.

Is Wan Flf Video safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Wan Flf Video use?

Wan Flf Video is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Wan Flf Video use?

About 5.1k tokens (SKILL.md is roughly 20k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.5k tokens, read only when the agent opens those files.

What are the alternatives to Wan Flf Video?

Skills that share tags, products or a category with Wan Flf Video: Gemma Trainer (google-gemma/gemma-skills, 1k stars), Hugging Face LLM Trainer (huggingface/skills, 11k stars), Unsloth Finetuning (sickn33/agentic-awesome-skills, 47k stars) and Quantized Export (wshobson/agents, 40k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Wan Flf Video?

artokun (a GitHub user) maintains it in artokun/comfyui-mcp, which has 803 GitHub stars. The repository holds 42 skills in this directory. The repository was last updated on October 5, 2026.

Source: artokun/comfyui-mcp on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.