Agent skill

AI Toolkit Trainer

by artokun in artokun/comfyui-mcp

Train custom LoRAs with ostris AI-Toolkit. An agent skill from artokun/comfyui-mcp.

MITAuto-check passedAI & LLM Engineering

Install AI Toolkit Trainer

skills CLI
$ npx skills add artokun/comfyui-mcp --skill ai-toolkit-trainer -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install artokun/comfyui-mcp ai-toolkit-trainer --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/artokun/comfyui-mcp.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugin/skills/ai-toolkit-trainer .claude/skills/ai-toolkit-trainer && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
ai-toolkit-trainer
GitHub stars
803
Token cost
~2.7k tokens
SKILL.md length
1,388 words
Files
1
Skills in repo
42
Repo updated
First seen
Licence
MIT

At a glance

Train custom LoRAs with ostris AI-Toolkit. An agent skill from artokun/comfyui-mcp.

  • Works in 3 steps: Copy .safetensors into ComfyUI… → Load with LoraLoaderModelOnly → Prompt using the trigger word or caption…
  • The user wants to train a WAN
  • SKILL.md covers Overview, Install, Launching the web UI and Dataset preparation, plus 6 more sections
  • Calls pip, python and npm; reaches download.pytorch.org and github.com

What it does

AI Toolkit Trainer is an agent skill from artokun/comfyui-mcp. Train custom LoRAs with ostris AI-Toolkit. Covers WAN 2.2/2.1 (people, styles, video motion) and Z-Image (Turbo & Base, low-VRAM image LoRAs). Use when the user wants to train a WAN or Z-Image LoRA; covers local + RunPod setup, dataset prep, key params, and using the result in a ComfyUI workflow.

Its SKILL.md is about 2.7k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering Fine-tuning and Diffusion and image models. It works with ComfyUI and Python. The repository describes itself as: Local-first, agent-native control plane for ComfyUI — MCP server + sidebar agent that generates images, video & audio, authors and runs workflows, and edits your live graph in… The licence is MIT.

When your agent uses it

  • The user wants to train a WAN
  • Covers local + RunPod setup
  • Using the result in a ComfyUI workflow

Example prompts

  • “/ai-toolkit-trainer”

Requirements

  • Python 3
  • Node.js

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. Copy .safetensors into ComfyUI models/loras/.
  2. Load with LoraLoaderModelOnly
  3. Prompt using the trigger word or caption style you trained with. For WAN motion LoRAs, describe the same camera or motion.

What it can do on your machine

Read from SKILL.md and the folder at commit 6ad6fc0. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pip
    • python
    • npm

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • download.pytorch.org
    • github.com
    • huggingface.co

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

AI Toolkit Trainer loads about 2.7k tokens when it runs. Until then it costs about 79 tokens; SKILL.md has 1,388 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~79
When it runs · the whole SKILL.md, loaded when a task matches
~2.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from artokun/comfyui-mcp at commit 6ad6fc0, republished under its MIT licence (© artokun). 1,388 words, ~2,729 tokens.

Download SKILL.mdSave it as .claude/skills/ai-toolkit-trainer/SKILL.md (or your agent's skills folder).
name
ai-toolkit-trainer
description
Train custom LoRAs with ostris AI-Toolkit. Covers WAN 2.2/2.1 (people, styles, video motion) and Z-Image (Turbo & Base, low-VRAM image LoRAs). Use when the user wants to train a WAN or Z-Image LoRA; covers local + RunPod setup, dataset prep, key params, and using the result in a ComfyUI workflow.
globs
**/*.json

AI-Toolkit LoRA Trainer (WAN 2.2 & Z-Image)

Overview

AI-Toolkit by ostris is an MIT-licensed trainer for finetuning diffusion models. It is a standalone trainer with its own web UI, not a ComfyUI custom node. It runs a Node.js UI front end over a Python (run.py) training backend and trains LoRAs for many model families. This skill covers the WAN 2.2 / 2.1 video models and Z-Image (Turbo & Base).

  • Repo: https://github.com/ostris/ai-toolkit (cloned by the installers).
  • Backend: python run.py config/<job>.yml. UI: a Node.js app under ui/ that schedules and monitors jobs. You do not have to keep the UI open while a job runs.
  • Output: a standard .safetensors LoRA you drop into ComfyUI models/loras/ and load with LoraLoaderModelOnly.

Best for:

  • WAN LoRAs. A person or character, an art style, or a specific camera or video motion, trained from image or video clip datasets. For using WAN see wan-t2v-video / wan-flf-video.
  • Z-Image LoRAs. Fast, very low-VRAM image LoRAs (faces, characters, outfits, styles) on the 6B Z-Image base/turbo. For using Z-Image see z-image-base / z-image-turbo, and the z-image-xy-plot pack to compare trained LoRAs.

For low-VRAM anime image LoRAs on a different stack (kohya sd-scripts), see the sibling anima-lora-trainer.

Two LoRA kinds for WAN. A WAN image LoRA trains on still images; it is cheaper (~24GB-class) and suits identity or style. A WAN video LoRA trains on short clips; it is heavier, best run on cloud, and suits motion. Z-Image is image-only.

Install

The installer comes in two generations. Both clone ostris/ai-toolkit, set up Torch for your GPU, and launch the web UI. Put it in a folder whose full path has no spaces (e.g. C:\AI-Toolkit).

  • V1, AI-TOOLKIT_AUTO_INSTALL.bat, expects Git, Python 3.10.x, and Node 18+ already in PATH.
  • V2, AI-TOOLKIT_AUTO_INSTALL-V2.bat (recommended), uses an embedded Python 3.10.11, auto-installs Git and Node, builds a clean PATH without your system Python, and adds aggressive pip/curl retries. It has far fewer prerequisites and fails less often. The Z-Image Turbo LoRA training release used it.

Both are CUDA-aware and select the Torch wheel by GPU generation:

ChoiceGPUCUDATorch indexTorch packages
1RTX 50-series (Blackwell)12.8https://download.pytorch.org/whl/cu128torch==2.7.0 torchvision==0.22.0
2RTX 40 / 30 / 20 and older12.6https://download.pytorch.org/whl/cu126torch==2.7.0 torchvision==0.22.0

Each then clones ostris/ai-toolkit, downloads two launcher scripts (LAUNCHER-TOOLKIT.bat, SECURE_LAUNCHER-TOOLKIT.bat, from https://huggingface.co/Aitrepreneur/FLX/resolve/main/), makes the venv, installs Torch from the chosen index, runs pip install -r requirements.txt, then cd ui && npm run build_and_start.

RunPod / Linux — AI-TOOLKIT_AUTO_INSTALL-RUNPOD.sh (and -V2.sh)

Installs into the persistent volume /workspace/ai-toolkit. It is idempotent; a re-run just relaunches the UI. Use RunPod's PyTorch 2.8.0 template and a 100GB disk. It installs apt deps, clones the repo, makes a venv, installs Torch (torchaudio included), installs nvm + Node 22, then builds and starts the UI.

ChoiceGPUStreamTorch spec
1RTX 5000-series (Blackwell)cu128torch==2.7.0+cu128 torchvision==0.22.0+cu128 torchaudio==2.7.0+cu128
2Ada / Hopper / Ampere, oldercu126torch==2.7.0 torchvision==0.22.0 torchaudio==2.7.0

The UI listens on 8675 and Jupyter on 8888. Set AI_TOOLKIT_AUTH (UI password) before launch. Reach it at https://${RUNPOD_POD_ID}-8675.proxy.runpod.net. Use an RTX 4090/5090 for image (WAN t2i/t2v, Z-Image) LoRAs and an RTX 6000 Pro (Blackwell) for heavy WAN video, high-res, or high-rank jobs.

Launching the web UI

  • On Windows, run LAUNCHER-TOOLKIT.bat (local) or SECURE_LAUNCHER-TOOLKIT.bat (password-protected) from the ai-toolkit folder.
  • On RunPod, rerun the .sh. It detects the install and starts the UI on :8675.

In the UI, create a Job, point it at a dataset folder, pick the model (WAN variant or Z-Image), set params, and start. Jobs run in the Python backend, so you can close the browser. To bypass the UI, copy a config/examples/*.yml, edit it, and run python run.py config/<job>.yml.

Dataset preparation

AI-Toolkit pairs each sample with a same-basename .txt caption and auto-resizes/buckets aspect ratios (no pre-cropping).

Image LoRA (WAN identity/style, or Z-Image)
my_dataset/
  001.png  001.txt
  002.jpg  002.txt
  • Captions are natural language. Include a unique trigger word for a person or character.
  • Use about 15 to 40 varied images for a person, more for a broad style.
Video LoRA (WAN motion only)

Short clips plus a .txt per clip; caption the motion or camera move. Set per-clip frames via the job's num_frames (e.g. 81). This is markedly heavier, so prefer cloud GPUs.

Key training params

WAN 2.2

WAN 2.2 14B is a Mixture-of-Experts with a high-noise expert (structure/motion) and a low-noise expert (detail). AI-Toolkit trains both via Multi-stage.

ParamDefaultNotes
Linear rank / dim1616 simple; 16–32 complex/cinematic
Learning rate5e-5 (identity)7e-5–1e-4 style; high LR → plasticky skin
Steps1500–2500stop before overbaking
Resolution512 (or 768)bucketed; 768 costs more VRAM
num_frames (video)81per-clip frame count
Multi-stageHigh + Low = ONtrains both experts
Switch Every10raise to 20–50 if offload swapping is slow
Optimizer / QuantAdamW8bit / 4-bit ARA or float8fits 14B on consumer cards
Show full SKILL.md (627 more words)Show less
Z-Image (Turbo & Base)

Z-Image is a ~6B single-stream model with no hi/lo multi-stage. Leave Multi-stage OFF; you train one model. It is the lightest target here. The headline of the Z-Image releases is training on very low VRAM.

ParamStarting pointNotes
Linear rank / dim16–3232 for detailed characters/styles
Learning rate1e-4lower (5e-5) for tighter identity
Steps1500–3000dataset-dependent
Resolution768 (or 1024)Z-Image's native range
Multi-stageOFFsingle-stream model, not WAN's MoE
Optimizer / QuantAdamW8bit / float8enables sub-12GB training

Train on Base, deploy anywhere. Z-Image Base is the finetuning-friendly model; a LoRA trained on Base generally applies to the Turbo workflow too. Use the z-image-xy-plot pack to grid-compare your trained LoRAs.

The param tables are aggregated starting points from community and training-guide sources, not read from the repo's config/examples/*.yml. Open the actual WAN / Z-Image example config in your clone and tune. See "Unverified".

VRAM / GPU guidance

  • Z-Image image LoRA is the lightest. It trains on modest consumer GPUs with quantization (the releases describe very-low-VRAM training); a 4090 is comfortable, and smaller cards work with float8 at 512 to 768 res.
  • WAN image LoRA (t2i/t2v) needs 24GB+ locally with quantization. Below that, use RunPod.
  • WAN video LoRA, high res, or high rank is heavier. Use cloud (RTX 5090, or RTX 6000 Pro Blackwell / H100).
  • Memory savers: quantization, batch size 1, 512 res, and (WAN) raising Switch Every.

Using the trained LoRA in ComfyUI

  1. Copy <your_lora>.safetensors into ComfyUI models/loras/.
  2. Load with LoraLoaderModelOnly:
    • WAN 2.2 is dual hi/lo. Apply the LoRA to both the HighNoise and LowNoise model branches (like lightning/concept LoRAs in wan-t2v-video). Typical strength 0.5 to 1.0.
    • Z-Image is a single model. Use one LoraLoaderModelOnly on the Z-Image model path (see the z-image-base / z-image-turbo packs). Strength 0.7 to 1.0.
    json
    { "class_type": "LoraLoaderModelOnly",
      "inputs": { "model": ["<base_model>", 0],
                  "lora_name": "<your_lora>.safetensors",
                  "strength_model": 1.0 } }
  3. Prompt using the trigger word or caption style you trained with. For WAN motion LoRAs, describe the same camera or motion.

Troubleshooting

  • No module named 'torchaudio' when starting a job (AI-Toolkit). The venv's Torch stack is mismatched. Activate the AI-Toolkit venv (venv\Scripts\activate), then pip uninstall torch torchaudio torchvision -y and pip install torch torchvision torchaudio --index-url https://download.pytorch.org/whl/cu121 (or your CUDA's index). This only affects the AI-Toolkit install, not ComfyUI.
  • self and mat2 must have the same dtype (ComfyUI-WanVideoWrapper, WAN usage). Re-clone ComfyUI-WanVideoWrapper in custom_nodes/ and reinstall its requirements.txt, then restart ComfyUI.
  • 5000-series (Blackwell) onnxruntime "QuickGelu" / CUDA error. pip install onnxruntime==1.20.1 in the affected venv.
  • Pascal/Maxwell GPUs (GTX 9xx/10xx). Recent Torch (cu128/cu130) dropped them. Reinstall the cu126 Torch build into the venv.
  • Path with spaces (Windows). Keep the install path space-free or the build/launch fails.
  • OOM during training. Quantization (4-bit ARA / float8), 512 res, batch 1, (WAN) raise Switch Every, or a bigger RunPod GPU.
  • RunPod UI won't load / asks for a password. Confirm AI_TOOLKIT_AUTH is set and you're on the 8675 proxy URL.

Unverified / verify before relying

  • The param tables (both WAN and Z-Image) are synthesized starting points, not read from the repo's config/examples/*.yml. Open the actual example config in your clone and adjust.
  • The release notes describe the Z-Image training VRAM floor only qualitatively ("very low VRAM"). Confirm against your card; quantization plus 512 to 768 res is the lever.
  • The Windows UI port is whatever the launcher binds (the installer doesn't print it; check the launcher window). RunPod 8675/8888 are per the template.
  • The launcher .bat files are downloaded from a third-party HuggingFace repo (Aitrepreneur/FLX); review before running on a security-sensitive machine.
  • Model weights are fetched at job time by AI-Toolkit/HF, not by the installer. Confirm the model selector lists your target WAN variant or Z-Image model before a long run.

Sources

© artokun, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in plugin/skills/ai-toolkit-trainer of artokun/comfyui-mcp.

Open the folder on GitHubat commit 6ad6fc0

Compare with similar skills

AI Toolkit Trainer next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

AI Toolkit Trainer compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
AI Toolkit Trainer this skillartokun/comfyui-mcp803—~2.7kAutomated safety check: PassMIT
Setupguaardvark/guaardvark257—~1.2kAutomated safety check: PassMIT
ComfyUI Custom Node BuilderConstantineB6/comfy-pilot230—~897Automated safety check: PassMIT
Add Comfyui NodeMooshieblob1/MooshieUI207—~936Automated safety check: PassAGPL-3.0
Edit Comfy Workflowpeteromallet/VibeComfy150—~2.2kAutomated safety check: PassMIT
ComfyUI Custom Node Basicsjtydhr88/comfyui-custom-node-skills296—~1.6kAutomated safety check: PassMIT

Similar skills

  • Setup

    guaardvark/guaardvark

    Connect this agent to a running Guaardvark (self-hosted AI studio) and check what it can do right now.

    257 GitHub stars~1.2k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • ComfyUI Custom Node Builder

    ConstantineB6/comfy-pilot

    Helps an agent write ComfyUI custom nodes in Python, including wrapping an existing script, mapping data types and handling image batches.

    230 GitHub stars~897 tokensUpdated 7 mo ago
    AI & LLM EngineeringAuto-check passed
  • Add Comfyui Node

    Mooshieblob1/MooshieUI

    Adds a custom ComfyUI Python node to MooshieUI — Python class in mooshienodes.py, Rust required-class registration, and optional workflow template chain hookup.

    207 GitHub stars~936 tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Edit Comfy Workflow

    peteromallet/VibeComfy

    Edit an existing VibeComfy or ComfyUI workflow, ready template, recipe, scratchpad, or target graph.

    150 GitHub stars~2.2k tokensUpdated 9 days ago
    AI & LLM EngineeringAuto-check passed
  • ComfyUI Custom Node Basics

    jtydhr88/comfyui-custom-node-skills

    Explains the V3 API for ComfyUI custom nodes: node classes, schema, inputs and outputs, registration and how it differs from the legacy V1 style.

    296 GitHub stars~1.6k tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check passed
  • Comfyui

    calesthio/OpenMontage

    A skill your agent uses when working with ComfyUI workflows in OpenMontage, including comfyuiimage/comfyuivideo/comfyuimusic, custom workflowjson/workflowpath inputs, outputnode selection, missing…

    66k GitHub stars~2k tokensUpdated 7 days ago
    AI & LLM EngineeringAuto-check passed

More from artokun/comfyui-mcp

All 42 skills in this repo
  • Anima Base

    artokun/comfyui-mcp

    Anime/illustration text-to-image (ANIMA 1.0, ~2B Cosmos DiT).

    803 GitHub stars~4k tokensUpdated 5 days ago
    Auto-check passed
  • Civitai

    artokun/comfyui-mcp

    Discover Civitai models with the BUILT-IN downloadmodel action:"searchcivitai" and install/generate them locally.

    803 GitHub stars~1.1k tokensUpdated 5 days ago
    Auto-check passed
  • Color Correction

    artokun/comfyui-mcp

    Diagnose and fix video/image color OBJECTIVELY with the getimage (action:"analyzecolor") tool (scopes/stats such as black/white points, contrast, saturation, clipping, cast) instead of eyeballing a…

    803 GitHub stars~2.4k tokensUpdated 5 days ago
    Auto-check passed
  • Comfyui Frontend Extensions

    artokun/comfyui-mcp

    Authoring ComfyUI v2 frontend extensions with @comfyorg/extension-api, covering defineNode/defineExtension/defineWidget, shell UI (sidebar tabs, commands, hotkeys), typed events, and handles.

    803 GitHub stars~5.4k tokensUpdated 5 days ago
    Auto-check passed
  • Comfyui Launch Flags

    artokun/comfyui-mcp

    Pick the right ComfyUI startup flags for VRAM, attention, caching, and speed.

    803 GitHub stars~3.1k tokensUpdated 5 days ago
    Auto-check passed
  • Installer Packs

    artokun/comfyui-mcp

    A skill your agent uses when installing a model family from an installer pack, or when building/deriving a new pack from an upstream installer or a workflow JSON.

    803 GitHub stars~952 tokensUpdated 5 days ago
    Auto-check passed

Works with

Questions about AI Toolkit Trainer

What does AI Toolkit Trainer do?

Train custom LoRAs with ostris AI-Toolkit. An agent skill from artokun/comfyui-mcp. AI Toolkit Trainer is an agent skill from artokun/comfyui-mcp. Train custom LoRAs with ostris AI-Toolkit.

When should I use AI Toolkit Trainer?

AI Toolkit Trainer fits situations like: the user wants to train a WAN; covers local + RunPod setup; using the result in a ComfyUI workflow.

How do I install AI Toolkit Trainer in Claude Code?

Run `npx skills add artokun/comfyui-mcp --skill ai-toolkit-trainer -a claude-code`. Or copy the skill folder (plugin/skills/ai-toolkit-trainer in artokun/comfyui-mcp) into .claude/skills/ai-toolkit-trainer in your project. Claude Code loads it when a task matches its description.

How do I install AI Toolkit Trainer in Codex?

Run `npx skills add artokun/comfyui-mcp --skill ai-toolkit-trainer -a codex`. Or copy the skill folder (plugin/skills/ai-toolkit-trainer in artokun/comfyui-mcp) into .agents/skills/ai-toolkit-trainer in your project. Codex loads it when a task matches its description.

Can I use AI Toolkit Trainer in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add artokun/comfyui-mcp --skill ai-toolkit-trainer -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ai-toolkit-trainer, .gemini/skills/ai-toolkit-trainer, .github/skills/ai-toolkit-trainer and .opencode/skills/ai-toolkit-trainer in your project.

What does AI Toolkit Trainer need to run?

Going by SKILL.md and its folder, AI Toolkit Trainer needs the command-line tools its instructions call (pip, python and npm). Our summary lists: Python 3; Node.js.

Does AI Toolkit Trainer access the network?

SKILL.md names 3 domains. In commands or code: download.pytorch.org, github.com and huggingface.co; the agent is likely to contact these when it follows the instructions. This is read from the text; nothing was executed.

Is AI Toolkit Trainer safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does AI Toolkit Trainer use?

AI Toolkit Trainer is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does AI Toolkit Trainer use?

About 2.7k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to AI Toolkit Trainer?

Skills that share tags, products or a category with AI Toolkit Trainer: Setup (guaardvark/guaardvark, 257 stars), ComfyUI Custom Node Builder (ConstantineB6/comfy-pilot, 230 stars), Add Comfyui Node (Mooshieblob1/MooshieUI, 207 stars) and Edit Comfy Workflow (peteromallet/VibeComfy, 150 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains AI Toolkit Trainer?

artokun (a GitHub user) maintains it in artokun/comfyui-mcp, which has 803 GitHub stars. The repository holds 42 skills in this directory. The repository was last updated on October 5, 2026.

Source: artokun/comfyui-mcp on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.