Agent skill

Z Image Txt2img

by artokun in artokun/comfyui-mcp

Build Z-Image txt2img workflows. An agent skill from artokun/comfyui-mcp.

MITAuto-check passedAI & LLM Engineering

Install Z Image Txt2img

skills CLI
$ npx skills add artokun/comfyui-mcp --skill z-image-txt2img -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install artokun/comfyui-mcp z-image-txt2img --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/artokun/comfyui-mcp.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugin/skills/z-image-txt2img .claude/skills/z-image-txt2img && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
z-image-txt2img
GitHub stars
803
Token cost
~2.8k tokens
SKILL.md length
796 words
Files
1
Skills in repo
42
Repo updated
First seen
Licence
MIT

At a glance

Build Z-Image txt2img workflows. An agent skill from artokun/comfyui-mcp.

  • Works in 2 steps: Z-Image Base (and RedCraft finetune).… → Z-Image Turbo. DMD-distilled, 8-10…
  • Tasks that involve Diffusion and image models
  • SKILL.md covers Overview, Models, Conditioning and Sampler Settings, plus 9 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Z Image Txt2img is an agent skill from artokun/comfyui-mcp. Build Z-Image txt2img workflows. RedCraft checkpoint, Z-Image Turbo/Base LoRAs, ControlNet, and sampler presets

Its SKILL.md is about 2.8k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering Diffusion and image models and Fine-tuning. It works with ComfyUI and Qwen. The repository describes itself as: Local-first, agent-native control plane for ComfyUI — MCP server + sidebar agent that generates images, video & audio, authors and runs workflows, and edits your live graph in… The licence is MIT.

When your agent uses it

  • Tasks that involve Diffusion and image models
  • Tasks that involve Fine-tuning

Example prompts

  • “/z-image-txt2img”

Workflow steps

2 steps, taken from the first numbered list in SKILL.md.

  1. Z-Image Base (and RedCraft finetune). Full model, supports negative prompts, LoRA training, ControlNet. 10-30 steps.
  2. Z-Image Turbo. DMD-distilled, 8-10 steps, no effective negative prompts (CFG baked in).

What it can do on your machine

Read from SKILL.md and the folder at commit 6ad6fc0. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are json).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Z Image Txt2img loads about 2.8k tokens when it runs. Until then it costs about 32 tokens; SKILL.md has 796 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~32
When it runs · the whole SKILL.md, loaded when a task matches
~2.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from artokun/comfyui-mcp at commit 6ad6fc0, republished under its MIT licence (© artokun). 796 words, ~2,849 tokens.

Download SKILL.mdSave it as .claude/skills/z-image-txt2img/SKILL.md (or your agent's skills folder).
name
z-image-txt2img
description
Build Z-Image txt2img workflows. RedCraft checkpoint, Z-Image Turbo/Base LoRAs, ControlNet, and sampler presets
globs
**/*.json

Z-Image Text-to-Image Workflows

Launch flag. Z-Image does not sample correctly under --use-sage-attention (black / garbled output). Launch ComfyUI with --use-pytorch-cross-attention for Z-Image. See comfyui-launch-flags.

Overview

Z-Image is a 6B-parameter image generation model from Alibaba's Tongyi Lab using a Scalable Single-Stream DiT (S3-DiT) architecture. It uses a Qwen text encoder (not CLIP-L/T5). Its VAE shares the Flux VAE architecture (same tensor shapes, so the file is the same 320MB size) but ships different weights. It is NOT byte-identical to Flux's ae.safetensors and must be kept as a separate file (z-image-ae.safetensors) to avoid clobbering the Flux VAE. Two variants:

  1. Z-Image Base (and RedCraft finetune). Full model, supports negative prompts, LoRA training, ControlNet. 10-30 steps.
  2. Z-Image Turbo. DMD-distilled, 8-10 steps, no effective negative prompts (CFG baked in).

Models

RedCraft Redzimage DX1 (Installed — Combined Checkpoint)
ComponentNodeModelNotes
CheckpointCheckpointLoaderSimpleredcraftRedzimageUpdatedJAN30_redzibDX1.safetensors17GB, bundles UNET+CLIP+VAE

RedCraft is a Z-Image Base finetune by the RedCraft team. Designed for faster inference than stock Z-Image Base. Uses CheckpointLoaderSimple since it's a combined checkpoint, so no separate loaders are needed.

Z-Image Turbo (Separate Components — May Need Download)
ComponentNodeModelNotes
UNETUNETLoaderz_image_turbo_bf16.safetensorsNot currently installed
CLIPCLIPLoader (type=qwen_image)qwen_3_4b.safetensorsNot currently installed
VAEVAELoaderz-image-ae.safetensors320MB. Flux VAE architecture but different weights — NOT the same file as Flux's ae.safetensors. From Comfy-Org/z_image_turbo (split_files/vae/ae.safetensors)
Z-Image Base (Separate Components — May Need Download)
ComponentNodeModelNotes
UNETUNETLoaderz_image_base_bf16.safetensorsNot currently installed
CLIPCLIPLoader (type=qwen_image)qwen_3_4b.safetensorsNot currently installed
VAEVAELoaderz-image-ae.safetensors320MB. Flux VAE architecture but different weights — NOT the same file as Flux's ae.safetensors

Conditioning

TextEncodeZImageOmni (Built-in)

For Z-Image separate component loading. Supports reference images via CLIP Vision:

Required Inputs:
  - clip: CLIP
  - prompt: STRING (multiline)
  - auto_resize_images: BOOLEAN (default true)

Optional Inputs:
  - image_encoder: CLIP_VISION (for reference images)
  - vae: VAE
  - image1-3: IMAGE (up to 3 reference images)

Outputs:
  [0] CONDITIONING
CLIPTextEncode (For RedCraft Checkpoint)

When using CheckpointLoaderSimple, standard CLIPTextEncode works since the checkpoint bundles the correct tokenizer:

json
{
  "class_type": "CLIPTextEncode",
  "inputs": { "clip": ["<checkpoint>", 1], "text": "<prompt>" }
}

Sampler Settings

RedCraft DX1
PresetStepsCFGSamplerSchedulerNotes
Distilled Fast101.0eulersimpleQuick iteration
Standard304.0eulersimpleFull quality
Z-Image Turbo
PresetStepsCFGSamplerSchedulerNotes
Author recommended141.0res_2ssimpleCopaxTimeless author pick
Beauty/fashion101.0euler_ancestralbetaSmooth skin, fashion photography
Sharpest101.0dpmpp_sdebetaSharpest, most natural (560-image test)
Z-Image Base (Two-Stage)

Stage 1, primary generation:

ParameterValue
Steps22
CFG4.0 (range 4–7)
Samplerres_2s
Schedulerbeta
Denoise1.0

Stage 2, detail refinement (optional img2img pass):

ParameterValue
Steps3
CFG4.0
Samplerres_2s
Schedulernormal
Denoise0.15

Negative Prompts

RedCraft / Z-Image Base

Supports negative prompts at CFG > 1.0:

3D, ai generated, semi realistic, illustrated, drawing, comic, digital painting, 3D model, blender, video game screenshot, screenshot, render, high-fidelity, smooth textures, CGI, masterpiece, text, writing, subtitle, watermark, logo, blurry, low quality, jpeg, artifacts, grainy
Z-Image Turbo

Negative prompts are not effective. CFG is baked in via distillation. Use the positive prompt to guide away from unwanted elements instead.

Recommended positive-side avoidance template:

over-smooth skin, plastic skin, doll face, anime, CGI, waxy texture, blurry face, fake pores, exaggerated makeup, over-sharpening, unrealistic symmetry, flat lighting, low detail skin, extra fingers, distorted anatomy

Resolutions

AspectResolutionNotes
Square1024x1024Standard
Square (native)1328x1328Higher quality at native resolution
Portrait 3:4896x1152
Portrait 5:8832x1216
Portrait 9:16768x1344
Landscape 16:91280x720

Dimensions must be divisible by 16.

LoRA System

Show full SKILL.md (332 more words)Show less
ZImageTurbo LoRAs

Located in loras/ZImageTurbo/ with subfolders:

  • style/: style LoRAs (e.g., TurboPussyZ_v2.safetensors)
  • concept/: concept LoRAs (e.g., body from below.safetensors, ZITnsfwLoRA.safetensors)
  • character/: character LoRAs (e.g., NSFW_master_ZIT_000008766.safetensors)
  • action/: action LoRAs

Use with Z-Image Turbo base model. Typical LoRA strength: 0.6 to 1.0.

ZImageBase LoRAs

Located in loras/ZImageBase/ with subfolders:

  • style/: style LoRAs (e.g., NSGIRL-Z-Image-LoRA-By-MM744.safetensors)
  • concept/: concept LoRAs

Use with Z-Image Base or RedCraft. Typical LoRA strength: 0.6 to 1.0.

Z-Image-Aesthetic-Base v1

General aesthetic improvement LoRA:

  • File: Z-Image-Aesthetic-Base v1.safetensors (352MB)
  • Settings: euler_ancestral + beta, 30 steps, CFG 4, strength 0.6 to 1.0
Applying LoRAs
json
{
  "class_type": "LoraLoader",
  "inputs": {
    "model": ["<checkpoint_or_unet>", 0],
    "clip": ["<checkpoint_or_clip>", 1],
    "lora_name": "ZImageTurbo\\style\\TurboPussyZ_v2.safetensors",
    "strength_model": 0.8,
    "strength_clip": 0.8
  }
}

When using CheckpointLoaderSimple for RedCraft, model output is index 0 and CLIP output is index 1. When stacking multiple LoRAs, chain them sequentially.

ControlNet

ZImageFunControlnet (Built-in)

Experimental built-in node for Z-Image ControlNet. Patches the model with a control signal:

Required Inputs:
  - model: MODEL
  - model_patch: MODEL_PATCH (from ControlNet loader)
  - vae: VAE
  - strength: FLOAT (default 1.0, range -10 to 10)

Optional Inputs:
  - image: IMAGE (reference/control image)
  - inpaint_image: IMAGE
  - mask: MASK

Outputs:
  [0] MODEL (patched)
Z-Image-Turbo-Fun-Controlnet-Union

A unified ControlNet supporting multiple condition types:

  • Canny, HED, Depth, Pose, MLSD
  • Strength: 0.65 to 0.80 (v2.1 recommended range)
  • Best paired with res_2s, res_5s, or res_2m samplers + beta57 scheduler

Complete Workflow: RedCraft DX1 (Fast, 10-Step)

json
{
  "1": { "class_type": "CheckpointLoaderSimple", "inputs": { "ckpt_name": "redcraftRedzimageUpdatedJAN30_redzibDX1.safetensors" }},
  "2": { "class_type": "CLIPTextEncode", "inputs": { "clip": ["1", 1], "text": "<positive prompt>" }, "_meta": { "title": "Positive" }},
  "3": { "class_type": "CLIPTextEncode", "inputs": { "clip": ["1", 1], "text": "" }, "_meta": { "title": "Negative" }},
  "4": { "class_type": "EmptyLatentImage", "inputs": { "width": 1024, "height": 1024, "batch_size": 1 }},
  "5": { "class_type": "KSampler", "inputs": {
    "model": ["1", 0],
    "positive": ["2", 0],
    "negative": ["3", 0],
    "latent_image": ["4", 0],
    "seed": 42, "steps": 10, "cfg": 1, "sampler_name": "euler", "scheduler": "simple", "denoise": 1
  }},
  "6": { "class_type": "VAEDecode", "inputs": { "samples": ["5", 0], "vae": ["1", 2] }},
  "7": { "class_type": "SaveImage", "inputs": { "images": ["6", 0], "filename_prefix": "redcraft" }}
}

Complete Workflow: RedCraft DX1 with LoRA Stack

json
{
  "1": { "class_type": "CheckpointLoaderSimple", "inputs": { "ckpt_name": "redcraftRedzimageUpdatedJAN30_redzibDX1.safetensors" }},
  "2": { "class_type": "LoraLoader", "inputs": {
    "model": ["1", 0], "clip": ["1", 1],
    "lora_name": "Z-Image-Aesthetic-Base v1.safetensors",
    "strength_model": 0.8, "strength_clip": 0.8
  }},
  "3": { "class_type": "LoraLoader", "inputs": {
    "model": ["2", 0], "clip": ["2", 1],
    "lora_name": "ZImageBase\\style\\NSGIRL-Z-Image-LoRA-By-MM744.safetensors",
    "strength_model": 0.7, "strength_clip": 0.7
  }},
  "4": { "class_type": "CLIPTextEncode", "inputs": { "clip": ["3", 1], "text": "<positive prompt>" }},
  "5": { "class_type": "CLIPTextEncode", "inputs": { "clip": ["3", 1], "text": "<negative prompt>" }},
  "6": { "class_type": "EmptyLatentImage", "inputs": { "width": 896, "height": 1152, "batch_size": 1 }},
  "7": { "class_type": "KSampler", "inputs": {
    "model": ["3", 0],
    "positive": ["4", 0],
    "negative": ["5", 0],
    "latent_image": ["6", 0],
    "seed": 42, "steps": 30, "cfg": 4, "sampler_name": "euler", "scheduler": "simple", "denoise": 1
  }},
  "8": { "class_type": "VAEDecode", "inputs": { "samples": ["7", 0], "vae": ["1", 2] }},
  "9": { "class_type": "SaveImage", "inputs": { "images": ["8", 0], "filename_prefix": "redcraft_lora" }}
}

Prompt Style

Natural language descriptions work best (uses Qwen LLM tokenizer, not CLIP):

Good: "Professional headshot of a confident businesswoman in her 30s, natural makeup, soft studio lighting, neutral gray background, sharp focus on eyes, Canon EOS R5"
Bad: "masterpiece, best quality, 1girl, businesswoman, studio"

VRAM Considerations

ConfigVRAMNotes
RedCraft DX1 checkpoint~17GBFits comfortably on RTX 4090
Z-Image Turbo separate~8GB UNET + CLIPVery lightweight
Z-Image Base separate~12GB
  • Always clear_vram before switching to Z-Image from another model family
  • RedCraft is one of the most VRAM-efficient quality models available

Tips

  1. RedCraft DX1 with 10 steps / CFG 1.0 is fast and high quality for quick iteration
  2. For maximum sharpness with Turbo LoRAs, use dpmpp_sde + beta scheduler
  3. The Z-Image-Aesthetic-Base v1 LoRA at 0.6 to 0.8 strength improves output quality across all Z-Image Base variants
  4. Z-Image is strong at photorealistic human generation and is the go-to for portrait and fashion photography
  5. When switching between Turbo and Base LoRAs, use the matching base model variant

Sources

  • Official: none found.
  • Empirical: sampler values, wiring, and prompt notes from working graphs in packs/ and observed renders; not a vendor prompting guide.

© artokun, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in plugin/skills/z-image-txt2img of artokun/comfyui-mcp.

Open the folder on GitHubat commit 6ad6fc0

Compare with similar skills

Z Image Txt2img next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Z Image Txt2img compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Z Image Txt2img this skillartokun/comfyui-mcp803—~2.8kAutomated safety check: PassMIT
Continuity Renderroadmaus/ComfyUI-Continuity133—~1.7kAutomated safety check: PassMIT
Setupguaardvark/guaardvark258—~1.2kAutomated safety check: PassMIT
Comfyuicalesthio/OpenMontage66k—~2kAutomated safety check: PassAGPL-3.0
Workflow Template BuilderMooshieblob1/MooshieUI207—~640Automated safety check: PassAGPL-3.0
Flux2 Lora TrainingAnastasiyaW/codex-claude-code-config154—~4.5kAutomated safety check: PassMIT

Similar skills

  • Continuity Render

    roadmaus/ComfyUI-Continuity

    Render videos and pictures on a ComfyUI server that has the Continuity node pack (MiniMax H3, LTX 2.5, Krea 2, Ideogram 4, Qwen Image, Flux 2 Klein) with one command, the render.py bundled in this…

    133 GitHub stars~1.7k tokensUpdated 3 days ago
    AI & LLM EngineeringAuto-check passed
  • Setup

    guaardvark/guaardvark

    Connect this agent to a running Guaardvark (self-hosted AI studio) and check what it can do right now.

    258 GitHub stars~1.2k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Comfyui

    calesthio/OpenMontage

    A skill your agent uses when working with ComfyUI workflows in OpenMontage, including comfyuiimage/comfyuivideo/comfyuimusic, custom workflowjson/workflowpath inputs, outputnode selection, missing…

    66k GitHub stars~2k tokensUpdated 7 days ago
    AI & LLM EngineeringAuto-check passed
  • Workflow Template Builder

    Mooshieblob1/MooshieUI

    Builds or modifies ComfyUI workflow JSON templates in MooshieUI's Rust backend (src-tauri/src/templates).

    207 GitHub stars~640 tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Flux2 Lora Training

    AnastasiyaW/codex-claude-code-config

    Plan or review LoRA and edit-training work specifically for FLUX.2 Klein or Qwen-Image-Edit, including paired datasets, trainer-version contracts, and held-out fidelity checks.

    154 GitHub stars~4.5k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Diffusion Engineering

    AnastasiyaW/codex-claude-code-config

    Практическая инженерия диффузионных моделей: архитектуры, обучение, инференс, оптимизация памяти.

    154 GitHub stars~1.2k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed

More from artokun/comfyui-mcp

All 42 skills in this repo
  • AI Toolkit Trainer

    artokun/comfyui-mcp

    Train custom LoRAs with ostris AI-Toolkit. An agent skill from artokun/comfyui-mcp.

    803 GitHub stars~2.7k tokensUpdated 6 days ago
    Auto-check passed
  • Anima Base

    artokun/comfyui-mcp

    Anime/illustration text-to-image (ANIMA 1.0, ~2B Cosmos DiT).

    803 GitHub stars~4k tokensUpdated 6 days ago
    Auto-check passed
  • Civitai

    artokun/comfyui-mcp

    Discover Civitai models with the BUILT-IN downloadmodel action:"searchcivitai" and install/generate them locally.

    803 GitHub stars~1.1k tokensUpdated 6 days ago
    Auto-check passed
  • Color Correction

    artokun/comfyui-mcp

    Diagnose and fix video/image color OBJECTIVELY with the getimage (action:"analyzecolor") tool (scopes/stats such as black/white points, contrast, saturation, clipping, cast) instead of eyeballing a…

    803 GitHub stars~2.4k tokensUpdated 6 days ago
    Auto-check passed
  • Comfyui Frontend Extensions

    artokun/comfyui-mcp

    Authoring ComfyUI v2 frontend extensions with @comfyorg/extension-api, covering defineNode/defineExtension/defineWidget, shell UI (sidebar tabs, commands, hotkeys), typed events, and handles.

    803 GitHub stars~5.4k tokensUpdated 6 days ago
    Auto-check passed
  • Comfyui Launch Flags

    artokun/comfyui-mcp

    Pick the right ComfyUI startup flags for VRAM, attention, caching, and speed.

    803 GitHub stars~3.1k tokensUpdated 6 days ago
    Auto-check passed

Works with

Questions about Z Image Txt2img

What does Z Image Txt2img do?

Build Z-Image txt2img workflows. An agent skill from artokun/comfyui-mcp. Z Image Txt2img is an agent skill from artokun/comfyui-mcp. Build Z-Image txt2img workflows.

When should I use Z Image Txt2img?

Z Image Txt2img fits situations like: tasks that involve Diffusion and image models; tasks that involve Fine-tuning.

How do I install Z Image Txt2img in Claude Code?

Run `npx skills add artokun/comfyui-mcp --skill z-image-txt2img -a claude-code`. Or copy the skill folder (plugin/skills/z-image-txt2img in artokun/comfyui-mcp) into .claude/skills/z-image-txt2img in your project. Claude Code loads it when a task matches its description.

How do I install Z Image Txt2img in Codex?

Run `npx skills add artokun/comfyui-mcp --skill z-image-txt2img -a codex`. Or copy the skill folder (plugin/skills/z-image-txt2img in artokun/comfyui-mcp) into .agents/skills/z-image-txt2img in your project. Codex loads it when a task matches its description.

Can I use Z Image Txt2img in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add artokun/comfyui-mcp --skill z-image-txt2img -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/z-image-txt2img, .gemini/skills/z-image-txt2img, .github/skills/z-image-txt2img and .opencode/skills/z-image-txt2img in your project.

What does Z Image Txt2img need to run?

SKILL.md names no scripts, command-line tools or credentials: Z Image Txt2img is instructions for the agent only.

Does Z Image Txt2img access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Z Image Txt2img safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Z Image Txt2img use?

Z Image Txt2img is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Z Image Txt2img use?

About 2.8k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Z Image Txt2img?

Skills that share tags, products or a category with Z Image Txt2img: Continuity Render (roadmaus/ComfyUI-Continuity, 133 stars), Setup (guaardvark/guaardvark, 258 stars), Comfyui (calesthio/OpenMontage, 66k stars) and Workflow Template Builder (Mooshieblob1/MooshieUI, 207 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Z Image Txt2img?

artokun (a GitHub user) maintains it in artokun/comfyui-mcp, which has 803 GitHub stars. The repository holds 42 skills in this directory. The repository was last updated on October 5, 2026.

Source: artokun/comfyui-mcp on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.