Agent skill

Ernie Image

by artokun in artokun/comfyui-mcp

Build Baidu ERNIE-Image / ERNIE-Image-Turbo workflows, primarily TEXT-TO-IMAGE.

MITAuto-check passedMedia & Creative

Install Ernie Image

skills CLI
$ npx skills add artokun/comfyui-mcp --skill ernie-image -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install artokun/comfyui-mcp ernie-image --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/artokun/comfyui-mcp.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugin/skills/ernie-image .claude/skills/ernie-image && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
ernie-image
GitHub stars
803
Token cost
~4.5k tokens
SKILL.md length
1,657 words
Files
1
Skills in repo
42
Repo updated
First seen
Licence
MIT

At a glance

Build Baidu ERNIE-Image / ERNIE-Image-Turbo workflows, primarily TEXT-TO-IMAGE.

  • Works in 5 steps: Lead with the literal text you want… → Use the prompt enhancer for short/lazy… → Use get_workflow (action:"analyze")… → …
  • Tasks that involve Diffusion and image models
  • SKILL.md covers What this is (read first), Separated packs…, Source of truth & a provenance… and Models, plus 9 more sections
  • Calls hf; reaches github.com and huggingface.co

What it does

Ernie Image is an agent skill from artokun/comfyui-mcp. Build Baidu ERNIE-Image / ERNIE-Image-Turbo workflows, primarily TEXT-TO-IMAGE. Pick ERNIE when you need precise multilingual text rendering, posters/signage, manga/anime multi-panel layouts, or strong instruction following for complex multi-object scenes. Also supports denoise-based image-to-image refine (NOT instruction-grounded editing; use Qwen-Image-Edit or Flux Kontext for "change X in this photo" edits).

Its SKILL.md is about 4.5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Media & Creative, covering Diffusion and image models, Comics and storyboards and Image generation. It works with Qwen. The repository describes itself as: Local-first, agent-native control plane for ComfyUI — MCP server + sidebar agent that generates images, video & audio, authors and runs workflows, and edits your live graph in… The licence is MIT.

When your agent uses it

  • Tasks that involve Diffusion and image models
  • Tasks that involve Comics and storyboards
  • Tasks that involve Image generation

Example prompts

  • “change X in this photo”
  • “/ernie-image”

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Lead with the literal text you want rendered, in quotes. That's ERNIE's headline strength.
  2. Use the prompt enhancer for short/lazy prompts; turn it off (ComfySwitchNode false) when you've written a detailed prompt yourself.
  3. Use get_workflow (action:"analyze") before executing the shipped graph. It has dozens of group-boxed variants gated by Fast Groups…
  4. Most groups are bypassed (mode 4) by default in the source file. Enable only the pipeline you want via the group bypasser, or build the…
  5. To choose a model: ERNIE for text/layout T2I, Z-Image Turbo for fast general T2I, Qwen-Image-Edit / Flux Kontext for actual editing.

What it can do on your machine

Read from SKILL.md and the folder at commit 6ad6fc0. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • hf

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • github.com
    • huggingface.co

    Also links to:

    • docs.comfy.org

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Ernie Image loads about 4.5k tokens when it runs. Until then it costs about 107 tokens; SKILL.md has 1,657 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~107
When it runs · the whole SKILL.md, loaded when a task matches
~4.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from artokun/comfyui-mcp at commit 6ad6fc0, republished under its MIT licence (© artokun). 1,657 words, ~4,545 tokens.

Download SKILL.mdSave it as .claude/skills/ernie-image/SKILL.md (or your agent's skills folder).
name
ernie-image
description
Build Baidu ERNIE-Image / ERNIE-Image-Turbo workflows, primarily TEXT-TO-IMAGE. Pick ERNIE when you need precise multilingual text rendering, posters/signage, manga/anime multi-panel layouts, or strong instruction following for complex multi-object scenes. Also supports denoise-based image-to-image refine (NOT instruction-grounded editing; use Qwen-Image-Edit or Flux Kontext for "change X in this photo" edits).
globs
**/*.json

ERNIE-Image / ERNIE-Image-Turbo Workflows

What this is (read first)

ERNIE-Image is Baidu's open-weight TEXT-TO-IMAGE model, an ~8B single-stream Diffusion Transformer (DiT), Apache-2.0, released April 2026 and repackaged for ComfyUI by Comfy-Org. It is not an instruction-based image editor.

  • ERNIE-Image (base): ~50 steps for peak quality.
  • ERNIE-Image-Turbo: distilled (Distribution Matching Distillation + RL), high-fidelity in ~8 steps at cfg 1. The downloaded pack uses Turbo (ernie-image-turbo-*.gguf).

Pick ERNIE when the job is precise text/typography rendering (multilingual, including Chinese), posters/signage/UI mockups, manga/anime storyboards and multi-panel layouts, or structured multi-object scenes from a complex prompt. Do not pick ERNIE for "edit this photo / change the shirt / swap the background". That is instruction-grounded editing, which ERNIE does not do. Use qwen-image-edit or Flux Kontext for those. ERNIE's "image-to-image" here is plain denoise-based refinement (style pass / detail pass), not reference-grounded editing.

Niche vs siblings. ERNIE is the best open-weight text rendering + layout T2I. Qwen-Image-Edit does instruction editing. Flux Kontext does reference editing. Z-Image Turbo does fast general T2I, and this same pack pairs the two; see Combo pipelines.

Separated packs (render-verified)

The original ernie monolith was a single toggle-template graph (every pipeline shipped bypassed; you activated one via the rgthree group toggles). It's now split into standalone, single-purpose packs, each a clean activated graph that renders headlessly with no group-toggling:

PackUseModelsVRAM
ernie-txt2imgtext-to-image (flagship)ERNIE only (4)<8GB
ernie-img2imgdenoise refine of a source imageERNIE only (4)<8GB
ernie-comboERNIE × Z-Image-Turbo combo pipelinesERNIE + Z-Image (7, ~32GB)12GB+

Working details verified live: the prompt-enhancer LLM is OFF by default (the ENHANCE PROMPT boolean is false; leave it off unless you want the 3B enhancer to rewrite the prompt). The grain/sharpen post-proc (FastFilmGrain/FastLaplacianSharpen, comfyui-vrgamedevgirl) needs librosa installed. In ernie-combo the Z-Image half's VAE is saved as z-image-ae.safetensors; its weights differ from Flux/ERNIE's ae.safetensors despite the same size, and the rename avoids a filename clash.

Source of truth & a provenance warning

This skill is derived from the actual pack files in C:\Users\Artokun\Downloads\:

  • ERNIE-IMAGE-ULTRA-WORKFLOW.json (authoritative; the ComfyUI graph)
  • ERNIE-IMAGE_ULTRA-MODELS-NODES_INSTALL.bat, ...-COMFYUI-MANAGER_AUTO_INSTALL.bat, ...-AUTO_INSTALL-RUNPOD.sh

Installer warning (verified). The three install scripts are copy-pasted from a Z-Image pack. Their headers literally say "Z-IMAGE-BASE"/"Z-IMAGE Base", and they download both ERNIE and Z-Image files. The model URLs/folders below are taken from those scripts but mirror this confusion. They pull z_image_turbo-*.gguf, Qwen3-4B-*.gguf, and ae.safetensors, which belong to the Z-Image half of the combo, not ERNIE. The ERNIE-only files are flagged below. All weights come from a third-party mirror huggingface.co/Aitrepreneur/FLX, not the official huggingface.co/Comfy-Org/ERNIE-Image (which hosts the same filenames; see Official sources).

Models

ERNIE-Image (the files ERNIE actually uses)

Confirmed from the workflow's virtual wires (Set_*/GetNode): the nodes tagged "ERNIE" resolve to these exact files.

ComponentNode (type)File (in workflow)FolderNotes
UNet (GGUF)UnetLoaderGGUFernie-image-turbo-Q8_0.ggufmodels/unet/Turbo DiT. Q5_K_S / Q6_K / Q8_0 quants offered by installer
Text encoderCLIPLoader (type=flux2)ministral-3-3b.safetensorsmodels/text_encoders/Ministral-3-3B is ERNIE's text encoder. Loaded with CLIP type flux2
VAEVAELoaderflux2-vae.safetensorsmodels/vae/ERNIE reuses the Flux 2 VAE
Prompt enhancerCLIPLoader (type=flux2) → TextGenerateernie-image-prompt-enhancer.safetensorsmodels/text_encoders/3B LLM that auto-expands a short prompt into a rich description (see Prompt enhancer). Optional, toggled per-pipeline

Quant guidance from the installer: Q5_K_S for GPUs <8 GB · Q6_K for 8 to 12 GB · Q8_0 for 12 to 16 GB+.

Z-Image Turbo (bundled in the same pack — the "ZIT" half)

The workflow also wires a parallel Z-Image Turbo pipeline for ERNIE→ZIT / ZIT→ERNIE combos. These files are Z-Image's, not ERNIE's. Do not confuse them:

ComponentNodeFileFolder
UNet (GGUF)UnetLoaderGGUFz_image_turbo-Q8_0.ggufmodels/unet/
Text encoderCLIPLoaderGGUF (type=lumina2)Qwen3-4B-UD-Q6_K_XL.ggufmodels/text_encoders/
VAEVAELoaderae.safetensorsmodels/vae/
LoRAs (referenced in the Power Lora Loader, off by default)

hirohiko-araki-style-ERNIE_000001250.safetensors and ernie-anime-v1.safetensors are community ERNIE style LoRAs, loaded via Power Lora Loader (rgthree) (both toggled off in the shipped graph). Not in the installer; user-supplied.

Upscalers / post (shared)

4x-ClearRealityV1.pth, RealESRGAN_x4plus_anime_6B.pth → models/upscale_models/.

Installation

Custom nodes (git clone into ComfyUI/custom_nodes/)

All three installers clone the same set:

Node packRepoWhy it's needed
ComfyUI-Managerhttps://github.com/ltdrdata/ComfyUI-Manager.gitmanagement
ComfyUI-GGUFhttps://github.com/city96/ComfyUI-GGUFUnetLoaderGGUF, CLIPLoaderGGUF
rgthree-comfyhttps://github.com/rgthree/rgthree-comfyPower Lora Loader, Label, Fast Groups Bypasser, Image Comparer
ComfyUI-Easy-Usehttps://github.com/yolain/ComfyUI-Easy-Useeasy cleanGpuUsed, easy clearCacheAll
ComfyUI-KJNodeshttps://github.com/kijai/ComfyUI-KJNodesutility nodes
ComfyUI_essentialshttps://github.com/cubiq/ComfyUI_essentialsImageResize+
wlsh_nodeshttps://github.com/wallish77/wlsh_nodesUpscale by Factor with Model (WLSH)
comfyui-vrgamedevgirlhttps://github.com/vrgamegirl19/comfyui-vrgamedevgirlFastFilmGrain, FastLaplacianSharpen
RES4LYFhttps://github.com/ClownsharkBatwing/RES4LYFadvanced samplers

The graph also uses TextGenerate, TextBox1, StringReplace, ComfySwitchNode, PreviewAny, SetNode/GetNode, PrimitiveBoolean, ModelSamplingAuraFlow, ConditioningZeroOut, EmptySD3LatentImage, EmptyFlux2LatentImage. Most are builtin or come from the packs above. SetNode/GetNode are from KJNodes. TextGenerate (runs the prompt-enhancer LLM) has an unverified pack origin; verify which pack provides it via ComfyUI-Manager if it shows as missing.

Model downloads (exact URLs from the installer)

Base URL HF = https://huggingface.co/Aitrepreneur/FLX/resolve/main (third-party mirror). !MODEL_VERSION! ∈ {Q5_K_S, Q6_K, Q8_0}.

# ERNIE (the files ERNIE actually uses)
unet/ernie-image-turbo-<Q>.gguf                <HF>/ernie-image-turbo-<Q>.gguf?download=true
text_encoders/ministral-3-3b.safetensors       <HF>/ministral-3-3b.safetensors?download=true
text_encoders/ernie-image-prompt-enhancer.safetensors  <HF>/ernie-image-prompt-enhancer.safetensors?download=true
vae/flux2-vae.safetensors                       <HF>/flux2-vae.safetensors?download=true

# Z-Image half (bundled; only needed for the ZIT combo pipelines)
unet/z_image_turbo-<Q>.gguf                      <HF>/z_image_turbo-<Q>.gguf?download=true
text_encoders/Qwen3-4B-UD-Q6_K_XL.gguf          <HF>/Qwen3-4B-UD-Q6_K_XL.gguf?download=true
vae/ae.safetensors                               <HF>/ae.safetensors?download=true

# Upscalers
upscale_models/4x-ClearRealityV1.pth            <HF>/4x-ClearRealityV1.pth?download=true
upscale_models/RealESRGAN_x4plus_anime_6B.pth   <HF>/RealESRGAN_x4plus_anime_6B.pth?download=true
Official sources (prefer these over the mirror)

The same filenames are hosted officially at huggingface.co/Comfy-Org/ERNIE-Image (unet|diffusion_models/, text_encoders/, vae/). Apache-2.0. Original model: github.com/baidu/ERNIE-Image. Comfy day-0 docs: docs.comfy.org/tutorials/image/ernie-image/ernie-image. The official repo also ships non-GGUF ernie-image.safetensors / ernie-image-turbo.safetensors (load with UNETLoader instead of UnetLoaderGGUF).

How the pipeline works (at a glance)

The big graph is a menu of group-boxed pipelines built from the same blocks. Core ERNIE-Image text-to-image flow:

UnetLoaderGGUF (ernie-image-turbo) ──► Power Lora Loader (rgthree) ──► ModelSamplingAuraFlow (shift=3.1) ──► MODEL
CLIPLoader (ministral-3-3b, type=flux2) ──► CLIP ──► CLIPTextEncode (positive)
                                              └──► ConditioningZeroOut  ──► negative   (cfg=1, so negative ≈ unused)
VAELoader (flux2-vae) ──► VAE
EmptySD3LatentImage (1920×1088) ──► LATENT
        │
KSampler (steps≈8–9, cfg=1, euler, simple, denoise=1) ──► VAEDecode ──► SaveImage / post

In the optional prompt enhancer path, the short user prompt + {width}/{height} are templated into a Chinese system prompt, fed to TextGenerate (which runs ernie-image-prompt-enhancer), and a ComfySwitchNode chooses raw prompt (switch=false) vs. enhanced prompt (switch=true) before CLIPTextEncode.

The image-to-image (refine) path is not editing. The graph's "ERNIE IMAGE TO IMAGE" groups take a LoadImage → ImageResize+ (1024, keep proportion, lanczos) and VAEEncode it, then run KSampler at low denoise (0.35 to 0.4) to refine or restyle. This is a denoise pass over a single source image; it does not follow edit instructions.

About the ~5 LoadImage + ~5 VAEEncode nodes: they are not multi-reference compositing. Each LoadImage feeds a separate pipeline variant (single-image img2img, or the combo refine stages). One source image per pipeline. The extra VAEEncodes are the encode steps for those independent img2img / two-pass refine chains.

The combo pipelines (ERNIE↔ZIT), with group titles ERNIE ---> ZIT COMBO, ZIT ---> ERNIE COMBO, TWO TIMES COMBO ..., chain ERNIE and Z-Image Turbo as a two-pass generate→refine, with optional film-grain (FastFilmGrain) or sharpening (FastLaplacianSharpen) finishing and SIMPLE UPSCALE.

Show full SKILL.md (657 more words)Show less

Settings (extracted from the shipped KSamplers)

PipelineStepsCFGSamplerSchedulerDenoiseShift
ERNIE text-to-image (Turbo)8–91eulersimple1.03.1
ERNIE image-to-image refine81eulersimple0.43.1
Combo refine pass (2nd stage)91eulersimple0.25–0.353.1
  • ModelSamplingAuraFlow shift = 3.1 is applied to the ERNIE model before sampling (flow-matching shift). Keep it.
  • CFG = 1 for Turbo, so negative conditioning is effectively inert; the graph still wires a ConditioningZeroOut as the negative.
  • The shipped latent is 1920×1088 (EmptySD3LatentImage). ERNIE is a high-res-capable DiT; 1024 to 2048 on the long edge is reasonable. Use EmptySD3LatentImage for ERNIE latents.
  • For base (non-Turbo) ernie-image, bump steps to ~50 and raise cfg (e.g. 3.5 to 5) since it is not distilled.

Prompt / instruction style

ERNIE rewards descriptive, structured natural-language prompts, and is unusually strong at literal text rendering. Write the exact text you want to appear in quotes.

A vintage travel poster of Kyoto in autumn, bold title text reading "KYOTO" at the top,
maple leaves, Mount fuji silhouette, clean vector layout, muted warm palette
A 3-panel manga page: panel 1 a samurai drawing his sword, panel 2 close-up of his eyes,
panel 3 a wide shot of cherry blossoms falling, black-and-white ink, speech bubble "参る"
  • For typography/signage, state the literal string ("a neon sign that says 'OPEN'"), placement, and font feel.
  • For layout, name the panel/grid structure and what goes in each region.
  • Multilingual prompts (incl. Chinese) work; the built-in enhancer's system prompt is Chinese.
  • This is txt2img phrasing, not edit phrasing. Do not write "change the…/remove the…" expecting grounded edits.

Complete API-format workflow (ERNIE-Image-Turbo text-to-image)

Derived from the source graph, flattened to API format (no subgraphs/virtual wires). The enhancer is omitted for clarity; CLIPTextEncode takes the prompt directly.

json
{
  "1": { "class_type": "UnetLoaderGGUF", "inputs": { "unet_name": "ernie-image-turbo-Q8_0.gguf" } },
  "2": { "class_type": "CLIPLoader", "inputs": { "clip_name": "ministral-3-3b.safetensors", "type": "flux2", "device": "default" } },
  "3": { "class_type": "VAELoader", "inputs": { "vae_name": "flux2-vae.safetensors" } },
  "4": { "class_type": "ModelSamplingAuraFlow", "inputs": { "model": ["1", 0], "shift": 3.1 } },
  "5": { "class_type": "CLIPTextEncode", "inputs": { "clip": ["2", 0], "text": "A vintage travel poster of Kyoto in autumn, bold title text reading \"KYOTO\" at the top, maple leaves, clean vector layout, muted warm palette" } },
  "6": { "class_type": "ConditioningZeroOut", "inputs": { "conditioning": ["5", 0] } },
  "7": { "class_type": "EmptySD3LatentImage", "inputs": { "width": 1920, "height": 1088, "batch_size": 1 } },
  "8": { "class_type": "KSampler", "inputs": {
    "model": ["4", 0], "positive": ["5", 0], "negative": ["6", 0], "latent_image": ["7", 0],
    "seed": 997032332094579, "steps": 9, "cfg": 1, "sampler_name": "euler", "scheduler": "simple", "denoise": 1
  } },
  "9": { "class_type": "VAEDecode", "inputs": { "samples": ["8", 0], "vae": ["3", 0] } },
  "10": { "class_type": "SaveImage", "inputs": { "images": ["9", 0], "filename_prefix": "ernie_image" } }
}
Image-to-image (refine) variant

Replace the empty latent with an encoded source image and lower denoise. This restyles or refines a single image; it is not instruction editing.

json
{
  "11": { "class_type": "LoadImage", "inputs": { "image": "source.png" } },
  "12": { "class_type": "ImageResize+", "inputs": { "image": ["11", 0], "width": 1024, "height": 1024, "interpolation": "lanczos", "method": "keep proportion", "condition": "always", "multiple_of": 0 } },
  "13": { "class_type": "VAEEncode", "inputs": { "pixels": ["12", 0], "vae": ["3", 0] } }
}

Then in the KSampler set "latent_image": ["13", 0] and "denoise": 0.4.

Adding LoRAs

Insert a Power Lora Loader (rgthree) between the UNet loader and ModelSamplingAuraFlow (model: ["1",0] → loader → ["4"].model). In API format you can substitute LoraLoaderModelOnly with lora_name: "ernie-anime-v1.safetensors", strength_model: 0.5.

VRAM

  • ERNIE-Image-Turbo GGUF: Q5_K_S for <8 GB · Q6_K for 8 to 12 GB · Q8_0 for 12 to 16 GB+ (installer's own guidance).
  • The Ministral-3-3B encoder + Flux2 VAE add a few GB. The graph includes easy cleanGpuUsed / easy clearCacheAll nodes between stages. Keep them for the combo/two-pass pipelines so VRAM is freed before swapping models.
  • Running the ERNIE↔ZIT combos loads two UNets; budget for both or run the single-model ERNIE group only.

Troubleshooting

  • UnetLoaderGGUF / CLIPLoaderGGUF missing. Install ComfyUI-GGUF (city96).
  • Power Lora Loader / Image Comparer / Label missing. Install rgthree-comfy.
  • ImageResize+ missing. Install ComfyUI_essentials.
  • TextGenerate missing (prompt enhancer). Install via ComfyUI-Manager search; pack origin unverified. If unavailable, set the ComfySwitchNode to use the raw prompt (switch=false) and skip enhancement.
  • CLIP type error on ministral. Ensure CLIPLoader type is flux2 (not qwen_image/lumina2). The lumina2 type belongs to the Z-Image (Qwen3) encoder, not ERNIE.
  • Wrong VAE artifacts. ERNIE must use flux2-vae.safetensors; ae.safetensors is the Z-Image VAE.
  • Blurry / undercooked output. Confirm ModelSamplingAuraFlow shift=3.1 is wired and steps ≥8 for Turbo; for base ernie-image use ~50 steps + higher cfg.
  • You wanted to EDIT a photo and it ignored the instruction. Expected. ERNIE is txt2img; use the qwen-image-edit skill or Flux Kontext for grounded edits.
  • Installer pulled Z-Image files too. Expected (the scripts are Z-Image-derived). Harmless; those files only feed the combo pipelines.

Tips

  1. Lead with the literal text you want rendered, in quotes. That's ERNIE's headline strength.
  2. Use the prompt enhancer for short/lazy prompts; turn it off (ComfySwitchNode false) when you've written a detailed prompt yourself.
  3. Use get_workflow (action:"analyze") before executing the shipped graph. It has dozens of group-boxed variants gated by Fast Groups Bypasser (rgthree); the analyzer summary is far easier than reading raw JSON.
  4. Most groups are bypassed (mode 4) by default in the source file. Enable only the pipeline you want via the group bypasser, or build the clean API workflow above.
  5. To choose a model: ERNIE for text/layout T2I, Z-Image Turbo for fast general T2I, Qwen-Image-Edit / Flux Kontext for actual editing.

Sources

© artokun, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in plugin/skills/ernie-image of artokun/comfyui-mcp.

Open the folder on GitHubat commit 6ad6fc0

Compare with similar skills

Ernie Image next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Ernie Image compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Ernie Image this skillartokun/comfyui-mcp803—~4.5kAutomated safety check: PassMIT
Flux2 Klein PromptingAnastasiyaW/codex-claude-code-config154—~2.8kAutomated safety check: PassMIT
Character Refseternityspring/shuohao-skills4.3k—~1.7kAutomated safety check: WarnApache-2.0
Qwen Image 2 1 Prompteriamyoki/qwen-image-2.1-skill157—~1.6kAutomated safety check: PassApache-2.0
Image Genopen-octo/octo-agent125—~3.1kAutomated safety check: NotesMIT
LoRA Space Builderhuggingface/skills11k2 repos~8.4kAutomated safety check: PassApache-2.0

Similar skills

  • Flux2 Klein Prompting

    AnastasiyaW/codex-claude-code-config

    Expert prompt engineering for FLUX.2 [klein] image generation and editing model.

    154 GitHub stars~2.8k tokensUpdated yesterday
    Media & CreativeAuto-check passed
  • Character Refs

    eternityspring/shuohao-skills

    给任何故事里的角色真出参考图(小说改编、自己原创的故事、单独设计一个角色都行,不需要小说原文): 一段话描述角色,拆成分层字段、补全后确认, 先出一张正面全身锚点,其余视图(大头照、90° 侧面、背面、细节、45° 大头照)都只参考这张锚点, 按需分档出图。每张图带标识、可单独重出,重出后自动标出哪些图过期。

    4.3k GitHub stars~1.7k tokensUpdated 2 days ago
    Media & CreativeAuto-check: warnings
  • Qwen Image 2 1 Prompter

    iamyoki/qwen-image-2.1-skill

    Optimize, rewrite, and craft image generation and editing prompts tailored specifically for Alibaba's Qwen-Image-2.1 diffusion model.

    157 GitHub stars~1.6k tokensUpdated 18 days ago
    Media & CreativeAuto-check passed
  • Image Gen

    open-octo/octo-agent

    Acquire images as files — generate them with an AI image model (14 providers: OpenAI/gpt-image, Gemini, Qwen, Zhipu, Volcengine, Stability, FLUX, Ideogram, MiniMax, and more), search openly-licensed…

    125 GitHub stars~3.1k tokensUpdated today
    Media & CreativeAuto-check: notes
  • LoRA Space Builder

    huggingface/skills

    Official

    Builds and publishes a Gradio demo on Hugging Face Spaces for a LoRA, with the pipeline, UI and settings chosen to match that LoRA's task and model card.

    11k GitHub starsUsed in 2 repos~8.4k tokens
    AI & LLM EngineeringAuto-check passed
  • Setup

    guaardvark/guaardvark

    Connect this agent to a running Guaardvark (self-hosted AI studio) and check what it can do right now.

    257 GitHub stars~1.2k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed

More from artokun/comfyui-mcp

All 42 skills in this repo
  • AI Toolkit Trainer

    artokun/comfyui-mcp

    Train custom LoRAs with ostris AI-Toolkit. An agent skill from artokun/comfyui-mcp.

    803 GitHub stars~2.7k tokensUpdated 5 days ago
    Auto-check passed
  • Anima Base

    artokun/comfyui-mcp

    Anime/illustration text-to-image (ANIMA 1.0, ~2B Cosmos DiT).

    803 GitHub stars~4k tokensUpdated 5 days ago
    Auto-check passed
  • Civitai

    artokun/comfyui-mcp

    Discover Civitai models with the BUILT-IN downloadmodel action:"searchcivitai" and install/generate them locally.

    803 GitHub stars~1.1k tokensUpdated 5 days ago
    Auto-check passed
  • Color Correction

    artokun/comfyui-mcp

    Diagnose and fix video/image color OBJECTIVELY with the getimage (action:"analyzecolor") tool (scopes/stats such as black/white points, contrast, saturation, clipping, cast) instead of eyeballing a…

    803 GitHub stars~2.4k tokensUpdated 5 days ago
    Auto-check passed
  • Comfyui Frontend Extensions

    artokun/comfyui-mcp

    Authoring ComfyUI v2 frontend extensions with @comfyorg/extension-api, covering defineNode/defineExtension/defineWidget, shell UI (sidebar tabs, commands, hotkeys), typed events, and handles.

    803 GitHub stars~5.4k tokensUpdated 5 days ago
    Auto-check passed
  • Comfyui Launch Flags

    artokun/comfyui-mcp

    Pick the right ComfyUI startup flags for VRAM, attention, caching, and speed.

    803 GitHub stars~3.1k tokensUpdated 5 days ago
    Auto-check passed

Works with

Questions about Ernie Image

What does Ernie Image do?

Build Baidu ERNIE-Image / ERNIE-Image-Turbo workflows, primarily TEXT-TO-IMAGE. Ernie Image is an agent skill from artokun/comfyui-mcp. Build Baidu ERNIE-Image / ERNIE-Image-Turbo workflows, primarily TEXT-TO-IMAGE.

When should I use Ernie Image?

Ernie Image fits situations like: tasks that involve Diffusion and image models; tasks that involve Comics and storyboards; tasks that involve Image generation.

How do I install Ernie Image in Claude Code?

Run `npx skills add artokun/comfyui-mcp --skill ernie-image -a claude-code`. Or copy the skill folder (plugin/skills/ernie-image in artokun/comfyui-mcp) into .claude/skills/ernie-image in your project. Claude Code loads it when a task matches its description.

How do I install Ernie Image in Codex?

Run `npx skills add artokun/comfyui-mcp --skill ernie-image -a codex`. Or copy the skill folder (plugin/skills/ernie-image in artokun/comfyui-mcp) into .agents/skills/ernie-image in your project. Codex loads it when a task matches its description.

Can I use Ernie Image in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add artokun/comfyui-mcp --skill ernie-image -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ernie-image, .gemini/skills/ernie-image, .github/skills/ernie-image and .opencode/skills/ernie-image in your project.

What does Ernie Image need to run?

Going by SKILL.md and its folder, Ernie Image needs the command-line tools its instructions call (hf).

Does Ernie Image access the network?

SKILL.md names 3 domains. In commands or code: github.com and huggingface.co; the agent is likely to contact these when it follows the instructions. As links in the text: docs.comfy.org. This is read from the text; nothing was executed.

Is Ernie Image safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Ernie Image use?

Ernie Image is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Ernie Image use?

About 4.5k tokens (SKILL.md is roughly 18k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Ernie Image?

Skills that share tags, products or a category with Ernie Image: Flux2 Klein Prompting (AnastasiyaW/codex-claude-code-config, 154 stars), Character Refs (eternityspring/shuohao-skills, 4.3k stars), Qwen Image 2 1 Prompter (iamyoki/qwen-image-2.1-skill, 157 stars) and Image Gen (open-octo/octo-agent, 125 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Ernie Image?

artokun (a GitHub user) maintains it in artokun/comfyui-mcp, which has 803 GitHub stars. The repository holds 42 skills in this directory. The repository was last updated on October 5, 2026.

Source: artokun/comfyui-mcp on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.