Agent skill

Comfyui Launch Flags

by artokun in artokun/comfyui-mcp

Pick the right ComfyUI startup flags for VRAM, attention, caching, and speed.

MITAuto-check passedAI & LLM Engineering

Install Comfyui Launch Flags

skills CLI
$ npx skills add artokun/comfyui-mcp --skill comfyui-launch-flags -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install artokun/comfyui-mcp comfyui-launch-flags --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/artokun/comfyui-mcp.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugin/skills/comfyui-launch-flags .claude/skills/comfyui-launch-flags && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
comfyui-launch-flags
GitHub stars
795
Token cost
~3.1k tokens
SKILL.md length
1,148 words
Files
1
Skills in repo
42
Repo updated
First seen
Licence
MIT

At a glance

Pick the right ComfyUI startup flags for VRAM, attention, caching, and speed.

  • Works in 2 steps: Z-Image + Sage = broken. Z-Image… → Sage black output on other models. If a…
  • A graph OOMs (especially long video like LTX 2 / WAN)
  • SKILL.md covers Overview, Decide first: which flag do…, VRAM strategy (mutually… and Attention backend (mutually…, plus 6 more sections
  • Calls python and uv

What it does

Comfyui Launch Flags is an agent skill from artokun/comfyui-mcp. Pick the right ComfyUI startup flags for VRAM, attention, caching, and speed. The full decision matrix for OOM (--novram / --cache-none / --disable-smart-memory), shared-VRAM creep on Windows (--reserve-vram N), model-switching with big text encoders (--cache-none), high-VRAM throughput (--gpu-only / --highvram), and attention-backend selection (--use-sage-attention for speed, --use-pytorch-cross-attention as the highest-quality / Z-Image-safe fallback). Also the acceleration-stack + Blackwell/RTX 5000 (sm120)…

Its SKILL.md is about 3.1k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering Diffusion and image models. It works with ComfyUI, PyTorch and Python. The repository describes itself as: Local-first, agent-native control plane for ComfyUI — MCP server + sidebar agent that generates images, video & audio, authors and runs workflows, and edits your live graph in… The licence is MIT.

When your agent uses it

  • A graph OOMs (especially long video like LTX 2 / WAN)
  • The GPU spills into shared VRAM and slows to a crawl
  • Switching between models eats all RAM
  • Z-Image produces black/garbled output under Sage

Example prompts

  • “/comfyui-launch-flags”

Requirements

  • Python 3

Workflow steps

2 steps, taken from the first numbered list in SKILL.md.

  1. Z-Image + Sage = broken. Z-Image (Turbo/Base) does not sample
  2. Sage black output on other models. If a model outputs black only with

What it can do on your machine

Read from SKILL.md and the folder at commit 6ad6fc0. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python
    • uv

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Comfyui Launch Flags loads about 3.1k tokens when it runs. Until then it costs about 223 tokens; SKILL.md has 1,148 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~223
When it runs · the whole SKILL.md, loaded when a task matches
~3.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from artokun/comfyui-mcp at commit 6ad6fc0, republished under its MIT licence (© artokun). 1,148 words, ~3,106 tokens.

Download SKILL.mdSave it as .claude/skills/comfyui-launch-flags/SKILL.md (or your agent's skills folder).
name
comfyui-launch-flags
description
Pick the right ComfyUI startup flags for VRAM, attention, caching, and speed. The full decision matrix for OOM (--novram / --cache-none / --disable-smart-memory), shared-VRAM creep on Windows (--reserve-vram N), model-switching with big text encoders (--cache-none), high-VRAM throughput (--gpu-only / --highvram), and attention-backend selection (--use-sage-attention for speed, --use-pytorch-cross-attention as the highest-quality / Z-Image-safe fallback). Also the acceleration-stack + Blackwell/RTX 5000 (sm_120) notes. Use when a graph OOMs (especially long video like LTX 2 / WAN), when the GPU spills into shared VRAM and slows to a crawl, when switching between models eats all RAM, when Z-Image produces black/garbled output under Sage, or when deciding which attention backend to launch with. Flag names verified against upstream comfy/cli_args.py; see Sources.
globs
**/*.json, **/packs/**

ComfyUI launch/performance flags

Overview

CLI flags passed to main.py control ComfyUI's runtime behavior (e.g. python main.py --reserve-vram 2 --use-sage-attention). The three that matter most for making a graph run rather than OOM or crawl are the VRAM strategy, the attention backend, and the cache mode. This skill is the decision matrix for choosing them.

⚠️ Verification note (August 2026). Every flag below was checked against upstream comfy/cli_args.py on current master. ComfyUI adds/renames flags often — when in doubt run python main.py --help in the target install and prefer that over this list. --enable-triton-backend / --disable-triton-backend ARE ComfyUI main.py flags on master (they used to be documented as SwarmUI-only; that is stale). --use-ck-attention is kitchen INT8 attention — no sageattention wheel. A June ComfyUI checkout still pins comfy-kitchen 0.2.10 and lacks --use-ck-attention; kitchen action:"status" reports ComfyUI-side flag support, not only the kitchen version. Use kitchen / panel_kitchen to see what this GPU can actually run.

How to apply today. The MCP's restart_comfyui (with action: "start") currently replays the exact argv of the previous run. It does not compose fresh flags. So set these when you launch ComfyUI yourself (the python main.py … line, a run.bat/shell alias, or the SwarmUI backend args box), and the tool will preserve them on restart. Injecting flags through the tool is a tracked follow-up.


Decide first: which flag do you need?

Symptom                                             ▶ Flag(s) to try
─────────────────────────────────────────────────────────────────────────────
CUDA out of memory, long video (LTX 2 / WAN)        ▶ --novram  (+ --cache-none)
OOM, still want models resident when they fit       ▶ --reserve-vram N  then --disable-smart-memory
GPU slows to a crawl, spills into "shared GPU        ▶ --reserve-vram 2..4
  memory" (Windows WDDM) mid-run
RAM blows up switching between models, or a huge     ▶ --cache-none
  text encoder (FLUX 2 / Mistral) won't unload
Plenty of VRAM (48GB+), want max throughput         ▶ --gpu-only  or  --highvram
Want faster sampling on NVIDIA                       ▶ --use-ck-attention if kitchen INT8 is available (skip the sage wheel); else --use-sage-attention
Z-Image produces BLACK / wrong output               ▶ --use-pytorch-cross-attention (NOT sage)
Sage gives black output on some models              ▶ --use-pytorch-cross-attention (or fix dtype)
ROCm, kitchen present, triton ≥ 3.7                  ▶ --enable-triton-backend

VRAM strategy and attention backend are each mutually exclusive groups, so pass at most one from each. You can combine one VRAM flag + one attention flag + one cache flag (e.g. --novram --use-sage-attention --cache-none).


VRAM strategy (mutually exclusive)

FlagWhat it doesUse when
--gpu-onlyKeep everything (incl. text encoders) on GPU48GB+ card, single model, max speed
--highvramKeep models resident in VRAM after useHigh-VRAM card, repeated runs of one model
(default)ComfyUI's smart offloadMost setups — try this first
--lowvramOffload text encoders / parts to CPUMid card OOMing on load
--novramExtreme offload — minimal VRAM footprintOOM on long video / huge models; pair with --cache-none
--cpuEverything on CPU (very slow)No usable CUDA GPU only

Modifiers (combine with the above):

  • --reserve-vram N reserves N GB for the OS and other apps. It is the fix for the Windows failure mode where the GPU quietly starts using shared VRAM and throughput collapses. Typical 2 to 4; bump to 10 for heavy video decode.
  • --disable-smart-memory forces aggressive offload to regular RAM instead of keeping models cached in VRAM. Reach for this when a run gets stuck or OOMs intermittently. Slightly slower, much more reliable.
  • --async-offload enables async weight offload streams (default on where supported); --disable-async-offload turns it off if it misbehaves.

Attention backend (mutually exclusive)

FlagNotes
--use-ck-attentionComfy Kitchen INT8 attention. No sageattention wheel. Needs comfy-kitchen present and int8_attention_is_available() on this GPU. Prefer this over the sage wheel-matching install when kitchen action:"status" says INT8 is available. Restart required.
--use-sage-attentionQuantized SageAttention kernel, ~20–40% faster sampling. Needs the sageattention package installed and version-matched — see triton-sageattention. Skip this dance when --use-ck-attention is available.
--use-flash-attentionFlashAttention kernels. Needs flash-attn built for your torch/CUDA.
--enable-triton-backend / --disable-triton-backendEnable or disable the comfy-kitchen triton backend. ComfyUI master flags (not SwarmUI-only). ROCm hosts with kitchen + triton ≥ 3.7 want --enable-triton-backend. Restart required.
--use-pytorch-cross-attentionPyTorch SDPA. Highest quality, always available, no extra deps. The safe default and the correct fallback.
--use-split-cross-attention / --use-quad-cross-attentionMemory-optimized math attention for older/low-VRAM cards.

Two gotchas worth memorizing:

  1. Z-Image + Sage = broken. Z-Image (Turbo/Base) does not sample correctly under --use-sage-attention; you get black or garbled output. Launch Z-Image with --use-pytorch-cross-attention instead. See z-image-txt2img.
  2. Sage black output on other models. If a model outputs black only with Sage, either switch to --use-pytorch-cross-attention, or (SwarmUI) set Advanced Sampling → Preferred DType = Default (16-bit). Sage-on vs Sage-off also produces slightly different images, so expect non-identical seeds.

When a graph hard-crashes with No module named 'sageattention' / triton: unavailable, the fix is the sdpa / no-compile fallback in triton-sageattention, not this flag.


Cache mode (mutually exclusive)

FlagEffect
(default --cache-ram)Cache results under RAM pressure
--cache-classicAggressive result caching
--cache-lru NKeep at most N node results (LRU)
--cache-noneCache nothing — re-executes every node; lowest RAM/VRAM. Essential when switching between dual models or when a giant text encoder (FLUX 2's Mistral) must fully unload.

Show full SKILL.md (453 more words)Show less

Speed / precision

  • --fast enables experimental, potentially quality-degrading optimizations. Accepts specific PerformanceFeature values: fp16_accumulation, fp8_matrix_mult, cublas_ops, autotune. Bare --fast turns them all on. Test output quality before committing to it.
  • UNet/VAE/text-encoder dtype casts exist too (--fp8_e4m3fn-unet, --fp16-unet, --bf16-unet, --fp32-unet, …) for forcing a compute precision. Usually the model or loader picks the right one, so only reach for these to work around a specific dtype error.

Long video OOM (LTX 2 / WAN, 24GB):   --novram --cache-none
                                      (add --disable-smart-memory if it stalls)
Windows shared-VRAM creep:            --reserve-vram 3
FLUX 2 / huge text-encoder swaps:     --cache-none
High-VRAM throughput (48GB+):         --gpu-only        (or --highvram)
Fast NVIDIA sampling (most models):   --use-ck-attention   (if kitchen INT8 is available)
                                      --use-sage-attention (otherwise; needs the wheel)
Z-Image (any):                        --use-pytorch-cross-attention
ROCm + kitchen + triton ≥ 3.7:        --enable-triton-backend

Cross-refs: video OOM specifics in ltxv2-video / wan-t2v-video; per-model VRAM math in troubleshooting and model-compatibility.


Acceleration stack & GPU coverage (context)

The attention/compile accelerators are version-locked to your exact torch + CUDA + Python. A mismatched wheel doesn't just fail to import; it can break the torch install. A known-good, mutually-compatible stack for late-2025 / 2026 NVIDIA (including Blackwell / RTX 5000, sm_120) looks like:

ComponentRoleNotes
Torch + CUDAbasee.g. Torch 2.9.x on CUDA 12.8/13; use the wheel index matching your driver
Tritontorch.compile / inductorWindows: triton-windows (woct0rdho)
SageAttention--use-sage-attentionwheel matched to torch/CUDA/python
FlashAttention--use-flash-attentionbuilt per torch/CUDA/python
xFormersmemory-efficient attentionoptional
InsightFaceFaceID / IP-Adapter / ReActoronnxruntime-gpu alongside

Operational facts worth carrying:

  • No system-wide CUDA toolkit is required to run ComfyUI. An up-to-date NVIDIA driver plus prebuilt wheels is enough. A full CUDA/MSVC/cuDNN toolchain is only needed to compile kernels yourself.
  • For broad arch coverage when building wheels, TORCH_CUDA_ARCH_LIST=7.5;8.0;8.6;8.9;9.0;10.0;12.0+PTX spans RTX 20xx→50xx and datacenter (A100/H100/B200). +PTX lets newer archs JIT.
  • DeepSpeed has no wheels for Python 3.13, and several accel wheels lag the newest Python. 3.10 to 3.12 is the safe range for the full stack.
  • Clear the Triton cache (~/.triton / %USERPROFILE%\.triton and temp) when you hit stale-kernel Triton errors after an upgrade.
  • Prefer uv pip install over pip for the venv. Resolves and downloads are dramatically faster. install_comfyui already supports this via preferUv.
  • A single bad custom node can crash all of ComfyUI at startup. Install and test acceleration and new node packs on a fresh/known-good install, not before a deadline. See troubleshooting.

Quantization quick take

  • FP8-scaled (per-tensor scaled) is markedly higher quality than plain base FP8, ~half the size of BF16, and usually faster.
  • Prefer FP8-scaled over GGUF when you have enough system RAM. ComfyUI's block-swap streams from RAM, so BF16/FP8 can run on 24GB GPUs given ample RAM. Fall back to GGUF (Q8→Q4) only when RAM is the constraint.
  • NVFP4 / NVFP8 are markedly faster on Blackwell (RTX 5000) at near-BF16 quality for supported models; LoRA support on NVFP4 is still partial.

Sources

  • Official: ComfyUI CLI args at https://github.com/comfyanonymous/ComfyUI/blob/master/comfy/cli_args.py (--use-ck-attention, --enable-triton-backend, --disable-triton-backend, --fast); hardware gates in comfy/model_management.py (supports_fp8_compute SM ≥ 8.9, supports_nvfp4_compute / supports_mxfp8_compute SM ≥ 10.0); kitchen backends in the comfy-kitchen README https://github.com/Comfy-Org/comfy-kitchen
  • Empirical: operational flag/stack recipes distilled from community auto-installer changelogs (SECourses); flags cross-checked against upstream above. The SwarmUI-only note for --enable-triton-backend is retracted as of ComfyUI master.

© artokun, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in plugin/skills/comfyui-launch-flags of artokun/comfyui-mcp.

Open the folder on GitHubat commit 6ad6fc0

Compare with similar skills

Comfyui Launch Flags next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Comfyui Launch Flags compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Comfyui Launch Flags this skillartokun/comfyui-mcp795—~3.1kAutomated safety check: PassMIT
ComfyUI Custom Node BuilderConstantineB6/comfy-pilot230—~897Automated safety check: PassMIT
Setupguaardvark/guaardvark2551 repos~1.2kAutomated safety check: PassMIT
Add Comfyui NodeMooshieblob1/MooshieUI207—~936Automated safety check: PassAGPL-3.0
Edit Comfy Workflowpeteromallet/VibeComfy150—~2.2kAutomated safety check: PassMIT
ComfyUI Custom Node Basicsjtydhr88/comfyui-custom-node-skills294—~1.6kAutomated safety check: PassMIT

Similar skills

  • ComfyUI Custom Node Builder

    ConstantineB6/comfy-pilot

    Helps an agent write ComfyUI custom nodes in Python, including wrapping an existing script, mapping data types and handling image batches.

    230 GitHub stars~897 tokensUpdated 7 mo ago
    AI & LLM EngineeringAuto-check passed
  • Setup

    guaardvark/guaardvark

    Connect this agent to a running Guaardvark (self-hosted AI studio) and check what it can do right now.

    255 GitHub starsUsed in 1 repo~1.2k tokens
    AI & LLM EngineeringAuto-check passed
  • Add Comfyui Node

    Mooshieblob1/MooshieUI

    Adds a custom ComfyUI Python node to MooshieUI — Python class in mooshienodes.py, Rust required-class registration, and optional workflow template chain hookup.

    207 GitHub stars~936 tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Edit Comfy Workflow

    peteromallet/VibeComfy

    Edit an existing VibeComfy or ComfyUI workflow, ready template, recipe, scratchpad, or target graph.

    150 GitHub stars~2.2k tokensUpdated 7 days ago
    AI & LLM EngineeringAuto-check passed
  • ComfyUI Custom Node Basics

    jtydhr88/comfyui-custom-node-skills

    Explains the V3 API for ComfyUI custom nodes: node classes, schema, inputs and outputs, registration and how it differs from the legacy V1 style.

    294 GitHub stars~1.6k tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check passed
  • ComfyUI Node Datatypes

    jtydhr88/comfyui-custom-node-skills

    Lists ComfyUI node data types, from IMAGE, MASK and LATENT tensors to model types, with their V3 classes and formats.

    294 GitHub stars~4.3k tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check passed

More from artokun/comfyui-mcp

All 42 skills in this repo
  • AI Toolkit Trainer

    artokun/comfyui-mcp

    Train custom LoRAs with ostris AI-Toolkit. An agent skill from artokun/comfyui-mcp.

    795 GitHub stars~2.7k tokensUpdated 3 days ago
    Auto-check passed
  • Anima Base

    artokun/comfyui-mcp

    Anime/illustration text-to-image (ANIMA 1.0, ~2B Cosmos DiT).

    795 GitHub stars~4k tokensUpdated 3 days ago
    Auto-check passed
  • Civitai

    artokun/comfyui-mcp

    Discover Civitai models with the BUILT-IN downloadmodel action:"searchcivitai" and install/generate them locally.

    795 GitHub stars~1.1k tokensUpdated 3 days ago
    Auto-check passed
  • Color Correction

    artokun/comfyui-mcp

    Diagnose and fix video/image color OBJECTIVELY with the getimage (action:"analyzecolor") tool (scopes/stats such as black/white points, contrast, saturation, clipping, cast) instead of eyeballing a…

    795 GitHub stars~2.4k tokensUpdated 3 days ago
    Auto-check passed
  • Comfyui Frontend Extensions

    artokun/comfyui-mcp

    Authoring ComfyUI v2 frontend extensions with @comfyorg/extension-api, covering defineNode/defineExtension/defineWidget, shell UI (sidebar tabs, commands, hotkeys), typed events, and handles.

    795 GitHub stars~5.4k tokensUpdated 3 days ago
    Auto-check passed
  • Installer Packs

    artokun/comfyui-mcp

    A skill your agent uses when installing a model family from an installer pack, or when building/deriving a new pack from an upstream installer or a workflow JSON.

    795 GitHub stars~952 tokensUpdated 3 days ago
    Auto-check passed

Questions about Comfyui Launch Flags

What does Comfyui Launch Flags do?

Pick the right ComfyUI startup flags for VRAM, attention, caching, and speed. Comfyui Launch Flags is an agent skill from artokun/comfyui-mcp. Pick the right ComfyUI startup flags for VRAM, attention, caching, and speed.

When should I use Comfyui Launch Flags?

Comfyui Launch Flags fits situations like: A graph OOMs (especially long video like LTX 2 / WAN); the GPU spills into shared VRAM and slows to a crawl; switching between models eats all RAM; Z-Image produces black/garbled output under Sage.

How do I install Comfyui Launch Flags in Claude Code?

Run `npx skills add artokun/comfyui-mcp --skill comfyui-launch-flags -a claude-code`. Or copy the skill folder (plugin/skills/comfyui-launch-flags in artokun/comfyui-mcp) into .claude/skills/comfyui-launch-flags in your project. Claude Code loads it when a task matches its description.

How do I install Comfyui Launch Flags in Codex?

Run `npx skills add artokun/comfyui-mcp --skill comfyui-launch-flags -a codex`. Or copy the skill folder (plugin/skills/comfyui-launch-flags in artokun/comfyui-mcp) into .agents/skills/comfyui-launch-flags in your project. Codex loads it when a task matches its description.

Can I use Comfyui Launch Flags in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add artokun/comfyui-mcp --skill comfyui-launch-flags -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/comfyui-launch-flags, .gemini/skills/comfyui-launch-flags, .github/skills/comfyui-launch-flags and .opencode/skills/comfyui-launch-flags in your project.

What does Comfyui Launch Flags need to run?

Going by SKILL.md and its folder, Comfyui Launch Flags needs the command-line tools its instructions call (python and uv). Our summary lists: Python 3.

Does Comfyui Launch Flags access the network?

SKILL.md names 1 domain. As links in the text: github.com. This is read from the text; nothing was executed.

Is Comfyui Launch Flags safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Comfyui Launch Flags use?

Comfyui Launch Flags is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Comfyui Launch Flags use?

About 3.1k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Comfyui Launch Flags?

Skills that share tags, products or a category with Comfyui Launch Flags: ComfyUI Custom Node Builder (ConstantineB6/comfy-pilot, 230 stars), Setup (guaardvark/guaardvark, 255 stars), Add Comfyui Node (Mooshieblob1/MooshieUI, 207 stars) and Edit Comfy Workflow (peteromallet/VibeComfy, 150 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Comfyui Launch Flags?

artokun (a GitHub user) maintains it in artokun/comfyui-mcp, which has 795 GitHub stars. The repository holds 42 skills in this directory. The repository was last updated on October 5, 2026.

Source: artokun/comfyui-mcp on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.