Validate Draw Things LoRA training end to end with draw-things-cli, including tiny-dataset training, loss and scaler checks, checkpoint sanity, and base-versus-LoRA generation comparison.

GPL-3.0Auto-check passedAI & LLM Engineering

Install Train Lora

skills CLI
$ npx skills add drawthingsai/draw-things-community --skill train-lora -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install drawthingsai/draw-things-community train-lora --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/drawthingsai/draw-things-community.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/train-lora .claude/skills/train-lora && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
train-lora
GitHub stars
579
Token cost
~2k tokens
SKILL.md length
809 words
Files
2
Skills in repo
8
Repo updated
First seen
Licence
GPL-3.0

At a glance

Validate Draw Things LoRA training end to end with draw-things-cli, including tiny-dataset training, loss and scaler checks, checkpoint sanity, and base-versus-LoRA generation comparison.

  • Works in 7 steps: Run a 1-step smoke test to confirm the… → Run a 20-step probe to confirm scale… → Run a 100-step check to see whether… → …
  • Tasks that involve Fine-tuning
  • SKILL.md covers Goal, Build Once, Runtime Notes and Dataset Setup, plus 8 more sections
  • Calls bazel and ffmpeg

What it does

Train Lora is an agent skill from drawthingsai/draw-things-community. Validate Draw Things LoRA training end to end with draw-things-cli, including tiny-dataset training, loss and scaler checks, checkpoint sanity, and base-versus-LoRA generation comparison.

Its SKILL.md is about 2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files (for example `agents/openai.yaml`).

It sits in AI & LLM Engineering, covering Fine-tuning and Diffusion and image models. The repository describes itself as: The community repository for the Draw Things app. The licence is GPL-3.0.

When your agent uses it

  • Tasks that involve Fine-tuning
  • Tasks that involve Diffusion and image models

Example prompts

  • “/train-lora”

Workflow steps

7 steps, taken from the first numbered list in SKILL.md.

  1. Run a 1-step smoke test to confirm the graph compiles, the loss is finite, and a checkpoint is written.
  2. Run a 20-step probe to confirm scale stays healthy and loss is not obviously blowing up.
  3. Run a 100-step check to see whether like-for-like timestep bands decline.
  4. Run a 500-step run before declaring the trainer healthy.
  5. If you are validating a new attention backend, run the 500-step check on that intended backend, not only on a fallback path.
  6. Generate a base image and a LoRA image with the same prompt, seed, and settings.
  7. Compose reference/base/LoRA into one image for visual review.

What it can do on your machine

Read from SKILL.md and the folder at commit 3cac075. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • bazel
    • ffmpeg

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Train Lora loads about 2k tokens when it runs. Until then it costs about 50 tokens; SKILL.md has 809 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~50
When it runs · the whole SKILL.md, loaded when a task matches
~2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from drawthingsai/draw-things-community at commit 3cac075, republished under its GPL-3.0 licence (© drawthingsai). 809 words, ~1,960 tokens.

Download SKILL.mdSave it as .claude/skills/train-lora/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
train-lora
description
Validate Draw Things LoRA training end to end with draw-things-cli, including tiny-dataset training, loss and scaler checks, checkpoint sanity, and base-versus-LoRA generation comparison.

Train LoRA Skill

Use this workflow to validate Draw Things LoRA training end to end with draw-things-cli.

Goal

Train on a tiny local dataset, watch loss and scaler health, then verify visually with a reference/base/LoRA comparison.

Build Once

Build the optimized CLI first:

sh
bazel build --compilation_mode=opt //Apps:DrawThingsCLI

Use the built binary for every run:

sh
bazel-bin/Apps/DrawThingsCLI

Do not switch between bazel run and bazel-bin/... during one validation cycle unless you need to. Reuse the same binary so compile/runtime behavior stays comparable and permission prompts stay predictable.

Runtime Notes

  • Graph compile can be quiet for a long time. With --compilation_mode=opt, 5 to 30 minutes is possible on heavy trainers. Do not assume a hang too early.
  • Training checkpoints are written into the models directory.
  • If the environment requires command approvals, ask once for a stable command shape and stable output/log names, then rename artifacts afterward.
  • For local unregistered LoRAs, prefer passing explicit loras[].version during generation instead of depending on custom_lora.json.

Dataset Setup

For a single-image reconstruction check:

  • Create a local dataset directory.
  • Put the image in that directory.
  • Add a matching .txt caption file beside it.
  • Keep the caption minimal for trigger-only tests, for example:
text
zimgdogref

Validation Ladder

Use this order:

  1. Run a 1-step smoke test to confirm the graph compiles, the loss is finite, and a checkpoint is written.
  2. Run a 20-step probe to confirm scale stays healthy and loss is not obviously blowing up.
  3. Run a 100-step check to see whether like-for-like timestep bands decline.
  4. Run a 500-step run before declaring the trainer healthy.
  5. If you are validating a new attention backend, run the 500-step check on that intended backend, not only on a fallback path.
  6. Generate a base image and a LoRA image with the same prompt, seed, and settings.
  7. Compose reference/base/LoRA into one image for visual review.

Baseline Train Command

Use this as the generic 512x512 single-image baseline:

sh
bazel-bin/Apps/DrawThingsCLI train lora \
  --models-dir /Users/liu/Library/Containers/com.liuliu.draw-things/Data/Documents/Models \
  --model MODEL.ckpt \
  --dataset /tmp/single_image_dataset \
  --output RUN_NAME \
  --name RUN_NAME \
  --steps 500 \
  --rank 32 \
  --scale 1 \
  --learning-rate 0:4e-4 \
  --gradient-accumulation 4 \
  --warmup-steps 20 \
  --save-every 100 \
  --width 512 \
  --height 512 \
  --seed 7 \
  --config-json '{"steps_between_restarts":200}' \
  --no-download-missing \
  --offline

Model Baselines

Scaler Rules
  • Healthy scale is architecture-dependent; choose it from the model's numeric contract.
  • Do not set scale lower than 1 to make a run stable. That hides overflow and can prevent useful learning.
  • If the model does not apply internal scaling that shrinks gradients, start from 32768.0.
  • If the model has explicit internal downscaling or projection compensation, use the validated lower scale for that model family and record why.
FLUX.1
  • Base model: flux_1_dev_q8p.ckpt
  • Validated guidance settings:
    • guidanceScale = 3.5
    • guidanceEmbed = 3.5
    • shift = 2
    • resolutionDependentShift = false
  • Healthy scale is typically 32768.0
Z Image Turbo
  • Base model: z_image_turbo_1.0_i8x.ckpt
  • The validated training baseline is the generic command above.
  • The validated trainer scale is 1024.0
  • For generation validation, use cfg = 1
Z Image Base
  • Base model: z_image_1.0_q8p.ckpt
  • The validated training baseline is the generic command above.
  • The validated trainer scale is 1024.0
  • For generation validation, keep the model’s recommended Base path:
json
{"sampler":17,"shift":1.8776105999999999,"resolutionDependentShift":true}
  • A validated comparison used cfg = 4
Qwen Image BF16
  • Base model: qwen_image_2512_bf16_i8x.ckpt
  • The validated training baseline is the generic command above.
  • Healthy scale is 32768.0
  • Use the exact training caption first before trying richer prompts
Show full SKILL.md (315 more words)Show less

What To Watch During Training

  • Raw loss is noisy because each step samples a different timestep. Do not expect monotonic decline step by step.
  • Compare like-for-like timestep bands instead.
  • Mid/high timestep bands should usually improve first.
  • Low timestep spikes can happen, but they should stay bounded.
  • For flow-style objectives, low timestep loss is not always the easiest band. If the target includes a full noise or velocity term that is weakly visible in the low-timestep input, low timestep bins can be intrinsically harder.
  • If scale steadily collapses, something is seriously wrong.
  • If scale collapses only on a new backend, compare against the known-stable backend before changing learning rate or dataset settings.
  • If scale collapses on a model with rotary applied through cmul, check whether trainer rotary constants are expanded to the real query/key head count.
  • Before blaming the optimizer, confirm the checkpoint is real:
    • nontrivial file size
    • lora_up tensors are not all zero

Generation Validation

Always compare base and LoRA with the exact same:

  • prompt
  • seed
  • width
  • height
  • steps
  • CFG
  • model-specific sampler/shift settings

Use the exact training caption first. If that fails, richer prompts are not useful for debugging.

For non-distilled base models, do not under-sample the generation validation. Use the model's real baseline settings, including enough steps, the correct CFG behavior, and the correct sampler family.

For local LoRAs, pass explicit version metadata:

json
"loras": [
  {
    "file": "RUN_NAME_500_lora_f32.ckpt",
    "version": "MODEL_VERSION",
    "weight": 1.0
  }
]

If the model also needs LoRA mode metadata, pass mode too.

Example Generate Commands

Z Image Turbo
sh
bazel-bin/Apps/DrawThingsCLI generate \
  --models-dir /Users/liu/Library/Containers/com.liuliu.draw-things/Data/Documents/Models \
  --model z_image_turbo_1.0_i8x.ckpt \
  --prompt zimgdogref \
  --width 512 \
  --height 512 \
  --steps 15 \
  --cfg 1 \
  --seed 7 \
  --config-json '{"loras":[{"file":"RUN_NAME_500_lora_f32.ckpt","version":"z_image","weight":1.0}]}' \
  --offline \
  --no-download-missing \
  --output /tmp/zimg_lora.png
Z Image Base
sh
bazel-bin/Apps/DrawThingsCLI generate \
  --models-dir /Users/liu/Library/Containers/com.liuliu.draw-things/Data/Documents/Models \
  --model z_image_1.0_q8p.ckpt \
  --prompt zimgdogref \
  --width 512 \
  --height 512 \
  --steps 20 \
  --cfg 4 \
  --seed 7 \
  --config-json '{"sampler":17,"shift":1.8776105999999999,"resolutionDependentShift":true,"loras":[{"file":"RUN_NAME_500_lora_f32.ckpt","version":"z_image","weight":1.0}]}' \
  --offline \
  --no-download-missing \
  --output /tmp/zimg_base_lora.png

Compose A Review Image

Use ffmpeg to compose reference, base, and LoRA side by side:

sh
ffmpeg -y \
  -i /tmp/single_image_dataset/dog.png \
  -i /tmp/base.png \
  -i /tmp/lora.png \
  -filter_complex hstack=inputs=3 \
  -frames:v 1 \
  /tmp/compare.png

Expected Outcomes

  • Base should usually look generic or unrelated to the exact training identity.
  • A healthy LoRA should pull noticeably toward the training subject by 100 to 500 steps.
  • For single-image dog tests, the correct check is not “perfect reconstruction”; it is whether the LoRA image is materially closer to the reference than the base image.

© drawthingsai, GPL-3.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in .agents/skills/train-lora of drawthingsai/draw-things-community.

  • SKILL.md
  • agents/openai.yaml

Open the folder on GitHubat commit 3cac075

Compare with similar skills

Train Lora next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Train Lora compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Train Lora this skilldrawthingsai/draw-things-community579—~2kAutomated safety check: PassGPL-3.0
Add Pipelineverl-project/verl-omni1.2k—~1kAutomated safety check: PassApache-2.0
Defect Image Generation with Cosmos AnomalyGenNVIDIA/skills3.5k—~5kAutomated safety check: NotesApache-2.0
Lora Managerartokun/comfyui-mcp793—~726Automated safety check: PassMIT
Comfyui ManagerShiroEirin/comfyui-good-anima476—~5.2kAutomated safety check: PassGPL-3.0
Debug Renderartokun/comfyui-mcp793—~1.7kAutomated safety check: PassMIT

Similar skills

  • Add Pipeline

    verl-project/verl-omni

    Router for adding a diffusion or omni pipeline to verl-omni.

    1.2k GitHub stars~1k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Orchestrates defect image generation for PCBA, metal surface and glass inspection with NVIDIA Cosmos AnomalyGen on OSMO, from cold-start Day 0 to real-photo Day 1 labeling.

    3.5k GitHub stars~5k tokensUpdated today
    AI & LLM EngineeringAuto-check: notes
  • Lora Manager

    artokun/comfyui-mcp

    Author ComfyUI-LoRA-Manager nodes from the panel. An agent skill from artokun/comfyui-mcp.

    793 GitHub stars~726 tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • Comfyui Manager

    ShiroEirin/comfyui-good-anima

    Manage ComfyUI server, models, workflows, LoRAs, queues, dependencies and CLI workflow execution via comfyui-skill.

    476 GitHub stars~5.2k tokensUpdated 3 mo ago
    AI & LLM EngineeringAuto-check passed
  • Debug Render

    artokun/comfyui-mcp

    Debug a WRONG or imperfect render (not a hard error) by inspecting inputs and intermediate steps with run-to-node.

    793 GitHub stars~1.7k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • Flux Txt2img

    artokun/comfyui-mcp

    Build Flux txt2img workflows with Flux.1 Dev (SRPO), Flux 2 Klein 9B, Turbo LoRAs, FluxGuidance, and DualCLIPLoader patterns

    793 GitHub stars~3k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed

More from drawthingsai/draw-things-community

All 8 skills in this repo
  • Run Longcat Avatar Video

    drawthingsai/draw-things-community

    Generate, benchmark, validate, and troubleshoot LongCat-Video-Avatar 1.5 videos with draw-things-cli.

    579 GitHub stars~2.8k tokensUpdated today
    Auto-check passed
  • Setup Cpu Proxy Server

    drawthingsai/draw-things-community

    Set up and verify a new Draw Things CPU proxy and Envoy server using the scripts in Scripts/ServerManagement/CPUScript.

    579 GitHub stars~3.8k tokensUpdated today
    Auto-check passed
  • Use Drawthings CLI

    drawthingsai/draw-things-community

    Use and troubleshoot an already installed Homebrew draw-things-cli for model discovery, authentication, local, cloud, or remote image generation, basic image-to-image and video generation, output…

    579 GitHub stars~3k tokensUpdated today
    Auto-check passed
  • Init GPU Server

    drawthingsai/draw-things-community

    Initialize a Draw Things GPU server with GPUScript, including script sync, Docker/CUDA/NVIDIA runtime setup, 7T data disk mounting, mergerfs, and end-to-end GPU verification.

    579 GitHub stars~2.2k tokensUpdated today
    Auto-check passed
  • New Model Integration

    drawthingsai/draw-things-community

    Add a new image or video generative model to the Draw Things app / CLI with a compile-first, end-to-end workflow across SwiftDiffusion, tokenizer plumbing, text encoder, fixed encoder, UNet / DiT…

    579 GitHub stars~4.8k tokensUpdated today
    Auto-check passed
  • Trainer Integration

    drawthingsai/draw-things-community

    Add or tighten Draw Things LoRA trainer support for generative models available in the Draw Things app / CLI, covering LoRA builders, trainer dispatch, tokenizer and fixed-encoder wiring, checkpoint…

    579 GitHub stars~2.6k tokensUpdated today
    Auto-check passed

Questions about Train Lora

What does Train Lora do?

Validate Draw Things LoRA training end to end with draw-things-cli, including tiny-dataset training, loss and scaler checks, checkpoint sanity, and base-versus-LoRA generation comparison. Train Lora is an agent skill from drawthingsai/draw-things-community. Validate Draw Things LoRA training end to end with draw-things-cli, including tiny-dataset training, loss and scaler checks, checkpoint sanity, and base-versus-LoRA generation comparison.

When should I use Train Lora?

Train Lora fits situations like: tasks that involve Fine-tuning; tasks that involve Diffusion and image models.

How do I install Train Lora in Claude Code?

Run `npx skills add drawthingsai/draw-things-community --skill train-lora -a claude-code`. Or copy the skill folder (.agents/skills/train-lora in drawthingsai/draw-things-community) into .claude/skills/train-lora in your project. Claude Code loads it when a task matches its description.

How do I install Train Lora in Codex?

Run `npx skills add drawthingsai/draw-things-community --skill train-lora -a codex`. Or copy the skill folder (.agents/skills/train-lora in drawthingsai/draw-things-community) into .agents/skills/train-lora in your project. Codex loads it when a task matches its description.

Can I use Train Lora in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add drawthingsai/draw-things-community --skill train-lora -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/train-lora, .gemini/skills/train-lora, .github/skills/train-lora and .opencode/skills/train-lora in your project.

What does Train Lora need to run?

Going by SKILL.md and its folder, Train Lora needs the command-line tools its instructions call (bazel and ffmpeg).

Does Train Lora access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Train Lora safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Train Lora use?

Train Lora is published under the GPL-3.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Train Lora use?

About 2k tokens (SKILL.md is roughly 7.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Train Lora?

Skills that share tags, products or a category with Train Lora: Add Pipeline (verl-project/verl-omni, 1.2k stars), Defect Image Generation with Cosmos AnomalyGen (NVIDIA/skills, 3.5k stars), Lora Manager (artokun/comfyui-mcp, 793 stars) and Comfyui Manager (ShiroEirin/comfyui-good-anima, 476 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Train Lora?

drawthingsai (a GitHub organization) maintains it in drawthingsai/draw-things-community, which has 579 GitHub stars. The repository holds 8 skills in this directory. The repository was last updated on October 6, 2026.

Source: drawthingsai/draw-things-community on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.