Agent skill

AI Image Creator

by centminmod in centminmod/my-claude-code-setup

Generate, edit-from-reference, or analyze images with AI via OpenRouter (Gemini, GPT Image, Seedream, Qwen, MAI, Grok, FLUX.2, Recraft, Muse, Riverflow; Cloudflare AI Gateway BYOK).

MITAuto-check: notesMedia & Creative

Install AI Image Creator

skills CLI
$ npx skills add centminmod/my-claude-code-setup --skill ai-image-creator -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install centminmod/my-claude-code-setup ai-image-creator --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/centminmod/my-claude-code-setup.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/ai-image-creator .claude/skills/ai-image-creator && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
ai-image-creator
GitHub stars
2.7k
Token cost
~8.1k tokens
SKILL.md length
3,154 words
Files
13 (incl. scripts, references)
Skills in repo
7
Repo updated
First seen
Licence
MIT

At a glance

Generate, edit-from-reference, or analyze images with AI via OpenRouter (Gemini, GPT Image, Seedream, Qwen, MAI, Grok, FLUX.2, Recraft, Muse, Riverflow; Cloudflare AI Gateway BYOK).

  • Works in 6 steps: Write Prompt → 5: Prompt Enhancement (Optional —… → Run Generation Script → …
  • The user asks to generate an image
  • SKILL.md covers Model Selection, Instructions, Parameters and Environment Variables, plus 8 more sections
  • Runs Python scripts from its folder; calls uv, magick and brew; reaches youtu.be; needs AI_IMG_CREATOR_CF_TOKEN and AI_IMG_CREATOR_OPENROUTER_KEY

What it does

AI Image Creator is an agent skill from centminmod/my-claude-code-setup. Generate, edit-from-reference, or analyze images with AI via OpenRouter (Gemini, GPT Image, Seedream, Qwen, MAI, Grok, FLUX.2, Recraft, Muse, Riverflow; Cloudflare AI Gateway BYOK). Also analyze a video (--analyze-video, read-only — no video generated) into a text description for video prompts. Use when the user asks to generate an image, create a PNG, make an icon, make it transparent, edit with a reference, design a logo/banner, describe/analyze/explain an image ("what's in this image"), or describe/analyze a…

Its SKILL.md is about 8.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 15 other files, including scripts and reference files (for example `references/analyze-reference.md`, `references/api-reference.md` and `references/composite-reference.md`). Compatibility notes: Requires uv (Python runner) and network access. Environment variables for CF AI Gateway or direct API keys must be configured in shell profile (~/.zshrc on…

It sits in Media & Creative, covering Image generation. It works with OpenRouter, Workers AI, Qwen and Google Gemini. The repository describes itself as: Shared starter template configuration and CLAUDE.md memory bank system for Claude Code. The licence is MIT.

When your agent uses it

  • The user asks to generate an image
  • Make it transparent
  • Edit with a reference
  • Design a logo/banner

Example prompts

  • “s in this image”
  • “what happens in this video”
  • “/ai-image-creator”

Requirements

  • Python 3
  • Compatibility (from SKILL.md): Requires uv (Python runner) and network access. Environment variables for CF AI Gateway or direct API keys must be configured in shell profile (~/.zshrc on macOS, ~/.bashrc on Linux, or System Environment Variables on Windows).
  • Pre-approved tools (allowed-tools): Bash, Read, Write

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Write Prompt
  2. 5: Prompt Enhancement (Optional — Progressive Disclosure)
  3. Run Generation Script
  4. Clean Up (if temp file used)
  5. Verify Output
  6. Post-Processing (optional)

What it can do on your machine

Read from SKILL.md and the folder at commit 7d5c374. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Bash
    • Read
    • Write

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 2 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • uv
    • magick
    • brew
    • apt
    • ffmpeg

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • youtu.be

    Also links to:

    • openrouter.ai

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • AI_IMG_CREATOR_CF_TOKEN
    • AI_IMG_CREATOR_OPENROUTER_KEY
    • AI_IMG_CREATOR_GEMINI_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Requires uv (Python runner) and network access. Environment variables for CF AI Gateway or direct API keys must be configured in shell profile (~/.zshrc on macOS, ~/.bashrc on Linux, or System Environment Variables on Windows).

    From compatibility in the SKILL.md frontmatter.

Context cost

AI Image Creator loads about 8.1k tokens when it runs, and up to ~25k if it reads all its reference files. Until then it costs about 143 tokens; SKILL.md has 3,154 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~143
When it runs · the whole SKILL.md, loaded when a task matches
~8.1k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~25k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Bash, Read, Write

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from centminmod/my-claude-code-setup at commit 7d5c374, republished under its MIT licence (© centminmod). 3,154 words, ~8,145 tokens.

Download SKILL.mdSave it as .claude/skills/ai-image-creator/SKILL.md (or your agent's skills folder). This skill also uses 12 other files; get the full folder from GitHub.
name
ai-image-creator
description
Generate, edit-from-reference, or analyze images with AI via OpenRouter (Gemini, GPT Image, Seedream, Qwen, MAI, Grok, FLUX.2, Recraft, Muse, Riverflow; Cloudflare AI Gateway BYOK). Also analyze a video (--analyze-video, read-only — no video generated) into a text description for video prompts. Use when the user asks to generate an image, create a PNG, make an icon, make it transparent, edit with a reference, design a logo/banner, describe/analyze/explain an image ("what's in this image"), or describe/analyze a video ("what happens in this video").
allowed-tools
Bash, Read, Write
compatibility
Requires uv (Python runner) and network access. Environment variables for CF AI Gateway or direct API keys must be configured in shell profile (~/.zshrc on macOS, ~/.bashrc on Linux, or System Environment Variables on Windows).
metadata.tags
image-generation, ai, openrouter, cloudflare, gemini, flux2, riverflow, seedream, gpt54, gpt-image, qwen, mai, grok, recraft, muse

AI Image Creator

Generate PNG images via multiple AI models, routed through Cloudflare AI Gateway BYOK or directly via OpenRouter/Google AI Studio.

Model Selection

When the user mentions a model keyword in their image request, use the corresponding --model flag:

KeywordModelUse When User SaysBest For (measured cost · time per image)
geminiGoogle Gemini 3.1 Flash"gemini", "nano banana 2"Previous default; still marginally closer on the Opus benchmark; up to 4K ($0.067 · 12s at 1K)
gemini-liteGoogle Gemini 3.1 Flash Lite"gemini lite", "nano banana lite", "fast draft"Fast cheap iteration; 1K only ($0.034 · 5s)
nano-banana-2.1Google Nano Banana 2.1 (default)"nano banana 2.1", "nb 2.1", "generate an image" (no model specified)Versatile default, Flash-tier successor to gemini: refs, --analyze, up to 4K; near-identical fidelity for ~44% less ($0.038 · 12s at 1K)
geminiproGoogle Gemini 3 Pro"geminipro", "gemini pro", "use gemini pro"Highest-quality Gemini (~$0.17 at 2K)
riverflowSourceful Riverflow v2 Pro"riverflow", "use riverflow"Artistic/illustration ($0.15)
flux2FLUX.2 Max"flux2", "flux", "use flux"Illustration, clean lines (~$0.07/MP)
seedreamByteDance Seedream 5.0 Lite"seedream", "use seedream"2K/4K only, web-connected knowledge, 14 refs ($0.035 · 40s at 2K)
gpt5.4OpenAI GPT-5.4 Image 2"gpt5.4", "gpt-5.4 image", "use gpt5.4"Multimodal GPT; also --analyze (token-billed)
gpt-sunburstOpenAI GPT Image 2.5 Sunburst"gpt image", "gpt image 2.5", "sunburst"Precision editing, 16 refs, --quality up to max, native -t ($0.015 · 26s at default quality; token-billed, rises with --quality)
gpt-flareOpenAI GPT Image 2.5 Flare"gpt flare", "fast gpt image"Same features as Sunburst, speed tier ($0.015 · 19s at default quality)
maiMicrosoft MAI-Image-2.6"mai", "microsoft image"Multi-reference compositing of people/products/styles, 5 refs ($0.041 · 25s)
mai-flashMicrosoft MAI-Image-2.6 Flash"mai flash"Faster MAI, same 5-ref editing ($0.020 · 15s)
grokxAI Grok Imagine Image 2.0"grok", "grok imagine"1K/2K, --quality low|medium, 3 refs ($0.060 · 66s — billed above its $0.04 list price; +$0.01/ref)
qwenQwen Image 3"qwen", "qwen image"Small legible text (posters, UI, infographics), 1K/2K, 4 refs ($0.030 · 63s)
qwen-proQwen Image 3 Pro"qwen pro"Qwen Image 3 with richer world knowledge ($0.040 · 57s)
museMeta Muse Image"muse", "meta image"Complex multi-part prompts; reasons first, may search the web; prompt only ($0.010 · 18s, 1600px)
recraft-flashRecraft V4.1 Flash"recraft", "cheapest", "quick draft"Cheapest/fastest ~1K drafts; no -r ($0.007 · 5s)

Models from seedream down to recraft-flash (except gpt5.4) use the OpenRouter Images API (/v1/images). Each accepts only the -a/-s/--quality/-r options it supports, and the script rejects anything else before calling the API. Run --list-models to see each model's limits.

Cost and time are real OpenRouter charges and end-to-end times (through the Cloudflare gateway) at default settings. They come from the Claude Opus robot benchmark, one sample per model on 2026-09-28 (nano-banana-2.1 on 2026-10-07). geminipro, riverflow, flux2 and gpt5.4 were not benchmarked; their figures are list prices or earlier cost-log values. For per-model output format, resolution and quality notes, read references/model-benchmarks.md. Re-run the benchmark with the ai-image-test-run skill.

Instructions

Routing check: If the user asks to describe, analyze, or explain an existing image (not generate a new one), skip directly to the Image Analysis (--analyze) section below. No prompt enhancement or output path needed.

Video routing: If the user asks to describe, analyze, or explain a video (or wants a text description of a clip to seed/extend a video prompt), skip directly to the Video Analysis (--analyze-video) section below.

Step 1: Write Prompt

For long or complex prompts (recommended), write to ${CLAUDE_SKILL_DIR}/tmp/prompt.txt using the Write tool:

Write prompt text to ${CLAUDE_SKILL_DIR}/tmp/prompt.txt

For short prompts (under 200 chars, no special characters), pass inline via --prompt.

CRITICAL — Prompt Quality Tips:

  • Be detailed and descriptive. Include style, colors, composition, background, and intended use.
  • Good: "A flat-design globe icon with vertical timezone band lines in blue and teal, white background, clean vector style, suitable for a web app at 512x512 pixels"
  • Bad: "globe icon"
  • Specify "transparent background" or "white background" explicitly.
  • For icons, mention the target size (e.g., "512x512", "favicon at 32x32").
  • For photos, describe lighting, camera angle, and mood.
Step 1.5: Prompt Enhancement (Optional — Progressive Disclosure)

Professional prompt patterns are available in 3 reference files. These are not loaded by default — only read them when the user's request matches a category or they explicitly ask for enhancement.

Category Detection — Match the user's request to a category:

If request mentions...CategoryAlso read
"product shot", "product photo", "hero image"product_heroprompt-core.md + prompt-categories.md § product_hero
"lifestyle", "in-use", "in context"lifestyleprompt-core.md + prompt-categories.md § lifestyle
"instagram", "social media", "tiktok", "pinterest"social_mediaprompt-core.md + prompt-platforms.md + prompt-categories.md § social_media
"banner", "ad", "email header"marketing_bannerprompt-core.md + prompt-platforms.md + prompt-categories.md § marketing_banner. Routing hint: If user has an existing logo and wants multiple standard sizes → use composite mode instead (see ## Composite Banners).
"website", "app", "logo", "ad format", "leaderboard", "skyscraper"web_appprompt-core.md + prompt-platforms.md + prompt-categories.md § web_app. Routing hint: For "logo banners" or "OG images with my logo" where user has existing logo → use composite-banners.py. For "design me a new logo" → use generate-image.py.
"brand kit", "logo banners", "banner sizes", "IAB sizes", "consistent banners" + user has existing logocompositeRead references/composite-reference.md, use composite-banners.py
"icon", "favicon", "app icon"icon_logoprompt-core.md + prompt-categories.md § icon_logo
"mascot", "character", "illustration", "artwork"illustrationprompt-core.md + prompt-categories.md § illustration
"food", "drink", "recipe", "restaurant"food_drinkprompt-core.md + prompt-categories.md § food_drink
"building", "interior", "room", "architecture"architectureprompt-core.md + prompt-categories.md § architecture
"chart", "infographic", "data", "diagram"infographicprompt-core.md + prompt-categories.md § infographic
"t-shirt", "mug design", "poster", "POD", "print-on-demand"pod_designprompt-core.md + prompt-platforms.md + prompt-categories.md § pod_design
"consistent character", "same character/product across frames", "comic strip", "storyboard", "frame set", "start and last frame", "panels", "before/after"frame_consistencyRead references/consistency-presets.md — keep people/objects/scenes consistent across a SET of frames (for video first/last frames or stitched comic strips)
"describe", "analyze", "what's in this image", "explain image"analyzeHandled by the top Routing check — read references/analyze-reference.md only for advanced/structured analysis patterns
No match / simple request—Skip patterns, generate directly

When to skip enhancement:

  • User's prompt is already detailed (150+ words with camera/lighting/composition specifics)
  • Simple/direct requests ("generate a blue circle on white background")
  • User says "no pattern" or provides a fully formed prompt

When to apply:

  • User says "use product_hero pattern" or "apply social_media pattern" (explicit)
  • Request clearly matches a category above (auto-detect)
  • User asks for "enhanced prompt" or "professional quality"

Reference files (in references/ directory):

  • prompt-core.md — Foundational rules: narrative prompting, camera/lens/lighting specs, text rendering rules, model recommendations
  • prompt-platforms.md — Social media ratios, IAB ad sizes, web dimensions, POD specs — all mapped to -a/-s flags
  • prompt-categories.md — 11 category formulas with templates and complete example prompts
Step 2: Run Generation Script
bash
uv run python ${CLAUDE_SKILL_DIR}/scripts/generate-image.py \
  -o "OUTPUT_PATH" \
  [--provider openrouter|google] \
  [-a "16:9"] \
  [-s "2K"] \
  [-m "model-id"] \
  [-r "ref-image.png"] \
  [--quality "high"] \
  [-t]

With a specific model:

bash
uv run python ${CLAUDE_SKILL_DIR}/scripts/generate-image.py \
  -o "OUTPUT_PATH" \
  -m riverflow \
  -p "A serene mountain lake at sunset"

With transparent background (requires ffmpeg + imagemagick):

bash
uv run python ${CLAUDE_SKILL_DIR}/scripts/generate-image.py \
  -o "mascot.png" \
  -t \
  -p "A friendly robot mascot character"

With reference image for editing/style transfer (see Reference Images for which models accept -r):

bash
uv run python ${CLAUDE_SKILL_DIR}/scripts/generate-image.py \
  -o "edited.png" \
  -r "original.png" \
  -p "Change the background to a sunset scene"

Or with inline prompt (default model):

bash
uv run python ${CLAUDE_SKILL_DIR}/scripts/generate-image.py \
  -o "OUTPUT_PATH" \
  -p "A simple blue circle on white background"
Step 3: Clean Up (if temp file used)
bash
rm -f ${CLAUDE_SKILL_DIR}/tmp/prompt.txt
Step 4: Verify Output
bash
file OUTPUT_PATH

Confirm it shows "PNG image data" and report the file path and size to the user.

Step 5: Post-Processing (optional)

If the user needs resizing, format conversion, or other manipulation, first detect available image tools, then use them. See Image Tools section below.

Parameters

ArgumentShortRequiredDefaultDescription
--output-oYes--Output file path (parent dirs auto-created). Saved in the format its extension names (.png, .jpg, .webp; anything else means PNG). Models that return another format (Gemini Flash Lite, Grok and Seedream send JPEG; Muse and Recraft send WebP; Nano Banana 2.1 also sends a non-PNG format) are converted with ImageMagick, or saved unconverted with a warning if it's missing. The result JSON's format field reports what was written
--prompt-pNo--Inline prompt text
--prompt-file--No../tmp/prompt.txtPath to prompt file
--provider--Noopenrouteropenrouter or google
--aspect-ratio-aNomodel defaultOpenRouter only: 1:1, 16:9, 9:16, 3:2, 2:3, 4:3, 3:4, 4:5, 5:4, 21:9
--image-size-sNomodel defaultOpenRouter only: 1K, 2K, 4K. Images-API models accept only their listed sizes (seedream 2K/4K; grok/qwen/qwen-pro 1K/2K; the rest none). gemini-lite is 1K only. 0.5K is accepted only on the Gemini 3.1 Flash preview build (-m google/gemini-3.1-flash-image-preview-20260226); every selectable keyword rejects it
--model-mNonano-banana-2.1Model keyword (see Model Selection or --list-models) or full model ID
--ref-rNo--Reference image file (repeatable). For editing/style transfer. See Reference Images for supported models and per-model limits
--quality--Nomodel defaultImages-API models only: gpt-sunburst/gpt-flare take auto, low, medium, high, xhigh, max; grok takes low, medium
--analyze--No--Analyze/describe a reference image (text-only output, no image generated). Requires -r. Multimodal chat models only (gemini, gemini-lite, nano-banana-2.1, geminipro, gpt5.4)
--analyze-video--No--Analyze/describe a video. Pass the video via -r (local file or URL). OpenRouter only. Choose a model/preset with -m (default gemini3.5-flash). Returns structured JSON by default
--prose--No--(--analyze-video only) Return free-text prose instead of the default structured JSON
--contact-sheet--No--(--analyze-video, local file only) Extract evenly-spaced keyframes with ffmpeg and save a labeled contact-sheet image to PATH — a human ground-truth reference. Skipped for URL sources / if ffmpeg is missing
--verify--No--(--analyze-video, local file only) Second pass that checks the analysis against extracted frames (no video re-sent) and classifies each claim supported/contradicted/not_visible. Adds a verification object. Costs one extra model call
--transparent-tNo--Generate with transparent background. Native on gpt-sunburst/gpt-flare (no extra tools); every other model requires ffmpeg + imagemagick
--costs--No--Display generation/cost history for this project and exit
--list-models--No--List available model keywords and exit

Environment Variables

VariableRequired ForDescription
AI_IMG_CREATOR_CF_ACCOUNT_IDGateway modeCloudflare account ID
AI_IMG_CREATOR_CF_GATEWAY_IDGateway modeAI Gateway name
AI_IMG_CREATOR_CF_TOKENGateway modeGateway auth token
AI_IMG_CREATOR_OPENROUTER_KEYDirect OpenRouterOpenRouter API key (sk-or-...)
AI_IMG_CREATOR_GEMINI_KEYDirect GoogleGoogle AI Studio API key

Gateway mode activates when all 3 CF_* vars are set. Falls back to direct mode if gateway fails.

For first-time setup, see references/setup-guide.md.

Transparent Mode (-t)

Generates images with transparent backgrounds using a 3-step pipeline:

  1. Green screen generation — Prompt is augmented to place subject on solid #00FF00 green
  2. FFmpeg chroma key — Removes green background + green fringe from edges
  3. ImageMagick auto-crop — Trims transparent padding

Requirements: brew install ffmpeg imagemagick

Native transparency: With -m gpt-sunburst or -m gpt-flare, -t sends background: transparent to the Images API instead. The model renders the alpha channel directly: there is no green-screen prompt, no chroma key and no ffmpeg/imagemagick requirement. Output must still be PNG or WebP.

Use cases: Game sprites, icons, logos, mascots, marketing assets with transparency.

bash
uv run python ${CLAUDE_SKILL_DIR}/scripts/generate-image.py \
  -o "sprite.png" -t -p "A pixel art treasure chest"

Reference Images (-r)

Send existing images alongside text prompts for editing, style transfer, or guided generation. Supports multiple references, up to each model's limit:

ModelsMax -r
gemini, gemini-lite, nano-banana-2.1, geminipro, gpt5.4 (chat)no script limit
gpt-sunburst, gpt-flare16
seedream14
mai, mai-flash5
qwen, qwen-pro4
grok3
riverflow, flux2, muse, recraft-flashnot supported (errors)
bash
# Edit an existing image
uv run python ${CLAUDE_SKILL_DIR}/scripts/generate-image.py \
  -o "edited.png" -r "photo.png" -p "Make the background white"

# Style transfer with multiple references
uv run python ${CLAUDE_SKILL_DIR}/scripts/generate-image.py \
  -o "combined.png" -r "style1.png" -r "content.png" -p "Apply the style of the first image to the second"

Supported formats: PNG, JPEG, WebP, GIF.

Image Analysis (--analyze)

Describe, analyze, or explain existing images using multimodal AI vision. Returns text-only output (no image generated). Multimodal chat models only (gemini, gemini-lite, nano-banana-2.1, geminipro, gpt5.4). Images-API models output images only and are rejected.

No -o output path needed. No prompt enhancement needed. The script outputs JSON to stdout with the model's analysis in the analysis field.

bash
# Analyze with default prompt (describes subject, style, colors, composition, mood, text)
uv run python ${CLAUDE_SKILL_DIR}/scripts/generate-image.py \
  --analyze -r "photo.png"

# Analyze with custom prompt
uv run python ${CLAUDE_SKILL_DIR}/scripts/generate-image.py \
  --analyze -r "photo.png" -p "Describe this image in plain text and also in JSON structured output"

# Analyze with a specific model
uv run python ${CLAUDE_SKILL_DIR}/scripts/generate-image.py \
  --analyze -r "photo.png" -m gpt5.4 -p "What text is visible in this image?"

# Analyze multiple images together
uv run python ${CLAUDE_SKILL_DIR}/scripts/generate-image.py \
  --analyze -r "before.png" -r "after.png" -p "Compare these two images and describe the differences"

JSON output format:

json
{"ok": true, "analyze": true, "analysis": "<model text>", "provider": "openrouter", "model": "...", "mode": "gateway", "elapsed_seconds": 3.2, "ref_images": 1}

Incompatible flags: --analyze cannot be combined with -t, -a, or -s. (-o is accepted but ignored in analyze mode, which returns text only.)

For advanced analysis prompt patterns (structured output, comparison, targeted analysis), read references/analyze-reference.md.

Show full SKILL.md (1,320 more words)Show less

Video Analysis (--analyze-video)

Describe or analyze a video using OpenRouter video-input LLMs (no image generated). Use this to turn an existing clip into a description you can feed back as a prompt to generate or extend a video (e.g. with the ai-video-creator skill).

Structured JSON is the default. All 15 video models support strict structured outputs (response_format json_schema, verified), so by default analysis is a structured object with these fields: summary, setting, subjects[] (each with role/appearance/confidence), shot_timeline[] (timestamp/action/camera), camera_techniques[], editing_stylization[], lighting, color_palette[], mood, uncertain_details[], and a distilled video_generation_prompt. The editing_stylization and uncertain_details fields specifically counter the two main failure modes (missed freeze-frame/black-and-white stylization, and confabulated details). Pass --prose for a free-text description instead. The envelope's structured field is true when JSON parsed cleanly.

Pass the video via -r — either a local file (mp4/mov/webm/mkv/avi; sent as a base64 data URL) or a URL (publicly accessible, including YouTube). OpenRouter only; no -o, prompt enhancement, or output path needed.

Model selection (-m) — three presets cover the common cases; or pick any model by keyword (see --list-models):

PresetResolves toWhen to use
video-default (or omit -m)gemini3.5-flash (Google Gemini 3.5 Flash)Default — best accuracy + fastest; reads audio. ~11× the cost of the cheap tier
video-cheapqwen3.5-flash (Qwen3.5 Flash)Rock-bottom cost for quick scene summaries (or mimo for a cheap, more detailed read)
video-qualitygemini3-pro (Google Gemini 3.1 Pro)Highest-accuracy reading when it matters most

All 15 video-capable models are selectable by keyword: qwen3.5-flash, seed-1.6-flash, seed-2.0-mini, mimo, qwen3.6-35b, qwen3.6-flash, step-3.7-flash, gemini3-flash-lite, seed-2.0-lite, seed-1.6, qwen3.5-plus, minimax-m3, qwen3.6-plus, gemini3.5-flash, gemini3-pro (cheapest → priciest). Run --list-models for IDs and per-1M-token pricing.

Bare family names are not keywords. -m gemini, -m seed, or -m qwen (the image-model families) are not valid --analyze-video selectors and error with "unknown video model". Use a preset (video-default/video-cheap/video-quality) or a full keyword from the list above (e.g. gemini3.5-flash, seed-1.6-flash).

bash
# Default model (gemini3.5-flash), structured JSON output
uv run python ${CLAUDE_SKILL_DIR}/scripts/generate-image.py \
  --analyze-video -r "clip.mp4"

# Free-text prose instead of JSON
uv run python ${CLAUDE_SKILL_DIR}/scripts/generate-image.py \
  --analyze-video -r "clip.mp4" --prose

# Rock-bottom cost preset
uv run python ${CLAUDE_SKILL_DIR}/scripts/generate-image.py \
  --analyze-video -r "clip.mp4" -m video-cheap

# Highest-accuracy preset on a YouTube URL with a custom focus
uv run python ${CLAUDE_SKILL_DIR}/scripts/generate-image.py \
  --analyze-video -r "https://youtu.be/VIDEO_ID" -m video-quality \
  -p "Focus on camera movement and lighting"

JSON output format (default — analysis is a structured object):

json
{"ok": true, "analyze": true, "analyze_video": true, "structured": true, "analysis": {"summary": "...", "setting": "...", "subjects": [{"role": "protagonist", "appearance": "...", "confidence": "high"}], "shot_timeline": [{"timestamp": "0:00", "action": "...", "camera": "..."}], "camera_techniques": ["..."], "editing_stylization": ["monochrome freeze-frame", "..."], "lighting": "...", "color_palette": ["..."], "mood": "...", "uncertain_details": ["..."], "video_generation_prompt": "..."}, "provider": "openrouter", "model": "google/gemini-3.5-flash", "mode": "gateway", "elapsed_seconds": 16.9, "video_source": "clip.mp4"}

With --prose, analysis is a plain text string and structured is false.

Frame grounding (--contact-sheet, --verify)

The model samples its own frames internally, but it can still slip a confabulation into a single shot (e.g. a "golden glowing eye" in the final beat that isn't there). Two opt-in, local-file-only aids ground the analysis against real pixels using ffmpeg-extracted keyframes:

  • --contact-sheet PATH — extracts ~12 evenly-spaced keyframes (always including first and last; capped uniform sampling, not scene-detect) and tiles them into one labeled image at PATH. This is the highest-leverage aid: a human (or you) can eyeball the whole clip at a glance to sanity-check the description. Built with ImageMagick montage (timestamp labels) or, if absent, ffmpeg's tile filter. The path is echoed back as contact_sheet in the JSON envelope.
  • --verify — runs a cheap second pass that sends the contact sheet + a few full keyframes (with timestamps) and the pass-1 analysis back to the same model, and asks it to classify each claim supported / contradicted / not_visible strictly from the frames. The video is not re-sent (that would just re-confabulate from the same pixels), and undiscernible details stay not_visible rather than being "resolved" into a guess. Adds a verification object: {claims[]{claim,verdict,evidence}, corrections[], overall_accuracy}.

Both are skipped with a warning (never a hard error) for URL/YouTube sources or if ffmpeg is missing — the analysis itself always proceeds. Extracted frames go to a temp dir that is cleaned up automatically; only the --contact-sheet image is kept.

bash
# Save a ground-truth contact sheet alongside the analysis, and verify the claims
uv run python ${CLAUDE_SKILL_DIR}/scripts/generate-image.py \
  --analyze-video -r "clip.mp4" \
  --contact-sheet "exports/clip_frames.png" --verify

Notes:

  • Incompatible flags: cannot be combined with --analyze, -t, -a, or -s, and requires --provider openrouter.
  • Large local files (>20 MB) trigger a warning — base64 payloads can be slow or rejected; prefer a hosted/YouTube URL or a shorter/lower-res clip.
  • Context limits: Seed/Step models cap at ~256K tokens (fine for short clips); the 1M-context models (Qwen, Gemini, MiMo, MiniMax) are safer for longer footage.

Cost Tracking (--costs)

Every generation is logged to .ai-image-creator/costs.json in your project directory. View history:

bash
uv run python ${CLAUDE_SKILL_DIR}/scripts/generate-image.py --costs

Shows per-model breakdown: generation count, total tokens, elapsed time, and recent entries. Security: Only non-sensitive data is logged (model, tokens, timing, file path). No API keys or credentials are ever stored.

Token totals may under-count. OpenRouter image-generation responses (and Cloudflare-gateway responses) often omit the usage block, so those entries log 0 tokens. Elapsed time and generation counts are always accurate; treat token totals as best-effort.

Consider adding .ai-image-creator/ to your .gitignore.

Composite Banners

Generate consistent logo banners across multiple sizes from a JSON config. Uses ImageMagick for offline compositing — no API calls, no network required. Composites an existing logo/mark onto branded backgrounds with text at standard dimensions.

Composite vs. AI Generation — Decision Rule

Use composite-banners.py when ALL of these are true:

  • User has an existing logo/mark they want to use as-is (provides or references a logo file)
  • User wants consistent branding across multiple standard sizes (not one creative image)
  • The output is logo + text on a solid/gradient background (not a photograph, illustration, or creative design)

Use generate-image.py (AI generation) when ANY of these are true:

  • User wants a creative/artistic banner design (describes a scene, mood, concept, or style)
  • User wants AI to design the visual content (product shots, illustrations, creative layouts)
  • User wants a single banner with artistic content, not a multi-size brand kit

When composite mode applies, read references/composite-reference.md for full config schema, preset dimensions, and font handling details.

Quick Start
  1. Init config: uv run python ${CLAUDE_SKILL_DIR}/scripts/composite-banners.py --init
  2. Edit banner-config.json — set logo path, brand text, colors, banner sizes
  3. Validate: uv run python ${CLAUDE_SKILL_DIR}/scripts/composite-banners.py --validate
  4. Generate: uv run python ${CLAUDE_SKILL_DIR}/scripts/composite-banners.py -c banner-config.json -o ./banners/
Composite Parameters
ArgumentShortDefaultDescription
--config-cbanner-config.jsonConfig JSON path
--output-dir-o.Output directory
--name-nallGenerate single banner by name
--format-fpngpng, webp, jpeg
--list-presetsList IAB/social/web size presets
--initGenerate starter config
--validateCheck config, exit 0 or 2
--dry-runPreview without rendering
--jsonStructured JSON to stdout
--verbose-vVerbose output

Requirements: ImageMagick 7 (brew install imagemagick or apt install imagemagick).

Workflow Hints

Starting composite mode:

  • Ask user for: logo file path, brand name, tagline text, brand colors (hex)
  • If user doesn't have a logo yet → use generate-image.py to create one first
  • Run --init to scaffold config, then help user fill in their brand values

During generation:

  • Always run --validate before generating to catch font/logo issues early
  • Use --name to iterate on one banner before generating the full set
  • Show user 3-4 representative sizes (hero, OG, square, leaderboard) for approval

After generation:

  • If user wants creative/artistic redesign of banner visuals → switch to generate-image.py (composite only does logo + text on gradient/solid backgrounds)
  • If banners look too plain → suggest AI-generating a textured or photographic background first, then compositing the logo onto it

Combined workflow (most powerful):

  1. Use generate-image.py to AI-create a hero background or textured pattern
  2. Use composite-banners.py to overlay the logo + text onto that background at all standard sizes This gives both creative AI visuals AND pixel-perfect logo consistency.

Image Tools

On first invocation, detect available image manipulation tools:

bash
which magick convert sips ffmpeg 2>/dev/null
Available Tools
ToolCheckKey Operations
ImageMagick 7 (magick)magick --versionResize, crop, convert, composite
ImageMagick 6 (convert)convert --versionSame ops, legacy command name
sips (macOS)sips --helpResize, format conversion
ffmpegffmpeg -versionConvert formats, resize
Common Post-Processing
bash
# Resize
magick output.png -resize 512x512 icon-512.png

# Multiple sizes (icons)
for s in 16 32 48 64 128 256 512; do magick output.png -resize ${s}x${s} icon-${s}.png; done

# Convert to WebP
magick output.png output.webp

# Maskable icon (add safe-zone padding)
magick output.png -gravity center -extent 120%x120% maskable.png

# macOS sips resize
sips --resampleWidth 512 --resampleHeight 512 output.png --out icon-512.png

CRITICAL: Check tool availability before using. Prefer magick (IM7) over convert (IM6). If no tools found, inform user: brew install imagemagick.

Common Issues

"No API credentials configured"

Cause: Environment variables not set or not exported. Fix: Add exports to ~/.zshrc and run source ~/.zshrc. See references/setup-guide.md.

"HTTP 401: Unauthorized"

Cause: Invalid or expired API key/token. Fix: Check AI_IMG_CREATOR_CF_TOKEN (gateway) or AI_IMG_CREATOR_OPENROUTER_KEY (direct). Regenerate if needed.

"No images in response"

Cause: Model returned text only (safety filter, unclear prompt, or unsupported request). Fix: Make the prompt more specific and descriptive. Avoid prohibited content.

"Connection error" / timeout

Cause: Network issue or image generation taking too long (300s timeout). Fix: Retry. If persistent, try --provider google as alternative. Check CF gateway status.

Detailed API Reference

For full API formats, response schemas, BYOK configuration, and curl examples: see references/api-reference.md

For first-time setup instructions: see references/setup-guide.md

© centminmod, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 12 other files (scripts, references) in .claude/skills/ai-image-creator of centminmod/my-claude-code-setup.

  • SKILL.md
  • references/analyze-reference.md
  • references/api-reference.md
  • references/composite-reference.md
  • references/consistency-presets.md
  • references/model-benchmarks.md
  • references/prompt-categories.md
  • references/prompt-core.md
  • references/prompt-platforms.md
  • references/setup-guide.md
  • scripts/composite-banners.py
  • scripts/generate-image.py
  • tmp/.gitkeep

Open the folder on GitHubat commit 7d5c374

Compare with similar skills

AI Image Creator next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

AI Image Creator compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
AI Image Creator this skillcentminmod/my-claude-code-setup2.7k—~8.1kAutomated safety check: NotesMIT
AI Image Creatorevolution-foundation/evo-nexus545—~5.1kAutomated safety check: NotesCustom licence
Openrouter Imgrtadewald/skills180—~800Automated safety check: NotesNone
Generate ImageK-Dense-AI/claude-scientific-writer2.4k1 repos~3.8kAutomated safety check: NotesMIT
AI Image Generationaiskillstore/marketplace4301 repos~7.3kAutomated safety check: PassMIT
AI Image Generation and Editingzhayujie/CowAgent47k—~1.3kAutomated safety check: PassMIT

Similar skills

  • AI Image Creator

    evolution-foundation/evo-nexus

    Generates PNG images through OpenRouter models, with transparent backgrounds and reference-image edits, and describes existing images with multimodal vision.

    545 GitHub stars~5.1k tokensUpdated 4 mo ago
    Media & CreativeAuto-check: notes
  • Openrouter Img

    rtadewald/skills

    Generate or edit images via OpenRouter Image API across top models (Nano Banana, GPT Image, FLUX.2, Seedream, Qwen, Grok).

    180 GitHub stars~800 tokensUpdated 10 days ago
    Media & CreativeAuto-check: notes
  • Generate Image

    K-Dense-AI/claude-scientific-writer

    Generate or edit images with AI models through the OpenRouter Image API (Gemini, Seedream, Recraft, GPT-Image, Riverflow).

    2.4k GitHub starsUsed in 1 repo~3.8k tokens
    Media & CreativeAuto-check: notes
  • AI Image Generation

    aiskillstore/marketplace

    Generate and edit images on RunComfy via the runcomfy CLI — a smart router across the full image-model catalog: FLUX 2 (Klein 9B/4B, Pro, Dev, Flash, Turbo, Max), Google Nano Banana 2 / Pro, OpenAI…

    430 GitHub starsUsed in 1 repo~7.3k tokens
    Media & CreativeAuto-check passed
  • Generates or edits images from text prompts through a Python script that picks an image backend based on which API keys are configured.

    47k GitHub stars~1.3k tokensUpdated yesterday
    Media & CreativeAuto-check passed
  • Generates images through a 9Router gateway's image endpoint, with model discovery, the request fields and per-provider quirks for OpenAI, Gemini, MiniMax and others.

    30k GitHub stars~830 tokensUpdated 7 days ago
    Media & CreativeAuto-check passed

More from centminmod/my-claude-code-setup

  • Audit Session Metrics

    centminmod/my-claude-code-setup

    Audit a session-metrics JSON export for token-usage waste and produce a plain-English findings report.

    2.7k GitHub stars~2.1k tokensUpdated yesterday
    Auto-check passed
  • Session Metrics

    centminmod/my-claude-code-setup

    Tally Claude Code session token usage and cost estimates from the raw JSONL conversation log.

    2.7k GitHub stars~8k tokensUpdated yesterday
    Auto-check passed
  • Claude Docs Consultant

    centminmod/my-claude-code-setup

    Consult official Claude Code documentation from code.claude.com using selective fetching.

    2.7k GitHub stars~959 tokensUpdated yesterday
    Auto-check passed
  • Consult Codex

    centminmod/my-claude-code-setup

    Dual-AI code analysis pairing OpenAI Codex with Claude code-searcher — the lightest consult variant, two citation-verified perspectives.

    2.7k GitHub stars~4.8k tokensUpdated yesterday
    Auto-check passed
  • Consult Zai

    centminmod/my-claude-code-setup

    Dual-AI code analysis pairing z.ai GLM 5.2 with Claude code-searcher — a lightweight two-model second opinion.

    2.7k GitHub stars~4.1k tokensUpdated yesterday
    Auto-check passed
  • Task Breakdown

    centminmod/my-claude-code-setup

    Group a session-metrics session's turns into higher-level SEMANTIC TASKS ("what was I actually trying to do") and render a Tasks companion page (tasks.html + tasks.md) with a worth-it / mixed /…

    2.7k GitHub stars~2.4k tokensUpdated yesterday
    Auto-check passed

Questions about AI Image Creator

What does AI Image Creator do?

Generate, edit-from-reference, or analyze images with AI via OpenRouter (Gemini, GPT Image, Seedream, Qwen, MAI, Grok, FLUX.2, Recraft, Muse, Riverflow; Cloudflare AI Gateway BYOK). AI Image Creator is an agent skill from centminmod/my-claude-code-setup.2, Recraft, Muse, Riverflow; Cloudflare AI Gateway BYOK).

When should I use AI Image Creator?

AI Image Creator fits situations like: the user asks to generate an image; make it transparent; edit with a reference; design a logo/banner.

How do I install AI Image Creator in Claude Code?

Run `npx skills add centminmod/my-claude-code-setup --skill ai-image-creator -a claude-code`. Or copy the skill folder (.claude/skills/ai-image-creator in centminmod/my-claude-code-setup) into .claude/skills/ai-image-creator in your project. Claude Code loads it when a task matches its description.

How do I install AI Image Creator in Codex?

Run `npx skills add centminmod/my-claude-code-setup --skill ai-image-creator -a codex`. Or copy the skill folder (.claude/skills/ai-image-creator in centminmod/my-claude-code-setup) into .agents/skills/ai-image-creator in your project. Codex loads it when a task matches its description.

Can I use AI Image Creator in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add centminmod/my-claude-code-setup --skill ai-image-creator -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ai-image-creator, .gemini/skills/ai-image-creator, .github/skills/ai-image-creator and .opencode/skills/ai-image-creator in your project.

What does AI Image Creator need to run?

Going by SKILL.md and its folder, AI Image Creator needs Python for the scripts in its folder, the command-line tools its instructions call (uv, magick, brew, apt and ffmpeg) and credentials named AI_IMG_CREATOR_CF_TOKEN, AI_IMG_CREATOR_OPENROUTER_KEY and AI_IMG_CREATOR_GEMINI_KEY. Our summary lists: Python 3. Its frontmatter pre-approves these tools: Bash, Read, Write. Compatibility (from SKILL.md): Requires uv (Python runner) and network access. Environment variables for CF AI Gateway or direct API keys must be configured in shell profile (~/.zshrc on macOS, ~/.bashrc on Linux, or System Environment Variables on Windows)..

Does AI Image Creator access the network?

SKILL.md names 2 domains. In commands or code: youtu.be; the agent is likely to contact it when it follows the instructions. As links in the text: openrouter.ai. This is read from the text; nothing was executed.

Is AI Image Creator safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does AI Image Creator use?

AI Image Creator is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does AI Image Creator use?

About 8.1k tokens (SKILL.md is roughly 33k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 17k tokens, read only when the agent opens those files.

What are the alternatives to AI Image Creator?

Skills that share tags, products or a category with AI Image Creator: AI Image Creator (evolution-foundation/evo-nexus, 545 stars), Openrouter Img (rtadewald/skills, 180 stars), Generate Image (K-Dense-AI/claude-scientific-writer, 2.4k stars) and AI Image Generation (aiskillstore/marketplace, 430 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains AI Image Creator?

centminmod (a GitHub user) maintains it in centminmod/my-claude-code-setup, which has 2,656 GitHub stars. The repository holds 7 skills in this directory. The repository was last updated on October 8, 2026.

Source: centminmod/my-claude-code-setup on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.