Generate character art and image variations using AI image generation (Google Gemini) with reference images for style and character consistency.

MITAuto-check passedMedia & Creative

Install Image Gen

skills CLI
$ npx skills add peterkrueck/Claude-Code-Development-Kit --skill image-gen -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install peterkrueck/Claude-Code-Development-Kit image-gen --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/peterkrueck/Claude-Code-Development-Kit.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/image-gen .claude/skills/image-gen && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
image-gen
GitHub stars
1.4k
Token cost
~1.4k tokens
SKILL.md length
644 words
Files
2 (incl. scripts)
Skills in repo
9
Repo updated
First seen
Licence
MIT

At a glance

Generate character art and image variations using AI image generation (Google Gemini) with reference images for style and character consistency.

  • Works in 6 steps: Understand what the user wants → Select reference images → Craft the prompt → …
  • The user asks to generate new character poses
  • SKILL.md covers Prerequisites, Workflow, Rate Limits and Troubleshooting
  • Runs TypeScript scripts from its folder; calls deno; needs GEMINI_API_KEY

What it does

Image Gen is an agent skill from peterkrueck/Claude-Code-Development-Kit. Generate character art and image variations using AI image generation (Google Gemini) with reference images for style and character consistency. Use this skill when the user asks to generate new character poses, mascot variations, art assets, illustrations, or any AI-generated images — especially when maintaining consistency with an existing character or style.

Its SKILL.md is about 1.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including scripts (for example `scripts/generate.ts`).

It sits in Media & Creative, covering Image generation. It works with Google Gemini. The repository describes itself as: Claude Code Workflow for beginners & intermediate users. Tutorial and Installer included. The licence is MIT.

When your agent uses it

  • The user asks to generate new character poses
  • Mascot variations
  • Any AI-generated images — especially when maintaining consistency with an existing character

Example prompts

  • “/image-gen”

Requirements

  • Node.js
  • A credential in GEMINI_API_KEY

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Understand what the user wants
  2. Select reference images
  3. Craft the prompt
  4. Generate variations
  5. Pick the best variant
  6. Post-process

What it can do on your machine

Read from SKILL.md and the folder at commit ba85375. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (TypeScript), which the agent can run.

    Shell commands in SKILL.md call:

    • deno

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • aistudio.google.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • GEMINI_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Image Gen loads about 1.4k tokens when it runs. Until then it costs about 93 tokens; SKILL.md has 644 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~93
When it runs · the whole SKILL.md, loaded when a task matches
~1.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from peterkrueck/Claude-Code-Development-Kit at commit ba85375, republished under its MIT licence (© peterkrueck). 644 words, ~1,369 tokens.

Download SKILL.mdSave it as .claude/skills/image-gen/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
image-gen
description
Generate character art and image variations using AI image generation (Google Gemini) with reference images for style and character consistency. Use this skill when the user asks to generate new character poses, mascot variations, art assets, illustrations, or any AI-generated images — especially when maintaining consistency with an existing character or style.
user_invocable
true

AI Image Generation (Gemini)

Generate image variations using Google's Gemini image generation model with reference images for style and character consistency. The model supports up to 14 reference images per request and can maintain consistency across multiple characters.

Prerequisites

  • GEMINI_API_KEY environment variable must be set
  • Deno runtime installed (for the generation script)

Workflow

Step 1 — Understand what the user wants

Clarify the subject, pose, expression, context, and where the asset will be used (app screen, social media, website, etc.). This context helps craft the right prompt and choose the right aspect ratio.

Step 2 — Select reference images

Always use 1-2 reference images for consistency:

  1. Primary reference (always first): The most canonical image of the character/subject. This anchors identity — face shape, color palette, defining features.

  2. Style/pose reference (second, optional): Pick the closest existing approved asset to the target pose. This anchors proportions and art style.

The primary reference anchors identity; the style reference anchors proportions. Both together produce the most consistent results.

Step 3 — Craft the prompt

Write a detailed prompt that describes the exact pose, expression, and style:

  1. Character/subject description — physical traits that define the character (so the model doesn't drift)
  2. Pose and expression — what the character is doing
  3. Style directives — art style, line style, shading approach
  4. Background — color, scene, or transparent
  5. Framing — full body, bust, three-quarter view, etc.

Prompt template:

[CHARACTER_DESCRIPTION]. [POSE_AND_EXPRESSION]. [STYLE_DIRECTIVES]. [BACKGROUND]. [VIEW/FRAMING].

Tips:

  • Be specific about what each hand/arm is doing — vague descriptions lead to random poses
  • Always specify the background explicitly
  • Include style keywords consistently (e.g., "flat color fills", "3D render", "watercolor")
Step 4 — Generate variations

Run the bundled generation script:

bash
deno run --allow-env --allow-read --allow-write --allow-net \
  .claude/skills/image-gen/scripts/generate.ts \
  --prompt "your prompt here" \
  --ref path/to/primary-reference.png \
  --ref path/to/style-reference.png \
  --output-dir /tmp/image-gen \
  --variants 4 \
  --aspect "<choose based on use case>" \
  --size "2K"

Parameters:

FlagDefaultOptions
--variants41-8 (each is a separate API call)
--aspect1:11:1, 3:4, 4:3, 9:16, 16:9, 2:3, 3:2
--size1K512, 1K, 2K, 4K

Always default to 2K for size — higher resolution gives better quality and can always be downscaled.

Choose aspect ratio based on use case:

Use CaseAspect Ratio
Full-body character poses3:4
App icons, avatars, social profiles1:1
Mobile screens, in-app cards9:16 or 3:4
Banner/header images, OG images16:9 or 4:3
Bust/upper-body portraits1:1 or 4:3

Cost: ~$0.10/image at 2K = ~$0.40 for 4 variants.

Show full SKILL.md (267 more words)Show less
Step 5 — Pick the best variant

Use the Read tool to visually inspect all generated images. Score each on:

Consistency (most important):

  • Does it match the reference images — face, proportions, colors, style?
  • Is the art style consistent (not drifting to photorealistic, 3D, etc.)?

Quality (tiebreaker):

  • Does the image have personality and visual appeal?
  • Would this work well as a production asset?

Pick the single best variant and copy it to the project's assets directory with a descriptive name. Briefly explain why you picked it.

If none are good enough, explain what went wrong and offer to regenerate with prompt adjustments.

Step 6 — Post-process

After picking the best variant:

  • Copy the chosen file to the appropriate assets directory
  • Clean up: delete the rejected variants and the temp output directory
  • Use the image-edit skill if the user needs a different crop or size

Rate Limits

If some variants fail with 429 errors: wait 60 seconds, then rerun with only the missing number of variants. Don't retry all — just fill in the gaps.

If all fail with 429: wait 60 seconds and try again. If it keeps failing, the daily quota may be exhausted — try later or enable billing for higher limits.

Troubleshooting

  • "GEMINI_API_KEY not set" — Get a key at https://aistudio.google.com/apikey
  • "Billing not enabled" or 403 — Enable billing in Google AI Studio for image generation
  • 429 rate limit — Wait 60 seconds and retry
  • Character looks wrong — Be more specific about physical traits, ensure both reference images are included
  • Style drifted — Reinforce style keywords more strongly in the prompt
  • Pose is wrong — Be extremely specific about what each arm/hand is doing

© peterkrueck, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (scripts) in skills/image-gen of peterkrueck/Claude-Code-Development-Kit.

  • SKILL.md
  • scripts/generate.ts

Open the folder on GitHubat commit ba85375

Compare with similar skills

Image Gen next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Image Gen compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Image Gen this skillpeterkrueck/Claude-Code-Development-Kit1.4k—~1.4kAutomated safety check: PassMIT
AI Image Generation and Editingzhayujie/CowAgent47k—~1.3kAutomated safety check: PassMIT
Gemini Web Reverse-Engineered ClientJimLiu/baoyu-skills26k6 repos~1.6kAutomated safety check: PassMIT
Logo Generatorop7418/logo-generator-skill2.2k—~1.8kAutomated safety check: NotesNone
NanobananaReScienceLab/opc-skills1.8k1 repos~1.3kAutomated safety check: PassApache-2.0
SEO Image GeneratorAgriciDaniel/claude-seo18k2 repos~2.1kAutomated safety check: PassMIT

Similar skills

  • Generates or edits images from text prompts through a Python script that picks an image backend based on which API keys are configured.

    47k GitHub stars~1.3k tokensUpdated today
    Media & CreativeAuto-check passed
  • Generates text and images through an unofficial, reverse-engineered Gemini Web API, supporting reference images and multi-turn conversations.

    26k GitHub starsUsed in 6 repos~1.6k tokens
    Media & CreativeAuto-check passed
  • Logo Generator

    op7418/logo-generator-skill

    Generate professional SVG logos and high-end showcase images.

    2.2k GitHub stars~1.8k tokensUpdated 5 mo ago
    Media & CreativeAuto-check: notes
  • Nanobanana

    ReScienceLab/opc-skills

    Generate and edit images using Google Gemini 3 Pro Image (Nano Banana Pro).

    1.8k GitHub starsUsed in 1 repo~1.3k tokens
    Media & CreativeAuto-check passed
  • SEO Image Generator

    AgriciDaniel/claude-seo

    Generates Open Graph previews, blog hero images, product photos and infographics for SEO use through Gemini image tools and the banana extension.

    18k GitHub starsUsed in 2 repos~2.1k tokens
    Media & CreativeAuto-check passed
  • BlockRun Image Generation

    BlockRunAI/ClawRouter

    Generates or edits images through ClawRouter's local image API, with a choice of models and sizes and payment handled automatically through x402.

    6.6k GitHub stars~2.1k tokensUpdated 3 days ago
    Media & CreativeAuto-check passed

More from peterkrueck/Claude-Code-Development-Kit

All 9 skills in this repo
  • Image Edit

    peterkrueck/Claude-Code-Development-Kit

    Edit images with precision — crop, resize, mirror, rotate, trim, and reframe.

    1.4k GitHub stars~1.2k tokensUpdated 2 mo ago
    Auto-check passed
  • Bg Remove

    peterkrueck/Claude-Code-Development-Kit

    Remove backgrounds from images using local AI (rembg). An agent skill from peterkrueck/Claude-Code-Development-Kit.

    1.4k GitHub stars~1.3k tokensUpdated 2 mo ago
    Auto-check passed
  • Context7 Guidance

    peterkrueck/Claude-Code-Development-Kit

    Fetch CURRENT library/framework/API/CLI documentation via Context7 instead of relying on training data.

    1.4k GitHub stars~671 tokensUpdated 2 mo ago
    Auto-check passed
  • Second Opinion

    peterkrueck/Claude-Code-Development-Kit

    Get a second opinion from OpenAI's Codex CLI running locally.

    1.4k GitHub stars~2.4k tokensUpdated 2 mo ago
    Auto-check: warnings
  • Deploy

    peterkrueck/Claude-Code-Development-Kit

    Test and deploy changes safely. An agent skill from peterkrueck/Claude-Code-Development-Kit.

    1.4k GitHub stars~3.3k tokensUpdated 2 mo ago
    Auto-check passed
  • Second Opinion Gemini

    peterkrueck/Claude-Code-Development-Kit

    Get a second opinion from Google's Gemini Pro via the locally installed Gemini CLI (defaults to gemini-3.1-pro-preview; override with the CLAUDESECONDOPINIONMODEL env var).

    1.4k GitHub stars~2.3k tokensUpdated 2 mo ago
    Auto-check: warnings

Works with

Questions about Image Gen

What does Image Gen do?

Generate character art and image variations using AI image generation (Google Gemini) with reference images for style and character consistency. Image Gen is an agent skill from peterkrueck/Claude-Code-Development-Kit. Generate character art and image variations using AI image generation (Google Gemini) with reference images for style and character consistency.

When should I use Image Gen?

Image Gen fits situations like: the user asks to generate new character poses; mascot variations; any AI-generated images — especially when maintaining consistency with an existing character.

How do I install Image Gen in Claude Code?

Run `npx skills add peterkrueck/Claude-Code-Development-Kit --skill image-gen -a claude-code`. Or copy the skill folder (skills/image-gen in peterkrueck/Claude-Code-Development-Kit) into .claude/skills/image-gen in your project. Claude Code loads it when a task matches its description.

How do I install Image Gen in Codex?

Run `npx skills add peterkrueck/Claude-Code-Development-Kit --skill image-gen -a codex`. Or copy the skill folder (skills/image-gen in peterkrueck/Claude-Code-Development-Kit) into .agents/skills/image-gen in your project. Codex loads it when a task matches its description.

Can I use Image Gen in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add peterkrueck/Claude-Code-Development-Kit --skill image-gen -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/image-gen, .gemini/skills/image-gen, .github/skills/image-gen and .opencode/skills/image-gen in your project.

What does Image Gen need to run?

Going by SKILL.md and its folder, Image Gen needs TypeScript for the scripts in its folder, the command-line tools its instructions call (deno) and credentials named GEMINI_API_KEY. Our summary lists: Node.js; A credential in GEMINI_API_KEY.

Does Image Gen access the network?

SKILL.md names 1 domain. As links in the text: aistudio.google.com. This is read from the text; nothing was executed.

Is Image Gen safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Image Gen use?

Image Gen is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Image Gen use?

About 1.4k tokens (SKILL.md is roughly 5.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Image Gen?

Skills that share tags, products or a category with Image Gen: AI Image Generation and Editing (zhayujie/CowAgent, 47k stars), Gemini Web Reverse-Engineered Client (JimLiu/baoyu-skills, 26k stars), Logo Generator (op7418/logo-generator-skill, 2.2k stars) and Nanobanana (ReScienceLab/opc-skills, 1.8k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Image Gen?

peterkrueck (a GitHub user) maintains it in peterkrueck/Claude-Code-Development-Kit, which has 1,385 GitHub stars. The repository holds 9 skills in this directory. The repository was last updated on July 22, 2026.

Source: peterkrueck/Claude-Code-Development-Kit on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.