Agent skill

Image Generation

by shinpr in shinpr/mcp-image

Optimizes image generation prompts using Subject-Context-Style structure.

MITAuto-check passedMedia & Creative

Install Image Generation

skills CLI
$ npx skills add shinpr/mcp-image --skill image-generation -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install shinpr/mcp-image image-generation --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/shinpr/mcp-image.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/image-generation .claude/skills/image-generation && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
image-generation
GitHub stars
172
Token cost
~1.5k tokens
SKILL.md length
699 words
Files
1
Skills in repo
1
Repo updated
First seen
Licence
MIT

At a glance

Optimizes image generation prompts using Subject-Context-Style structure.

  • Works in 3 steps: SUBJECT (What) → CONTEXT (Where/When) → STYLE (How)
  • Generating images
  • SKILL.md covers Prompt Structure, Core Principles, Output Format and Enhancement Patterns, plus 5 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Image Generation is an agent skill from shinpr/mcp-image. Optimizes image generation prompts using Subject-Context-Style structure. Use this skill when generating images, creating illustrations, photos, visual assets, editing images, or crafting prompts for any image generation model.

Its SKILL.md is about 1.5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Media & Creative, covering Image generation. It works with Model Context Protocol and Google Gemini. The repository describes itself as: MCP server for AI image generation and editing with automatic prompt optimization and quality presets. Supports Nano Banana (Gemini), OpenAI GPT Image, and BytePlus Seedream. The licence is MIT.

When your agent uses it

  • Generating images
  • Creating illustrations
  • Crafting prompts for any image generation model

Example prompts

  • “Use the image-generation skill to optimiz image generation prompts using Subject-Context-Style structure”
  • “/image-generation”

Workflow steps

3 steps, taken from the step headings in SKILL.md.

  1. SUBJECT (What)
  2. CONTEXT (Where/When)
  3. STYLE (How)

What it can do on your machine

Read from SKILL.md and the folder at commit a75308a. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Image Generation loads about 1.5k tokens when it runs. Until then it costs about 61 tokens; SKILL.md has 699 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~61
When it runs · the whole SKILL.md, loaded when a task matches
~1.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from shinpr/mcp-image at commit a75308a, republished under its MIT licence (© shinpr). 699 words, ~1,476 tokens.

Download SKILL.mdSave it as .claude/skills/image-generation/SKILL.md (or your agent's skills folder).
name
image-generation
description
Optimizes image generation prompts using Subject-Context-Style structure. Use this skill when generating images, creating illustrations, photos, visual assets, editing images, or crafting prompts for any image generation model.

Image Generation Prompt Best Practices

Prompt Structure

Enhance every image generation prompt around three core elements:

1. SUBJECT (What)
  • Physical characteristics: textures, materials, colors, scale
  • Actions, poses, expressions if applicable
  • Distinctive features that define the subject
2. CONTEXT (Where/When)
  • Setting, background, spatial relationships (foreground, midground, background)
  • Time of day, weather, atmospheric conditions
  • Mood and emotional tone of the scene
3. STYLE (How)
  • Artistic or photographic approach: reference specific artists, movements, or styles
  • Camera/lens choices: specify focal length, aperture, and shooting angle when photographic

Core Principles

  • Preserve intent — Add visual details (lighting, texture, composition) only in areas the user left unspecified; keep all user-specified elements unchanged
  • Positive descriptions only — Describe what should be present; rephrase any exclusion as an inclusion
  • Specific over vague — "golden hour sunlight at 15° angle" beats "nice lighting"
  • Natural flow — Weave elements into a single flowing description, not a bullet list

Output Format

Return the enhanced prompt as a single flowing paragraph. When the user provides multiple requests, return each as a separate enhanced prompt under a labeled heading.

Enhancement Patterns

Hyper-Specific Details

Add concrete visual details for any Subject/Context/Style element not specified by the user:

  • Lighting → direction, quality, color temperature, shadow behavior
  • Textures → surface materials, weathering, reflectivity
  • Atmosphere → particulates, humidity, depth haze
  • Scale → relative sizes, distances, proportions
Camera Control Terminology

When a photographic look is appropriate:

  • Lens type: "shot with 85mm portrait lens", "wide-angle 24mm"
  • Aperture: "shallow depth of field at f/1.8", "deep focus at f/11"
  • Angle: "low angle emphasizing height", "bird's eye view"
  • Motion: "motion blur on the paws", "frozen mid-action"
Atmospheric Enhancement

Convey mood through environmental details:

  • Emotional tone with visual indicators: "serene (soft diffused light, muted palette)", "ominous (low contrast, heavy shadows, desaturated)", "jubilant (high saturation, warm tones, dynamic motion)"
  • Weather/air: "morning mist", "dust particles in a sunbeam"
Text in Images

When the image should contain readable text (signs, labels, titles, typography):

  • Specify the exact text content in quotes: "OPEN 24 HOURS" in bold sans-serif
  • Describe visual treatment: font style, weight, size relative to the scene
  • Define placement and integration: "centered on the storefront awning", "hand-lettered on the chalkboard"

Feature Patterns

Character Consistency

When the same character must be recognizable across multiple images:

  • Include at least 3 recognizable visual markers (distinctive scar, signature clothing, unique hairstyle, characteristic accessory)
  • Use anchoring words: "distinctive", "signature", "always wears", "always has"
  • Be specific: "round tortoiseshell glasses" not just "glasses"
Compositional Integration

When combining multiple visual elements in one scene:

  • Define spatial relationships with proportions: "foreground (40% of frame)", "midground", "background"
  • Define how elements interact spatially and visually: overlap, reflection, shared lighting, color echo between foreground and background
  • Specify relative scale and interaction between elements
Show full SKILL.md (262 more words)Show less
Real-World Accuracy

When depicting real places, cultures, or historical elements:

  • Use specific terminology: "traditional Edo-period architecture", "authentic Moroccan zellige tilework"
  • Include culturally accurate details
  • Reference geographical or historical specifics
Purpose-Driven Enhancement

Tailor the prompt to the intended use:

PurposeEmphasis
Product photoClean background, studio lighting, commercial appeal
UI mockupFlat design elements, consistent spacing, screen-appropriate
Presentation slideBold composition, clear focal point, text-friendly layout
Social mediaEye-catching, vibrant, crop-friendly aspect ratio
Book/album coverTypography space, dramatic mood, symbolic elements

Image Editing

When modifying an existing image:

  • Preserve the original's core characteristics: color palette, lighting style, composition
  • Use anchoring phrases: "maintain the existing...", "preserve the original...", "keep the same..."
  • Be specific about what to change vs what to keep unchanged
  • Describe modifications relative to the existing image, not from scratch

Ambiguous Cases

  • When user intent is unclear between photographic and illustrative style, ask before enhancing
  • When enhancement would significantly change the user's concept, present the original interpretation alongside the enhanced version
  • When cultural or historical accuracy cannot be verified, flag the uncertainty rather than guessing

Scope

This skill covers static image prompt enhancement only. It does not cover video generation, 3D rendering, or image analysis/description.

Example

Input: "A happy dog in a park"

Enhanced: "Golden retriever mid-leap catching a red frisbee, ears flying, tongue out in joy, in a sunlit urban park. Soft morning light filtering through oak trees creates dappled shadows on emerald grass. Background shows families on picnic blankets, slightly out of focus. Shot from low angle emphasizing the dog's athletic movement, with motion blur on the paws suggesting speed."

© shinpr, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/image-generation of shinpr/mcp-image.

Open the folder on GitHubat commit a75308a

Compare with similar skills

Image Generation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Image Generation compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Image Generation this skillshinpr/mcp-image172—~1.5kAutomated safety check: PassMIT
Gemini Interactions APIAyuilos/Miffan192—~4.6kAutomated safety check: PassAGPL-3.0
Nano Banana Proswarmclawai/swarmclaw688—~481Automated safety check: PassMIT
Fal AI Mediamajiayu000/claude-skill-registry6665 repos~1.7kAutomated safety check: PassMIT
Fal AI Mediaaffaan-m/ECC275k2 repos~1.2kAutomated safety check: PassMIT
Fal AI Mediaaffaan-m/ECC275k—~1.4kAutomated safety check: PassMIT

Similar skills

  • A skill your agent uses when writing code that calls the Gemini API for text generation, multi-turn chat, multimodal understanding, image generation, video generation, streaming responses…

    192 GitHub stars~4.6k tokensUpdated 2 days ago
    Media & CreativeAuto-check passed
  • Nano Banana Pro

    swarmclawai/swarmclaw

    Generate or edit images via Gemini 3 Pro Image (Nano Banana Pro).

    688 GitHub stars~481 tokensUpdated 3 mo ago
    Media & CreativeAuto-check passed
  • Fal AI Media

    majiayu000/claude-skill-registry

    Unified media generation via fal.ai MCP — image, video, and audio.

    666 GitHub starsUsed in 5 repos~1.7k tokens
    Media & CreativeAuto-check passed
  • Fal AI Media

    affaan-m/ECC

    通过 fal.ai MCP 实现统一的媒体生成——图像、视频和音频。涵盖文本到图像(Nano Banana)、文本/图像到视频(Seedance、Kling、Veo 3)、文本到语音(CSM-1B),以及视频到音频(ThinkSound)。当用户想要使用 AI 生成图像、视频或音频时使用。

    275k GitHub starsUsed in 2 repos~1.2k tokens
    Media & CreativeAuto-check passed
  • Fal AI Media

    affaan-m/ECC

    fal.ai MCPによる統合メディア生成(画像、動画、音声)。テキストから画像(Nano Banana)、テキスト/画像から動画(Seedance、Kling、Veo 3)、テキストから音声(CSM-1B)、動画から音声(ThinkSound)をカバーします。ユーザーがAIで画像、動画、音声を生成したい場合に使用します。

    275k GitHub stars~1.4k tokensUpdated 3 days ago
    Media & CreativeAuto-check passed
  • Scenario Gemini Image

    scenario-labs/skills

    A skill your agent uses when generating or editing images with Google's Gemini image models (Nano Banana) on Scenario via MCP: text-to-image, natural-language instruction editing, identity locking…

    913 GitHub stars~1.5k tokensUpdated yesterday
    Media & CreativeAuto-check passed

Questions about Image Generation

What does Image Generation do?

Optimizes image generation prompts using Subject-Context-Style structure. Image Generation is an agent skill from shinpr/mcp-image. Optimizes image generation prompts using Subject-Context-Style structure.

When should I use Image Generation?

Image Generation fits situations like: generating images; creating illustrations; crafting prompts for any image generation model.

How do I install Image Generation in Claude Code?

Run `npx skills add shinpr/mcp-image --skill image-generation -a claude-code`. Or copy the skill folder (skills/image-generation in shinpr/mcp-image) into .claude/skills/image-generation in your project. Claude Code loads it when a task matches its description.

How do I install Image Generation in Codex?

Run `npx skills add shinpr/mcp-image --skill image-generation -a codex`. Or copy the skill folder (skills/image-generation in shinpr/mcp-image) into .agents/skills/image-generation in your project. Codex loads it when a task matches its description.

Can I use Image Generation in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add shinpr/mcp-image --skill image-generation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/image-generation, .gemini/skills/image-generation, .github/skills/image-generation and .opencode/skills/image-generation in your project.

What does Image Generation need to run?

SKILL.md names no scripts, command-line tools or credentials: Image Generation is instructions for the agent only.

Does Image Generation access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Image Generation safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Image Generation use?

Image Generation is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Image Generation use?

About 1.5k tokens (SKILL.md is roughly 5.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Image Generation?

Skills that share tags, products or a category with Image Generation: Gemini Interactions API (Ayuilos/Miffan, 192 stars), Nano Banana Pro (swarmclawai/swarmclaw, 688 stars), Fal AI Media (majiayu000/claude-skill-registry, 666 stars) and Fal AI Media (affaan-m/ECC, 275k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Image Generation?

shinpr (a GitHub user) maintains it in shinpr/mcp-image, which has 172 GitHub stars. The repository was last updated on October 8, 2026.

Source: shinpr/mcp-image on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.