Agent skill

Structured Image Generation

by bytedance in bytedance/deer-flow

Turns an image request into a structured JSON prompt and runs a bundled Python script to generate the picture, optionally guided by reference images.

MITAuto-check passedMedia & Creative

Install Structured Image Generation

skills CLI
$ npx skills add bytedance/deer-flow --skill image-generation -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install bytedance/deer-flow image-generation --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/bytedance/deer-flow.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/public/image-generation .claude/skills/image-generation && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
image-generation
GitHub stars
83k
Used in
5 other repos
Token cost
~2.9k tokens
SKILL.md length
786 words
Files
3 (incl. scripts)
Skills in repo
23
Repo updated
First seen
Licence
MIT

At a glance

Turns an image request into a structured JSON prompt and runs a bundled Python script to generate the picture, optionally guided by reference images.

  • Works in 3 steps: Understand Requirements → Create Structured Prompt → Execute Generation
  • Generating a character, scene or product image from a written description
  • SKILL.md covers Overview, Core Capabilities, Workflow and Character Generation Example, plus 6 more sections
  • Runs Python scripts from its folder; calls python; reaches api.openai.com and api.minimaxi.com; needs IMAGE_GENERATION_API_KEY and GEMINI_API_KEY

What it does

Before generating anything, the agent works out what the picture needs: the subject, style and mood, technical details such as aspect ratio, composition and lighting, and any reference images. It then writes a JSON prompt file into the workspace folder, named after the content, and calls `scripts/generate.py` with the prompt file, an output path and optionally reference images and an aspect ratio (16:9 by default).

Worked examples cover character design and scenes built from reference images, and the folder includes a `templates/doraemon.md` template. The agent is told to call the script with its parameters rather than read its source. Paths in the examples assume a sandbox that mounts user data and the skills folder under `/mnt`.

When your agent uses it

  • Generating a character, scene or product image from a written description
  • Creating an image that should follow the style or composition of reference pictures
  • Producing a visual with a specific aspect ratio, saved to a named output file

Example prompts

  • “Create a Tokyo street style woman character from the 1990s and save it as a JPG.”
  • “Generate a product shot of a ceramic teapot, using ./refs/teapot1.jpg and ./refs/teapot2.png for style.”
  • “Make a cyberpunk hacker portrait in a 9:16 ratio.”

Requirements

  • Python, to run `scripts/generate.py`

Workflow steps

3 steps, taken from the step headings in SKILL.md.

  1. Understand Requirements
  2. Create Structured Prompt
  3. Execute Generation

What it can do on your machine

Read from SKILL.md and the folder at commit be34cc4. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • api.openai.com
    • api.minimaxi.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • IMAGE_GENERATION_API_KEY
    • GEMINI_API_KEY
    • MINIMAX_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Structured Image Generation loads about 2.9k tokens when it runs. Until then it costs about 60 tokens; SKILL.md has 786 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~60
When it runs · the whole SKILL.md, loaded when a task matches
~2.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from bytedance/deer-flow at commit be34cc4, republished under its MIT licence (© bytedance). 786 words, ~2,859 tokens.

Download SKILL.mdSave it as .claude/skills/image-generation/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
image-generation
description
Use this skill when the user requests to generate, create, imagine, or visualize images including characters, scenes, products, or any visual content. Supports structured prompts and reference images for guided generation.

Image Generation Skill

Overview

This skill generates high-quality images using structured prompts and a Python script. The workflow includes creating JSON-formatted prompts and executing image generation with optional reference images.

Core Capabilities

  • Create structured JSON prompts for AIGC image generation
  • Support multiple reference images for style/composition guidance
  • Generate images through automated Python script execution
  • Handle various image generation scenarios (character design, scenes, products, etc.)

Workflow

Step 1: Understand Requirements

When a user requests image generation, identify:

  • Subject/content: What should be in the image
  • Style preferences: Art style, mood, color palette
  • Technical specs: Aspect ratio, composition, lighting
  • Reference images: Any images to guide generation
  • You don't need to check the folder under /mnt/user-data
Step 2: Create Structured Prompt

Generate a structured JSON file in /mnt/user-data/workspace/ with naming pattern: {descriptive-name}.json

Step 3: Execute Generation

Call the Python script:

bash
python /mnt/skills/public/image-generation/scripts/generate.py \
  --prompt-file /mnt/user-data/workspace/prompt-file.json \
  --reference-images /path/to/ref1.jpg /path/to/ref2.png \
  --output-file /mnt/user-data/outputs/generated-image.jpg
  --aspect-ratio 16:9

Parameters:

  • --prompt-file: Absolute path to JSON prompt file (required)
  • --reference-images: Absolute paths to reference images (optional, space-separated)
  • --output-file: Absolute path to output image file (required)
  • --aspect-ratio: Aspect ratio of the generated image (optional, default: 16:9)

[!NOTE] Do NOT read the python file, just call it with the parameters.

Character Generation Example

User request: "Create a Tokyo street style woman character in 1990s"

Create prompt file: /mnt/user-data/workspace/asian-woman.json

json
{
  "characters": [{
    "gender": "female",
    "age": "mid-20s",
    "ethnicity": "Japanese",
    "body_type": "slender, elegant",
    "facial_features": "delicate features, expressive eyes, subtle makeup with emphasis on lips, long dark hair partially wet from rain",
    "clothing": "stylish trench coat, designer handbag, high heels, contemporary Tokyo street fashion",
    "accessories": "minimal jewelry, statement earrings, leather handbag",
    "era": "1990s"
  }],
  "negative_prompt": "blurry face, deformed, low quality, overly sharp digital look, oversaturated colors, artificial lighting, studio setting, posed, selfie angle",
  "style": "Leica M11 street photography aesthetic, film-like rendering, natural color palette with slight warmth, bokeh background blur, analog photography feel",
  "composition": "medium shot, rule of thirds, subject slightly off-center, environmental context of Tokyo street visible, shallow depth of field isolating subject",
  "lighting": "neon lights from signs and storefronts, wet pavement reflections, soft ambient city glow, natural street lighting, rim lighting from background neons",
  "color_palette": "muted naturalistic tones, warm skin tones, cool blue and magenta neon accents, desaturated compared to digital photography, film grain texture"
}

Execute generation:

bash
python /mnt/skills/public/image-generation/scripts/generate.py \
  --prompt-file /mnt/user-data/workspace/cyberpunk-hacker.json \
  --output-file /mnt/user-data/outputs/cyberpunk-hacker-01.jpg \
  --aspect-ratio 2:3

With reference images:

json
{
  "characters": [{
    "gender": "based on [Image 1]",
    "age": "based on [Image 1]",
    "ethnicity": "human from [Image 1] adapted to Star Wars universe",
    "body_type": "based on [Image 1]",
    "facial_features": "matching [Image 1] with slight weathered look from space travel",
    "clothing": "Star Wars style outfit - worn leather jacket with utility vest, cargo pants with tactical pouches, scuffed boots, belt with holster",
    "accessories": "blaster pistol on hip, comlink device on wrist, goggles pushed up on forehead, satchel with supplies, personal vehicle based on [Image 2]",
    "era": "Star Wars universe, post-Empire era"
  }],
  "prompt": "Character inspired by [Image 1] standing next to a vehicle inspired by [Image 2] on a bustling alien planet street in Star Wars universe aesthetic. Character wearing worn leather jacket with utility vest, cargo pants with tactical pouches, scuffed boots, belt with blaster holster. The vehicle adapted to Star Wars aesthetic with weathered metal panels, repulsor engines, desert dust covering, parked on the street. Exotic alien marketplace street with multi-level architecture, weathered metal structures, hanging market stalls with colorful awnings, alien species walking by as background characters. Twin suns casting warm golden light, atmospheric dust particles in air, moisture vaporators visible in distance. Gritty lived-in Star Wars aesthetic, practical effects look, film grain texture, cinematic composition.",
  "negative_prompt": "clean futuristic look, sterile environment, overly CGI appearance, fantasy medieval elements, Earth architecture, modern city",
  "style": "Star Wars original trilogy aesthetic, lived-in universe, practical effects inspired, cinematic film look, slightly desaturated with warm tones",
  "composition": "medium wide shot, character in foreground with alien street extending into background, environmental storytelling, rule of thirds",
  "lighting": "warm golden hour lighting from twin suns, rim lighting on character, atmospheric haze, practical light sources from market stalls",
  "color_palette": "warm sandy tones, ochre and sienna, dusty blues, weathered metals, muted earth colors with pops of alien market colors",
  "technical": {
    "aspect_ratio": "9:16",
    "quality": "high",
    "detail_level": "highly detailed with film-like texture"
  }
}
bash
python /mnt/skills/public/image-generation/scripts/generate.py \
  --prompt-file /mnt/user-data/workspace/star-wars-scene.json \
  --reference-images /mnt/user-data/uploads/character-ref.jpg /mnt/user-data/uploads/vehicle-ref.jpg \
  --output-file /mnt/user-data/outputs/star-wars-scene-01.jpg \
  --aspect-ratio 16:9

Common Scenarios

Use different JSON schemas for different scenarios.

Character Design:

  • Physical attributes (gender, age, ethnicity, body type)
  • Facial features and expressions
  • Clothing and accessories
  • Historical era or setting
  • Pose and context

Scene Generation:

  • Environment description
  • Time of day, weather
  • Mood and atmosphere
  • Focal points and composition

Product Visualization:

  • Product details and materials
  • Lighting setup
  • Background and context
  • Presentation angle

Specific Templates

Read the following template file only when matching the user request.

Output Handling

After generation:

  • Images are typically saved in /mnt/user-data/outputs/
  • Share generated images with user using present_files tool
  • Provide brief description of the generation result
  • Offer to iterate if adjustments needed

Tips: Enhancing Generation with Reference Images

For scenarios where visual accuracy is critical, use the image_search tool first to find reference images before generation.

Recommended scenarios for using image_search tool:

  • Character/Portrait Generation: Search for similar poses, expressions, or styles to guide facial features and body proportions
  • Specific Objects or Products: Find reference images of real objects to ensure accurate representation
  • Architectural or Environmental Scenes: Search for location references to capture authentic details
  • Fashion and Clothing: Find style references to ensure accurate garment details and styling

Example workflow:

  1. Call the image_search tool to find suitable reference images:
    image_search(query="Japanese woman street photography 1990s", size="Large")
  2. Download the returned image URLs to local files
  3. Use the downloaded images as --reference-images parameter in the generation script

This approach significantly improves generation quality by providing the model with concrete visual guidance rather than relying solely on text descriptions.

Show full SKILL.md (334 more words)Show less

Providers (Gemini / MiniMax / OpenAI-compatible)

This skill auto-selects the provider by environment variables (no CLI change):

  • GEMINI_API_KEY set → use Gemini (default, unchanged).
  • Otherwise, MINIMAX_API_KEY set → use MiniMax (/v1/image_generation, model image-01).
  • Otherwise, IMAGE_GENERATION_API_KEY set → use an OpenAI-compatible Images API.
  • Force one explicitly with IMAGE_GENERATION_PROVIDER=gemini|minimax|openai. openai-compatible is also accepted as an alias for openai.

OpenAI-compatible settings:

  • IMAGE_GENERATION_API_KEY (required)
  • IMAGE_GENERATION_BASE_URL (default https://api.openai.com/v1)
  • IMAGE_GENERATION_MODEL (default gpt-image-2.5-flare)
  • IMAGE_GENERATION_SIZE (optional fixed size override)

Text-to-image calls use POST {base_url}/images/generations. Reference-image calls use multipart POST {base_url}/images/edits; a relay may support generation without supporting edits. Responses may contain base64 image data, a data URL, or a downloadable URL. Aspect ratios map to 1024x1024, 1536x1024, or 1024x1536 unless IMAGE_GENERATION_SIZE is set. The output extension selects the API output_format: .jpg/.jpeg uses jpeg, .webp uses webp, and all other extensions use png. When dall-e-2 or dall-e-3 is configured instead, the request uses the model's supported dimensions and response_format=b64_json; DALL-E output files must use a .png extension. Reference-image editing with DALL-E models is not supported by this skill; use the default GPT Image model for edits.

MiniMax optional overrides: MINIMAX_API_HOST (default https://api.minimaxi.com), MINIMAX_IMAGE_MODEL (default image-01). Reference images are sent as the MiniMax subject_reference character image. The CLI and --prompt-file / --reference-images / --output-file / --aspect-ratio arguments are identical for both providers.

MiniMax prompt handling (provider-internal). Authoring is provider-agnostic — write the same structured JSON regardless of which provider is active. MiniMax image-01 consumes a single text string, so the MiniMax path itself sends only the JSON prompt field (the other fields such as style / composition / negative_prompt apply to the Gemini path) and enables prompt_optimizer so MiniMax expands it server-side. MiniMax caps that prompt at 1500 characters; if the prompt field is longer, the script returns an error instead of calling the API. The Gemini path receives the full structured JSON.

Notes

  • Always use English for prompts regardless of user's language
  • JSON format ensures structured, parsable prompts
  • Reference images enhance generation quality significantly
  • Iterative refinement is normal for optimal results
  • For character generation, include the detailed character object plus a consolidated prompt field

© bytedance, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (scripts) in skills/public/image-generation of bytedance/deer-flow.

  • SKILL.md
  • scripts/generate.py
  • templates/doraemon.md

Open the folder on GitHubat commit be34cc4

Used in 5 other repositories

We found 5 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 5 other GitHub owners. This page covers the copy in bytedance/deer-flow, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Structured Image Generation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Structured Image Generation compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Structured Image Generation this skillbytedance/deer-flow83k5 repos~2.9kAutomated safety check: PassMIT
Stable Diffusion with DiffusersOrchestra-Research/AI-Research-SKILLs13k6 repos~3.2kAutomated safety check: PassMIT
Codex Imagebyungjunjang/slide-master277—~401Automated safety check: PassMIT
Scientific Image PromptingCitrus-bit/Anaxa120—~1.4kAutomated safety check: PassMIT
Poster Designeraipoch/medical-research-skills2k—~1.3kAutomated safety check: PassMIT
Forge Media Route Layer0x0funky/agent-sprite-forge4.3k—~2.2kAutomated safety check: PassMIT

Similar skills

  • Stable Diffusion with Diffusers

    Orchestra-Research/AI-Research-SKILLs

    Generates and edits images with Stable Diffusion through Hugging Face Diffusers, covering text-to-image, image-to-image, inpainting, SDXL and custom pipelines.

    13k GitHub starsUsed in 6 repos~3.2k tokens
    Media & CreativeAuto-check passed
  • Codex Image

    byungjunjang/slide-master

    Generate images via Codex CLI's built-in imagegen tool (gpt-image-2).

    277 GitHub stars~401 tokensUpdated 8 days ago
    Media & CreativeAuto-check passed
  • A skill your agent uses whenever the user asks for a graphical abstract, mechanism illustration, study design schematic, concept explainer, scientific cover art, or any non-data academic image that…

    120 GitHub stars~1.4k tokensUpdated 1 mo ago
    Media & CreativeAuto-check passed
  • Poster Designer

    aipoch/medical-research-skills

    Generate professional poster design concepts and optimized image-generation prompts, then automatically run a drawing script to produce the final poster image when a user needs a poster.

    2k GitHub stars~1.3k tokensUpdated 21 days ago
    Media & CreativeAuto-check passed
  • Forge Media Route Layer

    0x0funky/agent-sprite-forge

    Generates an image or an image-to-video clip through a configured provider API or a signed-in Codex or Grok CLI, and reports the route, file, hash and cost estimate.

    4.3k GitHub stars~2.2k tokensUpdated 2 days ago
    Media & CreativeAuto-check passed
  • Gemini Image

    tyrchen/geektime-bootcamp-ai

    Reference guide for using google-genai Python library to generate images with gemini-3-pro-image-preview model.

    236 GitHub starsUsed in 1 repo~1.1k tokens
    Media & CreativeAuto-check passed

More from bytedance/deer-flow

All 23 skills in this repo
  • Vercel Deploy

    bytedance/deer-flow

    Deploys a project to Vercel with one script and no login, then returns a live preview URL and a claim link for moving the deployment into your own Vercel account.

    83k GitHub starsUsed in 10 repos~797 tokens
    Auto-check passed
  • Chart Visualization

    bytedance/deer-flow

    Picks a suitable chart type from 26 options for your data, maps the data to that chart's parameters and generates a chart image through a JavaScript script.

    83k GitHub starsUsed in 2 repos~840 tokens
    Auto-check passed
  • GitHub Deep Research

    bytedance/deer-flow

    Researches a GitHub repository over four rounds using the GitHub API and web search, then writes a structured markdown report with timeline, metrics and Mermaid diagrams.

    83k GitHub starsUsed in 5 repos~1.3k tokens
    Auto-check passed
  • Excel and CSV Data Analysis

    bytedance/deer-flow

    Analyzes uploaded Excel and CSV files with SQL through DuckDB, producing schema inspections, statistical summaries and exports to CSV, JSON or Markdown.

    83k GitHub starsUsed in 4 repos~2.2k tokens
    Auto-check passed
  • DeerFlow Smoke Test

    bytedance/deer-flow

    Walks through an end-to-end smoke test of a DeerFlow deployment: pull the latest code, deploy with Docker or locally, verify services, run health checks and write a report.

    83k GitHub stars~2.5k tokensUpdated today
    Auto-check: notes
  • Video Generation

    bytedance/deer-flow

    Generates short videos from a structured JSON prompt, optionally guided by a reference image used as the first or last frame.

    83k GitHub starsUsed in 4 repos~1.4k tokens
    Auto-check passed

Works with

Questions about Structured Image Generation

What does Structured Image Generation do?

Turns an image request into a structured JSON prompt and runs a bundled Python script to generate the picture, optionally guided by reference images. Before generating anything, the agent works out what the picture needs: the subject, style and mood, technical details such as aspect ratio, composition and lighting, and any reference images.py` with the prompt file, an output path and optionally reference images and an aspect ratio (16:9 by default).

When should I use Structured Image Generation?

Structured Image Generation fits situations like: generating a character, scene or product image from a written description; creating an image that should follow the style or composition of reference pictures; producing a visual with a specific aspect ratio, saved to a named output file.

How do I install Structured Image Generation in Claude Code?

Run `npx skills add bytedance/deer-flow --skill image-generation -a claude-code`. Or copy the skill folder (skills/public/image-generation in bytedance/deer-flow) into .claude/skills/image-generation in your project. Claude Code loads it when a task matches its description.

How do I install Structured Image Generation in Codex?

Run `npx skills add bytedance/deer-flow --skill image-generation -a codex`. Or copy the skill folder (skills/public/image-generation in bytedance/deer-flow) into .agents/skills/image-generation in your project. Codex loads it when a task matches its description.

Can I use Structured Image Generation in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add bytedance/deer-flow --skill image-generation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/image-generation, .gemini/skills/image-generation, .github/skills/image-generation and .opencode/skills/image-generation in your project.

What does Structured Image Generation need to run?

Going by SKILL.md and its folder, Structured Image Generation needs Python for the scripts in its folder, the command-line tools its instructions call (python) and credentials named IMAGE_GENERATION_API_KEY, GEMINI_API_KEY and MINIMAX_API_KEY. Our summary lists: Python, to run `scripts/generate.py`.

Does Structured Image Generation access the network?

SKILL.md names 2 domains. In commands or code: api.openai.com and api.minimaxi.com; the agent is likely to contact these when it follows the instructions. This is read from the text; nothing was executed.

Is Structured Image Generation safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Structured Image Generation use?

Structured Image Generation is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Structured Image Generation use?

About 2.9k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Structured Image Generation?

Skills that share tags, products or a category with Structured Image Generation: Stable Diffusion with Diffusers (Orchestra-Research/AI-Research-SKILLs, 13k stars), Codex Image (byungjunjang/slide-master, 277 stars), Scientific Image Prompting (Citrus-bit/Anaxa, 120 stars) and Poster Designer (aipoch/medical-research-skills, 2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Structured Image Generation?

bytedance (a GitHub organization) maintains it in bytedance/deer-flow, which has 83,484 GitHub stars. The repository holds 23 skills in this directory. The repository was last updated on October 8, 2026.

Source: bytedance/deer-flow on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.