Stable Diffusion with Diffusers
Orchestra-Research/AI-Research-SKILLs
Generates and edits images with Stable Diffusion through Hugging Face Diffusers, covering text-to-image, image-to-image, inpainting, SDXL and custom pipelines.
Turns an image request into a structured JSON prompt and runs a bundled Python script to generate the picture, optionally guided by reference images.
$ npx skills add bytedance/deer-flow --skill image-generation -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install bytedance/deer-flow image-generation --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/bytedance/deer-flow.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/public/image-generation .claude/skills/image-generation && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "image-generation" agent skill from https://github.com/bytedance/deer-flow/tree/main/skills/public/image-generation into .claude/skills/image-generation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "image-generation", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/bytedance/deer-flow/tree/main/skills/public/image-generationType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add bytedance/deer-flow --skill image-generation -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install bytedance/deer-flow image-generation --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/bytedance/deer-flow.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/public/image-generation .agents/skills/image-generation && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "image-generation" agent skill from https://github.com/bytedance/deer-flow/tree/main/skills/public/image-generation into .agents/skills/image-generation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "image-generation", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add bytedance/deer-flow --skill image-generation -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install bytedance/deer-flow image-generation --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/bytedance/deer-flow.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/public/image-generation .cursor/skills/image-generation && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "image-generation" agent skill from https://github.com/bytedance/deer-flow/tree/main/skills/public/image-generation into .cursor/skills/image-generation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "image-generation", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/bytedance/deer-flow.git --path skills/public/image-generation--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add bytedance/deer-flow --skill image-generation -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install bytedance/deer-flow image-generation --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/bytedance/deer-flow.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/public/image-generation .gemini/skills/image-generation && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "image-generation" agent skill from https://github.com/bytedance/deer-flow/tree/main/skills/public/image-generation into .gemini/skills/image-generation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "image-generation", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install bytedance/deer-flow image-generationInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add bytedance/deer-flow --skill image-generation -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/bytedance/deer-flow.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/public/image-generation .github/skills/image-generation && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "image-generation" agent skill from https://github.com/bytedance/deer-flow/tree/main/skills/public/image-generation into .github/skills/image-generation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "image-generation", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add bytedance/deer-flow --skill image-generation -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install bytedance/deer-flow image-generation --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/bytedance/deer-flow.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/public/image-generation .opencode/skills/image-generation && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "image-generation" agent skill from https://github.com/bytedance/deer-flow/tree/main/skills/public/image-generation into .opencode/skills/image-generation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "image-generation", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
image-generationTurns an image request into a structured JSON prompt and runs a bundled Python script to generate the picture, optionally guided by reference images.
Before generating anything, the agent works out what the picture needs: the subject, style and mood, technical details such as aspect ratio, composition and lighting, and any reference images. It then writes a JSON prompt file into the workspace folder, named after the content, and calls `scripts/generate.py` with the prompt file, an output path and optionally reference images and an aspect ratio (16:9 by default).
Worked examples cover character design and scenes built from reference images, and the folder includes a `templates/doraemon.md` template. The agent is told to call the script with its parameters rather than read its source. Paths in the examples assume a sandbox that mounts user data and the skills folder under `/mnt`.
3 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit be34cc4. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 1 file in scripts/ (Python), which the agent can run.
Shell commands in SKILL.md call:
pythonFrom the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
api.openai.comapi.minimaxi.comFrom URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
IMAGE_GENERATION_API_KEYGEMINI_API_KEYMINIMAX_API_KEYFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Structured Image Generation loads about 2.9k tokens when it runs. Until then it costs about 60 tokens; SKILL.md has 786 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from bytedance/deer-flow at commit be34cc4, republished under its MIT licence (© bytedance). 786 words, ~2,859 tokens.
.claude/skills/image-generation/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.This skill generates high-quality images using structured prompts and a Python script. The workflow includes creating JSON-formatted prompts and executing image generation with optional reference images.
When a user requests image generation, identify:
/mnt/user-dataGenerate a structured JSON file in /mnt/user-data/workspace/ with naming pattern: {descriptive-name}.json
Call the Python script:
python /mnt/skills/public/image-generation/scripts/generate.py \
--prompt-file /mnt/user-data/workspace/prompt-file.json \
--reference-images /path/to/ref1.jpg /path/to/ref2.png \
--output-file /mnt/user-data/outputs/generated-image.jpg
--aspect-ratio 16:9Parameters:
--prompt-file: Absolute path to JSON prompt file (required)--reference-images: Absolute paths to reference images (optional, space-separated)--output-file: Absolute path to output image file (required)--aspect-ratio: Aspect ratio of the generated image (optional, default: 16:9)[!NOTE] Do NOT read the python file, just call it with the parameters.
User request: "Create a Tokyo street style woman character in 1990s"
Create prompt file: /mnt/user-data/workspace/asian-woman.json
{
"characters": [{
"gender": "female",
"age": "mid-20s",
"ethnicity": "Japanese",
"body_type": "slender, elegant",
"facial_features": "delicate features, expressive eyes, subtle makeup with emphasis on lips, long dark hair partially wet from rain",
"clothing": "stylish trench coat, designer handbag, high heels, contemporary Tokyo street fashion",
"accessories": "minimal jewelry, statement earrings, leather handbag",
"era": "1990s"
}],
"negative_prompt": "blurry face, deformed, low quality, overly sharp digital look, oversaturated colors, artificial lighting, studio setting, posed, selfie angle",
"style": "Leica M11 street photography aesthetic, film-like rendering, natural color palette with slight warmth, bokeh background blur, analog photography feel",
"composition": "medium shot, rule of thirds, subject slightly off-center, environmental context of Tokyo street visible, shallow depth of field isolating subject",
"lighting": "neon lights from signs and storefronts, wet pavement reflections, soft ambient city glow, natural street lighting, rim lighting from background neons",
"color_palette": "muted naturalistic tones, warm skin tones, cool blue and magenta neon accents, desaturated compared to digital photography, film grain texture"
}Execute generation:
python /mnt/skills/public/image-generation/scripts/generate.py \
--prompt-file /mnt/user-data/workspace/cyberpunk-hacker.json \
--output-file /mnt/user-data/outputs/cyberpunk-hacker-01.jpg \
--aspect-ratio 2:3With reference images:
{
"characters": [{
"gender": "based on [Image 1]",
"age": "based on [Image 1]",
"ethnicity": "human from [Image 1] adapted to Star Wars universe",
"body_type": "based on [Image 1]",
"facial_features": "matching [Image 1] with slight weathered look from space travel",
"clothing": "Star Wars style outfit - worn leather jacket with utility vest, cargo pants with tactical pouches, scuffed boots, belt with holster",
"accessories": "blaster pistol on hip, comlink device on wrist, goggles pushed up on forehead, satchel with supplies, personal vehicle based on [Image 2]",
"era": "Star Wars universe, post-Empire era"
}],
"prompt": "Character inspired by [Image 1] standing next to a vehicle inspired by [Image 2] on a bustling alien planet street in Star Wars universe aesthetic. Character wearing worn leather jacket with utility vest, cargo pants with tactical pouches, scuffed boots, belt with blaster holster. The vehicle adapted to Star Wars aesthetic with weathered metal panels, repulsor engines, desert dust covering, parked on the street. Exotic alien marketplace street with multi-level architecture, weathered metal structures, hanging market stalls with colorful awnings, alien species walking by as background characters. Twin suns casting warm golden light, atmospheric dust particles in air, moisture vaporators visible in distance. Gritty lived-in Star Wars aesthetic, practical effects look, film grain texture, cinematic composition.",
"negative_prompt": "clean futuristic look, sterile environment, overly CGI appearance, fantasy medieval elements, Earth architecture, modern city",
"style": "Star Wars original trilogy aesthetic, lived-in universe, practical effects inspired, cinematic film look, slightly desaturated with warm tones",
"composition": "medium wide shot, character in foreground with alien street extending into background, environmental storytelling, rule of thirds",
"lighting": "warm golden hour lighting from twin suns, rim lighting on character, atmospheric haze, practical light sources from market stalls",
"color_palette": "warm sandy tones, ochre and sienna, dusty blues, weathered metals, muted earth colors with pops of alien market colors",
"technical": {
"aspect_ratio": "9:16",
"quality": "high",
"detail_level": "highly detailed with film-like texture"
}
}python /mnt/skills/public/image-generation/scripts/generate.py \
--prompt-file /mnt/user-data/workspace/star-wars-scene.json \
--reference-images /mnt/user-data/uploads/character-ref.jpg /mnt/user-data/uploads/vehicle-ref.jpg \
--output-file /mnt/user-data/outputs/star-wars-scene-01.jpg \
--aspect-ratio 16:9Use different JSON schemas for different scenarios.
Character Design:
Scene Generation:
Product Visualization:
Read the following template file only when matching the user request.
After generation:
/mnt/user-data/outputs/For scenarios where visual accuracy is critical, use the image_search tool first to find reference images before generation.
Recommended scenarios for using image_search tool:
Example workflow:
image_search tool to find suitable reference images:image_search(query="Japanese woman street photography 1990s", size="Large")--reference-images parameter in the generation scriptThis approach significantly improves generation quality by providing the model with concrete visual guidance rather than relying solely on text descriptions.
This skill auto-selects the provider by environment variables (no CLI change):
GEMINI_API_KEY set → use Gemini (default, unchanged).MINIMAX_API_KEY set → use MiniMax (/v1/image_generation, model image-01).IMAGE_GENERATION_API_KEY set → use an OpenAI-compatible Images API.IMAGE_GENERATION_PROVIDER=gemini|minimax|openai.
openai-compatible is also accepted as an alias for openai.OpenAI-compatible settings:
IMAGE_GENERATION_API_KEY (required)IMAGE_GENERATION_BASE_URL (default https://api.openai.com/v1)IMAGE_GENERATION_MODEL (default gpt-image-2.5-flare)IMAGE_GENERATION_SIZE (optional fixed size override)Text-to-image calls use POST {base_url}/images/generations. Reference-image calls
use multipart POST {base_url}/images/edits; a relay may support generation without
supporting edits. Responses may contain base64 image data, a data URL, or a downloadable
URL. Aspect ratios map to 1024x1024, 1536x1024, or 1024x1536 unless
IMAGE_GENERATION_SIZE is set. The output extension selects the API output_format:
.jpg/.jpeg uses jpeg, .webp uses webp, and all other extensions use png.
When dall-e-2 or dall-e-3 is configured instead, the request uses the model's
supported dimensions and response_format=b64_json; DALL-E output files must use a
.png extension. Reference-image editing with DALL-E models is not supported by this
skill; use the default GPT Image model for edits.
MiniMax optional overrides: MINIMAX_API_HOST (default https://api.minimaxi.com),
MINIMAX_IMAGE_MODEL (default image-01). Reference images are sent as the MiniMax
subject_reference character image. The CLI and --prompt-file / --reference-images
/ --output-file / --aspect-ratio arguments are identical for both providers.
MiniMax prompt handling (provider-internal). Authoring is provider-agnostic — write
the same structured JSON regardless of which provider is active. MiniMax image-01
consumes a single text string, so the MiniMax path itself sends only the JSON prompt
field (the other fields such as style / composition / negative_prompt apply to the
Gemini path) and enables prompt_optimizer so MiniMax expands it server-side. MiniMax
caps that prompt at 1500 characters; if the prompt field is longer, the script returns
an error instead of calling the API. The Gemini path receives the full structured JSON.
© bytedance, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 2 other files (scripts) in skills/public/image-generation of bytedance/deer-flow.
Open the folder on GitHubat commit be34cc4
We found 5 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 5 other GitHub owners. This page covers the copy in bytedance/deer-flow, which our catalogue first saw on October 7, 2026.
Structured Image Generation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Structured Image Generation this skillbytedance/deer-flow | 83k | 5 repos | ~2.9k | Automated safety check: Pass | MIT | |
| Stable Diffusion with DiffusersOrchestra-Research/AI-Research-SKILLs | 13k | 6 repos | ~3.2k | Automated safety check: Pass | MIT | |
| Codex Imagebyungjunjang/slide-master | 277 | — | ~401 | Automated safety check: Pass | MIT | |
| Scientific Image PromptingCitrus-bit/Anaxa | 120 | — | ~1.4k | Automated safety check: Pass | MIT | |
| Poster Designeraipoch/medical-research-skills | 2k | — | ~1.3k | Automated safety check: Pass | MIT | |
| Forge Media Route Layer0x0funky/agent-sprite-forge | 4.3k | — | ~2.2k | Automated safety check: Pass | MIT |
Orchestra-Research/AI-Research-SKILLs
Generates and edits images with Stable Diffusion through Hugging Face Diffusers, covering text-to-image, image-to-image, inpainting, SDXL and custom pipelines.
byungjunjang/slide-master
Generate images via Codex CLI's built-in imagegen tool (gpt-image-2).
Citrus-bit/Anaxa
A skill your agent uses whenever the user asks for a graphical abstract, mechanism illustration, study design schematic, concept explainer, scientific cover art, or any non-data academic image that…
aipoch/medical-research-skills
Generate professional poster design concepts and optimized image-generation prompts, then automatically run a drawing script to produce the final poster image when a user needs a poster.
0x0funky/agent-sprite-forge
Generates an image or an image-to-video clip through a configured provider API or a signed-in Codex or Grok CLI, and reports the route, file, hash and cost estimate.
tyrchen/geektime-bootcamp-ai
Reference guide for using google-genai Python library to generate images with gemini-3-pro-image-preview model.
bytedance/deer-flow
Deploys a project to Vercel with one script and no login, then returns a live preview URL and a claim link for moving the deployment into your own Vercel account.
bytedance/deer-flow
Picks a suitable chart type from 26 options for your data, maps the data to that chart's parameters and generates a chart image through a JavaScript script.
bytedance/deer-flow
Researches a GitHub repository over four rounds using the GitHub API and web search, then writes a structured markdown report with timeline, metrics and Mermaid diagrams.
bytedance/deer-flow
Analyzes uploaded Excel and CSV files with SQL through DuckDB, producing schema inspections, statistical summaries and exports to CSV, JSON or Markdown.
bytedance/deer-flow
Walks through an end-to-end smoke test of a DeerFlow deployment: pull the latest code, deploy with Docker or locally, verify services, run health checks and write a report.
bytedance/deer-flow
Generates short videos from a structured JSON prompt, optionally guided by a reference image used as the first or last frame.
Works with
Categories
Turns an image request into a structured JSON prompt and runs a bundled Python script to generate the picture, optionally guided by reference images. Before generating anything, the agent works out what the picture needs: the subject, style and mood, technical details such as aspect ratio, composition and lighting, and any reference images.py` with the prompt file, an output path and optionally reference images and an aspect ratio (16:9 by default).
Structured Image Generation fits situations like: generating a character, scene or product image from a written description; creating an image that should follow the style or composition of reference pictures; producing a visual with a specific aspect ratio, saved to a named output file.
Run `npx skills add bytedance/deer-flow --skill image-generation -a claude-code`. Or copy the skill folder (skills/public/image-generation in bytedance/deer-flow) into .claude/skills/image-generation in your project. Claude Code loads it when a task matches its description.
Run `npx skills add bytedance/deer-flow --skill image-generation -a codex`. Or copy the skill folder (skills/public/image-generation in bytedance/deer-flow) into .agents/skills/image-generation in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add bytedance/deer-flow --skill image-generation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/image-generation, .gemini/skills/image-generation, .github/skills/image-generation and .opencode/skills/image-generation in your project.
Going by SKILL.md and its folder, Structured Image Generation needs Python for the scripts in its folder, the command-line tools its instructions call (python) and credentials named IMAGE_GENERATION_API_KEY, GEMINI_API_KEY and MINIMAX_API_KEY. Our summary lists: Python, to run `scripts/generate.py`.
SKILL.md names 2 domains. In commands or code: api.openai.com and api.minimaxi.com; the agent is likely to contact these when it follows the instructions. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Structured Image Generation is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.9k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Structured Image Generation: Stable Diffusion with Diffusers (Orchestra-Research/AI-Research-SKILLs, 13k stars), Codex Image (byungjunjang/slide-master, 277 stars), Scientific Image Prompting (Citrus-bit/Anaxa, 120 stars) and Poster Designer (aipoch/medical-research-skills, 2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
bytedance (a GitHub organization) maintains it in bytedance/deer-flow, which has 83,484 GitHub stars. The repository holds 23 skills in this directory. The repository was last updated on October 8, 2026.
Source: bytedance/deer-flow on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.