Baoyu Imagine
LeoYeAI/openclaw-master-skills
AI image generation with OpenAI, Azure OpenAI, Google, OpenRouter, DashScope, MiniMax, Jimeng, Seedream and Replicate APIs.
Two CLI tools for image generation + vision analysis using goclaw's provider-chain pattern.
$ npx skills add therichardngai-code/gpt-image-2-pro-max --skill media-tools -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install therichardngai-code/gpt-image-2-pro-max media-tools --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/therichardngai-code/gpt-image-2-pro-max.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/media-tools .claude/skills/media-tools && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "media-tools" agent skill from https://github.com/therichardngai-code/gpt-image-2-pro-max/tree/main/.claude/skills/media-tools into .claude/skills/media-tools/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "media-tools", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/therichardngai-code/gpt-image-2-pro-max/tree/main/.claude/skills/media-toolsType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add therichardngai-code/gpt-image-2-pro-max --skill media-tools -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install therichardngai-code/gpt-image-2-pro-max media-tools --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/therichardngai-code/gpt-image-2-pro-max.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.claude/skills/media-tools .agents/skills/media-tools && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "media-tools" agent skill from https://github.com/therichardngai-code/gpt-image-2-pro-max/tree/main/.claude/skills/media-tools into .agents/skills/media-tools/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "media-tools", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add therichardngai-code/gpt-image-2-pro-max --skill media-tools -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install therichardngai-code/gpt-image-2-pro-max media-tools --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/therichardngai-code/gpt-image-2-pro-max.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.claude/skills/media-tools .cursor/skills/media-tools && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "media-tools" agent skill from https://github.com/therichardngai-code/gpt-image-2-pro-max/tree/main/.claude/skills/media-tools into .cursor/skills/media-tools/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "media-tools", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/therichardngai-code/gpt-image-2-pro-max.git --path .claude/skills/media-tools--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add therichardngai-code/gpt-image-2-pro-max --skill media-tools -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install therichardngai-code/gpt-image-2-pro-max media-tools --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/therichardngai-code/gpt-image-2-pro-max.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.claude/skills/media-tools .gemini/skills/media-tools && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "media-tools" agent skill from https://github.com/therichardngai-code/gpt-image-2-pro-max/tree/main/.claude/skills/media-tools into .gemini/skills/media-tools/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "media-tools", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install therichardngai-code/gpt-image-2-pro-max media-toolsInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add therichardngai-code/gpt-image-2-pro-max --skill media-tools -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/therichardngai-code/gpt-image-2-pro-max.git skills-src && mkdir -p .github/skills && cp -r skills-src/.claude/skills/media-tools .github/skills/media-tools && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "media-tools" agent skill from https://github.com/therichardngai-code/gpt-image-2-pro-max/tree/main/.claude/skills/media-tools into .github/skills/media-tools/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "media-tools", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add therichardngai-code/gpt-image-2-pro-max --skill media-tools -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install therichardngai-code/gpt-image-2-pro-max media-tools --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/therichardngai-code/gpt-image-2-pro-max.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.claude/skills/media-tools .opencode/skills/media-tools && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "media-tools" agent skill from https://github.com/therichardngai-code/gpt-image-2-pro-max/tree/main/.claude/skills/media-tools into .opencode/skills/media-tools/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "media-tools", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
media-toolsTwo CLI tools for image generation + vision analysis using goclaw's provider-chain pattern.
Media Tools is an agent skill from therichardngai-code/gpt-image-2-pro-max. Two CLI tools for image generation + vision analysis using goclaw's provider-chain pattern. createimage generates images via 7-provider chain (ChatGPT-OAuth/Codex, OpenRouter, Gemini, OpenAI, MiniMax, DashScope, BytePlus) — first available credential wins (OAuth session OR API key), embeds prompt into PNG tEXt metadata, saves to date-folder workspace. The ChatGPT-OAuth provider reads from ~/.codex/auth.json so a ChatGPT Plus / Codex login works without any API key. readimage analyzes images via 4-provider vision…
Its SKILL.md is about 1.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 29 other files, including scripts (for example `README.md`, `scripts/chatgpt_oauth_login.py` and `scripts/create_image.py`).
It sits in Media & Creative, covering Image generation, Model routing and gateways and OAuth and OpenID Connect. It works with OpenAI, OpenRouter and MiniMax. The licence is MIT.
Read from SKILL.md and the folder at commit dd902b1. It shows what the files ask for, not the result of running them.
Pre-approves these tools, so the agent can use them without asking each time:
BashReadFrom allowed-tools in the SKILL.md frontmatter.
Ships 13 files in scripts/ (Python, from the files we listed), which the agent can run.
Shell commands in SKILL.md call:
pythonFrom the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
openrouter.aiapi.openai.comFrom URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
OPENROUTER_API_KEYGEMINI_API_KEYOPENAI_API_KEYMINIMAX_API_KEYDASHSCOPE_API_KEYBYTEPLUS_API_KEYANTHROPIC_API_KEYFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Media Tools loads about 1.5k tokens when it runs. Until then it costs about 221 tokens; SKILL.md has 264 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check noted patterns worth knowing about, such as sudo or a known installer.
/multi-line prompts, auto UTF-8 stdout, `.env` auto-load, and `--list-providers` diagnostic. Use when Claude needs to geallowed-tools: Bash, ReadAutomated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from therichardngai-code/gpt-image-2-pro-max at commit dd902b1, republished under its MIT licence (© therichardngai-code). 264 words, ~1,506 tokens.
.claude/skills/media-tools/SKILL.md (or your agent's skills folder). This skill also uses 26 other files; get the full folder from GitHub.Python ports of goclaw's create_image and read_image tools. Faithful provider-chain pattern: priority order, skip-if-no-key, first-success-wins, cascade-on-failure.
.claude/skills/.venv/Scripts/python.exe -m pip install -r .claude/skills/media-tools/requirements.txtSet whichever providers you have keys for. Chain skips providers without keys.
# Image generation
$env:OPENROUTER_API_KEY = "..." # OpenRouter (aggregator, OpenAI-compat)
$env:GEMINI_API_KEY = "..." # Google AI Studio
$env:OPENAI_API_KEY = "..." # OpenAI gpt-image
$env:MINIMAX_API_KEY = "..." # MiniMax image-01
$env:DASHSCOPE_API_KEY = "..." # Alibaba Qwen / Wan2.6
$env:BYTEPLUS_API_KEY = "..." # ByteDance Seedream
# Vision (some keys overlap with image gen)
$env:ANTHROPIC_API_KEY = "..." # Anthropic Claude vision
# Optional API base overrides (for proxies / regional endpoints)
$env:OPENROUTER_API_BASE = "https://openrouter.ai/api/v1"
$env:OPENAI_API_BASE = "https://api.openai.com/v1"
# etc — see lib/env_keys.py for full list.claude/skills/.venv/Scripts/python.exe .claude/skills/media-tools/scripts/create_image.py \
--prompt "a cyberpunk cat in neon rain, ukiyo-e style" \
--aspect-ratio 16:9 \
--filename-hint "neon-cat"Options:
--prompt (required) — text description--aspect-ratio — 1:1 (default) | 3:4 | 4:3 | 9:16 | 16:9--filename-hint — kebab-slug for output file (no extension)--workspace — output dir; default $WEB_TOOLS_WORKSPACE or OS temp--provider — force a specific provider (skip chain): chatgpt_oauth|openrouter|gemini|openai|minimax|dashscope|byteplus--provider-order — comma-separated chain override (default: chatgpt_oauth,openrouter,gemini,openai,minimax,dashscope,byteplus)--reference-image PATH — seed generation with a reference image (PNG/JPG/WEBP). Repeatable, max 4. Supported by chatgpt_oauth, gemini, openai; chain auto-skips others when refs are present.Output saved to <workspace>/generated/<YYYY-MM-DD>/<name>.png with prompt embedded in PNG tEXt metadata.
.claude/skills/.venv/Scripts/python.exe .claude/skills/media-tools/scripts/create_image.py \
--prompt "Same character, now wearing a red trench coat, neon Tokyo street at night" \
--reference-image ./character.png \
--aspect-ratio 9:16.claude/skills/.venv/Scripts/python.exe .claude/skills/media-tools/scripts/read_image.py \
--path "C:/path/to/image.png" \
--prompt "Describe this image in detail. What text appears?"Options:
--prompt (required) — what to ask about the image--path (required) — local image file (jpg/png/gif/webp/bmp, ≤10MB)--provider — force specific provider: openrouter|gemini|anthropic|dashscope--provider-order — comma-separated chain (default: openrouter,gemini,anthropic,dashscope)--max-tokens — response max tokens (default 1024)| Pattern | goclaw source | skill file |
|---|---|---|
| Provider chain (priority + skip-no-key + first-success) | media_provider_chain.go | lib/chain.py |
| OpenRouter image (OpenAI-compat + modalities) | create_image.go:callImageGenAPI | lib/providers/openrouter_image.py |
| Gemini image (native generateContent + responseModalities) | create_image.go:callGeminiNativeImageGen | lib/providers/gemini_image.py |
OpenAI image (/images/generations) | create_image.go:callStandardImageGenAPI | lib/providers/openai_image.py |
MiniMax image (/image_generation) | create_image_minimax.go | lib/providers/minimax_image.py |
| DashScope image (async task polling) | create_image_dashscope.go | lib/providers/dashscope_image.py |
BytePlus Seedream (/api/v3/images/generations + URL download) | create_image_byteplus.go | lib/providers/byteplus_image.py |
| Vision providers (4-way chain) | read_image.go + provider Chat() | lib/providers/*_vision.py |
| PNG tEXt prompt embedding | pngEmbedPrompt | lib/png_metadata.py |
| Date-folder workspace + sanitized filename | mediaFileName | lib/filename.py |
Each provider has its own size convention; the skill maps --aspect-ratio to each:
| aspect | dashscope | byteplus | minimax/openrouter/gemini |
|---|---|---|---|
| 1:1 | 1024*1024 | 1024x1024 | 1:1 (passed through) |
| 16:9 | 1280*720 | 1280x720 | 16:9 |
| 9:16 | 720*1280 | 720x1280 | 9:16 |
| 4:3 | 1024*768 | 1024x768 | 4:3 |
| 3:4 | 768*1024 | 768x1024 | 3:4 |
© therichardngai-code, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 26 other files (scripts) in .claude/skills/media-tools of therichardngai-code/gpt-image-2-pro-max.
Open the folder on GitHubat commit dd902b1
Media Tools next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Media Tools this skilltherichardngai-code/gpt-image-2-pro-max | 101 | — | ~1.5k | Automated safety check: Notes | MIT | |
| Baoyu ImagineLeoYeAI/openclaw-master-skills | 2.2k | — | ~5.1k | Automated safety check: Notes | MIT | |
| Baoyu Image GenJimLiu/baoyu-skills | 26k | 1 repos | ~5.3k | Automated safety check: Notes | MIT | |
| AI Image Creatorevolution-foundation/evo-nexus | 545 | — | ~5.1k | Automated safety check: Notes | Custom licence | |
| Generate ImageK-Dense-AI/claude-scientific-writer | 2.4k | 1 repos | ~3.8k | Automated safety check: Notes | MIT | |
| Baoyu Imagineguanyang/open-agent-hub | 975 | — | ~4.6k | Automated safety check: Notes | MIT |
LeoYeAI/openclaw-master-skills
AI image generation with OpenAI, Azure OpenAI, Google, OpenRouter, DashScope, MiniMax, Jimeng, Seedream and Replicate APIs.
JimLiu/baoyu-skills
AI image generation with OpenAI GPT Image 2.5, Azure OpenAI, Google, OpenRouter, DashScope, Z.AI GLM-Image, MiniMax, Jimeng, Seedream, Replicate and Agnes APIs.
evolution-foundation/evo-nexus
Generates PNG images through OpenRouter models, with transparent backgrounds and reference-image edits, and describes existing images with multimodal vision.
K-Dense-AI/claude-scientific-writer
Generate or edit images with AI models through the OpenRouter Image API (Gemini, Seedream, Recraft, GPT-Image, Riverflow).
guanyang/open-agent-hub
AI image generation with OpenAI GPT Image 2, Azure OpenAI, Google, OpenRouter, DashScope, Z.AI GLM-Image, MiniMax, Jimeng, Seedream and Replicate APIs.
Peiiii/nextclaw
Use the local aigen CLI to generate images through configured providers such as OpenRouter or OpenAI.
therichardngai-code/gpt-image-2-pro-max
Production prompt-engineering pipeline for GPT-Image-2 / OpenAI image generation.
Works with
Categories
Two CLI tools for image generation + vision analysis using goclaw's provider-chain pattern. Media Tools is an agent skill from therichardngai-code/gpt-image-2-pro-max. Two CLI tools for image generation + vision analysis using goclaw's provider-chain pattern.
Media Tools fits situations like: Claude needs to generate images; run vision analysis with a specific provider rather than relying on Claudes built-in image understanding.
Run `npx skills add therichardngai-code/gpt-image-2-pro-max --skill media-tools -a claude-code`. Or copy the skill folder (.claude/skills/media-tools in therichardngai-code/gpt-image-2-pro-max) into .claude/skills/media-tools in your project. Claude Code loads it when a task matches its description.
Run `npx skills add therichardngai-code/gpt-image-2-pro-max --skill media-tools -a codex`. Or copy the skill folder (.claude/skills/media-tools in therichardngai-code/gpt-image-2-pro-max) into .agents/skills/media-tools in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add therichardngai-code/gpt-image-2-pro-max --skill media-tools -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/media-tools, .gemini/skills/media-tools, .github/skills/media-tools and .opencode/skills/media-tools in your project.
Going by SKILL.md and its folder, Media Tools needs Python for the scripts in its folder, the command-line tools its instructions call (python) and credentials named OPENROUTER_API_KEY, GEMINI_API_KEY, OPENAI_API_KEY and MINIMAX_API_KEY. Our summary lists: Python 3; A credential in OPENROUTER_API_KEY; A credential in GEMINI_API_KEY. Its frontmatter pre-approves these tools: Bash, Read.
SKILL.md names 2 domains. In commands or code: openrouter.ai and api.openai.com; the agent is likely to contact these when it follows the instructions. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found notes only (mentions a .env file; pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Media Tools is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.5k tokens (SKILL.md is roughly 6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Media Tools: Baoyu Imagine (LeoYeAI/openclaw-master-skills, 2.2k stars), Baoyu Image Gen (JimLiu/baoyu-skills, 26k stars), AI Image Creator (evolution-foundation/evo-nexus, 545 stars) and Generate Image (K-Dense-AI/claude-scientific-writer, 2.4k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
therichardngai-code (a GitHub user) maintains it in therichardngai-code/gpt-image-2-pro-max, which has 101 GitHub stars. The repository holds 2 skills in this directory. The repository was last updated on May 21, 2026.
Source: therichardngai-code/gpt-image-2-pro-max on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.