Two CLI tools for image generation + vision analysis using goclaw's provider-chain pattern.

MITAuto-check: notesMedia & Creative

Install Media Tools

skills CLI
$ npx skills add therichardngai-code/gpt-image-2-pro-max --skill media-tools -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install therichardngai-code/gpt-image-2-pro-max media-tools --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/therichardngai-code/gpt-image-2-pro-max.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/media-tools .claude/skills/media-tools && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
media-tools
GitHub stars
101
Token cost
~1.5k tokens
SKILL.md length
264 words
Files
27 (incl. scripts)
Skills in repo
2
Repo updated
First seen
Licence
MIT

At a glance

Two CLI tools for image generation + vision analysis using goclaw's provider-chain pattern.

  • Claude needs to generate images
  • SKILL.md covers Setup, API keys (env vars), create_image and read_image, plus 2 more sections
  • Runs Python scripts from its folder; calls python; reaches openrouter.ai and api.openai.com; needs OPENROUTER_API_KEY and GEMINI_API_KEY
  • Run vision analysis with a specific provider rather than relying on Claudes built-in image understanding

What it does

Media Tools is an agent skill from therichardngai-code/gpt-image-2-pro-max. Two CLI tools for image generation + vision analysis using goclaw's provider-chain pattern. createimage generates images via 7-provider chain (ChatGPT-OAuth/Codex, OpenRouter, Gemini, OpenAI, MiniMax, DashScope, BytePlus) — first available credential wins (OAuth session OR API key), embeds prompt into PNG tEXt metadata, saves to date-folder workspace. The ChatGPT-OAuth provider reads from ~/.codex/auth.json so a ChatGPT Plus / Codex login works without any API key. readimage analyzes images via 4-provider vision…

Its SKILL.md is about 1.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 29 other files, including scripts (for example `README.md`, `scripts/chatgpt_oauth_login.py` and `scripts/create_image.py`).

It sits in Media & Creative, covering Image generation, Model routing and gateways and OAuth and OpenID Connect. It works with OpenAI, OpenRouter and MiniMax. The licence is MIT.

When your agent uses it

  • Claude needs to generate images
  • Run vision analysis with a specific provider rather than relying on Claudes built-in image understanding

Example prompts

  • “/media-tools”

Requirements

  • Python 3
  • A credential in OPENROUTER_API_KEY
  • A credential in GEMINI_API_KEY
  • Pre-approved tools (allowed-tools): Bash, Read

What it can do on your machine

Read from SKILL.md and the folder at commit dd902b1. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Bash
    • Read

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 13 files in scripts/ (Python, from the files we listed), which the agent can run.

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • openrouter.ai
    • api.openai.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • OPENROUTER_API_KEY
    • GEMINI_API_KEY
    • OPENAI_API_KEY
    • MINIMAX_API_KEY
    • DASHSCOPE_API_KEY
    • BYTEPLUS_API_KEY
    • ANTHROPIC_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Media Tools loads about 1.5k tokens when it runs. Until then it costs about 221 tokens; SKILL.md has 264 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~221
When it runs · the whole SKILL.md, loaded when a task matches
~1.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteMentions a .env fileSKILL.md:3
    /multi-line prompts, auto UTF-8 stdout, `.env` auto-load, and `--list-providers` diagnostic. Use when Claude needs to ge
  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Bash, Read

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from therichardngai-code/gpt-image-2-pro-max at commit dd902b1, republished under its MIT licence (© therichardngai-code). 264 words, ~1,506 tokens.

Download SKILL.mdSave it as .claude/skills/media-tools/SKILL.md (or your agent's skills folder). This skill also uses 26 other files; get the full folder from GitHub.
name
media-tools
description
Two CLI tools for image generation + vision analysis using goclaw's provider-chain pattern. `create_image` generates images via 7-provider chain (ChatGPT-OAuth/Codex, OpenRouter, Gemini, OpenAI, MiniMax, DashScope, BytePlus) — first available credential wins (OAuth session OR API key), embeds prompt into PNG tEXt metadata, saves to date-folder workspace. The ChatGPT-OAuth provider reads from `~/.codex/auth.json` so a ChatGPT Plus / Codex login works without any API key. `read_image` analyzes images via 4-provider vision chain (OpenRouter, Gemini, Anthropic, DashScope) — accepts file path. Supports `--prompt-file` for long/multi-line prompts, auto UTF-8 stdout, `.env` auto-load, and `--list-providers` diagnostic. Use when Claude needs to generate images or run vision analysis with a specific provider rather than relying on Claude's built-in image understanding.
allowed-tools
Bash, Read
license
MIT

Media Tools

Python ports of goclaw's create_image and read_image tools. Faithful provider-chain pattern: priority order, skip-if-no-key, first-success-wins, cascade-on-failure.

Setup

bash
.claude/skills/.venv/Scripts/python.exe -m pip install -r .claude/skills/media-tools/requirements.txt

API keys (env vars)

Set whichever providers you have keys for. Chain skips providers without keys.

powershell
# Image generation
$env:OPENROUTER_API_KEY = "..."   # OpenRouter (aggregator, OpenAI-compat)
$env:GEMINI_API_KEY     = "..."   # Google AI Studio
$env:OPENAI_API_KEY     = "..."   # OpenAI gpt-image
$env:MINIMAX_API_KEY    = "..."   # MiniMax image-01
$env:DASHSCOPE_API_KEY  = "..."   # Alibaba Qwen / Wan2.6
$env:BYTEPLUS_API_KEY   = "..."   # ByteDance Seedream

# Vision (some keys overlap with image gen)
$env:ANTHROPIC_API_KEY  = "..."   # Anthropic Claude vision

# Optional API base overrides (for proxies / regional endpoints)
$env:OPENROUTER_API_BASE = "https://openrouter.ai/api/v1"
$env:OPENAI_API_BASE     = "https://api.openai.com/v1"
# etc — see lib/env_keys.py for full list

create_image

bash
.claude/skills/.venv/Scripts/python.exe .claude/skills/media-tools/scripts/create_image.py \
  --prompt "a cyberpunk cat in neon rain, ukiyo-e style" \
  --aspect-ratio 16:9 \
  --filename-hint "neon-cat"

Options:

  • --prompt (required) — text description
  • --aspect-ratio — 1:1 (default) | 3:4 | 4:3 | 9:16 | 16:9
  • --filename-hint — kebab-slug for output file (no extension)
  • --workspace — output dir; default $WEB_TOOLS_WORKSPACE or OS temp
  • --provider — force a specific provider (skip chain): chatgpt_oauth|openrouter|gemini|openai|minimax|dashscope|byteplus
  • --provider-order — comma-separated chain override (default: chatgpt_oauth,openrouter,gemini,openai,minimax,dashscope,byteplus)
  • --reference-image PATH — seed generation with a reference image (PNG/JPG/WEBP). Repeatable, max 4. Supported by chatgpt_oauth, gemini, openai; chain auto-skips others when refs are present.

Output saved to <workspace>/generated/<YYYY-MM-DD>/<name>.png with prompt embedded in PNG tEXt metadata.

Reference-image example (image-to-image / restyle / character-consistency)
bash
.claude/skills/.venv/Scripts/python.exe .claude/skills/media-tools/scripts/create_image.py \
  --prompt "Same character, now wearing a red trench coat, neon Tokyo street at night" \
  --reference-image ./character.png \
  --aspect-ratio 9:16

read_image

bash
.claude/skills/.venv/Scripts/python.exe .claude/skills/media-tools/scripts/read_image.py \
  --path "C:/path/to/image.png" \
  --prompt "Describe this image in detail. What text appears?"

Options:

  • --prompt (required) — what to ask about the image
  • --path (required) — local image file (jpg/png/gif/webp/bmp, ≤10MB)
  • --provider — force specific provider: openrouter|gemini|anthropic|dashscope
  • --provider-order — comma-separated chain (default: openrouter,gemini,anthropic,dashscope)
  • --max-tokens — response max tokens (default 1024)

Patterns ported from goclaw

Patterngoclaw sourceskill file
Provider chain (priority + skip-no-key + first-success)media_provider_chain.golib/chain.py
OpenRouter image (OpenAI-compat + modalities)create_image.go:callImageGenAPIlib/providers/openrouter_image.py
Gemini image (native generateContent + responseModalities)create_image.go:callGeminiNativeImageGenlib/providers/gemini_image.py
OpenAI image (/images/generations)create_image.go:callStandardImageGenAPIlib/providers/openai_image.py
MiniMax image (/image_generation)create_image_minimax.golib/providers/minimax_image.py
DashScope image (async task polling)create_image_dashscope.golib/providers/dashscope_image.py
BytePlus Seedream (/api/v3/images/generations + URL download)create_image_byteplus.golib/providers/byteplus_image.py
Vision providers (4-way chain)read_image.go + provider Chat()lib/providers/*_vision.py
PNG tEXt prompt embeddingpngEmbedPromptlib/png_metadata.py
Date-folder workspace + sanitized filenamemediaFileNamelib/filename.py

Aspect-ratio mapping

Each provider has its own size convention; the skill maps --aspect-ratio to each:

aspectdashscopebyteplusminimax/openrouter/gemini
1:11024*10241024x10241:1 (passed through)
16:91280*7201280x72016:9
9:16720*1280720x12809:16
4:31024*7681024x7684:3
3:4768*1024768x10243:4

© therichardngai-code, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 26 other files (scripts) in .claude/skills/media-tools of therichardngai-code/gpt-image-2-pro-max.

  • SKILL.md
  • .env.example
  • .gitignore
  • README.md
  • requirements.txt
  • scripts/chatgpt_oauth_login.py
  • scripts/create_image.py
  • scripts/lib/__init__.py
  • scripts/lib/chain.py
  • scripts/lib/chatgpt_oauth_token.py
  • scripts/lib/dotenv_loader.py
  • scripts/lib/env_keys.py
  • scripts/lib/filename.py
  • scripts/lib/png_metadata.py
  • scripts/lib/providers/__init__.py
  • scripts/lib/providers/anthropic_vision.py
  • scripts/lib/providers/byteplus_image.py
  • scripts/lib/providers/chatgpt_oauth_image.py
  • … and 9 more

Open the folder on GitHubat commit dd902b1

Compare with similar skills

Media Tools next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Media Tools compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Media Tools this skilltherichardngai-code/gpt-image-2-pro-max101—~1.5kAutomated safety check: NotesMIT
Baoyu ImagineLeoYeAI/openclaw-master-skills2.2k—~5.1kAutomated safety check: NotesMIT
Baoyu Image GenJimLiu/baoyu-skills26k1 repos~5.3kAutomated safety check: NotesMIT
AI Image Creatorevolution-foundation/evo-nexus545—~5.1kAutomated safety check: NotesCustom licence
Generate ImageK-Dense-AI/claude-scientific-writer2.4k1 repos~3.8kAutomated safety check: NotesMIT
Baoyu Imagineguanyang/open-agent-hub975—~4.6kAutomated safety check: NotesMIT

Similar skills

  • Baoyu Imagine

    LeoYeAI/openclaw-master-skills

    AI image generation with OpenAI, Azure OpenAI, Google, OpenRouter, DashScope, MiniMax, Jimeng, Seedream and Replicate APIs.

    2.2k GitHub stars~5.1k tokensUpdated 2 mo ago
    Media & CreativeAuto-check: notes
  • Baoyu Image Gen

    JimLiu/baoyu-skills

    AI image generation with OpenAI GPT Image 2.5, Azure OpenAI, Google, OpenRouter, DashScope, Z.AI GLM-Image, MiniMax, Jimeng, Seedream, Replicate and Agnes APIs.

    26k GitHub starsUsed in 1 repo~5.3k tokens
    Media & CreativeAuto-check: notes
  • AI Image Creator

    evolution-foundation/evo-nexus

    Generates PNG images through OpenRouter models, with transparent backgrounds and reference-image edits, and describes existing images with multimodal vision.

    545 GitHub stars~5.1k tokensUpdated 4 mo ago
    Media & CreativeAuto-check: notes
  • Generate Image

    K-Dense-AI/claude-scientific-writer

    Generate or edit images with AI models through the OpenRouter Image API (Gemini, Seedream, Recraft, GPT-Image, Riverflow).

    2.4k GitHub starsUsed in 1 repo~3.8k tokens
    Media & CreativeAuto-check: notes
  • Baoyu Imagine

    guanyang/open-agent-hub

    AI image generation with OpenAI GPT Image 2, Azure OpenAI, Google, OpenRouter, DashScope, Z.AI GLM-Image, MiniMax, Jimeng, Seedream and Replicate APIs.

    975 GitHub stars~4.6k tokensUpdated yesterday
    Media & CreativeAuto-check: notes
  • Aigen Image Generation

    Peiiii/nextclaw

    Use the local aigen CLI to generate images through configured providers such as OpenRouter or OpenAI.

    260 GitHub stars~1.3k tokensUpdated yesterday
    Media & CreativeAuto-check passed

More from therichardngai-code/gpt-image-2-pro-max

  • Gpt Image 2 Pro Max

    therichardngai-code/gpt-image-2-pro-max

    Production prompt-engineering pipeline for GPT-Image-2 / OpenAI image generation.

    101 GitHub stars~1.4k tokensUpdated 4 mo ago
    Auto-check passed

Questions about Media Tools

What does Media Tools do?

Two CLI tools for image generation + vision analysis using goclaw's provider-chain pattern. Media Tools is an agent skill from therichardngai-code/gpt-image-2-pro-max. Two CLI tools for image generation + vision analysis using goclaw's provider-chain pattern.

When should I use Media Tools?

Media Tools fits situations like: Claude needs to generate images; run vision analysis with a specific provider rather than relying on Claudes built-in image understanding.

How do I install Media Tools in Claude Code?

Run `npx skills add therichardngai-code/gpt-image-2-pro-max --skill media-tools -a claude-code`. Or copy the skill folder (.claude/skills/media-tools in therichardngai-code/gpt-image-2-pro-max) into .claude/skills/media-tools in your project. Claude Code loads it when a task matches its description.

How do I install Media Tools in Codex?

Run `npx skills add therichardngai-code/gpt-image-2-pro-max --skill media-tools -a codex`. Or copy the skill folder (.claude/skills/media-tools in therichardngai-code/gpt-image-2-pro-max) into .agents/skills/media-tools in your project. Codex loads it when a task matches its description.

Can I use Media Tools in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add therichardngai-code/gpt-image-2-pro-max --skill media-tools -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/media-tools, .gemini/skills/media-tools, .github/skills/media-tools and .opencode/skills/media-tools in your project.

What does Media Tools need to run?

Going by SKILL.md and its folder, Media Tools needs Python for the scripts in its folder, the command-line tools its instructions call (python) and credentials named OPENROUTER_API_KEY, GEMINI_API_KEY, OPENAI_API_KEY and MINIMAX_API_KEY. Our summary lists: Python 3; A credential in OPENROUTER_API_KEY; A credential in GEMINI_API_KEY. Its frontmatter pre-approves these tools: Bash, Read.

Does Media Tools access the network?

SKILL.md names 2 domains. In commands or code: openrouter.ai and api.openai.com; the agent is likely to contact these when it follows the instructions. This is read from the text; nothing was executed.

Is Media Tools safe to install?

Our automated static check of SKILL.md found notes only (mentions a .env file; pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Media Tools use?

Media Tools is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Media Tools use?

About 1.5k tokens (SKILL.md is roughly 6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Media Tools?

Skills that share tags, products or a category with Media Tools: Baoyu Imagine (LeoYeAI/openclaw-master-skills, 2.2k stars), Baoyu Image Gen (JimLiu/baoyu-skills, 26k stars), AI Image Creator (evolution-foundation/evo-nexus, 545 stars) and Generate Image (K-Dense-AI/claude-scientific-writer, 2.4k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Media Tools?

therichardngai-code (a GitHub user) maintains it in therichardngai-code/gpt-image-2-pro-max, which has 101 GitHub stars. The repository holds 2 skills in this directory. The repository was last updated on May 21, 2026.

Source: therichardngai-code/gpt-image-2-pro-max on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.