Agent skill

Chatgpt Image

by ManiacMike in ManiacMike/chatgpt-endless-canvas

Generate images through the ChatGPT web app (chatgpt.com) via a local infinite-canvas board — queued jobs, scheduled batches, annotate-to-edit, style-reference generation, lineage.

MITAuto-check passedMedia & Creative

Install Chatgpt Image

skills CLI
$ npx skills add ManiacMike/chatgpt-endless-canvas --skill chatgpt-image -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install ManiacMike/chatgpt-endless-canvas chatgpt-image --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/ManiacMike/chatgpt-endless-canvas.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/chatgpt-image .claude/skills/chatgpt-image && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
chatgpt-image
GitHub stars
116
Token cost
~1.8k tokens
SKILL.md length
763 words
Files
1
Skills in repo
1
Repo updated
First seen
Licence
MIT

At a glance

Generate images through the ChatGPT web app (chatgpt.com) via a local infinite-canvas board — queued jobs, scheduled batches, annotate-to-edit, style-reference generation, lineage.

  • Works in 4 steps: Debug Chrome ready? → Board server running? Use the idempotent… → Submit (POST /api/generate, pick by… → …
  • The user asks to generate images (生图
  • SKILL.md covers Standard flow (every request), Boards, Board UI (tell the user when… and Fallback: direct script…, plus 1 more section
  • Calls curl, bash and git

What it does

Chatgpt Image is an agent skill from ManiacMike/chatgpt-endless-canvas. Generate images through the ChatGPT web app (chatgpt.com) via a local infinite-canvas board — queued jobs, scheduled batches, annotate-to-edit, style-reference generation, lineage. Use when the user asks to generate images ("生图", "draw with ChatGPT", "open the board", "new canvas"), or wants images consistent with a reference style. Requires a debug Chrome logged into chatgpt.com (CDP port 9222).

Its SKILL.md is about 1.8k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Media & Creative, covering Image generation. It works with OpenAI. The repository describes itself as: Infinite-canvas image board driven by the ChatGPT web app. The licence is MIT.

When your agent uses it

  • The user asks to generate images (生图
  • Draw with ChatGPT
  • Wants images consistent with a reference style

Example prompts

  • “draw with ChatGPT”
  • “open the board”
  • “new canvas”
  • “/chatgpt-image”

Requirements

  • Python 3

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Debug Chrome ready?
  2. Board server running? Use the idempotent launcher (safe from any
  3. Submit (POST /api/generate, pick by scenario)
  4. Await results: poll GET /api/state → jobs until the job is done

What it can do on your machine

Read from SKILL.md and the folder at commit 7783f02. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • curl
    • bash
    • git
    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use curl and git, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Chatgpt Image loads about 1.8k tokens when it runs. Until then it costs about 103 tokens; SKILL.md has 763 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~103
When it runs · the whole SKILL.md, loaded when a task matches
~1.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from ManiacMike/chatgpt-endless-canvas at commit 7783f02, republished under its MIT licence (© ManiacMike). 763 words, ~1,751 tokens.

Download SKILL.mdSave it as .claude/skills/chatgpt-image/SKILL.md (or your agent's skills folder).
name
chatgpt-image
description
Generate images through the ChatGPT web app (chatgpt.com) via a local infinite-canvas board — queued jobs, scheduled batches, annotate-to-edit, style-reference generation, lineage. Use when the user asks to generate images ("生图", "draw with ChatGPT", "open the board", "new canvas"), or wants images consistent with a reference style. Requires a debug Chrome logged into chatgpt.com (CDP port 9222).

ChatGPT image generation (board-first)

Project home: ${CHATGPT_IMAGE_GEN_HOME:-~/Workspace/chatgpt-endless-canvas} (referred to as $PROJ below — resolve it once per session). Data root: ~/Documents/chatgpt-endless-image-gen/.

Bootstrap (skill install ≠ project install). This file only teaches the workflow; the project code + venv must exist at $PROJ. Check once per session, and set it up if missing:

bash
[ -f "$PROJ/board_server.py" ] || {
  git clone https://github.com/ManiacMike/chatgpt-endless-canvas.git "$PROJ" &&
  python3 -m venv "$PROJ/.venv" &&
  "$PROJ/.venv/bin/pip" install -r "$PROJ/requirements.txt"
}

Core rule: submit ALL generation requests through the board server API (POST /api/generate) instead of invoking the Python script directly — the server serializes jobs, spaces batches with random pauses (rate-limit protection), streams results onto the canvas, and records lineage. Call the script directly only as a fallback when the server cannot run, or for --grab-only recovery.

Standard flow (every request)

  1. Debug Chrome ready?

    bash
    curl -s --max-time 2 http://127.0.0.1:9222/json/version

    On failure → ask the user to run bash $PROJ/launch-chrome-debug.sh, log into chatgpt.com in the opened window, and keep it open. Do NOT run it for them in the background — the window needs to stay interactive.

  2. Board server running? Use the idempotent launcher (safe from any session; already-running → no-op; starts detached via nohup so it outlives this session):

    bash
    bash $PROJ/start.sh

    It prints the ACTUAL url (board started: http://127.0.0.1:<port>) — the port may NOT be 8090: ports occupied by other programs (including stale pre-rename copies of this project) are skipped automatically. Use the printed url as $BOARD for ALL API calls below, and open $BOARD if the user doesn't have the board open. To re-discover it later, GET /api/health must return "app": "chatgpt-endless-canvas" (also in ~/Documents/chatgpt-endless-image-gen/server.json alongside pid/port) — a port that answers /api/state but not /api/health is NOT this server. Do NOT run board_server.py with run_in_background — that ties the server to this session. Log: ~/Documents/chatgpt-endless-image-gen/board.log. Unfinished jobs survive a server restart (persisted + requeued), so if the environment kills the detached process, just rerun start.sh.

  3. Submit (POST /api/generate, pick by scenario):

    • Plain: {"prompt": "specific English description"}
    • Multiple images: {"prompts": ["...", ...]} (≤50; the server queues them with random 30–120 s gaps and runs up to 3 in parallel (BOARD_WORKERS), each in its own dedicated ChatGPT tab — ALWAYS use this for batches, never loop yourself. The debug Chrome holding several chatgpt.com tabs is expected; don't close them.)
    • Style reference from a board image: add "name": "<filename>"
    • Local file as reference: first POST /api/upload?name=<file> with raw bytes (curl --data-binary @file.png), then generate with the returned name
    • Edit by annotations: POST /api/regenerate {"name":"<image>"} — needs pending annotations (you may write them for the user via POST /api/annotations: {image: [{id,x,y,w,h,note,status:"pending"}]}, coords normalized 0-1, point marks have w=h=0)
    • Delete from the board: POST /api/delete {"name":"<image>"} — removes the file and its layout/annotation/lineage entries (409 while a job is still using the image; confirm with the user before deleting)
  4. Await results: poll GET /api/state → jobs until the job is done (output = filename) or error (reason included). ~1-5 min per image. For batches don't block — tell the user images will appear on the board as they finish. Files land in the active board dir (state.board.dir). Each job records conversationId — the ChatGPT chat uuid it ran in (its stable identity; interrupted jobs are auto-recovered from that conversation on server restart instead of regenerating).

Show full SKILL.md (262 more words)Show less

Boards

  • One board = one self-contained directory (images + layout/annotations/ lineage JSON). Registry: ~/Documents/chatgpt-endless-image-gen/boards.json.
  • GET /api/boards lists; POST /api/boards {"action":"create","name":"test"} → dir <data root>/test (named boards use the name, unnamed use a timestamp; dir may point anywhere — an existing directory is "opened"); {"action":"open","id":"..."} switches.
  • When the user says "new canvas / switch canvas / open directory X" → call the API; subsequent generations follow the new active board automatically.

Board UI (tell the user when relevant)

Wheel zoom, drag-empty pan, drag cards, double-click full size, drag local images in. Card buttons: 标注 (annotate: box + note) → 改图 (regenerate), 参考生图 (style-reference; one prompt per line — multiple lines become a scheduled batch), 删除 (delete, with confirm). Lineage: solid "↳ 改自" (edit), dashed "☆ 参考…风格" (ref).

Fallback: direct script (server unusable only)

bash
PY="$PROJ/.venv/bin/python"; [ -x "$PY" ] || PY=python3   # needs playwright
"$PY" $PROJ/generate_chatgpt_image.py \
  --prompt "..." --output /abs/path/out.png \
  [--reference ref.png] [--timeout 240] [--grab-only]

Progress on stderr, exit 0 = success; set Bash timeout ≥ --timeout + 60 s. If the script timed out but the image finished on the ChatGPT page, re-grab with --grab-only (pass any placeholder --prompt); add --conversation <uuid> (from the job's conversationId) to grab from that specific chat.

Troubleshooting

  • "ChatGPT not reachable" → debug Chrome not running; step 1.
  • "could not find the ChatGPT message box" → not logged into chatgpt.com.
  • "stale CDP state" → quit the debug Chrome fully, rerun the launch script (login persists).
  • Job error / prompt sent but no reply at all → ChatGPT silent rate limiting; retry later. The batch scheduler's random gaps exist to avoid this.
  • All selectors failing → ChatGPT web redesign; update _find_composer / _submit / _IMAGE_JS in generate_chatgpt_image.py.

Env knobs: BOARD_PORT (8090, auto-increments if taken), BOARD_PORT_TRIES (20), BOARD_WORKERS (3 parallel generations), IMAGE_GEN_DATA (data root), CHATGPT_CDP_URL (9222), BATCH_INTERVAL ("30-120"), BOARD_DIR (single-board mode), CHATGPT_IMAGE_GEN_HOME (project location).

© ManiacMike, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/chatgpt-image of ManiacMike/chatgpt-endless-canvas.

Open the folder on GitHubat commit 7783f02

Compare with similar skills

Chatgpt Image next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Chatgpt Image compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Chatgpt Image this skillManiacMike/chatgpt-endless-canvas116—~1.8kAutomated safety check: PassMIT
AI Image Generation and Editingzhayujie/CowAgent47k—~1.3kAutomated safety check: PassMIT
GPT Image Generation CLIwuyoscar/GPT-Image2-Skill5.7k—~2.5kAutomated safety check: NotesMIT
Imagegentheowenyoung/home1154 repos~4.8kAutomated safety check: PassApache-2.0
Openai Image Gentrpc-group/trpc-agent-go1.9k12 repos~843Automated safety check: PassApache-2.0
Image Generationonyx-dot-app/onyx32k1 repos~1.7kAutomated safety check: PassCustom licence

Similar skills

  • Generates or edits images from text prompts through a Python script that picks an image backend based on which API keys are configured.

    47k GitHub stars~1.3k tokensUpdated today
    Media & CreativeAuto-check passed
  • GPT Image Generation CLI

    wuyoscar/GPT-Image2-Skill

    Generates and edits images with GPT Image 2 or 2.5 through a packaged CLI and a prompt gallery, after settling which model fits the request.

    5.7k GitHub stars~2.5k tokensUpdated 10 days ago
    Media & CreativeAuto-check: notes
  • Imagegen

    theowenyoung/home

    Generate or edit raster images when the task benefits from AI-created bitmap visuals such as photos, illustrations, textures, sprites, mockups, or transparent-background cutouts.

    115 GitHub starsUsed in 4 repos~4.8k tokens
    Media & CreativeAuto-check passed
  • Openai Image Gen

    trpc-group/trpc-agent-go

    Batch-generate images via OpenAI Images API. An agent skill from trpc-group/trpc-agent-go.

    1.9k GitHub starsUsed in 12 repos~843 tokens
    Media & CreativeAuto-check passed
  • Image Generation

    onyx-dot-app/onyx

    Generate or edit raster images (photos, illustrations, textures, sprites, mockups, logos, infographics) using the workspace's configured image-generation provider via onyx-cli image.

    32k GitHub starsUsed in 1 repo~1.7k tokens
    Media & CreativeAuto-check passed
  • BlockRun Image Generation

    BlockRunAI/ClawRouter

    Generates or edits images through ClawRouter's local image API, with a choice of models and sizes and payment handled automatically through x402.

    6.6k GitHub stars~2.1k tokensUpdated 4 days ago
    Media & CreativeAuto-check passed

Works with

Questions about Chatgpt Image

What does Chatgpt Image do?

Generate images through the ChatGPT web app (chatgpt.com) via a local infinite-canvas board — queued jobs, scheduled batches, annotate-to-edit, style-reference generation, lineage. Chatgpt Image is an agent skill from ManiacMike/chatgpt-endless-canvas.com) via a local infinite-canvas board — queued jobs, scheduled batches, annotate-to-edit, style-reference generation, lineage.

When should I use Chatgpt Image?

Chatgpt Image fits situations like: the user asks to generate images (生图; draw with ChatGPT; wants images consistent with a reference style.

How do I install Chatgpt Image in Claude Code?

Run `npx skills add ManiacMike/chatgpt-endless-canvas --skill chatgpt-image -a claude-code`. Or copy the skill folder (skills/chatgpt-image in ManiacMike/chatgpt-endless-canvas) into .claude/skills/chatgpt-image in your project. Claude Code loads it when a task matches its description.

How do I install Chatgpt Image in Codex?

Run `npx skills add ManiacMike/chatgpt-endless-canvas --skill chatgpt-image -a codex`. Or copy the skill folder (skills/chatgpt-image in ManiacMike/chatgpt-endless-canvas) into .agents/skills/chatgpt-image in your project. Codex loads it when a task matches its description.

Can I use Chatgpt Image in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ManiacMike/chatgpt-endless-canvas --skill chatgpt-image -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/chatgpt-image, .gemini/skills/chatgpt-image, .github/skills/chatgpt-image and .opencode/skills/chatgpt-image in your project.

What does Chatgpt Image need to run?

Going by SKILL.md and its folder, Chatgpt Image needs the command-line tools its instructions call (curl, bash, git and python3). Our summary lists: Python 3.

Does Chatgpt Image access the network?

SKILL.md contains no URLs. Its commands use curl and git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Chatgpt Image safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Chatgpt Image use?

Chatgpt Image is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Chatgpt Image use?

About 1.8k tokens (SKILL.md is roughly 7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Chatgpt Image?

Skills that share tags, products or a category with Chatgpt Image: AI Image Generation and Editing (zhayujie/CowAgent, 47k stars), GPT Image Generation CLI (wuyoscar/GPT-Image2-Skill, 5.7k stars), Imagegen (theowenyoung/home, 115 stars) and Openai Image Gen (trpc-group/trpc-agent-go, 1.9k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Chatgpt Image?

ManiacMike (a GitHub user) maintains it in ManiacMike/chatgpt-endless-canvas, which has 116 GitHub stars. The repository was last updated on July 18, 2026.

Source: ManiacMike/chatgpt-endless-canvas on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.