Agent skill

AI Image Generation

by 0xsline in 0xsline/OpenChatCut

Generates still images through the submit_image tool, choosing among Fal.ai, gpt-image-2, nano-banana, MiniMax image-01 and Grok Imagine by configured keys.

AGPL-3.0Auto-check passedMedia & Creative

Install AI Image Generation

skills CLI
$ npx skills add 0xsline/OpenChatCut --skill image-gen -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install 0xsline/OpenChatCut image-gen --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/0xsline/OpenChatCut.git skills-src && mkdir -p .claude/skills && cp -r skills-src/src/agent/skills/image-gen .claude/skills/image-gen && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
image-gen
GitHub stars
2.2k
Token cost
~1.3k tokens
SKILL.md length
437 words
Files
6 (incl. references)
Skills in repo
31
Repo updated
First seen
Licence
AGPL-3.0

At a glance

Generates still images through the submit_image tool, choosing among Fal.ai, gpt-image-2, nano-banana, MiniMax image-01 and Grok Imagine by configured keys.

  • Generating a single image or still from a text prompt
  • SKILL.md covers Model Selection, Tool Params, Defaults and Ask Before Submit, plus 3 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • Making an image that follows reference pictures closely

What it does

The skill generates images through a submit_image tool, and only with providers whose keys are configured. It offers Fal.ai with a chosen Fal model, gpt-image-2, nano-banana, MiniMax image-01 and xAI Grok Imagine, each with its own reference file that the agent must read before generating. gpt-image-2 is the default when its key is on, reference-heavy requests go to nano-banana, and a request that names MiniMax, or a setup with only the MiniMax key, uses image-01.

Tool parameters include aspect ratio (16:9 by default), image size from 512px up to 4K depending on the model, width and height, a quality setting for gpt-image-2, reference asset ids, a library name for the result and a count of images. The agent should make one clear still per request unless you ask for variants. A table lists how many reference images each model accepts, and the Fal limits depend on the model chosen.

When your agent uses it

  • Generating a single image or still from a text prompt
  • Making an image that follows reference pictures closely
  • Choosing between image models according to the configured providers
  • Producing images in a specific aspect ratio and size

Example prompts

  • “Generate a 16:9 still of a lighthouse at dusk for the video intro.”
  • “Create an image of our mascot using the reference images in the project library.”
  • “Make a poster with the headline Open Studio in bold type.”
  • “Use MiniMax to generate a live-style still of a night market.”

Requirements

  • An API key for at least one of the supported image providers

What it can do on your machine

Read from SKILL.md and the folder at commit 2e6f4a2. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are typescript).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

AI Image Generation loads about 1.3k tokens when it runs, and up to ~3.5k if it reads all its reference files. Until then it costs about 45 tokens; SKILL.md has 437 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~45
When it runs · the whole SKILL.md, loaded when a task matches
~1.3k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~3.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from 0xsline/OpenChatCut at commit 2e6f4a2, republished under its AGPL-3.0 licence (© 0xsline). 437 words, ~1,349 tokens.

Download SKILL.mdSave it as .claude/skills/image-gen/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.
name
image-gen
description
AI image generation via Fal.ai, gpt-image-2, nano-banana, MiniMax image-01, and xAI Grok Imagine. Use when the user wants to generate or create an image / picture / still.
user-invocable
true

Image Gen

Generate AI images via submit_image (configured provider keys only). Prefer one clear still per request unless the user asked for variants.

Model Selection

ModelReferenceStrengthsMax refs
fal + falModelreferences/fal.mdExplicit Fal catalog; see tool schema for per-model limitsModel-specific
gpt-image-2references/gpt-image-2.mdBest text rendering, strongest prompt adherence16
nano-bananareferences/nano-banana.mdStrongest reference-image fidelity14
image-01references/image-01.mdMiniMax stills / live style; one subject reference via R21
grok-imaginereferences/grok-imagine.mdxAI Grok Imagine; text-to-image, ≤4 outputs, 1K/2K0
  • If Fal.ai is selected or requested, use model: "fal" and the requested falModel or saved Fal default from capabilities. Ask if none is selected. Native-provider defaults and controls below do not apply to Fal.
  • Default: gpt-image-2 when that key is on.
  • Reference-heavy → nano-banana.
  • User named MiniMax / only MiniMax image key on → image-01.
  • Respect capabilities: do not call a model whose vendor is not configured.

IMPORTANT: Before generating, READ the chosen model's reference.

Tool Params

ParamValuesDefault
aspectRatio1:1, 16:9, 9:16, 4:3, 3:4, 3:2, 2:3, 4:5, 5:4, 21:916:9
imageSize512px, 1K, 2K, 4K (model-specific)1K
width / heightGPT Image: 512–3840, /16; MiniMax: 512–2048, /8—
qualitylow, medium, high, auto (gpt-image-2 only)high
referenceAssetIdsArray of project asset ids — backend resolves bytes server-side—
nameShort descriptive asset name shown in the library—
countNumber of images to generate (1–10; image-01 max 9)1
promptOptimizerMiniMax image-01 only — prompt_optimizerfalse
seedMiniMax image-01 only—
maskAssetId, background, moderation, inputFidelityGPT Image edit/output controls—
outputFormat, outputCompressionGPT Image PNG/JPEG/WebP controlsPNG

Defaults

  • Aspect ratio: 16:9. If the project composition is not 16:9, ASK the user which aspect ratio they want before generating.
  • Size: 1K.
Show full SKILL.md (166 more words)Show less

Ask Before Submit

  • Never auto-upgrade size.
  • Only pass imageSize: "2K" or "4K" when the user explicitly asks. Warn that 2K/4K are EXPERIMENTAL and may be slower.

Reference Images

Use when the user provides source material to edit, blend, or use as visual guidance (e.g. "change the background", "combine these into a poster").

  • Pass project asset ids via referenceAssetIds. The backend fetches and encodes them server-side — never pull the asset bytes yourself.
  • When the user @-references an image asset, pass its id directly in referenceAssetIds.
  • Formats accepted by backend: png, jpeg, webp, svg (auto-rasterized to png), heic, heif. Each ≤ 50MB.

Run

ts
// Basic generation
submit_image({
  model: "gpt-image-2",
  prompt: "a cute orange cat",
  name: "Cat",
});

// With quality (gpt-image-2 only)
submit_image({
  model: "gpt-image-2",
  prompt: "hero poster with bold title",
  quality: "high",
  name: "Hero Poster",
});

// With reference images — pass project asset ids; backend resolves bytes
submit_image({
  model: "gpt-image-2",
  prompt: "change background to beach",
  referenceAssetIds: ["<assetId>"],
  name: "Beach Edit",
});

// Reference-heavy with nano-banana
submit_image({
  model: "nano-banana",
  prompt: "composite poster",
  referenceAssetIds: ["<id1>", "<id2>"],
  name: "Composite",
});

// Multiple images
submit_image({
  model: "gpt-image-2",
  prompt: "product shots",
  count: 3,
  name: "Product",
});

// MiniMax (optional single subject reference; R2 must be configured for refs)
submit_image({
  model: "image-01",
  prompt: "matte product bottle on marble, soft studio light",
  name: "Bottle still",
  promptOptimizer: false,
});

OpenChatCut’s submit_image may return completed pool assets synchronously depending on the provider path. If a jobId is returned, use track_progress; otherwise treat the asset ids in the result as done.

Rules

  • Always provide name with a short descriptive asset name.
  • Before submitting, briefly tell the user what you're about to generate — especially when generating multiple images.
  • Only call models whose vendor key is configured (capabilities prompt).

© 0xsline, AGPL-3.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 5 other files (references) in src/agent/skills/image-gen of 0xsline/OpenChatCut.

  • SKILL.md
  • references/fal.md
  • references/gpt-image-2.md
  • references/grok-imagine.md
  • references/image-01.md
  • references/nano-banana.md

Open the folder on GitHubat commit 2e6f4a2

Compare with similar skills

AI Image Generation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

AI Image Generation compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
AI Image Generation this skill0xsline/OpenChatCut2.2k—~1.3kAutomated safety check: PassAGPL-3.0
9Router Image Generationdecolua/9router30k—~830Automated safety check: PassMIT
Keirouter Imagemydisha/keirouter147—~691Automated safety check: PassMIT
AI Image Generation and Editingzhayujie/CowAgent47k—~1.3kAutomated safety check: PassMIT
Forge Media Route Layer0x0funky/agent-sprite-forge4.4k—~2.2kAutomated safety check: PassMIT
Video-to-Sprite Animation Generator0x0funky/agent-sprite-forge4.4k—~3.8kAutomated safety check: PassMIT

Similar skills

  • Generates images through a 9Router gateway's image endpoint, with model discovery, the request fields and per-provider quirks for OpenAI, Gemini, MiniMax and others.

    30k GitHub stars~830 tokensUpdated yesterday
    Media & CreativeAuto-check passed
  • Keirouter Image

    mydisha/keirouter

    Generate images via KeiRouter /v1/images/generations using OpenAI DALL-E / Gemini Imagen / FLUX / MiniMax / Stability AI / Fal.ai models.

    147 GitHub stars~691 tokensUpdated 28 days ago
    Media & CreativeAuto-check passed
  • Generates or edits images from text prompts through a Python script that picks an image backend based on which API keys are configured.

    47k GitHub stars~1.3k tokensUpdated today
    Media & CreativeAuto-check passed
  • Forge Media Route Layer

    0x0funky/agent-sprite-forge

    Generates an image or an image-to-video clip through a configured provider API or a signed-in Codex or Grok CLI, and reports the route, file, hash and cost estimate.

    4.4k GitHub stars~2.2k tokensUpdated 3 days ago
    Media & CreativeAuto-check passed
  • Video-to-Sprite Animation Generator

    0x0funky/agent-sprite-forge

    Turns one approved master still into a full set of animated action clips for a character, one image-to-video take per action, then packages them for a game engine.

    4.4k GitHub stars~3.8k tokensUpdated 3 days ago
    Media & CreativeAuto-check passed
  • Media Production

    leon-ai/leon

    Generates images, audio and video through Leon's media tools, joins them with FFmpeg, checks the output and attaches playable files for the owner.

    18k GitHub stars~1k tokensUpdated today
    Media & CreativeAuto-check passed

More from 0xsline/OpenChatCut

All 31 skills in this repo
  • OpenChatCut Video Editing

    0xsline/OpenChatCut

    Connects an MCP-capable agent to the local OpenChatCut video editor to inspect and edit projects through draft edit sessions, with manual approval by default.

    2.2k GitHub starsUsed in 1 repo~655 tokens
    Auto-check passed
  • Video Shader Generator

    0xsline/OpenChatCut

    Generates WebGL shaders for video effects, transitions, masks and color grades in the OpenChatCut editor, trying built-in catalog effects such as zoom before making anything new.

    2.2k GitHub starsUsed in 1 repo~3.2k tokens
    Auto-check passed
  • Livestream to Clips

    0xsline/OpenChatCut

    Cuts a livestream recording into evidence-backed, platform-ready clips by combining transcript, visual, audio and genre-specific signals.

    2.2k GitHub stars~2.7k tokensUpdated 2 days ago
    Auto-check passed
  • Music Generation

    0xsline/OpenChatCut

    Generates instrumentals, songs, soundtracks and covers through Mureka, MiniMax, Atlas Cloud or Sonilo using the `submit_music` tool.

    2.2k GitHub stars~1.1k tokensUpdated 2 days ago
    Auto-check passed
  • AI Video Generation

    0xsline/OpenChatCut

    Submits AI video generation jobs to Fal.ai, Seedance, Kling, MiniMax Hailuo, xAI Grok Imagine or OFox for text-to-video, image-to-video, transitions and clip extension.

    2.2k GitHub stars~4.3k tokensUpdated 2 days ago
    Auto-check passed
  • Generates text-to-speech narration and custom sound effects for a video timeline, keeping existing voiceover in sync after visual retiming edits.

    2.2k GitHub stars~4.4k tokensUpdated 2 days ago
    Auto-check passed

Questions about AI Image Generation

What does AI Image Generation do?

Generates still images through the submit_image tool, choosing among Fal.ai, gpt-image-2, nano-banana, MiniMax image-01 and Grok Imagine by configured keys. The skill generates images through a submit_image tool, and only with providers whose keys are configured.ai with a chosen Fal model, gpt-image-2, nano-banana, MiniMax image-01 and xAI Grok Imagine, each with its own reference file that the agent must read before generating.

When should I use AI Image Generation?

AI Image Generation fits situations like: generating a single image or still from a text prompt; making an image that follows reference pictures closely; choosing between image models according to the configured providers; producing images in a specific aspect ratio and size.

How do I install AI Image Generation in Claude Code?

Run `npx skills add 0xsline/OpenChatCut --skill image-gen -a claude-code`. Or copy the skill folder (src/agent/skills/image-gen in 0xsline/OpenChatCut) into .claude/skills/image-gen in your project. Claude Code loads it when a task matches its description.

How do I install AI Image Generation in Codex?

Run `npx skills add 0xsline/OpenChatCut --skill image-gen -a codex`. Or copy the skill folder (src/agent/skills/image-gen in 0xsline/OpenChatCut) into .agents/skills/image-gen in your project. Codex loads it when a task matches its description.

Can I use AI Image Generation in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add 0xsline/OpenChatCut --skill image-gen -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/image-gen, .gemini/skills/image-gen, .github/skills/image-gen and .opencode/skills/image-gen in your project.

What does AI Image Generation need to run?

SKILL.md names no scripts, command-line tools or credentials: AI Image Generation is instructions for the agent only. Our summary lists: An API key for at least one of the supported image providers.

Does AI Image Generation access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is AI Image Generation safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does AI Image Generation use?

AI Image Generation is published under the AGPL-3.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does AI Image Generation use?

About 1.3k tokens (SKILL.md is roughly 5.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.2k tokens, read only when the agent opens those files.

What are the alternatives to AI Image Generation?

Skills that share tags, products or a category with AI Image Generation: 9Router Image Generation (decolua/9router, 30k stars), Keirouter Image (mydisha/keirouter, 147 stars), AI Image Generation and Editing (zhayujie/CowAgent, 47k stars) and Forge Media Route Layer (0x0funky/agent-sprite-forge, 4.4k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains AI Image Generation?

0xsline (a GitHub user) maintains it in 0xsline/OpenChatCut, which has 2,211 GitHub stars. The repository holds 31 skills in this directory. The repository was last updated on October 7, 2026.

Source: 0xsline/OpenChatCut on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.