Agent skill

Image Gen

by Sma1lboy in Sma1lboy/rove

Generate a single image from a text prompt using the MiniMax image generation API.

MITAuto-check passedMedia & Creative

Install Image Gen

skills CLI
$ npx skills add Sma1lboy/rove --skill image-gen -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Sma1lboy/rove image-gen --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Sma1lboy/rove.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/image-gen .claude/skills/image-gen && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
image-gen
GitHub stars
160
Token cost
~927 tokens
SKILL.md length
264 words
Files
2 (incl. scripts)
Skills in repo
13
Repo updated
First seen
Licence
MIT

At a glance

Generate a single image from a text prompt using the MiniMax image generation API.

  • Works in 4 steps: Santorini paper-cut collage (UI style,… → Square UI icon (1:1) → Landscape banner (16:9) → …
  • The user asks to create
  • SKILL.md covers Output, Prompt writing tips and Examples
  • Runs Python scripts from its folder; calls python3

What it does

Image Gen is an agent skill from Sma1lboy/rove. Generate a single image from a text prompt using the MiniMax image generation API. Use when the user asks to create, generate, or render an image from a prompt and explicit width/height.

Its SKILL.md is about 930 tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including scripts (for example `scripts/image-gen.py`).

It sits in Media & Creative, covering Image generation. It works with MiniMax. The repository describes itself as: Rove — the agent multiplexer for your terminal. Run coding agents on parallel tasks with isolated worktrees and persistent sessions. The licence is MIT.

When your agent uses it

  • The user asks to create
  • Render an image from a prompt and explicit width/height

Example prompts

  • “/image-gen”

Requirements

  • Python 3

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Santorini paper-cut collage (UI style, non-photorealistic)
  2. Square UI icon (1:1)
  3. Landscape banner (16:9)
  4. Portrait poster (2:3)

What it can do on your machine

Read from SKILL.md and the folder at commit 139a8cb. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Image Gen loads about 927 tokens when it runs. Until then it costs about 49 tokens; SKILL.md has 264 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~49
When it runs · the whole SKILL.md, loaded when a task matches
~927

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from Sma1lboy/rove at commit 139a8cb, republished under its MIT licence (© Sma1lboy). 264 words, ~927 tokens.

Download SKILL.mdSave it as .claude/skills/image-gen/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
image-gen
description
Generate a single image from a text prompt using the MiniMax image generation API. Use when the user asks to create, generate, or render an image from a prompt and explicit width/height.

image-gen

Generate one image from a text prompt via MiniMax image-01 model. The output file is written to the filename you pass as the 4th argument (relative to the caller's cwd).

The script path is resolved relative to this SKILL.md (not the caller's cwd), so it works from any directory. It accepts exactly four positional arguments: prompt, width, height, output_filename.

bash
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
python3 "$SCRIPT_DIR/scripts/image-gen.py" "<prompt>" <width> <height> <output_filename>
ArgumentTypeConstraints
promptstring≤ 1500 chars; describe subject + style + mood
widthint512–2048, must be a multiple of 8
heightint512–2048, must be a multiple of 8
output_filenamestringFilename to write the image to, relative to cwd (e.g. img1.png). The file extension is preserved as-is.

Output

  • File: <output_filename> (relative to cwd — the directory the command is invoked from)
  • Format: raw image bytes base64-decoded from the API response; extension is whatever you pass (e.g. .png, .jpeg)
  • Always overwrites an existing file at the same path

After running, confirm the file exists with ls -lh <output_filename> and report the absolute path to the user.

Prompt writing tips

Be explicit about style to avoid photorealistic defaults. Example style keywords:

  • Flat / UI / vector: UI style, flat illustration, vector art, minimalist, clean cut-and-paste aesthetic, no photography no realistic textures
  • Collage / cut-and-paste: paper cut collage, layered flat cut-out shapes, torn paper edges, mixed media
  • Vintage / retro: vintage paper collage, sun-faded tones, retro travel poster
  • Photorealistic: photorealistic, documentary photography, film grain

Combine subject + style + palette. Example structure: "<style keywords>: <subject>, <key visual elements>, <color palette>, <negative constraints>"

Examples

1. Santorini paper-cut collage (UI style, non-photorealistic)
bash
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
python3 "$SCRIPT_DIR/scripts/image-gen.py" "UI style paper cut collage of Santorini, Greece: whitewashed houses with blue domes, ancient Greek columns, olive branches, terracotta rooftops, layered flat cut-out shapes with torn paper edges, clean vector illustration, minimalist cut-and-paste aesthetic, pastel color palette, no photography no realistic textures" 512 512 santorini.png
2. Square UI icon (1:1)
bash
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
python3 "$SCRIPT_DIR/scripts/image-gen.py" "flat vector icon of a paper airplane, minimal geometric shapes, soft blue and white palette, clean UI illustration, no shadows no gradients" 512 512 airplane.png
3. Landscape banner (16:9)
bash
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
python3 "$SCRIPT_DIR/scripts/image-gen.py" "minimalist flat illustration of mountains at sunset, layered geometric shapes, warm orange and purple palette, clean vector style, no texture no realism" 1280 720 mountains.png
4. Portrait poster (2:3)
bash
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
python3 "$SCRIPT_DIR/scripts/image-gen.py" "vintage paper collage portrait of a woman in 1920s fashion, torn paper layers, muted sepia and teal tones, art deco cut-out shapes, mixed media aesthetic" 832 1248 portrait.png

© Sma1lboy, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (scripts) in .claude/skills/image-gen of Sma1lboy/rove.

  • SKILL.md
  • scripts/image-gen.py

Open the folder on GitHubat commit 139a8cb

Compare with similar skills

Image Gen next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Image Gen compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Image Gen this skillSma1lboy/rove160—~927Automated safety check: PassMIT
AI Image Generation and Editingzhayujie/CowAgent47k—~1.3kAutomated safety check: PassMIT
Novel Artsuihe1/short-drama-production180—~1.3kAutomated safety check: NotesApache-2.0
ComfyUI Local DriverSlavaSexton/ComfyUI-Agent-Kit105—~12kAutomated safety check: PassApache-2.0
Mmx CLIagentscope-ai/OpenJudge870—~655Automated safety check: PassApache-2.0
9Router Image Generationdecolua/9router30k—~830Automated safety check: PassMIT

Similar skills

  • Generates or edits images from text prompts through a Python script that picks an image backend based on which API keys are configured.

    47k GitHub stars~1.3k tokensUpdated today
    Media & CreativeAuto-check passed
  • Novel Art

    suihe1/short-drama-production

    为 AI 短剧设计场景和叙事道具母本、光照与状态变体,产出 art.json、Markdown、评审报告和可选设定图。用于美术设定、场景与道具一致性;不代替导演调度或成片人流设计。

    180 GitHub stars~1.3k tokensUpdated 24 days ago
    Media & CreativeAuto-check: notes
  • ComfyUI Local Driver

    SlavaSexton/ComfyUI-Agent-Kit

    Drives a local ComfyUI install over its HTTP API to generate and edit images, video and audio, with per-model prompt recipes and workflow guidance.

    105 GitHub stars~12k tokensUpdated 1 mo ago
    Media & CreativeAuto-check passed
  • Mmx CLI

    agentscope-ai/OpenJudge

    Generate text, images, video, speech, and music via the MiniMax AI platform.

    870 GitHub stars~655 tokensUpdated 28 days ago
    Media & CreativeAuto-check passed
  • Generates images through a 9Router gateway's image endpoint, with model discovery, the request fields and per-provider quirks for OpenAI, Gemini, MiniMax and others.

    30k GitHub stars~830 tokensUpdated yesterday
    Media & CreativeAuto-check passed
  • Media Production

    leon-ai/leon

    Generates images, audio and video through Leon's media tools, joins them with FFmpeg, checks the output and attaches playable files for the owner.

    18k GitHub stars~1k tokensUpdated today
    Media & CreativeAuto-check passed

More from Sma1lboy/rove

All 13 skills in this repo
  • Hyperframes Design

    Sma1lboy/rove

    Non-animation creative direction for HyperFrames videos. An agent skill from Sma1lboy/rove.

    160 GitHub stars~1.2k tokensUpdated yesterday
    Auto-check passed
  • Changelog Generator

    Sma1lboy/rove

    Draft Rove release notes as Changesets. An agent skill from Sma1lboy/rove.

    160 GitHub stars~1.5k tokensUpdated yesterday
    Auto-check passed
  • File Issue

    Sma1lboy/rove

    Turn a rough idea, a bug, or a batch of unsolved problems into well-structured GitHub issue(s) and file them with gh, auto-classifying type + labels from the content (recommend, then confirm) and…

    160 GitHub stars~2.4k tokensUpdated yesterday
    Auto-check passed
  • Rove

    Sma1lboy/rove

    A skill your agent uses when controlling Rove tasks, parallel coding attempts, hosted agent sessions, task lifecycle, or the daemon-owned issue tracker from a shell.

    160 GitHub stars~8.9k tokensUpdated yesterday
    Auto-check passed
  • Unslop

    Sma1lboy/rove

    Strip AI writing tells from Rove's user-facing prose — README, the docs/ pages that sync to docs.rove.run, landing copy, and release notes.

    160 GitHub stars~1.3k tokensUpdated yesterday
    Auto-check passed
  • Why

    Sma1lboy/rove

    A skill your agent uses for 'why does X work this way', 'why we picked Y', '为什么这么设计', design rationale, regressions, or where a magic number came from.

    160 GitHub stars~5.6k tokensUpdated yesterday
    Auto-check passed

Works with

Questions about Image Gen

What does Image Gen do?

Generate a single image from a text prompt using the MiniMax image generation API. Image Gen is an agent skill from Sma1lboy/rove. Generate a single image from a text prompt using the MiniMax image generation API.

When should I use Image Gen?

Image Gen fits situations like: the user asks to create; render an image from a prompt and explicit width/height.

How do I install Image Gen in Claude Code?

Run `npx skills add Sma1lboy/rove --skill image-gen -a claude-code`. Or copy the skill folder (.claude/skills/image-gen in Sma1lboy/rove) into .claude/skills/image-gen in your project. Claude Code loads it when a task matches its description.

How do I install Image Gen in Codex?

Run `npx skills add Sma1lboy/rove --skill image-gen -a codex`. Or copy the skill folder (.claude/skills/image-gen in Sma1lboy/rove) into .agents/skills/image-gen in your project. Codex loads it when a task matches its description.

Can I use Image Gen in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Sma1lboy/rove --skill image-gen -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/image-gen, .gemini/skills/image-gen, .github/skills/image-gen and .opencode/skills/image-gen in your project.

What does Image Gen need to run?

Going by SKILL.md and its folder, Image Gen needs Python for the scripts in its folder and the command-line tools its instructions call (python3). Our summary lists: Python 3.

Does Image Gen access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Image Gen safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Image Gen use?

Image Gen is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Image Gen use?

About 927 tokens (SKILL.md is roughly 3.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Image Gen?

Skills that share tags, products or a category with Image Gen: AI Image Generation and Editing (zhayujie/CowAgent, 47k stars), Novel Art (suihe1/short-drama-production, 180 stars), ComfyUI Local Driver (SlavaSexton/ComfyUI-Agent-Kit, 105 stars) and Mmx CLI (agentscope-ai/OpenJudge, 870 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Image Gen?

Sma1lboy (a GitHub user) maintains it in Sma1lboy/rove, which has 160 GitHub stars. The repository holds 13 skills in this directory. The repository was last updated on October 8, 2026.

Source: Sma1lboy/rove on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.