Agent skill

Multi Agent Image

by kangarooking in kangarooking/kangarooking-skills

Standalone multi-agent image generation skill for Hermes. An agent skill from kangarooking/kangarooking-skills.

MITAuto-check passedMedia & Creative

Install Multi Agent Image

skills CLI
$ npx skills add kangarooking/kangarooking-skills --skill multi-agent-image -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install kangarooking/kangarooking-skills multi-agent-image --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/kangarooking/kangarooking-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/multi-agent-image .claude/skills/multi-agent-image && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
multi-agent-image
GitHub stars
656
Token cost
~2.6k tokens
SKILL.md length
895 words
Files
20 (incl. scripts, references)
Skills in repo
19
Repo updated
First seen
Licence
MIT

At a glance

Standalone multi-agent image generation skill for Hermes. An agent skill from kangarooking/kangarooking-skills.

  • Works in 8 steps: Deploy the skill → Install Python dependencies → Set API key → …
  • Tasks that involve Image generation
  • SKILL.md covers When to Use, Architecture, Setup and Core Components, plus 4 more sections
  • Runs Python scripts from its folder; calls python3 and pip; needs OPENAI_API_KEY

What it does

Multi Agent Image is an agent skill from kangarooking/kangarooking-skills. Standalone multi-agent image generation skill for Hermes. Includes an internal design compiler, GPT-Image-2 generation via apimart.ai, case library reuse, interactive reference selection, batch workflows, and style-consistent series generation.

Its SKILL.md is about 2.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 22 other files, including scripts and reference files (for example `README.md`, `references/agents/image_generator.md` and `references/agents/metadata_manager.md`).

It sits in Media & Creative, covering Image generation and Multi-agent orchestration. The repository describes itself as: My custom AI Agent skills. The licence is MIT.

When your agent uses it

  • Tasks that involve Image generation
  • Tasks that involve Multi-agent orchestration

Example prompts

  • “/multi-agent-image”

Requirements

  • Python 3
  • A credential in OPENAI_API_KEY

Workflow steps

8 steps, taken from the step headings in SKILL.md.

  1. Deploy the skill
  2. Install Python dependencies
  3. Set API key
  4. design_reasoning
  5. compiled_brief
  6. prompt
  7. prepare text for the user
  8. execute after the user chooses

What it can do on your machine

Read from SKILL.md and the folder at commit 08bbee0. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 11 files in scripts/ (Python, from the files we listed), which the agent can run.

    Shell commands in SKILL.md call:

    • python3
    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • OPENAI_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Multi Agent Image loads about 2.6k tokens when it runs, and up to ~5.3k if it reads all its reference files. Until then it costs about 66 tokens; SKILL.md has 895 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~66
When it runs · the whole SKILL.md, loaded when a task matches
~2.6k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~5.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from kangarooking/kangarooking-skills at commit 08bbee0, republished under its MIT licence (© kangarooking). 895 words, ~2,551 tokens.

Download SKILL.mdSave it as .claude/skills/multi-agent-image/SKILL.md (or your agent's skills folder). This skill also uses 19 other files; get the full folder from GitHub.
name
multi-agent-image
description
Standalone multi-agent image generation skill for Hermes. Includes an internal design compiler, GPT-Image-2 generation via apimart.ai, case library reuse, interactive reference selection, batch workflows, and style-consistent series generation.
version
3.0.0
author
Hermes Agent
license
MIT

Multi-Agent Image

multi-agent-image is a standalone Hermes skill for image generation workflows.

It is designed for cases where a simple one-line prompt is not enough. Instead of sending raw user input directly to an image model, this skill:

  1. analyzes the request,
  2. compiles it into a design-aware prompt,
  3. generates through gpt-image-2,
  4. archives the result,
  5. and optionally reuses successful outputs as future style references.

This skill is independent at runtime. The design compiler is built into this repository and does not require an external skill.

When to Use

Use this skill when the user wants one or more of the following:

  • Design-oriented poster generation
  • Product images or ad visuals
  • PPT cover visuals or chapter art
  • Infographic-like or teaching/demo visuals
  • Style reference reuse from prior generations
  • Interactive “show examples first, then generate” flow
  • Batch generation for multiple directions or aspect ratios
  • Series generation where multiple images should share one visual language

Do not use this skill for:

  • pixel-accurate UI recreation
  • editable charts
  • exact typography output inside the image
  • tasks that require vector, HTML, or PPT-native assets rather than raster images

Architecture

text
User Request
    ↓
[Prompt Engineer]
    ↓
[Style Scout]
    ↓
[Internal Design Compiler]
    ↓
[GPT-Image-2 Generation]
    ↓
[QA + Archive]
    ↓
[Case Library]

Optional layers on top of the main path:

  • Interactive reference selection
  • Batch generation
  • Series generation

Setup

1. Deploy the skill

The skill source lives in:

bash
~/.hermes/skills/multi-agent-image/

Install runtime files into the working directory:

bash
python3 ~/.hermes/skills/multi-agent-image/scripts/install.py

This prepares:

  • ~/.hermes/agents/multi-agent-image/output/
  • ~/.hermes/agents/multi-agent-image/case_library/
  • agent role folders and memory files
  • local runtime scripts copied from the skill
2. Install Python dependencies
bash
pip install openai requests
3. Set API key
bash
export OPENAI_API_KEY="sk-..."

This key is used with the apimart-compatible GPT-Image-2 endpoints in this skill.

Core Components

scripts/design_compiler.py

Internal prompt compiler.

Responsibilities:

  • detect task type
  • choose defaults for aspect and quality
  • build design_reasoning
  • compress it into compiled_brief
  • produce the final generation prompt

This is the core logic that makes the skill independent.

scripts/design_image.py

CLI entrypoint for the internal compiler.

Use it when you want:

  • prompt-only output
  • a local design compilation test
  • direct generation without the full multi-agent workflow

Example:

bash
cd ~/.hermes/agents/multi-agent-image
python3 design_image.py \
  --task poster \
  --brief "AI训练营招生海报,强调速度、增长、实战" \
  --direction balanced \
  --aspect 3:4 \
  --prompt-only

It prints:

  • design_reasoning
  • compiled_brief
  • prompt
  • settings
scripts/orchestrator_v2.py

Main workflow entrypoint.

Responsibilities:

  • run prompt analysis
  • choose task and generation parameters
  • optionally select a reference from the case library
  • call the internal compiler
  • call GPT-Image-2
  • archive outputs
  • auto-save successful results into the case library
scripts/gpt_image2_generator.py

Low-level GPT-Image-2 client.

Responsibilities:

  • submit async generation tasks
  • poll task status
  • download image results

Use this when you want direct API access without the full workflow.

scripts/case_library.py

Persistent library of past generations.

Responsibilities:

  • save outputs by task type
  • store metadata and rating
  • search by brief, prompt, or tags
  • return image paths for reuse as references
scripts/case_selector.py

Interactive helper for Hermes dialogue flows.

Responsibilities:

  • render user-facing selection text
  • parse replies like 1, n, case_001, or 搜索蓝色
scripts/interactive_run.py

Two-phase dialogue wrapper.

Use it when the workflow needs to ask the user before generating.

scripts/batch_generator_v2.py

Batch generation entrypoint.

Supports:

  • same brief, multiple directions
  • same brief, multiple aspect ratios
  • multiple briefs in one run
scripts/series_generator.py

Style-consistent series generator.

Workflow:

  1. generate a master image
  2. extract style signals from its compiled brief
  3. generate child images that follow the same visual system
templates/linear_batch.py

Editable template for resumable sequential runs.

Useful when you want:

  • explicit scene lists
  • filesystem-based progress monitoring
  • style propagation from the first generated image

Internal Design Compiler

The internal compiler produces three layers:

Show full SKILL.md (365 more words)Show less
1. design_reasoning

This captures design intent before generation.

Typical fields:

  • task
  • communication_goal
  • audience
  • channel
  • visual_system
  • hierarchy_strategy
  • safe_zone_strategy
  • lighting_strategy
  • palette_strategy
  • anti_filler_rules
  • anti_slop_rules
2. compiled_brief

This is a compressed design brief for generation.

It includes:

  • what the image is for
  • what should dominate visually
  • what space should remain available
  • what to avoid
3. prompt

Final model-facing prompt used for GPT-Image-2.

The prompt is generated from design logic, not just from a list of style keywords.

Supported Tasks

The built-in compiler understands these task classes:

  • poster
  • product
  • ppt
  • infographic
  • teaching
  • auto

Default aspect assumptions:

  • poster → 3:4
  • product → 1:1
  • ppt → 16:9
  • infographic → 4:3
  • teaching → 16:9

Direction modes:

  • conservative
  • balanced
  • bold

Quality modes:

  • draft
  • final
  • premium

Current generation channel:

  • gpt-image-2

Usage

Quick start
bash
cd ~/.hermes/agents/multi-agent-image
python3 quick_start.py "AI训练营招生海报,强调速度、增长、实战"
Prompt-only compilation
bash
cd ~/.hermes/agents/multi-agent-image
python3 design_image.py \
  --task product \
  --brief "高端陶瓷咖啡杯电商首图,温暖晨光,突出釉面质感" \
  --prompt-only
Full orchestrated generation
python
from orchestrator_v2 import run

run("AI训练营招生海报,强调速度增长实战")
Force task and visual settings
python
from orchestrator_v2 import run

run(
    "高端咖啡杯商品图",
    task="product",
    direction="balanced",
    aspect="1:1",
    quality="final",
    use_reference=False,
)

Interactive Workflow

Use the two-phase pattern when Hermes should ask before generating.

Phase 1: prepare text for the user
python
from interactive_run import prepare

text = prepare("帮我做张 AI 训练营海报", task="poster")
print(text)
Phase 2: execute after the user chooses
python
from interactive_run import execute

result = execute("帮我做张 AI 训练营海报", user_choice="1", task="poster")

Supported reply patterns:

  • 1, 2, 3
  • n
  • y
  • case_001
  • 搜索蓝色

Batch Generation

Same brief, multiple directions
python
from batch_generator_v2 import batch_styles

batch_styles("AI训练营海报", task="poster")
Same brief, multiple aspect ratios
python
from batch_generator_v2 import batch_aspects

batch_aspects("AI训练营海报", task="poster", aspects=["1:1", "16:9", "9:16"])
Multiple briefs
python
from batch_generator_v2 import batch_briefs

batch_briefs(["海报A", "海报B", "海报C"], task="poster")

Series Generation

Use this when several outputs should feel like the same campaign or product family.

python
from series_generator import SeriesGenerator

sg = SeriesGenerator()
sg.create_series(
    master_brief="AI训练营系列视觉,科技蓝,专业商务感",
    items=[
        {"name": "主海报", "brief": "AI训练营招生主海报", "aspect": "3:4"},
        {"name": "Banner", "brief": "官网 Banner", "aspect": "16:9"},
        {"name": "朋友圈", "brief": "朋友圈推广方形图", "aspect": "1:1"},
    ],
    task="poster",
    direction="balanced",
)

Case Library

Case library directory:

text
~/.hermes/agents/multi-agent-image/case_library/

Output directory:

text
~/.hermes/agents/multi-agent-image/output/

Typical case structure:

text
case_library/
├── poster/
│   └── case_001_example/
│       ├── image.png
│       └── metadata.json

Typical metadata fields:

  • case_id
  • task
  • brief
  • prompt
  • params
  • tags
  • rating

Validation Guidance

Before generating at scale, test prompt quality first:

bash
python3 design_image.py \
  --task poster \
  --brief "AI训练营招生海报,强调速度、增长、实战" \
  --direction balanced \
  --aspect 3:4 \
  --prompt-only

What to check:

  • Does design_reasoning state a clear communication goal?
  • Is there an explicit safe zone?
  • Is hierarchy obvious?
  • Do anti_slop_rules remove HUD overlays, fog, and generic clutter?
  • Does the prompt describe a single strong visual idea rather than a pile of elements?

Current Limits

  • Current image provider is centered on gpt-image-2
  • QA scoring is intentionally lightweight
  • Series generation is heavier than one-off generation
  • The skill is optimized for raster outputs, not editable assets
  • Some reference documents remain longer than necessary, but the main runtime path is consistent

Version History

  • v1.0.0 Initial multi-agent workflow for GPT-Image-2 generation
  • v2.0.0 Added case library, interactive reference selection, and image-to-image style reuse
  • v2.1.0 Added stronger download retry logic, batch workflows, and series generation
  • v2.2.0 Packaged as a reusable Hermes skill with install script and runtime layout
  • v3.0.0 Internalized the design compiler and removed external runtime dependency

© kangarooking, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 19 other files (scripts, references) in multi-agent-image of kangarooking/kangarooking-skills.

  • SKILL.md
  • README.md
  • references/agents/image_generator.md
  • references/agents/metadata_manager.md
  • references/agents/prompt_engineer.md
  • references/agents/qa_bot.md
  • references/agents/style_scout.md
  • scripts/batch_generator_v2.py
  • scripts/case_library.py
  • scripts/case_selector.py
  • scripts/design_compiler.py
  • scripts/design_image.py
  • scripts/gpt_image2_generator.py
  • scripts/install.py
  • scripts/interactive_run.py
  • scripts/orchestrator_v2.py
  • scripts/quick_start.py
  • scripts/series_generator.py
  • … and 2 more

Open the folder on GitHubat commit 08bbee0

Compare with similar skills

Multi Agent Image next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Multi Agent Image compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Multi Agent Image this skillkangarooking/kangarooking-skills656—~2.6kAutomated safety check: PassMIT
Gauntlet Loopduolahypercho/gauntlet-loop163—~689Automated safety check: PassMIT
Agent Fleetyan-labs/yan-skills208—~2.3kAutomated safety check: PassMIT
Design With Imagesdzhng/skills1k—~1.1kAutomated safety check: PassMIT
Design ReviewPrismer-AI/PrismerCloud1.6k—~1.3kAutomated safety check: NotesMIT
Cullglebis/claude-skills388—~1.9kAutomated safety check: PassMIT

Similar skills

  • Gauntlet Loop

    duolahypercho/gauntlet-loop

    GAME skill. An agent skill from duolahypercho/gauntlet-loop.

    163 GitHub stars~689 tokensUpdated 2 mo ago
    Media & CreativeAuto-check passed
  • Agent Fleet

    yan-labs/yan-skills

    使用本机 fleet 分派 Codex GPT-6、Gemini、Grok 或 JEV 任务,或调用 Kollab 图片/视频/音频/多模态能力时使用;包括用户点名 agent-fleet、Nano Banana、nanobanana、香蕉、便宜模型、多模型并行,用户说“让 Codex 或 GPT-6 做某事”的编码、调研与 review 派单,以及按全局 CLAUDE.md §2…

    208 GitHub stars~2.3k tokensUpdated 3 days ago
    Media & CreativeAuto-check passed
  • Design With Images

    dzhng/skills

    Explore UI and visual design with image generation, then iterate the real implementation against the user-selected concept.

    1k GitHub stars~1.1k tokensUpdated yesterday
    Media & CreativeAuto-check passed
  • Design Review

    Prismer-AI/PrismerCloud

    Five-dimension design audit (frontend UI/UX · server data-model & flow · endpoint spec · daemon-runtime/SDK & local-cache-first · built-in skill/agent-role), executed by a NON-implementing agent.

    1.6k GitHub stars~1.3k tokensUpdated 7 days ago
    Media & CreativeAuto-check: notes
  • Cull

    glebis/claude-skills

    This skill should be used when the user wants to view, review, rate, organize, search, or export images / AI-art generations with the Cull app.

    388 GitHub stars~1.9k tokensUpdated 11 days ago
    Media & CreativeAuto-check passed
  • Structured Image Generation

    bytedance/deer-flow

    Turns an image request into a structured JSON prompt and runs a bundled Python script to generate the picture, optionally guided by reference images.

    83k GitHub starsUsed in 5 repos~2.9k tokens
    Media & CreativeAuto-check passed

More from kangarooking/kangarooking-skills

All 19 skills in this repo
  • Scroll Promo Site Builder

    kangarooking/kangarooking-skills

    Create a scroll-controlled cinematic product website with rich motion (动效网站) from product materials, reference pages or videos, and brand assets.

    656 GitHub stars~2.6k tokensUpdated 1 mo ago
    Auto-check passed
  • Twitter Monitor

    kangarooking/kangarooking-skills

    Fetch recent posts from one or more X/Twitter accounts through twitterapi.io, output structured JSON/CSV records, optionally sync records to Feishu/Lark Bitable through feishu-cli, and optionally…

    656 GitHub stars~798 tokensUpdated 1 mo ago
    Auto-check passed
  • Apimart Image Gen

    kangarooking/kangarooking-skills

    Generate and download images through APIMart's asynchronous GPT-Image-2 API.

    656 GitHub stars~745 tokensUpdated 1 mo ago
    Auto-check passed
  • Bilibili Viral Topic

    kangarooking/kangarooking-skills

    Find low-follower viral Bilibili videos for account-growth topic selection.

    656 GitHub stars~756 tokensUpdated 1 mo ago
    Auto-check passed
  • Viral Title

    kangarooking/kangarooking-skills

    Generate high-potential viral title candidates for content across WeChat public account articles, X/Twitter posts, YouTube videos, Bilibili videos, and similar content platforms.

    656 GitHub stars~1.7k tokensUpdated 1 mo ago
    Auto-check passed
  • Wechat Viral Topic

    kangarooking/kangarooking-skills

    Find WeChat public-account articles that break out above the account's average reads for account-growth topic selection.

    656 GitHub stars~914 tokensUpdated 1 mo ago
    Auto-check passed

Questions about Multi Agent Image

What does Multi Agent Image do?

Standalone multi-agent image generation skill for Hermes. An agent skill from kangarooking/kangarooking-skills. Multi Agent Image is an agent skill from kangarooking/kangarooking-skills. Standalone multi-agent image generation skill for Hermes.

When should I use Multi Agent Image?

Multi Agent Image fits situations like: tasks that involve Image generation; tasks that involve Multi-agent orchestration.

How do I install Multi Agent Image in Claude Code?

Run `npx skills add kangarooking/kangarooking-skills --skill multi-agent-image -a claude-code`. Or copy the skill folder (multi-agent-image in kangarooking/kangarooking-skills) into .claude/skills/multi-agent-image in your project. Claude Code loads it when a task matches its description.

How do I install Multi Agent Image in Codex?

Run `npx skills add kangarooking/kangarooking-skills --skill multi-agent-image -a codex`. Or copy the skill folder (multi-agent-image in kangarooking/kangarooking-skills) into .agents/skills/multi-agent-image in your project. Codex loads it when a task matches its description.

Can I use Multi Agent Image in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add kangarooking/kangarooking-skills --skill multi-agent-image -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/multi-agent-image, .gemini/skills/multi-agent-image, .github/skills/multi-agent-image and .opencode/skills/multi-agent-image in your project.

What does Multi Agent Image need to run?

Going by SKILL.md and its folder, Multi Agent Image needs Python for the scripts in its folder, the command-line tools its instructions call (python3 and pip) and credentials named OPENAI_API_KEY. Our summary lists: Python 3; A credential in OPENAI_API_KEY.

Does Multi Agent Image access the network?

SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Multi Agent Image safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Multi Agent Image use?

Multi Agent Image is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Multi Agent Image use?

About 2.6k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.7k tokens, read only when the agent opens those files.

What are the alternatives to Multi Agent Image?

Skills that share tags, products or a category with Multi Agent Image: Gauntlet Loop (duolahypercho/gauntlet-loop, 163 stars), Agent Fleet (yan-labs/yan-skills, 208 stars), Design With Images (dzhng/skills, 1k stars) and Design Review (Prismer-AI/PrismerCloud, 1.6k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Multi Agent Image?

kangarooking (a GitHub user) maintains it in kangarooking/kangarooking-skills, which has 656 GitHub stars. The repository holds 19 skills in this directory. The repository was last updated on September 7, 2026.

Source: kangarooking/kangarooking-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.