Agent skill

Ideogram 4 Prompt Builder

by digitalsamba in digitalsamba/claude-code-video-toolkit

Turns a casual image request into the structured JSON caption Ideogram 4 needs for legible on-image text, exact brand colors and controlled layout.

MITAuto-check: notesMedia & Creative

Install Ideogram 4 Prompt Builder

skills CLI
$ npx skills add digitalsamba/claude-code-video-toolkit --skill ideogram4 -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install digitalsamba/claude-code-video-toolkit ideogram4 --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/digitalsamba/claude-code-video-toolkit.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/ideogram4 .claude/skills/ideogram4 && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
ideogram4
GitHub stars
2.2k
Used in
1 other repo
Token cost
~1.3k tokens
SKILL.md length
491 words
Files
3
Skills in repo
11
Repo updated
First seen
Licence
MIT

At a glance

Turns a casual image request into the structured JSON caption Ideogram 4 needs for legible on-image text, exact brand colors and controlled layout.

  • Generating a title card or thumbnail with legible on-image text
  • SKILL.md covers When to Use This Skill, The One Thing to Get Right, Quick Reference —… and Key Files, plus 1 more section
  • Calls uv; needs IDEOGRAM_API_KEY
  • Producing an image that must match exact brand hex colors

What it does

Ideogram 4's advantage over larger models is rendering legible signage, logos, captions and multi-line text, plus exact color-palette and bounding-box control, but that advantage is locked behind a structured JSON caption format; a plain-text prompt only gets lesser results and misses the point of using this model. The skill has the agent act as the prompt expander, turning a casual request into that JSON caption, since the model is trained exclusively on captions that name every element explicitly, and the skill argues this expansion beats Ideogram's own free hosted magic-prompt, whose documentation says the shipped version differs from the one used in production.

The toolkit calls Ideogram's hosted v4 API rather than self-hosted weights, because the self-hostable weights are non-commercial while the paid API plans include a commercial license. The skill reaches for Ideogram 4 specifically for title cards, thumbnails, signage, logos, quote cards and calls to action with a baked-in headline, exact brand hex colors, bounding-box layout, or multilingual on-image text, and reaches for a different model instead when there is no critical text or the need is just a fast atmospheric background. Full schema details, key-ordering rules and the bounding-box coordinate system live in a prompting reference file.

When your agent uses it

  • Generating a title card or thumbnail with legible on-image text
  • Producing an image that must match exact brand hex colors
  • Placing text or objects in specific regions with bounding-box layout
  • Choosing between Ideogram 4 and another model for an image request

Example prompts

  • “Generate a YouTube thumbnail with the headline 'Episode 12' in our brand colors.”
  • “Make a quote card with this exact text placed in the lower third.”
  • “I need a logo mockup with the company name rendered legibly. Use Ideogram.”

Requirements

  • Access to Ideogram's hosted v4 API

What it can do on your machine

Read from SKILL.md and the folder at commit 2c99460. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • uv

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use uv, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • IDEOGRAM_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Ideogram 4 Prompt Builder loads about 1.3k tokens when it runs. Until then it costs about 118 tokens; SKILL.md has 491 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~118
When it runs · the whole SKILL.md, loaded when a task matches
~1.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteMentions a .env fileSKILL.md:55
    ted v4 API. Needs `IDEOGRAM_API_KEY` in `.env`

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from digitalsamba/claude-code-video-toolkit at commit 2c99460, republished under its MIT licence (© digitalsamba). 491 words, ~1,339 tokens.

Download SKILL.mdSave it as .claude/skills/ideogram4/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
ideogram4
description
Prompting patterns for Ideogram 4 text-to-image — best-in-class in-image text rendering and exact color/layout control via structured JSON captions. Use when generating images that need legible on-image text (title cards, thumbnails, logos, signage, CTAs), precise brand colors, or controlled spatial layout. Triggers include title slide image, thumbnail with text, on-image text, legible text in image, brand color palette image, bounding-box layout, Ideogram.

Ideogram 4 Skill

Text-to-image generation with Ideogram 4 (9.3B, open-weight, released June 2026). Its superpower is best-in-class in-image text rendering — it beats much larger models (FLUX.2 dev 32B, Qwen-Image 20B, Hunyuan 80B) at rendering legible signage, logos, captions, and multi-line text — plus exact color-palette and bounding-box control.

That advantage is locked behind a structured JSON caption format. A plain-text prompt gets you FLUX-level results and misses the entire point of using this model. This skill teaches Claude to act as the "magic prompt" expander — turning a user's casual request into the JSON caption Ideogram 4 was trained on.

Backend: The toolkit uses Ideogram's hosted v4 API (not self-hosted weights). The API accepts a structured json_prompt, so everything this skill teaches applies directly — Claude builds the caption, the tool posts it as json_prompt. Paid API plans include a commercial license, which the self-hostable weights (non-commercial) do not — that's why we use the API. Cost is ~$0.03/image (turbo) to ~$0.09/image (quality).

When to Use This Skill

Reach for Ideogram 4 (over FLUX.2) when the image needs:

  • Legible on-image text — title cards, thumbnails, lower-thirds backgrounds, signage, logos, quote cards, CTAs with a headline baked in
  • Exact brand colors — hex color-palette conditioning, per-element
  • Controlled layout — bounding boxes place text/objects in specific regions
  • Multilingual text in the image

Use FLUX.2 instead when: the image has no critical text, you need commercial-licensed output, or you just want a fast atmospheric background. FLUX takes plain natural-language prompts; Ideogram wants JSON. See tools/flux2.py.

The One Thing to Get Right

Always emit a structured JSON caption, not a plain sentence. The model is trained exclusively on JSON captions that name every element explicitly. Claude is a better expander than Ideogram's free hosted magic-prompt (their own docs note the shipped one "is not the same used in production"), so build the caption yourself using this skill rather than passing raw text.

Minimal valid caption:

json
{"high_level_description":"A sailboat at sunset on calm water.","style_description":{"aesthetics":"serene, warm, golden hour","lighting":"golden hour backlighting","photo":"wide angle, f/8","medium":"photograph","color_palette":["#FF6B35","#F7C59F","#004E89"]},"compositional_deconstruction":{"background":"Calm ocean at low horizon with orange-pink sky.","elements":[{"type":"obj","desc":"White triangular sail silhouetted against the setting sun."}]}}

Full schema, strict key-ordering rules, and the bbox coordinate system are in prompting.md. Worked title-card / thumbnail / quote-card examples are in examples.md.

Show full SKILL.md (156 more words)Show less

Quick Reference — tools/ideogram4.py

Thin wrapper over Ideogram's hosted v4 API. Needs IDEOGRAM_API_KEY in .env (key from developer.ideogram.ai). --json posts the caption as the API's json_prompt field (no server-side magic prompt — Claude is the expander); --prompt posts text_prompt.

bash
# Hand-authored JSON caption (the recommended path for text/layout) — Claude writes caption.json
uv run tools/ideogram4.py --json caption.json --output title.png

# Caption from stdin (Claude can pipe it directly)
cat caption.json | uv run tools/ideogram4.py --json - --output title.png

# Plain prompt — Ideogram's server-side magic prompt expands it (weaker; prefer --json)
uv run tools/ideogram4.py --prompt "Title card: 'AI ENGINEERING REVIEW' bold white on dark" --output title.png

# Inject brand hex colors into the caption's palette (JSON mode)
uv run tools/ideogram4.py --json caption.json --brand digital-samba --output cta.png

# Quality tier + resolution
uv run tools/ideogram4.py --json caption.json --speed QUALITY --resolution 2048x2048 --output slide.png

Key Files

  • prompting.md — full JSON schema, strict key ordering, bbox coordinate system, palette rules
  • examples.md — worked captions for title cards, thumbnails, quote cards, brand CTAs

Video Production Fit

Ideogram 4's niche in the toolkit is slides and thumbnails with baked-in text, where FLUX and LTX-2 fail (both render garbled text). Natural pairings:

Use caseWhy Ideogram 4
Title-card / CTA background with headline textLegible text + exact brand hex colors in one pass
YouTube/social thumbnail with a punchy phraseBig readable text is its strongest suit
Quote card / stat cardMulti-line text + layout control via bboxes
Signage/logos inside a product-demo sceneIn-image text other models can't render

Then feed the still into Remotion (<OffthreadVideo>/Img) or animate it with tools/ltx2.py --input.

© digitalsamba, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files in .claude/skills/ideogram4 of digitalsamba/claude-code-video-toolkit.

  • SKILL.md
  • examples.md
  • prompting.md

Open the folder on GitHubat commit 2c99460

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in digitalsamba/claude-code-video-toolkit, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Ideogram 4 Prompt Builder next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Ideogram 4 Prompt Builder compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Ideogram 4 Prompt Builder this skilldigitalsamba/claude-code-video-toolkit2.2k1 repos~1.3kAutomated safety check: NotesMIT
Logo Generatorop7418/logo-generator-skill2.2k—~1.8kAutomated safety check: NotesNone
Web Asset Generatoralonw0/web-asset-generator5121 repos~6.6kAutomated safety check: PassMIT
Image Generationonyx-dot-app/onyx32k1 repos~1.7kAutomated safety check: PassCustom licence
Gemini Image Generatordair-ai/dair-academy-plugins6142 repos~3.5kAutomated safety check: NotesMIT
Image Prompt ReverseLunarXuan/image-prompt-reverse464—~678Automated safety check: PassGPL-3.0

Similar skills

  • Logo Generator

    op7418/logo-generator-skill

    Generate professional SVG logos and high-end showcase images.

    2.2k GitHub stars~1.8k tokensUpdated 5 mo ago
    Media & CreativeAuto-check: notes
  • Web Asset Generator

    alonw0/web-asset-generator

    Generate web assets including favicons, app icons (PWA), and social media meta images (Open Graph) for Facebook, Twitter, WhatsApp, and LinkedIn.

    512 GitHub starsUsed in 1 repo~6.6k tokens
    Media & CreativeAuto-check passed
  • Image Generation

    onyx-dot-app/onyx

    Generate or edit raster images (photos, illustrations, textures, sprites, mockups, logos, infographics) using the workspace's configured image-generation provider via onyx-cli image.

    32k GitHub starsUsed in 1 repo~1.7k tokens
    Media & CreativeAuto-check passed
  • Gemini Image Generator

    dair-ai/dair-academy-plugins

    Generates and edits images with Google's Gemini Nano Banana Pro model through the Gemini API, including photo edits and multi-image composition.

    614 GitHub starsUsed in 2 repos~3.5k tokens
    Media & CreativeAuto-check: notes
  • Image Prompt Reverse

    LunarXuan/image-prompt-reverse

    Analyze user-provided reference images and reverse-engineer high-fidelity AI image-generation prompts.

    464 GitHub stars~678 tokensUpdated 1 mo ago
    Media & CreativeAuto-check passed
  • Generate Youtube Thumbnail

    krusemediallc/arcads-claude-code

    Generate high-CTR YouTube thumbnails using Nano Banana 2 via the Arcads external API.

    1.6k GitHub stars~2.4k tokensUpdated 14 days ago
    Media & CreativeAuto-check: notes

More from digitalsamba/claude-code-video-toolkit

All 11 skills in this repo
  • FFmpeg for Video Production

    digitalsamba/claude-code-video-toolkit

    Command recipes for converting, resizing, compressing, trimming and extracting audio from video with FFmpeg, including settings for Remotion projects.

    2.2k GitHub starsUsed in 3 repos~3.3k tokens
    Auto-check passed
  • LTX-2.3 Video Generation

    digitalsamba/claude-code-video-toolkit

    Generates roughly five-second video clips from a text prompt or a still image with the LTX-2.3 22B model, run through a Modal endpoint by `tools/ltx2.py`.

    2.2k GitHub starsUsed in 2 repos~2.4k tokens
    Auto-check: notes
  • ElevenLabs Voiceover Generator

    digitalsamba/claude-code-video-toolkit

    Generates narration, sound effects and cloned voices through the ElevenLabs API, with model and setting choices tuned to the content's style.

    2.2k GitHub starsUsed in 1 repo~2.7k tokens
    Auto-check: notes
  • Playwright Browser Demo Recording

    digitalsamba/claude-code-video-toolkit

    Records browser interactions as video with Playwright, covering viewport sizing, cursor highlighting, and converting output for Remotion.

    2.2k GitHub starsUsed in 1 repo~3.2k tokens
    Auto-check passed
  • Qwen Edit

    digitalsamba/claude-code-video-toolkit

    AI image editing prompting patterns for Qwen-Image-Edit. An agent skill from digitalsamba/claude-code-video-toolkit.

    2.2k GitHub starsUsed in 1 repo~711 tokens
    Auto-check passed
  • ACE-Step Music Generation

    digitalsamba/claude-code-video-toolkit

    Generates background music, vocal tracks, covers and stems with ACE-Step 1.5 through a bundled music_gen.py tool, using cloud or self-hosted providers.

    2.2k GitHub stars~3.3k tokensUpdated yesterday
    Auto-check: notes

Questions about Ideogram 4 Prompt Builder

What does Ideogram 4 Prompt Builder do?

Turns a casual image request into the structured JSON caption Ideogram 4 needs for legible on-image text, exact brand colors and controlled layout. Ideogram 4's advantage over larger models is rendering legible signage, logos, captions and multi-line text, plus exact color-palette and bounding-box control, but that advantage is locked behind a structured JSON caption format; a plain-text prompt only gets lesser results and misses the point of using this model. The skill has the agent act as the prompt expander, turning a casual request into that JSON caption, since the model is trained exclusively on captions that name every element explicitly, and the skill argues this expansion beats Ideogram's own free hosted magic-prompt, whose documentation says the shipped version differs from the one used in production.

When should I use Ideogram 4 Prompt Builder?

Ideogram 4 Prompt Builder fits situations like: generating a title card or thumbnail with legible on-image text; producing an image that must match exact brand hex colors; placing text or objects in specific regions with bounding-box layout; choosing between Ideogram 4 and another model for an image request.

How do I install Ideogram 4 Prompt Builder in Claude Code?

Run `npx skills add digitalsamba/claude-code-video-toolkit --skill ideogram4 -a claude-code`. Or copy the skill folder (.claude/skills/ideogram4 in digitalsamba/claude-code-video-toolkit) into .claude/skills/ideogram4 in your project. Claude Code loads it when a task matches its description.

How do I install Ideogram 4 Prompt Builder in Codex?

Run `npx skills add digitalsamba/claude-code-video-toolkit --skill ideogram4 -a codex`. Or copy the skill folder (.claude/skills/ideogram4 in digitalsamba/claude-code-video-toolkit) into .agents/skills/ideogram4 in your project. Codex loads it when a task matches its description.

Can I use Ideogram 4 Prompt Builder in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add digitalsamba/claude-code-video-toolkit --skill ideogram4 -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ideogram4, .gemini/skills/ideogram4, .github/skills/ideogram4 and .opencode/skills/ideogram4 in your project.

What does Ideogram 4 Prompt Builder need to run?

Going by SKILL.md and its folder, Ideogram 4 Prompt Builder needs the command-line tools its instructions call (uv) and credentials named IDEOGRAM_API_KEY. Our summary lists: Access to Ideogram's hosted v4 API.

Does Ideogram 4 Prompt Builder access the network?

SKILL.md contains no URLs. Its commands use uv, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Ideogram 4 Prompt Builder safe to install?

Our automated static check of SKILL.md found notes only (mentions a .env file), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Ideogram 4 Prompt Builder use?

Ideogram 4 Prompt Builder is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Ideogram 4 Prompt Builder use?

About 1.3k tokens (SKILL.md is roughly 5.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Ideogram 4 Prompt Builder?

Skills that share tags, products or a category with Ideogram 4 Prompt Builder: Logo Generator (op7418/logo-generator-skill, 2.2k stars), Web Asset Generator (alonw0/web-asset-generator, 512 stars), Image Generation (onyx-dot-app/onyx, 32k stars) and Gemini Image Generator (dair-ai/dair-academy-plugins, 614 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Ideogram 4 Prompt Builder?

digitalsamba (a GitHub organization) maintains it in digitalsamba/claude-code-video-toolkit, which has 2,172 GitHub stars. The repository holds 11 skills in this directory. The repository was last updated on October 5, 2026.

Source: digitalsamba/claude-code-video-toolkit on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.