Openai Image Gen
swarmclawai/swarmclaw
Generate images via OpenAI Images API (GPT Image, DALL-E 3, DALL-E 2).
A skill your agent uses when the user mentions GPT Image 2.0 or GPT Image 2.5, gpt-image-2, gpt-image-2.5, gptimage25, GPT-Image-2 prompts, the Flare / Sunburst variants, a transparent-background…
$ npx skills add OSideMedia/higgsfield-ai-prompt-skill --skill higgsfield-gpt-image-2 -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install OSideMedia/higgsfield-ai-prompt-skill higgsfield-gpt-image-2 --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/OSideMedia/higgsfield-ai-prompt-skill.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/higgsfield-gpt-image-2 .claude/skills/higgsfield-gpt-image-2 && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "higgsfield-gpt-image-2" agent skill from https://github.com/OSideMedia/higgsfield-ai-prompt-skill/tree/main/skills/higgsfield-gpt-image-2 into .claude/skills/higgsfield-gpt-image-2/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "higgsfield-gpt-image-2", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/OSideMedia/higgsfield-ai-prompt-skill/tree/main/skills/higgsfield-gpt-image-2Type this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add OSideMedia/higgsfield-ai-prompt-skill --skill higgsfield-gpt-image-2 -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install OSideMedia/higgsfield-ai-prompt-skill higgsfield-gpt-image-2 --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/OSideMedia/higgsfield-ai-prompt-skill.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/higgsfield-gpt-image-2 .agents/skills/higgsfield-gpt-image-2 && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "higgsfield-gpt-image-2" agent skill from https://github.com/OSideMedia/higgsfield-ai-prompt-skill/tree/main/skills/higgsfield-gpt-image-2 into .agents/skills/higgsfield-gpt-image-2/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "higgsfield-gpt-image-2", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add OSideMedia/higgsfield-ai-prompt-skill --skill higgsfield-gpt-image-2 -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install OSideMedia/higgsfield-ai-prompt-skill higgsfield-gpt-image-2 --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/OSideMedia/higgsfield-ai-prompt-skill.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/higgsfield-gpt-image-2 .cursor/skills/higgsfield-gpt-image-2 && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "higgsfield-gpt-image-2" agent skill from https://github.com/OSideMedia/higgsfield-ai-prompt-skill/tree/main/skills/higgsfield-gpt-image-2 into .cursor/skills/higgsfield-gpt-image-2/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "higgsfield-gpt-image-2", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/OSideMedia/higgsfield-ai-prompt-skill.git --path skills/higgsfield-gpt-image-2--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add OSideMedia/higgsfield-ai-prompt-skill --skill higgsfield-gpt-image-2 -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install OSideMedia/higgsfield-ai-prompt-skill higgsfield-gpt-image-2 --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/OSideMedia/higgsfield-ai-prompt-skill.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/higgsfield-gpt-image-2 .gemini/skills/higgsfield-gpt-image-2 && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "higgsfield-gpt-image-2" agent skill from https://github.com/OSideMedia/higgsfield-ai-prompt-skill/tree/main/skills/higgsfield-gpt-image-2 into .gemini/skills/higgsfield-gpt-image-2/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "higgsfield-gpt-image-2", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install OSideMedia/higgsfield-ai-prompt-skill higgsfield-gpt-image-2Installs for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add OSideMedia/higgsfield-ai-prompt-skill --skill higgsfield-gpt-image-2 -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/OSideMedia/higgsfield-ai-prompt-skill.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/higgsfield-gpt-image-2 .github/skills/higgsfield-gpt-image-2 && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "higgsfield-gpt-image-2" agent skill from https://github.com/OSideMedia/higgsfield-ai-prompt-skill/tree/main/skills/higgsfield-gpt-image-2 into .github/skills/higgsfield-gpt-image-2/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "higgsfield-gpt-image-2", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add OSideMedia/higgsfield-ai-prompt-skill --skill higgsfield-gpt-image-2 -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install OSideMedia/higgsfield-ai-prompt-skill higgsfield-gpt-image-2 --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/OSideMedia/higgsfield-ai-prompt-skill.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/higgsfield-gpt-image-2 .opencode/skills/higgsfield-gpt-image-2 && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "higgsfield-gpt-image-2" agent skill from https://github.com/OSideMedia/higgsfield-ai-prompt-skill/tree/main/skills/higgsfield-gpt-image-2 into .opencode/skills/higgsfield-gpt-image-2/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "higgsfield-gpt-image-2", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
higgsfield-gpt-image-2A skill your agent uses when the user mentions GPT Image 2.0 or GPT Image 2.5, gpt-image-2, gpt-image-2.5, gptimage25, GPT-Image-2 prompts, the Flare / Sunburst variants, a transparent-background…
Higgsfield Gpt Image 2 is an agent skill from OSideMedia/higgsfield-ai-prompt-skill. Use when the user mentions GPT Image 2.0 or GPT Image 2.5, gpt-image-2, gpt-image-2.5, gptimage25, GPT-Image-2 prompts, the Flare / Sunburst variants, a transparent-background image, or wants to generate an image with GPT Image 2.0 or 2.5. Covers when to prefer GPT Image 2.5 over 2.0 (transparent background, xhigh/max quality, extra aspect ratios — from platform and OpenAI facts), the CLI-vs-MCP disagreement on gptimage2, the three-format prompt taxonomy (Format A structured JSON for UI mockups and layout-dense…
Its SKILL.md is about 6.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files (for example `reference-sheet-workflow.md` and `static-ads-workflow.md`).
It sits in Media & Creative, covering Image generation. It works with Model Context Protocol and OpenAI. The repository describes itself as: Claude AI skill for cinematic Higgsfield AI prompts — 32 sub-skills covering Seedance 2.5 (omni-reference, video edit + extend) and 2.0, the Hell Grind feature-film pipeline, an… The licence is MIT.
11 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 7075497. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are json).
From the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
platform.openai.comFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Higgsfield Gpt Image 2 loads about 6.5k tokens when it runs. Until then it costs about 238 tokens; SKILL.md has 3,100 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from OSideMedia/higgsfield-ai-prompt-skill at commit 7075497, republished under its MIT licence (© OSideMedia). 3,100 words, ~6,503 tokens.
.claude/skills/higgsfield-gpt-image-2/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.A prompt director for GPT Image 2.0. Converts plain-text concepts into production-ready prompts that route by output type: structured JSON for layout-dense images (UI mockups, infographics, character sheets, multi-panel posters), dense cinematic prose for single-subject scenes (portraits, photographs, landscapes), or auto-derive meta-prompts for theme-only concepts where the model self-generates the composition.
Translated from Adil Aliyev's gpt-image-2-director source corpus per the v3.7.13 / v3.7.15 translation precedent. Two companion satellites extend this sub-skill: static-ads-workflow.md covers the ad-recreation workflow that uses GPT Image 2.0 as its generation engine, and reference-sheet-workflow.md covers the Automatic Product Reference Sheet + Automatic Prompt Creator workflow (one product image → a multi-view identity-locked reference sheet for high-consistency generation).
Sections 1–11 were written for GPT Image 2.0. GPT Image 2.5 — its surface, when to prefer it, and what is and is not documented about prompting it — is the next section.
New on Higgsfield in the 2026-09-26 catalog as gpt_image_2_5 — not yet field-rated. Facts
below are [OFFICIAL — platform, snapshot 2026-09-26] (MCP and CLI agree on all of them) or
[OFFICIAL — OpenAI docs] from the image generation guide
and the Sunburst /
Flare model pages (fetched 2026-09-26).
| Param | Values (default first) |
|---|---|
variant | flare · sunburst |
quality | low · medium · high · xhigh · max |
resolution | 1k · 2k · 4k |
background | omit (model default) · auto · opaque · transparent |
| aspect ratio | auto, 1:1, 3:2, 2:3, 4:3, 3:4, 16:9, 9:16, 21:9, 27:16, 16:27, 9:8, 8:9, 4:5, 5:4 |
| media | image_references (generation and editing) |
A bare call is flare / low / 1k — draft settings. OpenAI: "Use quality: "low" for quick
drafts. For final assets, compare higher quality settings to find the right balance of detail,
latency, and cost." State quality and resolution explicitly in every delivered settings header.
Only platform and vendor facts — neither model has been field-rated against the other here, so no quality ranking is claimed.
Prefer GPT Image 2.5 when:
background: transparent. It is the first
transparent-background route in the MCP catalog, and the only one both Higgsfield surfaces
agree on. OpenAI's API needs PNG or WebP output for transparency; Higgsfield exposes no
output-format parameter, so confirm the downloaded file actually carries alpha before
building on it.high is wanted — xhigh / max. OpenAI: earlier GPT Image models
"support quality settings up to high".ms_image is the other image model that does).variant: sunburst ("for workflows where editing
precision matters most"). Fast everyday generation → variant: flare ("fast, high-quality
everyday image generation").Keep GPT Image 2.0 when:
../../image-models.md § Routing by Asset Class) and the Soul Cinema → GPT Image 2
refinement pipeline — was gathered on 2.0 and does not transfer to 2.5 automatically.gpt_image_2 — e.g. the MCP's bundled thumbnail-generation
workflow routes its 3D-logo step to gpt_image_2.mask / is_inpaint are CLI-listed only (next table).gpt_image_2The two official Higgsfield surfaces describe gpt_image_2 differently on 2026-09-26. Recorded as
a disagreement — neither surface is declared the winner:
| Field | MCP models_explore | CLI higgsfield model get gpt_image_2 |
|---|---|---|
| Aspect ratios | 1:1, 4:3, 3:4, 16:9, 21:9, 9:16, 3:2, 2:3 | same plus auto, 4:5, 5:4 |
resolution default | 1k | 2k |
quality default | low | high |
background | not listed | auto / opaque / transparent |
mask / is_inpaint | not listed | listed |
Until one is verified on a real generation: set quality and resolution explicitly on 2.0, and
do not promise background, masks, or auto / 4:5 / 5:4 on gpt_image_2 to an MCP user.
The Format A / B / C taxonomy in §§ 2–9 was built and exercised on GPT Image 2.0. No GPT Image 2.5 prompting guide was found in OpenAI's image-generation guide or the Sunburst / Flare model pages (searched 2026-09-26), and the taxonomy has not been validated on 2.5 — the 2.5 prompting dialect is not yet documented here. If a user asks for a 2.5 prompt, the §§ 2–5 formats may be used only with an explicit note that they are 2.0 craft, unvalidated on 2.5. What OpenAI does state for its GPT Image models: complex prompts can take up to 2 minutes; text rendering "can still struggle with precise text placement and clarity"; and consistency for recurring characters or brand elements may drift across generations.
GPT Image 2.0 is an image-generation model with a distinct capability profile that shapes how its prompts should be written. Four properties drive format choice across the three prompt taxonomies in §§ 2–5 below:
Granular layout precision. GPT Image 2.0 honors granular layout instructions — top-left panel shows X, mid-right shows Y, N icons in a row labeled A/B/C — in a way other models don't reliably match. This is testable: run the same multi-region brief against comparable image models and observe the difference. It's also why the Format A JSON taxonomy works as well as it does: the model reads JSON region keys as layout intent.
Text rendering. Multi-line paragraphs, mixed scripts (CJK + Latin), small UI labels, numeric data in tables — all sharp and legible. This is one of the model's distinctive strengths over comparable image generators. Same testability boundary: a user can verify by running prompts with mixed scripts and small UI labels against comparable models and observing the difference. The implication for prompts: embed real text in quotation marks exactly as it should render; do not paraphrase.
Design and UI as sweet spot. Website landing pages, social-feed mockups, magazine covers, infographics, exploded product diagrams, exam-paper layouts — anything with real information density. Lean prompts into the strengths.
Cinematic photorealism is the weakness. Human faces often go plasticky on realism-flagged prompts. Lean into stylized, illustrated, or editorial aesthetics rather than hyperreal skin. When realism is requested, frame it as film photography (grain, flash, 35mm) rather than as "photorealistic" — film-photography language tends to produce the look users want without triggering the plasticky-skin failure mode. Cross-reference: vocab.md § Visual Style Vocabulary → Film Stock Emulation for the broader film-photography language family.
Pick one based on the user's concept. If the concept fits multiple, pick the one best suited to the subject — don't hedge.
| Format | Use when | Output type |
|---|---|---|
| A — Structured JSON | Output has discrete regions, labeled parts, UI chrome, multi-panel grids, or information hierarchy | UI mockups, landing pages, infographics, exploded diagrams, character reference sheets, social-media post mockups, magazine layouts, editorial document renders, multi-panel posters, comic / manga pages, brand-identity boards, design-system boards, card grids |
| B — Dense cinematic prose | Output is one scene, one frame, one subject with no chrome or layout regions | portraits, cinematic scenes, concept art, illustrations, landscapes, fashion shots, character moments |
| C — Auto-derive meta-prompt | User gives a theme and wants the model to self-generate the whole composition | concept posters from a single topic, character relationship diagrams, encyclopedia-style infographics |
Each format has its own craft patterns in §§ 3–5 below. The routing decision is consolidated in § 6.
When in genuine doubt between A and B (e.g., "a character with some labels around them") — default to A. Layout precision is GPT Image 2.0's primary differentiator and prompts should reach for it.
Write a single JSON object describing every visible region. GPT Image 2.0 reads this as a layout spec.
type — one-line description of what this image is ("infographic poster", "landing page mockup", "exploded view diagram", "anime character reference sheet", "social media app interface mockup")style — the visual style ("cute flat vector illustration, cozy, warm, soft shading", "clean high-tech 3D render, studio lighting, glowing accents", "GTA V cover art style, cel-shaded, thick black panel borders")subject or character — the main entity, with specific visual attributeslayout — nested objects for regions: header, centerpiece, sections, footer, left_side, right_side, grid_panels, top_header, bottom_bar, etc. This is where precision matters most.background — color, texture, or sceneCount-and-label pattern. When there are multiple similar items (buttons, icons, chat messages, panels, callouts), give a count and a parallel labels array:
"messages": {
"count": 7,
"items": ["user1: hello", "user2: hi there", "..."]
}Position-scoped regions. Explicitly name positions: top-left, top-center, mid-right, bottom-center-right. GPT Image 2.0 respects these.
Section objects with title, position, count, labels. For infographics with multiple zones:
{
"title": "衣装・装備詳細",
"position": "bottom-left",
"count": 9,
"labels": ["胸当て", "肩当て", "腕甲", "..."]
}Templateable slots with {argument name="x" default="y"}. When a user explicitly wants a reusable template, slot notation like {argument name="city" default="Tokyo"} works. As a heuristic, don't add slots by default on one-off prompts — keep the prompt concrete. When slots are used, the default value should be a realistic value, not a placeholder string.
Inline typography callouts. When typography matters, include it inline: "title in large serif font", "11px Inter Regular", "Space Grotesk Bold Caps".
{
"type": "landing page mockup",
"style": "clean modern e-commerce, soft pastel palette, generous whitespace",
"layout": {
"header": {
"logo": "small black wordmark 'AURA'",
"nav": ["Shop", "About", "Journal", "Contact"],
"cart_icon": "top-right"
},
"hero": {
"left_side": "large product photo of amber glass serum bottle on marble",
"right_side": {
"headline": "Skin, restored.",
"subheadline": "A 7-day reset ritual. Clinically tested.",
"cta_button": "black pill button 'Shop the ritual'"
}
},
"below_hero": {
"ingredient_grid": {"count": 4, "labels": ["Vitamin C", "Niacinamide", "Peptides", "Hyaluronic Acid"]}
}
}
}For the ad-recreation variant of structured layout — fractional-coordinate zone notation, safe-zone discipline, brand-vs-structure separation, wireframe intermediation — see the companion static-ads-workflow.md § 3 (Layout zones + safe zones) and § 4 (Brand-vs-structure separation).
Write one continuous paragraph. Order the information roughly as: image type / medium → main subject with specific visual details → pose or action → background / setting → environmental details → lighting → color palette / film stock / texture → mood descriptor.
Roughly because GPT Image 2.0 handles slight ordering deviation gracefully — the order is a scaffolding, not a HARD RULE. The point is to put the most concrete visual anchors (medium, subject, props) early enough that the model commits to them before downstream details modify.
Specific over atmospheric. "White ribbed tank top and a loose beige knit cardigan slipping off one shoulder" beats "casual outfit". GPT Image 2.0 executes specificity better than it interprets mood.
Concrete props and objects. Reference exact things: "a white vintage Toyota Levin hatchback with red taillights", "an open notebook, a pen, and a pink flower on a desk". See vocab.md § Composition Vocabulary for the broader spatial-anchoring vocabulary that pairs with concrete-prop direction.
Camera and film language. "35mm film photograph", "direct camera flash", "low-angle dynamic perspective", "aerial drone shot", "shallow depth of field" — these actually steer the output toward the look named.
Embedded text in quotation marks. When text appears in the image, put it in quotes exactly as it should render: elegant vertical Japanese text that reads "都会の夜に溶けていく". This pattern leans into the text-rendering capability documented in § 1.
Avoid "photorealistic" when faces are in frame. Use "cinematic", "film photograph", "35mm", "editorial portrait" instead — these bias toward a look GPT Image 2.0 actually nails, rather than triggering its plasticky-skin failure mode. Cross-reference: § 1 capability framing + vocab.md § Visual Style Vocabulary → Film Stock Emulation.
A cinematic, moody photograph of an Asian woman looking back over her shoulder at the viewer on a rainy night in a bustling street. She has wet, stringy black hair plastered to her face and a melancholic expression, wearing a loose, oversized greyish-green jacket. The street is wet, reflecting the blurred, glowing neon signs and traffic lights of the city. Parked on the wet asphalt to her left is a white vintage Toyota Levin hatchback with its red taillights illuminated. On the top left side of the image, elegant vertical Japanese text reads "都会の夜に溶けていく" in a large serif font. The overall aesthetic is atmospheric and cinematic, 35mm film texture, muted warm palette, capturing a quiet introspective moment amidst urban chaos.
Adapted: this example is otherwise preserved verbatim from the source corpus (Adil's
gpt-image-2-director source, not held in this repo); one age word before "Asian woman" is removed. The Seedance
engine rule does not govern GPT Image 2 — this repo keeps age words out of every example
it ships [HOUSE] (../higgsfield-seedance/ENGINE-RULES.md rule 1 states the reason).
What this example demonstrates: cinematic-medium opening ("A cinematic, moody photograph"), specific subject detail (wet stringy hair, oversized jacket), concrete prop (white vintage Toyota Levin hatchback with red taillights), embedded text in original CJK script with size + font direction ("elegant vertical Japanese text reads … in a large serif font"), film-stock language ("35mm film texture"), mood closing line. The CJK text rendering is the simplest demonstration of § 1's text-rendering capability — preserve or substitute embedded text as appropriate to the user's concept.
The user gives a theme. You write instructions for the model to self-generate the full composition.
Please automatically generate a [output type] centered around [THEME].
Require the AI to automatically derive and uniformly design the entire following visual system based on this theme, without my extra specification:
- [list of derivations the model should make — core subject, supporting structure, hovering elements, color hierarchy, material contrast, lighting, typography, etc.]
[Overall Style]
[specific style direction — "cel-shaded illustration", "ultra-realistic 3D commercial CGI rendering", "watercolor and ink hand-drawn illustration", etc.]
[Composition Rules]
- [rules about premium quality, central order, negative space, hierarchy]
[Visual Quality]
- [rules about detail level, lighting, materials]
[Typography System]
- [ratio of visual to text, title/subtitle generation, font temperament]
[Signature]
Naturally add the signature "[NAME]" in the [position].The user gives only a theme ("Chinese emperors", "Demon Slayer character map", "the psychology of procrastination") and wants a rich, self-derived output. If they give specific layout details, use Format A instead — the auto-derive meta-prompt cedes layout control to the model, which is the wrong choice when layout precision was specified.
Scan the user's concept and pick based on what they describe:
When in genuine doubt between A and B (e.g., "a character with some labels around them"), default to A per § 2 tie-break. The cost of routing to A when B was acceptable is small (extra layout precision); the cost of routing to B when A was needed is larger (lost layout anchoring).
Return only the finished prompt in a code block. No preamble, no explanation, no "here's your prompt:", no format-choice justification. The user pastes it into GPT Image 2.0 directly.
This matches root SKILL.md HARD RULES item 7 (output the prompt; no preamble) — same production discipline applied at this surface.
Code-block conventions:
```json code block.``` code block.``` code block.Multiple variations: when the user asks for multiple variations, return them as separate code blocks with a one-line label before each (e.g., **Variant A — magazine layout:**).
Before outputting the prompt, scan it against this 6-item check. The pass takes 30–60 seconds; the savings compound across iteration loops.
This is the surface-specific operationalization of the cross-cutting pre-delivery discipline. See DISCIPLINE.md § Pre-Delivery Discipline for the underlying pattern (the v3.7.5 + v3.7.7 audit-corpus discipline applied across multiple sub-skills).
Six paired examples showing the routing decision in practice.
| User concept | Format | Reasoning |
|---|---|---|
| "make me a landing page for a matcha tea startup called Kori, emphasis on clean Japanese minimalism" | A (JSON) | Landing page has discrete regions (header, hero, product grid, footer) — Format A's primary use case. |
| "a woman sitting alone at a bus stop at dusk, 90s vibe" | B (prose) | Single scene, no layout. |
| "make a poster about the history of the samurai" | C (meta-prompt) | Theme only; no specifics on layout. |
| "a character reference sheet for a cyberpunk bounty hunter named Iris, show front/side/back views and 4 expressions" | A (JSON) | Explicit layout regions ("front/side/back views and 4 expressions"). The 'character sheet' label is a strong Format A signal. |
| "photo of a grey-bearded man fixing a vintage arcade machine, lit by the machine's screen" | B (prose) | One framed photograph. |
| "infographic about the types of clouds, make it look like a vintage encyclopedia page" | A or C | If the user lists the cloud types and what to show for each → A. If they just say "types of clouds" and expect you to fill it in → C. |
Two user concepts (the bus stop, the arcade machine) are adapted from this table's earlier
wording to keep age words out of examples [HOUSE].
GPT Image 2.0 is one image-side surface in the Higgsfield platform ecosystem. Two cross-surface paths are worth naming:
Companion: ad-recreation workflow. When the user uploads a winning ad format and wants to recreate it with their brand's products and copy, route to the companion static-ads-workflow.md. That doc covers the 10-step workflow that uses GPT Image 2.0 as its generation engine plus the production discipline (fractional-coordinate zone notation, safe-zone top/bottom-10% rule, brand-vs-structure separation, wireframe intermediation) that Adil's source corpus documented.
Higgsfield-native alternative: ms_image / "DTC Ads". Higgsfield's platform exposes a native image-generation surface for ad work — ms_image (display name "DTC Ads") — that is brand-kit-aware (accepts brand_kit_id), ad-format curated (accepts required style_id), supports batch generation up to 20 images per call, and accepts up to 14 reference media. For brand-kit-consistency-across-many-images use cases or batch generation of 5–20 ads per call, ms_image is the more integrated option. Adil's source corpus doesn't document ms_image (source-corpus reconciliation #12) — treat GPT Image 2.0 as the default for image-side work, and reach for ms_image when its specific differentiators apply. See the Marketing Studio cross-surface-workflow.md § 3 for the full ms_image coverage.
GPT Image 2.0 coverage in this sub-skill is translated from source material by Adil Aliyev — Higgsfield-team-adjacent author of gpt-image-2-director.md (sibling to marketing-studio-director.md + higgsfield-content-factory.md translated in v3.7.13, and cinematic-motion-language.md translated in v3.7.15). Source corpus: gpt-image-2-director source skill (SKILL.md, ~14 KB, 206 lines).
Per the v3.7.13 author-signature calibration carried through v3.7.15: format-routing taxonomy (Format A / B / C) is Adil's craft synthesis (translated freely to our voice); per-format vocabulary and craft patterns are standard prompt-engineering and cinematography (adopted close to verbatim); capability claims about GPT Image 2.0's prompt-following and text-rendering strengths (translated under the v3.7.15 Spatial Zoning testability exception — testable directional claims about prompt-side effects pass the testability boundary; universalizing metaphysical claims about model cognition would not).
Per-source disposition: the v3.7.16 Phase 0 verification notes (internal build notes) § VERIFY 0.1 (gpt-image-2-director source read + disposition class TRANSLATE-WITH-VERIFICATION LOWER-FRICTION SUB-CLASS).
Per-element translation rules: the v3.7.16 Phase 1 inventory notes (internal build notes) § 1B-A (per-element ADOPT / DOWNGRADE table).
Companion satellite: static-ads-workflow.md covers the static-ads.md source corpus translation per § 1B-B.
© OSideMedia, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 2 other files in skills/higgsfield-gpt-image-2 of OSideMedia/higgsfield-ai-prompt-skill.
Open the folder on GitHubat commit 7075497
Higgsfield Gpt Image 2 next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Higgsfield Gpt Image 2 this skillOSideMedia/higgsfield-ai-prompt-skill | 713 | — | ~6.5k | Automated safety check: Pass | MIT | |
| Openai Image Genswarmclawai/swarmclaw | 689 | — | ~705 | Automated safety check: Pass | MIT | |
| Scenario Gpt Imagescenario-labs/skills | 946 | — | ~2.7k | Automated safety check: Pass | MIT | |
| Ag2 Use Builtin Toolsag2ai/build-with-ag2 | 252 | — | ~1.3k | Automated safety check: Pass | Apache-2.0 | |
| Nano Banana Pro Prompts Recommend SkillYouMind-OpenLab/nano-banana-pro-prompts-recommend-skill | 1.9k | 1 repos | ~4.1k | Automated safety check: Pass | None | |
| Md2wechatgeekjourneyx/md2wechat-skill | 3.7k | — | ~3.8k | Automated safety check: Pass | Custom licence |
swarmclawai/swarmclaw
Generate images via OpenAI Images API (GPT Image, DALL-E 3, DALL-E 2).
scenario-labs/skills
A skill your agent uses when generating or editing images with OpenAI's GPT Image models on Scenario via MCP: text-to-image, edits from reference images, inpainting with an alpha mask, in-image text…
ag2ai/build-with-ag2
Wire AG2 beta's shipped tools into an Agent — both provider-native server-side tools (web search, web fetch, code execution, MCP, image generation, memory) and locally-executed common toolkits…
YouMind-OpenLab/nano-banana-pro-prompts-recommend-skill
Recommend suitable prompts from 10,000+ Nano Banana Pro image generation prompts based on user needs.
geekjourneyx/md2wechat-skill
Convert Markdown to WeChat Official Account HTML. An agent skill from geekjourneyx/md2wechat-skill.
op7418/guizang-yingzao-skill
Transform real Chinese architecture and place-based cultural photos into art-directed editorial posters, integrated multi-photo scenes, and optional source comparisons.
OSideMedia/higgsfield-ai-prompt-skill
A skill your agent uses whenever the user asks anything about Higgsfield AI — writing or refining video/image prompts, choosing a model (Kling, Veo, Wan, Seedance, Minimax Hailuo, DoP, Soul, Nano…
OSideMedia/higgsfield-ai-prompt-skill
A skill your agent uses when the user asks about Higgsfield Assist (the built-in GPT-5 copilot), how to use the platform's native AI assistant, credit optimization strategies, plan selection, how to…
OSideMedia/higgsfield-ai-prompt-skill
A skill your agent uses when the user wants to generate a cinematic still image on Higgsfield, asks about shot framing, camera angle, or composition for image prompts, needs a specific shot type…
OSideMedia/higgsfield-ai-prompt-skill
A skill your agent uses when the user asks about Mixed Media, wants to apply artistic preset styles to an image (Noir, Sketch, Paper, Canvas, Particles, Neon, etc.), combine multiple artistic…
OSideMedia/higgsfield-ai-prompt-skill
A skill your agent uses when the user asks about Moodboard, building a moodboard from reference images, curated moodboard presets, Soul Hex color transfer, applying a visual style direction to…
OSideMedia/higgsfield-ai-prompt-skill
A skill your agent uses when the user wants to apply a named Higgsfield motion preset, asks about VFX presets, transformation effects, elemental effects, or transition presets.
Works with
Categories
A skill your agent uses when the user mentions GPT Image 2.0 or GPT Image 2.5, gpt-image-2, gpt-image-2.5, gptimage25, GPT-Image-2 prompts, the Flare / Sunburst variants, a transparent-background…. Higgsfield Gpt Image 2 is an agent skill from OSideMedia/higgsfield-ai-prompt-skill.5.
Higgsfield Gpt Image 2 fits situations like: the user mentions GPT Image 2.0; GPT-Image-2 prompts; the Flare / Sunburst variants; A transparent-background image.
Run `npx skills add OSideMedia/higgsfield-ai-prompt-skill --skill higgsfield-gpt-image-2 -a claude-code`. Or copy the skill folder (skills/higgsfield-gpt-image-2 in OSideMedia/higgsfield-ai-prompt-skill) into .claude/skills/higgsfield-gpt-image-2 in your project. Claude Code loads it when a task matches its description.
Run `npx skills add OSideMedia/higgsfield-ai-prompt-skill --skill higgsfield-gpt-image-2 -a codex`. Or copy the skill folder (skills/higgsfield-gpt-image-2 in OSideMedia/higgsfield-ai-prompt-skill) into .agents/skills/higgsfield-gpt-image-2 in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add OSideMedia/higgsfield-ai-prompt-skill --skill higgsfield-gpt-image-2 -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/higgsfield-gpt-image-2, .gemini/skills/higgsfield-gpt-image-2, .github/skills/higgsfield-gpt-image-2 and .opencode/skills/higgsfield-gpt-image-2 in your project.
SKILL.md names no scripts, command-line tools or credentials: Higgsfield Gpt Image 2 is instructions for the agent only.
SKILL.md names 1 domain. As links in the text: platform.openai.com. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Higgsfield Gpt Image 2 is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 6.5k tokens (SKILL.md is roughly 26k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Higgsfield Gpt Image 2: Openai Image Gen (swarmclawai/swarmclaw, 689 stars), Scenario Gpt Image (scenario-labs/skills, 946 stars), Ag2 Use Builtin Tools (ag2ai/build-with-ag2, 252 stars) and Nano Banana Pro Prompts Recommend Skill (YouMind-OpenLab/nano-banana-pro-prompts-recommend-skill, 1.9k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
OSideMedia (a GitHub user) maintains it in OSideMedia/higgsfield-ai-prompt-skill, which has 713 GitHub stars. The repository holds 33 skills in this directory. The repository was last updated on September 27, 2026.
Source: OSideMedia/higgsfield-ai-prompt-skill on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.