Agent skill

Higgsfield Gpt Image 2

by OSideMedia in OSideMedia/higgsfield-ai-prompt-skill

A skill your agent uses when the user mentions GPT Image 2.0 or GPT Image 2.5, gpt-image-2, gpt-image-2.5, gptimage25, GPT-Image-2 prompts, the Flare / Sunburst variants, a transparent-background…

MITAuto-check passedMedia & Creative

Install Higgsfield Gpt Image 2

skills CLI
$ npx skills add OSideMedia/higgsfield-ai-prompt-skill --skill higgsfield-gpt-image-2 -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install OSideMedia/higgsfield-ai-prompt-skill higgsfield-gpt-image-2 --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/OSideMedia/higgsfield-ai-prompt-skill.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/higgsfield-gpt-image-2 .claude/skills/higgsfield-gpt-image-2 && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
higgsfield-gpt-image-2
GitHub stars
713
Token cost
~6.5k tokens
SKILL.md length
3,100 words
Files
3
Skills in repo
33
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when the user mentions GPT Image 2.0 or GPT Image 2.5, gpt-image-2, gpt-image-2.5, gptimage25, GPT-Image-2 prompts, the Flare / Sunburst variants, a transparent-background…

  • Works in 11 steps: What GPT Image 2.0 is → Three prompt formats → Format A — Structured JSON → …
  • The user mentions GPT Image 2.0
  • SKILL.md covers GPT Image 2.5, 1. What GPT Image 2.0 is, 2. Three prompt formats and 3. Format A — Structured JSON, plus 8 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Higgsfield Gpt Image 2 is an agent skill from OSideMedia/higgsfield-ai-prompt-skill. Use when the user mentions GPT Image 2.0 or GPT Image 2.5, gpt-image-2, gpt-image-2.5, gptimage25, GPT-Image-2 prompts, the Flare / Sunburst variants, a transparent-background image, or wants to generate an image with GPT Image 2.0 or 2.5. Covers when to prefer GPT Image 2.5 over 2.0 (transparent background, xhigh/max quality, extra aspect ratios — from platform and OpenAI facts), the CLI-vs-MCP disagreement on gptimage2, the three-format prompt taxonomy (Format A structured JSON for UI mockups and layout-dense…

Its SKILL.md is about 6.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files (for example `reference-sheet-workflow.md` and `static-ads-workflow.md`).

It sits in Media & Creative, covering Image generation. It works with Model Context Protocol and OpenAI. The repository describes itself as: Claude AI skill for cinematic Higgsfield AI prompts — 32 sub-skills covering Seedance 2.5 (omni-reference, video edit + extend) and 2.0, the Hell Grind feature-film pipeline, an… The licence is MIT.

When your agent uses it

  • The user mentions GPT Image 2.0
  • GPT-Image-2 prompts
  • The Flare / Sunburst variants
  • A transparent-background image

Example prompts

  • “/higgsfield-gpt-image-2”

Workflow steps

11 steps, taken from the step headings in SKILL.md.

  1. What GPT Image 2.0 is
  2. Three prompt formats
  3. Format A — Structured JSON
  4. Format B — Dense cinematic prose
  5. Format C — Auto-derive meta-prompt
  6. Routing decision
  7. Output format conventions
  8. Pre-delivery checklist
  9. Example routings
  10. Cross-surface workflow context
  11. Source acknowledgment

What it can do on your machine

Read from SKILL.md and the folder at commit 7075497. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are json).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • platform.openai.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Higgsfield Gpt Image 2 loads about 6.5k tokens when it runs. Until then it costs about 238 tokens; SKILL.md has 3,100 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~238
When it runs · the whole SKILL.md, loaded when a task matches
~6.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from OSideMedia/higgsfield-ai-prompt-skill at commit 7075497, republished under its MIT licence (© OSideMedia). 3,100 words, ~6,503 tokens.

Download SKILL.mdSave it as .claude/skills/higgsfield-gpt-image-2/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
higgsfield-gpt-image-2
description
Use when the user mentions GPT Image 2.0 or GPT Image 2.5, gpt-image-2, gpt-image-2.5, gpt_image_2_5, GPT-Image-2 prompts, the Flare / Sunburst variants, a transparent-background image, or wants to generate an image with GPT Image 2.0 or 2.5. Covers when to prefer GPT Image 2.5 over 2.0 (transparent background, xhigh/max quality, extra aspect ratios — from platform and OpenAI facts), the CLI-vs-MCP disagreement on gpt_image_2, the three-format prompt taxonomy (Format A structured JSON for UI mockups and layout-dense images; Format B dense cinematic prose for single-subject scenes; Format C auto-derive meta-prompt for theme-only concepts), per-format craft patterns, output conventions, the 6-item pre-delivery checklist, and cross-surface workflow context (companion static-ads-workflow.md for ad recreation; higgsfield-marketing-studio cross-surface-workflow.md §3 for ms_image / DTC Ads Higgsfield-native alternative).
user-invocable
true
metadata.tags
higgsfield, gpt-image-2, gpt-image-2.5, flare, sunburst, transparent-background, prompt-director, image, json, prose, meta-prompt, layout, mockup…
metadata.version
1.3.1
metadata.updated
2026-09-26
metadata.parent
higgsfield

Higgsfield GPT Image 2.0

A prompt director for GPT Image 2.0. Converts plain-text concepts into production-ready prompts that route by output type: structured JSON for layout-dense images (UI mockups, infographics, character sheets, multi-panel posters), dense cinematic prose for single-subject scenes (portraits, photographs, landscapes), or auto-derive meta-prompts for theme-only concepts where the model self-generates the composition.

Translated from Adil Aliyev's gpt-image-2-director source corpus per the v3.7.13 / v3.7.15 translation precedent. Two companion satellites extend this sub-skill: static-ads-workflow.md covers the ad-recreation workflow that uses GPT Image 2.0 as its generation engine, and reference-sheet-workflow.md covers the Automatic Product Reference Sheet + Automatic Prompt Creator workflow (one product image → a multi-view identity-locked reference sheet for high-consistency generation).

Sections 1–11 were written for GPT Image 2.0. GPT Image 2.5 — its surface, when to prefer it, and what is and is not documented about prompting it — is the next section.


GPT Image 2.5

New on Higgsfield in the 2026-09-26 catalog as gpt_image_2_5 — not yet field-rated. Facts below are [OFFICIAL — platform, snapshot 2026-09-26] (MCP and CLI agree on all of them) or [OFFICIAL — OpenAI docs] from the image generation guide and the Sunburst / Flare model pages (fetched 2026-09-26).

Surface
ParamValues (default first)
variantflare · sunburst
qualitylow · medium · high · xhigh · max
resolution1k · 2k · 4k
backgroundomit (model default) · auto · opaque · transparent
aspect ratioauto, 1:1, 3:2, 2:3, 4:3, 3:4, 16:9, 9:16, 21:9, 27:16, 16:27, 9:8, 8:9, 4:5, 5:4
mediaimage_references (generation and editing)

A bare call is flare / low / 1k — draft settings. OpenAI: "Use quality: "low" for quick drafts. For final assets, compare higher quality settings to find the right balance of detail, latency, and cost." State quality and resolution explicitly in every delivered settings header.

When to prefer 2.5 vs 2.0

Only platform and vendor facts — neither model has been field-rated against the other here, so no quality ranking is claimed.

Prefer GPT Image 2.5 when:

  • The asset needs a transparent background — background: transparent. It is the first transparent-background route in the MCP catalog, and the only one both Higgsfield surfaces agree on. OpenAI's API needs PNG or WebP output for transparency; Higgsfield exposes no output-format parameter, so confirm the downloaded file actually carries alpha before building on it.
  • A quality tier above high is wanted — xhigh / max. OpenAI: earlier GPT Image models "support quality settings up to high".
  • The frame needs 27:16, 16:27, 9:8, or 8:9 — 2.5 has them, 2.0 does not (DTC Ads / ms_image is the other image model that does).
  • The job is editing a reference precisely → variant: sunburst ("for workflows where editing precision matters most"). Fast everyday generation → variant: flare ("fast, high-quality everyday image generation").
  • It is a new integration — OpenAI: "For new integrations, use one of the GPT Image 2.5 models."

Keep GPT Image 2.0 when:

  • Continuing work already made on 2.0. The repo's field evidence — clothing and wardrobe on GPT Image 2 (../../image-models.md § Routing by Asset Class) and the Soul Cinema → GPT Image 2 refinement pipeline — was gathered on 2.0 and does not transfer to 2.5 automatically.
  • A Higgsfield workflow names gpt_image_2 — e.g. the MCP's bundled thumbnail-generation workflow routes its 3D-logo step to gpt_image_2.
  • You need CLI mask inpainting on 2.0 — mask / is_inpaint are CLI-listed only (next table).
CLI ↔ MCP disagreement on gpt_image_2

The two official Higgsfield surfaces describe gpt_image_2 differently on 2026-09-26. Recorded as a disagreement — neither surface is declared the winner:

FieldMCP models_exploreCLI higgsfield model get gpt_image_2
Aspect ratios1:1, 4:3, 3:4, 16:9, 21:9, 9:16, 3:2, 2:3same plus auto, 4:5, 5:4
resolution default1k2k
quality defaultlowhigh
backgroundnot listedauto / opaque / transparent
mask / is_inpaintnot listedlisted

Until one is verified on a real generation: set quality and resolution explicitly on 2.0, and do not promise background, masks, or auto / 4:5 / 5:4 on gpt_image_2 to an MCP user.

Prompting status on 2.5

The Format A / B / C taxonomy in §§ 2–9 was built and exercised on GPT Image 2.0. No GPT Image 2.5 prompting guide was found in OpenAI's image-generation guide or the Sunburst / Flare model pages (searched 2026-09-26), and the taxonomy has not been validated on 2.5 — the 2.5 prompting dialect is not yet documented here. If a user asks for a 2.5 prompt, the §§ 2–5 formats may be used only with an explicit note that they are 2.0 craft, unvalidated on 2.5. What OpenAI does state for its GPT Image models: complex prompts can take up to 2 minutes; text rendering "can still struggle with precise text placement and clarity"; and consistency for recurring characters or brand elements may drift across generations.


1. What GPT Image 2.0 is

GPT Image 2.0 is an image-generation model with a distinct capability profile that shapes how its prompts should be written. Four properties drive format choice across the three prompt taxonomies in §§ 2–5 below:

Granular layout precision. GPT Image 2.0 honors granular layout instructions — top-left panel shows X, mid-right shows Y, N icons in a row labeled A/B/C — in a way other models don't reliably match. This is testable: run the same multi-region brief against comparable image models and observe the difference. It's also why the Format A JSON taxonomy works as well as it does: the model reads JSON region keys as layout intent.

Text rendering. Multi-line paragraphs, mixed scripts (CJK + Latin), small UI labels, numeric data in tables — all sharp and legible. This is one of the model's distinctive strengths over comparable image generators. Same testability boundary: a user can verify by running prompts with mixed scripts and small UI labels against comparable models and observing the difference. The implication for prompts: embed real text in quotation marks exactly as it should render; do not paraphrase.

Design and UI as sweet spot. Website landing pages, social-feed mockups, magazine covers, infographics, exploded product diagrams, exam-paper layouts — anything with real information density. Lean prompts into the strengths.

Cinematic photorealism is the weakness. Human faces often go plasticky on realism-flagged prompts. Lean into stylized, illustrated, or editorial aesthetics rather than hyperreal skin. When realism is requested, frame it as film photography (grain, flash, 35mm) rather than as "photorealistic" — film-photography language tends to produce the look users want without triggering the plasticky-skin failure mode. Cross-reference: vocab.md § Visual Style Vocabulary → Film Stock Emulation for the broader film-photography language family.


2. Three prompt formats

Pick one based on the user's concept. If the concept fits multiple, pick the one best suited to the subject — don't hedge.

FormatUse whenOutput type
A — Structured JSONOutput has discrete regions, labeled parts, UI chrome, multi-panel grids, or information hierarchyUI mockups, landing pages, infographics, exploded diagrams, character reference sheets, social-media post mockups, magazine layouts, editorial document renders, multi-panel posters, comic / manga pages, brand-identity boards, design-system boards, card grids
B — Dense cinematic proseOutput is one scene, one frame, one subject with no chrome or layout regionsportraits, cinematic scenes, concept art, illustrations, landscapes, fashion shots, character moments
C — Auto-derive meta-promptUser gives a theme and wants the model to self-generate the whole compositionconcept posters from a single topic, character relationship diagrams, encyclopedia-style infographics

Each format has its own craft patterns in §§ 3–5 below. The routing decision is consolidated in § 6.

Tie-break

When in genuine doubt between A and B (e.g., "a character with some labels around them") — default to A. Layout precision is GPT Image 2.0's primary differentiator and prompts should reach for it.


3. Format A — Structured JSON

Write a single JSON object describing every visible region. GPT Image 2.0 reads this as a layout spec.

Core fields to reach for
  • type — one-line description of what this image is ("infographic poster", "landing page mockup", "exploded view diagram", "anime character reference sheet", "social media app interface mockup")
  • style — the visual style ("cute flat vector illustration, cozy, warm, soft shading", "clean high-tech 3D render, studio lighting, glowing accents", "GTA V cover art style, cel-shaded, thick black panel borders")
  • subject or character — the main entity, with specific visual attributes
  • layout — nested objects for regions: header, centerpiece, sections, footer, left_side, right_side, grid_panels, top_header, bottom_bar, etc. This is where precision matters most.
  • background — color, texture, or scene
  • Text content embedded in quoted strings. Keep real text if the user provided it — don't paraphrase. CJK and other non-Latin scripts stay in their original form.
Key patterns that make JSON prompts work

Count-and-label pattern. When there are multiple similar items (buttons, icons, chat messages, panels, callouts), give a count and a parallel labels array:

json
"messages": {
  "count": 7,
  "items": ["user1: hello", "user2: hi there", "..."]
}

Position-scoped regions. Explicitly name positions: top-left, top-center, mid-right, bottom-center-right. GPT Image 2.0 respects these.

Section objects with title, position, count, labels. For infographics with multiple zones:

json
{
  "title": "衣装・装備詳細",
  "position": "bottom-left",
  "count": 9,
  "labels": ["胸当て", "肩当て", "腕甲", "..."]
}

Templateable slots with {argument name="x" default="y"}. When a user explicitly wants a reusable template, slot notation like {argument name="city" default="Tokyo"} works. As a heuristic, don't add slots by default on one-off prompts — keep the prompt concrete. When slots are used, the default value should be a realistic value, not a placeholder string.

Inline typography callouts. When typography matters, include it inline: "title in large serif font", "11px Inter Regular", "Space Grotesk Bold Caps".

Worked example — minimal shape
json
{
  "type": "landing page mockup",
  "style": "clean modern e-commerce, soft pastel palette, generous whitespace",
  "layout": {
    "header": {
      "logo": "small black wordmark 'AURA'",
      "nav": ["Shop", "About", "Journal", "Contact"],
      "cart_icon": "top-right"
    },
    "hero": {
      "left_side": "large product photo of amber glass serum bottle on marble",
      "right_side": {
        "headline": "Skin, restored.",
        "subheadline": "A 7-day reset ritual. Clinically tested.",
        "cta_button": "black pill button 'Shop the ritual'"
      }
    },
    "below_hero": {
      "ingredient_grid": {"count": 4, "labels": ["Vitamin C", "Niacinamide", "Peptides", "Hyaluronic Acid"]}
    }
  }
}

For the ad-recreation variant of structured layout — fractional-coordinate zone notation, safe-zone discipline, brand-vs-structure separation, wireframe intermediation — see the companion static-ads-workflow.md § 3 (Layout zones + safe zones) and § 4 (Brand-vs-structure separation).


4. Format B — Dense cinematic prose

Write one continuous paragraph. Order the information roughly as: image type / medium → main subject with specific visual details → pose or action → background / setting → environmental details → lighting → color palette / film stock / texture → mood descriptor.

Roughly because GPT Image 2.0 handles slight ordering deviation gracefully — the order is a scaffolding, not a HARD RULE. The point is to put the most concrete visual anchors (medium, subject, props) early enough that the model commits to them before downstream details modify.

What makes prose prompts work

Specific over atmospheric. "White ribbed tank top and a loose beige knit cardigan slipping off one shoulder" beats "casual outfit". GPT Image 2.0 executes specificity better than it interprets mood.

Concrete props and objects. Reference exact things: "a white vintage Toyota Levin hatchback with red taillights", "an open notebook, a pen, and a pink flower on a desk". See vocab.md § Composition Vocabulary for the broader spatial-anchoring vocabulary that pairs with concrete-prop direction.

Camera and film language. "35mm film photograph", "direct camera flash", "low-angle dynamic perspective", "aerial drone shot", "shallow depth of field" — these actually steer the output toward the look named.

Embedded text in quotation marks. When text appears in the image, put it in quotes exactly as it should render: elegant vertical Japanese text that reads "都会の夜に溶けていく". This pattern leans into the text-rendering capability documented in § 1.

Avoid "photorealistic" when faces are in frame. Use "cinematic", "film photograph", "35mm", "editorial portrait" instead — these bias toward a look GPT Image 2.0 actually nails, rather than triggering its plasticky-skin failure mode. Cross-reference: § 1 capability framing + vocab.md § Visual Style Vocabulary → Film Stock Emulation.

Show full SKILL.md (1,310 more words)Show less
Worked example

A cinematic, moody photograph of an Asian woman looking back over her shoulder at the viewer on a rainy night in a bustling street. She has wet, stringy black hair plastered to her face and a melancholic expression, wearing a loose, oversized greyish-green jacket. The street is wet, reflecting the blurred, glowing neon signs and traffic lights of the city. Parked on the wet asphalt to her left is a white vintage Toyota Levin hatchback with its red taillights illuminated. On the top left side of the image, elegant vertical Japanese text reads "都会の夜に溶けていく" in a large serif font. The overall aesthetic is atmospheric and cinematic, 35mm film texture, muted warm palette, capturing a quiet introspective moment amidst urban chaos.

Adapted: this example is otherwise preserved verbatim from the source corpus (Adil's gpt-image-2-director source, not held in this repo); one age word before "Asian woman" is removed. The Seedance engine rule does not govern GPT Image 2 — this repo keeps age words out of every example it ships [HOUSE] (../higgsfield-seedance/ENGINE-RULES.md rule 1 states the reason).

What this example demonstrates: cinematic-medium opening ("A cinematic, moody photograph"), specific subject detail (wet stringy hair, oversized jacket), concrete prop (white vintage Toyota Levin hatchback with red taillights), embedded text in original CJK script with size + font direction ("elegant vertical Japanese text reads … in a large serif font"), film-stock language ("35mm film texture"), mood closing line. The CJK text rendering is the simplest demonstration of § 1's text-rendering capability — preserve or substitute embedded text as appropriate to the user's concept.


5. Format C — Auto-derive meta-prompt

The user gives a theme. You write instructions for the model to self-generate the full composition.

Structure
Please automatically generate a [output type] centered around [THEME].

Require the AI to automatically derive and uniformly design the entire following visual system based on this theme, without my extra specification:
- [list of derivations the model should make — core subject, supporting structure, hovering elements, color hierarchy, material contrast, lighting, typography, etc.]

[Overall Style]
[specific style direction — "cel-shaded illustration", "ultra-realistic 3D commercial CGI rendering", "watercolor and ink hand-drawn illustration", etc.]

[Composition Rules]
- [rules about premium quality, central order, negative space, hierarchy]

[Visual Quality]
- [rules about detail level, lighting, materials]

[Typography System]
- [ratio of visual to text, title/subtitle generation, font temperament]

[Signature]
Naturally add the signature "[NAME]" in the [position].
When to use

The user gives only a theme ("Chinese emperors", "Demon Slayer character map", "the psychology of procrastination") and wants a rich, self-derived output. If they give specific layout details, use Format A instead — the auto-derive meta-prompt cedes layout control to the model, which is the wrong choice when layout precision was specified.


6. Routing decision

Scan the user's concept and pick based on what they describe:

  • Mockup, UI, landing page, infographic, poster with panels, character sheet, magazine layout, grid, dashboard, diagram, social feed, exam paper, technical document → Format A (JSON).
  • One scene, one subject, one frame with no discrete regions — a portrait, a cinematic shot, a landscape, a character moment, an illustration, a photograph → Format B (prose).
  • A theme only with the user asking you to design the whole thing — "make a poster about X", "relationship diagram of X", "encyclopedia page for X" without specifying the layout → Format C (meta-prompt).

When in genuine doubt between A and B (e.g., "a character with some labels around them"), default to A per § 2 tie-break. The cost of routing to A when B was acceptable is small (extra layout precision); the cost of routing to B when A was needed is larger (lost layout anchoring).


7. Output format conventions

Return only the finished prompt in a code block. No preamble, no explanation, no "here's your prompt:", no format-choice justification. The user pastes it into GPT Image 2.0 directly.

This matches root SKILL.md HARD RULES item 7 (output the prompt; no preamble) — same production discipline applied at this surface.

Code-block conventions:

  • JSON prompts: wrap in a ```json code block.
  • Prose prompts: wrap in a plain ``` code block.
  • Meta-prompts: wrap in a plain ``` code block.

Multiple variations: when the user asks for multiple variations, return them as separate code blocks with a one-line label before each (e.g., **Variant A — magazine layout:**).


8. Pre-delivery checklist

Before outputting the prompt, scan it against this 6-item check. The pass takes 30–60 seconds; the savings compound across iteration loops.

  1. Region coverage — does every distinct visible region have a named location or field?
  2. Counts and labels — are counts and labels explicit where there are multiple similar items?
  3. Real text preserved — is real text kept in its original language and in quotes, ready to render?
  4. Realism framing — for face-heavy prompts, have you avoided "photorealistic" in favor of film / cinematic language (per § 1, § 4)?
  5. Style specificity — is the style line specific enough to produce a recognizable aesthetic, or is it generic adjective stacking?
  6. JSON validity (Format A only) — is the JSON valid (check braces, commas, escaping of quotes inside strings)?

This is the surface-specific operationalization of the cross-cutting pre-delivery discipline. See DISCIPLINE.md § Pre-Delivery Discipline for the underlying pattern (the v3.7.5 + v3.7.7 audit-corpus discipline applied across multiple sub-skills).


9. Example routings

Six paired examples showing the routing decision in practice.

User conceptFormatReasoning
"make me a landing page for a matcha tea startup called Kori, emphasis on clean Japanese minimalism"A (JSON)Landing page has discrete regions (header, hero, product grid, footer) — Format A's primary use case.
"a woman sitting alone at a bus stop at dusk, 90s vibe"B (prose)Single scene, no layout.
"make a poster about the history of the samurai"C (meta-prompt)Theme only; no specifics on layout.
"a character reference sheet for a cyberpunk bounty hunter named Iris, show front/side/back views and 4 expressions"A (JSON)Explicit layout regions ("front/side/back views and 4 expressions"). The 'character sheet' label is a strong Format A signal.
"photo of a grey-bearded man fixing a vintage arcade machine, lit by the machine's screen"B (prose)One framed photograph.
"infographic about the types of clouds, make it look like a vintage encyclopedia page"A or CIf the user lists the cloud types and what to show for each → A. If they just say "types of clouds" and expect you to fill it in → C.

Two user concepts (the bus stop, the arcade machine) are adapted from this table's earlier wording to keep age words out of examples [HOUSE].


10. Cross-surface workflow context

GPT Image 2.0 is one image-side surface in the Higgsfield platform ecosystem. Two cross-surface paths are worth naming:

Companion: ad-recreation workflow. When the user uploads a winning ad format and wants to recreate it with their brand's products and copy, route to the companion static-ads-workflow.md. That doc covers the 10-step workflow that uses GPT Image 2.0 as its generation engine plus the production discipline (fractional-coordinate zone notation, safe-zone top/bottom-10% rule, brand-vs-structure separation, wireframe intermediation) that Adil's source corpus documented.

Higgsfield-native alternative: ms_image / "DTC Ads". Higgsfield's platform exposes a native image-generation surface for ad work — ms_image (display name "DTC Ads") — that is brand-kit-aware (accepts brand_kit_id), ad-format curated (accepts required style_id), supports batch generation up to 20 images per call, and accepts up to 14 reference media. For brand-kit-consistency-across-many-images use cases or batch generation of 5–20 ads per call, ms_image is the more integrated option. Adil's source corpus doesn't document ms_image (source-corpus reconciliation #12) — treat GPT Image 2.0 as the default for image-side work, and reach for ms_image when its specific differentiators apply. See the Marketing Studio cross-surface-workflow.md § 3 for the full ms_image coverage.


11. Source acknowledgment

GPT Image 2.0 coverage in this sub-skill is translated from source material by Adil Aliyev — Higgsfield-team-adjacent author of gpt-image-2-director.md (sibling to marketing-studio-director.md + higgsfield-content-factory.md translated in v3.7.13, and cinematic-motion-language.md translated in v3.7.15). Source corpus: gpt-image-2-director source skill (SKILL.md, ~14 KB, 206 lines).

Per-claim translation calibration

Per the v3.7.13 author-signature calibration carried through v3.7.15: format-routing taxonomy (Format A / B / C) is Adil's craft synthesis (translated freely to our voice); per-format vocabulary and craft patterns are standard prompt-engineering and cinematography (adopted close to verbatim); capability claims about GPT Image 2.0's prompt-following and text-rendering strengths (translated under the v3.7.15 Spatial Zoning testability exception — testable directional claims about prompt-side effects pass the testability boundary; universalizing metaphysical claims about model cognition would not).

Verification trail

Per-source disposition: the v3.7.16 Phase 0 verification notes (internal build notes) § VERIFY 0.1 (gpt-image-2-director source read + disposition class TRANSLATE-WITH-VERIFICATION LOWER-FRICTION SUB-CLASS).

Per-element translation rules: the v3.7.16 Phase 1 inventory notes (internal build notes) § 1B-A (per-element ADOPT / DOWNGRADE table).

Companion satellite: static-ads-workflow.md covers the static-ads.md source corpus translation per § 1B-B.

© OSideMedia, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files in skills/higgsfield-gpt-image-2 of OSideMedia/higgsfield-ai-prompt-skill.

  • SKILL.md
  • reference-sheet-workflow.md
  • static-ads-workflow.md

Open the folder on GitHubat commit 7075497

Compare with similar skills

Higgsfield Gpt Image 2 next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Higgsfield Gpt Image 2 compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Higgsfield Gpt Image 2 this skillOSideMedia/higgsfield-ai-prompt-skill713—~6.5kAutomated safety check: PassMIT
Openai Image Genswarmclawai/swarmclaw689—~705Automated safety check: PassMIT
Scenario Gpt Imagescenario-labs/skills946—~2.7kAutomated safety check: PassMIT
Ag2 Use Builtin Toolsag2ai/build-with-ag2252—~1.3kAutomated safety check: PassApache-2.0
Nano Banana Pro Prompts Recommend SkillYouMind-OpenLab/nano-banana-pro-prompts-recommend-skill1.9k1 repos~4.1kAutomated safety check: PassNone
Md2wechatgeekjourneyx/md2wechat-skill3.7k—~3.8kAutomated safety check: PassCustom licence

Similar skills

  • Openai Image Gen

    swarmclawai/swarmclaw

    Generate images via OpenAI Images API (GPT Image, DALL-E 3, DALL-E 2).

    689 GitHub stars~705 tokensUpdated 3 mo ago
    Media & CreativeAuto-check passed
  • Scenario Gpt Image

    scenario-labs/skills

    A skill your agent uses when generating or editing images with OpenAI's GPT Image models on Scenario via MCP: text-to-image, edits from reference images, inpainting with an alpha mask, in-image text…

    946 GitHub stars~2.7k tokensUpdated yesterday
    Media & CreativeAuto-check passed
  • Ag2 Use Builtin Tools

    ag2ai/build-with-ag2

    Wire AG2 beta's shipped tools into an Agent — both provider-native server-side tools (web search, web fetch, code execution, MCP, image generation, memory) and locally-executed common toolkits…

    252 GitHub stars~1.3k tokensUpdated 1 mo ago
    Productivity & AutomationAuto-check passed
  • Nano Banana Pro Prompts Recommend Skill

    YouMind-OpenLab/nano-banana-pro-prompts-recommend-skill

    Recommend suitable prompts from 10,000+ Nano Banana Pro image generation prompts based on user needs.

    1.9k GitHub starsUsed in 1 repo~4.1k tokens
    Media & CreativeAuto-check passed
  • Md2wechat

    geekjourneyx/md2wechat-skill

    Convert Markdown to WeChat Official Account HTML. An agent skill from geekjourneyx/md2wechat-skill.

    3.7k GitHub stars~3.8k tokensUpdated 17 days ago
    Media & CreativeAuto-check passed
  • Yingzao

    op7418/guizang-yingzao-skill

    Transform real Chinese architecture and place-based cultural photos into art-directed editorial posters, integrated multi-photo scenes, and optional source comparisons.

    496 GitHub stars~1.1k tokensUpdated 1 mo ago
    Media & CreativeAuto-check passed

More from OSideMedia/higgsfield-ai-prompt-skill

All 33 skills in this repo
  • Higgsfield

    OSideMedia/higgsfield-ai-prompt-skill

    A skill your agent uses whenever the user asks anything about Higgsfield AI — writing or refining video/image prompts, choosing a model (Kling, Veo, Wan, Seedance, Minimax Hailuo, DoP, Soul, Nano…

    713 GitHub stars~9.1k tokensUpdated 14 days ago
    Auto-check passed
  • Higgsfield Assist

    OSideMedia/higgsfield-ai-prompt-skill

    A skill your agent uses when the user asks about Higgsfield Assist (the built-in GPT-5 copilot), how to use the platform's native AI assistant, credit optimization strategies, plan selection, how to…

    713 GitHub stars~2.9k tokensUpdated 14 days ago
    Auto-check passed
  • Higgsfield Image Shots

    OSideMedia/higgsfield-ai-prompt-skill

    A skill your agent uses when the user wants to generate a cinematic still image on Higgsfield, asks about shot framing, camera angle, or composition for image prompts, needs a specific shot type…

    713 GitHub stars~5.1k tokensUpdated 14 days ago
    Auto-check passed
  • Higgsfield Mixed Media

    OSideMedia/higgsfield-ai-prompt-skill

    A skill your agent uses when the user asks about Mixed Media, wants to apply artistic preset styles to an image (Noir, Sketch, Paper, Canvas, Particles, Neon, etc.), combine multiple artistic…

    713 GitHub stars~2.3k tokensUpdated 14 days ago
    Auto-check passed
  • Higgsfield Moodboard

    OSideMedia/higgsfield-ai-prompt-skill

    A skill your agent uses when the user asks about Moodboard, building a moodboard from reference images, curated moodboard presets, Soul Hex color transfer, applying a visual style direction to…

    713 GitHub stars~2.3k tokensUpdated 14 days ago
    Auto-check passed
  • Higgsfield Motion

    OSideMedia/higgsfield-ai-prompt-skill

    A skill your agent uses when the user wants to apply a named Higgsfield motion preset, asks about VFX presets, transformation effects, elemental effects, or transition presets.

    713 GitHub stars~5k tokensUpdated 14 days ago
    Auto-check passed

Questions about Higgsfield Gpt Image 2

What does Higgsfield Gpt Image 2 do?

A skill your agent uses when the user mentions GPT Image 2.0 or GPT Image 2.5, gpt-image-2, gpt-image-2.5, gptimage25, GPT-Image-2 prompts, the Flare / Sunburst variants, a transparent-background…. Higgsfield Gpt Image 2 is an agent skill from OSideMedia/higgsfield-ai-prompt-skill.5.

When should I use Higgsfield Gpt Image 2?

Higgsfield Gpt Image 2 fits situations like: the user mentions GPT Image 2.0; GPT-Image-2 prompts; the Flare / Sunburst variants; A transparent-background image.

How do I install Higgsfield Gpt Image 2 in Claude Code?

Run `npx skills add OSideMedia/higgsfield-ai-prompt-skill --skill higgsfield-gpt-image-2 -a claude-code`. Or copy the skill folder (skills/higgsfield-gpt-image-2 in OSideMedia/higgsfield-ai-prompt-skill) into .claude/skills/higgsfield-gpt-image-2 in your project. Claude Code loads it when a task matches its description.

How do I install Higgsfield Gpt Image 2 in Codex?

Run `npx skills add OSideMedia/higgsfield-ai-prompt-skill --skill higgsfield-gpt-image-2 -a codex`. Or copy the skill folder (skills/higgsfield-gpt-image-2 in OSideMedia/higgsfield-ai-prompt-skill) into .agents/skills/higgsfield-gpt-image-2 in your project. Codex loads it when a task matches its description.

Can I use Higgsfield Gpt Image 2 in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add OSideMedia/higgsfield-ai-prompt-skill --skill higgsfield-gpt-image-2 -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/higgsfield-gpt-image-2, .gemini/skills/higgsfield-gpt-image-2, .github/skills/higgsfield-gpt-image-2 and .opencode/skills/higgsfield-gpt-image-2 in your project.

What does Higgsfield Gpt Image 2 need to run?

SKILL.md names no scripts, command-line tools or credentials: Higgsfield Gpt Image 2 is instructions for the agent only.

Does Higgsfield Gpt Image 2 access the network?

SKILL.md names 1 domain. As links in the text: platform.openai.com. This is read from the text; nothing was executed.

Is Higgsfield Gpt Image 2 safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Higgsfield Gpt Image 2 use?

Higgsfield Gpt Image 2 is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Higgsfield Gpt Image 2 use?

About 6.5k tokens (SKILL.md is roughly 26k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Higgsfield Gpt Image 2?

Skills that share tags, products or a category with Higgsfield Gpt Image 2: Openai Image Gen (swarmclawai/swarmclaw, 689 stars), Scenario Gpt Image (scenario-labs/skills, 946 stars), Ag2 Use Builtin Tools (ag2ai/build-with-ag2, 252 stars) and Nano Banana Pro Prompts Recommend Skill (YouMind-OpenLab/nano-banana-pro-prompts-recommend-skill, 1.9k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Higgsfield Gpt Image 2?

OSideMedia (a GitHub user) maintains it in OSideMedia/higgsfield-ai-prompt-skill, which has 713 GitHub stars. The repository holds 33 skills in this directory. The repository was last updated on September 27, 2026.

Source: OSideMedia/higgsfield-ai-prompt-skill on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.