Agent skill

Visual Prompt Builder

by UfukNode in UfukNode/Noustiny

Rewrite a narrative beat into a rich photoreal visual description suitable for text-to-image models.

MITAuto-check passedAI & LLM Engineering

Install Visual Prompt Builder

skills CLI
$ npx skills add UfukNode/Noustiny --skill visual-prompt-builder -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install UfukNode/Noustiny visual-prompt-builder --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/UfukNode/Noustiny.git skills-src && mkdir -p .claude/skills && cp -r skills-src/hermes-additions/skills/creative/visual-prompt-builder .claude/skills/visual-prompt-builder && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
visual-prompt-builder
GitHub stars
181
Token cost
~3.1k tokens
SKILL.md length
1,370 words
Files
1
Skills in repo
13
Repo updated
First seen
Licence
MIT

At a glance

Rewrite a narrative beat into a rich photoreal visual description suitable for text-to-image models.

  • Tasks that involve Diffusion and image models
  • SKILL.md covers Input — single-beat OR batch, Output — STRICT JSON, no…, Rules and Story register table — pick…, plus 3 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • Tasks that involve Image generation

What it does

Visual Prompt Builder is an agent skill from UfukNode/Noustiny. Rewrite a narrative beat into a rich photoreal visual description suitable for text-to-image models. Strips every named intellectual-property reference and replaces each character / prop / place with a precise physical description. Output is strict JSON — one ≤60-word prompt (≤90 for climax beats), zero IP names, built for FLUX / Nano Banana / SeeDream.

Its SKILL.md is about 3.1k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering Diffusion and image models, Image generation and Intellectual property. It works with Google Gemini. The repository describes itself as: An agent native video creation pipeline that runs on top of Hermes Agent. The licence is MIT.

When your agent uses it

  • Tasks that involve Diffusion and image models
  • Tasks that involve Image generation
  • Tasks that involve Intellectual property

Example prompts

  • “/visual-prompt-builder”

What it can do on your machine

Read from SKILL.md and the folder at commit 09a0c82. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are json).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Visual Prompt Builder loads about 3.1k tokens when it runs. Until then it costs about 95 tokens; SKILL.md has 1,370 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~95
When it runs · the whole SKILL.md, loaded when a task matches
~3.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from UfukNode/Noustiny at commit 09a0c82, republished under its MIT licence (© UfukNode). 1,370 words, ~3,071 tokens.

Download SKILL.mdSave it as .claude/skills/visual-prompt-builder/SKILL.md (or your agent's skills folder).
name
visual-prompt-builder
description
Rewrite a narrative beat into a rich photoreal visual description suitable for text-to-image models. Strips every named intellectual-property reference and replaces each character / prop / place with a precise physical description. Output is strict JSON — one ≤60-word prompt (≤90 for climax beats), zero IP names, built for FLUX / Nano Banana / SeeDream.
author
Noustiny
license
MIT
version
1.1.0

Visual Prompt Builder

Convert a narrative beat into a frame-specific visual description. The output is eaten by FLUX / SeeDream / Nano Banana as their user prompt. They are eyes without minds — give them physical specifics (8 feet tall, violet skin, rune-engraved hammer), not narrative shorthand ("Thor").

Input — single-beat OR batch

Single beat:

json
{
  "title": "string",
  "body":  "string",
  "seed":  "string (optional — story logline)",
  "canonBeats": [{ "title": "...", "body": "..." }],
  "characters": { "Thor": "a tall blond warrior god..." },
  "franchise": "avatar-airbender | marvel | null",
  "allow_ip_names": true
}

Batch — same fields except title+body are replaced by a beats array:

json
{
  "beats": [{ "title": "...", "body": "..." }, { "title": "...", "body": "..." }],
  "seed":  "string",
  "canonBeats": [...],
  "characters": {...},
  "franchise": "avatar-airbender",
  "allow_ip_names": true
}
  • franchise — slug from the detector (marvel, avatar-airbender, null, …).
  • allow_ip_names — when true, prompts MAY include the IP character name in addition to the physical description. When false, prompts must stay fully IP-free (the existing sanitised form). Let allow_ip_names control whether "Katara" appears or whether it must be rewritten as "fourteen-year-old brown-skinned Water Tribe girl".

Output — STRICT JSON, no prose, no code fence

Single beat → one object:

json
{
  "storyRegister": "<register line — see table>",
  "prompt":        "<≤60 words (≤90 for climax), ends with storyRegister verbatim>",
  "characters_seen": [{ "source_name": "Thor", "visual_description": "..." }]
}

Batch → JSON array in input order, same storyRegister for every entry:

json
[
  { "storyRegister": "...", "prompt": "...", "characters_seen": [...] },
  { "storyRegister": "...", "prompt": "...", "characters_seen": [...] }
]

First char { or [, last char } or ]. Nothing else. No step-by-step reasoning, no craft notes, no markdown fences.

Rules

  • Fidelity to body — hardest rule. The prompt must render the body's literal stage action. Every concrete physical fact in body (spatial relationship, posture, gesture stop-point, object state, who is where relative to whom) must appear in the prompt.

    • body "Sokka's finger stops short of touching the blue arrow — he doesn't need to." → prompt MUST say "a stocky arctic polar boy with a warrior's ponytail hovers his index finger a centimetre above a glowing blue arrow tattoo, not touching it" ✓
    • prompt "a boy pointing at a glowing forehead" ✗ (drops the "stops short / not touching" beat — the whole reason this is the beat) If the body says a gesture was withheld, the prompt shows the withholding, not the completed version. If the body says someone is trapped inside ice, the prompt shows them inside ice, not on it. Treat the body like a stage direction, not a summary.
  • Carry canon state forward — the single biggest continuity failure mode. Before writing your prompt, read the last canonBeats entry (title + body + imagePrompt when present) and note every visible state change it introduced — these are permanent unless the current beat's body explicitly reverses them. Never let a destroyed/altered object silently revert to its pristine state just because the current body doesn't re-mention it. Apply this generically across any franchise:

    What counts as a persistent visual state change (all genres):

    • Objects damaged / destroyed: shattered glass, cracked ice, collapsed wall, split hull, torn cloak → STAY damaged in the next frame. No auto-repair.
    • Objects moved / consumed / lost: a sword dropped into water, a document burned, a phone thrown off a cliff, a pill swallowed, a key used and left in the lock → the object is where it last was, not back in the character's hand.
    • Characters released / captured / killed / wounded: freed from bonds, pulled from wreckage, bleeding, unconscious, dead → the state carries forward. A character freed from ice in beat N is NOT encased in ice again in beat N+1. A dead character stays dead.
    • Environment altered: doors now open, room on fire, hallway flooded, mask removed, armor dented, face scarred → the alteration persists.
    • Spatial position: a character who climbed a ladder is now on the upper level; a vehicle that crashed is wrecked on the ground.

    Examples of the failure this rule prevents (any franchise, any medium):

    • ❌ Beat N: "Katara shatters the ice block and pulls Aang free." · Beat N+1: body says "Aang drifts in the cold cavern." · Prompt shows Aang back inside an intact iceberg with Katara and Sokka peering in from outside — WRONG. The ice is already shattered; Aang is already out. Prompt MUST show broken ice shards on the cavern floor with Aang standing amid them, not a new intact prison.
    • ❌ Beat N: "Tony crushes the gauntlet; the stones go dark." · Beat N+1: body says "Tony kneels in the rubble." · Prompt shows a glowing intact gauntlet on Tony's hand — WRONG. The gauntlet is destroyed; prompt MUST show burnt/twisted metal on his arm or no gauntlet at all.
    • ❌ Beat N: "Walt drops the phone into the river." · Beat N+1: body says "Walt walks down the dirt road." · Prompt shows a phone in Walt's hand — WRONG. The phone is gone; his hand is empty.

    Rule of thumb: if beat N's image shows the world visibly changed (something broken, someone freed, fire lit, person wounded), beat N+1's prompt must inherit that changed world visually, not reset it. Read canonBeats[canonBeats.length - 1].imagePrompt when present — it tells you what the last frame actually depicted.

  • Prompt MUST open with the established setting. First phrase of the prompt is the location/environment inherited from the most recent canonBeats entry. Do NOT lead with emotion or gesture and trust the model to infer location — if the parent beat is inside a cracking iceberg at the South Pole, the child prompt starts "Inside the cracking South Pole iceberg, …" even if the child body doesn't re-mention it. Models latch onto the first phrase; an abstract opener ("Overwhelmed by sudden faces, Aang gasps…") makes them invent a new setting (desert ruins, stone statues) instead of reusing the iceberg.

    • canon "Aang opens his eyes inside a cracking iceberg" + child body "Overwhelmed by the sudden faces, Aang gasps and airbends, shooting upward" → prompt: "Inside the cracking South Pole iceberg, Aang, a twelve-year-old bald monk boy with a blue arrow tattoo, looks up through the fractured ice at Katara and Sokka; he gasps and airbends violently, a swirling vortex of wind propelling him upward through the shattered shell, …" ✓
    • prompt "Overwhelmed by sudden faces beneath him, Aang gasps and airbends, shoots upward above stone temple ruins, …" ✗ (drops iceberg, drops Katara/Sokka, invents "stone ruins")
  • Resolve ambiguous nouns against canon. The word "faces" in a body usually means the human faces already in the scene (Katara, Sokka). Never interpret it as decorative motifs, stone carvings, or crowds the canon hasn't established.

  • IP naming rule — conditional on allow_ip_names:

    • allow_ip_names: true → prompts MAY keep canonical IP names in addition to rich physical descriptions. "Katara kneels beside the fractured ice…" is fine; back it up with the registry description so a model that doesn't recognise the name still renders faithfully.
    • allow_ip_names: false → prompts MUST be IP-free. Thor → thunder-god warrior, Mjolnir → rune-engraved war hammer, Stormbreaker → short-handled thunder axe, Infinity Gauntlet → six-gem golden armor glove, Thanos → eight-foot violet-skinned titan warlord, Na'vi → ten-foot blue-skinned forest humanoid, Aang → bald monk boy with a blue arrow tattoo.
  • Reuse characters[name] verbatim when the beat names that character. If a character is on-frame but missing from the map, synthesise one description and include it in characters_seen.

  • Every pronoun resolved to its noun phrase. No "he / she / it" in the prompt.

  • storyRegister is derived from seed once and appended verbatim to the prompt. Consistent storyRegister across every beat locks the storyboard's look.

  • Climax beats (mood = climax): budget 90 words, may add one intensifier clause after the register.

  • Non-climax beats: 60 words, no intensifier.

Show full SKILL.md (245 more words)Show less

Story register table — pick one row from the seed

Seed fingerprintstoryRegister
Superhero finale, ensemble (Marvel/DC vibe)superhero-finale, operatic cosmic, IMAX 70mm color grade, high dynamic range, heroic silhouettes
Space opera, starships, alien worldsspace-opera, practical-effects cinematic, anamorphic lens flares, deep navy shadows, amber cockpit glow
Anime / bending / cel-shade (Airbender, Naruto, One Piece)animated-feature, cel-shaded 2D animation, crisp ink outlines, vibrant saturated palette, Studio Mir-style atmospheric shading
High fantasy, sword and sorceryhigh-fantasy epic, oil-painting detail, volumetric cathedral light, earth-and-ember palette
Noir / neo-noir crimeneo-noir, chiaroscuro lighting, desaturated navy and sodium yellow, wet asphalt reflections
Horror, supernatural dreadgothic horror, candle-flame key light, cold mist, near-black shadows, shallow focus terror
Domestic drama, literary interiorquiet literary drama, natural window light, muted earth tones, medium-format stills
Coming-of-age, YAwarm coming-of-age, golden-hour sun, pastel sky, soft grain
Period piece / historicalperiod drama, candle-lit interiors, hand-dyed fabric tones, academy 1.66 framing
Cyberpunk megacitycyberpunk neon-noir, magenta and teal key lights, wet neon reflections, low-angle anamorphic
Unclear / defaultcinematic, balanced naturalism, 35mm film

Same row for every beat in the story — do not vary mid-storyboard.

Prompt composition order

primary subject → their posture/action → what they hold → secondary subjects → setting → lighting → storyRegister (verbatim).

Anti-patterns

  • Never name a copyrighted character, place, faction, or artifact.
  • Never break the 60/90 word budget (storyRegister counts toward it).
  • Never narrate internal state ("feeling scared") — show the body.
  • Never output reasoning steps or "Step 1…" sections. JSON only.
  • Never wrap JSON in triple backticks.

Example

Input:

json
{
  "title": "Thor takes it, burns with it",
  "body":  "Thor clamps the gauntlet shut; the stones flood his arm. Skin goes white, then blackens.",
  "seed":  "In the final battle against Thanos, every Avenger must choose what to sacrifice.",
  "characters": {"Thor": "a tall blond bearded warrior god, scuffed dark plate armor with crimson cape"}
}

Output:

json
{
  "storyRegister": "superhero-finale, operatic cosmic, IMAX 70mm color grade, high dynamic range, heroic silhouettes",
  "prompt": "A tall blond bearded warrior god clamps a six-gem golden armor glove on his left fist, skin bleaching white then blackening as radiant light of six stones floods his arm; standing on a ruined temple plaza, dust and embers, low amber rim-light, superhero-finale, operatic cosmic, IMAX 70mm color grade, high dynamic range, heroic silhouettes.",
  "characters_seen": [
    {"source_name": "Thor", "visual_description": "a tall blond bearded warrior god, scuffed dark plate armor with crimson cape"}
  ]
}

© UfukNode, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in hermes-additions/skills/creative/visual-prompt-builder of UfukNode/Noustiny.

Open the folder on GitHubat commit 09a0c82

Compare with similar skills

Visual Prompt Builder next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Visual Prompt Builder compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Visual Prompt Builder this skillUfukNode/Noustiny181—~3.1kAutomated safety check: PassMIT
9Router Image Generationdecolua/9router30k—~830Automated safety check: PassMIT
AI Image Prompts SkillLeoYeAI/openclaw-master-skills2.2k—~4.3kAutomated safety check: PassMIT
ImageNexus-JPF/note-companion8692 repos~3.9kAutomated safety check: PassMIT
Higgsfield Image ShotsOSideMedia/higgsfield-ai-prompt-skill697—~5.1kAutomated safety check: PassMIT
LoRA Space Builderhuggingface/skills11k2 repos~8.4kAutomated safety check: PassApache-2.0

Similar skills

  • Generates images through a 9Router gateway's image endpoint, with model discovery, the request fields and per-provider quirks for OpenAI, Gemini, MiniMax and others.

    30k GitHub stars~830 tokensUpdated 6 days ago
    Media & CreativeAuto-check passed
  • AI Image Prompts Skill

    LeoYeAI/openclaw-master-skills

    Recommend curated prompts from a 10,000+ real-world image generation prompt library.

    2.2k GitHub stars~4.3k tokensUpdated 2 mo ago
    Media & CreativeAuto-check passed
  • Image

    Nexus-JPF/note-companion

    When the user wants to create, generate, edit, or optimize images for marketing — blog heroes, social graphics, product mockups, profile banners, listing visuals, or brand assets.

    869 GitHub starsUsed in 2 repos~3.9k tokens
    Media & CreativeAuto-check passed
  • Higgsfield Image Shots

    OSideMedia/higgsfield-ai-prompt-skill

    A skill your agent uses when the user wants to generate a cinematic still image on Higgsfield, asks about shot framing, camera angle, or composition for image prompts, needs a specific shot type…

    697 GitHub stars~5.1k tokensUpdated 10 days ago
    Media & CreativeAuto-check passed
  • LoRA Space Builder

    huggingface/skills

    Official

    Builds and publishes a Gradio demo on Hugging Face Spaces for a LoRA, with the pipeline, UI and settings chosen to match that LoRA's task and model card.

    11k GitHub starsUsed in 2 repos~8.4k tokens
    AI & LLM EngineeringAuto-check passed
  • Qwen Txt2img

    artokun/comfyui-mcp

    Build Qwen Image 2512 text-to-image workflows with QwenImageIntegratedKSampler, separate component loading, lightning LoRAs, and fine-tuned model variants

    793 GitHub stars~3.3k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed

More from UfukNode/Noustiny

All 13 skills in this repo
  • Grades whether a downstream story beat still holds after an upstream insertion, returning a strict JSON verdict of still_valid, needs_rewrite or must_delete.

    181 GitHub stars~2.3k tokensUpdated 5 mo ago
    Auto-check passed
  • Narrative Writer

    UfukNode/Noustiny

    Turns an ordered list of canon story beats into one continuous piece of present-tense prose with tonal continuity and motif carry-through, never JSON or tool calls.

    181 GitHub stars~1.9k tokensUpdated 5 mo ago
    Auto-check passed
  • Expands a one-line author intent into a full story beat that fits between two known nodes of a canon path, returned as strict JSON.

    181 GitHub stars~3.1k tokensUpdated 5 mo ago
    Auto-check passed
  • Story Scene Composer

    UfukNode/Noustiny

    Groups every node of an existing branching story tree into scenes, titles each scene, assigns an act and a one-word motif, and returns the result as JSON.

    181 GitHub stars~2.3k tokensUpdated 5 mo ago
    Auto-check passed
  • Story Copyright Detector

    UfukNode/Noustiny

    Classifies a story seed as known, inspired or original IP and returns strict JSON naming the franchise and the image model to use.

    181 GitHub stars~1.7k tokensUpdated 5 mo ago
    Auto-check passed
  • Chooses the pace, color palette, transition and length of the opening montage for a Noustiny audiobook storybook when the user leaves those settings open.

    181 GitHub stars~3.2k tokensUpdated 5 mo ago
    Auto-check passed

Works with

Questions about Visual Prompt Builder

What does Visual Prompt Builder do?

Rewrite a narrative beat into a rich photoreal visual description suitable for text-to-image models. Visual Prompt Builder is an agent skill from UfukNode/Noustiny. Rewrite a narrative beat into a rich photoreal visual description suitable for text-to-image models.

When should I use Visual Prompt Builder?

Visual Prompt Builder fits situations like: tasks that involve Diffusion and image models; tasks that involve Image generation; tasks that involve Intellectual property.

How do I install Visual Prompt Builder in Claude Code?

Run `npx skills add UfukNode/Noustiny --skill visual-prompt-builder -a claude-code`. Or copy the skill folder (hermes-additions/skills/creative/visual-prompt-builder in UfukNode/Noustiny) into .claude/skills/visual-prompt-builder in your project. Claude Code loads it when a task matches its description.

How do I install Visual Prompt Builder in Codex?

Run `npx skills add UfukNode/Noustiny --skill visual-prompt-builder -a codex`. Or copy the skill folder (hermes-additions/skills/creative/visual-prompt-builder in UfukNode/Noustiny) into .agents/skills/visual-prompt-builder in your project. Codex loads it when a task matches its description.

Can I use Visual Prompt Builder in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add UfukNode/Noustiny --skill visual-prompt-builder -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/visual-prompt-builder, .gemini/skills/visual-prompt-builder, .github/skills/visual-prompt-builder and .opencode/skills/visual-prompt-builder in your project.

What does Visual Prompt Builder need to run?

SKILL.md names no scripts, command-line tools or credentials: Visual Prompt Builder is instructions for the agent only.

Does Visual Prompt Builder access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Visual Prompt Builder safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Visual Prompt Builder use?

Visual Prompt Builder is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Visual Prompt Builder use?

About 3.1k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Visual Prompt Builder?

Skills that share tags, products or a category with Visual Prompt Builder: 9Router Image Generation (decolua/9router, 30k stars), AI Image Prompts Skill (LeoYeAI/openclaw-master-skills, 2.2k stars), Image (Nexus-JPF/note-companion, 869 stars) and Higgsfield Image Shots (OSideMedia/higgsfield-ai-prompt-skill, 697 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Visual Prompt Builder?

UfukNode (a GitHub user) maintains it in UfukNode/Noustiny, which has 181 GitHub stars. The repository holds 13 skills in this directory. The repository was last updated on May 4, 2026.

Source: UfukNode/Noustiny on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.