Agent skill

Cursor Image Generation

by tmcfarlane in tmcfarlane/oh-my-cursor

Generate and iterate images in Cursor using the built-in image model and strong prompts.

MITAuto-check passedMedia & Creative

Install Cursor Image Generation

skills CLI
$ npx skills add tmcfarlane/oh-my-cursor --skill cursor-image-generation -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install tmcfarlane/oh-my-cursor cursor-image-generation --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/tmcfarlane/oh-my-cursor.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/cursor-image-generation .claude/skills/cursor-image-generation && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
cursor-image-generation
GitHub stars
110
Token cost
~1.8k tokens
SKILL.md length
758 words
Files
1
Skills in repo
14
Repo updated
First seen
Licence
MIT

At a glance

Generate and iterate images in Cursor using the built-in image model and strong prompts.

  • Works in 3 steps: Infer or ask for missing constraints… → Rewrite the ask into one structured… → Call GenerateImage with the rewritten…
  • Marketing visuals
  • SKILL.md covers Rough prompt in, strong prompt…, When to use this skill, Core principles (Nano Banana /… and Anti-patterns, plus 4 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Cursor Image Generation is an agent skill from tmcfarlane/oh-my-cursor. Generate and iterate images in Cursor using the built-in image model and strong prompts. Use when creating icons, illustrations, UI mockups, diagrams, marketing visuals, or any raster asset from a text description or reference image.

Its SKILL.md is about 1.8k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Media & Creative, covering Image generation, Icons and illustration and UI design. It works with Google Gemini. The repository describes itself as: Like “oh-my-opencode”, but for Cursor IDE. Multi-agent orchestration, natively, using nothing but a few config files. The licence is MIT.

When your agent uses it

  • Marketing visuals
  • Any raster asset from a text description
  • Reference image

Example prompts

  • “/cursor-image-generation”

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. Infer or ask for missing constraints (medium, aspect ratio, style, brand colors, text to render).
  2. Rewrite the ask into one structured prompt (or a tight second pass) using the principles below.
  3. Call GenerateImage with the rewritten prompt only.

What it can do on your machine

Read from SKILL.md and the folder at commit 5bad458. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • blog.google
    • cloud.google.com
    • deepmind.google

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Cursor Image Generation loads about 1.8k tokens when it runs. Until then it costs about 64 tokens; SKILL.md has 758 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~64
When it runs · the whole SKILL.md, loaded when a task matches
~1.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from tmcfarlane/oh-my-cursor at commit 5bad458, republished under its MIT licence (© tmcfarlane). 758 words, ~1,762 tokens.

Download SKILL.mdSave it as .claude/skills/cursor-image-generation/SKILL.md (or your agent's skills folder).
name
cursor-image-generation
description
Generate and iterate images in Cursor using the built-in image model and strong prompts. Use when creating icons, illustrations, UI mockups, diagrams, marketing visuals, or any raster asset from a text description or reference image.
metadata.author
oh-my-cursor
metadata.version
1.0.0

Cursor image generation (Nano Banana Pro)

Generate images in the Cursor agent using the GenerateImage tool. Image generation is backed by Google Nano Banana Pro. Previews are saved under assets/ by default unless you specify otherwise.

This skill is about prompting and workflow, not about replacing Figma or vector code (use other skills for those).

Rough prompt in, strong prompt out

The user may give a short or vague request (“a hero for the login page”, “cyberpunk icon”). Do not pass that string raw to GenerateImage when it lacks the layers this skill describes. Instead:

  1. Infer or ask for missing constraints (medium, aspect ratio, style, brand colors, text to render).
  2. Rewrite the ask into one structured prompt (or a tight second pass) using the principles below.
  3. Call GenerateImage with the rewritten prompt only.

The skill is the contract: the agent’s job is to expand and sharpen the user’s intent before generation, then iterate with deltas.

When to use this skill

  • User asks for an image, icon, hero visual, diagram look, mockup still, or iteration on an existing generated image.
  • You need text in the image (titles, labels, buttons in a mockup).
  • User uploads a reference image and wants a variation or edit-style direction.

Core principles (Nano Banana / Gemini image family)

The bullets below are a condensed synthesis of common guidance for Nano Banana Pro / Gemini image models — not verbatim quotes. For authoritative wording and edge cases, use the References below.

They align with the spirit of Google’s public guides (prompt tips, Google Cloud guide, DeepMind prompt guide):

  1. Brief a human artist — Use clear, grammatical sentences. Avoid keyword soup ("cyber, 4k, hdr, epic") unless you deliberately want a tag-like aesthetic.
  2. Layer the description — Subject → action/pose → environment → camera (wide shot, isometric, macro) → lighting (soft window light, neon rim, overcast) → materials (brushed aluminum, matte paper, glass) → style (editorial photo, flat illustration, low-poly 3D render).
  3. Text in images — Put exact wording in double quotes and specify typography feel (e.g. "bold geometric sans", "narrow serif for headlines"). Ask for legibility and high contrast if the text is important.
  4. Aspect ratio and framing — State orientation (square, 16:9 landscape, 9:16 story) and safe margins if the asset will be cropped (e.g. app icon: centered subject, padding).
  5. Edit, don't always re-roll — If the image is roughly right, ask for specific changes ("change the background to warm beige", "make the logo 20% larger", "remove the extra person on the left") instead of a full new prompt.
  6. Reference images — When the user supplies a reference, describe what to keep (palette, mood, composition) and what to change so the model does not drift.
Show full SKILL.md (317 more words)Show less

Anti-patterns

  • Vague superlatives with no visual anchor: “make it more beautiful / premium / modern” — always attach concrete cues (materials, palette, era, reference).
  • Contradictory constraints in one shot: “minimalist” + “dense infographic” + “single hero object” — split into steps or iterations.
  • Ignoring the use case: icon vs hero vs print — state target size or viewing distance when it matters.

Workflow

  1. Capture — Accept the user’s brief even if it is one line; note gaps.
  2. Clarify — Output medium (icon, social, slide, mockup), rough dimensions or aspect ratio, brand colors if any, and must-have vs nice-to-have (ask only when blocking).
  3. Rewrite — Produce the full prompt using the layering order above (this is the “good practices” step).
  4. Generate — Call GenerateImage with the rewritten description. Prefer saving to assets/ with a descriptive filename (e.g. assets/hero-spring-campaign.png).
  5. Iterate — If close: issue a delta prompt; if wrong: adjust the layer that failed (camera, lighting, style) before scrapping everything.
  6. Report — Return file path(s), the final prompt (or summary), and optional next iterations.

Prompting patterns (copy and adapt)

Square brackets [like-this] mark placeholders to fill in. Double quotes "..." mark exact text the model should render in the image.

App icon (square, legible at small size)

text
Square app icon, centered symbol of [subject], flat vector style with subtle depth, limited palette [colors], 10% safe margin from edges, no tiny text, crisp edges, high contrast on [background tone].

UI mockup still (marketing)

text
Photorealistic product screenshot of a [mobile/web] app, [screen name] view, centered device, soft studio lighting, neutral background, clean sans UI. Render the following text exactly: headline "[headline text]", button label "[CTA text]". Modern SaaS aesthetic.

Illustration (not photo)

text
Editorial illustration of [subject], [mood], limited palette [colors], visible brush texture or clean vector shapes (pick one), generous whitespace, no photorealistic faces unless requested.

Diagram / concept

text
Isometric diagram of [system], simple shapes, light grid, high contrast lines, no clutter, presentation slide style. Label the zones exactly: "[zone A label]", "[zone B label]".

Examples: weak vs stronger

WeakStronger
"A nice logo for my app"Minimal wordmark for a productivity app, lowercase sans-serif feel, single accent color #2563EB on white, generous letter-spacing, no icon, horizontal logo lockup.
"Cyberpunk city"Wide 16:9 cinematic shot of a rainy cyberpunk street at night, neon reflections on wet asphalt, single vanishing point, shallow depth of field, no readable text, teal and magenta accents.
"Fix the image"Keep the same subject and composition; change only the background to soft gradient from #0f172a to #1e293b; leave lighting on the subject unchanged.

References

© tmcfarlane, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/cursor-image-generation of tmcfarlane/oh-my-cursor.

Open the folder on GitHubat commit 5bad458

Compare with similar skills

Cursor Image Generation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Cursor Image Generation compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Cursor Image Generation this skilltmcfarlane/oh-my-cursor110—~1.8kAutomated safety check: PassMIT
Imagensanjay3290/ai-skills4317 repos~657Automated safety check: PassApache-2.0
Sf Diagram NanobananaproJaganpro/sf-skills424—~1.6kAutomated safety check: PassMIT
Imagennexu-io/open-design100k—~279Automated safety check: PassApache-2.0
Nano Bananakkoppenhaver/cc-nano-banana3781 repos~1.4kAutomated safety check: PassMIT
FigureMuuuun/luxas1.2k—~1.2kAutomated safety check: PassMIT

Similar skills

  • Imagen

    sanjay3290/ai-skills

    Generate images using Google Gemini's image generation capabilities.

    431 GitHub starsUsed in 7 repos~657 tokens
    Media & CreativeAuto-check passed
  • Sf Diagram Nanobananapro

    Jaganpro/sf-skills

    AI-powered image generation for Salesforce visuals via Nano Banana Pro.

    424 GitHub stars~1.6k tokensUpdated 5 mo ago
    Media & CreativeAuto-check passed
  • Imagen

    nexu-io/open-design

    Generate images using Google Gemini's image generation API for UI mockups, icons, illustrations, and visual assets.

    100k GitHub stars~279 tokensUpdated today
    Media & CreativeAuto-check passed
  • Nano Banana

    kkoppenhaver/cc-nano-banana

    REQUIRED for all image generation requests. An agent skill from kkoppenhaver/cc-nano-banana.

    378 GitHub starsUsed in 1 repo~1.4k tokens
    Media & CreativeAuto-check passed
  • Figure

    Muuuun/luxas

    Hybrid figure pipeline (Nano Banana raster + rembg background removal + TikZ vector assembly).

    1.2k GitHub stars~1.2k tokensUpdated 1 mo ago
    Media & CreativeAuto-check passed
  • AI Multimodal

    Microck/ordinary-claude-skills

    Process and generate multimedia content using Google Gemini API.

    403 GitHub starsUsed in 1 repo~2.7k tokens
    Media & CreativeAuto-check: notes

More from tmcfarlane/oh-my-cursor

All 14 skills in this repo
  • Docs Write

    tmcfarlane/oh-my-cursor

    Write documentation following Metabase's conversational, clear, and user-focused style.

    110 GitHub starsUsed in 3 repos~716 tokens
    Auto-check: notes
  • Debugging

    tmcfarlane/oh-my-cursor

    Systematic 4-phase debugging with root cause investigation. An agent skill from tmcfarlane/oh-my-cursor.

    110 GitHub stars~4.4k tokensUpdated 3 mo ago
    Auto-check passed
  • Documentation Engineer

    tmcfarlane/oh-my-cursor

    Technical documentation expert for creating clear, comprehensive documentation.

    110 GitHub stars~881 tokensUpdated 3 mo ago
    Auto-check passed
  • Planning

    tmcfarlane/oh-my-cursor

    Technical implementation planning and architecture design. An agent skill from tmcfarlane/oh-my-cursor.

    110 GitHub stars~815 tokensUpdated 3 mo ago
    Auto-check passed
  • Codebase Search

    tmcfarlane/oh-my-cursor

    Search and navigate large codebases efficiently. An agent skill from tmcfarlane/oh-my-cursor.

    110 GitHub starsUsed in 1 repo~2.8k tokens
    Auto-check: notes
  • Design Patterns Implementation

    tmcfarlane/oh-my-cursor

    Apply appropriate design patterns (Singleton, Factory, Observer, Strategy, etc.) to solve architectural problems.

    110 GitHub stars~2.2k tokensUpdated 3 mo ago
    Auto-check passed

Works with

Questions about Cursor Image Generation

What does Cursor Image Generation do?

Generate and iterate images in Cursor using the built-in image model and strong prompts. Cursor Image Generation is an agent skill from tmcfarlane/oh-my-cursor. Generate and iterate images in Cursor using the built-in image model and strong prompts.

When should I use Cursor Image Generation?

Cursor Image Generation fits situations like: marketing visuals; any raster asset from a text description; reference image.

How do I install Cursor Image Generation in Claude Code?

Run `npx skills add tmcfarlane/oh-my-cursor --skill cursor-image-generation -a claude-code`. Or copy the skill folder (skills/cursor-image-generation in tmcfarlane/oh-my-cursor) into .claude/skills/cursor-image-generation in your project. Claude Code loads it when a task matches its description.

How do I install Cursor Image Generation in Codex?

Run `npx skills add tmcfarlane/oh-my-cursor --skill cursor-image-generation -a codex`. Or copy the skill folder (skills/cursor-image-generation in tmcfarlane/oh-my-cursor) into .agents/skills/cursor-image-generation in your project. Codex loads it when a task matches its description.

Can I use Cursor Image Generation in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add tmcfarlane/oh-my-cursor --skill cursor-image-generation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/cursor-image-generation, .gemini/skills/cursor-image-generation, .github/skills/cursor-image-generation and .opencode/skills/cursor-image-generation in your project.

What does Cursor Image Generation need to run?

SKILL.md names no scripts, command-line tools or credentials: Cursor Image Generation is instructions for the agent only.

Does Cursor Image Generation access the network?

SKILL.md names 3 domains. As links in the text: blog.google, cloud.google.com and deepmind.google. This is read from the text; nothing was executed.

Is Cursor Image Generation safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Cursor Image Generation use?

Cursor Image Generation is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Cursor Image Generation use?

About 1.8k tokens (SKILL.md is roughly 7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Cursor Image Generation?

Skills that share tags, products or a category with Cursor Image Generation: Imagen (sanjay3290/ai-skills, 431 stars), Sf Diagram Nanobananapro (Jaganpro/sf-skills, 424 stars), Imagen (nexu-io/open-design, 100k stars) and Nano Banana (kkoppenhaver/cc-nano-banana, 378 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Cursor Image Generation?

tmcfarlane (a GitHub user) maintains it in tmcfarlane/oh-my-cursor, which has 110 GitHub stars. The repository holds 14 skills in this directory. The repository was last updated on July 2, 2026.

Source: tmcfarlane/oh-my-cursor on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.