Agent skill

Grok Imagine

by nikships in nikships/droidproxy

Generate or edit images via Grok Imagine Image 2.0 through DroidProxy (no XAIAPIKEY).

MITAuto-check passedBackend & APIs

Install Grok Imagine

skills CLI
$ npx skills add nikships/droidproxy --skill grok-imagine -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install nikships/droidproxy grok-imagine --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/nikships/droidproxy.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/grok-imagine .claude/skills/grok-imagine && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
grok-imagine
GitHub stars
122
Token cost
~2.2k tokens
SKILL.md length
1,052 words
Files
1
Skills in repo
4
Repo updated
First seen
Licence
MIT

At a glance

Generate or edit images via Grok Imagine Image 2.0 through DroidProxy (no XAIAPIKEY).

  • Works in 4 steps: Generate on a flat, uniform background… → Key it out locally (PIL, ImageMagick). → Recolor the subject for the background… → …
  • The user asks to generate
  • SKILL.md covers Prompting, Generate, Edit and Parameters, plus 3 more sections
  • Calls jq and curl; reaches imgen.x.ai; needs XAI_API_KEY

What it does

Grok Imagine is an agent skill from nikships/droidproxy. Generate or edit images via Grok Imagine Image 2.0 through DroidProxy (no XAIAPIKEY). Use when the user asks to generate, create, draw, imagine, or edit an image, illustration, icon, mockup, or similar visual asset and DroidProxy Grok OAuth is available. Prefer this over inventing image URLs or base64.

Its SKILL.md is about 2.2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Backend & APIs, covering Image editing and OAuth and OpenID Connect. It works with xAI Grok. The repository describes itself as: macOS menu bar app that proxies Claude/Codex/Grok + More subscriptions for use with Factory Droid CLI. Actively maintained and updated with new model releases! The licence is MIT.

When your agent uses it

  • The user asks to generate
  • Similar visual asset and DroidProxy Grok OAuth is available

Example prompts

  • “/grok-imagine”

Requirements

  • A credential in XAI_API_KEY

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Generate on a flat, uniform background that contrasts with the subject.
  2. Key it out locally (PIL, ImageMagick).
  3. Recolor the subject for the background they named, before you finish.
  4. Say what you did, and mention the ink color you chose so they can ask for

What it can do on your machine

Read from SKILL.md and the folder at commit 9fdd443. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • jq
    • curl

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • imgen.x.ai

    Also links to:

    • docs.x.ai

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • XAI_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Grok Imagine loads about 2.2k tokens when it runs. Until then it costs about 80 tokens; SKILL.md has 1,052 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~80
When it runs · the whole SKILL.md, loaded when a task matches
~2.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from nikships/droidproxy at commit 9fdd443, republished under its MIT licence (© nikships). 1,052 words, ~2,212 tokens.

Download SKILL.mdSave it as .claude/skills/grok-imagine/SKILL.md (or your agent's skills folder).
name
grok-imagine
description
Generate or edit images via Grok Imagine Image 2.0 through DroidProxy (no XAI_API_KEY). Use when the user asks to generate, create, draw, imagine, or edit an image, illustration, icon, mockup, or similar visual asset and DroidProxy Grok OAuth is available. Prefer this over inventing image URLs or base64.
version
3.0.0

Grok Imagine Image 2.0 (via DroidProxy)

Generate and edit images with grok-imagine-image-2.0 using the SuperGrok / X Premium+ OAuth session that powers DroidProxy chat. This is the current Imagine image model (text or image → image). Do not use the older grok-imagine-image / grok-imagine-image-quality ids unless the user explicitly asks for them.

Two things you can't guess and that everything else depends on:

  • The endpoint is http://localhost:8317, not api.x.ai. DroidProxy TLS-forwards any request whose model starts with grok- and injects the OAuth bearer.
  • Auth is Authorization: Bearer dummy-not-used. The real token comes from the proxy. Never ask for or use an XAI_API_KEY.

Use for generating or editing visual assets. Not for understanding an image the user already attached — that's chat multimodal, not this skill.

Prompting

xAI rewrites the prompt with an upsampler before generation, so write a clear scene rather than a tag soup. If the user gives a detailed prompt or asks you to use theirs, use it verbatim. Otherwise craft 2–5 sentences of natural prose:

subject → action/pose → setting → style → composition → lighting/mood → key details

  • Front-load the subject. Give strong high-level direction for mood, composition, lighting, and style without specifying every pixel.
  • State what to include, not what to exclude. No negative prompts.
  • One coherent scene per prompt. Match aspect_ratio to the use case: icons 1:1, banners 16:9, stories 9:16.
  • For edits, describe only what changes and what must stay the same.
  • For multi-image edits, address sources as <IMAGE_0>, <IMAGE_1>, <IMAGE_2> in the prompt (max 3). Combining subjects, transferring a style, or composing a scene from references is the intended use.
  • Ground real people, brands, places, and “current/latest” facts with a web search first and put the verified details in the prompt. For a named real person, edit from a real reference photo — do not generate the likeness from text alone.

Generate

bash
RESP=$(curl -sS http://localhost:8317/v1/images/generations \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer dummy-not-used" \
  -d "$(jq -n --arg p "A red balloon on a wooden table, soft natural light" \
    '{model:"grok-imagine-image-2.0", prompt:$p, aspect_ratio:"1:1",
      resolution:"1k", response_format:"b64_json"}')")

if B64=$(printf '%s' "$RESP" | jq -er '.data[0].b64_json' 2>/dev/null); then
  printf '%s' "$B64" | base64 -d > out.jpg
else
  printf '%s\n' "$RESP" >&2   # errors are plain text, not JSON
  exit 1
fi

Two things that bite here:

  • Check .data[0].b64_json before decoding. On failure there is no b64 field, and piping the error body through base64 -d writes a corrupt file that passes an ls check and only reveals itself when someone opens it.
  • Print the raw body on failure, not jq '.error.message'. Errors come back as text/plain, so parsing them as JSON discards the useful part. The raw text is specific and worth reading — a bad aspect_ratio replies unknown variant `21:9`, expected one of `1:1`, `3:4`, ..., which tells you exactly what to send.

Edit

POST /v1/images/edits, same auth and routing. Pass either image (single) or images (up to 3 references, addressed as <IMAGE_0>, <IMAGE_1>, <IMAGE_2> in the prompt) — never both. Each source is JPEG, PNG, or WebP.

bash
curl -sS http://localhost:8317/v1/images/edits \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer dummy-not-used" \
  -d "$(jq -n --arg p "Make it blue-tinted studio lighting" --arg u "$SRC" \
    '{model:"grok-imagine-image-2.0", prompt:$p,
      image:{url:$u, type:"image_url"},
      response_format:"b64_json"}')"

The source url takes a public HTTPS URL, an xAI file_id (on the image object, mutually exclusive with url), or a data URL. For a local file there's no upload step — inline it as a data URL:

bash
SRC="data:image/jpeg;base64,$(base64 -i ./photo.jpg)"

Omit aspect_ratio on single-image edits; it follows the source, and setting it can crop or letterbox. On multi-image edits it follows the first source unless you override it.

Chain edits by feeding each output back as the next image. That is how you refine, restyle, or correct without starting over.

Parameters

Source of truth: Image Generation, Image Editing, Multi-Image Editing, and the Images REST API.

FieldValuesNotes
modelgrok-imagine-image-2.0Must start with grok- or the proxy won't route it. Always send this id.
promptstringSee Prompting. Required on edits.
aspect_ratio1:1, 3:4, 4:3, 9:16, 16:9, 2:3, 3:2, 9:19.5, 19.5:9, 9:20, 20:9, 1:2, 2:1, autoDefault auto (model picks the ratio). Icons 1:1, banners 16:9, stories 9:16.
resolution1k, 2kDefault 1k. Stay there unless the user asks for print/hi-dpi — 2k returns much larger files.
qualitylow, medium2.0 only. Default medium when omitted. Do not send high, hd, or standard. Omit unless the user asks for low.
response_formaturl, b64_jsonurl is the default but expires fast — prefer b64_json when saving to disk.
n1–10Default 1. Same-prompt variations belong in one request via n, not parallel calls.
image / imagesobject or array (max 3)Edits only. Mutually exclusive. Each item is {url, type:"image_url"} or {file_id}.
storage_options{filename, expires_after?, public_url?}Persists to the xAI Files API; response gains file_output: {file_id, filename}. expires_after is seconds, max 2592000.
userstringOptional end-user id for abuse monitoring. Don't invent one.

Anything outside these enums returns 422, and OpenAI-only fields (size, style, background, output_format, quality:"hd") return 400 or are silently ignored. Don't reach for them.

Show full SKILL.md (328 more words)Show less

Response

json
{"data": [{"b64_json": "...", "mime_type": "image/jpeg", "url": "https://imgen.x.ai/..."}]}

Pick the file extension from mime_type, not from the prompt or the filename the user asked for — 1k usually returns JPEG and 2k often returns PNG, so a .jpg name on PNG bytes is a real outcome. If that contradicts the name they gave, save the correct extension and say why. mime_type may also be image/webp.

With response_format: "url" the URL is temporary (imgen.x.ai) — download it immediately rather than handing it to the user.

No transparency

Imagine has no alpha channel; output is always opaque JPEG, PNG, or WebP. There is no API field that changes this, so don't send one.

When someone asks for a transparent PNG:

  1. Generate on a flat, uniform background that contrasts with the subject.
  2. Key it out locally (PIL, ImageMagick).
  3. Recolor the subject for the background they named, before you finish. Someone asking for transparency has told you where it's going — "over our dark header," "on a white card." Keying preserves whatever ink Imagine happened to render, which is usually near-black, so a mark destined for a dark background comes out invisible. Compare the subject's luminance to the target background and invert it if they collide. That's the difference between a usable asset and one the user has to send back.
  4. Say what you did, and mention the ink color you chose so they can ask for the other one.

Failures

Most errors explain themselves; these three don't:

SymptomFix
Connection refusedDroidProxy isn't running — tell the user to launch it
401, or 401 after a long idleOAuth session dead — Settings → Connect Grok
403 permission / quotaSubscription tier gates API OAuth. A BYOK XAI_API_KEY is the only fallback

On a moderation / safety block, stop. Don't retry and don't paraphrase the prompt to evade the filter. Tell the user it was blocked and offer a different direction.

Report the API's own message rather than guessing, and never invent image content or a URL for a request that failed.

© nikships, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/grok-imagine of nikships/droidproxy.

Open the folder on GitHubat commit 9fdd443

Compare with similar skills

Grok Imagine next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Grok Imagine compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Grok Imagine this skillnikships/droidproxy122—~2.2kAutomated safety check: PassMIT
Grok Mediacalesthio/OpenMontage66k—~950Automated safety check: PassAGPL-3.0
Setup API Keysonilo-ai/skills115—~3kAutomated safety check: PassMIT
Youtube APIericrisco/rsc-harness180—~4kAutomated safety check: PassMIT
Peertubemagnus919/agent-skills115—~3.1kAutomated safety check: PassMIT
Fortify Developmentcoollabsio/coolify63k4 repos~1.9kAutomated safety check: PassMIT

Similar skills

  • Grok Media

    calesthio/OpenMontage

    xAI Grok image and video generation guide covering authentication, endpoints, prompt structure, image editing, reference-image video, and async polling.

    66k GitHub stars~950 tokensUpdated 7 days ago
    Media & CreativeAuto-check passed
  • Setup API Key

    sonilo-ai/skills

    Guides users through connecting to Sonilo — signing in with sonilo login (OAuth, no key), the remote OAuth plugin for Claude Code or Muse Code, or an API key for CI and headless use.

    115 GitHub stars~3k tokensUpdated 2 days ago
    Backend & APIsAuto-check passed
  • Youtube API

    ericrisco/rsc-harness

    A skill your agent uses when wiring code to a real YouTube channel: user OAuth 2.0, resumable videos.insert uploads, editing metadata after publish, pulling views/watch time/retention/traffic from…

    180 GitHub stars~4k tokensUpdated yesterday
    Backend & APIsAuto-check passed
  • Peertube

    magnus919/agent-skills

    Browse PeerTube federated video from the terminal — instance stats, latest videos, video detail, comment threads, channels, accounts, instance-local search, and OAuth2 login with per-instance token…

    115 GitHub stars~3.1k tokensUpdated yesterday
    Backend & APIsAuto-check passed
  • Fortify Development

    coollabsio/coolify

    ACTIVATE when the user works on authentication in Laravel. An agent skill from coollabsio/coolify.

    63k GitHub starsUsed in 4 repos~1.9k tokens
    Backend & APIsAuto-check passed
  • OmniRoute Provider Management

    diegosouzapw/OmniRoute

    Manages AI provider connections, API keys, OAuth flows and connection tests through OmniRoute's REST API across its 327-provider catalog.

    75k GitHub stars~2.4k tokensUpdated today
    Backend & APIsAuto-check passed

More from nikships/droidproxy

  • Droid Model Setup

    nikships/droidproxy

    Configure any third-party model as a custom model in Factory Droid (~/.factory/settings.json): resolve specs (context window, max output tokens, reasoning efforts), write a safe config entry, set…

    122 GitHub stars~2k tokensUpdated yesterday
    Auto-check passed
  • Gpt Image

    nikships/droidproxy

    Generate or edit images via GPT Image 2.5 Flare or Sunburst through DroidProxy Codex OAuth (no OPENAIAPIKEY).

    122 GitHub stars~2.5k tokensUpdated yesterday
    Auto-check passed
  • Release

    nikships/droidproxy

    Build, sign, notarize, and publish a new DroidProxy release.

    122 GitHub stars~1.1k tokensUpdated yesterday
    Auto-check passed

Works with

Questions about Grok Imagine

What does Grok Imagine do?

Generate or edit images via Grok Imagine Image 2.0 through DroidProxy (no XAIAPIKEY). Grok Imagine is an agent skill from nikships/droidproxy.0 through DroidProxy (no XAIAPIKEY).

When should I use Grok Imagine?

Grok Imagine fits situations like: the user asks to generate; similar visual asset and DroidProxy Grok OAuth is available.

How do I install Grok Imagine in Claude Code?

Run `npx skills add nikships/droidproxy --skill grok-imagine -a claude-code`. Or copy the skill folder (skills/grok-imagine in nikships/droidproxy) into .claude/skills/grok-imagine in your project. Claude Code loads it when a task matches its description.

How do I install Grok Imagine in Codex?

Run `npx skills add nikships/droidproxy --skill grok-imagine -a codex`. Or copy the skill folder (skills/grok-imagine in nikships/droidproxy) into .agents/skills/grok-imagine in your project. Codex loads it when a task matches its description.

Can I use Grok Imagine in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add nikships/droidproxy --skill grok-imagine -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/grok-imagine, .gemini/skills/grok-imagine, .github/skills/grok-imagine and .opencode/skills/grok-imagine in your project.

What does Grok Imagine need to run?

Going by SKILL.md and its folder, Grok Imagine needs the command-line tools its instructions call (jq and curl) and credentials named XAI_API_KEY. Our summary lists: A credential in XAI_API_KEY.

Does Grok Imagine access the network?

SKILL.md names 2 domains. In commands or code: imgen.x.ai; the agent is likely to contact it when it follows the instructions. As links in the text: docs.x.ai. This is read from the text; nothing was executed.

Is Grok Imagine safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Grok Imagine use?

Grok Imagine is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Grok Imagine use?

About 2.2k tokens (SKILL.md is roughly 8.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Grok Imagine?

Skills that share tags, products or a category with Grok Imagine: Grok Media (calesthio/OpenMontage, 66k stars), Setup API Key (sonilo-ai/skills, 115 stars), Youtube API (ericrisco/rsc-harness, 180 stars) and Peertube (magnus919/agent-skills, 115 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Grok Imagine?

nikships (a GitHub user) maintains it in nikships/droidproxy, which has 122 GitHub stars. The repository holds 4 skills in this directory. The repository was last updated on October 9, 2026.

Source: nikships/droidproxy on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.