Agent skill

Gemini Image

by redf0x1 in redf0x1/camofox-browser

Automate image generation on Google Gemini (gemini.google.com) using CamoFox CLI.

MITAuto-check passedMedia & Creative

Install Gemini Image

skills CLI
$ npx skills add redf0x1/camofox-browser --skill gemini-image -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install redf0x1/camofox-browser gemini-image --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/redf0x1/camofox-browser.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/gemini-image .claude/skills/gemini-image && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
gemini-image
GitHub stars
412
Token cost
~1.6k tokens
SKILL.md length
377 words
Files
2 (incl. references)
Skills in repo
5
Repo updated
First seen
Licence
MIT

At a glance

Automate image generation on Google Gemini (gemini.google.com) using CamoFox CLI.

  • Works in 8 steps: Prerequisites → Quick Start (Fast Path — Authenticated) → Authentication Flow → …
  • The user asks to generate image
  • SKILL.md covers 1. Prerequisites, 2. Quick Start (Fast Path —…, 3. Authentication Flow and 4. Image Generation, plus 4 more sections
  • Calls curl; reaches gemini.google.com and accounts.google.com

What it does

Gemini Image is an agent skill from redf0x1/camofox-browser. Automate image generation on Google Gemini (gemini.google.com) using CamoFox CLI. Use when the user asks to "generate image", "create image with Gemini", "AI image generation", "tạo hình ảnh", or needs automated Gemini workflows.

Its SKILL.md is about 1.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including reference files (for example `references/workflow-reference.md`).

It sits in Media & Creative, covering Image generation. It works with Google Gemini. The repository describes itself as: Anti-detection browser server for AI agents — REST API wrapping Camoufox engine with OpenClaw plugin support. The licence is MIT.

When your agent uses it

  • The user asks to generate image
  • Create image with Gemini
  • AI image generation
  • Needs automated Gemini workflows

Example prompts

  • “generate image”
  • “create image with Gemini”
  • “AI image generation”
  • “/gemini-image”

Requirements

  • Pre-approved tools (allowed-tools): Bash(camofox:*)

Workflow steps

8 steps, taken from the step headings in SKILL.md.

  1. Prerequisites
  2. Quick Start (Fast Path — Authenticated)
  3. Authentication Flow
  4. Image Generation
  5. Download Generated Image
  6. Tips & Best Practices
  7. Error Recovery
  8. Verified Test Results

What it can do on your machine

Read from SKILL.md and the folder at commit 4f23dcf. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Bash(camofox:*)

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • curl

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • gemini.google.com
    • accounts.google.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Gemini Image loads about 1.6k tokens when it runs, and up to ~1.7k if it reads all its reference files. Until then it costs about 61 tokens; SKILL.md has 377 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~61
When it runs · the whole SKILL.md, loaded when a task matches
~1.6k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~1.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from redf0x1/camofox-browser at commit 4f23dcf, republished under its MIT licence (© redf0x1). 377 words, ~1,557 tokens.

Download SKILL.mdSave it as .claude/skills/gemini-image/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
gemini-image
description
Automate image generation on Google Gemini (gemini.google.com) using CamoFox CLI. Use when the user asks to "generate image", "create image with Gemini", "AI image generation", "tạo hình ảnh", or needs automated Gemini workflows.
allowed-tools
Bash(camofox:*)

Gemini Image Generation Skill (CLI)

Use this skill for Gemini image generation flows with CamoFox CLI commands only. Use curl only for download file retrieval from the local CamoFox download API.

1. Prerequisites

  • CamoFox server must be running on port 9377.
  • Google account with Gemini access.
  • First-time login requires email/password and 2FA approval on phone.
  • After first login, save and reuse a session profile.

2. Quick Start (Fast Path — Authenticated)

When a valid Gemini session is already saved:

All commands below assume --user gemini. Add it explicitly in multi-session scenarios.

bash
# Load saved session and open Gemini
camofox open https://gemini.google.com/app --user gemini

# Load saved profile
camofox session load <profile-name> --user gemini

# Verify auth (check for profile avatar, no "Sign in" button)
camofox snapshot --user gemini

# Type image prompt
camofox type <prompt-ref> "your image description here" --user gemini

# Submit prompt
camofox press Enter --user gemini

# Wait for image generation to complete (15-180 seconds)
camofox wait 'mat-icon[fonticon="download"]' --timeout 180000 --user gemini

# Download the generated image
camofox click 'mat-icon[fonticon="download"]' --user gemini

# Cleanup
camofox close --user gemini

3. Authentication Flow

Step 3.1: Check Auth State
bash
camofox open https://gemini.google.com/app --user gemini
camofox snapshot --user gemini
# Look for: "Sign in" button -> NOT authenticated
# Look for: Greeting text "Xin chào..." -> authenticated
Step 3.2: Login (if needed)
bash
# Navigate to Google login
camofox navigate https://accounts.google.com --user gemini

# Enter email
camofox fill '[e1]="email@gmail.com"' --user gemini
# OR
camofox type e1 "email@gmail.com" --user gemini
camofox click e4 --user gemini  # Next button

# Wait for password page
camofox wait 'input[type="password"]' --user gemini
camofox snapshot --user gemini

# Enter password
camofox type e2 "password" --user gemini
camofox click e4 --user gemini  # Next button

Security Tip: For production workflows, store credentials in CamoFox Auth Vault:

bash
camofox auth save google-gemini
# Follow interactive prompts to save credentials securely
camofox auth load google-gemini --user gemini
Step 3.3: 2FA (Push Notification)

After entering password, Google may prompt 2FA:

  • Push notification: user taps "Yes" on phone app.
  • Wait for redirect:
bash
camofox wait 'a[href*="myaccount.google.com"]' --timeout 90000 --user gemini
  • Or wait for URL change away from accounts.google.com:
bash
camofox wait navigation --timeout 90000 --user gemini
Step 3.4: Save Session
bash
camofox session save <profile-name> --user gemini
# Reuse next time with: camofox session load <profile-name> --user gemini

4. Image Generation

bash
# Navigate to Gemini (if not already there)
camofox navigate https://gemini.google.com/app --user gemini

# Snapshot to find prompt input ref
camofox snapshot --user gemini
# Prompt textbox is typically the main textarea element

# Type your image description
camofox type <prompt-ref> "Generate an image of a cute red fox wearing sunglasses in digital art style" --user gemini

# Submit the prompt
camofox press Enter --user gemini

# Wait for generation to complete (15-180 seconds)
# The download button appears when generation is done
camofox wait 'mat-icon[fonticon="download"]' --timeout 180000 --user gemini

# Verify the result
camofox snapshot --user gemini
# Key refs after generation:
# - Download button (mat-icon fonticon="download")
# - Like (mat-icon fonticon="thumb_up")
# - Dislike (mat-icon fonticon="thumb_down")
# - Redo
# - Share

5. Download Generated Image

bash
# Click download button (locale-independent CSS selector)
camofox click 'mat-icon[fonticon="download"]' --user gemini

# Check download status
camofox downloads --format json --user gemini
# Parse `downloadId` from JSON output (for example with `jq -r '.[0].id'`)

# Retrieve the file via REST API
# The click response or downloads command shows the download ID
curl -o "generated-image.png" "http://localhost:9377/downloads/<download-id>/content?userId=gemini"

# Cleanup
camofox close --user gemini

6. Tips & Best Practices

  • Session reuse: Always save profile after login. Load before each run to skip login.
  • Locale-agnostic selectors: Use mat-icon[fonticon="..."] instead of text labels.
  • Wait instead of sleep: Use camofox wait with CSS selectors, not fixed delays.
  • Re-snapshot after actions: Element refs (eN) shift after navigation/interaction.
  • Image prompts: Be descriptive; Gemini performs better with detailed prompts.
  • Rate limits: Free tier has daily image generation limits. If limit is reached, wait or switch account.
Show full SKILL.md (128 more words)Show less

7. Error Recovery

ErrorSolution
"Sign in" still visible after session loadSession expired (24-48h). Re-login with email/password + 2FA
wait times out for download buttonImage generation may have failed. Run camofox snapshot and inspect error text
"can't create it right now"Not authenticated. Re-check auth state and login again
"Image Generation Limit Reached"Daily limit reached. Wait 24h or use another account
TAB_NOT_FOUNDTab was invalidated. Open a new tab with camofox open
type does not input textGemini may use custom editor. Fallback: camofox eval 'document.querySelector(".ql-editor").innerHTML = "prompt"'

8. Verified Test Results

  • Account: user@gmail.com (User)
  • Profile: gemini-sky (52 cookies)
  • Test prompt: Generate an image of a cute red fox wearing sunglasses in digital art style
  • Result: 8.1MB PNG (Gemini_Generated_Image_ecu3ujecu3ujecu3.png)
  • Generation time: ~30 seconds
  • Download method: click mat-icon[fonticon="download"]

© redf0x1, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (references) in skills/gemini-image of redf0x1/camofox-browser.

  • SKILL.md
  • references/workflow-reference.md

Open the folder on GitHubat commit 4f23dcf

Compare with similar skills

Gemini Image next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Gemini Image compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Gemini Image this skillredf0x1/camofox-browser412—~1.6kAutomated safety check: PassMIT
AI Image Generation and Editingzhayujie/CowAgent47k—~1.3kAutomated safety check: PassMIT
Logo Generatorop7418/logo-generator-skill2.2k—~1.8kAutomated safety check: NotesNone
Gemini Web Reverse-Engineered ClientJimLiu/baoyu-skills26k5 repos~1.6kAutomated safety check: PassMIT
SEO Image GeneratorAgriciDaniel/claude-seo19k2 repos~2.1kAutomated safety check: PassMIT
BlockRun Image GenerationBlockRunAI/ClawRouter6.6k—~2.1kAutomated safety check: PassMIT

Similar skills

  • Generates or edits images from text prompts through a Python script that picks an image backend based on which API keys are configured.

    47k GitHub stars~1.3k tokensUpdated today
    Media & CreativeAuto-check passed
  • Logo Generator

    op7418/logo-generator-skill

    Generate professional SVG logos and high-end showcase images.

    2.2k GitHub stars~1.8k tokensUpdated 5 mo ago
    Media & CreativeAuto-check: notes
  • Generates text and images through an unofficial, reverse-engineered Gemini Web API, supporting reference images and multi-turn conversations.

    26k GitHub starsUsed in 5 repos~1.6k tokens
    Media & CreativeAuto-check passed
  • SEO Image Generator

    AgriciDaniel/claude-seo

    Generates Open Graph previews, blog hero images, product photos and infographics for SEO use through Gemini image tools and the banana extension.

    19k GitHub starsUsed in 2 repos~2.1k tokens
    Media & CreativeAuto-check passed
  • BlockRun Image Generation

    BlockRunAI/ClawRouter

    Generates or edits images through ClawRouter's local image API, with a choice of models and sizes and payment handled automatically through x402.

    6.6k GitHub stars~2.1k tokensUpdated 3 days ago
    Media & CreativeAuto-check passed
  • Antigravity Gemini Image

    uluckyXH/OpenMOSS

    Generate or edit images using the Antigravity-hosted Gemini image model via the local gateway.

    1.3k GitHub stars~730 tokensUpdated 3 mo ago
    Media & CreativeAuto-check: notes

More from redf0x1/camofox-browser

  • Camofox Browser

    redf0x1/camofox-browser

    Anti-detection browser automation for AI agents. An agent skill from redf0x1/camofox-browser.

    412 GitHub stars~4.6k tokensUpdated 17 days ago
    Auto-check passed
  • Camofox CLI

    redf0x1/camofox-browser

    CamoFox CLI — 50 commands for anti-detection browser automation from the terminal.

    412 GitHub stars~1.3k tokensUpdated 17 days ago
    Auto-check passed
  • Dogfood

    redf0x1/camofox-browser

    QA testing workflow for CamoFox Browser — systematic testing with console capture, error detection, and Playwright tracing.

    412 GitHub stars~1.2k tokensUpdated 17 days ago
    Auto-check passed
  • Reddit

    redf0x1/camofox-browser

    Reddit automation with CamoFox CLI — login, browse, post, comment, reply with anti-detection.

    412 GitHub stars~1.5k tokensUpdated 17 days ago
    Auto-check passed

Works with

Questions about Gemini Image

What does Gemini Image do?

Automate image generation on Google Gemini (gemini.google.com) using CamoFox CLI. Gemini Image is an agent skill from redf0x1/camofox-browser.com) using CamoFox CLI.

When should I use Gemini Image?

Gemini Image fits situations like: the user asks to generate image; create image with Gemini; AI image generation; needs automated Gemini workflows.

How do I install Gemini Image in Claude Code?

Run `npx skills add redf0x1/camofox-browser --skill gemini-image -a claude-code`. Or copy the skill folder (skills/gemini-image in redf0x1/camofox-browser) into .claude/skills/gemini-image in your project. Claude Code loads it when a task matches its description.

How do I install Gemini Image in Codex?

Run `npx skills add redf0x1/camofox-browser --skill gemini-image -a codex`. Or copy the skill folder (skills/gemini-image in redf0x1/camofox-browser) into .agents/skills/gemini-image in your project. Codex loads it when a task matches its description.

Can I use Gemini Image in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add redf0x1/camofox-browser --skill gemini-image -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/gemini-image, .gemini/skills/gemini-image, .github/skills/gemini-image and .opencode/skills/gemini-image in your project.

What does Gemini Image need to run?

Going by SKILL.md and its folder, Gemini Image needs the command-line tools its instructions call (curl). Its frontmatter pre-approves these tools: Bash(camofox:*).

Does Gemini Image access the network?

SKILL.md names 2 domains. In commands or code: gemini.google.com and accounts.google.com; the agent is likely to contact these when it follows the instructions. This is read from the text; nothing was executed.

Is Gemini Image safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Gemini Image use?

Gemini Image is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Gemini Image use?

About 1.6k tokens (SKILL.md is roughly 6.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 139 tokens, read only when the agent opens those files.

What are the alternatives to Gemini Image?

Skills that share tags, products or a category with Gemini Image: AI Image Generation and Editing (zhayujie/CowAgent, 47k stars), Logo Generator (op7418/logo-generator-skill, 2.2k stars), Gemini Web Reverse-Engineered Client (JimLiu/baoyu-skills, 26k stars) and SEO Image Generator (AgriciDaniel/claude-seo, 19k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Gemini Image?

redf0x1 (a GitHub user) maintains it in redf0x1/camofox-browser, which has 412 GitHub stars. The repository holds 5 skills in this directory. The repository was last updated on September 22, 2026.

Source: redf0x1/camofox-browser on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.