Agent skill

Gemini Vision

by einverne in einverne/dotfiles

Guide for implementing Google Gemini API image understanding - analyze images with captioning, classification, visual QA, object detection, segmentation, and multi-image comparison.

MITAuto-check: notesAI & LLM Engineering

Install Gemini Vision

skills CLI
$ npx skills add einverne/dotfiles --skill gemini-vision -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install einverne/dotfiles gemini-vision --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/einverne/dotfiles.git skills-src && mkdir -p .claude/skills && cp -r skills-src/claude/skills/gemini-vision .claude/skills/gemini-vision && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
gemini-vision
GitHub stars
121
Token cost
~1.6k tokens
SKILL.md length
586 words
Files
8 (incl. scripts, references)
Skills in repo
39
Repo updated
First seen
Licence
MIT

At a glance

Guide for implementing Google Gemini API image understanding - analyze images with captioning, classification, visual QA, object detection, segmentation, and multi-image comparison.

  • Works in 2 steps: Get API Key: Obtain from Google AI Studio → Install SDK: pip install google-genai…
  • Analyzing images
  • SKILL.md covers Quick Start, Core Capabilities, Supported Formats and Available Models, plus 8 more sections
  • Runs Python scripts from its folder; calls python and pip; needs GEMINI_API_KEY

What it does

Gemini Vision is an agent skill from einverne/dotfiles. Guide for implementing Google Gemini API image understanding - analyze images with captioning, classification, visual QA, object detection, segmentation, and multi-image comparison. Use when analyzing images, answering visual questions, detecting objects, or processing documents with vision.

Its SKILL.md is about 1.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 9 other files, including scripts and reference files (for example `README.md`, `references/api-reference.md` and `references/best-practices.md`).

It sits in AI & LLM Engineering, covering Computer vision and Visual regression testing. It works with Google Gemini. The repository describes itself as: my personal dotfiles managed by dotbot, zinit. The licence is MIT.

When your agent uses it

  • Analyzing images
  • Answering visual questions
  • Detecting objects
  • Processing documents with vision

Example prompts

  • “/gemini-vision”

Requirements

  • Python 3
  • A credential in GEMINI_API_KEY
  • Pre-approved tools (allowed-tools): Bash, Read, Write, Edit

Workflow steps

2 steps, taken from the first numbered list in SKILL.md.

  1. Get API Key: Obtain from Google AI Studio
  2. Install SDK: pip install google-genai (Python 3.9+)

What it can do on your machine

Read from SKILL.md and the folder at commit c6c0686. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Bash
    • Read
    • Write
    • Edit

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 3 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python
    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • aistudio.google.com
    • ai.google.dev

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • GEMINI_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Gemini Vision loads about 1.6k tokens when it runs, and up to ~11k if it reads all its reference files. Until then it costs about 77 tokens; SKILL.md has 586 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~77
When it runs · the whole SKILL.md, loaded when a task matches
~1.6k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~11k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteMentions a .env fileSKILL.md:32
    rectory**: `.claude/skills/gemini-vision/.env`
  • NoteMentions a .env fileSKILL.md:37
    3. **Project directory**: `.env` or `.gemini_api_key` in project root
  • NoteMentions a .env fileSKILL.md:39
    commit API keys to version control. Add `.env` to `.gitignore`.
  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Bash, Read, Write, Edit

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from einverne/dotfiles at commit c6c0686, republished under its MIT licence (© einverne). 586 words, ~1,577 tokens.

Download SKILL.mdSave it as .claude/skills/gemini-vision/SKILL.md (or your agent's skills folder). This skill also uses 7 other files; get the full folder from GitHub.
name
gemini-vision
description
Guide for implementing Google Gemini API image understanding - analyze images with captioning, classification, visual QA, object detection, segmentation, and multi-image comparison. Use when analyzing images, answering visual questions, detecting objects, or processing documents with vision.
allowed-tools
Bash, Read, Write, Edit
license
MIT

Gemini Vision API Skill

This skill enables Claude to use Google's Gemini API for advanced image understanding tasks including captioning, classification, visual question answering, object detection, segmentation, and multi-image analysis.

Quick Start

Prerequisites
  1. Get API Key: Obtain from Google AI Studio
  2. Install SDK: pip install google-genai (Python 3.9+)
API Key Configuration

The skill checks for GEMINI_API_KEY in this order:

  1. Process environment variable (recommended)

    bash
    export GEMINI_API_KEY="your-api-key"
  2. Skill directory: .claude/skills/gemini-vision/.env

    GEMINI_API_KEY=your-api-key
  3. Project directory: .env or .gemini_api_key in project root

Security: Never commit API keys to version control. Add .env to .gitignore.

Core Capabilities

Image Analysis
  • Captioning: Generate descriptive text for images
  • Classification: Categorize and identify image content
  • Visual QA: Answer questions about image content
  • Multi-image: Compare and analyze up to 3,600 images
Advanced Features (Model-Specific)
  • Object Detection: Identify and locate objects with bounding boxes (Gemini 2.0+)
  • Segmentation: Create pixel-level masks for objects (Gemini 2.5+)
  • Document Understanding: Process PDFs with vision (up to 1,000 pages)

Supported Formats

  • Images: PNG, JPEG, WEBP, HEIC, HEIF
  • Documents: PDF (up to 1,000 pages)
  • Size Limits:
    • Inline: 20MB max total request size
    • File API: For larger files
    • Max images: 3,600 per request

Available Models

  • gemini-2.5-pro: Most capable, segmentation + detection
  • gemini-2.5-flash: Fast, efficient, segmentation + detection
  • gemini-2.5-flash-lite: Lightweight, segmentation + detection
  • gemini-2.0-flash: Object detection support
  • gemini-1.5-pro/flash: Previous generation

Usage Examples

Basic Image Analysis
bash
# Analyze a local image
python scripts/analyze-image.py path/to/image.jpg "What's in this image?"

# Analyze from URL
python scripts/analyze-image.py https://example.com/image.jpg "Describe this"

# Specify model
python scripts/analyze-image.py image.jpg "Caption this" --model gemini-2.5-pro
Object Detection (2.0+)
bash
python scripts/analyze-image.py image.jpg "Detect all objects" --model gemini-2.0-flash
Multi-Image Comparison
bash
python scripts/analyze-image.py img1.jpg img2.jpg "What's different between these?"
File Upload (for large files or reuse)
bash
# Upload file
python scripts/upload-file.py path/to/large-image.jpg

# Use uploaded file
python scripts/analyze-image.py file://file-id "Caption this"
File Management
bash
# List uploaded files
python scripts/manage-files.py list

# Get file info
python scripts/manage-files.py get file-id

# Delete file
python scripts/manage-files.py delete file-id

Token Costs

Images consume tokens based on size:

  • Small (≤384px both dimensions): 258 tokens
  • Large: Tiled into 768×768 chunks, 258 tokens each

Token Formula:

crop_unit = floor(min(width, height) / 1.5)
tiles = (width / crop_unit) × (height / crop_unit)
total_tokens = tiles × 258

Example: 960×540 image = 6 tiles = 1,548 tokens

Rate Limits

Limits vary by tier (Free, Tier 1, 2, 3):

  • Measured in RPM (requests/min), TPM (tokens/min), RPD (requests/day)
  • Applied per project, not per API key
  • RPD resets at midnight Pacific

Best Practices

Image Quality
  • Use clear, non-blurry images
  • Verify correct image rotation
  • Consider token costs when sizing
Prompting
  • Be specific in instructions
  • Place text after image for single-image prompts
  • Use few-shot examples for better accuracy
  • Specify output format (JSON, markdown, etc.)
File Management
  • Use File API for files >20MB
  • Use File API for repeated usage (saves tokens)
  • Files auto-delete after 48 hours
  • Clean up manually when done
Show full SKILL.md (229 more words)Show less
Security
  • Never expose API keys in code
  • Use environment variables
  • Add API key restrictions in Google Cloud Console
  • Monitor usage regularly
  • Rotate keys periodically

Error Handling

Common errors:

  • 401: Invalid API key
  • 429: Rate limit exceeded
  • 400: Invalid request (check file size, format)
  • 403: Permission denied (check API key restrictions)

Additional Resources

See the references/ directory for:

  • api-reference.md: Detailed API methods and endpoints
  • examples.md: Comprehensive code examples
  • best-practices.md: Advanced tips and optimization strategies

Implementation Guide

When implementing Gemini vision features:

  1. Check API key availability using the 3-step lookup
  2. Choose appropriate model based on requirements:
    • Need segmentation? Use 2.5+ models
    • Need detection? Use 2.0+ models
    • Need speed? Use Flash variants
    • Need quality? Use Pro variants
  3. Validate inputs:
    • Check file format (PNG, JPEG, WEBP, HEIC, HEIF, PDF)
    • Verify file size (<20MB for inline, >20MB use File API)
    • Count images (max 3,600)
  4. Handle responses appropriately:
    • Parse structured output if requested
    • Extract bounding boxes for object detection
    • Process segmentation masks if applicable
  5. Manage files efficiently:
    • Upload large files via File API
    • Reuse uploaded files when possible
    • Clean up after use

Scripts Overview

All scripts support the 3-step API key lookup:

  • analyze-image.py: Main script for image analysis, supports inline and File API
  • upload-file.py: Upload files to Gemini File API
  • manage-files.py: List, get metadata, and delete uploaded files

Run any script with --help for detailed usage instructions.


Official Documentation: https://ai.google.dev/gemini-api/docs/image-understanding

© einverne, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 7 other files (scripts, references) in claude/skills/gemini-vision of einverne/dotfiles.

  • SKILL.md
  • README.md
  • references/api-reference.md
  • references/best-practices.md
  • references/examples.md
  • scripts/analyze-image.py
  • scripts/manage-files.py
  • scripts/upload-file.py

Open the folder on GitHubat commit c6c0686

Compare with similar skills

Gemini Vision next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Gemini Vision compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Gemini Vision this skilleinverne/dotfiles121—~1.6kAutomated safety check: NotesMIT
ModLens Image Vision Bridgeliustack/modlens4.2k—~1.3kAutomated safety check: NotesMIT
Vision Model Retry Policyskuramatata/my-pi-agent114—~1kAutomated safety check: PassNone
Multimodal Media Feature ExtractionTyrealQ/q-skills108—~2kAutomated safety check: NotesMIT
Gemini Video Understandingbenchflow-ai/skillsbench1.8k—~2.4kAutomated safety check: PassApache-2.0
Thumbnail Creatormohitagw15856/pm-claude-skills1.4k—~6.3kAutomated safety check: PassMIT

Similar skills

  • Gives text-only models sight by running the modlens CLI on an image path or URL and returning structured JSON evidence with transcribed text, layout and semantics.

    4.2k GitHub stars~1.3k tokensUpdated 5 days ago
    AI & LLM EngineeringAuto-check: notes
  • Vision Model Retry Policy

    skuramatata/my-pi-agent

    A skill your agent uses when image reading, screenshot analysis, visual diff, design mockup analysis, pasted image handling, or Z.AI Vision MCP fails, times out, returns 429/Too Many Requests, or…

    114 GitHub stars~1k tokensUpdated 3 mo ago
    AI & LLM EngineeringAuto-check passed
  • Extracts pixel, video-frame, speech, music and visual-semantic features from image, video and audio files for research datasets, using local tools or the Gemini API.

    108 GitHub stars~2k tokensUpdated 16 days ago
    Data & AnalyticsAuto-check: notes
  • Gemini Video Understanding

    benchflow-ai/skillsbench

    Analyze videos with Google Gemini API (summaries, Q&A, transcription with timestamps + visual context, scene/timeline detection, video clipping, FPS control, multi-video comparison, and YouTube URL…

    1.8k GitHub stars~2.4k tokensUpdated 2 mo ago
    Media & CreativeAuto-check passed
  • Thumbnail Creator

    mohitagw15856/pm-claude-skills

    Generate article or newsletter thumbnail candidates using the Gemini API from inside Claude Code.

    1.4k GitHub stars~6.3k tokensUpdated yesterday
    Media & CreativeAuto-check passed
  • Segment Anything Model Guide

    Orchestra-Research/AI-Research-SKILLs

    Guide to using Meta's Segment Anything Model for zero-shot image segmentation with point, box or mask prompts, or automatic mask generation.

    13k GitHub starsUsed in 8 repos~3.3k tokens
    AI & LLM EngineeringAuto-check passed

More from einverne/dotfiles

All 39 skills in this repo
  • DOCX

    einverne/dotfiles

    Comprehensive document creation, editing, and analysis with support for tracked changes, comments, formatting preservation, and text extraction.

    121 GitHub starsUsed in 35 repos~2.5k tokens
    Auto-check: notes
  • Chrome Devtools

    einverne/dotfiles

    Browser automation, debugging, and performance analysis using Puppeteer CLI scripts.

    121 GitHub starsUsed in 1 repo~1.6k tokens
    Auto-check: notes
  • PDF

    einverne/dotfiles

    Comprehensive PDF manipulation toolkit for extracting text and tables, creating new PDFs, merging/splitting documents, and handling forms.

    121 GitHub starsUsed in 47 repos~1.8k tokens
    Auto-check passed
  • Gemini Audio

    einverne/dotfiles

    Guide for implementing Google Gemini API audio capabilities - analyze audio with transcription, summarization, and understanding (up to 9.5 hours), plus generate speech with controllable TTS.

    121 GitHub stars~2k tokensUpdated 1 mo ago
    Auto-check: notes
  • Guide for implementing Google Gemini API document processing - analyze PDFs with native vision to extract text, images, diagrams, charts, and tables.

    121 GitHub stars~1.6k tokensUpdated 1 mo ago
    Auto-check: notes
  • PPTX

    einverne/dotfiles

    Presentation creation, editing, and analysis. An agent skill from einverne/dotfiles.

    121 GitHub starsUsed in 38 repos~6.4k tokens
    Auto-check: notes

Works with

Questions about Gemini Vision

What does Gemini Vision do?

Guide for implementing Google Gemini API image understanding - analyze images with captioning, classification, visual QA, object detection, segmentation, and multi-image comparison. Gemini Vision is an agent skill from einverne/dotfiles. Guide for implementing Google Gemini API image understanding - analyze images with captioning, classification, visual QA, object detection, segmentation, and multi-image comparison.

When should I use Gemini Vision?

Gemini Vision fits situations like: analyzing images; answering visual questions; detecting objects; processing documents with vision.

How do I install Gemini Vision in Claude Code?

Run `npx skills add einverne/dotfiles --skill gemini-vision -a claude-code`. Or copy the skill folder (claude/skills/gemini-vision in einverne/dotfiles) into .claude/skills/gemini-vision in your project. Claude Code loads it when a task matches its description.

How do I install Gemini Vision in Codex?

Run `npx skills add einverne/dotfiles --skill gemini-vision -a codex`. Or copy the skill folder (claude/skills/gemini-vision in einverne/dotfiles) into .agents/skills/gemini-vision in your project. Codex loads it when a task matches its description.

Can I use Gemini Vision in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add einverne/dotfiles --skill gemini-vision -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/gemini-vision, .gemini/skills/gemini-vision, .github/skills/gemini-vision and .opencode/skills/gemini-vision in your project.

What does Gemini Vision need to run?

Going by SKILL.md and its folder, Gemini Vision needs Python for the scripts in its folder, the command-line tools its instructions call (python and pip) and credentials named GEMINI_API_KEY. Our summary lists: Python 3; A credential in GEMINI_API_KEY. Its frontmatter pre-approves these tools: Bash, Read, Write, Edit.

Does Gemini Vision access the network?

SKILL.md names 2 domains. As links in the text: aistudio.google.com and ai.google.dev. This is read from the text; nothing was executed.

Is Gemini Vision safe to install?

Our automated static check of SKILL.md found notes only (mentions a .env file; pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Gemini Vision use?

Gemini Vision is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Gemini Vision use?

About 1.6k tokens (SKILL.md is roughly 6.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 9.8k tokens, read only when the agent opens those files.

What are the alternatives to Gemini Vision?

Skills that share tags, products or a category with Gemini Vision: ModLens Image Vision Bridge (liustack/modlens, 4.2k stars), Vision Model Retry Policy (skuramatata/my-pi-agent, 114 stars), Multimodal Media Feature Extraction (TyrealQ/q-skills, 108 stars) and Gemini Video Understanding (benchflow-ai/skillsbench, 1.8k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Gemini Vision?

einverne (a GitHub user) maintains it in einverne/dotfiles, which has 121 GitHub stars. The repository holds 39 skills in this directory. The repository was last updated on September 9, 2026.

Source: einverne/dotfiles on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.