Agent skill

Gemini Document Processing

by einverne in einverne/dotfiles

Guide for implementing Google Gemini API document processing - analyze PDFs with native vision to extract text, images, diagrams, charts, and tables.

GPL-3.0Auto-check: notesDocuments & Office

Install Gemini Document Processing

skills CLI
$ npx skills add einverne/dotfiles --skill gemini-document-processing -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install einverne/dotfiles gemini-document-processing --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/einverne/dotfiles.git skills-src && mkdir -p .claude/skills && cp -r skills-src/claude/skills/gemini-document-processing .claude/skills/gemini-document-processing && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
gemini-document-processing
GitHub stars
121
Token cost
~1.6k tokens
SKILL.md length
401 words
Files
9 (incl. scripts, references)
Skills in repo
39
Repo updated
First seen
Licence
GPL-3.0

At a glance

Guide for implementing Google Gemini API document processing - analyze PDFs with native vision to extract text, images, diagrams, charts, and tables.

  • Works in 7 steps: API Key Configuration → Install Dependencies → Extract Structured Data from PDF → …
  • Processing documents
  • SKILL.md covers Core Capabilities, When to Use This Skill, Quick Setup and Common Use Cases, plus 7 more sections
  • Runs Shell and Python scripts from its folder; calls python and pip; needs GEMINI_API_KEY

What it does

Gemini Document Processing is an agent skill from einverne/dotfiles. Guide for implementing Google Gemini API document processing - analyze PDFs with native vision to extract text, images, diagrams, charts, and tables. Use when processing documents, extracting structured data, summarizing PDFs, answering questions about document content, or converting documents to structured formats. (project)

Its SKILL.md is about 1.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 10 other files, including scripts and reference files (for example `CREATION_SUMMARY.md`, `README.md` and `references/code-examples.md`).

It sits in Documents & Office, covering Schema markup, Diagrams and PDF. It works with Google Gemini. The repository describes itself as: my personal dotfiles managed by dotbot, zinit. The licence is GPL-3.0.

When your agent uses it

  • Processing documents
  • Extracting structured data
  • Summarizing PDFs
  • Answering questions about document content

Example prompts

  • “/gemini-document-processing”

Requirements

  • Python 3
  • A Bash shell
  • A credential in GEMINI_API_KEY

Workflow steps

7 steps, taken from the step headings in SKILL.md.

  1. API Key Configuration
  2. Install Dependencies
  3. Extract Structured Data from PDF
  4. Summarize Long Document
  5. Answer Questions About Document
  6. Process with Python SDK
  7. Structured Output with JSON Schema

What it can do on your machine

Read from SKILL.md and the folder at commit c6c0686. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 2 files in scripts/ (Shell and Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python
    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • aistudio.google.com
    • ai.google.dev

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • GEMINI_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Gemini Document Processing loads about 1.6k tokens when it runs, and up to ~13k if it reads all its reference files. Until then it costs about 89 tokens; SKILL.md has 401 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~89
When it runs · the whole SKILL.md, loaded when a task matches
~1.6k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~13k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteMentions a .env fileSKILL.md:36
    2. `.env` file in skill directory (`.claude/skills/gemini-document-processing/.env`)
  • NoteMentions a .env fileSKILL.md:37
    3. `.env` file in project root
  • NoteMentions a .env fileSKILL.md:49
    cho "GEMINI_API_KEY=your-api-key-here" > .env
  • NoteMentions a .env fileSKILL.md:54
    cho "GEMINI_API_KEY=your-api-key-here" > .env

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from einverne/dotfiles at commit c6c0686, republished under its GPL-3.0 licence (© einverne). 401 words, ~1,627 tokens.

Download SKILL.mdSave it as .claude/skills/gemini-document-processing/SKILL.md (or your agent's skills folder). This skill also uses 8 other files; get the full folder from GitHub.
name
gemini-document-processing
description
Guide for implementing Google Gemini API document processing - analyze PDFs with native vision to extract text, images, diagrams, charts, and tables. Use when processing documents, extracting structured data, summarizing PDFs, answering questions about document content, or converting documents to structured formats. (project)

Gemini Document Processing

Process and analyze PDF documents using Google Gemini's native vision capabilities. Extract structured information, summarize content, answer questions, and understand complex documents with text, images, diagrams, charts, and tables.

Core Capabilities

  • PDF Vision Processing: Native understanding of PDFs up to 1,000 pages (258 tokens/page)
  • Multimodal Analysis: Process text, images, diagrams, charts, and tables
  • Structured Extraction: Output to JSON with schema validation
  • Document Q&A: Answer questions based on document content
  • Summarization: Generate summaries preserving context
  • Format Conversion: Transcribe to HTML while preserving layout

When to Use This Skill

Use this skill when you need to:

  • Extract structured data from PDF documents (invoices, resumes, forms)
  • Summarize long documents or reports
  • Answer questions about PDF content
  • Analyze documents with complex layouts, charts, or diagrams
  • Convert PDFs to structured formats (JSON, HTML)
  • Process multiple documents in batch
  • Build document processing pipelines

Quick Setup

1. API Key Configuration

The skill checks for GEMINI_API_KEY in this priority order:

  1. Process environment variable
  2. .env file in skill directory (.claude/skills/gemini-document-processing/.env)
  3. .env file in project root

Get your API key: https://aistudio.google.com/apikey

Option A: Environment Variable (Recommended)

bash
export GEMINI_API_KEY="your-api-key-here"

Option B: Skill Directory

bash
cd .claude/skills/gemini-document-processing
echo "GEMINI_API_KEY=your-api-key-here" > .env

Option C: Project Root

bash
echo "GEMINI_API_KEY=your-api-key-here" > .env
2. Install Dependencies
bash
pip install google-genai python-dotenv

Common Use Cases

1. Extract Structured Data from PDF
python
# Use the provided script
python .claude/skills/gemini-document-processing/scripts/process-document.py \
  --file invoice.pdf \
  --prompt "Extract invoice details as JSON" \
  --format json
2. Summarize Long Document
python
# Process and summarize
python .claude/skills/gemini-document-processing/scripts/process-document.py \
  --file report.pdf \
  --prompt "Provide a concise executive summary"
3. Answer Questions About Document
python
# Q&A on document content
python .claude/skills/gemini-document-processing/scripts/process-document.py \
  --file contract.pdf \
  --prompt "What are the key terms and conditions?"
4. Process with Python SDK
python
from google import genai

client = genai.Client()

# Read PDF
with open('document.pdf', 'rb') as f:
    pdf_data = f.read()

# Process document
response = client.models.generate_content(
    model='gemini-2.5-flash',
    contents=[
        'Extract key information from this document',
        genai.types.Part.from_bytes(
            data=pdf_data,
            mime_type='application/pdf'
        )
    ]
)

print(response.text)
5. Structured Output with JSON Schema
python
from google import genai
from pydantic import BaseModel

class InvoiceData(BaseModel):
    invoice_number: str
    date: str
    total: float
    vendor: str

client = genai.Client()

response = client.models.generate_content(
    model='gemini-2.5-flash',
    contents=[
        'Extract invoice details',
        genai.types.Part.from_bytes(
            data=open('invoice.pdf', 'rb').read(),
            mime_type='application/pdf'
        )
    ],
    config=genai.types.GenerateContentConfig(
        response_mime_type='application/json',
        response_schema=InvoiceData
    )
)

invoice_data = InvoiceData.model_validate_json(response.text)
Show full SKILL.md (179 more words)Show less

Key Constraints

  • Format: Only PDFs get vision processing (TXT, HTML, Markdown are text-only)
  • Size: < 20MB use inline encoding, > 20MB use File API
  • Pages: Max 1,000 pages per document
  • Storage: File API stores for 48 hours only
  • Cost: 258 tokens per page (fixed, regardless of content density)

Performance Tips

  1. Use Inline Encoding for PDFs < 20MB (simpler, single request)
  2. Use File API for larger files or repeated queries (enables context caching)
  3. Place Prompt After PDF for single-page documents
  4. Use Context Caching when querying same PDF multiple times
  5. Process in Parallel for multiple independent documents
  6. Use gemini-2.5-flash for best price/performance ratio

Decision Guide

PDF < 20MB?
├─ Yes → Use inline base64 encoding
└─ No  → Use File API

Need structured JSON output?
├─ Yes → Define response_schema with Pydantic
└─ No  → Get text response

Multiple queries on same PDF?
├─ Yes → Use File API + Context Caching
└─ No  → Inline encoding is sufficient

Script Reference

The skill includes a ready-to-use processing script:

bash
# Basic usage
python scripts/process-document.py --file document.pdf --prompt "Your prompt"

# With JSON output
python scripts/process-document.py --file document.pdf --prompt "Extract data" --format json

# With File API (for large files)
python scripts/process-document.py --file large-document.pdf --prompt "Summarize" --use-file-api

# Multiple prompts
python scripts/process-document.py --file document.pdf --prompt "Question 1" --prompt "Question 2"

References

For comprehensive documentation, see:

  • references/gemini-document-processing-report.md - Complete API reference
  • references/quick-reference.md - Quick lookup guide
  • references/code-examples.md - Additional code patterns

Troubleshooting

API Key Not Found:

bash
# Check API key is set
./scripts/check-api-key.sh

File Too Large:

  • Use File API for files > 20MB
  • Add --use-file-api flag to the script

Vision Not Working:

  • Ensure file is PDF format
  • Other formats (TXT, HTML) don't support vision processing

Support

© einverne, GPL-3.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 8 other files (scripts, references) in claude/skills/gemini-document-processing of einverne/dotfiles.

  • SKILL.md
  • .env.example
  • CREATION_SUMMARY.md
  • README.md
  • references/code-examples.md
  • references/gemini-document-processing-report.md
  • references/quick-reference.md
  • scripts/check-api-key.sh
  • scripts/process-document.py

Open the folder on GitHubat commit c6c0686

Compare with similar skills

Gemini Document Processing next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Gemini Document Processing compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Gemini Document Processing this skilleinverne/dotfiles121—~1.6kAutomated safety check: NotesGPL-3.0
AI MultimodalMicrock/ordinary-claude-skills4011 repos~2.7kAutomated safety check: NotesMIT
Nsfc RoadmapInternScience/DrClaw172—~2.5kAutomated safety check: NotesNone
Nsfc SchematicInternScience/DrClaw172—~3.7kAutomated safety check: NotesNone
Document ConverterBlackBeltTechnology/pi-agent-dashboard316—~999Automated safety check: PassMIT
Harness Book Best Practicewquguru/harness-books3.2k—~4.1kAutomated safety check: PassNone

Similar skills

  • AI Multimodal

    Microck/ordinary-claude-skills

    Process and generate multimedia content using Google Gemini API.

    401 GitHub starsUsed in 1 repo~2.7k tokens
    Media & CreativeAuto-check: notes
  • Nsfc Roadmap

    InternScience/DrClaw

    当用户明确要求"生成 NSFC 技术路线图/技术路线图绘制/roadmap/flowchart"或需要把标书研究内容转成"可打印、A4 可读"的技术路线图时使用。默认输出可编辑源文件(.drawio)与可嵌入文档的渲染结果(.svg/.png/.pdf);当用户主动提及 Nano Banana/Gemini 图片模型时,可切换为 PNG-only 模式。⚠️…

    172 GitHub stars~2.5k tokensUpdated 6 mo ago
    Documents & OfficeAuto-check: notes
  • Nsfc Schematic

    InternScience/DrClaw

    当用户明确要求"生成 NSFC 原理图/机制图/schematic diagram/mechanism diagram"或需要把标书中的研究机制、算法架构、模块关系转成"可编辑 + 可嵌入文档"的图示时使用。默认输出可编辑源文件(.drawio)与渲染文件(.pdf/.svg/.png);当用户主动提及 Nano Banana/Gemini 图片模型时,可切换为 PNG-only 模式。⚠️…

    172 GitHub stars~3.7k tokensUpdated 6 mo ago
    Documents & OfficeAuto-check: notes
  • Document Converter

    BlackBeltTechnology/pi-agent-dashboard

    Convert documents bidirectionally via the pi-doc-engine facade: ingest PDF/DOCX/PPTX/XLSX to provenance-stamped Markdown (with OCR), and produce templated DOCX/PDF from Markdown with diagrams, TOC…

    316 GitHub stars~999 tokensUpdated yesterday
    Documents & OfficeAuto-check passed
  • Harness Book Best Practice

    wquguru/harness-books

    Best practices for working on the Harness books repo. An agent skill from wquguru/harness-books.

    3.2k GitHub stars~4.1k tokensUpdated 5 mo ago
    Documents & OfficeAuto-check passed
  • Markitdown

    jimmc414/Kosmos

    Convert various file formats (PDF, Office documents, images, audio, web content, structured data) to Markdown optimized for LLM processing.

    594 GitHub starsUsed in 2 repos~1.7k tokens
    Documents & OfficeAuto-check passed

More from einverne/dotfiles

All 39 skills in this repo
  • DOCX

    einverne/dotfiles

    Comprehensive document creation, editing, and analysis with support for tracked changes, comments, formatting preservation, and text extraction.

    121 GitHub starsUsed in 34 repos~2.5k tokens
    Auto-check: notes
  • Chrome Devtools

    einverne/dotfiles

    Browser automation, debugging, and performance analysis using Puppeteer CLI scripts.

    121 GitHub starsUsed in 1 repo~1.6k tokens
    Auto-check: notes
  • PDF

    einverne/dotfiles

    Comprehensive PDF manipulation toolkit for extracting text and tables, creating new PDFs, merging/splitting documents, and handling forms.

    121 GitHub starsUsed in 46 repos~1.8k tokens
    Auto-check passed
  • Gemini Audio

    einverne/dotfiles

    Guide for implementing Google Gemini API audio capabilities - analyze audio with transcription, summarization, and understanding (up to 9.5 hours), plus generate speech with controllable TTS.

    121 GitHub starsUsed in 1 repo~2k tokens
    Auto-check: notes
  • Gemini Image Gen

    einverne/dotfiles

    Guide for implementing Google Gemini API image generation - create high-quality images from text prompts using gemini-2.5-flash-image model.

    121 GitHub starsUsed in 1 repo~1.8k tokens
    Auto-check: notes
  • Gemini Vision

    einverne/dotfiles

    Guide for implementing Google Gemini API image understanding - analyze images with captioning, classification, visual QA, object detection, segmentation, and multi-image comparison.

    121 GitHub starsUsed in 1 repo~1.6k tokens
    Auto-check: notes

Works with

Questions about Gemini Document Processing

What does Gemini Document Processing do?

Guide for implementing Google Gemini API document processing - analyze PDFs with native vision to extract text, images, diagrams, charts, and tables. Gemini Document Processing is an agent skill from einverne/dotfiles. Guide for implementing Google Gemini API document processing - analyze PDFs with native vision to extract text, images, diagrams, charts, and tables.

When should I use Gemini Document Processing?

Gemini Document Processing fits situations like: processing documents; extracting structured data; summarizing PDFs; answering questions about document content.

How do I install Gemini Document Processing in Claude Code?

Run `npx skills add einverne/dotfiles --skill gemini-document-processing -a claude-code`. Or copy the skill folder (claude/skills/gemini-document-processing in einverne/dotfiles) into .claude/skills/gemini-document-processing in your project. Claude Code loads it when a task matches its description.

How do I install Gemini Document Processing in Codex?

Run `npx skills add einverne/dotfiles --skill gemini-document-processing -a codex`. Or copy the skill folder (claude/skills/gemini-document-processing in einverne/dotfiles) into .agents/skills/gemini-document-processing in your project. Codex loads it when a task matches its description.

Can I use Gemini Document Processing in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add einverne/dotfiles --skill gemini-document-processing -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/gemini-document-processing, .gemini/skills/gemini-document-processing, .github/skills/gemini-document-processing and .opencode/skills/gemini-document-processing in your project.

What does Gemini Document Processing need to run?

Going by SKILL.md and its folder, Gemini Document Processing needs a shell and Python for the scripts in its folder, the command-line tools its instructions call (python and pip) and credentials named GEMINI_API_KEY. Our summary lists: Python 3; A Bash shell; A credential in GEMINI_API_KEY.

Does Gemini Document Processing access the network?

SKILL.md names 2 domains. As links in the text: aistudio.google.com and ai.google.dev. This is read from the text; nothing was executed.

Is Gemini Document Processing safe to install?

Our automated static check of SKILL.md found notes only (mentions a .env file), nothing it rates as a warning. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Gemini Document Processing use?

Gemini Document Processing is published under the GPL-3.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Gemini Document Processing use?

About 1.6k tokens (SKILL.md is roughly 6.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 11k tokens, read only when the agent opens those files.

What are the alternatives to Gemini Document Processing?

Skills that share tags, products or a category with Gemini Document Processing: AI Multimodal (Microck/ordinary-claude-skills, 401 stars), Nsfc Roadmap (InternScience/DrClaw, 172 stars), Nsfc Schematic (InternScience/DrClaw, 172 stars) and Document Converter (BlackBeltTechnology/pi-agent-dashboard, 316 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Gemini Document Processing?

einverne (a GitHub user) maintains it in einverne/dotfiles, which has 121 GitHub stars. The repository holds 39 skills in this directory. The repository was last updated on September 9, 2026.

Source: einverne/dotfiles on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.