Agent skill

AI Multimodal

by Microck in Microck/ordinary-claude-skills

Process and generate multimedia content using Google Gemini API.

MITAuto-check: notesMedia & Creative

Install AI Multimodal

skills CLI
$ npx skills add Microck/ordinary-claude-skills --skill ai-multimodal -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Microck/ordinary-claude-skills ai-multimodal --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Microck/ordinary-claude-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills_all/ai-multimodal .claude/skills/ai-multimodal && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
ai-multimodal
GitHub stars
403
Used in
1 other repo
Token cost
~2.7k tokens
SKILL.md length
875 words
Files
2
Skills in repo
91
Repo updated
First seen
Licence
MIT

At a glance

Process and generate multimedia content using Google Gemini API.

  • Works in 5 steps: Process environment: export… → Project root: .env → .claude/.env → …
  • Working with audio/video files
  • SKILL.md covers Core Capabilities, Capability Matrix, Model Selection Guide and Quick Start, plus 3 more sections
  • Calls python and pip; needs GEMINI_API_KEY

What it does

AI Multimodal is an agent skill from Microck/ordinary-claude-skills. Process and generate multimedia content using Google Gemini API. Capabilities include analyze audio files (transcription with timestamps, summarization, speech understanding, music/sound analysis up to 9.5 hours), understand images (captioning, object detection, OCR, visual Q&A, segmentation), process videos (scene detection, Q&A, temporal analysis, YouTube URLs, up to 6 hours), extract from documents (PDF tables, forms, charts, diagrams, multi-page), generate images (text-to-image, editing, composition…

Its SKILL.md is about 2.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 1 other file (for example `metadata.json`).

It sits in Media & Creative, covering Image generation, Computer vision and Context engineering. It works with Google Gemini and YouTube. The repository describes itself as: An unappealing collection of Claude Skills and resources. The licence is MIT.

When your agent uses it

  • Working with audio/video files
  • Analyzing images
  • Processing PDF documents
  • Extracting structured data from media

Example prompts

  • “/ai-multimodal”

Requirements

  • Python 3
  • A credential in GEMINI_API_KEY
  • Pre-approved tools (allowed-tools): Bash, Read, Write, Edit

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Process environment: export GEMINI_API_KEY="your-key"
  2. Project root: .env
  3. .claude/.env
  4. .claude/skills/.env
  5. .claude/skills/ai-multimodal/.env

What it can do on your machine

Read from SKILL.md and the folder at commit 1056d29. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Bash
    • Read
    • Write
    • Edit

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python
    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • ai.google.dev
    • aistudio.google.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • GEMINI_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

AI Multimodal loads about 2.7k tokens when it runs. Until then it costs about 208 tokens; SKILL.md has 875 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~208
When it runs · the whole SKILL.md, loaded when a task matches
~2.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteMentions a .env fileSKILL.md:104
    2. Project root: `.env`
  • NoteMentions a .env fileSKILL.md:105
    3. `.claude/.env`
  • NoteMentions a .env fileSKILL.md:106
    4. `.claude/skills/.env`
  • NoteMentions a .env fileSKILL.md:107
    5. `.claude/skills/ai-multimodal/.env`
  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Bash, Read, Write, Edit

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Microck/ordinary-claude-skills at commit 1056d29, republished under its MIT licence (© Microck). 875 words, ~2,663 tokens.

Download SKILL.mdSave it as .claude/skills/ai-multimodal/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
ai-multimodal
description
Process and generate multimedia content using Google Gemini API. Capabilities include analyze audio files (transcription with timestamps, summarization, speech understanding, music/sound analysis up to 9.5 hours), understand images (captioning, object detection, OCR, visual Q&A, segmentation), process videos (scene detection, Q&A, temporal analysis, YouTube URLs, up to 6 hours), extract from documents (PDF tables, forms, charts, diagrams, multi-page), generate images (text-to-image, editing, composition, refinement). Use when working with audio/video files, analyzing images or screenshots, processing PDF documents, extracting structured data from media, creating images from text prompts, or implementing multimodal AI features. Supports multiple models (Gemini 2.5/2.0) with context windows up to 2M tokens.
allowed-tools
Bash, Read, Write, Edit
license
MIT

AI Multimodal Processing Skill

Process audio, images, videos, documents, and generate images using Google Gemini's multimodal API. Unified interface for all multimedia content understanding and generation.

Core Capabilities

Audio Processing
  • Transcription with timestamps (up to 9.5 hours)
  • Audio summarization and analysis
  • Speech understanding and speaker identification
  • Music and environmental sound analysis
  • Text-to-speech generation with controllable voice
Image Understanding
  • Image captioning and description
  • Object detection with bounding boxes (2.0+)
  • Pixel-level segmentation (2.5+)
  • Visual question answering
  • Multi-image comparison (up to 3,600 images)
  • OCR and text extraction
Video Analysis
  • Scene detection and summarization
  • Video Q&A with temporal understanding
  • Transcription with visual descriptions
  • YouTube URL support
  • Long video processing (up to 6 hours)
  • Frame-level analysis
Document Extraction
  • Native PDF vision processing (up to 1,000 pages)
  • Table and form extraction
  • Chart and diagram analysis
  • Multi-page document understanding
  • Structured data output (JSON schema)
  • Format conversion (PDF to HTML/JSON)
Image Generation
  • Text-to-image generation
  • Image editing and modification
  • Multi-image composition (up to 3 images)
  • Iterative refinement
  • Multiple aspect ratios (1:1, 16:9, 9:16, 4:3, 3:4)
  • Controllable style and quality

Capability Matrix

TaskAudioImageVideoDocumentGeneration
Transcription✓-✓--
Summarization✓✓✓✓-
Q&A✓✓✓✓-
Object Detection-✓✓--
Text Extraction-✓-✓-
Structured Output✓✓✓✓-
CreationTTS---✓
Timestamps✓-✓--
Segmentation-✓---

Model Selection Guide

  • gemini-2.5-pro: Highest quality, all features, 1M-2M context
  • gemini-2.5-flash: Best balance, all features, 1M-2M context
  • gemini-2.5-flash-lite: Lightweight, segmentation support
  • gemini-2.5-flash-image: Image generation only
Gemini 2.0 Series
  • gemini-2.0-flash: Fast processing, object detection
  • gemini-2.0-flash-lite: Lightweight option
Feature Requirements
  • Segmentation: Requires 2.5+ models
  • Object Detection: Requires 2.0+ models
  • Multi-video: Requires 2.5+ models
  • Image Generation: Requires flash-image model
Context Windows
  • 2M tokens: ~6 hours video (low-res) or ~2 hours (default)
  • 1M tokens: ~3 hours video (low-res) or ~1 hour (default)
  • Audio: 32 tokens/second (1 min = 1,920 tokens)
  • PDF: 258 tokens/page (fixed)
  • Image: 258-1,548 tokens based on size

Quick Start

Prerequisites

API Key Setup: Supports both Google AI Studio and Vertex AI.

The skill checks for GEMINI_API_KEY in this order:

  1. Process environment: export GEMINI_API_KEY="your-key"
  2. Project root: .env
  3. .claude/.env
  4. .claude/skills/.env
  5. .claude/skills/ai-multimodal/.env

Get API key: https://aistudio.google.com/apikey

For Vertex AI:

bash
export GEMINI_USE_VERTEX=true
export VERTEX_PROJECT_ID=your-gcp-project-id
export VERTEX_LOCATION=us-central1  # Optional

Install SDK:

bash
pip install google-genai python-dotenv pillow
Common Patterns

Transcribe Audio:

bash
python scripts/gemini_batch_process.py \
  --files audio.mp3 \
  --task transcribe \
  --model gemini-2.5-flash

Analyze Image:

bash
python scripts/gemini_batch_process.py \
  --files image.jpg \
  --task analyze \
  --prompt "Describe this image" \
  --output docs/assets/<output-name>.md \
  --model gemini-2.5-flash

Process Video:

bash
python scripts/gemini_batch_process.py \
  --files video.mp4 \
  --task analyze \
  --prompt "Summarize key points with timestamps" \
  --output docs/assets/<output-name>.md \
  --model gemini-2.5-flash

Extract from PDF:

bash
python scripts/gemini_batch_process.py \
  --files document.pdf \
  --task extract \
  --prompt "Extract table data as JSON" \
  --output docs/assets/<output-name>.md \
  --format json

Generate Image:

bash
python scripts/gemini_batch_process.py \
  --task generate \
  --prompt "A futuristic city at sunset" \
  --output docs/assets/<output-file-name> \
  --model gemini-2.5-flash-image \
  --aspect-ratio 16:9

Optimize Media:

bash
# Prepare large video for processing
python scripts/media_optimizer.py \
  --input large-video.mp4 \
  --output docs/assets/<output-file-name> \
  --target-size 100MB

# Batch optimize multiple files
python scripts/media_optimizer.py \
  --input-dir ./videos \
  --output-dir docs/assets/optimized \
  --quality 85

Convert Documents to Markdown:

bash
# Convert to PDF
python scripts/document_converter.py \
  --input document.docx \
  --output docs/assets/document.md

# Extract pages
python scripts/document_converter.py \
  --input large.pdf \
  --output docs/assets/chapter1.md \
  --pages 1-20

Supported Formats

Audio
  • WAV, MP3, AAC, FLAC, OGG Vorbis, AIFF
  • Max 9.5 hours per request
  • Auto-downsampled to 16 Kbps mono
Images
  • PNG, JPEG, WEBP, HEIC, HEIF
  • Max 3,600 images per request
  • Resolution: ≤384px = 258 tokens, larger = tiled
Video
  • MP4, MPEG, MOV, AVI, FLV, MPG, WebM, WMV, 3GPP
  • Max 6 hours (low-res) or 2 hours (default)
  • YouTube URLs supported (public only)
Documents
  • PDF only for vision processing
  • Max 1,000 pages
  • TXT, HTML, Markdown supported (text-only)
Size Limits
  • Inline: <20MB total request
  • File API: 2GB per file, 20GB project quota
  • Retention: 48 hours auto-delete

Reference Navigation

For detailed implementation guidance, see:

Audio Processing
  • references/audio-processing.md - Transcription, analysis, TTS
    • Timestamp handling and segment analysis
    • Multi-speaker identification
    • Non-speech audio analysis
    • Text-to-speech generation
Image Understanding
  • references/vision-understanding.md - Captioning, detection, OCR
    • Object detection and localization
    • Pixel-level segmentation
    • Visual question answering
    • Multi-image comparison
Video Analysis
  • references/video-analysis.md - Scene detection, temporal understanding
    • YouTube URL processing
    • Timestamp-based queries
    • Video clipping and FPS control
    • Long video optimization
Document Extraction
  • references/document-extraction.md - PDF processing, structured output
    • Table and form extraction
    • Chart and diagram analysis
    • JSON schema validation
    • Multi-page handling
Show full SKILL.md (346 more words)Show less
Image Generation
  • references/image-generation.md - Text-to-image, editing
    • Prompt engineering strategies
    • Image editing and composition
    • Aspect ratio selection
    • Safety settings

Cost Optimization

Token Costs

Input Pricing:

  • Gemini 2.5 Flash: $1.00/1M input, $0.10/1M output
  • Gemini 2.5 Pro: $3.00/1M input, $12.00/1M output
  • Gemini 1.5 Flash: $0.70/1M input, $0.175/1M output

Token Rates:

  • Audio: 32 tokens/second (1 min = 1,920 tokens)
  • Video: ~300 tokens/second (default) or ~100 (low-res)
  • PDF: 258 tokens/page (fixed)
  • Image: 258-1,548 tokens based on size

TTS Pricing:

  • Flash TTS: $10/1M tokens
  • Pro TTS: $20/1M tokens
Best Practices
  1. Use gemini-2.5-flash for most tasks (best price/performance)
  2. Use File API for files >20MB or repeated queries
  3. Optimize media before upload (see media_optimizer.py)
  4. Process specific segments instead of full videos
  5. Use lower FPS for static content
  6. Implement context caching for repeated queries
  7. Batch process multiple files in parallel

Rate Limits

Free Tier:

  • 10-15 RPM (requests per minute)
  • 1M-4M TPM (tokens per minute)
  • 1,500 RPD (requests per day)

YouTube Limits:

  • Free tier: 8 hours/day
  • Paid tier: No length limits
  • Public videos only

Storage Limits:

  • 20GB per project
  • 2GB per file
  • 48-hour retention

Error Handling

Common errors and solutions:

  • 400: Invalid format/size - validate before upload
  • 401: Invalid API key - check configuration
  • 403: Permission denied - verify API key restrictions
  • 404: File not found - ensure file uploaded and active
  • 429: Rate limit exceeded - implement exponential backoff
  • 500: Server error - retry with backoff

Scripts Overview

All scripts support unified API key detection and error handling:

gemini_batch_process.py: Batch process multiple media files

  • Supports all modalities (audio, image, video, PDF)
  • Progress tracking and error recovery
  • Output formats: JSON, Markdown, CSV
  • Rate limiting and retry logic
  • Dry-run mode

media_optimizer.py: Prepare media for Gemini API

  • Compress videos/audio for size limits
  • Resize images appropriately
  • Split long videos into chunks
  • Format conversion
  • Quality vs size optimization

document_converter.py: Convert documents to PDF

  • Convert DOCX, XLSX, PPTX to PDF
  • Extract page ranges
  • Optimize PDFs for Gemini
  • Extract images from PDFs
  • Batch conversion support

Run any script with --help for detailed usage.

Resources

© Microck, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in skills_all/ai-multimodal of Microck/ordinary-claude-skills.

  • SKILL.md
  • metadata.json

Open the folder on GitHubat commit 1056d29

Used in 1 other repository

We found 15 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in Microck/ordinary-claude-skills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

AI Multimodal next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

AI Multimodal compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
AI Multimodal this skillMicrock/ordinary-claude-skills4031 repos~2.7kAutomated safety check: NotesMIT
Gemini Video Understandingeinverne/dotfiles1211 repos~2.6kAutomated safety check: NotesMIT
Nano Bananakkoppenhaver/cc-nano-banana3781 repos~1.4kAutomated safety check: PassMIT
Nano Bananajh941213/my-cc-harness126—~1kAutomated safety check: PassNone
Gemini Video Understandingbenchflow-ai/skillsbench1.8k—~2.4kAutomated safety check: PassApache-2.0
Nsfc RoadmapInternScience/DrClaw172—~2.5kAutomated safety check: NotesNone

Similar skills

  • Analyze videos using Google's Gemini API - describe content, answer questions, transcribe audio with visual descriptions, reference timestamps, clip videos, and process YouTube URLs.

    121 GitHub starsUsed in 1 repo~2.6k tokens
    AI & LLM EngineeringAuto-check: notes
  • Nano Banana

    kkoppenhaver/cc-nano-banana

    REQUIRED for all image generation requests. An agent skill from kkoppenhaver/cc-nano-banana.

    378 GitHub starsUsed in 1 repo~1.4k tokens
    Media & CreativeAuto-check passed
  • Nano Banana

    jh941213/my-cc-harness

    REQUIRED for all image generation requests. An agent skill from jh941213/my-cc-harness.

    126 GitHub stars~1k tokensUpdated 2 mo ago
    Media & CreativeAuto-check passed
  • Gemini Video Understanding

    benchflow-ai/skillsbench

    Analyze videos with Google Gemini API (summaries, Q&A, transcription with timestamps + visual context, scene/timeline detection, video clipping, FPS control, multi-video comparison, and YouTube URL…

    1.8k GitHub stars~2.4k tokensUpdated 2 mo ago
    Media & CreativeAuto-check passed
  • Nsfc Roadmap

    InternScience/DrClaw

    当用户明确要求"生成 NSFC 技术路线图/技术路线图绘制/roadmap/flowchart"或需要把标书研究内容转成"可打印、A4 可读"的技术路线图时使用。默认输出可编辑源文件(.drawio)与可嵌入文档的渲染结果(.svg/.png/.pdf);当用户主动提及 Nano Banana/Gemini 图片模型时,可切换为 PNG-only 模式。⚠️…

    172 GitHub stars~2.5k tokensUpdated 6 mo ago
    Documents & OfficeAuto-check: notes
  • Nsfc Schematic

    InternScience/DrClaw

    当用户明确要求"生成 NSFC 原理图/机制图/schematic diagram/mechanism diagram"或需要把标书中的研究机制、算法架构、模块关系转成"可编辑 + 可嵌入文档"的图示时使用。默认输出可编辑源文件(.drawio)与渲染文件(.pdf/.svg/.png);当用户主动提及 Nano Banana/Gemini 图片模型时,可切换为 PNG-only 模式。⚠️…

    172 GitHub stars~3.7k tokensUpdated 6 mo ago
    Documents & OfficeAuto-check: notes

More from Microck/ordinary-claude-skills

All 91 skills in this repo
  • Every Style Editor

    Microck/ordinary-claude-skills

    This skill should be used when reviewing or editing copy to ensure adherence to Every's style guide.

    403 GitHub starsUsed in 3 repos~1.2k tokens
    Auto-check passed
  • Gemini Imagegen

    Microck/ordinary-claude-skills

    Generate and edit images using the Gemini API (Nano Banana Pro).

    403 GitHub starsUsed in 3 repos~1.6k tokens
    Auto-check passed
  • Andrew Kane Gem Writer

    Microck/ordinary-claude-skills

    Write Ruby gems following Andrew Kane's proven patterns and philosophy.

    403 GitHub starsUsed in 2 repos~1.2k tokens
    Auto-check passed
  • Codex

    Microck/ordinary-claude-skills

    Execute Codex CLI for code analysis, refactoring, and automated code changes.

    403 GitHub starsUsed in 2 repos~2.8k tokens
    Auto-check passed
  • Docs Review

    Microck/ordinary-claude-skills

    Review documentation changes for compliance with the Metabase writing style guide.

    403 GitHub starsUsed in 2 repos~1.8k tokens
    Auto-check: notes
  • File Todos

    Microck/ordinary-claude-skills

    This skill should be used when managing the file-based todo tracking system in the todos/ directory.

    403 GitHub starsUsed in 2 repos~1.9k tokens
    Auto-check passed

Questions about AI Multimodal

What does AI Multimodal do?

Process and generate multimedia content using Google Gemini API. AI Multimodal is an agent skill from Microck/ordinary-claude-skills. Process and generate multimedia content using Google Gemini API.

When should I use AI Multimodal?

AI Multimodal fits situations like: working with audio/video files; analyzing images; processing PDF documents; extracting structured data from media.

How do I install AI Multimodal in Claude Code?

Run `npx skills add Microck/ordinary-claude-skills --skill ai-multimodal -a claude-code`. Or copy the skill folder (skills_all/ai-multimodal in Microck/ordinary-claude-skills) into .claude/skills/ai-multimodal in your project. Claude Code loads it when a task matches its description.

How do I install AI Multimodal in Codex?

Run `npx skills add Microck/ordinary-claude-skills --skill ai-multimodal -a codex`. Or copy the skill folder (skills_all/ai-multimodal in Microck/ordinary-claude-skills) into .agents/skills/ai-multimodal in your project. Codex loads it when a task matches its description.

Can I use AI Multimodal in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Microck/ordinary-claude-skills --skill ai-multimodal -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ai-multimodal, .gemini/skills/ai-multimodal, .github/skills/ai-multimodal and .opencode/skills/ai-multimodal in your project.

What does AI Multimodal need to run?

Going by SKILL.md and its folder, AI Multimodal needs the command-line tools its instructions call (python and pip) and credentials named GEMINI_API_KEY. Our summary lists: Python 3; A credential in GEMINI_API_KEY. Its frontmatter pre-approves these tools: Bash, Read, Write, Edit.

Does AI Multimodal access the network?

SKILL.md names 2 domains. As links in the text: ai.google.dev and aistudio.google.com. This is read from the text; nothing was executed.

Is AI Multimodal safe to install?

Our automated static check of SKILL.md found notes only (mentions a .env file; pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does AI Multimodal use?

AI Multimodal is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does AI Multimodal use?

About 2.7k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to AI Multimodal?

Skills that share tags, products or a category with AI Multimodal: Gemini Video Understanding (einverne/dotfiles, 121 stars), Nano Banana (kkoppenhaver/cc-nano-banana, 378 stars), Nano Banana (jh941213/my-cc-harness, 126 stars) and Gemini Video Understanding (benchflow-ai/skillsbench, 1.8k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains AI Multimodal?

Microck (a GitHub user) maintains it in Microck/ordinary-claude-skills, which has 403 GitHub stars. The repository holds 91 skills in this directory. The repository was last updated on September 6, 2026.

Source: Microck/ordinary-claude-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.