Gemini Video Understanding
einverne/dotfiles
Analyze videos using Google's Gemini API - describe content, answer questions, transcribe audio with visual descriptions, reference timestamps, clip videos, and process YouTube URLs.
Process and generate multimedia content using Google Gemini API.
$ npx skills add Microck/ordinary-claude-skills --skill ai-multimodal -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install Microck/ordinary-claude-skills ai-multimodal --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/Microck/ordinary-claude-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills_all/ai-multimodal .claude/skills/ai-multimodal && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "ai-multimodal" agent skill from https://github.com/Microck/ordinary-claude-skills/tree/main/skills_all/ai-multimodal into .claude/skills/ai-multimodal/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ai-multimodal", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/Microck/ordinary-claude-skills/tree/main/skills_all/ai-multimodalType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add Microck/ordinary-claude-skills --skill ai-multimodal -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install Microck/ordinary-claude-skills ai-multimodal --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Microck/ordinary-claude-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills_all/ai-multimodal .agents/skills/ai-multimodal && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "ai-multimodal" agent skill from https://github.com/Microck/ordinary-claude-skills/tree/main/skills_all/ai-multimodal into .agents/skills/ai-multimodal/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ai-multimodal", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Microck/ordinary-claude-skills --skill ai-multimodal -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install Microck/ordinary-claude-skills ai-multimodal --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Microck/ordinary-claude-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills_all/ai-multimodal .cursor/skills/ai-multimodal && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "ai-multimodal" agent skill from https://github.com/Microck/ordinary-claude-skills/tree/main/skills_all/ai-multimodal into .cursor/skills/ai-multimodal/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ai-multimodal", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/Microck/ordinary-claude-skills.git --path skills_all/ai-multimodal--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add Microck/ordinary-claude-skills --skill ai-multimodal -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install Microck/ordinary-claude-skills ai-multimodal --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Microck/ordinary-claude-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills_all/ai-multimodal .gemini/skills/ai-multimodal && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "ai-multimodal" agent skill from https://github.com/Microck/ordinary-claude-skills/tree/main/skills_all/ai-multimodal into .gemini/skills/ai-multimodal/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ai-multimodal", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install Microck/ordinary-claude-skills ai-multimodalInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add Microck/ordinary-claude-skills --skill ai-multimodal -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/Microck/ordinary-claude-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills_all/ai-multimodal .github/skills/ai-multimodal && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "ai-multimodal" agent skill from https://github.com/Microck/ordinary-claude-skills/tree/main/skills_all/ai-multimodal into .github/skills/ai-multimodal/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ai-multimodal", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Microck/ordinary-claude-skills --skill ai-multimodal -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install Microck/ordinary-claude-skills ai-multimodal --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Microck/ordinary-claude-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills_all/ai-multimodal .opencode/skills/ai-multimodal && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "ai-multimodal" agent skill from https://github.com/Microck/ordinary-claude-skills/tree/main/skills_all/ai-multimodal into .opencode/skills/ai-multimodal/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ai-multimodal", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
ai-multimodalProcess and generate multimedia content using Google Gemini API.
AI Multimodal is an agent skill from Microck/ordinary-claude-skills. Process and generate multimedia content using Google Gemini API. Capabilities include analyze audio files (transcription with timestamps, summarization, speech understanding, music/sound analysis up to 9.5 hours), understand images (captioning, object detection, OCR, visual Q&A, segmentation), process videos (scene detection, Q&A, temporal analysis, YouTube URLs, up to 6 hours), extract from documents (PDF tables, forms, charts, diagrams, multi-page), generate images (text-to-image, editing, composition…
Its SKILL.md is about 2.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 1 other file (for example `metadata.json`).
It sits in Media & Creative, covering Image generation, Computer vision and Context engineering. It works with Google Gemini and YouTube. The repository describes itself as: An unappealing collection of Claude Skills and resources. The licence is MIT.
5 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 1056d29. It shows what the files ask for, not the result of running them.
Pre-approves these tools, so the agent can use them without asking each time:
BashReadWriteEditFrom allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
pythonpipFrom the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
ai.google.devaistudio.google.comFrom URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
GEMINI_API_KEYFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
AI Multimodal loads about 2.7k tokens when it runs. Until then it costs about 208 tokens; SKILL.md has 875 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check noted patterns worth knowing about, such as sudo or a known installer.
2. Project root: `.env`3. `.claude/.env`4. `.claude/skills/.env`5. `.claude/skills/ai-multimodal/.env`allowed-tools: Bash, Read, Write, EditAutomated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from Microck/ordinary-claude-skills at commit 1056d29, republished under its MIT licence (© Microck). 875 words, ~2,663 tokens.
.claude/skills/ai-multimodal/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.Process audio, images, videos, documents, and generate images using Google Gemini's multimodal API. Unified interface for all multimedia content understanding and generation.
| Task | Audio | Image | Video | Document | Generation |
|---|---|---|---|---|---|
| Transcription | ✓ | - | ✓ | - | - |
| Summarization | ✓ | ✓ | ✓ | ✓ | - |
| Q&A | ✓ | ✓ | ✓ | ✓ | - |
| Object Detection | - | ✓ | ✓ | - | - |
| Text Extraction | - | ✓ | - | ✓ | - |
| Structured Output | ✓ | ✓ | ✓ | ✓ | - |
| Creation | TTS | - | - | - | ✓ |
| Timestamps | ✓ | - | ✓ | - | - |
| Segmentation | - | ✓ | - | - | - |
API Key Setup: Supports both Google AI Studio and Vertex AI.
The skill checks for GEMINI_API_KEY in this order:
export GEMINI_API_KEY="your-key".env.claude/.env.claude/skills/.env.claude/skills/ai-multimodal/.envGet API key: https://aistudio.google.com/apikey
For Vertex AI:
export GEMINI_USE_VERTEX=true
export VERTEX_PROJECT_ID=your-gcp-project-id
export VERTEX_LOCATION=us-central1 # OptionalInstall SDK:
pip install google-genai python-dotenv pillowTranscribe Audio:
python scripts/gemini_batch_process.py \
--files audio.mp3 \
--task transcribe \
--model gemini-2.5-flashAnalyze Image:
python scripts/gemini_batch_process.py \
--files image.jpg \
--task analyze \
--prompt "Describe this image" \
--output docs/assets/<output-name>.md \
--model gemini-2.5-flashProcess Video:
python scripts/gemini_batch_process.py \
--files video.mp4 \
--task analyze \
--prompt "Summarize key points with timestamps" \
--output docs/assets/<output-name>.md \
--model gemini-2.5-flashExtract from PDF:
python scripts/gemini_batch_process.py \
--files document.pdf \
--task extract \
--prompt "Extract table data as JSON" \
--output docs/assets/<output-name>.md \
--format jsonGenerate Image:
python scripts/gemini_batch_process.py \
--task generate \
--prompt "A futuristic city at sunset" \
--output docs/assets/<output-file-name> \
--model gemini-2.5-flash-image \
--aspect-ratio 16:9Optimize Media:
# Prepare large video for processing
python scripts/media_optimizer.py \
--input large-video.mp4 \
--output docs/assets/<output-file-name> \
--target-size 100MB
# Batch optimize multiple files
python scripts/media_optimizer.py \
--input-dir ./videos \
--output-dir docs/assets/optimized \
--quality 85Convert Documents to Markdown:
# Convert to PDF
python scripts/document_converter.py \
--input document.docx \
--output docs/assets/document.md
# Extract pages
python scripts/document_converter.py \
--input large.pdf \
--output docs/assets/chapter1.md \
--pages 1-20For detailed implementation guidance, see:
references/audio-processing.md - Transcription, analysis, TTSreferences/vision-understanding.md - Captioning, detection, OCRreferences/video-analysis.md - Scene detection, temporal understandingreferences/document-extraction.md - PDF processing, structured outputreferences/image-generation.md - Text-to-image, editingInput Pricing:
Token Rates:
TTS Pricing:
gemini-2.5-flash for most tasks (best price/performance)media_optimizer.py)Free Tier:
YouTube Limits:
Storage Limits:
Common errors and solutions:
All scripts support unified API key detection and error handling:
gemini_batch_process.py: Batch process multiple media files
media_optimizer.py: Prepare media for Gemini API
document_converter.py: Convert documents to PDF
Run any script with --help for detailed usage.
© Microck, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 1 other file in skills_all/ai-multimodal of Microck/ordinary-claude-skills.
Open the folder on GitHubat commit 1056d29
We found 15 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in Microck/ordinary-claude-skills, which our catalogue first saw on October 7, 2026.
AI Multimodal next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| AI Multimodal this skillMicrock/ordinary-claude-skills | 403 | 1 repos | ~2.7k | Automated safety check: Notes | MIT | |
| Gemini Video Understandingeinverne/dotfiles | 121 | 1 repos | ~2.6k | Automated safety check: Notes | MIT | |
| Nano Bananakkoppenhaver/cc-nano-banana | 378 | 1 repos | ~1.4k | Automated safety check: Pass | MIT | |
| Nano Bananajh941213/my-cc-harness | 126 | — | ~1k | Automated safety check: Pass | None | |
| Gemini Video Understandingbenchflow-ai/skillsbench | 1.8k | — | ~2.4k | Automated safety check: Pass | Apache-2.0 | |
| Nsfc RoadmapInternScience/DrClaw | 172 | — | ~2.5k | Automated safety check: Notes | None |
einverne/dotfiles
Analyze videos using Google's Gemini API - describe content, answer questions, transcribe audio with visual descriptions, reference timestamps, clip videos, and process YouTube URLs.
kkoppenhaver/cc-nano-banana
REQUIRED for all image generation requests. An agent skill from kkoppenhaver/cc-nano-banana.
jh941213/my-cc-harness
REQUIRED for all image generation requests. An agent skill from jh941213/my-cc-harness.
benchflow-ai/skillsbench
Analyze videos with Google Gemini API (summaries, Q&A, transcription with timestamps + visual context, scene/timeline detection, video clipping, FPS control, multi-video comparison, and YouTube URL…
InternScience/DrClaw
当用户明确要求"生成 NSFC 技术路线图/技术路线图绘制/roadmap/flowchart"或需要把标书研究内容转成"可打印、A4 可读"的技术路线图时使用。默认输出可编辑源文件(.drawio)与可嵌入文档的渲染结果(.svg/.png/.pdf);当用户主动提及 Nano Banana/Gemini 图片模型时,可切换为 PNG-only 模式。⚠️…
InternScience/DrClaw
当用户明确要求"生成 NSFC 原理图/机制图/schematic diagram/mechanism diagram"或需要把标书中的研究机制、算法架构、模块关系转成"可编辑 + 可嵌入文档"的图示时使用。默认输出可编辑源文件(.drawio)与渲染文件(.pdf/.svg/.png);当用户主动提及 Nano Banana/Gemini 图片模型时,可切换为 PNG-only 模式。⚠️…
Microck/ordinary-claude-skills
This skill should be used when reviewing or editing copy to ensure adherence to Every's style guide.
Microck/ordinary-claude-skills
Generate and edit images using the Gemini API (Nano Banana Pro).
Microck/ordinary-claude-skills
Write Ruby gems following Andrew Kane's proven patterns and philosophy.
Microck/ordinary-claude-skills
Execute Codex CLI for code analysis, refactoring, and automated code changes.
Microck/ordinary-claude-skills
Review documentation changes for compliance with the Metabase writing style guide.
Microck/ordinary-claude-skills
This skill should be used when managing the file-based todo tracking system in the todos/ directory.
Works with
Categories
Process and generate multimedia content using Google Gemini API. AI Multimodal is an agent skill from Microck/ordinary-claude-skills. Process and generate multimedia content using Google Gemini API.
AI Multimodal fits situations like: working with audio/video files; analyzing images; processing PDF documents; extracting structured data from media.
Run `npx skills add Microck/ordinary-claude-skills --skill ai-multimodal -a claude-code`. Or copy the skill folder (skills_all/ai-multimodal in Microck/ordinary-claude-skills) into .claude/skills/ai-multimodal in your project. Claude Code loads it when a task matches its description.
Run `npx skills add Microck/ordinary-claude-skills --skill ai-multimodal -a codex`. Or copy the skill folder (skills_all/ai-multimodal in Microck/ordinary-claude-skills) into .agents/skills/ai-multimodal in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Microck/ordinary-claude-skills --skill ai-multimodal -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ai-multimodal, .gemini/skills/ai-multimodal, .github/skills/ai-multimodal and .opencode/skills/ai-multimodal in your project.
Going by SKILL.md and its folder, AI Multimodal needs the command-line tools its instructions call (python and pip) and credentials named GEMINI_API_KEY. Our summary lists: Python 3; A credential in GEMINI_API_KEY. Its frontmatter pre-approves these tools: Bash, Read, Write, Edit.
SKILL.md names 2 domains. As links in the text: ai.google.dev and aistudio.google.com. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found notes only (mentions a .env file; pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.
AI Multimodal is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.7k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with AI Multimodal: Gemini Video Understanding (einverne/dotfiles, 121 stars), Nano Banana (kkoppenhaver/cc-nano-banana, 378 stars), Nano Banana (jh941213/my-cc-harness, 126 stars) and Gemini Video Understanding (benchflow-ai/skillsbench, 1.8k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
Microck (a GitHub user) maintains it in Microck/ordinary-claude-skills, which has 403 GitHub stars. The repository holds 91 skills in this directory. The repository was last updated on September 6, 2026.
Source: Microck/ordinary-claude-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.