Agent skill

Multimodal LLM

by yonatangross in yonatangross/orchestkit

Vision, audio, video generation, and multimodal LLM integration patterns.

MITAuto-check passedMedia & Creative

Install Multimodal LLM

skills CLI
$ npx skills add yonatangross/orchestkit --skill multimodal-llm -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install yonatangross/orchestkit multimodal-llm --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/yonatangross/orchestkit.git skills-src && mkdir -p .claude/skills && cp -r skills-src/src/skills/multimodal-llm .claude/skills/multimodal-llm && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
multimodal-llm
GitHub stars
292
Token cost
~2.2k tokens
SKILL.md length
806 words
Files
13
Skills in repo
108
Repo updated
First seen
Licence
MIT

At a glance

Vision, audio, video generation, and multimodal LLM integration patterns.

  • Works in 10 steps: Not setting max_tokens on vision… → Sending oversized images without… → Using high detail level for simple… → …
  • Processing images
  • SKILL.md covers Quick Reference, Vision: Image Analysis, Vision: Document Understanding and Vision: Model Selection, plus 10 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Multimodal LLM is an agent skill from yonatangross/orchestkit. Vision, audio, video generation, and multimodal LLM integration patterns. Use when processing images, transcribing audio, generating speech, generating AI video (Kling v3, Sora 2, Veo 3.1 std/lite/fast, Runway Gen-4.5 via gen4turbo), or building multimodal AI pipelines.

Its SKILL.md is about 2.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 13 other files (for example `rules/_sections.md`, `rules/_template.md` and `rules/audio-models.md`). Compatibility notes: Claude Code 2.1.277+.

It sits in Media & Creative, covering AI video generation. It works with Google Veo. The repository describes itself as: The Complete AI Development Toolkit for Claude Code. 106 skills, 36 agents, 171 hooks. Install ork for stable (v9.x), or ork-alpha for the v10 line, which ships daily. The licence is MIT.

When your agent uses it

  • Processing images
  • Transcribing audio
  • Generating speech
  • Generating AI video (Kling v3

Example prompts

  • “/multimodal-llm”

Requirements

  • Python 3
  • Compatibility (from SKILL.md): Claude Code 2.1.277+.
  • Pre-approved tools (allowed-tools): Read, Glob, Grep, WebFetch, WebSearch

Workflow steps

10 steps, taken from the first numbered list in SKILL.md.

  1. Not setting max_tokens on vision requests (responses truncated)
  2. Sending oversized images without resizing (>2048px)
  3. Using high detail level for simple yes/no classification
  4. Using STT+LLM+TTS pipeline instead of native speech-to-speech
  5. Not leveraging barge-in support for natural voice conversations
  6. Using deprecated models (GPT-4V, Whisper-1)
  7. Ignoring rate limits on vision and audio endpoints
  8. Calling video generation APIs synchronously (they're async — poll or use callbacks)
  9. Generating separate clips without character elements (characters look different each time)
  10. Using Sora for high-volume social content (expensive, slow — use Kling Standard instead)

What it can do on your machine

Read from SKILL.md and the folder at commit e4ff8d9. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Glob
    • Grep
    • WebFetch
    • WebSearch

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Claude Code 2.1.277+.

    From compatibility in the SKILL.md frontmatter.

Context cost

Multimodal LLM loads about 2.2k tokens when it runs. Until then it costs about 72 tokens; SKILL.md has 806 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~72
When it runs · the whole SKILL.md, loaded when a task matches
~2.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from yonatangross/orchestkit at commit e4ff8d9, republished under its MIT licence (© yonatangross). 806 words, ~2,246 tokens.

Download SKILL.mdSave it as .claude/skills/multimodal-llm/SKILL.md (or your agent's skills folder). This skill also uses 12 other files; get the full folder from GitHub.
name
multimodal-llm
description
Vision, audio, video generation, and multimodal LLM integration patterns. Use when processing images, transcribing audio, generating speech, generating AI video (Kling v3, Sora 2, Veo 3.1 std/lite/fast, Runway Gen-4.5 via `gen4_turbo`), or building multimodal AI pipelines.
allowed-tools
Read, Glob, Grep, WebFetch, WebSearch
compatibility
Claude Code 2.1.277+.
license
MIT
user-invocable
false
disable-model-invocation
true
effort
high
metadata.owner-agent
multimodal-specialist
metadata.category
mcp-enhancement
metadata.version
2.1.1
metadata.author
OrchestKit
metadata.complexity
high
metadata.tags
vision, audio, video, multimodal, image, speech, transcription, tts, kling, sora, veo, video-generation

Multimodal LLM Patterns

Integrate vision, audio, and video generation capabilities from leading multimodal models. Covers image analysis, document understanding, real-time voice agents, speech-to-text, text-to-speech, and AI video generation (Kling v3, Sora 2, Veo 3.1 std/lite/fast tiers, Runway Gen-4.5 via gen4_turbo).

Canonical model IDs (Google IDs checked against ai.google.dev/gemini-api/docs/models on 2026-09-23):

ProviderModel IDs
Anthropicclaude-opus-5-5 (recommended, the default Opus since CC 2.1.280, same 2,576 px budget as Opus 5), claude-opus-5, claude-opus-4-8, claude-opus-4-7, claude-opus-4-6, claude-sonnet-5-5, claude-sonnet-4-6, claude-haiku-4-5-20251001. claude-fable-5 is Anthropic's frontier tier above Opus (GA 2026-07). Premium cost, never auto-pin it; the fable-spend-consent gate requires explicit user consent before any Fable spend
OpenAIgpt-5.5 (current flagship)
Googlegemini-3.1-pro-preview (flagship), gemini-3.1-flash-lite (cost)
Veoveo-3.1-generate-preview / veo-3.1-lite-generate-preview / veo-3.1-fast-generate-preview
Klingkling-v3 (model_name field in Kling API)
Runwaygen4_turbo (product label: Gen-4.5)

Quick Reference

CategoryRulesImpactWhen to Use
Vision: Image Analysis1HIGHImage captioning, VQA, multi-image comparison, object detection
Vision: Document Understanding1HIGHOCR, chart/diagram analysis, PDF processing, table extraction
Vision: Model Selection1MEDIUMChoosing provider, cost optimization, image size limits
Audio: Speech-to-Text1HIGHTranscription, speaker diarization, long-form audio
Audio: Text-to-Speech1MEDIUMVoice synthesis, expressive TTS, multi-speaker dialogue
Audio: Model Selection1MEDIUMReal-time voice agents, provider comparison, pricing
Video: Model Selection1HIGHChoosing video gen provider (Kling, Sora, Veo, Runway)
Video: API Patterns1HIGHAsync task polling, SDK integration, webhook callbacks
Video: Multi-Shot1HIGHStoryboarding, character elements, scene consistency

Total: 9 rules across 3 categories (Vision, Audio, Video Generation)

Vision: Image Analysis

Send images to multimodal LLMs for captioning, visual QA, and object detection. Always set max_tokens and resize images before encoding.

RuleFileKey Pattern
Image Analysisrules/vision-image-analysis.mdBase64 encoding, multi-image, bounding boxes

Vision: Document Understanding

Extract structured data from documents, charts, and PDFs using vision models.

RuleFileKey Pattern
Document Visionrules/vision-document.mdPDF page ranges, detail levels, OCR strategies

Vision: Model Selection

Choose the right vision provider based on accuracy, cost, and context window needs.

RuleFileKey Pattern
Vision Modelsrules/vision-models.mdProvider comparison, token costs, image limits

Audio: Speech-to-Text

Convert audio to text with speaker diarization, timestamps, and sentiment analysis.

RuleFileKey Pattern
Speech-to-Textrules/audio-speech-to-text.mdGemini long-form, GPT-4o-Transcribe, AssemblyAI features

Audio: Text-to-Speech

Generate natural speech from text with voice selection and expressive cues.

RuleFileKey Pattern
Text-to-Speechrules/audio-text-to-speech.mdGemini TTS, voice config, auditory cues

Audio: Model Selection

Select the right audio/voice provider for real-time, transcription, or TTS use cases.

RuleFileKey Pattern
Audio Modelsrules/audio-models.mdReal-time voice comparison, STT benchmarks, pricing

Video: Model Selection

Choose the right video generation provider based on use case, duration, and budget.

RuleFileKey Pattern
Video Modelsrules/video-generation-models.mdKling vs Sora vs Veo vs Runway, pricing, capabilities

Video: API Patterns

Integrate video generation APIs with proper async polling, SDKs, and webhook callbacks.

RuleFileKey Pattern
API Integrationrules/video-generation-patterns.mdKling REST, fal.ai SDK, Vercel AI SDK, task polling
Show full SKILL.md (330 more words)Show less

Video: Multi-Shot

Generate multi-scene videos with consistent characters using storyboarding and character elements.

RuleFileKey Pattern
Multi-Shotrules/video-multi-shot.mdKling v3 character elements, 6-shot storyboards, identity binding

Key Decisions

DecisionRecommendation
High accuracy visionclaude-opus-5-5 (default Opus since CC 2.1.280, 2,576 px vision budget, 3× what Opus 4.6 allotted; give it crop/analyze tools rather than more thinking, which is the cheaper lever on this model). (claude-fable-5 is the frontier SOTA option, GA 2026-07 — premium cost, use only with explicit consent via the fable-spend-consent gate)
Long documentsgemini-3.1-pro-preview (1M+ context)
Cost-efficient visiongemini-3.1-flash-lite (GA successor of the Flash-Lite preview, which was shut down 2026-05-25; shutdown scheduled 2027-05-07)
Video analysisgemini-3.1-pro-preview (native video, supersedes 2.5 Pro)
Voice assistantGrok Voice Agent on Grok 4.20 (fastest, <1s)
Emotional voice AIGemini Live API
Long audio transcriptiongemini-3.1-pro-preview (9.5hr)
Speaker diarizationAssemblyAI or Gemini
Self-hosted STTWhisper Large V3
Character-consistent videokling-v3 (Character Elements 3.0)
Narrative video / storytellingSora 2 (best cause-and-effect coherence)
Cinematic B-rollveo-3.1-generate-preview (camera control + polished motion)
Budget draftsveo-3.1-lite-generate-preview (~$0.05/s, 720/1080p)
Mid-tier fast rendersveo-3.1-fast-generate-preview
Professional VFXRunway gen4_turbo (Act-Two motion transfer)
High-volume social videokling-v3 Standard (~$0.20/video)
Open-source video genWan 2.6 or LTX-2
Lip-sync / avatar videokling-v3 (native lip-sync API)

Example

python
import anthropic, base64

client = anthropic.Anthropic()
with open("image.png", "rb") as f:
    b64 = base64.standard_b64encode(f.read()).decode("utf-8")

response = client.messages.create(
    model="claude-opus-5-5",
    max_tokens=1024,
    messages=[{"role": "user", "content": [
        {"type": "image", "source": {"type": "base64", "media_type": "image/png", "data": b64}},
        {"type": "text", "text": "Describe this image"}
    ]}]
)

Common Mistakes

  1. Not setting max_tokens on vision requests (responses truncated)
  2. Sending oversized images without resizing (>2048px)
  3. Using high detail level for simple yes/no classification
  4. Using STT+LLM+TTS pipeline instead of native speech-to-speech
  5. Not leveraging barge-in support for natural voice conversations
  6. Using deprecated models (GPT-4V, Whisper-1)
  7. Ignoring rate limits on vision and audio endpoints
  8. Calling video generation APIs synchronously (they're async — poll or use callbacks)
  9. Generating separate clips without character elements (characters look different each time)
  10. Using Sora for high-volume social content (expensive, slow — use Kling Standard instead)
  • ork:rag-retrieval - Multimodal RAG with image + text retrieval
  • ork:llm-integration - General LLM function calling patterns
  • streaming-api-patterns - WebSocket patterns for real-time audio
  • ork:demo-producer - Terminal demo videos (VHS, asciinema) — not AI video gen

© yonatangross, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 12 other files in src/skills/multimodal-llm of yonatangross/orchestkit.

  • SKILL.md
  • rules/_sections.md
  • rules/_template.md
  • rules/audio-models.md
  • rules/audio-speech-to-text.md
  • rules/audio-text-to-speech.md
  • rules/video-generation-models.md
  • rules/video-generation-patterns.md
  • rules/video-multi-shot.md
  • rules/vision-document.md
  • rules/vision-image-analysis.md
  • rules/vision-models.md
  • test-cases.json

Open the folder on GitHubat commit e4ff8d9

Compare with similar skills

Multimodal LLM next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Multimodal LLM compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Multimodal LLM this skillyonatangross/orchestkit292—~2.2kAutomated safety check: PassMIT
Veo Usecnemri/google-genai-skills127—~625Automated safety check: PassMIT
Veo Buildcnemri/google-genai-skills127—~513Automated safety check: PassMIT
Video Gen UsageOtoDock/oto-dock190—~1.4kAutomated safety check: PassCustom licence
Capcutsocial-media-skills/skills134—~2kAutomated safety check: PassMIT
Klingsocial-media-skills/skills134—~1.3kAutomated safety check: PassMIT

Similar skills

  • Veo Use

    cnemri/google-genai-skills

    Create and edit videos using Google's Veo 2 and Veo 3 models.

    127 GitHub stars~625 tokensUpdated 8 mo ago
    Media & CreativeAuto-check passed
  • Veo Build

    cnemri/google-genai-skills

    Create and edit videos using Google's Veo 2 and Veo 3 models.

    127 GitHub stars~513 tokensUpdated 8 mo ago
    Media & CreativeAuto-check passed
  • Video Gen Usage

    OtoDock/oto-dock

    Generate AI videos, transitions between clips, and AI video edits.

    190 GitHub stars~1.4k tokensUpdated 2 days ago
    Media & CreativeAuto-check passed
  • Capcut

    social-media-skills/skills

    The CapCut craft skill — edit short-form social video (TikTok, Reels, Shorts) fast and safely: retention-paced cuts, auto-captions, beat-synced sound, exports.

    134 GitHub stars~2k tokensUpdated 8 days ago
    Media & CreativeAuto-check passed
  • Kling

    social-media-skills/skills

    The generative-video producer for native 4K, multi-shot storyboarding, and motion-transfer (Kling-led).

    134 GitHub stars~1.3k tokensUpdated 8 days ago
    Media & CreativeAuto-check passed
  • Runway

    social-media-skills/skills

    Prompt and direct Runway (Gen-4.5, Aleph) — the control-grade generative video tool for camera moves, character consistency, and edit-grade shots.

    134 GitHub stars~1.3k tokensUpdated 8 days ago
    Media & CreativeAuto-check passed

More from yonatangross/orchestkit

All 108 skills in this repo
  • API Design

    yonatangross/orchestkit

    API contract design for REST and GraphQL, covering resource shape, URL and header versioning with deprecation windows, RFC 9457 Problem Details error handling, and OpenAPI specs.

    292 GitHub stars~2.9k tokensUpdated today
    Auto-check passed
  • Architecture Decision Record

    yonatangross/orchestkit

    ADR templates in the Nygard format with context, decision, consequences, and alternatives.

    292 GitHub stars~2k tokensUpdated today
    Auto-check passed
  • Audit Full

    yonatangross/orchestkit

    Single-pass codebase analysis leveraging a 1M-token context window for comprehensive security scanning, architecture review, and dependency auditing.

    292 GitHub stars~3.5k tokensUpdated today
    Auto-check: notes
  • Code Review Playbook

    yonatangross/orchestkit

    Structured review processes, conventional comments, language-specific checklists, and feedback templates.

    292 GitHub stars~2.2k tokensUpdated today
    Auto-check passed
  • Create PR

    yonatangross/orchestkit

    Creates GitHub pull requests with pre-flight validation, conventional title formatting, and structured summary generation.

    292 GitHub stars~4.5k tokensUpdated today
    Auto-check: notes
  • Explore

    yonatangross/orchestkit

    Multi-angle codebase exploration spawning 3-5 parallel agents for code structure, data flow, architecture patterns, and health assessment.

    292 GitHub stars~3.9k tokensUpdated today
    Auto-check: notes

Works with

Questions about Multimodal LLM

What does Multimodal LLM do?

Vision, audio, video generation, and multimodal LLM integration patterns. Multimodal LLM is an agent skill from yonatangross/orchestkit. Vision, audio, video generation, and multimodal LLM integration patterns.

When should I use Multimodal LLM?

Multimodal LLM fits situations like: processing images; transcribing audio; generating speech; generating AI video (Kling v3.

How do I install Multimodal LLM in Claude Code?

Run `npx skills add yonatangross/orchestkit --skill multimodal-llm -a claude-code`. Or copy the skill folder (src/skills/multimodal-llm in yonatangross/orchestkit) into .claude/skills/multimodal-llm in your project. Claude Code loads it when a task matches its description.

How do I install Multimodal LLM in Codex?

Run `npx skills add yonatangross/orchestkit --skill multimodal-llm -a codex`. Or copy the skill folder (src/skills/multimodal-llm in yonatangross/orchestkit) into .agents/skills/multimodal-llm in your project. Codex loads it when a task matches its description.

Can I use Multimodal LLM in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add yonatangross/orchestkit --skill multimodal-llm -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/multimodal-llm, .gemini/skills/multimodal-llm, .github/skills/multimodal-llm and .opencode/skills/multimodal-llm in your project.

What does Multimodal LLM need to run?

SKILL.md names no scripts, command-line tools or credentials: Multimodal LLM is instructions for the agent only. Our summary lists: Python 3. Its frontmatter pre-approves these tools: Read, Glob, Grep, WebFetch, WebSearch. Compatibility (from SKILL.md): Claude Code 2.1.277+..

Does Multimodal LLM access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Multimodal LLM safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Multimodal LLM use?

Multimodal LLM is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Multimodal LLM use?

About 2.2k tokens (SKILL.md is roughly 9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Multimodal LLM?

Skills that share tags, products or a category with Multimodal LLM: Veo Use (cnemri/google-genai-skills, 127 stars), Veo Build (cnemri/google-genai-skills, 127 stars), Video Gen Usage (OtoDock/oto-dock, 190 stars) and Capcut (social-media-skills/skills, 134 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Multimodal LLM?

yonatangross (a GitHub user) maintains it in yonatangross/orchestkit, which has 292 GitHub stars. The repository holds 108 skills in this directory. The repository was last updated on October 10, 2026.

Source: yonatangross/orchestkit on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.