Agent skill

Ag2 Multimodal Input

by ag2ai in ag2ai/build-with-ag2

Send images, audio, video, or documents into an AG2 beta Agent alongside text.

Apache-2.0Auto-check passedDocuments & Office

Install Ag2 Multimodal Input

skills CLI
$ npx skills add ag2ai/build-with-ag2 --skill ag2-multimodal-input -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install ag2ai/build-with-ag2 ag2-multimodal-input --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/ag2ai/build-with-ag2.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/ag2-multimodal-input .claude/skills/ag2-multimodal-input && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
ag2-multimodal-input
GitHub stars
252
Token cost
~1.7k tokens
SKILL.md length
350 words
Files
1
Skills in repo
16
Repo updated
First seen
Licence
Apache-2.0

At a glance

Send images, audio, video, or documents into an AG2 beta Agent alongside text.

  • The user wants the agent to process non-text input — describe a photo
  • SKILL.md covers When to use, 60-second recipe, Input factories and Provider matrix, plus 4 more sections
  • Reaches youtube.com
  • Transcribe audio

What it does

Ag2 Multimodal Input is an agent skill from ag2ai/build-with-ag2. Send images, audio, video, or documents into an AG2 beta Agent alongside text. Pass ImageInput, AudioInput, VideoInput, or DocumentInput as positional args to agent.ask(...). Use when the user wants the agent to process non-text input — describe a photo, transcribe audio, summarise a PDF, analyse a video. Covers per-provider support matrix, the four ways to source data (URL / path / bytes / fileid), Gemini-specific YouTube + media-resolution + clipping, OpenAI image-detail, Anthropic prompt-caching on…

Its SKILL.md is about 1.7k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Documents & Office, covering LLM cost and token optimization, Transcription and PDF. It works with OpenAI and YouTube. The repository describes itself as: Sample code and application showcases to get you going with AG2 (formally AutoGen). The licence is Apache-2.0.

When your agent uses it

  • The user wants the agent to process non-text input — describe a photo
  • Transcribe audio
  • Summarise a PDF
  • Analyse a video

Example prompts

  • “/ag2-multimodal-input”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit 29eeac3. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • youtube.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Ag2 Multimodal Input loads about 1.7k tokens when it runs. Until then it costs about 148 tokens; SKILL.md has 350 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~148
When it runs · the whole SKILL.md, loaded when a task matches
~1.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from ag2ai/build-with-ag2 at commit 29eeac3, republished under its Apache-2.0 licence (© ag2ai). 350 words, ~1,710 tokens.

Download SKILL.mdSave it as .claude/skills/ag2-multimodal-input/SKILL.md (or your agent's skills folder).
name
ag2-multimodal-input
description
Send images, audio, video, or documents into an AG2 beta `Agent` alongside text. Pass `ImageInput`, `AudioInput`, `VideoInput`, or `DocumentInput` as positional args to `agent.ask(...)`. Use when the user wants the agent to process non-text input — describe a photo, transcribe audio, summarise a PDF, analyse a video. Covers per-provider support matrix, the four ways to source data (URL / path / bytes / file_id), Gemini-specific YouTube + media-resolution + clipping, OpenAI image-detail, Anthropic prompt-caching on attachments, and `FilesAPI` for upload lifecycle.
license
Apache-2.0

Multimodal inputs

When to use

The user wants the agent to process non-text input: an image to describe, audio to transcribe, video to summarise, or a PDF / document to extract from. The same factory pattern works across providers; per-provider support varies.

60-second recipe

python
from autogen.beta import Agent
from autogen.beta.config import GeminiConfig
from autogen.beta.events import ImageInput

agent = Agent(
    "vision",
    "You describe images.",
    config=GeminiConfig(model="gemini-3-flash-preview"),
)

image = ImageInput("https://example.com/photo.jpg")
reply = await agent.ask("Describe this image in detail.", image)
print(reply.body)

Multiple inputs in one ask are fine:

python
reply = await agent.ask(
    "Compare these two images.",
    ImageInput("https://example.com/before.jpg"),
    ImageInput("https://example.com/after.jpg"),
)

Input factories

FactoryFormats
ImageInput(...)JPEG, PNG, GIF, WebP
AudioInput(...)WAV, MP3, OGG, FLAC, AAC
VideoInput(...)MP4, WebM, MOV, MKV, MPEG
DocumentInput(...)PDF, TXT, HTML, Markdown, CSV, JSON, Office formats

Each accepts the same four data sources:

python
from autogen.beta.events import ImageInput

ImageInput("https://example.com/photo.jpg")     # URL
ImageInput(path="photo.jpg")                    # local file
ImageInput(data=raw_bytes, media_type="image/png")  # bytes
ImageInput(file_id="file-abc123")               # provider-uploaded

Provider matrix

Input typeOpenAIOpenAI ResponsesGeminiAnthropic
Text✓✓✓✓
Image (URL)✓✓✓✓
Image (binary)✓✓✓✓
Audio (URL)––✓–
Audio (binary)✓–✓–
Video (URL)––✓–
Video (binary)––✓–
Document (URL)–✓✓✓
Document (binary)––✓✓
File ID–✓–✓

Unsupported combinations raise UnsupportedInputError with a clear message.

Gemini has the broadest multimodal support. If you don't know which provider to pick for a multimodal task, start there.

Provider-specific niceties

Gemini — YouTube URLs work directly
python
from autogen.beta.events import VideoInput

video = VideoInput("https://www.youtube.com/watch?v=dQw4w9WgXcQ")
reply = await agent.ask("Summarize this video.", video)
Gemini — large files (> 20MB) via Google Files API
python
from google import genai
from autogen.beta.events import VideoInput
import time

client = genai.Client()
uploaded = client.files.upload(file="large_video.mp4")
while uploaded.state.name == "PROCESSING":
    time.sleep(2)
    uploaded = client.files.get(name=uploaded.name)

video = VideoInput(uploaded.uri)
Gemini — vendor_metadata
KeyPurpose
media_resolutionMEDIA_RESOLUTION_LOW/MEDIUM/HIGH/ULTRA_HIGH — token vs cost
video_metadataClipping (start_offset, end_offset) and fps
display_nameDisplay name for the file
python
ImageInput(data=raw, media_type="image/jpeg", vendor_metadata={"media_resolution": "MEDIA_RESOLUTION_LOW"})

VideoInput(path="lecture.mp4", vendor_metadata={
    "video_metadata": {"start_offset": "60s", "end_offset": "120s", "fps": 0.5},
})
OpenAI — image detail
python
ImageInput(data=raw, media_type="image/png", vendor_metadata={"detail": "low"})  # "low" | "high" | "auto"
Anthropic — File ID + prompt caching
python
import anthropic
from autogen.beta.events import ImageInput, DocumentInput

client = anthropic.Anthropic()
uploaded = client.beta.files.upload(file=("photo.jpg", open("photo.jpg", "rb"), "image/jpeg"))

# filename determines block type (image vs document)
image = ImageInput(file_id=uploaded.id, filename="photo.jpg")

# Cache an attachment so subsequent turns skip re-uploading
doc = DocumentInput(path="report.pdf", vendor_metadata={"cache_control": {"type": "ephemeral"}})

FilesAPI — upload lifecycle, provider-agnostic

For any provider that has a file API (OpenAIConfig, OpenAIResponsesConfig, AnthropicConfig, GeminiConfig):

python
from autogen.beta import FilesAPI
from autogen.beta.config import OpenAIResponsesConfig

files = FilesAPI(OpenAIResponsesConfig(model="gpt-5-mini"))

uploaded = await files.upload(path="report.pdf", purpose="assistants")
print(uploaded.file_id)

# Or from bytes (filename required)
uploaded = await files.upload(data=b"...", filename="hello.txt", purpose="assistants")

# List, read, delete
all_files = await files.list()
data = await files.read(uploaded.file_id)        # NotImplementedError on Gemini
await files.delete(uploaded.file_id)

Pass the file_id to DocumentInput, ImageInput, etc.:

python
from autogen.beta.events import DocumentInput

doc = DocumentInput(file_id=uploaded.file_id)
reply = await agent.ask("Summarize this report.", doc)
Show full SKILL.md (140 more words)Show less

Going deeper

  • website/docs/beta/inputs/inputs.mdx — full provider matrix and vendor_metadata reference.
  • website/docs/beta/advanced/files.mdx — FilesAPI reference (upload / list / read / delete).
  • For tools that return images / binary back to the LLM, see ag2-add-custom-tool (ImageInput, BinaryInput, ToolResult).

Common pitfalls

  • Picking a provider that doesn't support your input type — silently you'll get UnsupportedInputError. Check the matrix; Gemini is broadest.
  • FilesAPI.read() on Gemini — raises NotImplementedError. Gemini doesn't expose download.
  • Calling files.upload(data=...) without filename= — raises ValueError. Filename is required for in-memory uploads.
  • Providing path= and data= to the same factory — pick one source. Same for file_id=.
  • Anthropic ImageInput(file_id=...) without filename= — Anthropic decides block type (image vs document) by filename extension. Pass it.
  • Gemini vendor_metadata keys are nested — video_metadata itself takes a dict. Check the doc table for shape.
  • Forgetting to wait for Gemini file processing — large uploads have a PROCESSING state. Poll client.files.get(name=...) until ready before referencing the URI.

© ag2ai, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/ag2-multimodal-input of ag2ai/build-with-ag2.

Open the folder on GitHubat commit 29eeac3

Compare with similar skills

Ag2 Multimodal Input next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Ag2 Multimodal Input compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Ag2 Multimodal Input this skillag2ai/build-with-ag2252—~1.7kAutomated safety check: PassApache-2.0
MarkitdownImCa0/just-laws78214 repos~3.2kAutomated safety check: NotesMIT
Markitdownjimmc414/Kosmos5952 repos~1.7kAutomated safety check: PassNone
Watchlesschenzixin1/watchless144—~3.7kAutomated safety check: WarnMIT
Lecture To Notesysyecust/lecture-to-notes273—~14kAutomated safety check: NotesCustom licence
Youtube SummarizerBrianRWagner/ai-marketing-claude-code-skills4401 repos~3.8kAutomated safety check: PassNone

Similar skills

  • Markitdown

    ImCa0/just-laws

    Convert files and office documents to Markdown. An agent skill from ImCa0/just-laws.

    782 GitHub starsUsed in 14 repos~3.2k tokens
    Documents & OfficeAuto-check: notes
  • Markitdown

    jimmc414/Kosmos

    Convert various file formats (PDF, Office documents, images, audio, web content, structured data) to Markdown optimized for LLM processing.

    595 GitHub starsUsed in 2 repos~1.7k tokens
    Documents & OfficeAuto-check passed
  • Watchless

    chenzixin1/watchless

    A skill your agent uses when turning a YouTube URL or local presentation, explainer, interview, podcast, or product-demo video into complete screenshot-led notes, faithful light-polished text, HTML…

    144 GitHub stars~3.7k tokensUpdated 23 days ago
    Media & CreativeAuto-check: warnings
  • Lecture To Notes

    ysyecust/lecture-to-notes

    A skill your agent uses when users provide YouTube, Bilibili, or X/Twitter lecture URLs and want reader-first Chinese LaTeX/PDF notes with source-faithful claims, fluent authored prose, and verified…

    273 GitHub stars~14k tokensUpdated 6 days ago
    Documents & OfficeAuto-check: notes
  • Youtube Summarizer

    BrianRWagner/ai-marketing-claude-code-skills

    Automatically fetch YouTube video transcripts, generate structured summaries, and send full transcripts to messaging platforms.

    440 GitHub starsUsed in 1 repo~3.8k tokens
    Productivity & AutomationAuto-check passed
  • Video Transcribe

    wendy7756/AI-Video-Transcriber

    Transcribe and summarize a video or podcast from a URL (YouTube, TikTok, Bilibili, Apple Podcasts, SoundCloud, 30+ platforms) or from a local media/.txt file.

    3.3k GitHub stars~937 tokensUpdated 24 days ago
    Media & CreativeAuto-check: notes

More from ag2ai/build-with-ag2

All 16 skills in this repo
  • Ag2 Add Custom Tool

    ag2ai/build-with-ag2

    Add a custom Python tool to an AG2 beta Agent using the @tool decorator.

    252 GitHub stars~1.9k tokensUpdated 1 mo ago
    Auto-check passed
  • Ag2 Middleware

    ag2ai/build-with-ag2

    Intercept the AG2 beta agent loop with BaseMiddleware — wrap full turns (onturn), each LLM call (onllmcall), each tool execution (ontoolexecution), or each human-input request (onhumaninput).

    252 GitHub stars~1.9k tokensUpdated 1 mo ago
    Auto-check passed
  • Ag2 Use Builtin Tools

    ag2ai/build-with-ag2

    Wire AG2 beta's shipped tools into an Agent — both provider-native server-side tools (web search, web fetch, code execution, MCP, image generation, memory) and locally-executed common toolkits…

    252 GitHub stars~1.3k tokensUpdated 1 mo ago
    Auto-check passed
  • Ag2 Knowledge And Memory

    ag2ai/build-with-ag2

    Persist agent state across runs, shape what the LLM sees per turn, and cap history to fit a context window.

    252 GitHub stars~2.9k tokensUpdated 1 mo ago
    Auto-check passed
  • Ag2 Observers And Alerts

    ag2ai/build-with-ag2

    Monitor an AG2 beta agent's stream — log events, detect repeated tool calls, track token spend, build trigger-driven observers, route observer alerts to the model, and halt on FATAL conditions.

    252 GitHub stars~2.5k tokensUpdated 1 mo ago
    Auto-check passed
  • Ag2 Quickstart

    ag2ai/build-with-ag2

    Build a minimal AG2 beta Agent end to end — pick a model provider, set a prompt, call agent.ask(), then continue the conversation with reply.ask() (multi-turn).

    252 GitHub stars~1.7k tokensUpdated 1 mo ago
    Auto-check: notes

Works with

Questions about Ag2 Multimodal Input

What does Ag2 Multimodal Input do?

Send images, audio, video, or documents into an AG2 beta Agent alongside text. Ag2 Multimodal Input is an agent skill from ag2ai/build-with-ag2. Send images, audio, video, or documents into an AG2 beta Agent alongside text.

When should I use Ag2 Multimodal Input?

Ag2 Multimodal Input fits situations like: the user wants the agent to process non-text input — describe a photo; transcribe audio; summarise a PDF; analyse a video.

How do I install Ag2 Multimodal Input in Claude Code?

Run `npx skills add ag2ai/build-with-ag2 --skill ag2-multimodal-input -a claude-code`. Or copy the skill folder (.agents/skills/ag2-multimodal-input in ag2ai/build-with-ag2) into .claude/skills/ag2-multimodal-input in your project. Claude Code loads it when a task matches its description.

How do I install Ag2 Multimodal Input in Codex?

Run `npx skills add ag2ai/build-with-ag2 --skill ag2-multimodal-input -a codex`. Or copy the skill folder (.agents/skills/ag2-multimodal-input in ag2ai/build-with-ag2) into .agents/skills/ag2-multimodal-input in your project. Codex loads it when a task matches its description.

Can I use Ag2 Multimodal Input in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ag2ai/build-with-ag2 --skill ag2-multimodal-input -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ag2-multimodal-input, .gemini/skills/ag2-multimodal-input, .github/skills/ag2-multimodal-input and .opencode/skills/ag2-multimodal-input in your project.

What does Ag2 Multimodal Input need to run?

SKILL.md names no scripts, command-line tools or credentials: Ag2 Multimodal Input is instructions for the agent only. Our summary lists: Python 3.

Does Ag2 Multimodal Input access the network?

SKILL.md names 1 domain. In commands or code: youtube.com; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is Ag2 Multimodal Input safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Ag2 Multimodal Input use?

Ag2 Multimodal Input is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Ag2 Multimodal Input use?

About 1.7k tokens (SKILL.md is roughly 6.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Ag2 Multimodal Input?

Skills that share tags, products or a category with Ag2 Multimodal Input: Markitdown (ImCa0/just-laws, 782 stars), Markitdown (jimmc414/Kosmos, 595 stars), Watchless (chenzixin1/watchless, 144 stars) and Lecture To Notes (ysyecust/lecture-to-notes, 273 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Ag2 Multimodal Input?

ag2ai (a GitHub organization) maintains it in ag2ai/build-with-ag2, which has 252 GitHub stars. The repository holds 16 skills in this directory. The repository was last updated on September 6, 2026.

Source: ag2ai/build-with-ag2 on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.