Markitdown
ImCa0/just-laws
Convert files and office documents to Markdown. An agent skill from ImCa0/just-laws.
Send images, audio, video, or documents into an AG2 beta Agent alongside text.
$ npx skills add ag2ai/build-with-ag2 --skill ag2-multimodal-input -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install ag2ai/build-with-ag2 ag2-multimodal-input --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/ag2ai/build-with-ag2.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/ag2-multimodal-input .claude/skills/ag2-multimodal-input && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "ag2-multimodal-input" agent skill from https://github.com/ag2ai/build-with-ag2/tree/main/.agents/skills/ag2-multimodal-input into .claude/skills/ag2-multimodal-input/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ag2-multimodal-input", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/ag2ai/build-with-ag2/tree/main/.agents/skills/ag2-multimodal-inputType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add ag2ai/build-with-ag2 --skill ag2-multimodal-input -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install ag2ai/build-with-ag2 ag2-multimodal-input --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ag2ai/build-with-ag2.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.agents/skills/ag2-multimodal-input .agents/skills/ag2-multimodal-input && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "ag2-multimodal-input" agent skill from https://github.com/ag2ai/build-with-ag2/tree/main/.agents/skills/ag2-multimodal-input into .agents/skills/ag2-multimodal-input/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ag2-multimodal-input", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add ag2ai/build-with-ag2 --skill ag2-multimodal-input -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install ag2ai/build-with-ag2 ag2-multimodal-input --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ag2ai/build-with-ag2.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.agents/skills/ag2-multimodal-input .cursor/skills/ag2-multimodal-input && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "ag2-multimodal-input" agent skill from https://github.com/ag2ai/build-with-ag2/tree/main/.agents/skills/ag2-multimodal-input into .cursor/skills/ag2-multimodal-input/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ag2-multimodal-input", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/ag2ai/build-with-ag2.git --path .agents/skills/ag2-multimodal-input--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add ag2ai/build-with-ag2 --skill ag2-multimodal-input -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install ag2ai/build-with-ag2 ag2-multimodal-input --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ag2ai/build-with-ag2.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.agents/skills/ag2-multimodal-input .gemini/skills/ag2-multimodal-input && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "ag2-multimodal-input" agent skill from https://github.com/ag2ai/build-with-ag2/tree/main/.agents/skills/ag2-multimodal-input into .gemini/skills/ag2-multimodal-input/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ag2-multimodal-input", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install ag2ai/build-with-ag2 ag2-multimodal-inputInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add ag2ai/build-with-ag2 --skill ag2-multimodal-input -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/ag2ai/build-with-ag2.git skills-src && mkdir -p .github/skills && cp -r skills-src/.agents/skills/ag2-multimodal-input .github/skills/ag2-multimodal-input && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "ag2-multimodal-input" agent skill from https://github.com/ag2ai/build-with-ag2/tree/main/.agents/skills/ag2-multimodal-input into .github/skills/ag2-multimodal-input/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ag2-multimodal-input", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add ag2ai/build-with-ag2 --skill ag2-multimodal-input -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install ag2ai/build-with-ag2 ag2-multimodal-input --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ag2ai/build-with-ag2.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.agents/skills/ag2-multimodal-input .opencode/skills/ag2-multimodal-input && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "ag2-multimodal-input" agent skill from https://github.com/ag2ai/build-with-ag2/tree/main/.agents/skills/ag2-multimodal-input into .opencode/skills/ag2-multimodal-input/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ag2-multimodal-input", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
ag2-multimodal-inputSend images, audio, video, or documents into an AG2 beta Agent alongside text.
Ag2 Multimodal Input is an agent skill from ag2ai/build-with-ag2. Send images, audio, video, or documents into an AG2 beta Agent alongside text. Pass ImageInput, AudioInput, VideoInput, or DocumentInput as positional args to agent.ask(...). Use when the user wants the agent to process non-text input — describe a photo, transcribe audio, summarise a PDF, analyse a video. Covers per-provider support matrix, the four ways to source data (URL / path / bytes / fileid), Gemini-specific YouTube + media-resolution + clipping, OpenAI image-detail, Anthropic prompt-caching on…
Its SKILL.md is about 1.7k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Documents & Office, covering LLM cost and token optimization, Transcription and PDF. It works with OpenAI and YouTube. The repository describes itself as: Sample code and application showcases to get you going with AG2 (formally AutoGen). The licence is Apache-2.0.
Read from SKILL.md and the folder at commit 29eeac3. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are python).
From the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
youtube.comFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Ag2 Multimodal Input loads about 1.7k tokens when it runs. Until then it costs about 148 tokens; SKILL.md has 350 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from ag2ai/build-with-ag2 at commit 29eeac3, republished under its Apache-2.0 licence (© ag2ai). 350 words, ~1,710 tokens.
.claude/skills/ag2-multimodal-input/SKILL.md (or your agent's skills folder).The user wants the agent to process non-text input: an image to describe, audio to transcribe, video to summarise, or a PDF / document to extract from. The same factory pattern works across providers; per-provider support varies.
from autogen.beta import Agent
from autogen.beta.config import GeminiConfig
from autogen.beta.events import ImageInput
agent = Agent(
"vision",
"You describe images.",
config=GeminiConfig(model="gemini-3-flash-preview"),
)
image = ImageInput("https://example.com/photo.jpg")
reply = await agent.ask("Describe this image in detail.", image)
print(reply.body)Multiple inputs in one ask are fine:
reply = await agent.ask(
"Compare these two images.",
ImageInput("https://example.com/before.jpg"),
ImageInput("https://example.com/after.jpg"),
)| Factory | Formats |
|---|---|
ImageInput(...) | JPEG, PNG, GIF, WebP |
AudioInput(...) | WAV, MP3, OGG, FLAC, AAC |
VideoInput(...) | MP4, WebM, MOV, MKV, MPEG |
DocumentInput(...) | PDF, TXT, HTML, Markdown, CSV, JSON, Office formats |
Each accepts the same four data sources:
from autogen.beta.events import ImageInput
ImageInput("https://example.com/photo.jpg") # URL
ImageInput(path="photo.jpg") # local file
ImageInput(data=raw_bytes, media_type="image/png") # bytes
ImageInput(file_id="file-abc123") # provider-uploaded| Input type | OpenAI | OpenAI Responses | Gemini | Anthropic |
|---|---|---|---|---|
| Text | ✓ | ✓ | ✓ | ✓ |
| Image (URL) | ✓ | ✓ | ✓ | ✓ |
| Image (binary) | ✓ | ✓ | ✓ | ✓ |
| Audio (URL) | – | – | ✓ | – |
| Audio (binary) | ✓ | – | ✓ | – |
| Video (URL) | – | – | ✓ | – |
| Video (binary) | – | – | ✓ | – |
| Document (URL) | – | ✓ | ✓ | ✓ |
| Document (binary) | – | – | ✓ | ✓ |
| File ID | – | ✓ | – | ✓ |
Unsupported combinations raise UnsupportedInputError with a clear message.
Gemini has the broadest multimodal support. If you don't know which provider to pick for a multimodal task, start there.
from autogen.beta.events import VideoInput
video = VideoInput("https://www.youtube.com/watch?v=dQw4w9WgXcQ")
reply = await agent.ask("Summarize this video.", video)from google import genai
from autogen.beta.events import VideoInput
import time
client = genai.Client()
uploaded = client.files.upload(file="large_video.mp4")
while uploaded.state.name == "PROCESSING":
time.sleep(2)
uploaded = client.files.get(name=uploaded.name)
video = VideoInput(uploaded.uri)vendor_metadata| Key | Purpose |
|---|---|
media_resolution | MEDIA_RESOLUTION_LOW/MEDIUM/HIGH/ULTRA_HIGH — token vs cost |
video_metadata | Clipping (start_offset, end_offset) and fps |
display_name | Display name for the file |
ImageInput(data=raw, media_type="image/jpeg", vendor_metadata={"media_resolution": "MEDIA_RESOLUTION_LOW"})
VideoInput(path="lecture.mp4", vendor_metadata={
"video_metadata": {"start_offset": "60s", "end_offset": "120s", "fps": 0.5},
})ImageInput(data=raw, media_type="image/png", vendor_metadata={"detail": "low"}) # "low" | "high" | "auto"import anthropic
from autogen.beta.events import ImageInput, DocumentInput
client = anthropic.Anthropic()
uploaded = client.beta.files.upload(file=("photo.jpg", open("photo.jpg", "rb"), "image/jpeg"))
# filename determines block type (image vs document)
image = ImageInput(file_id=uploaded.id, filename="photo.jpg")
# Cache an attachment so subsequent turns skip re-uploading
doc = DocumentInput(path="report.pdf", vendor_metadata={"cache_control": {"type": "ephemeral"}})FilesAPI — upload lifecycle, provider-agnosticFor any provider that has a file API (OpenAIConfig, OpenAIResponsesConfig, AnthropicConfig, GeminiConfig):
from autogen.beta import FilesAPI
from autogen.beta.config import OpenAIResponsesConfig
files = FilesAPI(OpenAIResponsesConfig(model="gpt-5-mini"))
uploaded = await files.upload(path="report.pdf", purpose="assistants")
print(uploaded.file_id)
# Or from bytes (filename required)
uploaded = await files.upload(data=b"...", filename="hello.txt", purpose="assistants")
# List, read, delete
all_files = await files.list()
data = await files.read(uploaded.file_id) # NotImplementedError on Gemini
await files.delete(uploaded.file_id)Pass the file_id to DocumentInput, ImageInput, etc.:
from autogen.beta.events import DocumentInput
doc = DocumentInput(file_id=uploaded.file_id)
reply = await agent.ask("Summarize this report.", doc)website/docs/beta/inputs/inputs.mdx — full provider matrix and vendor_metadata reference.website/docs/beta/advanced/files.mdx — FilesAPI reference (upload / list / read / delete).ag2-add-custom-tool (ImageInput, BinaryInput, ToolResult).UnsupportedInputError. Check the matrix; Gemini is broadest.FilesAPI.read() on Gemini — raises NotImplementedError. Gemini doesn't expose download.files.upload(data=...) without filename= — raises ValueError. Filename is required for in-memory uploads.path= and data= to the same factory — pick one source. Same for file_id=.ImageInput(file_id=...) without filename= — Anthropic decides block type (image vs document) by filename extension. Pass it.vendor_metadata keys are nested — video_metadata itself takes a dict. Check the doc table for shape.PROCESSING state. Poll client.files.get(name=...) until ready before referencing the URI.© ag2ai, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .agents/skills/ag2-multimodal-input of ag2ai/build-with-ag2.
Open the folder on GitHubat commit 29eeac3
Ag2 Multimodal Input next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Ag2 Multimodal Input this skillag2ai/build-with-ag2 | 252 | — | ~1.7k | Automated safety check: Pass | Apache-2.0 | |
| MarkitdownImCa0/just-laws | 782 | 14 repos | ~3.2k | Automated safety check: Notes | MIT | |
| Markitdownjimmc414/Kosmos | 595 | 2 repos | ~1.7k | Automated safety check: Pass | None | |
| Watchlesschenzixin1/watchless | 144 | — | ~3.7k | Automated safety check: Warn | MIT | |
| Lecture To Notesysyecust/lecture-to-notes | 273 | — | ~14k | Automated safety check: Notes | Custom licence | |
| Youtube SummarizerBrianRWagner/ai-marketing-claude-code-skills | 440 | 1 repos | ~3.8k | Automated safety check: Pass | None |
ImCa0/just-laws
Convert files and office documents to Markdown. An agent skill from ImCa0/just-laws.
jimmc414/Kosmos
Convert various file formats (PDF, Office documents, images, audio, web content, structured data) to Markdown optimized for LLM processing.
chenzixin1/watchless
A skill your agent uses when turning a YouTube URL or local presentation, explainer, interview, podcast, or product-demo video into complete screenshot-led notes, faithful light-polished text, HTML…
ysyecust/lecture-to-notes
A skill your agent uses when users provide YouTube, Bilibili, or X/Twitter lecture URLs and want reader-first Chinese LaTeX/PDF notes with source-faithful claims, fluent authored prose, and verified…
BrianRWagner/ai-marketing-claude-code-skills
Automatically fetch YouTube video transcripts, generate structured summaries, and send full transcripts to messaging platforms.
wendy7756/AI-Video-Transcriber
Transcribe and summarize a video or podcast from a URL (YouTube, TikTok, Bilibili, Apple Podcasts, SoundCloud, 30+ platforms) or from a local media/.txt file.
ag2ai/build-with-ag2
Add a custom Python tool to an AG2 beta Agent using the @tool decorator.
ag2ai/build-with-ag2
Intercept the AG2 beta agent loop with BaseMiddleware — wrap full turns (onturn), each LLM call (onllmcall), each tool execution (ontoolexecution), or each human-input request (onhumaninput).
ag2ai/build-with-ag2
Wire AG2 beta's shipped tools into an Agent — both provider-native server-side tools (web search, web fetch, code execution, MCP, image generation, memory) and locally-executed common toolkits…
ag2ai/build-with-ag2
Persist agent state across runs, shape what the LLM sees per turn, and cap history to fit a context window.
ag2ai/build-with-ag2
Monitor an AG2 beta agent's stream — log events, detect repeated tool calls, track token spend, build trigger-driven observers, route observer alerts to the model, and halt on FATAL conditions.
ag2ai/build-with-ag2
Build a minimal AG2 beta Agent end to end — pick a model provider, set a prompt, call agent.ask(), then continue the conversation with reply.ask() (multi-turn).
Send images, audio, video, or documents into an AG2 beta Agent alongside text. Ag2 Multimodal Input is an agent skill from ag2ai/build-with-ag2. Send images, audio, video, or documents into an AG2 beta Agent alongside text.
Ag2 Multimodal Input fits situations like: the user wants the agent to process non-text input — describe a photo; transcribe audio; summarise a PDF; analyse a video.
Run `npx skills add ag2ai/build-with-ag2 --skill ag2-multimodal-input -a claude-code`. Or copy the skill folder (.agents/skills/ag2-multimodal-input in ag2ai/build-with-ag2) into .claude/skills/ag2-multimodal-input in your project. Claude Code loads it when a task matches its description.
Run `npx skills add ag2ai/build-with-ag2 --skill ag2-multimodal-input -a codex`. Or copy the skill folder (.agents/skills/ag2-multimodal-input in ag2ai/build-with-ag2) into .agents/skills/ag2-multimodal-input in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ag2ai/build-with-ag2 --skill ag2-multimodal-input -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ag2-multimodal-input, .gemini/skills/ag2-multimodal-input, .github/skills/ag2-multimodal-input and .opencode/skills/ag2-multimodal-input in your project.
SKILL.md names no scripts, command-line tools or credentials: Ag2 Multimodal Input is instructions for the agent only. Our summary lists: Python 3.
SKILL.md names 1 domain. In commands or code: youtube.com; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Ag2 Multimodal Input is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.7k tokens (SKILL.md is roughly 6.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Ag2 Multimodal Input: Markitdown (ImCa0/just-laws, 782 stars), Markitdown (jimmc414/Kosmos, 595 stars), Watchless (chenzixin1/watchless, 144 stars) and Lecture To Notes (ysyecust/lecture-to-notes, 273 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
ag2ai (a GitHub organization) maintains it in ag2ai/build-with-ag2, which has 252 GitHub stars. The repository holds 16 skills in this directory. The repository was last updated on September 6, 2026.
Source: ag2ai/build-with-ag2 on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.