Agent Builder
shareAI-lab/learn-claude-code
Design and build AI agents for any domain. An agent skill from shareAI-lab/learn-claude-code.
Ingest and process media files (video, audio, image). An agent skill from vellum-ai/vellum-assistant.
$ npx skills add vellum-ai/vellum-assistant --skill media-processing -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install vellum-ai/vellum-assistant media-processing --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/vellum-ai/vellum-assistant.git skills-src && mkdir -p .claude/skills && cp -r skills-src/assistant/src/config/bundled-skills/media-processing .claude/skills/media-processing && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "media-processing" agent skill from https://github.com/vellum-ai/vellum-assistant/tree/main/assistant/src/config/bundled-skills/media-processing into .claude/skills/media-processing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "media-processing", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/vellum-ai/vellum-assistant/tree/main/assistant/src/config/bundled-skills/media-processingType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add vellum-ai/vellum-assistant --skill media-processing -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install vellum-ai/vellum-assistant media-processing --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/vellum-ai/vellum-assistant.git skills-src && mkdir -p .agents/skills && cp -r skills-src/assistant/src/config/bundled-skills/media-processing .agents/skills/media-processing && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "media-processing" agent skill from https://github.com/vellum-ai/vellum-assistant/tree/main/assistant/src/config/bundled-skills/media-processing into .agents/skills/media-processing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "media-processing", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add vellum-ai/vellum-assistant --skill media-processing -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install vellum-ai/vellum-assistant media-processing --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/vellum-ai/vellum-assistant.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/assistant/src/config/bundled-skills/media-processing .cursor/skills/media-processing && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "media-processing" agent skill from https://github.com/vellum-ai/vellum-assistant/tree/main/assistant/src/config/bundled-skills/media-processing into .cursor/skills/media-processing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "media-processing", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/vellum-ai/vellum-assistant.git --path assistant/src/config/bundled-skills/media-processing--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add vellum-ai/vellum-assistant --skill media-processing -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install vellum-ai/vellum-assistant media-processing --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/vellum-ai/vellum-assistant.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/assistant/src/config/bundled-skills/media-processing .gemini/skills/media-processing && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "media-processing" agent skill from https://github.com/vellum-ai/vellum-assistant/tree/main/assistant/src/config/bundled-skills/media-processing into .gemini/skills/media-processing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "media-processing", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install vellum-ai/vellum-assistant media-processingInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add vellum-ai/vellum-assistant --skill media-processing -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/vellum-ai/vellum-assistant.git skills-src && mkdir -p .github/skills && cp -r skills-src/assistant/src/config/bundled-skills/media-processing .github/skills/media-processing && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "media-processing" agent skill from https://github.com/vellum-ai/vellum-assistant/tree/main/assistant/src/config/bundled-skills/media-processing into .github/skills/media-processing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "media-processing", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add vellum-ai/vellum-assistant --skill media-processing -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install vellum-ai/vellum-assistant media-processing --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/vellum-ai/vellum-assistant.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/assistant/src/config/bundled-skills/media-processing .opencode/skills/media-processing && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "media-processing" agent skill from https://github.com/vellum-ai/vellum-assistant/tree/main/assistant/src/config/bundled-skills/media-processing into .opencode/skills/media-processing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "media-processing", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
media-processingIngest and process media files (video, audio, image). An agent skill from vellum-ai/vellum-assistant.
Media Processing is an agent skill from vellum-ai/vellum-assistant. Ingest and process media files (video, audio, image)
Its SKILL.md is about 4.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 26 other files (for example `TOOLS.json`, `__tests__/audio-transcribe.test.ts` and `__tests__/concurrency-pool.test.ts`). Compatibility notes: Designed for Vellum personal assistants
It sits in AI & LLM Engineering. The repository describes itself as: An AI Assistant that’s easy to setup, does your work 24/7, knows your preferences and gets better over time. The licence is MIT.
5 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit c92ead1. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships script files (TypeScript, from the files we listed), which the agent can run.
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Designed for Vellum personal assistants
From compatibility in the SKILL.md frontmatter.
Media Processing loads about 4.2k tokens when it runs. Until then it costs about 17 tokens; SKILL.md has 1,846 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from vellum-ai/vellum-assistant at commit c92ead1, republished under its MIT licence (© vellum-ai). 1,846 words, ~4,152 tokens.
.claude/skills/media-processing/SKILL.md (or your agent's skills folder). This skill also uses 24 other files; get the full folder from GitHub.Ingest and track processing of media files (video, audio, images) through a configurable 3-phase pipeline.
The processing pipeline follows a sequential 3-phase flow:
ingest_media) - Register a media file, detect MIME type, extract duration, deduplicate by content hash.extract_keyframes) - Detect dead time, segment the video into windows, extract downscaled keyframes, build a subject registry, and write a pipeline manifest.analyze_keyframes) - Send each segment's frames to the configured Gemini vision model with assistant-provided extraction instructions and a JSON Schema for guaranteed structured output. Supports concurrency pooling, cost tracking, resumability, and automatic retries.query_media) - Send all map output to Claude for intelligent analysis and Q&A. Supports arbitrary natural language queries about video content.generate_clip) - Extract video clips around specific moments.The processing pipeline service (services/processing-pipeline.ts) orchestrates phases 2-4 automatically with retries, resumability, and cancellation support.
Register a media file for processing. Accepts an absolute file path, validates the file exists, detects MIME type, extracts duration (for video/audio via ffprobe), and registers the asset with content-hash deduplication.
Query the processing status of a media asset. Returns the asset metadata along with per-stage progress details. Use this to monitor pipeline progress.
Preprocess a video asset: detect dead time via mpdecimate, segment the video into windows, extract downscaled keyframes at regular intervals, build a subject registry, and write a pipeline manifest.
Parameters:
asset_id (required) - ID of the media asset.interval_seconds - Interval between keyframes (default: 1s). Use 0.5s for sports/action content where frame density matters.segment_duration - Duration of each segment window (default: 15s).dead_time_threshold - Sensitivity for dead-time detection (default: 0.02).section_config - Path to a JSON file with manual section boundaries.detect_dead_time - Whether to detect and skip dead time (default: false). Dead-time detection can be too aggressive for continuous action video like sports - it may incorrectly skip live play. Enable only for content with clear idle periods (e.g., lectures, surveillance footage).short_edge - Short edge resolution for downscaled frames in pixels (default: 480).include_audio - Whether to extract and transcribe audio for each segment (default: false). When enabled, each segment's audio is transcribed using the configured STT service and stored alongside visual frames.Map video segments through Gemini's structured output API. Supports two modes:
keyframes (default) - Reads frames from the preprocess manifest, sends each segment's images to Gemini. Requires extract_keyframes to be run first. Best for longer videos (> 1 hour) or when you need fine-grained control over frame selection (interval, segment duration, dead-time skipping).direct_video - Uploads the video file directly to Gemini's Files API. Gemini sees actual motion and temporal context instead of static frames. Best for shorter videos (< 1 hour) where temporal context matters (detecting actions, transitions, motion patterns). Has a 2 GB file size limit. Does not require extract_keyframes preprocessing.Both modes produce the same MapOutput format, so query_media works identically regardless of which mode was used.
Parameters:
asset_id (required) - ID of the media asset.system_prompt (required) - Extraction instructions for Gemini.output_schema (required) - JSON Schema for structured output.mode - Analysis mode: 'keyframes' (default) or 'direct_video'.context - Additional context to include in the prompt.model - Gemini model to use (defaults to the Gemini provider's recommended vision model).concurrency - Maximum concurrent API requests (default: 10, keyframes mode only).max_retries - Retry attempts per segment on failure (default: 3).Query video analysis data using natural language. Sends map output (from analyze_keyframes) to Claude for intelligent analysis and Q&A. Supports arbitrary questions about video content.
Parameters:
asset_id (required) - ID of the media asset.query (required) - Natural language query about the video data.system_prompt - Optional system prompt for Claude.model - LLM model to use. Defaults to the assistant's active chat model.Extract a video clip from a media asset using ffmpeg. Applies configurable pre/post-roll padding (clamped to file boundaries), outputs the clip as a temporary file.
Orchestrates the full processing pipeline with reliability features:
cancelled and the pipeline stops between stages.Handles dead-time detection, video segmentation, keyframe extraction, and subject registry building. Writes a pipeline manifest consumed by the Map phase.
Sends video segments to the configured Gemini vision model with structured output schemas. Handles concurrency pooling, cost tracking, resumability, and retries.
Sends Map output to Claude as text for analysis. Two modes:
Limits concurrent API calls during the Map phase to avoid rate limiting.
Tracks estimated API costs during pipeline execution.
When include_audio is enabled on extract_keyframes, the pipeline transcribes each segment's audio track using the configured STT service and attaches the transcript to the segment data. During the Map phase (analyze_keyframes), Gemini receives both the visual frames and the audio transcript for each segment, enabling multimodal analysis that combines what is seen with what is said.
This is useful for:
Audio transcription uses the STT service configured in assistant settings. If no STT service is configured or transcription fails for a segment (no audio track, service errors), the segment gracefully degrades to visual-only analysis.
The single most important insight: always use a broad, descriptive map prompt instead of a targeted one.
A targeted prompt like "find turnovers" locks you into one topic. If the user later wants to ask about defense, formations, or specific players, you'd need to reprocess the entire video. Instead, run a general-purpose descriptive prompt that captures everything visible, creating a rich, reusable dataset. Then all follow-up questions can be handled via query_media with no reprocessing.
One map run, many queries.
The map output will be larger (more tokens per segment), but Gemini Flash is cheap enough that this is a good tradeoff. Only use a targeted prompt if the user explicitly asks for something narrow.
Use this as a starting point for the system_prompt parameter in analyze_keyframes:
You are analyzing keyframes from a video. For each segment, describe everything you can observe:
- People visible: count, positions, identifying features (jersey numbers, clothing, names if visible)
- Actions and movements: what people are doing, direction of movement, interactions
- Objects of interest: ball location, equipment, vehicles, on-screen graphics
- Environment: setting, lighting, weather if outdoors
- Text on screen: scores, captions, titles, signs, timestamps
- Scene composition: camera angle, zoom level, any transitions between shots
- Any stoppages, pauses, or changes in activity
Be specific and factual. Describe what you see, not what you infer happened between frames.{
"type": "object",
"properties": {
"scene_description": { "type": "string" },
"people": {
"type": "array",
"items": {
"type": "object",
"properties": {
"description": { "type": "string" },
"position": { "type": "string" },
"action": { "type": "string" }
}
}
},
"objects_of_interest": { "type": "array", "items": { "type": "string" } },
"on_screen_text": { "type": "array", "items": { "type": "string" } },
"camera": { "type": "string" },
"notable_events": { "type": "array", "items": { "type": "string" } }
}
}The generate_clip tool automatically opens clips in the user's default video player after extraction (handled internally - do not run open via host_bash). Clips are saved persistently in the asset's pipeline directory (pipeline/<assetId>/clips/), falling back to a temp directory when the source location is read-only. Each clip gets a unique filename so concurrent or repeated extractions at the same range never collide. The clipPath field in the tool response contains the absolute file path.
The tool handles high-bitrate and incompatible codec sources automatically - it tries stream copy first for speed, then falls back to H.264 re-encoding if needed. Always use generate_clip rather than manual ffmpeg commands.
Always provide a descriptive title parameter (e.g. "snow-dive-closeup", "goal-celebration") so clips get meaningful filenames instead of timestamp-based names.
Gemini performs well at spatial/descriptive analysis from static keyframes:
Gemini hallucinates when asked to detect fast temporal events from static frames (keyframes mode), regardless of frame density:
The model is good at describing what is there but bad at detecting what happened from static frames. For content where temporal context matters, consider using mode: 'direct_video' which lets Gemini see actual motion. For keyframes mode, structure your map prompts and queries accordingly - ask the model to describe scenes, then use query_media (Claude) to reason about patterns and events across the descriptive data.
Use media_status to check the current state of any asset:
The response includes per-stage progress (0-100%) so you can see exactly where processing stands.
Use media_status to check processing stages:
stages array for any stage with status: "failed".lastError field for that stage to understand what went wrong.durationMs to see if a stage timed out or ran unusually long.After fixing the root cause, re-run the failed stage. The pipeline is resumable - it picks up from where it left off.
The Map phase (Gemini) is the primary cost driver - it scales with video duration and keyframe interval. The Q&A phase (Claude) is negligible per query.
ingest_media call processes one file. Batch ingestion is not yet supported.| Symptom | Likely Cause | Fix |
|---|---|---|
| "No keyframes found" | extract_keyframes not run or failed | Check preprocess stage status; re-run if needed |
| "No map output found" | analyze_keyframes not run | Run analyze_keyframes with appropriate system_prompt and output_schema |
| "No LLM provider available" | API key not configured | Add one in Settings |
| Map phase slow | Large video, small interval | Increase interval_seconds or reduce concurrency |
| Gemini returns errors | Rate limits or schema issues | Check max_retries setting; simplify output_schema if needed |
| Pipeline stuck at "processing" | Stage crashed without updating status | Use media_status to check stage progress; re-run manually |
ingest_media tool requires an absolute path to a local file.analyze_keyframes tool is marked as medium risk because it makes external API calls to Gemini, which incur costs.© vellum-ai, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 24 other files in assistant/src/config/bundled-skills/media-processing of vellum-ai/vellum-assistant.
Open the folder on GitHubat commit c92ead1
Media Processing next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Media Processing this skillvellum-ai/vellum-assistant | 1.4k | — | ~4.2k | Automated safety check: Pass | MIT | |
| Agent BuildershareAI-lab/learn-claude-code | 78k | 6 repos | ~1.2k | Automated safety check: Pass | MIT | |
| Add Uint Supportpytorch/pytorch | 104k | 2 repos | ~2.3k | Automated safety check: Pass | Custom licence | |
| Peft Fine TuningOrchestra-Research/AI-Research-SKILLs | 13k | 9 repos | ~3.1k | Automated safety check: Pass | MIT | |
| Segment Anything Model GuideOrchestra-Research/AI-Research-SKILLs | 13k | 9 repos | ~3.3k | Automated safety check: Pass | MIT | |
| 1passwordtrpc-group/trpc-agent-go | 1.8k | 15 repos | ~656 | Automated safety check: Pass | Apache-2.0 |
shareAI-lab/learn-claude-code
Design and build AI agents for any domain. An agent skill from shareAI-lab/learn-claude-code.
pytorch/pytorch
Add unsigned integer (uint) type support to PyTorch operators by updating ATDISPATCH macros.
Orchestra-Research/AI-Research-SKILLs
Parameter-efficient fine-tuning for LLMs using LoRA, QLoRA, and 25+ methods.
Orchestra-Research/AI-Research-SKILLs
Guide to using Meta's Segment Anything Model for zero-shot image segmentation with point, box or mask prompts, or automatic mask generation.
trpc-group/trpc-agent-go
Set up and use 1Password CLI (op). An agent skill from trpc-group/trpc-agent-go.
Orchestra-Research/AI-Research-SKILLs
Shows how to store documents and embeddings in Chroma, query them by similarity with metadata filters, and persist them to disk for RAG and semantic search projects.
vellum-ai/vellum-assistant
Create and configure a GitHub App so the assistant can push commits, open PRs, and comment under its own bot identity.
vellum-ai/vellum-assistant
Connect a Discord bot to the assistant via the Discord Gateway with guided application creation and intent configuration
vellum-ai/vellum-assistant
Create and configure a Sentry internal integration so the assistant can manage issues, alerts, and releases under its own identity
vellum-ai/vellum-assistant
Ingest a large dataset into memory as a skimmed map. An agent skill from vellum-ai/vellum-assistant.
vellum-ai/vellum-assistant
A skill your agent uses when the user wants to build, scaffold, ship, or edit a Vellum plugin that bundles multiple surfaces (hooks, tools, skills, and more) into one installable package.
vellum-ai/vellum-assistant
Connect a Slack app to the Vellum Assistant via Socket Mode.
Categories
Ingest and process media files (video, audio, image). An agent skill from vellum-ai/vellum-assistant. Media Processing is an agent skill from vellum-ai/vellum-assistant.
Media Processing fits situations like: AI & LLM Engineering work in your project.
Run `npx skills add vellum-ai/vellum-assistant --skill media-processing -a claude-code`. Or copy the skill folder (assistant/src/config/bundled-skills/media-processing in vellum-ai/vellum-assistant) into .claude/skills/media-processing in your project. Claude Code loads it when a task matches its description.
Run `npx skills add vellum-ai/vellum-assistant --skill media-processing -a codex`. Or copy the skill folder (assistant/src/config/bundled-skills/media-processing in vellum-ai/vellum-assistant) into .agents/skills/media-processing in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add vellum-ai/vellum-assistant --skill media-processing -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/media-processing, .gemini/skills/media-processing, .github/skills/media-processing and .opencode/skills/media-processing in your project.
Going by SKILL.md and its folder, Media Processing needs TypeScript for the scripts in its folder. Our summary lists: Node.js. Compatibility (from SKILL.md): Designed for Vellum personal assistants.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Media Processing is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 4.2k tokens (SKILL.md is roughly 17k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Media Processing: Agent Builder (shareAI-lab/learn-claude-code, 78k stars), Add Uint Support (pytorch/pytorch, 104k stars), Peft Fine Tuning (Orchestra-Research/AI-Research-SKILLs, 13k stars) and Segment Anything Model Guide (Orchestra-Research/AI-Research-SKILLs, 13k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
vellum-ai (a GitHub organization) maintains it in vellum-ai/vellum-assistant, which has 1,397 GitHub stars. The repository holds 108 skills in this directory. The repository was last updated on October 7, 2026.
Source: vellum-ai/vellum-assistant on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.