Agent skill

Media Processing

by vellum-ai in vellum-ai/vellum-assistant

Ingest and process media files (video, audio, image). An agent skill from vellum-ai/vellum-assistant.

MITAuto-check passedAI & LLM Engineering

Install Media Processing

skills CLI
$ npx skills add vellum-ai/vellum-assistant --skill media-processing -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install vellum-ai/vellum-assistant media-processing --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/vellum-ai/vellum-assistant.git skills-src && mkdir -p .claude/skills && cp -r skills-src/assistant/src/config/bundled-skills/media-processing .claude/skills/media-processing && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
media-processing
GitHub stars
1.4k
Token cost
~4.2k tokens
SKILL.md length
1,846 words
Files
25
Skills in repo
108
Repo updated
First seen
Licence
MIT

At a glance

Ingest and process media files (video, audio, image). An agent skill from vellum-ai/vellum-assistant.

  • Works in 5 steps: Ingest (ingest_media) - Register a media… → Preprocess (extract_keyframes) - Detect… → Map (analyze_keyframes) - Send each… → …
  • AI & LLM Engineering work in your project
  • SKILL.md covers End-to-End Workflow, Tools, Services and Audio + Vision Multimodal…, plus 4 more sections
  • Runs TypeScript scripts from its folder

What it does

Media Processing is an agent skill from vellum-ai/vellum-assistant. Ingest and process media files (video, audio, image)

Its SKILL.md is about 4.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 26 other files (for example `TOOLS.json`, `__tests__/audio-transcribe.test.ts` and `__tests__/concurrency-pool.test.ts`). Compatibility notes: Designed for Vellum personal assistants

It sits in AI & LLM Engineering. The repository describes itself as: An AI Assistant that’s easy to setup, does your work 24/7, knows your preferences and gets better over time. The licence is MIT.

When your agent uses it

  • AI & LLM Engineering work in your project

Example prompts

  • “/media-processing”

Requirements

  • Node.js
  • Compatibility (from SKILL.md): Designed for Vellum personal assistants

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Ingest (ingest_media) - Register a media file, detect MIME type, extract duration, deduplicate by content hash.
  2. Preprocess (extract_keyframes) - Detect dead time, segment the video into windows, extract downscaled keyframes, build a subject registry…
  3. Map (analyze_keyframes) - Send each segment's frames to the configured Gemini vision model with assistant-provided extraction instructions…
  4. Reduce / Query (query_media) - Send all map output to Claude for intelligent analysis and Q&A. Supports arbitrary natural language queries…
  5. Clip (generate_clip) - Extract video clips around specific moments.

What it can do on your machine

Read from SKILL.md and the folder at commit c92ead1. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (TypeScript, from the files we listed), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Designed for Vellum personal assistants

    From compatibility in the SKILL.md frontmatter.

Context cost

Media Processing loads about 4.2k tokens when it runs. Until then it costs about 17 tokens; SKILL.md has 1,846 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~17
When it runs · the whole SKILL.md, loaded when a task matches
~4.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from vellum-ai/vellum-assistant at commit c92ead1, republished under its MIT licence (© vellum-ai). 1,846 words, ~4,152 tokens.

Download SKILL.mdSave it as .claude/skills/media-processing/SKILL.md (or your agent's skills folder). This skill also uses 24 other files; get the full folder from GitHub.
name
media-processing
description
Ingest and process media files (video, audio, image)
compatibility
Designed for Vellum personal assistants
metadata.emoji
🎬

Ingest and track processing of media files (video, audio, images) through a configurable 3-phase pipeline.

End-to-End Workflow

The processing pipeline follows a sequential 3-phase flow:

  1. Ingest (ingest_media) - Register a media file, detect MIME type, extract duration, deduplicate by content hash.
  2. Preprocess (extract_keyframes) - Detect dead time, segment the video into windows, extract downscaled keyframes, build a subject registry, and write a pipeline manifest.
  3. Map (analyze_keyframes) - Send each segment's frames to the configured Gemini vision model with assistant-provided extraction instructions and a JSON Schema for guaranteed structured output. Supports concurrency pooling, cost tracking, resumability, and automatic retries.
  4. Reduce / Query (query_media) - Send all map output to Claude for intelligent analysis and Q&A. Supports arbitrary natural language queries about video content.
  5. Clip (generate_clip) - Extract video clips around specific moments.

The processing pipeline service (services/processing-pipeline.ts) orchestrates phases 2-4 automatically with retries, resumability, and cancellation support.

Tools

ingest_media

Register a media file for processing. Accepts an absolute file path, validates the file exists, detects MIME type, extracts duration (for video/audio via ffprobe), and registers the asset with content-hash deduplication.

media_status

Query the processing status of a media asset. Returns the asset metadata along with per-stage progress details. Use this to monitor pipeline progress.

extract_keyframes

Preprocess a video asset: detect dead time via mpdecimate, segment the video into windows, extract downscaled keyframes at regular intervals, build a subject registry, and write a pipeline manifest.

Parameters:

  • asset_id (required) - ID of the media asset.
  • interval_seconds - Interval between keyframes (default: 1s). Use 0.5s for sports/action content where frame density matters.
  • segment_duration - Duration of each segment window (default: 15s).
  • dead_time_threshold - Sensitivity for dead-time detection (default: 0.02).
  • section_config - Path to a JSON file with manual section boundaries.
  • detect_dead_time - Whether to detect and skip dead time (default: false). Dead-time detection can be too aggressive for continuous action video like sports - it may incorrectly skip live play. Enable only for content with clear idle periods (e.g., lectures, surveillance footage).
  • short_edge - Short edge resolution for downscaled frames in pixels (default: 480).
  • include_audio - Whether to extract and transcribe audio for each segment (default: false). When enabled, each segment's audio is transcribed using the configured STT service and stored alongside visual frames.
analyze_keyframes

Map video segments through Gemini's structured output API. Supports two modes:

  • keyframes (default) - Reads frames from the preprocess manifest, sends each segment's images to Gemini. Requires extract_keyframes to be run first. Best for longer videos (> 1 hour) or when you need fine-grained control over frame selection (interval, segment duration, dead-time skipping).
  • direct_video - Uploads the video file directly to Gemini's Files API. Gemini sees actual motion and temporal context instead of static frames. Best for shorter videos (< 1 hour) where temporal context matters (detecting actions, transitions, motion patterns). Has a 2 GB file size limit. Does not require extract_keyframes preprocessing.

Both modes produce the same MapOutput format, so query_media works identically regardless of which mode was used.

Parameters:

  • asset_id (required) - ID of the media asset.
  • system_prompt (required) - Extraction instructions for Gemini.
  • output_schema (required) - JSON Schema for structured output.
  • mode - Analysis mode: 'keyframes' (default) or 'direct_video'.
  • context - Additional context to include in the prompt.
  • model - Gemini model to use (defaults to the Gemini provider's recommended vision model).
  • concurrency - Maximum concurrent API requests (default: 10, keyframes mode only).
  • max_retries - Retry attempts per segment on failure (default: 3).
query_media

Query video analysis data using natural language. Sends map output (from analyze_keyframes) to Claude for intelligent analysis and Q&A. Supports arbitrary questions about video content.

Parameters:

  • asset_id (required) - ID of the media asset.
  • query (required) - Natural language query about the video data.
  • system_prompt - Optional system prompt for Claude.
  • model - LLM model to use. Defaults to the assistant's active chat model.
generate_clip

Extract a video clip from a media asset using ffmpeg. Applies configurable pre/post-roll padding (clamped to file boundaries), outputs the clip as a temporary file.

Services

Processing Pipeline (services/processing-pipeline.ts)

Orchestrates the full processing pipeline with reliability features:

  • Sequential execution: preprocess, map, reduce.
  • Retries: Each stage is retried with exponential backoff and jitter (configurable max retries and base delay).
  • Resumability: Checks processing_stages to find the last completed stage and resumes from there. Safe to restart after crashes.
  • Cancellation: Cooperative cancellation via asset status. Set asset status to cancelled and the pipeline stops between stages.
  • Idempotency: Re-ingesting the same file hash is a no-op. Re-running a fully completed pipeline is also a no-op.
  • Graceful degradation: If a stage fails mid-batch (e.g., Gemini API errors), partial results are saved. The stage is marked as failed with the error details, and the pipeline stops without losing work.
Preprocess (services/preprocess.ts)

Handles dead-time detection, video segmentation, keyframe extraction, and subject registry building. Writes a pipeline manifest consumed by the Map phase.

Gemini Map (services/gemini-map.ts)

Sends video segments to the configured Gemini vision model with structured output schemas. Handles concurrency pooling, cost tracking, resumability, and retries.

Reduce (services/reduce.ts)

Sends Map output to Claude as text for analysis. Two modes:

  • One-shot merge: assembles all Map results and sends to Claude with a system prompt.
  • Interactive Q&A: loads existing map output + user query, sends to Claude.
Concurrency Pool (services/concurrency-pool.ts)

Limits concurrent API calls during the Map phase to avoid rate limiting.

Cost Tracker (services/cost-tracker.ts)

Tracks estimated API costs during pipeline execution.

Audio + Vision Multimodal Analysis

When include_audio is enabled on extract_keyframes, the pipeline transcribes each segment's audio track using the configured STT service and attaches the transcript to the segment data. During the Map phase (analyze_keyframes), Gemini receives both the visual frames and the audio transcript for each segment, enabling multimodal analysis that combines what is seen with what is said.

This is useful for:

  • Lectures and presentations: Correlate slide content (visual) with speaker narration (audio).
  • Sports broadcasts: Combine on-screen action with commentary for richer event detection.
  • Meetings and interviews: Pair facial expressions and gestures with spoken dialogue.
  • Tutorials and demos: Link on-screen actions with verbal instructions.

Audio transcription uses the STT service configured in assistant settings. If no STT service is configured or transcription fails for a segment (no audio track, service errors), the segment gracefully degrades to visual-only analysis.

Best Practices

Map Prompt Strategy: Go Broad, Not Targeted

The single most important insight: always use a broad, descriptive map prompt instead of a targeted one.

A targeted prompt like "find turnovers" locks you into one topic. If the user later wants to ask about defense, formations, or specific players, you'd need to reprocess the entire video. Instead, run a general-purpose descriptive prompt that captures everything visible, creating a rich, reusable dataset. Then all follow-up questions can be handled via query_media with no reprocessing.

One map run, many queries.

The map output will be larger (more tokens per segment), but Gemini Flash is cheap enough that this is a good tradeoff. Only use a targeted prompt if the user explicitly asks for something narrow.

Show full SKILL.md (723 more words)Show less
Sample General-Purpose Map Prompt

Use this as a starting point for the system_prompt parameter in analyze_keyframes:

You are analyzing keyframes from a video. For each segment, describe everything you can observe:

- People visible: count, positions, identifying features (jersey numbers, clothing, names if visible)
- Actions and movements: what people are doing, direction of movement, interactions
- Objects of interest: ball location, equipment, vehicles, on-screen graphics
- Environment: setting, lighting, weather if outdoors
- Text on screen: scores, captions, titles, signs, timestamps
- Scene composition: camera angle, zoom level, any transitions between shots
- Any stoppages, pauses, or changes in activity

Be specific and factual. Describe what you see, not what you infer happened between frames.
Sample Output Schema
json
{
  "type": "object",
  "properties": {
    "scene_description": { "type": "string" },
    "people": {
      "type": "array",
      "items": {
        "type": "object",
        "properties": {
          "description": { "type": "string" },
          "position": { "type": "string" },
          "action": { "type": "string" }
        }
      }
    },
    "objects_of_interest": { "type": "array", "items": { "type": "string" } },
    "on_screen_text": { "type": "array", "items": { "type": "string" } },
    "camera": { "type": "string" },
    "notable_events": { "type": "array", "items": { "type": "string" } }
  }
}
Clip Delivery

The generate_clip tool automatically opens clips in the user's default video player after extraction (handled internally - do not run open via host_bash). Clips are saved persistently in the asset's pipeline directory (pipeline/<assetId>/clips/), falling back to a temp directory when the source location is read-only. Each clip gets a unique filename so concurrent or repeated extractions at the same range never collide. The clipPath field in the tool response contains the absolute file path.

The tool handles high-bitrate and incompatible codec sources automatically - it tries stream copy first for speed, then falls back to H.264 re-encoding if needed. Always use generate_clip rather than manual ffmpeg commands.

Always provide a descriptive title parameter (e.g. "snow-dive-closeup", "goal-celebration") so clips get meaningful filenames instead of timestamp-based names.

Known Limitations - Vision Analysis

Gemini performs well at spatial/descriptive analysis from static keyframes:

  • Player positions, formations, and spacing
  • Jersey numbers and identifying features
  • Ball location and which team has possession
  • Score and on-screen text
  • Camera angles and scene composition

Gemini hallucinates when asked to detect fast temporal events from static frames (keyframes mode), regardless of frame density:

  • Turnovers, steals, fouls, and specific plays
  • Fast transitions and split-second actions
  • Causality between frames (what "happened" vs. what's visible)

The model is good at describing what is there but bad at detecting what happened from static frames. For content where temporal context matters, consider using mode: 'direct_video' which lets Gemini see actual motion. For keyframes mode, structure your map prompts and queries accordingly - ask the model to describe scenes, then use query_media (Claude) to reason about patterns and events across the descriptive data.

Operator Runbook

Monitoring Progress

Use media_status to check the current state of any asset:

  • registered - Ingested but not yet processed.
  • processing - Pipeline is running.
  • indexed - All stages completed successfully.
  • failed - A stage failed. Check stage details for the error.

The response includes per-stage progress (0-100%) so you can see exactly where processing stands.

Diagnosing Failures

Use media_status to check processing stages:

  1. Check the stages array for any stage with status: "failed".
  2. Read the lastError field for that stage to understand what went wrong.
  3. Check durationMs to see if a stage timed out or ran unusually long.
  4. Common failure causes:
    • preprocess: ffmpeg not installed, corrupt video file, disk full.
    • map: Gemini API key not configured, API rate limits, network errors.
    • reduce: No LLM provider configured, no map output exists.

After fixing the root cause, re-run the failed stage. The pipeline is resumable - it picks up from where it left off.

Cost Expectations

The Map phase (Gemini) is the primary cost driver - it scales with video duration and keyframe interval. The Q&A phase (Claude) is negligible per query.

Known Limitations
  • ffmpeg required: Keyframe extraction and clip generation require ffmpeg to be installed on the host.
  • Single-file ingestion: Each ingest_media call processes one file. Batch ingestion is not yet supported.
  • Gemini rate limits: The Map phase uses concurrency pooling (default 10) to stay within API limits. Reduce concurrency if you hit 429 errors.
  • No real-time processing: The pipeline processes pre-recorded media files. Live/streaming video is not supported.
Troubleshooting
SymptomLikely CauseFix
"No keyframes found"extract_keyframes not run or failedCheck preprocess stage status; re-run if needed
"No map output found"analyze_keyframes not runRun analyze_keyframes with appropriate system_prompt and output_schema
"No LLM provider available"API key not configuredAdd one in Settings
Map phase slowLarge video, small intervalIncrease interval_seconds or reduce concurrency
Gemini returns errorsRate limits or schema issuesCheck max_retries setting; simplify output_schema if needed
Pipeline stuck at "processing"Stage crashed without updating statusUse media_status to check stage progress; re-run manually

Usage Notes

  • The ingest_media tool requires an absolute path to a local file.
  • Supported media types: video (mp4, mov, avi, mkv, webm, etc.), audio (mp3, wav, m4a, etc.), and images (png, jpg, gif, webp, etc.).
  • For video and audio files, duration is automatically extracted via ffprobe (requires ffmpeg to be installed).
  • Duplicate files are detected by content hash and return the existing asset record.
  • The analyze_keyframes tool is marked as medium risk because it makes external API calls to Gemini, which incur costs.
  • All schema tables, services, and tool interfaces are media-generic. Domain-specific interpretation belongs in the system_prompt and output_schema parameters.

© vellum-ai, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 24 other files in assistant/src/config/bundled-skills/media-processing of vellum-ai/vellum-assistant.

  • SKILL.md
  • TOOLS.json
  • __tests__/audio-transcribe.test.ts
  • __tests__/concurrency-pool.test.ts
  • __tests__/cost-tracker.test.ts
  • __tests__/extract-keyframes.test.ts
  • __tests__/media-analysis-default-model.test.ts
  • __tests__/preprocess-audio.test.ts
  • __tests__/preprocess.test.ts
  • services/audio-transcribe.ts
  • services/concurrency-pool.ts
  • services/cost-tracker.ts
  • services/gemini-map.ts
  • services/gemini-video.ts
  • services/media-analysis-model.ts
  • services/preprocess.ts
  • services/processing-pipeline.ts
  • services/reduce.ts
  • tools
  • … and 6 more

Open the folder on GitHubat commit c92ead1

Compare with similar skills

Media Processing next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Media Processing compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Media Processing this skillvellum-ai/vellum-assistant1.4k—~4.2kAutomated safety check: PassMIT
Agent BuildershareAI-lab/learn-claude-code78k6 repos~1.2kAutomated safety check: PassMIT
Add Uint Supportpytorch/pytorch104k2 repos~2.3kAutomated safety check: PassCustom licence
Peft Fine TuningOrchestra-Research/AI-Research-SKILLs13k9 repos~3.1kAutomated safety check: PassMIT
Segment Anything Model GuideOrchestra-Research/AI-Research-SKILLs13k9 repos~3.3kAutomated safety check: PassMIT
1passwordtrpc-group/trpc-agent-go1.8k15 repos~656Automated safety check: PassApache-2.0

Similar skills

  • Agent Builder

    shareAI-lab/learn-claude-code

    Design and build AI agents for any domain. An agent skill from shareAI-lab/learn-claude-code.

    78k GitHub starsUsed in 6 repos~1.2k tokens
    AI & LLM EngineeringAuto-check passed
  • Add Uint Support

    pytorch/pytorch

    Add unsigned integer (uint) type support to PyTorch operators by updating ATDISPATCH macros.

    104k GitHub starsUsed in 2 repos~2.3k tokens
    AI & LLM EngineeringAuto-check passed
  • Peft Fine Tuning

    Orchestra-Research/AI-Research-SKILLs

    Parameter-efficient fine-tuning for LLMs using LoRA, QLoRA, and 25+ methods.

    13k GitHub starsUsed in 9 repos~3.1k tokens
    AI & LLM EngineeringAuto-check passed
  • Segment Anything Model Guide

    Orchestra-Research/AI-Research-SKILLs

    Guide to using Meta's Segment Anything Model for zero-shot image segmentation with point, box or mask prompts, or automatic mask generation.

    13k GitHub starsUsed in 9 repos~3.3k tokens
    AI & LLM EngineeringAuto-check passed
  • 1password

    trpc-group/trpc-agent-go

    Set up and use 1Password CLI (op). An agent skill from trpc-group/trpc-agent-go.

    1.8k GitHub starsUsed in 15 repos~656 tokens
    AI & LLM EngineeringAuto-check passed
  • Chroma Vector Database

    Orchestra-Research/AI-Research-SKILLs

    Shows how to store documents and embeddings in Chroma, query them by similarity with metadata filters, and persist them to disk for RAG and semantic search projects.

    13k GitHub starsUsed in 8 repos~2.3k tokens
    AI & LLM EngineeringAuto-check passed

More from vellum-ai/vellum-assistant

All 108 skills in this repo
  • Vellum GitHub App Setup

    vellum-ai/vellum-assistant

    Create and configure a GitHub App so the assistant can push commits, open PRs, and comment under its own bot identity.

    1.4k GitHub stars~3.1k tokensUpdated yesterday
    Auto-check passed
  • Discord App Setup

    vellum-ai/vellum-assistant

    Connect a Discord bot to the assistant via the Discord Gateway with guided application creation and intent configuration

    1.4k GitHub stars~4.2k tokensUpdated yesterday
    Auto-check passed
  • Sentry App Setup

    vellum-ai/vellum-assistant

    Create and configure a Sentry internal integration so the assistant can manage issues, alerts, and releases under its own identity

    1.4k GitHub stars~1.3k tokensUpdated yesterday
    Auto-check passed
  • Memory Corpus Ingest

    vellum-ai/vellum-assistant

    Ingest a large dataset into memory as a skimmed map. An agent skill from vellum-ai/vellum-assistant.

    1.4k GitHub stars~3k tokensUpdated yesterday
    Auto-check: notes
  • Plugin Builder

    vellum-ai/vellum-assistant

    A skill your agent uses when the user wants to build, scaffold, ship, or edit a Vellum plugin that bundles multiple surfaces (hooks, tools, skills, and more) into one installable package.

    1.4k GitHub stars~3.1k tokensUpdated yesterday
    Auto-check passed
  • Slack App Setup

    vellum-ai/vellum-assistant

    Connect a Slack app to the Vellum Assistant via Socket Mode.

    1.4k GitHub stars~2.5k tokensUpdated yesterday
    Auto-check: warnings

Questions about Media Processing

What does Media Processing do?

Ingest and process media files (video, audio, image). An agent skill from vellum-ai/vellum-assistant. Media Processing is an agent skill from vellum-ai/vellum-assistant.

When should I use Media Processing?

Media Processing fits situations like: AI & LLM Engineering work in your project.

How do I install Media Processing in Claude Code?

Run `npx skills add vellum-ai/vellum-assistant --skill media-processing -a claude-code`. Or copy the skill folder (assistant/src/config/bundled-skills/media-processing in vellum-ai/vellum-assistant) into .claude/skills/media-processing in your project. Claude Code loads it when a task matches its description.

How do I install Media Processing in Codex?

Run `npx skills add vellum-ai/vellum-assistant --skill media-processing -a codex`. Or copy the skill folder (assistant/src/config/bundled-skills/media-processing in vellum-ai/vellum-assistant) into .agents/skills/media-processing in your project. Codex loads it when a task matches its description.

Can I use Media Processing in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add vellum-ai/vellum-assistant --skill media-processing -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/media-processing, .gemini/skills/media-processing, .github/skills/media-processing and .opencode/skills/media-processing in your project.

What does Media Processing need to run?

Going by SKILL.md and its folder, Media Processing needs TypeScript for the scripts in its folder. Our summary lists: Node.js. Compatibility (from SKILL.md): Designed for Vellum personal assistants.

Does Media Processing access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Media Processing safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Media Processing use?

Media Processing is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Media Processing use?

About 4.2k tokens (SKILL.md is roughly 17k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Media Processing?

Skills that share tags, products or a category with Media Processing: Agent Builder (shareAI-lab/learn-claude-code, 78k stars), Add Uint Support (pytorch/pytorch, 104k stars), Peft Fine Tuning (Orchestra-Research/AI-Research-SKILLs, 13k stars) and Segment Anything Model Guide (Orchestra-Research/AI-Research-SKILLs, 13k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Media Processing?

vellum-ai (a GitHub organization) maintains it in vellum-ai/vellum-assistant, which has 1,397 GitHub stars. The repository holds 108 skills in this directory. The repository was last updated on October 7, 2026.

Source: vellum-ai/vellum-assistant on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.