Agent skill

Multimodal Media Feature Extraction

by TyrealQ in TyrealQ/q-skills

Extracts pixel, video-frame, speech, music and visual-semantic features from image, video and audio files for research datasets, using local tools or the Gemini API.

MITAuto-check: notesData & Analytics

Install Multimodal Media Feature Extraction

skills CLI
$ npx skills add TyrealQ/q-skills --skill q-multimodal -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install TyrealQ/q-skills q-multimodal --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/TyrealQ/q-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/q-scholar/q-multimodal .claude/skills/q-multimodal && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
q-multimodal
GitHub stars
108
Token cost
~2k tokens
SKILL.md length
775 words
Files
27 (incl. scripts, references)
Skills in repo
16
Repo updated
First seen
Licence
MIT

At a glance

Extracts pixel, video-frame, speech, music and visual-semantic features from image, video and audio files for research datasets, using local tools or the Gemini API.

  • Extracting image or video-frame features for a media dataset
  • SKILL.md covers Setup (first time in a project), References, Dependencies and Pipelines, plus 2 more sections
  • Runs Python scripts from its folder; calls python; needs GOOGLE_API_KEY
  • Computing speech or music features from audio recordings

What it does

Two kinds of pipelines run over a folder of media. The local ones are driven from the command line and pull low-level features: pixel features with Pillow, video frames with FFmpeg and Pillow, speech and audio features with openSMILE, and music features with librosa. The Gemini ones do mid and high-level visual semantic analysis through the Gemini API, in either batch or standard mode.

On first use in a project the agent locates your input table, media directory, `.env` file with the Gemini API keys and system prompt, leaves them where they are, and copies only the scripts and an output folder into the project base directory. It scans the input columns and confirms with you which id columns and feature categories to extract. The Gemini batch scripts are numbered steps for uploading media, building JSONL, submitting jobs, checking status, retrying errors and reviewing results, and they read settings from a copy of `pipeline_config.py` passed with `--config`. Reference files cover checkpoints and rotating several API keys.

When your agent uses it

  • Extracting image or video-frame features for a media dataset
  • Computing speech or music features from audio recordings
  • Running large Gemini batch jobs over media files and retrying failed requests
  • Coding visual content semantically with Gemini across a whole dataset

Example prompts

  • “Extract pixel features from the images in ./media/posts and keep post_id from posts.xlsx.”
  • “Run openSMILE speech features on the interview recordings in ./audio.”
  • “Submit the Gemini batch job for the ad videos and show which requests failed.”
  • “Pull librosa music features from the tracks in ./playlists.”

Requirements

  • Python with Pillow, openSMILE and librosa for the local pipelines
  • FFmpeg for video frames
  • Google API keys for Gemini in a .env file

What it can do on your machine

Read from SKILL.md and the folder at commit d8aaee7. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 8 files in scripts/ (Python, from the files we listed), which the agent can run.

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • GOOGLE_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Multimodal Media Feature Extraction loads about 2k tokens when it runs, and up to ~12k if it reads all its reference files. Until then it costs about 63 tokens; SKILL.md has 775 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~63
When it runs · the whole SKILL.md, loaded when a task matches
~2k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~12k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteMentions a .env fileSKILL.md:12
    put/` — user assets (input data, media, `.env`, system prompt) can stay wherever they already live; point the scripts at
  • NoteMentions a .env fileSKILL.md:18
    - `.env` with `GOOGLE_API_KEY1`-`4` (Gemini only; check project root, home directory, common locations)
  • NoteMentions a .env fileSKILL.md:56
    gle-genai`, `python-dotenv` (+ above) | `.env` with `GOOGLE_API_KEY1`-`4` |

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from TyrealQ/q-skills at commit d8aaee7, republished under its MIT licence (© TyrealQ). 775 words, ~2,033 tokens.

Download SKILL.mdSave it as .claude/skills/q-multimodal/SKILL.md (or your agent's skills folder). This skill also uses 26 other files; get the full folder from GitHub.
name
q-multimodal
description
Extract visual, video, and audio features from media. Use for pixel features (Pillow), video frames (FFmpeg+Pillow), speech/audio features (openSMILE), music features (librosa), and visual semantic analysis (Gemini API batch or standard).

Q-Multimodal

Multimodal media analysis: local low-level features (Pillow, openSMILE), mid/high-level visual semantic analysis (Gemini API). Local pipelines are fully generic and CLI-driven. Gemini pipelines are config-driven: copy scripts/gemini/pipeline_config.py to your project, customize, and run with --config <path>.

Setup (first time in a project)

Do this once when adopting the skill in a new project. The canonical layout is a target only for scripts/ and output/ — user assets (input data, media, .env, system prompt) can stay wherever they already live; point the scripts at them via absolute paths.

  • Identify <BASE_DIR>: read the project's CLAUDE.md if it exists. If BASE_DIR isn't defined, ask the user which directory is the project root.
  • Locate existing user assets — do not move them. Search, confirm each location with the user before proceeding:
    • Input dataset file (xlsx/csv/json/parquet)
    • Media directory (grouping structure — one subfolder per subject is ideal)
    • .env with GOOGLE_API_KEY1-4 (Gemini only; check project root, home directory, common locations)
    • System prompt file (Gemini only)
  • Default: point at files in place. Set pipeline_config.py fields or CLI --input / --base-dir arguments to the absolute paths you found. Never move user data without explicit confirmation.
  • Materialize only scripts/ and output/ under <BASE_DIR>. Copy the pipelines actually being used from ${SKILL_DIR}/scripts/ into <BASE_DIR>/scripts/:
    • Local pipelines: pillow/, opensmile/, librosa/, common.py
    • Gemini pipelines: gemini/batch/, gemini/standard/, gemini/pipeline_config.py (template → adapt in place or copy to <BASE_DIR>/scripts/pipeline_config.py)
    • output/ is auto-created by scripts on first run
  • Scan input columns (adapt reader to file format): python -c "import pandas as pd; print(list(pd.read_<FORMAT>('INPUT', nrows=1).columns))"
  • Confirm --id-cols with the user. The file column (--file-col) is always retained in output — every row identifies its exact media asset — and --id-cols adds further columns (e.g. a post id) carried through checkpoints and merges.
  • Confirm --features with the user as well, so the feature categories a pipeline extracts match what the analysis needs.

References

Read the relevant reference file before executing a pipeline. These contain all flags, output column definitions, edge cases, and validation rules.

Local pipelines:

  • references/image-visual-features.md — all feature categories, column definitions, computation notes
  • references/video-visual-features.md — frame extraction, aggregation logic, dual output format
  • references/audio-features.md — openSMILE feature sets, interpretable scores, feature levels
  • references/music-features.md — librosa feature sets, tier-1 music scores, raw tonal/timbre block

Gemini pipelines:

  • references/gemini-batch-workflow.md — full 6-step batch pipeline, retry workflow, error handling
  • references/gemini-standard.md — standard pipeline details, model config, adapting for new projects
  • references/multi-key-management.md — multi-key quota strategy, retry threshold decision table

Shared:

  • references/checkpoint-format.md — column order, validation rules, output directory structure

Dependencies

PipelinePython packagesSystem
Image visualPillow, numpy, pandas, tqdm, openpyxl—
Video visual(same as image) + scenedetect[opencv]ffmpeg on PATH (for --extractor ffmpeg)
Audioopensmile, pandas, tqdm, openpyxlffmpeg + ffprobe on PATH (preflight-checked; both ship with any FFmpeg install)
Musiclibrosa, numpy, scipy, soundfile, pandas, tqdm, openpyxlffmpeg on PATH (compressed/video formats, via audioread)
Geminigoogle-genai, python-dotenv (+ above).env with GOOGLE_API_KEY1-4

Pipelines

Script path = ${SKILL_DIR}/scripts/<path>. Read the pipeline's reference file before running.

Show full SKILL.md (308 more words)Show less
Local Pipelines (generic, CLI-driven)
ScriptInputOutputReference
pillow/visual_features.pyImages47 pixel features (color, texture, spatial, quality)image-visual-features.md
pillow/video_features.pyVideosFrame-level + video-level aggregated features (scene-based extraction by default, FFmpeg fixed-interval optional)video-visual-features.md
opensmile/audio_features.pyVideo/audio8 interpretable scores + raw openSMILE features + stream/signal diagnostics (audio_status, configurable silence threshold)audio-features.md
librosa/music_features.pyAudio/video13 music-native scores + raw librosa featuresmusic-features.md

librosa/music_features.py complements opensmile/audio_features.py: openSMILE covers speech/prosody, librosa covers music-native features (tempo, key/mode, harmony, timbre).

Shared utilities: common.py — read_input(), save_excel(), derive_subject(), merge_checkpoints()

Command pattern: python <script> --input <file> --base-dir <root> [--features ...] [--id-cols ...] [--subjects ...] [--preview] [--merge]

After a run finishes, check the output for rows with ok=False and confirm the failure count is zero or has an understood cause before treating the extraction as complete.

Gemini Pipelines (config-driven)

Both pipelines read a pipeline_config.py file that defines paths, schema, metadata formatting, and validation rules. Copy scripts/gemini/pipeline_config.py to your project and customize.

Standard (gemini/standard/gemini_standard.py): inline media, 25 workers, auto-retry. See gemini-standard.md.

Batch (gemini/batch/[0-5]*.py + utils.py): 6-step pipeline, 50% discount. URIs expire after 48 hours. See gemini-batch-workflow.md.

bash
python 0uploadMedia.py --config /path/to/config.py --submit --max-batch-gb 2 --key 1
python 3checkStatus.py --config /path/to/config.py --poll
python 4retryErrors.py --config /path/to/config.py --preview
# >500 failures: batch retry
python 4retryErrors.py --config /path/to/config.py --submit
python 3checkStatus.py --config /path/to/config.py --poll
python 4retryErrors.py --config /path/to/config.py --collect
# <=500 failures or after batch retries: live fallback
python 4retryErrors.py --config /path/to/config.py --standard
python 5review.py --config /path/to/config.py --merge

Decision: >5 GB or >10 subjects and not time-sensitive → batch. Otherwise → standard. See gemini-standard.md.

Multi-key: Each GOOGLE_API_KEY{N} = 20 GB quota. See multi-key-management.md.

Adapting for New Projects

Local pipelines (Pillow, openSMILE, librosa): No modification needed. All project-specific values come from CLI args.

Gemini pipelines: Config-driven, no script modification needed. Scripts are copied to the project in step 2 above, then:

  • Adapt <BASE_DIR>/scripts/pipeline_config.py (already copied from template in step 2)
  • Set BASE_DIR, INPUT_PATH, SYSTEM_PROMPT_PATH to your project paths (SYSTEM_PROMPT_PATH relative to BASE_DIR, e.g., scripts/<prompt>.txt)
  • Set GROUP_COL, FILE_COL, ANALYSIS_FIELDS to match your input schema and system prompt
  • Implement subject_id() and format_metadata() for your domain
  • Optionally implement validate_row() for field-specific validation rules
  • Run any script with --config <BASE_DIR>/scripts/pipeline_config.py

Scope

Include: Image/video/audio feature extraction, Gemini visual semantic analysis, batch job management, checkpoint merging, multi-key quota management.

Exclude: ML model training, deep learning inference, real-time streaming analysis.

© TyrealQ, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 26 other files (scripts, references) in skills/q-scholar/q-multimodal of TyrealQ/q-skills.

  • SKILL.md
  • references/audio-features.md
  • references/checkpoint-format.md
  • references/gemini-batch-workflow.md
  • references/gemini-standard.md
  • references/image-visual-features.md
  • references/multi-key-management.md
  • references/music-features.md
  • references/video-visual-features.md
  • scripts/common.py
  • scripts/gemini/batch/0uploadMedia.py
  • scripts/gemini/batch/1buildJsonl.py
  • scripts/gemini/batch/2submitJobs.py
  • scripts/gemini/batch/3checkStatus.py
  • scripts/gemini/batch/4retryErrors.py
  • scripts/gemini/batch/5review.py
  • scripts/gemini/batch/utils.py
  • … and 10 more

Open the folder on GitHubat commit d8aaee7

Compare with similar skills

Multimodal Media Feature Extraction next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Multimodal Media Feature Extraction compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Multimodal Media Feature Extraction this skillTyrealQ/q-skills108—~2kAutomated safety check: NotesMIT
Transforming Dataancoleman/ai-design-components525—~3kAutomated safety check: PassMIT
Python Pipelinejamditis/claude-skills-journalism416—~4.8kAutomated safety check: PassMIT
Gemini Video Understandingeinverne/dotfiles121—~2.6kAutomated safety check: NotesMIT
Watch Video Q&Abradautomates/claude-video18k—~4.3kAutomated safety check: NotesMIT
ModLens Image Vision Bridgeliustack/modlens4.2k—~1.3kAutomated safety check: NotesMIT

Similar skills

  • Transforming Data

    ancoleman/ai-design-components

    Transform raw data into analytical assets using ETL/ELT patterns, SQL (dbt), Python (pandas/polars/PySpark), and orchestration (Airflow).

    525 GitHub stars~3k tokensUpdated 10 mo ago
    Data & AnalyticsAuto-check passed
  • Python Pipeline

    jamditis/claude-skills-journalism

    Python data pipelines with modular architecture. An agent skill from jamditis/claude-skills-journalism.

    416 GitHub stars~4.8k tokensUpdated 5 days ago
    Data & AnalyticsAuto-check passed
  • Analyze videos using Google's Gemini API - describe content, answer questions, transcribe audio with visual descriptions, reference timestamps, clip videos, and process YouTube URLs.

    121 GitHub stars~2.6k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check: notes
  • Watch Video Q&A

    bradautomates/claude-video

    Lets the agent answer questions about a video from a URL or local file by downloading it, extracting frames and a transcript, or by sending it to Gemini's video model.

    18k GitHub stars~4.3k tokensUpdated 15 days ago
    Media & CreativeAuto-check: notes
  • Gives text-only models sight by running the modlens CLI on an image path or URL and returning structured JSON evidence with transcribed text, layout and semantics.

    4.2k GitHub stars~1.3k tokensUpdated 6 days ago
    AI & LLM EngineeringAuto-check: notes
  • Gemini API Dev

    google-gemini/gemini-skills

    Official

    A skill your agent uses when writing code that calls the Gemini API for text generation, multi-turn chat, multimodal understanding, image generation, video generation, speech generation (TTS), voice…

    4.3k GitHub stars~5.1k tokensUpdated 3 days ago
    AI & LLM EngineeringAuto-check passed

More from TyrealQ/q-skills

All 16 skills in this repo
  • Q-Infographics

    TyrealQ/q-skills

    Converts a report or other document into a business story and then an infographic image, pausing for your review after each step.

    108 GitHub stars~814 tokensUpdated 16 days ago
    Auto-check: notes
  • Runs exploratory data analysis on tabular data after you confirm each column's measurement level, then writes CSV tables and a narrative summary.

    108 GitHub stars~1.1k tokensUpdated 16 days ago
    Auto-check passed
  • Consolidates BERTopic, LDA or NMF topic output into a theory-driven classification framework and writes the final labels back to an Excel file.

    108 GitHub stars~1k tokensUpdated 16 days ago
    Auto-check passed
  • Generates branded slide deck images from written content, with a content analysis step, a layout catalog and scripts that merge the slides into PowerPoint or PDF.

    108 GitHub stars~1.1k tokensUpdated 16 days ago
    Auto-check passed
  • Audits a repository's file layout and project documentation against a written convention file, then proposes moves, deletions and doc fixes as an approved plan before touching anything.

    108 GitHub stars~4.4k tokensUpdated 16 days ago
    Auto-check: notes
  • Commit

    TyrealQ/q-skills

    Stage and commit uncommitted changes with conventional commit messages.

    108 GitHub stars~1.1k tokensUpdated 16 days ago
    Auto-check: notes

Questions about Multimodal Media Feature Extraction

What does Multimodal Media Feature Extraction do?

Extracts pixel, video-frame, speech, music and visual-semantic features from image, video and audio files for research datasets, using local tools or the Gemini API. Two kinds of pipelines run over a folder of media. The local ones are driven from the command line and pull low-level features: pixel features with Pillow, video frames with FFmpeg and Pillow, speech and audio features with openSMILE, and music features with librosa.

When should I use Multimodal Media Feature Extraction?

Multimodal Media Feature Extraction fits situations like: extracting image or video-frame features for a media dataset; computing speech or music features from audio recordings; running large Gemini batch jobs over media files and retrying failed requests; coding visual content semantically with Gemini across a whole dataset.

How do I install Multimodal Media Feature Extraction in Claude Code?

Run `npx skills add TyrealQ/q-skills --skill q-multimodal -a claude-code`. Or copy the skill folder (skills/q-scholar/q-multimodal in TyrealQ/q-skills) into .claude/skills/q-multimodal in your project. Claude Code loads it when a task matches its description.

How do I install Multimodal Media Feature Extraction in Codex?

Run `npx skills add TyrealQ/q-skills --skill q-multimodal -a codex`. Or copy the skill folder (skills/q-scholar/q-multimodal in TyrealQ/q-skills) into .agents/skills/q-multimodal in your project. Codex loads it when a task matches its description.

Can I use Multimodal Media Feature Extraction in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add TyrealQ/q-skills --skill q-multimodal -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/q-multimodal, .gemini/skills/q-multimodal, .github/skills/q-multimodal and .opencode/skills/q-multimodal in your project.

What does Multimodal Media Feature Extraction need to run?

Going by SKILL.md and its folder, Multimodal Media Feature Extraction needs Python for the scripts in its folder, the command-line tools its instructions call (python) and credentials named GOOGLE_API_KEY. Our summary lists: Python with Pillow, openSMILE and librosa for the local pipelines; FFmpeg for video frames; Google API keys for Gemini in a .env file.

Does Multimodal Media Feature Extraction access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Multimodal Media Feature Extraction safe to install?

Our automated static check of SKILL.md found notes only (mentions a .env file), nothing it rates as a warning. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Multimodal Media Feature Extraction use?

Multimodal Media Feature Extraction is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Multimodal Media Feature Extraction use?

About 2k tokens (SKILL.md is roughly 8.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 9.8k tokens, read only when the agent opens those files.

What are the alternatives to Multimodal Media Feature Extraction?

Skills that share tags, products or a category with Multimodal Media Feature Extraction: Transforming Data (ancoleman/ai-design-components, 525 stars), Python Pipeline (jamditis/claude-skills-journalism, 416 stars), Gemini Video Understanding (einverne/dotfiles, 121 stars) and Watch Video Q&A (bradautomates/claude-video, 18k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Multimodal Media Feature Extraction?

TyrealQ (a GitHub user) maintains it in TyrealQ/q-skills, which has 108 GitHub stars. The repository holds 16 skills in this directory. The repository was last updated on September 23, 2026.

Source: TyrealQ/q-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.