Transforming Data
ancoleman/ai-design-components
Transform raw data into analytical assets using ETL/ELT patterns, SQL (dbt), Python (pandas/polars/PySpark), and orchestration (Airflow).
Extracts pixel, video-frame, speech, music and visual-semantic features from image, video and audio files for research datasets, using local tools or the Gemini API.
$ npx skills add TyrealQ/q-skills --skill q-multimodal -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install TyrealQ/q-skills q-multimodal --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/TyrealQ/q-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/q-scholar/q-multimodal .claude/skills/q-multimodal && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "q-multimodal" agent skill from https://github.com/TyrealQ/q-skills/tree/main/skills/q-scholar/q-multimodal into .claude/skills/q-multimodal/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "q-multimodal", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/TyrealQ/q-skills/tree/main/skills/q-scholar/q-multimodalType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add TyrealQ/q-skills --skill q-multimodal -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install TyrealQ/q-skills q-multimodal --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/TyrealQ/q-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/q-scholar/q-multimodal .agents/skills/q-multimodal && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "q-multimodal" agent skill from https://github.com/TyrealQ/q-skills/tree/main/skills/q-scholar/q-multimodal into .agents/skills/q-multimodal/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "q-multimodal", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add TyrealQ/q-skills --skill q-multimodal -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install TyrealQ/q-skills q-multimodal --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/TyrealQ/q-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/q-scholar/q-multimodal .cursor/skills/q-multimodal && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "q-multimodal" agent skill from https://github.com/TyrealQ/q-skills/tree/main/skills/q-scholar/q-multimodal into .cursor/skills/q-multimodal/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "q-multimodal", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/TyrealQ/q-skills.git --path skills/q-scholar/q-multimodal--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add TyrealQ/q-skills --skill q-multimodal -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install TyrealQ/q-skills q-multimodal --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/TyrealQ/q-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/q-scholar/q-multimodal .gemini/skills/q-multimodal && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "q-multimodal" agent skill from https://github.com/TyrealQ/q-skills/tree/main/skills/q-scholar/q-multimodal into .gemini/skills/q-multimodal/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "q-multimodal", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install TyrealQ/q-skills q-multimodalInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add TyrealQ/q-skills --skill q-multimodal -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/TyrealQ/q-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/q-scholar/q-multimodal .github/skills/q-multimodal && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "q-multimodal" agent skill from https://github.com/TyrealQ/q-skills/tree/main/skills/q-scholar/q-multimodal into .github/skills/q-multimodal/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "q-multimodal", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add TyrealQ/q-skills --skill q-multimodal -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install TyrealQ/q-skills q-multimodal --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/TyrealQ/q-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/q-scholar/q-multimodal .opencode/skills/q-multimodal && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "q-multimodal" agent skill from https://github.com/TyrealQ/q-skills/tree/main/skills/q-scholar/q-multimodal into .opencode/skills/q-multimodal/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "q-multimodal", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
q-multimodalExtracts pixel, video-frame, speech, music and visual-semantic features from image, video and audio files for research datasets, using local tools or the Gemini API.
Two kinds of pipelines run over a folder of media. The local ones are driven from the command line and pull low-level features: pixel features with Pillow, video frames with FFmpeg and Pillow, speech and audio features with openSMILE, and music features with librosa. The Gemini ones do mid and high-level visual semantic analysis through the Gemini API, in either batch or standard mode.
On first use in a project the agent locates your input table, media directory, `.env` file with the Gemini API keys and system prompt, leaves them where they are, and copies only the scripts and an output folder into the project base directory. It scans the input columns and confirms with you which id columns and feature categories to extract. The Gemini batch scripts are numbered steps for uploading media, building JSONL, submitting jobs, checking status, retrying errors and reviewing results, and they read settings from a copy of `pipeline_config.py` passed with `--config`. Reference files cover checkpoints and rotating several API keys.
Read from SKILL.md and the folder at commit d8aaee7. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 8 files in scripts/ (Python, from the files we listed), which the agent can run.
Shell commands in SKILL.md call:
pythonFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
GOOGLE_API_KEYFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Multimodal Media Feature Extraction loads about 2k tokens when it runs, and up to ~12k if it reads all its reference files. Until then it costs about 63 tokens; SKILL.md has 775 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check noted patterns worth knowing about, such as sudo or a known installer.
put/` — user assets (input data, media, `.env`, system prompt) can stay wherever they already live; point the scripts at- `.env` with `GOOGLE_API_KEY1`-`4` (Gemini only; check project root, home directory, common locations)gle-genai`, `python-dotenv` (+ above) | `.env` with `GOOGLE_API_KEY1`-`4` |Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from TyrealQ/q-skills at commit d8aaee7, republished under its MIT licence (© TyrealQ). 775 words, ~2,033 tokens.
.claude/skills/q-multimodal/SKILL.md (or your agent's skills folder). This skill also uses 26 other files; get the full folder from GitHub.Multimodal media analysis: local low-level features (Pillow, openSMILE), mid/high-level visual semantic analysis (Gemini API). Local pipelines are fully generic and CLI-driven. Gemini pipelines are config-driven: copy scripts/gemini/pipeline_config.py to your project, customize, and run with --config <path>.
Do this once when adopting the skill in a new project. The canonical layout is a target only for scripts/ and output/ — user assets (input data, media, .env, system prompt) can stay wherever they already live; point the scripts at them via absolute paths.
<BASE_DIR>: read the project's CLAUDE.md if it exists. If BASE_DIR isn't defined, ask the user which directory is the project root..env with GOOGLE_API_KEY1-4 (Gemini only; check project root, home directory, common locations)pipeline_config.py fields or CLI --input / --base-dir arguments to the absolute paths you found. Never move user data without explicit confirmation.scripts/ and output/ under <BASE_DIR>. Copy the pipelines actually being used from ${SKILL_DIR}/scripts/ into <BASE_DIR>/scripts/:pillow/, opensmile/, librosa/, common.pygemini/batch/, gemini/standard/, gemini/pipeline_config.py (template → adapt in place or copy to <BASE_DIR>/scripts/pipeline_config.py)output/ is auto-created by scripts on first runpython -c "import pandas as pd; print(list(pd.read_<FORMAT>('INPUT', nrows=1).columns))"--id-cols with the user. The file column (--file-col) is always retained in output — every row identifies its exact media asset — and --id-cols adds further columns (e.g. a post id) carried through checkpoints and merges.--features with the user as well, so the feature categories a pipeline extracts match what the analysis needs.Read the relevant reference file before executing a pipeline. These contain all flags, output column definitions, edge cases, and validation rules.
Local pipelines:
references/image-visual-features.md — all feature categories, column definitions, computation notesreferences/video-visual-features.md — frame extraction, aggregation logic, dual output formatreferences/audio-features.md — openSMILE feature sets, interpretable scores, feature levelsreferences/music-features.md — librosa feature sets, tier-1 music scores, raw tonal/timbre blockGemini pipelines:
references/gemini-batch-workflow.md — full 6-step batch pipeline, retry workflow, error handlingreferences/gemini-standard.md — standard pipeline details, model config, adapting for new projectsreferences/multi-key-management.md — multi-key quota strategy, retry threshold decision tableShared:
references/checkpoint-format.md — column order, validation rules, output directory structure| Pipeline | Python packages | System |
|---|---|---|
| Image visual | Pillow, numpy, pandas, tqdm, openpyxl | — |
| Video visual | (same as image) + scenedetect[opencv] | ffmpeg on PATH (for --extractor ffmpeg) |
| Audio | opensmile, pandas, tqdm, openpyxl | ffmpeg + ffprobe on PATH (preflight-checked; both ship with any FFmpeg install) |
| Music | librosa, numpy, scipy, soundfile, pandas, tqdm, openpyxl | ffmpeg on PATH (compressed/video formats, via audioread) |
| Gemini | google-genai, python-dotenv (+ above) | .env with GOOGLE_API_KEY1-4 |
Script path = ${SKILL_DIR}/scripts/<path>. Read the pipeline's reference file before running.
| Script | Input | Output | Reference |
|---|---|---|---|
pillow/visual_features.py | Images | 47 pixel features (color, texture, spatial, quality) | image-visual-features.md |
pillow/video_features.py | Videos | Frame-level + video-level aggregated features (scene-based extraction by default, FFmpeg fixed-interval optional) | video-visual-features.md |
opensmile/audio_features.py | Video/audio | 8 interpretable scores + raw openSMILE features + stream/signal diagnostics (audio_status, configurable silence threshold) | audio-features.md |
librosa/music_features.py | Audio/video | 13 music-native scores + raw librosa features | music-features.md |
librosa/music_features.py complements opensmile/audio_features.py: openSMILE covers speech/prosody, librosa covers music-native features (tempo, key/mode, harmony, timbre).
Shared utilities: common.py — read_input(), save_excel(), derive_subject(), merge_checkpoints()
Command pattern: python <script> --input <file> --base-dir <root> [--features ...] [--id-cols ...] [--subjects ...] [--preview] [--merge]
After a run finishes, check the output for rows with ok=False and confirm the failure count is zero or has an understood cause before treating the extraction as complete.
Both pipelines read a pipeline_config.py file that defines paths, schema, metadata formatting, and validation rules. Copy scripts/gemini/pipeline_config.py to your project and customize.
Standard (gemini/standard/gemini_standard.py): inline media, 25 workers, auto-retry. See gemini-standard.md.
Batch (gemini/batch/[0-5]*.py + utils.py): 6-step pipeline, 50% discount. URIs expire after 48 hours. See gemini-batch-workflow.md.
python 0uploadMedia.py --config /path/to/config.py --submit --max-batch-gb 2 --key 1
python 3checkStatus.py --config /path/to/config.py --poll
python 4retryErrors.py --config /path/to/config.py --preview
# >500 failures: batch retry
python 4retryErrors.py --config /path/to/config.py --submit
python 3checkStatus.py --config /path/to/config.py --poll
python 4retryErrors.py --config /path/to/config.py --collect
# <=500 failures or after batch retries: live fallback
python 4retryErrors.py --config /path/to/config.py --standard
python 5review.py --config /path/to/config.py --mergeDecision: >5 GB or >10 subjects and not time-sensitive → batch. Otherwise → standard. See gemini-standard.md.
Multi-key: Each GOOGLE_API_KEY{N} = 20 GB quota. See multi-key-management.md.
Local pipelines (Pillow, openSMILE, librosa): No modification needed. All project-specific values come from CLI args.
Gemini pipelines: Config-driven, no script modification needed. Scripts are copied to the project in step 2 above, then:
<BASE_DIR>/scripts/pipeline_config.py (already copied from template in step 2)BASE_DIR, INPUT_PATH, SYSTEM_PROMPT_PATH to your project paths (SYSTEM_PROMPT_PATH relative to BASE_DIR, e.g., scripts/<prompt>.txt)GROUP_COL, FILE_COL, ANALYSIS_FIELDS to match your input schema and system promptsubject_id() and format_metadata() for your domainvalidate_row() for field-specific validation rules--config <BASE_DIR>/scripts/pipeline_config.pyInclude: Image/video/audio feature extraction, Gemini visual semantic analysis, batch job management, checkpoint merging, multi-key quota management.
Exclude: ML model training, deep learning inference, real-time streaming analysis.
© TyrealQ, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 26 other files (scripts, references) in skills/q-scholar/q-multimodal of TyrealQ/q-skills.
Open the folder on GitHubat commit d8aaee7
Multimodal Media Feature Extraction next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Multimodal Media Feature Extraction this skillTyrealQ/q-skills | 108 | — | ~2k | Automated safety check: Notes | MIT | |
| Transforming Dataancoleman/ai-design-components | 525 | — | ~3k | Automated safety check: Pass | MIT | |
| Python Pipelinejamditis/claude-skills-journalism | 416 | — | ~4.8k | Automated safety check: Pass | MIT | |
| Gemini Video Understandingeinverne/dotfiles | 121 | — | ~2.6k | Automated safety check: Notes | MIT | |
| Watch Video Q&Abradautomates/claude-video | 18k | — | ~4.3k | Automated safety check: Notes | MIT | |
| ModLens Image Vision Bridgeliustack/modlens | 4.2k | — | ~1.3k | Automated safety check: Notes | MIT |
ancoleman/ai-design-components
Transform raw data into analytical assets using ETL/ELT patterns, SQL (dbt), Python (pandas/polars/PySpark), and orchestration (Airflow).
jamditis/claude-skills-journalism
Python data pipelines with modular architecture. An agent skill from jamditis/claude-skills-journalism.
einverne/dotfiles
Analyze videos using Google's Gemini API - describe content, answer questions, transcribe audio with visual descriptions, reference timestamps, clip videos, and process YouTube URLs.
bradautomates/claude-video
Lets the agent answer questions about a video from a URL or local file by downloading it, extracting frames and a transcript, or by sending it to Gemini's video model.
liustack/modlens
Gives text-only models sight by running the modlens CLI on an image path or URL and returning structured JSON evidence with transcribed text, layout and semantics.
google-gemini/gemini-skills
A skill your agent uses when writing code that calls the Gemini API for text generation, multi-turn chat, multimodal understanding, image generation, video generation, speech generation (TTS), voice…
TyrealQ/q-skills
Converts a report or other document into a business story and then an infographic image, pausing for your review after each step.
TyrealQ/q-skills
Runs exploratory data analysis on tabular data after you confirm each column's measurement level, then writes CSV tables and a narrative summary.
TyrealQ/q-skills
Consolidates BERTopic, LDA or NMF topic output into a theory-driven classification framework and writes the final labels back to an Excel file.
TyrealQ/q-skills
Generates branded slide deck images from written content, with a content analysis step, a layout catalog and scripts that merge the slides into PowerPoint or PDF.
TyrealQ/q-skills
Audits a repository's file layout and project documentation against a written convention file, then proposes moves, deletions and doc fixes as an approved plan before touching anything.
TyrealQ/q-skills
Stage and commit uncommitted changes with conventional commit messages.
Works with
Categories
Extracts pixel, video-frame, speech, music and visual-semantic features from image, video and audio files for research datasets, using local tools or the Gemini API. Two kinds of pipelines run over a folder of media. The local ones are driven from the command line and pull low-level features: pixel features with Pillow, video frames with FFmpeg and Pillow, speech and audio features with openSMILE, and music features with librosa.
Multimodal Media Feature Extraction fits situations like: extracting image or video-frame features for a media dataset; computing speech or music features from audio recordings; running large Gemini batch jobs over media files and retrying failed requests; coding visual content semantically with Gemini across a whole dataset.
Run `npx skills add TyrealQ/q-skills --skill q-multimodal -a claude-code`. Or copy the skill folder (skills/q-scholar/q-multimodal in TyrealQ/q-skills) into .claude/skills/q-multimodal in your project. Claude Code loads it when a task matches its description.
Run `npx skills add TyrealQ/q-skills --skill q-multimodal -a codex`. Or copy the skill folder (skills/q-scholar/q-multimodal in TyrealQ/q-skills) into .agents/skills/q-multimodal in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add TyrealQ/q-skills --skill q-multimodal -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/q-multimodal, .gemini/skills/q-multimodal, .github/skills/q-multimodal and .opencode/skills/q-multimodal in your project.
Going by SKILL.md and its folder, Multimodal Media Feature Extraction needs Python for the scripts in its folder, the command-line tools its instructions call (python) and credentials named GOOGLE_API_KEY. Our summary lists: Python with Pillow, openSMILE and librosa for the local pipelines; FFmpeg for video frames; Google API keys for Gemini in a .env file.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found notes only (mentions a .env file), nothing it rates as a warning. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Multimodal Media Feature Extraction is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2k tokens (SKILL.md is roughly 8.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 9.8k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Multimodal Media Feature Extraction: Transforming Data (ancoleman/ai-design-components, 525 stars), Python Pipeline (jamditis/claude-skills-journalism, 416 stars), Gemini Video Understanding (einverne/dotfiles, 121 stars) and Watch Video Q&A (bradautomates/claude-video, 18k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
TyrealQ (a GitHub user) maintains it in TyrealQ/q-skills, which has 108 GitHub stars. The repository holds 16 skills in this directory. The repository was last updated on September 23, 2026.
Source: TyrealQ/q-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.