Music Video Subtitle Generator
tl2012tl/comfyUI-llama-TE
For musicians, video creators, and social-media editors producing AI music videos or emotional short films with lyric typography.
Plan, write, audit, or produce music videos with coordinated music, imagery, camera motion, optional performance, and lyric typography.
$ npx skills add vllm-project/vllm-omni --skill music-video-subtitle-generator -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install vllm-project/vllm-omni music-video-subtitle-generator --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/vllm-project/vllm-omni.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/music-video-subtitle-generator .claude/skills/music-video-subtitle-generator && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "music-video-subtitle-generator" agent skill from https://github.com/vllm-project/vllm-omni/tree/main/.agents/skills/music-video-subtitle-generator into .claude/skills/music-video-subtitle-generator/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "music-video-subtitle-generator", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/vllm-project/vllm-omni/tree/main/.agents/skills/music-video-subtitle-generatorType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add vllm-project/vllm-omni --skill music-video-subtitle-generator -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install vllm-project/vllm-omni music-video-subtitle-generator --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/vllm-project/vllm-omni.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.agents/skills/music-video-subtitle-generator .agents/skills/music-video-subtitle-generator && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "music-video-subtitle-generator" agent skill from https://github.com/vllm-project/vllm-omni/tree/main/.agents/skills/music-video-subtitle-generator into .agents/skills/music-video-subtitle-generator/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "music-video-subtitle-generator", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add vllm-project/vllm-omni --skill music-video-subtitle-generator -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install vllm-project/vllm-omni music-video-subtitle-generator --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/vllm-project/vllm-omni.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.agents/skills/music-video-subtitle-generator .cursor/skills/music-video-subtitle-generator && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "music-video-subtitle-generator" agent skill from https://github.com/vllm-project/vllm-omni/tree/main/.agents/skills/music-video-subtitle-generator into .cursor/skills/music-video-subtitle-generator/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "music-video-subtitle-generator", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/vllm-project/vllm-omni.git --path .agents/skills/music-video-subtitle-generator--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add vllm-project/vllm-omni --skill music-video-subtitle-generator -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install vllm-project/vllm-omni music-video-subtitle-generator --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/vllm-project/vllm-omni.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.agents/skills/music-video-subtitle-generator .gemini/skills/music-video-subtitle-generator && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "music-video-subtitle-generator" agent skill from https://github.com/vllm-project/vllm-omni/tree/main/.agents/skills/music-video-subtitle-generator into .gemini/skills/music-video-subtitle-generator/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "music-video-subtitle-generator", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install vllm-project/vllm-omni music-video-subtitle-generatorInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add vllm-project/vllm-omni --skill music-video-subtitle-generator -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/vllm-project/vllm-omni.git skills-src && mkdir -p .github/skills && cp -r skills-src/.agents/skills/music-video-subtitle-generator .github/skills/music-video-subtitle-generator && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "music-video-subtitle-generator" agent skill from https://github.com/vllm-project/vllm-omni/tree/main/.agents/skills/music-video-subtitle-generator into .github/skills/music-video-subtitle-generator/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "music-video-subtitle-generator", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add vllm-project/vllm-omni --skill music-video-subtitle-generator -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install vllm-project/vllm-omni music-video-subtitle-generator --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/vllm-project/vllm-omni.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.agents/skills/music-video-subtitle-generator .opencode/skills/music-video-subtitle-generator && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "music-video-subtitle-generator" agent skill from https://github.com/vllm-project/vllm-omni/tree/main/.agents/skills/music-video-subtitle-generator into .opencode/skills/music-video-subtitle-generator/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "music-video-subtitle-generator", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
music-video-subtitle-generatorPlan, write, audit, or produce music videos with coordinated music, imagery, camera motion, optional performance, and lyric typography.
Music Video Subtitle Generator is an agent skill from vllm-project/vllm-omni. Plan, write, audit, or produce music videos with coordinated music, imagery, camera motion, optional performance, and lyric typography. Use for beat-aware MVs and instrumental visual narratives, not ordinary subtitle cleanup or unrelated video edits.
Its SKILL.md is about 1.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files, including reference files (for example `agents/openai.yaml` and `references/music-timeline.md`).
It sits in Media & Creative, covering Transcription and Typography. The repository describes itself as: A framework for efficient model inference with omni-modality models. The licence is Apache-2.0.
Read from SKILL.md and the folder at commit 096988d. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md.
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Music Video Subtitle Generator loads about 1.1k tokens when it runs, and up to ~1.6k if it reads all its reference files. Until then it costs about 70 tokens; SKILL.md has 533 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from vllm-project/vllm-omni at commit 096988d, republished under its Apache-2.0 licence (© vllm-project). 533 words, ~1,087 tokens.
.claude/skills/music-video-subtitle-generator/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.Read the portable workflow and, when building a timeline, music timeline. Deliver the requested scope: audit, prompt, storyboard, or complete video.
Identify duration, ratio, output resolution, music source/excerpt, visual style, character/scene references, performance mode, and whether text is wanted. Treat a supplied song as the master track unless the user requests replacement. Preserve provided lyrics exactly; do not translate or rewrite them unasked. For instrumental work, do not invent singing, lyrics, a performer, or typography. Generate original lyrics only when requested or clearly within an authorized songwriting brief.
Measure supplied audio duration and inspect/listen for beat, phrase, breath, drop, and energy changes using available tools. Label estimated or planned timing when measurement is unavailable. If the requested excerpt exceeds the audio, resolve the discrepancy without silently stretching or looping it.
Do not apply a preset's grain, darkness, hard cuts, or constant close-ups to an incompatible brief. A bright nighttime scene remains night. A continuous flat vector morph keeps smooth color fields and uses elastic shape motion for accents. Use hard cuts, slow holds, or dissolves only when appropriate to the user's intent.
Map visual actions, camera moves, scene changes, and optional text to the music's global time. Build a causal visual progression instead of defaulting to one person singing throughout. Allow room to read a shape, gesture, or lyric before changing it.
For longer work, choose a supported continuation or clip-assembly approach through the portable workflow. Keep visual shot boundaries distinct from generation windows. Local window prompts must account for overlap without replaying the full previous action. A new scene may require a new reference rather than inheritance of the old scene. Record entry/exit states for each planned segment.
Compile the timeline into H3's base or Ref2VA fields, preserving dialogue, lyrics, and visible text verbatim. Include concrete instrumentation and rhythmic accents, but treat requested BPM and beat alignment as targets until checked in the output.
When requested, give each shot one main readable text event. Keep words away from eyes and critical lip motion. Match sung text to the active phrase, and distinguish spatial typography from subtitle bars. Reduce text complexity if the model cannot render it, or use a disclosed overlay consistent with the requested deliverable.
For supplied audio, assemble clips to one global master track, maintaining phrase and lip continuity; avoid cuts inside a vowel. For native audio, inspect actual continuity across generation windows. Preserve useful SFX and avoid layered duplicate music. Retiming, grain, grading, interpolation, and upscaling are deliberate editing choices, not automatic finishing requirements.
Verify transformation/story order, identities, scene changes, requested text, musical continuity, timing, and full media decode. Label raw and processed versions. Deliver the final file and current prompt/timeline files, with measured properties and remaining limitations rather than claims of perfect synchronization.
© vllm-project, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 2 other files (references) in .agents/skills/music-video-subtitle-generator of vllm-project/vllm-omni.
Open the folder on GitHubat commit 096988d
Music Video Subtitle Generator next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Music Video Subtitle Generator this skillvllm-project/vllm-omni | 7.1k | — | ~1.1k | Automated safety check: Pass | Apache-2.0 | |
| Music Video Subtitle Generatortl2012tl/comfyUI-llama-TE | 241 | 4 repos | ~3.6k | Automated safety check: Pass | None | |
| Paper Collage Explainer Generatortl2012tl/comfyUI-llama-TE | 241 | 4 repos | ~5.2k | Automated safety check: Pass | None | |
| Anthropic Brand Stylinganthropics/skills | 180k | 30 repos | ~559 | Automated safety check: Pass | Apache-2.0 | |
| Video Understandcalesthio/OpenMontage | 66k | — | ~841 | Automated safety check: Pass | AGPL-3.0 | |
| Minimal Zine Poster GeneratorLiamGvchi/gc-minimal-zine-poster | 7.3k | — | ~2.9k | Automated safety check: Pass | MIT |
tl2012tl/comfyUI-llama-TE
For musicians, video creators, and social-media editors producing AI music videos or emotional short films with lyric typography.
tl2012tl/comfyUI-llama-TE
For creators, educators, and social-video editors who need a tactile paper-collage language for narration, knowledge points, opinions, or abstract topics.
anthropics/skills
Applies Anthropic's brand colors and fonts to artifacts such as PowerPoint slides, using fixed hex values for text and accents, Poppins headings and Lora body text.
calesthio/OpenMontage
Understand video content locally using ffmpeg frame extraction and Whisper transcription.
LiamGvchi/gc-minimal-zine-poster
Creates or analyzes quiet, paper-texture zine posters with big negative space, one color accent and experimental type, returning an image prompt and the generated poster.
heygen-com/hyperframes
Finds, generates and edits media for HyperFrames video projects: music, sound effects, images, icons, logos, voiceovers, captions and color grades.
vllm-project/vllm-omni
Diagnose and optimize vLLM Omni diffusion workloads, especially Wan/Qwen/Flux-style image and video generation.
vllm-project/vllm-omni
Self-check your branch before creating a PR — catch dead code, prevent new model-specific Python examples, verify accuracy/perf claims, validate PR title format, and confirm merge readiness.
vllm-project/vllm-omni
Work on vLLM-Omni quantization for diffusion, autoregressive, omni, or multi-stage models.
vllm-project/vllm-omni
Review pull requests and local branches for vllm-project/vllm-omni with a frozen snapshot, module-design ownership, feature-design overlays, targeted validation, and concise evidence-backed findings.
vllm-project/vllm-omni
Write MiniMax H3 video generation prompts for T2VA, I2VA, FL2VA, L2VA, and Ref2VA.
vllm-project/vllm-omni
Add a new diffusion model (text-to-image, text-to-video, image-to-video, text-to-audio, image editing) to vLLM-Omni, including native non-Diffusers ports, reference-parity validation, Cache-DiT…
Categories
Plan, write, audit, or produce music videos with coordinated music, imagery, camera motion, optional performance, and lyric typography. Music Video Subtitle Generator is an agent skill from vllm-project/vllm-omni. Plan, write, audit, or produce music videos with coordinated music, imagery, camera motion, optional performance, and lyric typography.
Music Video Subtitle Generator fits situations like: beat-aware MVs and instrumental visual narratives; not ordinary subtitle cleanup; unrelated video edits.
Run `npx skills add vllm-project/vllm-omni --skill music-video-subtitle-generator -a claude-code`. Or copy the skill folder (.agents/skills/music-video-subtitle-generator in vllm-project/vllm-omni) into .claude/skills/music-video-subtitle-generator in your project. Claude Code loads it when a task matches its description.
Run `npx skills add vllm-project/vllm-omni --skill music-video-subtitle-generator -a codex`. Or copy the skill folder (.agents/skills/music-video-subtitle-generator in vllm-project/vllm-omni) into .agents/skills/music-video-subtitle-generator in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add vllm-project/vllm-omni --skill music-video-subtitle-generator -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/music-video-subtitle-generator, .gemini/skills/music-video-subtitle-generator, .github/skills/music-video-subtitle-generator and .opencode/skills/music-video-subtitle-generator in your project.
SKILL.md names no scripts, command-line tools or credentials: Music Video Subtitle Generator is instructions for the agent only.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Music Video Subtitle Generator is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.1k tokens (SKILL.md is roughly 4.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 513 tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Music Video Subtitle Generator: Music Video Subtitle Generator (tl2012tl/comfyUI-llama-TE, 241 stars), Paper Collage Explainer Generator (tl2012tl/comfyUI-llama-TE, 241 stars), Anthropic Brand Styling (anthropics/skills, 180k stars) and Video Understand (calesthio/OpenMontage, 66k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
vllm-project (a GitHub organization) maintains it in vllm-project/vllm-omni, which has 7,119 GitHub stars. The repository holds 20 skills in this directory. The repository was last updated on October 11, 2026.
Source: vllm-project/vllm-omni on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.