Paper Collage Explainer Generator
tl2012tl/comfyUI-llama-TE
For creators, educators, and social-video editors who need a tactile paper-collage language for narration, knowledge points, opinions, or abstract topics.
Provider-independent production guidance for translating measured audio features into deterministic video timing and motion.
$ npx skills add calesthio/generative-media-skills --skill audio-reactive-video-composition -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install calesthio/generative-media-skills audio-reactive-video-composition --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/calesthio/generative-media-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/production/runtime-assembly/audio-reactive-video-composition .claude/skills/audio-reactive-video-composition && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "audio-reactive-video-composition" agent skill from https://github.com/calesthio/generative-media-skills/tree/main/skills/production/runtime-assembly/audio-reactive-video-composition into .claude/skills/audio-reactive-video-composition/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "audio-reactive-video-composition", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/calesthio/generative-media-skills/tree/main/skills/production/runtime-assembly/audio-reactive-video-compositionType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add calesthio/generative-media-skills --skill audio-reactive-video-composition -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install calesthio/generative-media-skills audio-reactive-video-composition --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/calesthio/generative-media-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/production/runtime-assembly/audio-reactive-video-composition .agents/skills/audio-reactive-video-composition && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "audio-reactive-video-composition" agent skill from https://github.com/calesthio/generative-media-skills/tree/main/skills/production/runtime-assembly/audio-reactive-video-composition into .agents/skills/audio-reactive-video-composition/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "audio-reactive-video-composition", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add calesthio/generative-media-skills --skill audio-reactive-video-composition -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install calesthio/generative-media-skills audio-reactive-video-composition --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/calesthio/generative-media-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/production/runtime-assembly/audio-reactive-video-composition .cursor/skills/audio-reactive-video-composition && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "audio-reactive-video-composition" agent skill from https://github.com/calesthio/generative-media-skills/tree/main/skills/production/runtime-assembly/audio-reactive-video-composition into .cursor/skills/audio-reactive-video-composition/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "audio-reactive-video-composition", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/calesthio/generative-media-skills.git --path skills/production/runtime-assembly/audio-reactive-video-composition--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add calesthio/generative-media-skills --skill audio-reactive-video-composition -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install calesthio/generative-media-skills audio-reactive-video-composition --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/calesthio/generative-media-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/production/runtime-assembly/audio-reactive-video-composition .gemini/skills/audio-reactive-video-composition && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "audio-reactive-video-composition" agent skill from https://github.com/calesthio/generative-media-skills/tree/main/skills/production/runtime-assembly/audio-reactive-video-composition into .gemini/skills/audio-reactive-video-composition/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "audio-reactive-video-composition", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install calesthio/generative-media-skills audio-reactive-video-compositionInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add calesthio/generative-media-skills --skill audio-reactive-video-composition -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/calesthio/generative-media-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/production/runtime-assembly/audio-reactive-video-composition .github/skills/audio-reactive-video-composition && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "audio-reactive-video-composition" agent skill from https://github.com/calesthio/generative-media-skills/tree/main/skills/production/runtime-assembly/audio-reactive-video-composition into .github/skills/audio-reactive-video-composition/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "audio-reactive-video-composition", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add calesthio/generative-media-skills --skill audio-reactive-video-composition -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install calesthio/generative-media-skills audio-reactive-video-composition --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/calesthio/generative-media-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/production/runtime-assembly/audio-reactive-video-composition .opencode/skills/audio-reactive-video-composition && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "audio-reactive-video-composition" agent skill from https://github.com/calesthio/generative-media-skills/tree/main/skills/production/runtime-assembly/audio-reactive-video-composition into .opencode/skills/audio-reactive-video-composition/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "audio-reactive-video-composition", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
audio-reactive-video-compositionProvider-independent production guidance for translating measured audio features into deterministic video timing and motion.
Audio Reactive Video Composition is an agent skill from calesthio/generative-media-skills. Provider-independent production guidance for translating measured audio features into deterministic video timing and motion. Use for beat-, onset-, phrase-, energy-, silence-, or spectrum-reactive visualizers, edits, typography, and generated compositions; not for music-video concept direction, audio generation, or a runtime-specific recipe.
Its SKILL.md is about 2.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 1 other file (for example `EVAL.md`).
It sits in Media & Creative, covering Music and audio generation, Translation and Typography. The repository describes itself as: Research-backed agent skills and tools for premium image, video, audio, voice, and generative media production across AI coding assistants. The licence is MIT.
9 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 8c85352. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are json).
From the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
w3.orglibrosa.orgessentia.upf.edumir-eval.readthedocs.ioee.columbia.eduffmpeg.orgtech.ebu.chcopyright.govFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Audio Reactive Video Composition loads about 2.9k tokens when it runs. Until then it costs about 94 tokens; SKILL.md has 1,353 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from calesthio/generative-media-skills at commit 8c85352, republished under its MIT licence (© calesthio). 1,353 words, ~2,861 tokens.
.claude/skills/audio-reactive-video-composition/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.Use this skill to turn an approved audio source into an auditable cue map and then into bounded visual behavior. The central contract is:
$$ \text{source audio} \rightarrow \text{measured features} \rightarrow \text{confidence-bearing cues} \rightarrow \text{reviewed anchors} \rightarrow \text{deterministic visual timeline} $$
Do not equate a detector output with editorial meaning. An onset is not automatically a beat, cut, drop, or lyric accent; an acoustic cluster is not automatically a verse or chorus.
Analysis algorithms, defaults, and model behavior are volatile. Facts were verified 2026-07-12. Pin the decoder, library, model, parameters, and random seeds used for each production.
This skill owns source custody, analysis policy, confidence interpretation, rhythmic/non-rhythmic routing, cue promotion, feature-to-visual mappings, rational frame alignment, accessibility, and render QA.
It does not own music-video narrative or artist branding, music generation, mastering, source separation, transcription, lyric writing, or a HyperFrames/Remotion/FFmpeg-specific implementation.
Record before analysis:
Use sample indices as the primary analysis clock where possible. Derived seconds should not replace exact source positions.
Documented facts:
librosa.beat.beat_track estimates tempo from onset strength and selects beat positions consistent with it; it does not return calibrated beat confidence.librosa.onset.onset_detect peak-picks an onset-strength envelope. Onsets estimate event attacks, not semantic accents.Keep raw candidates distinct from human-promoted anchors.
Each cue should include:
{
"id": "accent-014",
"type": "onset",
"time_samples": 417312,
"time_seconds": 9.46399,
"value": 0.82,
"units": "normalized-onset-strength",
"analyzer": "record exact library and algorithm",
"confidence": null,
"confidence_semantics": "not supplied by analyzer",
"profile_hash": "sha256:...",
"status": "promoted-anchor",
"editorial_role": "single visual accent",
"reason": "reviewed strong transient before section change"
}For interval features, include exact end samples/times. Record rejected candidates rather than deleting evidence. Keep acoustic section names neutral until lyrics, metadata, or human review supports functional labels.
Use when the selected tracker has useful evidence, local tempo is stable enough, and beat/onset results agree. Periodic motion may follow beats; selected strong accents or reviewed section changes may motivate cuts.
Use reliable rhythmic windows locally. Elsewhere shift to onsets, transcript phrases, energy, silence, or manual anchors. Never extrapolate a grid through a failed interval.
Use reviewed speech/lyric phrases, silence, energy contour, spectral change, and manually confirmed macro anchors. This suits rubato, ambient work, spoken word, sparse recordings, and free improvisation.
Production heuristics:
Avoid quantitative-looking mappings that imply false measurement. If energy maps to scale, define the exact bounded range and do not call it emotion.
Keep exact rates such as $30000/1001$ rational. For each cue, define a rounding policy. A conservative visual-response policy assigns the cue to the first output frame whose presentation time is at or after the audio event:
$$ f = \left\lceil t_{event} \cdot \frac{fps_{num}}{fps_{den}} \right\rceil $$
Record both the original event time and mapped frame. Verify actual frame presentation timestamps after encoding because an encoder may duplicate or drop frames to satisfy constant-frame-rate output.
The visual state at frame $f$ must derive from cue data and frame time, not wall-clock playback. Seed procedural mappings and freeze analyzer output before distributed rendering.
Treat transcript or lyric timing as a separate evidence stream. Do not use pitch voicing as vocal detection. Human-review names, lyrics, line boundaries, and timing before they control typography.
Prerecorded synchronized media with meaningful speech needs accurate captions. Captions should include meaningful non-speech sound where needed. A stylized lyric layer does not automatically replace accessible captions or transcript.
WCAG 2.2 SC 2.3.1 limits flashing above three times in one second unless below general/red-flash thresholds. Test loops while looping and at the largest intended scale. Reduced motion does not make unsafe flashing safe.
Provide a lower-motion version when large displacement, zoom, shake, or dense event response may cause discomfort. Reduce event density and travel, not merely output FPS.
The musical work, lyrics, and sound recording can carry separate rights. Possession of a file does not grant synchronization, adaptation, or distribution rights. Record source and transformation provenance; do not claim metadata proves authenticity or permission.
This is a complete example, not a mandatory formula.
Intent: 30-second 9:16 and 16:9 visualizer from an authorized 44.1 kHz stereo instrumental at 30 fps.
Approach: confirm stable tempo with a documented analyzer and inspect local pulse. Retain all beat/onset candidates. Promote every fourth beat for medium choreography, reviewed top-strength transients for sparse accents, and reviewed recurrence changes for macro transitions. Beat phase drives scale only from 1.000 to 1.035; low-band energy controls bounded depth; centroid controls a narrow texture-density range. A six-beat low-energy span holds title copy.
Map all anchors to the first frame at or after their sample time. Both aspect variants use identical cue IDs. QA click tracks, rerun hashes, PTS, duration, final-size text, and full-screen flashing.
Likely failure: half/double-tempo ambiguity. Repair by documenting the metrical level selected for visual periodicity without rewriting the raw detections.
This is a complete example, not a mandatory formula.
Intent: 75-second poem with ambient bed at 24 fps.
Approach: global beat trackers disagree, so use the non-rhythmic route. Reviewed transcript line starts/ends are primary anchors; pauses reset the field; sustained loudness rises control subtle expansion; acoustic-change candidates remain review prompts. Each stanza establishes a stable visual field, and punctuation settles motion. No beat cuts or word-by-word scaling.
The vertical variant reflows text but retains timing. QA muted-caption comprehension, audio-only clarity, every line boundary, breath-adjacent silence, reduced motion, flashing, and separate text/recording rights.
Likely failure: an acoustic boundary lands inside a sentence. Keep it as a low-confidence event and reject it as an editorial anchor.
Verified 2026-07-12:
mir_eval beat, onset, and segment metrics: https://mir-eval.readthedocs.io/© calesthio, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 1 other file in skills/production/runtime-assembly/audio-reactive-video-composition of calesthio/generative-media-skills.
Open the folder on GitHubat commit 8c85352
Audio Reactive Video Composition next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Audio Reactive Video Composition this skillcalesthio/generative-media-skills | 197 | — | ~2.9k | Automated safety check: Pass | MIT | |
| Paper Collage Explainer Generatortl2012tl/comfyUI-llama-TE | 241 | 4 repos | ~5.2k | Automated safety check: Pass | None | |
| Wonder BlocksKhan/wonder-blocks | 163 | — | ~3.2k | Automated safety check: Pass | MIT | |
| Vibe Matchercuriositech/some_claude_skills | 244 | — | ~3.3k | Automated safety check: Pass | MIT | |
| Design Tokenshashgraph-online/awesome-codex-plugins | 1.3k | — | ~1.2k | Automated safety check: Pass | Apache-2.0 | |
| Anthropic Brand Stylinganthropics/skills | 180k | 30 repos | ~559 | Automated safety check: Pass | Apache-2.0 |
tl2012tl/comfyUI-llama-TE
For creators, educators, and social-video editors who need a tactile paper-collage language for narration, knowledge points, opinions, or abstract topics.
Khan/wonder-blocks
Implements user interfaces using the Wonder Blocks (WB) design system — Khan Academy's React component library.
curiositech/some_claude_skills
Synesthete designer that translates emotional vibes and brand keywords into concrete visual DNA (colors, typography, layouts, interactions).
hashgraph-online/awesome-codex-plugins
A skill your agent uses when the user asks for design tokens, DTCG tokens, theme systems, color/typography/spacing/radius/elevation/motion tokens, translating tokens, exporting tokens, making a…
anthropics/skills
Applies Anthropic's brand colors and fonts to artifacts such as PowerPoint slides, using fixed hex values for text and accents, Poppins headings and Lora body text.
LiamGvchi/gc-minimal-zine-poster
Creates or analyzes quiet, paper-texture zine posters with big negative space, one color accent and experimental type, returning an image prompt and the generated poster.
calesthio/generative-media-skills
A skill your agent uses to turn generated, captured, scanned, or modeled 3D output into production-ready standalone assets for DCC, real-time engine, web, or interchange delivery.
calesthio/generative-media-skills
Provider-independent audio mixing and mastering direction for AI agents finishing generated videos, ads, trailers, explainers, podcasts, recuts, avatar clips, music videos, documentaries, and social…
calesthio/generative-media-skills
Provider-independent captions and media accessibility direction for AI agents producing or finishing generated videos, ads, social clips, explainers, avatar videos, documentaries, podcasts/video…
calesthio/generative-media-skills
Provider-independent production workflow for agents assembling, auditing, executing, and handing off ComfyUI node-graph workflows for image, video, upscale, inpaint, conditioning, and batch media…
calesthio/generative-media-skills
Provider-independent FFmpeg finishing workflow for AI agents preparing generated or edited media deliverables.
calesthio/generative-media-skills
Provider-independent quality assurance for AI-generated and AI-assisted media.
Categories
Provider-independent production guidance for translating measured audio features into deterministic video timing and motion. Audio Reactive Video Composition is an agent skill from calesthio/generative-media-skills. Provider-independent production guidance for translating measured audio features into deterministic video timing and motion.
Audio Reactive Video Composition fits situations like: spectrum-reactive visualizers; generated compositions; not for music-video concept direction; audio generation.
Run `npx skills add calesthio/generative-media-skills --skill audio-reactive-video-composition -a claude-code`. Or copy the skill folder (skills/production/runtime-assembly/audio-reactive-video-composition in calesthio/generative-media-skills) into .claude/skills/audio-reactive-video-composition in your project. Claude Code loads it when a task matches its description.
Run `npx skills add calesthio/generative-media-skills --skill audio-reactive-video-composition -a codex`. Or copy the skill folder (skills/production/runtime-assembly/audio-reactive-video-composition in calesthio/generative-media-skills) into .agents/skills/audio-reactive-video-composition in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add calesthio/generative-media-skills --skill audio-reactive-video-composition -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/audio-reactive-video-composition, .gemini/skills/audio-reactive-video-composition, .github/skills/audio-reactive-video-composition and .opencode/skills/audio-reactive-video-composition in your project.
SKILL.md names no scripts, command-line tools or credentials: Audio Reactive Video Composition is instructions for the agent only.
SKILL.md names 8 domains. As links in the text: w3.org, librosa.org, essentia.upf.edu, mir-eval.readthedocs.io, ee.columbia.edu, ffmpeg.org, tech.ebu.ch and copyright.gov. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Audio Reactive Video Composition is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.9k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Audio Reactive Video Composition: Paper Collage Explainer Generator (tl2012tl/comfyUI-llama-TE, 241 stars), Wonder Blocks (Khan/wonder-blocks, 163 stars), Vibe Matcher (curiositech/some_claude_skills, 244 stars) and Design Tokens (hashgraph-online/awesome-codex-plugins, 1.3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
calesthio (a GitHub user) maintains it in calesthio/generative-media-skills, which has 197 GitHub stars. The repository holds 26 skills in this directory. The repository was last updated on July 14, 2026.
Source: calesthio/generative-media-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.