Agent skill

Audio Reactive Video Composition

by calesthio in calesthio/generative-media-skills

Provider-independent production guidance for translating measured audio features into deterministic video timing and motion.

MITAuto-check passedMedia & Creative

Install Audio Reactive Video Composition

skills CLI
$ npx skills add calesthio/generative-media-skills --skill audio-reactive-video-composition -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install calesthio/generative-media-skills audio-reactive-video-composition --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/calesthio/generative-media-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/production/runtime-assembly/audio-reactive-video-composition .claude/skills/audio-reactive-video-composition && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
audio-reactive-video-composition
GitHub stars
197
Token cost
~2.9k tokens
SKILL.md length
1,353 words
Files
2
Skills in repo
26
Repo updated
First seen
Licence
MIT

At a glance

Provider-independent production guidance for translating measured audio features into deterministic video timing and motion.

  • Works in 9 steps: Re-run identical input/configuration and… → Audition click-marked candidate and… → Inspect tempo-level ambiguity and… → …
  • Spectrum-reactive visualizers
  • SKILL.md covers Evidence stance, Scope, Freeze the audio contract and Interpret features…, plus 10 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Audio Reactive Video Composition is an agent skill from calesthio/generative-media-skills. Provider-independent production guidance for translating measured audio features into deterministic video timing and motion. Use for beat-, onset-, phrase-, energy-, silence-, or spectrum-reactive visualizers, edits, typography, and generated compositions; not for music-video concept direction, audio generation, or a runtime-specific recipe.

Its SKILL.md is about 2.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 1 other file (for example `EVAL.md`).

It sits in Media & Creative, covering Music and audio generation, Translation and Typography. The repository describes itself as: Research-backed agent skills and tools for premium image, video, audio, voice, and generative media production across AI coding assistants. The licence is MIT.

When your agent uses it

  • Spectrum-reactive visualizers
  • Generated compositions
  • Not for music-video concept direction
  • Audio generation

Example prompts

  • “/audio-reactive-video-composition”

Workflow steps

9 steps, taken from the first numbered list in SKILL.md.

  1. Re-run identical input/configuration and compare canonical cue-map hashes.
  2. Audition click-marked candidate and promoted anchors.
  3. Inspect tempo-level ambiguity and failed-confidence windows.
  4. Compare source timing against decoded/trimmed master and non-zero timestamps.
  5. Inspect frames immediately before, on, and after every macro anchor.
  6. Verify aspect variants share cue IDs and frame assignments.
  7. Probe output frame PTS, duration, and audio sync.
  8. Review captions/lyrics, reduced motion, and flash safety.
  9. Preserve source hash, profile, cue map, mapping table, render versions, and approvals.

What it can do on your machine

Read from SKILL.md and the folder at commit 8c85352. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are json).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • w3.org
    • librosa.org
    • essentia.upf.edu
    • mir-eval.readthedocs.io
    • ee.columbia.edu
    • ffmpeg.org
    • tech.ebu.ch
    • copyright.gov

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Audio Reactive Video Composition loads about 2.9k tokens when it runs. Until then it costs about 94 tokens; SKILL.md has 1,353 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~94
When it runs · the whole SKILL.md, loaded when a task matches
~2.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from calesthio/generative-media-skills at commit 8c85352, republished under its MIT licence (© calesthio). 1,353 words, ~2,861 tokens.

Download SKILL.mdSave it as .claude/skills/audio-reactive-video-composition/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
audio-reactive-video-composition
description
Provider-independent production guidance for translating measured audio features into deterministic video timing and motion. Use for beat-, onset-, phrase-, energy-, silence-, or spectrum-reactive visualizers, edits, typography, and generated compositions; not for music-video concept direction, audio generation, or a runtime-specific recipe.

Audio-reactive video composition

Use this skill to turn an approved audio source into an auditable cue map and then into bounded visual behavior. The central contract is:

$$ \text{source audio} \rightarrow \text{measured features} \rightarrow \text{confidence-bearing cues} \rightarrow \text{reviewed anchors} \rightarrow \text{deterministic visual timeline} $$

Do not equate a detector output with editorial meaning. An onset is not automatically a beat, cut, drop, or lyric accent; an acoustic cluster is not automatically a verse or chorus.

Evidence stance

  • Documented fact: behavior stated by official analysis libraries, standards, or cited research.
  • Production heuristic: a practical mapping that must be tested on the track and audience.
  • Empirical observation: a measured result from the supplied audio, analyzer run, render, or playback.

Analysis algorithms, defaults, and model behavior are volatile. Facts were verified 2026-07-12. Pin the decoder, library, model, parameters, and random seeds used for each production.

Scope

This skill owns source custody, analysis policy, confidence interpretation, rhythmic/non-rhythmic routing, cue promotion, feature-to-visual mappings, rational frame alignment, accessibility, and render QA.

It does not own music-video narrative or artist branding, music generation, mastering, source separation, transcription, lyric writing, or a HyperFrames/Remotion/FFmpeg-specific implementation.

Freeze the audio contract

Record before analysis:

  • source path/URI, SHA-256, acquisition source, rights basis, and restrictions;
  • selected stream, codec, native sample rate, channels/layout, start timestamp, and duration;
  • decoder/resampler and versions;
  • canonical PCM format, channel/downmix policy, analysis sample rate, window/hop sizes, centering, and padding;
  • analyzer/library/model versions, priors, thresholds, and seeds;
  • trim offsets and source time origin;
  • master output frame rate as a rational number;
  • transcript/lyrics source, language, timing provenance, review status, and separate rights basis.

Use sample indices as the primary analysis clock where possible. Derived seconds should not replace exact source positions.

Interpret features conservatively

Documented facts:

  • librosa.beat.beat_track estimates tempo from onset strength and selects beat positions consistent with it; it does not return calibrated beat confidence.
  • librosa.onset.onset_detect peak-picks an onset-strength envelope. Onsets estimate event attacks, not semantic accents.
  • Predominant local pulse can model changing tempo, but remains an estimate.
  • Essentia confidence values are algorithm-specific; they are not universal probabilities, and some routes return an unusable zero confidence.
  • Beat evaluation commonly accounts for half/double-tempo metrical ambiguity and uses tolerances rather than exact equality.
  • Structural boundaries and structural labels are separate tasks. Cluster labels such as A/B/C do not establish verse/chorus meaning.
  • RMS is an energy measure. EBU R128 loudness uses defined weighting and windows. Neither is a direct emotion score.
  • Spectral centroid, bandwidth, and contrast describe spectrum distribution; they do not mean happiness, tension, or quality.
  • Silence detection is relative to a declared reference and threshold.
  • pYIN voicing means pitch periodicity, not proof that a person is speaking or singing.

Keep raw candidates distinct from human-promoted anchors.

Cue-map contract

Each cue should include:

json
{
  "id": "accent-014",
  "type": "onset",
  "time_samples": 417312,
  "time_seconds": 9.46399,
  "value": 0.82,
  "units": "normalized-onset-strength",
  "analyzer": "record exact library and algorithm",
  "confidence": null,
  "confidence_semantics": "not supplied by analyzer",
  "profile_hash": "sha256:...",
  "status": "promoted-anchor",
  "editorial_role": "single visual accent",
  "reason": "reviewed strong transient before section change"
}

For interval features, include exact end samples/times. Record rejected candidates rather than deleting evidence. Keep acoustic section names neutral until lyrics, metadata, or human review supports functional labels.

Choose a routing mode

Rhythmic

Use when the selected tracker has useful evidence, local tempo is stable enough, and beat/onset results agree. Periodic motion may follow beats; selected strong accents or reviewed section changes may motivate cuts.

Mixed confidence

Use reliable rhythmic windows locally. Elsewhere shift to onsets, transcript phrases, energy, silence, or manual anchors. Never extrapolate a grid through a failed interval.

Non-rhythmic

Use reviewed speech/lyric phrases, silence, energy contour, spectral change, and manually confirmed macro anchors. This suits rubato, ambient work, spoken word, sparse recordings, and free improvisation.

Map features to visual behavior

Production heuristics:

  • Give each feature a limited visual responsibility.
  • Normalize continuous values with robust per-track or per-section statistics, not one extreme maximum.
  • Smooth noise and use hysteresis before switching visual states.
  • Clamp every parameter and define neutral behavior for missing/invalid values.
  • Use macro section anchors, medium phrase anchors, and sparse accents rather than cutting on every event.
  • Reserve low-energy or silent spans for holds, reading, resets, and visual breathing room.
  • Protect lyric lines and important vocal phrases from competing cuts and overlays.
  • Preserve one cue map across aspect ratios; recompose layout, not timing.

Avoid quantitative-looking mappings that imply false measurement. If energy maps to scale, define the exact bounded range and do not call it emotion.

Deterministic frame timing

Keep exact rates such as $30000/1001$ rational. For each cue, define a rounding policy. A conservative visual-response policy assigns the cue to the first output frame whose presentation time is at or after the audio event:

$$ f = \left\lceil t_{event} \cdot \frac{fps_{num}}{fps_{den}} \right\rceil $$

Record both the original event time and mapped frame. Verify actual frame presentation timestamps after encoding because an encoder may duplicate or drop frames to satisfy constant-frame-rate output.

The visual state at frame $f$ must derive from cue data and frame time, not wall-clock playback. Seed procedural mappings and freeze analyzer output before distributed rendering.

Show full SKILL.md (567 more words)Show less

Lyrics, vocals, and captions

Treat transcript or lyric timing as a separate evidence stream. Do not use pitch voicing as vocal detection. Human-review names, lyrics, line boundaries, and timing before they control typography.

Prerecorded synchronized media with meaningful speech needs accurate captions. Captions should include meaningful non-speech sound where needed. A stylized lyric layer does not automatically replace accessible captions or transcript.

Safety, rights, and provenance

WCAG 2.2 SC 2.3.1 limits flashing above three times in one second unless below general/red-flash thresholds. Test loops while looping and at the largest intended scale. Reduced motion does not make unsafe flashing safe.

Provide a lower-motion version when large displacement, zoom, shake, or dense event response may cause discomfort. Reduce event density and travel, not merely output FPS.

The musical work, lyrics, and sound recording can carry separate rights. Possession of a file does not grant synchronization, adaptation, or distribution rights. Record source and transformation provenance; do not claim metadata proves authenticity or permission.

QA

  1. Re-run identical input/configuration and compare canonical cue-map hashes.
  2. Audition click-marked candidate and promoted anchors.
  3. Inspect tempo-level ambiguity and failed-confidence windows.
  4. Compare source timing against decoded/trimmed master and non-zero timestamps.
  5. Inspect frames immediately before, on, and after every macro anchor.
  6. Verify aspect variants share cue IDs and frame assignments.
  7. Probe output frame PTS, duration, and audio sync.
  8. Review captions/lyrics, reduced motion, and flash safety.
  9. Preserve source hash, profile, cue map, mapping table, render versions, and approvals.

Example 1: constant-tempo electronic visualizer

This is a complete example, not a mandatory formula.

Intent: 30-second 9:16 and 16:9 visualizer from an authorized 44.1 kHz stereo instrumental at 30 fps.

Approach: confirm stable tempo with a documented analyzer and inspect local pulse. Retain all beat/onset candidates. Promote every fourth beat for medium choreography, reviewed top-strength transients for sparse accents, and reviewed recurrence changes for macro transitions. Beat phase drives scale only from 1.000 to 1.035; low-band energy controls bounded depth; centroid controls a narrow texture-density range. A six-beat low-energy span holds title copy.

Map all anchors to the first frame at or after their sample time. Both aspect variants use identical cue IDs. QA click tracks, rerun hashes, PTS, duration, final-size text, and full-screen flashing.

Likely failure: half/double-tempo ambiguity. Repair by documenting the metrical level selected for visual periodicity without rewriting the raw detections.

Example 2: rubato spoken word

This is a complete example, not a mandatory formula.

Intent: 75-second poem with ambient bed at 24 fps.

Approach: global beat trackers disagree, so use the non-rhythmic route. Reviewed transcript line starts/ends are primary anchors; pauses reset the field; sustained loudness rises control subtle expansion; acoustic-change candidates remain review prompts. Each stanza establishes a stable visual field, and punctuation settles motion. No beat cuts or word-by-word scaling.

The vertical variant reflows text but retains timing. QA muted-caption comprehension, audio-only clarity, every line boundary, breath-adjacent silence, reduced motion, flashing, and separate text/recording rights.

Likely failure: an acoustic boundary lands inside a sentence. Keep it as a low-confidence event and reject it as an editorial anchor.

Sources

Verified 2026-07-12:

© calesthio, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in skills/production/runtime-assembly/audio-reactive-video-composition of calesthio/generative-media-skills.

  • SKILL.md
  • EVAL.md

Open the folder on GitHubat commit 8c85352

Compare with similar skills

Audio Reactive Video Composition next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Audio Reactive Video Composition compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Audio Reactive Video Composition this skillcalesthio/generative-media-skills197—~2.9kAutomated safety check: PassMIT
Paper Collage Explainer Generatortl2012tl/comfyUI-llama-TE2414 repos~5.2kAutomated safety check: PassNone
Wonder BlocksKhan/wonder-blocks163—~3.2kAutomated safety check: PassMIT
Vibe Matchercuriositech/some_claude_skills244—~3.3kAutomated safety check: PassMIT
Design Tokenshashgraph-online/awesome-codex-plugins1.3k—~1.2kAutomated safety check: PassApache-2.0
Anthropic Brand Stylinganthropics/skills180k30 repos~559Automated safety check: PassApache-2.0

Similar skills

  • Paper Collage Explainer Generator

    tl2012tl/comfyUI-llama-TE

    For creators, educators, and social-video editors who need a tactile paper-collage language for narration, knowledge points, opinions, or abstract topics.

    241 GitHub starsUsed in 4 repos~5.2k tokens
    Media & CreativeAuto-check passed
  • Wonder Blocks

    Khan/wonder-blocks

    Implements user interfaces using the Wonder Blocks (WB) design system — Khan Academy's React component library.

    163 GitHub stars~3.2k tokensUpdated today
    Frontend & DesignAuto-check passed
  • Vibe Matcher

    curiositech/some_claude_skills

    Synesthete designer that translates emotional vibes and brand keywords into concrete visual DNA (colors, typography, layouts, interactions).

    244 GitHub stars~3.3k tokensUpdated 1 mo ago
    Frontend & DesignAuto-check passed
  • Design Tokens

    hashgraph-online/awesome-codex-plugins

    A skill your agent uses when the user asks for design tokens, DTCG tokens, theme systems, color/typography/spacing/radius/elevation/motion tokens, translating tokens, exporting tokens, making a…

    1.3k GitHub stars~1.2k tokensUpdated today
    Frontend & DesignAuto-check passed
  • Anthropic Brand Styling

    anthropics/skills

    Official

    Applies Anthropic's brand colors and fonts to artifacts such as PowerPoint slides, using fixed hex values for text and accents, Poppins headings and Lora body text.

    180k GitHub starsUsed in 30 repos~559 tokens
    Media & CreativeAuto-check passed
  • Minimal Zine Poster Generator

    LiamGvchi/gc-minimal-zine-poster

    Creates or analyzes quiet, paper-texture zine posters with big negative space, one color accent and experimental type, returning an image prompt and the generated poster.

    7.3k GitHub stars~2.9k tokensUpdated 1 mo ago
    Media & CreativeAuto-check passed

More from calesthio/generative-media-skills

All 26 skills in this repo
  • 3D Asset Production

    calesthio/generative-media-skills

    A skill your agent uses to turn generated, captured, scanned, or modeled 3D output into production-ready standalone assets for DCC, real-time engine, web, or interchange delivery.

    197 GitHub stars~9.7k tokensUpdated 2 mo ago
    Auto-check passed
  • Audio Mixing Mastering

    calesthio/generative-media-skills

    Provider-independent audio mixing and mastering direction for AI agents finishing generated videos, ads, trailers, explainers, podcasts, recuts, avatar clips, music videos, documentaries, and social…

    197 GitHub stars~7k tokensUpdated 2 mo ago
    Auto-check passed
  • Captions Media Accessibility

    calesthio/generative-media-skills

    Provider-independent captions and media accessibility direction for AI agents producing or finishing generated videos, ads, social clips, explainers, avatar videos, documentaries, podcasts/video…

    197 GitHub stars~6.9k tokensUpdated 2 mo ago
    Auto-check passed
  • Comfyui Media Workflows

    calesthio/generative-media-skills

    Provider-independent production workflow for agents assembling, auditing, executing, and handing off ComfyUI node-graph workflows for image, video, upscale, inpaint, conditioning, and batch media…

    197 GitHub stars~8.5k tokensUpdated 2 mo ago
    Auto-check passed
  • Ffmpeg Media Finishing

    calesthio/generative-media-skills

    Provider-independent FFmpeg finishing workflow for AI agents preparing generated or edited media deliverables.

    197 GitHub stars~8.2k tokensUpdated 2 mo ago
    Auto-check passed
  • Generated Media QA

    calesthio/generative-media-skills

    Provider-independent quality assurance for AI-generated and AI-assisted media.

    197 GitHub stars~8k tokensUpdated 2 mo ago
    Auto-check passed

Questions about Audio Reactive Video Composition

What does Audio Reactive Video Composition do?

Provider-independent production guidance for translating measured audio features into deterministic video timing and motion. Audio Reactive Video Composition is an agent skill from calesthio/generative-media-skills. Provider-independent production guidance for translating measured audio features into deterministic video timing and motion.

When should I use Audio Reactive Video Composition?

Audio Reactive Video Composition fits situations like: spectrum-reactive visualizers; generated compositions; not for music-video concept direction; audio generation.

How do I install Audio Reactive Video Composition in Claude Code?

Run `npx skills add calesthio/generative-media-skills --skill audio-reactive-video-composition -a claude-code`. Or copy the skill folder (skills/production/runtime-assembly/audio-reactive-video-composition in calesthio/generative-media-skills) into .claude/skills/audio-reactive-video-composition in your project. Claude Code loads it when a task matches its description.

How do I install Audio Reactive Video Composition in Codex?

Run `npx skills add calesthio/generative-media-skills --skill audio-reactive-video-composition -a codex`. Or copy the skill folder (skills/production/runtime-assembly/audio-reactive-video-composition in calesthio/generative-media-skills) into .agents/skills/audio-reactive-video-composition in your project. Codex loads it when a task matches its description.

Can I use Audio Reactive Video Composition in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add calesthio/generative-media-skills --skill audio-reactive-video-composition -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/audio-reactive-video-composition, .gemini/skills/audio-reactive-video-composition, .github/skills/audio-reactive-video-composition and .opencode/skills/audio-reactive-video-composition in your project.

What does Audio Reactive Video Composition need to run?

SKILL.md names no scripts, command-line tools or credentials: Audio Reactive Video Composition is instructions for the agent only.

Does Audio Reactive Video Composition access the network?

SKILL.md names 8 domains. As links in the text: w3.org, librosa.org, essentia.upf.edu, mir-eval.readthedocs.io, ee.columbia.edu, ffmpeg.org, tech.ebu.ch and copyright.gov. This is read from the text; nothing was executed.

Is Audio Reactive Video Composition safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Audio Reactive Video Composition use?

Audio Reactive Video Composition is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Audio Reactive Video Composition use?

About 2.9k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Audio Reactive Video Composition?

Skills that share tags, products or a category with Audio Reactive Video Composition: Paper Collage Explainer Generator (tl2012tl/comfyUI-llama-TE, 241 stars), Wonder Blocks (Khan/wonder-blocks, 163 stars), Vibe Matcher (curiositech/some_claude_skills, 244 stars) and Design Tokens (hashgraph-online/awesome-codex-plugins, 1.3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Audio Reactive Video Composition?

calesthio (a GitHub user) maintains it in calesthio/generative-media-skills, which has 197 GitHub stars. The repository holds 26 skills in this directory. The repository was last updated on July 14, 2026.

Source: calesthio/generative-media-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.