Short Form Edit
nateherkai/hyperframes-student-kit
Turn talking-head footage into a finished reel, YouTube Short, or short advertisement with curiosity-led openings, earned payoffs, story-driven cuts, transcript-synced motion graphics, moving…
Mixes audio already placed in a HyperFrames composition: fades, gain, ducking under a voiceover, effect chains, automation and shared submix buses.
$ npx skills add heygen-com/hyperframes --skill hyperframes-audio -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install heygen-com/hyperframes hyperframes-audio --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/heygen-com/hyperframes.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/hyperframes-audio .claude/skills/hyperframes-audio && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "hyperframes-audio" agent skill from https://github.com/heygen-com/hyperframes/tree/main/skills/hyperframes-audio into .claude/skills/hyperframes-audio/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "hyperframes-audio", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/heygen-com/hyperframes/tree/main/skills/hyperframes-audioType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add heygen-com/hyperframes --skill hyperframes-audio -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install heygen-com/hyperframes hyperframes-audio --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/heygen-com/hyperframes.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/hyperframes-audio .agents/skills/hyperframes-audio && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "hyperframes-audio" agent skill from https://github.com/heygen-com/hyperframes/tree/main/skills/hyperframes-audio into .agents/skills/hyperframes-audio/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "hyperframes-audio", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add heygen-com/hyperframes --skill hyperframes-audio -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install heygen-com/hyperframes hyperframes-audio --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/heygen-com/hyperframes.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/hyperframes-audio .cursor/skills/hyperframes-audio && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "hyperframes-audio" agent skill from https://github.com/heygen-com/hyperframes/tree/main/skills/hyperframes-audio into .cursor/skills/hyperframes-audio/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "hyperframes-audio", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/heygen-com/hyperframes.git --path skills/hyperframes-audio--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add heygen-com/hyperframes --skill hyperframes-audio -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install heygen-com/hyperframes hyperframes-audio --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/heygen-com/hyperframes.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/hyperframes-audio .gemini/skills/hyperframes-audio && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "hyperframes-audio" agent skill from https://github.com/heygen-com/hyperframes/tree/main/skills/hyperframes-audio into .gemini/skills/hyperframes-audio/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "hyperframes-audio", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install heygen-com/hyperframes hyperframes-audioInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add heygen-com/hyperframes --skill hyperframes-audio -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/heygen-com/hyperframes.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/hyperframes-audio .github/skills/hyperframes-audio && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "hyperframes-audio" agent skill from https://github.com/heygen-com/hyperframes/tree/main/skills/hyperframes-audio into .github/skills/hyperframes-audio/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "hyperframes-audio", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add heygen-com/hyperframes --skill hyperframes-audio -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install heygen-com/hyperframes hyperframes-audio --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/heygen-com/hyperframes.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/hyperframes-audio .opencode/skills/hyperframes-audio && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "hyperframes-audio" agent skill from https://github.com/heygen-com/hyperframes/tree/main/skills/hyperframes-audio into .opencode/skills/hyperframes-audio/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "hyperframes-audio", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
hyperframes-audioMixes audio already placed in a HyperFrames composition: fades, gain, ducking under a voiceover, effect chains, automation and shared submix buses.
This skill handles the mix for audio that is already in a HyperFrames video composition: fade-in and fade-out, crossfades, track gain, volume automation, ducking, and carving a music bed out of the way of a voiceover. It treats a mix as a set of relationships between tracks, not a stack of processors, and looks for what two tracks compete over before simply turning one down.
Everything is expressed through three attributes on the audio or video element: data-fx-chain for the effects in signal order, data-automation for envelopes on volume or effect parameters, and data-fx-carve for the carve's own settings so it can be re-derived. Effects include gain, EQ, compressor, limiter, gate, saturation, delay, reverb, chorus, phaser and bitcrush. Preview and render run the same Web Audio graph, so what you hear while scrubbing is what gets written. Several tracks can share one submix bus that carries a chain, a fader and an automation clock.
It does not source or generate audio, which belongs to /media-use, and it does not handle clip timing or track layout, which belongs to /hyperframes-core. References cover attributes, diagnosis, the effects registry and presets, and there is a carve.mjs script with a test file. HyperFrames offers no automatic waveform sync or drift correction.
Read from SKILL.md and the folder at commit f6b3821. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 2 files in scripts/ (JavaScript), which the agent can run.
Shell commands in SKILL.md call:
nodenpxFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use npx, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
HyperFrames Audio loads about 6.4k tokens when it runs, and up to ~16k if it reads all its reference files. Until then it costs about 175 tokens; SKILL.md has 3,492 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from heygen-com/hyperframes at commit f6b3821, republished under its Apache-2.0 licence (© heygen-com). 3,492 words, ~6,370 tokens.
.claude/skills/hyperframes-audio/SKILL.md (or your agent's skills folder). This skill also uses 6 other files; get the full folder from GitHub.Plugin installs: Before setup or freshness commands, follow plugin execution rules when this skill is inside a HyperFrames plugin. Standalone installs keep the update instructions below.
A mix is a set of relationships, not a stack of processors. Two tracks that each sound right alone can be unlistenable together, and the fix is almost never "turn one down" — it is finding what they are fighting over and giving it to whichever one needs it. Every tool here exists to express one of those relationships.
Effects live on the element as data-fx-chain, and preview and render run the
same Web Audio graph — the studio in a live context, the engine in an offline one
inside the browser it already drives. There is one implementation of each effect,
so what you hear while scrubbing is what gets written. You never tune twice.
Clip timing remains /hyperframes-core: audio/video trims and source ranges use
data-start, data-duration, and data-media-start, and crossfades overlap
clips on different tracks. This skill owns placed-track fade-in/fade-out,
crossfade envelopes, track gain/track volume, volume and effect automation,
ducking/voiceover carve, and the effect chain. /media-use owns sourcing,
generation, and preprocessing.
Constant data-playback-rate (0.1..10) is render-safe for picture and
pitch-preserved sound when matching audio/video elements use the same timing,
source offset, and rate. A speed ramp is a rate lane in data-automation
(see docs/reference/speed-ramps); it wins over the constant and keeps pitch
in preview and render. HyperFrames does not
provide automatic waveform sync or drift correction.
For copyable cut/crossfade/retime recipes, use /hyperframes-core → references/creator-editing-recipes.md.
Three attributes carry everything, on the audio/video element itself — or, for
the first two, on an <hf-audio-group> bus (see "One bus for many tracks"):
| Attribute | Holds |
|---|---|
data-fx-chain | the effects, in signal order |
data-automation | envelopes on this track's volume or its effect parameters |
data-fx-carve | the carve's own settings, so it can be re-derived |
The shipped effect families are gain, EQ (highpass, lowpass, peaking, shelves), compressor, limiter, gate, saturate, delay, reverb, chorus, phaser, and bitcrush.
Exact JSON for each, and the rules a lane must satisfy: references/attributes.md.
Every effect with its parameters, ranges and units: references/fx-registry.md.
How to work out what is wrong with a file you cannot hear:
references/diagnosis.md.
Presets, named jobs and one-knob profiles, plus a symptom-to-fix table:
references/presets.md — read that before hand-building a chain, because one
of the presets or named jobs usually already names the problem.
Two authoring surfaces write those attributes; two runtimes read them through the same builders. That shared middle is why preview predicts the render.
flowchart TB
voice["voice track<br/>media file"]
bed["music bed<br/>media file"]
subgraph AUTHOR["Authoring — the only things that write attributes"]
panel["Studio<br/>Voiceover carve control"]
script["scripts/carve.mjs<br/>detects the pair"]
analysis["core/audioCarve.ts<br/>carveProfile · analyseCarveBands<br/>analyseCarveDuck · analyseCarveDynamics"]
panel --> analysis
script --> analysis
end
voice --> analysis
bed --> analysis
subgraph ATTRS["Written onto the bed element"]
carveAttr["data-fx-carve<br/>sources · strength"]
chainAttr["data-fx-chain<br/>peaking xN + gain, tagged fromCarve"]
autoAttr["data-automation<br/>a lane per carved parameter"]
end
analysis --> carveAttr
analysis --> chainAttr
analysis --> autoAttr
subgraph SHARED["One implementation, read by both"]
build["audioFxGraph.ts · buildFxChain"]
sched["audioFxAutomation.ts · scheduleChainAutomation"]
end
chainAttr --> build
autoAttr --> sched
build --> preview["Preview<br/>live AudioContext<br/>attachElementFxChain"]
sched --> preview
build --> render["Render<br/>OfflineAudioContext in the headless browser<br/>applyAudioFxChain"]
sched --> render
preview --> heard["what you hear while scrubbing"]
render --> wav["processed WAV<br/>+ chainTailSeconds so the mix lets the tail through"]
wav --> mix["engine · audioMixer<br/>volume lane baked into the PCM here, not in the graph"]
mix --> out["the rendered mix"]
edit["editing the attribute mid-playback"] -.->|MutationObserver| previewThe carve's own settings are never read at playback — the chain and lanes it
produced are what play. data-fx-carve exists so strength can be changed on an
existing carve instead of guessed back out of the filters.
Inside a carved bed the signal runs through the dips first, then the level match, then anything you built yourself — which is why a limiter you add still acts as the last ceiling:
flowchart LR
src["decoded bed"] --> p1["peaking<br/>400 Hz"]
p1 --> p2["peaking<br/>1 kHz"]
p2 --> p3["peaking<br/>1.6 kHz"]
p3 --> g["gain<br/>level match"]
g --> hand["your own effects<br/>e.g. limiter"]
hand --> dest["track gain, then out"]
l1["lane fx.n1.gain"] -.->|"envelope of the voice's<br/>level in that band"| p1
l4["lane fx.n4.gain"] -.->|"how far the bed<br/>ducks overall"| gThe table below starts from "it sounds boomy" — which presumes somebody already listened and said so. Handed a file and "fix this", you have no such sentence and you cannot listen, so you have to measure. One rule governs all of it:
The absolute spectrum of a single unknown voice cannot be diagnosed. Formants are ±10 dB, fundamentals run 85–255 Hz, and sentences decline 5–6 dB as they end. Every one of those reads as a defect on its own, and every one of them is the speaker.
So compare, and compare against something inside the same file: the clean original if it exists, otherwise the pauses — whatever is audible in a gap is additive, and the gap's spectrum is the channel rather than the voice. Comparing against a published average spectrum or a synthesised control voice does not work: two speakers differ by more than most defects, and both wrong answers in the evaluation behind this guidance came from exactly that.
When there is no original and no usable silence, a static tonal defect is genuinely under-determined. Say so and offer the readings that fit, rather than picking one and building a chain on it.
Commands, traps and worked recipes: references/diagnosis.md. Read it
before diagnosing a file nobody has described.
Once you know the band and the kind, name what is wrong with the audio. Most bad audio is one or two of these, and each has a shipped answer:
| It sounds like | Reach for |
|---|---|
| Hum or thump underneath | rumble-cut, or a highpass at 80 Hz |
| Boomy, chesty | Tame Boominess job (200 Hz) |
| Muffled, behind cardboard | Reduce Mud job (250 Hz) |
| Words hard to make out | Add Clarity job (3 kHz), or carve the bed |
| Harsh and tiring | Soften Harshness job (3.2 kHz) |
| Some words much louder than others | Evenness on a compressor, or Even Out Levels |
| Room tone between sentences | room-gate |
| Voice and music fighting | Voiceover carve — not an EQ on either |
| Dry, recorded nowhere | room-tight or room-natural |
| Just "amateur" | voice-clean, which is four of the above in order |
Full catalogue, what each preset contains, the band vocabulary, and what is
deliberately NOT covered (de-essing, noise removal, tone match):
references/presets.md.
Subtract before you add, level after you filter, relationships after level, character and ceiling last. Each step changes what the next one hears — a compressor set before a high-pass spends its time chasing rumble.
Filters (highpass, lowpass, peaking, lowshelf, highshelf) decide
which frequencies a track is allowed to occupy. This is the first tool for two
sources colliding, because collisions happen in bands: a bed and a voice both
want 1–3 kHz, and taking that from the bed costs the bed far less than turning
the whole thing down costs the mix. A high-pass on a voice is the standard fix
for rumble; a low-pass darkens or muffles deliberately.
Dynamics (gain, compressor, limiter, gate) decide how a track's level
behaves over time. Compression narrows the distance between loud and quiet so the
quiet parts can come up. A limiter is a ceiling — it does not shape anything, it
guarantees nothing gets past. A gate removes what is below a threshold, which is
how you silence room tone between phrases. gain is a plain level stage, and it
is what an automation lane rides when a track has to move out of the way.
Nonlinear (saturate, bitcrush) changes the waveform's shape, which adds
harmonics that were not there. Reach for it when a track needs character or
grit rather than correction — and remember it is generative: it makes a thin
source denser, not cleaner.
Time (delay, reverb, chorus, phaser) puts a track in a space or gives
it width. These are the ones that most easily wreck a mix, because a tail or a
detuned copy occupies the same room a voice needs. Use them on the thing that
should sit behind something else, and keep the wet amount lower than sounds
right in isolation.
The chain is serial: each effect processes what the one before it produced. So corrective filtering goes early, character in the middle, and a limiter last where it can actually act as a ceiling.
The problem it solves. A music bed under a voice makes the voice hard to follow. The reflex is to duck the whole bed, which works and costs the bed all of its presence — the music goes limp for the entire voiceover. But the voice does not need the whole spectrum. It needs the few bands it actually occupies. Carve takes only those, and the bed keeps its low end and its top, so it is still music while the voice is still intelligible.
It is a relationship, not an effect. The settings live on the bed — the track that gets processed — and they name the voices to listen to, exactly as a sidechain compressor does: you select the track that gets quieter and pick what makes it quieter. Never put a carve on a voice track. A voice carved against itself is a bug, not a subtle mix choice.
Every voice, not one of them. sources is a list, because a bed usually runs
under a whole sequence — a narrator, an interview answer, a second presenter. They
are summed onto the bed's own clock before anything is measured (mixCarveSources),
so one analysis covers all of them: the bands come from all the speech there is, and
the envelopes rise wherever any of it is happening. Voices that never play while the
bed does are left out; they cannot mask it.
A carve against more than one clip id is wrong. Group the clips and carve
against the group. This is an invariant, not a tip. Naming clips one by one has
to be exhaustively right and stays right only until the next edit — a fourth
narration clip added later plays outside the carve's awareness, and the bed
fails to duck under it silently. Naming the group instead resolves membership at
analysis time, so a clip added to the group later is covered without touching
sources at all:
<!-- group the narration, then carve the bed against the group -->
<audio id="vo-intro" data-audio-group="voiceover" …></audio>
<audio id="vo-middle" data-audio-group="voiceover" …></audio>
<audio id="vo-outro" data-audio-group="voiceover" …></audio>
<audio id="music" data-fx-carve='{"enabled":true,"sources":["voiceover"],"strength":0.8}' …></audio>A sources list naming two or more plain clip ids instead of a group is caught
by the audio_carve_ungrouped_sources lint rule — it still works, but it is the
version that silently rots when a clip is added.
Keep the carve group a voice group: no bed, no SFX, no music. A group id in
sources resolves to every current member on every analysis, so the group
you name is the group you get later — not the tracks that were measured when it
was written. Two ways that bites:
Both are invisible at the moment the carve is written: the analysis sums the
voices it detected and never round-trips through group resolution, so the first
pass is genuinely correct and only the next one is wrong. So give each role its
own group — music for the bed, voiceover for the narration, sfx for the
hits — and keep the group named in sources holding nothing but voices.
carve.mjs refuses to write the group form when it sees either case, records
clip ids, and says on stderr which member blocked it. The
audio_carve_ungrouped_sources rule then points at the arrangement instead of
the CLI quietly persisting a wider carve than it measured.
A voice that this run left out is not one of these cases and does not block
the group form: carve.mjs only analyses voices that overlap the bed, and
picking up a clip that plays later without an edit to sources is the whole
reason to name the group.
Membership alone is enough to carve against, as above — but add an
<hf-audio-group> element with that id and the group becomes a real submix bus:
one chain, one fader, one automation clock for every member.
<hf-audio-group
id="voiceover"
data-label="Voiceover"
data-volume="0.9"
data-fx-chain='{"version":1,"nodes":[
{"type":"compressor","id":"g1","params":{"threshold":-18,"ratio":3}},
{"type":"peaking","id":"g2","params":{"frequency":3000,"gain":2,"q":1}}]}'
></hf-audio-group>
<audio id="vo-intro" data-audio-group="voiceover" …></audio>
<audio id="vo-middle" data-audio-group="voiceover" …></audio>Reach for the bus when the same treatment belongs on several tracks. Four narration clips that each want the same compressor is four chains to keep in step, and they drift the moment one is edited; on the bus it is one chain, and the compressor sees the whole voice rather than each clip in isolation — which is the point, since a compressor cannot ride a sequence it only hears a third of. Per-clip chains remain right for what is genuinely per-clip: one noisy take that needs its own de-esser.
| On the bus | Does |
|---|---|
data-fx-chain | one chain over the summed members |
data-automation | envelopes on the bus, in COMPOSITION time |
data-volume | one fader for every member (default 1) |
data-label | the display name; falls back to the id |
data-hidden | drops every member from the mix |
Group automation is composition time, not clip time. A bus has no
data-start — members are already at their composition positions when they
reach it — so t: 0 in a group lane is the start of the composition, not of any
clip. A lane on a clip is clip-local; the same numbers mean different instants on
the two, which is the one thing to get right when moving an envelope from a clip
up onto its bus.
A carve stays on the clip. data-fx-carve is not a group attribute. The bed
being carved is a single track, and it is that track which carries
data-fx-carve — pointed AT a group, per the rule above. Group and carve meet in
sources, not on one element. A carve written onto a bus is half an effect
applied twice: the level half measures the bed's own audio, which a bus has none
of, so only the filters survive — and a bus and its members are one signal path,
so the bed then runs through the bus's filters AND its own. The
audio_group_carve_attr lint rule catches it.
One clip is not a bus. A group exists to give several tracks one chain, one
fader and one clock. Wrapping a single clip in a bus buys nothing the clip's own
data-fx-chain does not already do, and it doubles the places a later edit has
to land. The one reason to do it anyway: a bus's automation clock is composition
time, so a single-member bus is how a lane on that clip gets composition-time
timing.
One knob. strength is 0..1 and derives everything: how deep to cut, how
many bands, how wide, how far to favour intelligibility over raw voice energy,
how far the level may drop, how far under the voice to aim. Those six move
together in any real mix — a gentle carve is a shallow cut in few bands with
little ducking, a hard one is deeper in more bands with more — so they are one
relationship written once, in carveProfile. carve.mjs defaults to 0.8 —
six bands from 250 Hz to 2.5 kHz cut about 7 dB each and 15 dB at 1.6 kHz, with
19 dB of level room — because a bed under narration has to get out of the way
first and be music second; 0.25 (a 6 dB dip in three bands, 6 dB of room) kept
the bed present but still let it fight the voice, and was judged too weak in
practice. At 0.5 the dip reaches 10 dB, which is where a carve starts being
heard as an effect rather than as room for the voice. Drop the strength when the
bed is the point and the voice is sparse. 0 is spectral only — one band, no
level match at all.
Carve by default — required whenever music plays under a voice. A bed
under any voice track (narration, avatar speech, interview, voiceover) gets a
carve as part of finishing the mix, not as a polish step to get to if there is
time. Place both tracks, run the command below (default strength 0.8; add
--bed / --voice when detection picks wrong), confirm the written
data-fx-carve, data-fx-chain and data-automation with npx hyperframes check,
and only then render. A volume duck on its own is not a finished mix: it leaves
the voice and the bed fighting in the 1–3 kHz band and costs the bed all of its
presence for the whole voiceover. Skip the carve only when there is no voice for
the music to sit under — a music video, a title card, a montage cut to the track.
It always follows the voice. There is no static mode: a fixed depth thins the bed through every pause, and once you have heard both there is no reason to want it. Every value becomes an envelope of the speech's own level — silence leaves the bed alone, a loud passage pushes the carve to full depth — written as ordinary automation, which is why the lanes show up in the timeline and can be edited afterwards.
Level matching is part of it. Spectral carving cannot fix a bed that is
simply louder than the voice. So the carve also measures how far over the voice
the bed sits and writes a gain stage driven by an envelope. That envelope releases slowly on
purpose — music that snaps back to full the instant a word ends sounds like a
machine doing it.
Running it. In Studio the carve is one module at the top of a track's effect rack — voice, strength, and the analysis it produced, in one card. It is there whenever another track could be the voice, and a bed with exactly one candidate above it is carved by default, at the default strength: that is what a bed under narration wants, and the module is where you change or switch it off. Several candidates leaves the picker waiting rather than guessing. Headless — which is the path when you are authoring a composition rather than editing one:
node <SKILL_DIR>/scripts/carve.mjs --comp index.htmlThat is the whole command. It finds the voice and the bed itself, carves at the default strength, and prints what it decided:
bed music-bed (name looks like music)
voice narration (only track left)
carve strength 0.8, 1 voice
bands 250Hz -7.4dB q2.06, 400Hz -7.4dB q2.06, 630Hz -7.4dB q2.06, 1000Hz -7.4dB q2.06, 1600Hz -14.8dB q2.06, 2500Hz -7.4dB q2.06
level 273-point envelope, floor -19.2 dBName the tracks with --bed / --voice (repeatable) when the automatic choice is
wrong, --strength to push it, --dry-run to see that report and write nothing.
How it picks the tracks. Names first, because that is what you already told it
and the answer is explainable — classifyAudioName in core, the same classifier
Studio's own picker uses, so the two cannot disagree. A track whose id or filename
looks like music (music, bgm, bed, score…) is the bed; everything else that
plays over it and is not SFX-shaped is a voice. Audio elements are preferred: video
counts only when no audio track is left to be the voice, or every B-roll clip in the
composition would read as somebody talking. It refuses when it cannot tell which
track is the bed rather than carving the wrong one — typing one id is cheap.
Same analysis functions as the panel, so the result is identical. Needs ffmpeg
on PATH and @hyperframes/core installed in the project (npm i -D @hyperframes/core) — the CLI inlines core rather than shipping it, so it cannot
be borrowed from there.
What it writes is an ordinary chain of peaking filters plus a gain stage,
tagged fromCarve. That tagging is the whole trick: a re-run replaces the
previous carve and leaves every effect you built by hand — and every lane you
drew by hand — exactly where it was. So re-carving at a new strength is safe and
repeatable, and data-fx-carve exists so the settings can be read back rather
than guessed from the filters.
A lane is a set of breakpoints on one parameter: {t, v} in clip-local seconds
and the parameter's own units. Targets are volume for the track's level, or
fx.<nodeId>.<param> for an effect's knob.
Only some parameters can be automated, and a lane on the others is silently
inert. A knob is automatable when a Web Audio AudioParam backs it. The four
worklet-based effects — compressor, limiter, gate, bitcrush — expose
none at all, so no lane on any of their parameters will ever move: to make a
compressor's behaviour change over time, automate a gain stage before it
instead. references/fx-registry.md marks every parameter.
Almost no static gate covers the mix. The linter reads data-automation for
exactly one conflict — audio_volume_double_automation, a volume lane on a track
that also has a GSAP tween on volume, where the lane wins and the tween is
ignored — plus audio_volume_tween_overrides_gain, an authored data-volume
on a track whose volume is tweened, where the tween's values are absolute and
replace that gain instead of scaling it. Nothing validates the
chain or the effect lanes at all. What
enforces those is the render: a chain it cannot parse fails the whole mix rather
than quietly writing the dry signal, because a mix that sounds plausible and is
wrong is worse than a refusal. Preview is the opposite by design: an unreadable
chain plays dry so the composition stays workable.
A lane pointing at a node the chain does not have is pruned on read, not an
error — so a typo'd nodeId costs you the envelope silently. Read the ids back
out of the chain rather than assuming what was minted.
Effects with a tail (reverb, delay) make the rendered track longer than
its source, and the mix is told how much by the chain. So a bed with reverb no
longer ends exactly at its data-duration; that is expected, not a bug.
Beyond that, a mix is verified by rendering and listening. For a carve: the voice should be legible without the bed sounding hollowed, and the bed should come back up between phrases rather than staying flat. If the bed sounds notched rather than simply quieter under the voice, the strength is too high — that is the one failure mode with an obvious sound.
© heygen-com, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 6 other files (scripts, references) in skills/hyperframes-audio of heygen-com/hyperframes.
Open the folder on GitHubat commit f6b3821
We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in heygen-com/hyperframes, which our catalogue first saw on October 7, 2026.
HyperFrames Audio next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| HyperFrames Audio this skillheygen-com/hyperframes | 58k | 1 repos | ~6.4k | Automated safety check: Pass | Apache-2.0 | |
| Short Form Editnateherkai/hyperframes-student-kit | 1.2k | — | ~5.3k | Automated safety check: Pass | Custom licence | |
| Noti Tiktok Full Textnotivn/AIEV | 126 | — | ~4.8k | Automated safety check: Pass | MIT | |
| Strudel Live-Coding Video Producerheygen-com/hyperframes-community-skills | 175 | — | ~1.3k | Automated safety check: Pass | Apache-2.0 | |
| Noti Tiktok Vnnotivn/AIEV | 126 | — | ~5.4k | Automated safety check: Pass | MIT | |
| Noti Youtube Editnotivn/AIEV | 126 | — | ~5.6k | Automated safety check: Pass | MIT |
nateherkai/hyperframes-student-kit
Turn talking-head footage into a finished reel, YouTube Short, or short advertisement with curiosity-led openings, earned payoffs, story-driven cuts, transcript-synced motion graphics, moving…
notivn/AIEV
Build a Vietnamese vertical TikTok explainer in the "MỔ XẺ PAPER AI" (AI paper dissection) format with HyperFrames (HTML/CSS/GSAP → MP4), Noti.vn style.
heygen-com/hyperframes-community-skills
Turns a Strudel live-coding music track into a 1080x1080 HyperFrames video where the code itself is the picture and each line highlights on the note it plays.
notivn/AIEV
Edit a Vietnamese vertical TikTok video (9:16) with HyperFrames following the Noti.vn/GĐT standard - talking-head + kinetic typography + karaoke captions + zoom/punch-in camera + timestamp-synced…
notivn/AIEV
Build a Vietnamese landscape 16:9 YouTube video (1920×1080) with HyperFrames (HTML/CSS/GSAP → MP4), keeping the Noti.vn/GĐT branding inherited from noti-tiktok-vn.
latent-spaces/brag
Turns the current project website or app into a short, shareable launch video with Hyperframes, planned from the project code with options for tone, format and voiceover.
heygen-com/hyperframes
Collects motion rules, scene blueprints, transitions and runtime adapters for HyperFrames video compositions, with GSAP as the default animation runtime.
heygen-com/hyperframes
Finds, generates and edits media for HyperFrames video projects: music, sound effects, images, icons, logos, voiceovers, captions and color grades.
heygen-com/hyperframes
Adds captions to a single-subject talking-head video without editing the footage, from plain subtitles to cinematic text placed behind the speaker.
heygen-com/hyperframes
Turns a weekly changelog markdown file into a branded HyperFrames video with voiceover, animated mock-UI scenes and captions, using fonts, background and scripts bundled in the skill.
heygen-com/hyperframes
Turns an article, notes or a topic brief into an explainer video whose visuals are invented per scene, built frame by frame in HyperFrames with no footage.
heygen-com/hyperframes
Imports Figma assets, brand tokens, components and motion into a HyperFrames video composition, using the Figma REST API with a connector or native export for shaders.
Works with
Categories
Mixes audio already placed in a HyperFrames composition: fades, gain, ducking under a voiceover, effect chains, automation and shared submix buses. This skill handles the mix for audio that is already in a HyperFrames video composition: fade-in and fade-out, crossfades, track gain, volume automation, ducking, and carving a music bed out of the way of a voiceover. It treats a mix as a set of relationships between tracks, not a stack of processors, and looks for what two tracks compete over before simply turning one down.
HyperFrames Audio fits situations like: fading music in and out under a scene in a HyperFrames video; ducking a music bed so a voiceover stays clear; adding EQ, compression or reverb to a single track; routing several tracks through one shared submix bus with its own fader.
Run `npx skills add heygen-com/hyperframes --skill hyperframes-audio -a claude-code`. Or copy the skill folder (skills/hyperframes-audio in heygen-com/hyperframes) into .claude/skills/hyperframes-audio in your project. Claude Code loads it when a task matches its description.
Run `npx skills add heygen-com/hyperframes --skill hyperframes-audio -a codex`. Or copy the skill folder (skills/hyperframes-audio in heygen-com/hyperframes) into .agents/skills/hyperframes-audio in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add heygen-com/hyperframes --skill hyperframes-audio -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/hyperframes-audio, .gemini/skills/hyperframes-audio, .github/skills/hyperframes-audio and .opencode/skills/hyperframes-audio in your project.
Going by SKILL.md and its folder, HyperFrames Audio needs JavaScript for the scripts in its folder and the command-line tools its instructions call (node and npx). Our summary lists: A HyperFrames composition with audio tracks already placed.
SKILL.md contains no URLs. Its commands use npx, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
HyperFrames Audio is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 6.4k tokens (SKILL.md is roughly 25k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 9.6k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with HyperFrames Audio: Short Form Edit (nateherkai/hyperframes-student-kit, 1.2k stars), Noti Tiktok Full Text (notivn/AIEV, 126 stars), Strudel Live-Coding Video Producer (heygen-com/hyperframes-community-skills, 175 stars) and Noti Tiktok Vn (notivn/AIEV, 126 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
heygen-com (a GitHub organization) maintains it in heygen-com/hyperframes, which has 58,020 GitHub stars. The repository holds 32 skills in this directory. The repository was last updated on October 7, 2026.
Source: heygen-com/hyperframes on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.