Agent skill

Hyperframes Media

by boraoztunc in boraoztunc/skills

Asset preprocessing for HyperFrames compositions — text-to-speech narration (Kokoro), audio/video transcription (Whisper), and background removal for transparent overlays (u2net).

Apache-2.0Auto-check passedMedia & Creative

Install Hyperframes Media

skills CLI
$ npx skills add boraoztunc/skills --skill hyperframes-media -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install boraoztunc/skills hyperframes-media --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/boraoztunc/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/hyperframes-media .claude/skills/hyperframes-media && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
hyperframes-media
GitHub stars
393
Used in
2 other repos
Token cost
~3.7k tokens
SKILL.md length
1,181 words
Files
1
Skills in repo
52
Repo updated
First seen
Licence
Apache-2.0

At a glance

Asset preprocessing for HyperFrames compositions — text-to-speech narration (Kokoro), audio/video transcription (Whisper), and background removal for transparent overlays (u2net).

  • Works in 3 steps: Language known and non-English → --model… → Language known and English → --model… → Language unknown → --model small (no…
  • Generating voiceover from text
  • SKILL.md covers Text-to-Speech (tts), Transcription (transcribe), Background Removal… and TTS → Transcribe → Captions
  • Calls npx, brew and apt-get

What it does

Hyperframes Media is an agent skill from boraoztunc/skills. Asset preprocessing for HyperFrames compositions — text-to-speech narration (Kokoro), audio/video transcription (Whisper), and background removal for transparent overlays (u2net). Use when generating voiceover from text, transcribing speech for captions, removing the background from a video or image to use as a transparent overlay, choosing a TTS voice or whisper model, or chaining these (TTS → transcribe → captions). Each command downloads its own model on first run.

Its SKILL.md is about 3.7k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Media & Creative, covering Text to speech and voice, Transcription and Motion graphics. It works with HeyGen. The repository describes itself as: Claude Code skills for copywriting, SEO, design, and more. The licence is Apache-2.0.

When your agent uses it

  • Generating voiceover from text
  • Transcribing speech for captions
  • Removing the background from a video
  • Image to use as a transparent overlay

Example prompts

  • “/hyperframes-media”

Requirements

  • Python 3
  • Node.js

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. Language known and non-English → --model small --language (no .en suffix)
  2. Language known and English → --model small.en
  3. Language unknown → --model small (no .en, no --language) — whisper auto-detects

What it can do on your machine

Read from SKILL.md and the folder at commit 645553c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • npx
    • brew
    • apt-get
    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use npx and pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Hyperframes Media loads about 3.7k tokens when it runs. Until then it costs about 123 tokens; SKILL.md has 1,181 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~123
When it runs · the whole SKILL.md, loaded when a task matches
~3.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from boraoztunc/skills at commit 645553c, republished under its Apache-2.0 licence (© boraoztunc). 1,181 words, ~3,655 tokens.

Download SKILL.mdSave it as .claude/skills/hyperframes-media/SKILL.md (or your agent's skills folder).
name
hyperframes-media
description
Asset preprocessing for HyperFrames compositions — text-to-speech narration (Kokoro), audio/video transcription (Whisper), and background removal for transparent overlays (u2net). Use when generating voiceover from text, transcribing speech for captions, removing the background from a video or image to use as a transparent overlay, choosing a TTS voice or whisper model, or chaining these (TTS → transcribe → captions). Each command downloads its own model on first run.

HyperFrames Media Preprocessing

Three CLI commands that produce assets for compositions: tts (speech), transcribe (timestamps), and remove-background (transparent video). Each downloads a model on first run and caches it under ~/.cache/hyperframes/. Drop the output into the project, then reference it from the composition HTML — see the hyperframes skill for the audio/video element conventions.

Text-to-Speech (tts)

Generate speech audio locally with Kokoro-82M. No API key.

bash
npx hyperframes tts "Text here" --voice af_nova --output narration.wav
npx hyperframes tts script.txt --voice bf_emma --output narration.wav
npx hyperframes tts --list                       # all 54 voices
Voice Selection

Match voice to content. Default is af_heart.

Content typeVoiceWhy
Product demoaf_heart/af_novaWarm, professional
Tutorial / how-toam_adam/bf_emmaNeutral, easy to follow
Marketing / promoaf_sky/am_michaelEnergetic or authoritative
Documentationbf_emma/bm_georgeClear British English, formal
Casual / socialaf_heart/af_skyApproachable, natural
Multilingual

Voice IDs encode language in the first letter: a=American English, b=British English, e=Spanish, f=French, h=Hindi, i=Italian, j=Japanese, p=Brazilian Portuguese, z=Mandarin. The CLI auto-detects the phonemizer locale from the prefix — no --lang needed when the voice matches the text.

bash
npx hyperframes tts "La reunión empieza a las nueve" --voice ef_dora --output es.wav
npx hyperframes tts "今日はいい天気ですね" --voice jf_alpha --output ja.wav

Use --lang only to override auto-detection (stylized accents). Valid codes: en-us, en-gb, es, fr-fr, hi, it, pt-br, ja, zh. Non-English phonemization requires espeak-ng system-wide (brew install espeak-ng / apt-get install espeak-ng).

Speed
  • 0.7-0.8 — tutorial, complex content, accessibility
  • 1.0 — natural pace (default)
  • 1.1-1.2 — intros, transitions, upbeat content
  • 1.5+ — rarely appropriate; test carefully
Long Scripts

For more than a few paragraphs, write to a .txt file and pass the path. Inputs over ~5 minutes of speech may benefit from splitting into segments.

Requirements

Python 3.8+ with kokoro-onnx and soundfile (pip install kokoro-onnx soundfile). Model downloads on first use (~311 MB + ~27 MB voices, cached in ~/.cache/hyperframes/tts/).

Transcription (transcribe)

Produce a normalized transcript.json with word-level timestamps.

bash
npx hyperframes transcribe audio.mp3
npx hyperframes transcribe video.mp4 --model small --language es
npx hyperframes transcribe subtitles.srt          # import existing
npx hyperframes transcribe subtitles.vtt
npx hyperframes transcribe openai-response.json
Language Rule (Non-Negotiable)

Never use .en models unless the user explicitly states the audio is English. .en models (small.en, medium.en) translate non-English audio into English instead of transcribing it. This silently destroys the original language.

  1. Language known and non-English → --model small --language <code> (no .en suffix)
  2. Language known and English → --model small.en
  3. Language unknown → --model small (no .en, no --language) — whisper auto-detects

Default model is small, not small.en.

Model Sizes
ModelSizeSpeedWhen to use
tiny75 MBFastestQuick previews, testing pipeline
base142 MBFastShort clips, clear audio
small466 MBModerateDefault — most content
medium1.5 GBSlowImportant content, noisy audio, music
large-v33.1 GBSlowestProduction quality

Music with vocals: start at medium minimum; produced tracks often need manual SRT/VTT import. For caption-quality checks (mandatory after every transcription), the cleaning JS, retry rules, and the OpenAI/Groq API import path, see hyperframes/references/transcript-guide.md.

Output Shape

Compositions consume a flat array of word objects. The id field (w0, w1, ...) is added during normalization for stable references in caption overrides; it's optional for backwards compatibility.

json
[
  { "id": "w0", "text": "Hello", "start": 0.0, "end": 0.5 },
  { "id": "w1", "text": "world.", "start": 0.6, "end": 1.2 }
]

Background Removal (remove-background)

Remove the background from a video or image so the subject (typically a person — avatar, presenter, talking head) sits as a transparent overlay in a composition.

bash
npx hyperframes remove-background subject.mp4 -o transparent.webm  # default: VP9 alpha WebM
npx hyperframes remove-background subject.mp4 -o transparent.mov   # ProRes 4444 (editing)
npx hyperframes remove-background portrait.jpg -o cutout.png       # single-image cutout
npx hyperframes remove-background subject.mp4 -o subject.webm \
  --background-output plate.webm                                   # both layers in one pass
npx hyperframes remove-background subject.mp4 -o transparent.webm --device cpu
npx hyperframes remove-background --info                           # detected providers

Uses u2net_human_seg (MIT). First run downloads ~168 MB of weights to ~/.cache/hyperframes/background-removal/models/.

Layer separation (--background-output)

Pass --background-output (or -b) to emit a second transparent video alongside the cutout: same source RGB, alpha is 255 − mask instead of mask. The cutout is the subject with a transparent background; the plate is the original surroundings with a transparent hole where the subject was.

FileAlpha is…Use it for
-o subject.webmThe mask — subject opaque, background transparentForeground layer, place on top
--background-output plate.webmInverse — surroundings opaque, subject region transparentBottom layer; put text or graphics between this and the subject

Both outputs share the same --quality preset and run from a single inference pass — encode cost roughly doubles, segmentation cost stays the same. Only valid for video inputs and .webm/.mov outputs.

Hole-cut plate, not an inpainted clean plate. The subject region in plate.webm is fully transparent — composite something opaque under it to fill the hole. The single test for whether --background-output is the right tool: will anything ever be visible through the subject's silhouette where the subject used to be?

Use caseRight tool
Text/graphics between the cutout and the plate (this command's reason for existing)Hole-cut (--background-output)
Subject onto an unrelated sceneJust subject.webm; ignore the plate
Show the room without the person, alone over no other contentClean plate — needs an inpainter (LaMa, ProPainter, E2FGVI). Not this command.
Replace the subject with a different subjectClean plate — same as above

If a user asks for "the room with the person removed" and intends to display it standalone, do not reach for --background-output. Tell them they need an inpainter.

Typical layered composition (the canonical hole-cut use case):

html
<!-- z=1 the inverse-alpha plate fills everything except the subject region -->
<video
  src="plate.webm"
  data-start="0"
  data-duration="6"
  data-track-index="0"
  muted
  playsinline
></video>

<!-- z=2 graphics / text live between the two layers -->
<h1 id="headline" style="z-index:2; ...">MAKE IT IN HYPERFRAMES</h1>

<!-- z=3 the cutout floats the subject back over the headline -->
<div class="cutout-wrap" style="position:absolute;inset:0;z-index:3">
  <video
    src="subject.webm"
    data-start="0"
    data-duration="6"
    data-track-index="1"
    muted
    playsinline
  ></video>
</div>

This is functionally equivalent to the text-behind-subject pattern below, but you don't need the original presenter.mp4 in the project — the plate replaces it. Useful when you want to ship just the two transparent layers and let the user drop arbitrary content between them.

Show full SKILL.md (394 more words)Show less
Output Format
FormatWhen
.webm (VP9 + alpha)Default. Compositions play this directly via <video>.
.mov (ProRes 4444)Editing in DaVinci/Premiere/FCP. Large files.
.pngSingle-image cutout (still subject, layered over a backdrop).

Chrome decodes VP9 alpha natively, so the .webm plugs into a composition like any other muted-autoplay video — see the hyperframes skill for the <video> track conventions.

Quality presets

--quality fast|balanced|best controls only the VP9 encoder's CRF — segmentation quality is fixed.

PresetCRFWhen
fast30Iterating, smaller file, looser color match
balanced18Default. Visually identical for most uses
best12Master / final delivery. Largest file, tightest match
Compositing patterns — pick the right one

The cutout webm is a re-encoded copy of the source mp4's RGB. That choice has consequences depending on what you put behind it:

PatternWhat's behind the cutoutResult
Cutout over a different scene (most common)Static image, gradient, or unrelated videoLooks great. The cutout's RGB is the only source of the subject — no doubling, no edge halo. This is what remove-background is built for.
Cutout over its own source mp4 (text-behind-subject)Same mp4 the cutout was generated fromTwo RGB sources for the same person. At default --quality balanced (crf 18) the doubling is barely visible; at --quality fast (crf 30) you'll see a faint color shift / edge halo. Use --quality best (crf 12) for masters.
Cutout over a different take of the same personFootage of the same subjectWill look like two separate people overlapping. Don't do this.

Text-behind-subject (headline behind a presenter):

html
<video
  src="presenter.mp4"
  id="bg"
  data-start="0"
  data-duration="6"
  data-track-index="0"
  muted
  playsinline
></video>
<h1 id="headline" style="z-index:2; ...">MAKE IT IN HYPERFRAMES</h1>
<div class="cutout-wrap" style="position:absolute;inset:0;z-index:3;opacity:0">
  <video
    src="presenter.webm"
    data-start="0"
    data-duration="6"
    data-track-index="1"
    muted
    playsinline
  ></video>
</div>

Two key rules:

  1. Wrap the cutout video in a non-timed <div> and animate the wrapper's opacity, not the video element's. The framework forces opacity:1 on active clips (any element with data-start/data-duration), so animating the video's opacity directly is silently overridden. The wrapper has no data-* attributes, so it's owned by your CSS/GSAP.
  2. Both videos use data-start="0" and data-media-start="0" so the framework decodes them in sync from t=0. Late-mounting the cutout (data-start=3.3) introduces a seek + warm-up that lands a frame off the base mp4 — visible as one frame of misalignment at the cut.

Then GSAP-flip the wrapper opacity at the cut: tl.set(cutoutWrap, { opacity: 1 }, 3.3).

TTS → Transcribe → Captions

When there's no pre-recorded voiceover, generate one and transcribe it back to get word-level timestamps for captions:

bash
npx hyperframes tts script.txt --voice af_heart --output narration.wav
npx hyperframes transcribe narration.wav   # → transcript.json

Whisper extracts precise word boundaries from the generated audio, so caption timing matches delivery without hand-tuning.

© boraoztunc, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in hyperframes-media of boraoztunc/skills.

Open the folder on GitHubat commit 645553c

Used in 2 other repositories

We found 3 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 2 other GitHub owners. This page covers the copy in boraoztunc/skills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Hyperframes Media next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Hyperframes Media compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Hyperframes Media this skillboraoztunc/skills3932 repos~3.7kAutomated safety check: PassApache-2.0
Hyperframesscott-fryxell/brayness124—~3.2kAutomated safety check: PassMIT
Hyperframes Mediacosmicstack-labs/mercury-agent-skills476—~1.7kAutomated safety check: PassMIT
Hyperframes CLIaiskillstore/marketplace4301 repos~2.7kAutomated safety check: PassNone
HyperFrames Media Useheygen-com/hyperframes58k2 repos~2.1kAutomated safety check: PassApache-2.0
Hyperframes Mediachmonitor/chmonitor2981 repos~2.8kAutomated safety check: NotesGPL-3.0

Similar skills

  • Hyperframes

    scott-fryxell/brayness

    Create video compositions, animations, title cards, overlays, captions, voiceovers, audio-reactive visuals, and scene transitions in HyperFrames HTML.

    124 GitHub stars~3.2k tokensUpdated 9 days ago
    Media & CreativeAuto-check passed
  • Hyperframes Media

    cosmicstack-labs/mercury-agent-skills

    Asset preprocessing for HyperFrames compositions — local text-to-speech narration (Kokoro-82M, no API key), audio/video transcription (Whisper), and background removal for transparent overlays…

    476 GitHub stars~1.7k tokensUpdated 1 mo ago
    Media & CreativeAuto-check passed
  • Hyperframes CLI

    aiskillstore/marketplace

    Use the HyperFrames CLI development loop: init, add, catalog, capture, lint, check, snapshot, compare, grade-compare, preview, play, present, beats, keyframes, single or batch render, publish…

    430 GitHub starsUsed in 1 repo~2.7k tokens
    Media & CreativeAuto-check passed
  • HyperFrames Media Use

    heygen-com/hyperframes

    Finds, generates and edits media for HyperFrames video projects: music, sound effects, images, icons, logos, voiceovers, captions and color grades.

    58k GitHub starsUsed in 2 repos~2.1k tokens
    Media & CreativeAuto-check passed
  • Hyperframes Media

    chmonitor/chmonitor

    Audio and media assets for HyperFrames compositions, produced by one shared audio engine (scripts/audio.mjs) — multi-provider TTS (HeyGen / ElevenLabs / Kokoro local), background music + sound…

    298 GitHub starsUsed in 1 repo~2.8k tokens
    Media & CreativeAuto-check: notes
  • Hyperframes CLI

    nateherkai/hyperframes-student-kit

    HyperFrames CLI tool — hyperframes init, lint, preview, render, transcribe, tts, doctor, browser, info, upgrade, compositions, docs, benchmark.

    1.2k GitHub starsUsed in 3 repos~1.2k tokens
    Media & CreativeAuto-check passed

More from boraoztunc/skills

All 52 skills in this repo
  • Gsap

    boraoztunc/skills

    GSAP animation reference for HyperFrames. An agent skill from boraoztunc/skills.

    393 GitHub starsUsed in 5 repos~1.9k tokens
    Auto-check passed
  • Hyperframes

    boraoztunc/skills

    Create video compositions, animations, title cards, overlays, captions, voiceovers, audio-reactive visuals, and scene transitions in HyperFrames HTML.

    393 GitHub starsUsed in 9 repos~7.6k tokens
    Auto-check passed
  • Remotion To Hyperframes

    boraoztunc/skills

    Translate an existing Remotion (React-based) video composition into a HyperFrames HTML composition.

    393 GitHub stars~2.2k tokensUpdated 1 mo ago
    Auto-check passed
  • Animejs

    boraoztunc/skills

    Anime.js adapter patterns for HyperFrames. An agent skill from boraoztunc/skills.

    393 GitHub starsUsed in 2 repos~828 tokens
    Auto-check passed
  • Beam Glow States

    boraoztunc/skills

    Create React loading, processing, selected, current, focus, and pressed states with the border-beam package's animated edge glow.

    393 GitHub starsUsed in 2 repos~3.2k tokens
    Auto-check passed
  • Minimal Zine Poster

    boraoztunc/skills

    Compile a theme, sentence, object, mood, article idea, or photo into a quiet Japanese/Korean zine-style editorial poster — tall aged paper, large negative space, one small image anchor, experimental…

    393 GitHub stars~2.5k tokensUpdated 1 mo ago
    Auto-check passed

Works with

Questions about Hyperframes Media

What does Hyperframes Media do?

Asset preprocessing for HyperFrames compositions — text-to-speech narration (Kokoro), audio/video transcription (Whisper), and background removal for transparent overlays (u2net). Hyperframes Media is an agent skill from boraoztunc/skills. Asset preprocessing for HyperFrames compositions — text-to-speech narration (Kokoro), audio/video transcription (Whisper), and background removal for transparent overlays (u2net).

When should I use Hyperframes Media?

Hyperframes Media fits situations like: generating voiceover from text; transcribing speech for captions; removing the background from a video; image to use as a transparent overlay.

How do I install Hyperframes Media in Claude Code?

Run `npx skills add boraoztunc/skills --skill hyperframes-media -a claude-code`. Or copy the skill folder (hyperframes-media in boraoztunc/skills) into .claude/skills/hyperframes-media in your project. Claude Code loads it when a task matches its description.

How do I install Hyperframes Media in Codex?

Run `npx skills add boraoztunc/skills --skill hyperframes-media -a codex`. Or copy the skill folder (hyperframes-media in boraoztunc/skills) into .agents/skills/hyperframes-media in your project. Codex loads it when a task matches its description.

Can I use Hyperframes Media in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add boraoztunc/skills --skill hyperframes-media -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/hyperframes-media, .gemini/skills/hyperframes-media, .github/skills/hyperframes-media and .opencode/skills/hyperframes-media in your project.

What does Hyperframes Media need to run?

Going by SKILL.md and its folder, Hyperframes Media needs the command-line tools its instructions call (npx, brew, apt-get and pip). Our summary lists: Python 3; Node.js.

Does Hyperframes Media access the network?

SKILL.md contains no URLs. Its commands use npx and pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Hyperframes Media safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Hyperframes Media use?

Hyperframes Media is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Hyperframes Media use?

About 3.7k tokens (SKILL.md is roughly 15k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Hyperframes Media?

Skills that share tags, products or a category with Hyperframes Media: Hyperframes (scott-fryxell/brayness, 124 stars), Hyperframes Media (cosmicstack-labs/mercury-agent-skills, 476 stars), Hyperframes CLI (aiskillstore/marketplace, 430 stars) and HyperFrames Media Use (heygen-com/hyperframes, 58k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Hyperframes Media?

boraoztunc (a GitHub user) maintains it in boraoztunc/skills, which has 393 GitHub stars. The repository holds 52 skills in this directory. The repository was last updated on August 15, 2026.

Source: boraoztunc/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.