Asset preprocessing for HyperFrames compositions — local text-to-speech narration (Kokoro-82M, no API key), audio/video transcription (Whisper), and background removal for transparent overlays…

MITAuto-check passedMedia & Creative

Install Hyperframes Media

skills CLI
$ npx skills add cosmicstack-labs/mercury-agent-skills --skill hyperframes-media -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install cosmicstack-labs/mercury-agent-skills hyperframes-media --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/cosmicstack-labs/mercury-agent-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/categories/development/hyperframes-media .claude/skills/hyperframes-media && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
hyperframes-media
GitHub stars
476
Token cost
~1.7k tokens
SKILL.md length
536 words
Files
1
Skills in repo
12
Repo updated
First seen
Licence
MIT

At a glance

Asset preprocessing for HyperFrames compositions — local text-to-speech narration (Kokoro-82M, no API key), audio/video transcription (Whisper), and background removal for transparent overlays…

  • Works in 3 steps: Language known and non-English → --model… → Language known and English → --model… → Language unknown → --model small (no…
  • Generating voiceover from text
  • SKILL.md covers Text-to-Speech (tts), Transcription (transcribe), Background Removal… and TTS -> Transcribe -> Captions…, plus 1 more section
  • Calls npx and pip

What it does

Hyperframes Media is an agent skill from cosmicstack-labs/mercury-agent-skills. Asset preprocessing for HyperFrames compositions — local text-to-speech narration (Kokoro-82M, no API key), audio/video transcription (Whisper), and background removal for transparent overlays (u2net). Use when generating voiceover from text, transcribing speech for captions, removing background from video/images, choosing TTS voices or whisper models, or chaining TTS - transcribe - captions. Each command downloads its own model on first run.

Its SKILL.md is about 1.7k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Media & Creative, covering Text to speech and voice, Transcription and Motion graphics. It works with HeyGen. The repository describes itself as: A curated registry of reusable Mercury Agent, Open Claw or Hermes Agent skills designed for real developer workflows, persistent memory, and token-efficient execution. The licence is MIT.

When your agent uses it

  • Generating voiceover from text
  • Transcribing speech for captions
  • Removing background from video/images
  • Choosing TTS voices

Example prompts

  • “/hyperframes-media”

Requirements

  • Python 3
  • Node.js

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. Language known and non-English → --model small --language (no .en suffix)
  2. Language known and English → --model small.en
  3. Language unknown → --model small (no .en, no --language) — whisper auto-detects

What it can do on your machine

Read from SKILL.md and the folder at commit 30392fb. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • npx
    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use npx and pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Hyperframes Media loads about 1.7k tokens when it runs. Until then it costs about 117 tokens; SKILL.md has 536 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~117
When it runs · the whole SKILL.md, loaded when a task matches
~1.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from cosmicstack-labs/mercury-agent-skills at commit 30392fb, republished under its MIT licence (© cosmicstack-labs). 536 words, ~1,690 tokens.

Download SKILL.mdSave it as .claude/skills/hyperframes-media/SKILL.md (or your agent's skills folder).
name
hyperframes-media
description
Asset preprocessing for HyperFrames compositions — local text-to-speech narration (Kokoro-82M, no API key), audio/video transcription (Whisper), and background removal for transparent overlays (u2net). Use when generating voiceover from text, transcribing speech for captions, removing background from video/images, choosing TTS voices or whisper models, or chaining TTS -> transcribe -> captions. Each command downloads its own model on first run.
metadata.author
cosmicstack-labs
metadata.version
1.0.0
metadata.category
development
metadata.tags
hyperframes, media, tts, transcription, background-removal, kokoro, whisper, u2net, captions

HyperFrames Media Preprocessing

Three CLI commands that produce assets for compositions: tts (speech), transcribe (timestamps), and remove-background (transparent video). Each downloads a model on first run and caches it under ~/.cache/hyperframes/.


Text-to-Speech (tts)

Generate speech audio locally with Kokoro-82M. No API key required.

bash
npx hyperframes tts "Text here" --voice af_nova --output narration.wav
npx hyperframes tts script.txt --voice bf_emma --output narration.wav
npx hyperframes tts --list                       # list all 54 voices
Voice Selection
Content TypeRecommended VoicesWhy
Product demoaf_heart / af_novaWarm, professional
Tutorial / how-toam_adam / bf_emmaNeutral, easy to follow
Marketing / promoaf_sky / am_michaelEnergetic or authoritative
Documentationbf_emma / bm_georgeClear British English, formal
Casual / socialaf_heart / af_skyApproachable, natural
Multilingual

Voice IDs encode language in the first letter:

  • a = American English, b = British English, e = Spanish
  • f = French, h = Hindi, i = Italian, j = Japanese
  • p = Brazilian Portuguese, z = Mandarin

The CLI auto-detects the phonemizer locale from the prefix — no --lang needed when the voice matches the text.

bash
npx hyperframes tts "La reunión empieza a las nueve" --voice ef_dora --output es.wav
npx hyperframes tts "今日はいい天気ですね" --voice jf_alpha --output ja.wav

Use --lang only to override auto-detection (stylized accents). Valid codes: en-us, en-gb, es, fr-fr, hi, it, pt-br, ja, zh.

Speed
SpeedUse Case
0.7-0.8Tutorial, complex content, accessibility
1.0Natural pace (default)
1.1-1.2Intros, transitions, upbeat content
1.5+Rarely appropriate; test carefully
Long Scripts

Write to a .txt file and pass the path. Inputs over ~5 minutes may benefit from splitting into segments.

Requirements

Python 3.8+ with kokoro-onnx and soundfile (pip install kokoro-onnx soundfile). Model downloads on first use (~311 MB + ~27 MB voices, cached in ~/.cache/hyperframes/tts/).


Transcription (transcribe)

Produce a normalized transcript.json with word-level timestamps.

bash
npx hyperframes transcribe audio.mp3
npx hyperframes transcribe video.mp4 --model small --language es
npx hyperframes transcribe subtitles.srt          # import existing
npx hyperframes transcribe subtitles.vtt
npx hyperframes transcribe openai-response.json
Critical Language Rule

Never use .en models unless the user explicitly states the audio is English. .en models (small.en, medium.en) translate non-English audio into English instead of transcribing it. This silently destroys the original language.

  1. Language known and non-English → --model small --language <code> (no .en suffix)
  2. Language known and English → --model small.en
  3. Language unknown → --model small (no .en, no --language) — whisper auto-detects

Default model is small, not small.en.

Show full SKILL.md (237 more words)Show less
Model Sizes
ModelSizeSpeedWhen to use
tiny75 MBFastestQuick previews, testing pipeline
base142 MBFastShort clips, clear audio
small466 MBModerateDefault — most content
medium1.5 GBSlowImportant content, noisy audio, music
large-v33.1 GBSlowestProduction quality

Music with vocals: start at medium minimum.

Output Shape
json
[
  { "id": "w0", "text": "Hello", "start": 0.0, "end": 0.5 },
  { "id": "w1", "text": "world.", "start": 0.6, "end": 1.2 }
]

Background Removal (remove-background)

Remove the background from a video or image so the subject sits as a transparent overlay.

bash
npx hyperframes remove-background subject.mp4 -o transparent.webm  # VP9 alpha WebM
npx hyperframes remove-background subject.mp4 -o transparent.mov   # ProRes 4444
npx hyperframes remove-background portrait.jpg -o cutout.png       # single-image cutout
npx hyperframes remove-background subject.mp4 -o subject.webm \
  --background-output plate.webm                                   # both layers
npx hyperframes remove-background --info                           # detected providers

Uses u2net_human_seg (MIT). First run downloads ~168 MB of weights.

Layer Separation (--background-output)

Pass --background-output (or -b) to emit a second transparent video with the inverse alpha:

FileAlpha is...Use it for
-o subject.webmThe mask — subject opaque, bg transparentForeground layer
--background-output plate.webmInverse — bg opaque, subject transparentBottom layer; put text/graphics between

Both share the same quality preset and run from a single inference pass.

Output Format
FormatWhen
.webm (VP9 + alpha)Default. Compositions play directly via <video>.
.mov (ProRes 4444)Editing in DaVinci/Premiere/FCP. Large files.
.pngSingle-image cutout.
Quality Presets
PresetCRFWhen
fast30Iterating, smaller file
balanced18Default. Visually identical for most uses
best12Master / final delivery

TTS -> Transcribe -> Captions Pipeline

Generate voiceover, get word-level timestamps, and create captions:

bash
npx hyperframes tts script.txt --voice af_heart --output narration.wav
npx hyperframes transcribe narration.wav   # -> transcript.json

Whisper extracts precise word boundaries from the generated audio, so caption timing matches delivery without hand-tuning.


SkillPurpose
hyperframesComposition authoring (HTML, GSAP, captions, variables)
hyperframes-cliCLI dev loop (init, lint, preview, render, doctor)

© cosmicstack-labs, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in categories/development/hyperframes-media of cosmicstack-labs/mercury-agent-skills.

Open the folder on GitHubat commit 30392fb

Compare with similar skills

Hyperframes Media next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Hyperframes Media compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Hyperframes Media this skillcosmicstack-labs/mercury-agent-skills476—~1.7kAutomated safety check: PassMIT
Hyperframesscott-fryxell/brayness124—~3.2kAutomated safety check: PassMIT
Hyperframes Mediaboraoztunc/skills3932 repos~3.7kAutomated safety check: PassApache-2.0
Hyperframes CLIaiskillstore/marketplace4301 repos~2.7kAutomated safety check: PassNone
HyperFrames Media Useheygen-com/hyperframes58k2 repos~2.1kAutomated safety check: PassApache-2.0
Hyperframesboraoztunc/skills3939 repos~7.6kAutomated safety check: PassApache-2.0

Similar skills

  • Hyperframes

    scott-fryxell/brayness

    Create video compositions, animations, title cards, overlays, captions, voiceovers, audio-reactive visuals, and scene transitions in HyperFrames HTML.

    124 GitHub stars~3.2k tokensUpdated 9 days ago
    Media & CreativeAuto-check passed
  • Hyperframes Media

    boraoztunc/skills

    Asset preprocessing for HyperFrames compositions — text-to-speech narration (Kokoro), audio/video transcription (Whisper), and background removal for transparent overlays (u2net).

    393 GitHub starsUsed in 2 repos~3.7k tokens
    Media & CreativeAuto-check passed
  • Hyperframes CLI

    aiskillstore/marketplace

    Use the HyperFrames CLI development loop: init, add, catalog, capture, lint, check, snapshot, compare, grade-compare, preview, play, present, beats, keyframes, single or batch render, publish…

    430 GitHub starsUsed in 1 repo~2.7k tokens
    Media & CreativeAuto-check passed
  • HyperFrames Media Use

    heygen-com/hyperframes

    Finds, generates and edits media for HyperFrames video projects: music, sound effects, images, icons, logos, voiceovers, captions and color grades.

    58k GitHub starsUsed in 2 repos~2.1k tokens
    Media & CreativeAuto-check passed
  • Hyperframes

    boraoztunc/skills

    Create video compositions, animations, title cards, overlays, captions, voiceovers, audio-reactive visuals, and scene transitions in HyperFrames HTML.

    393 GitHub starsUsed in 9 repos~7.6k tokens
    Media & CreativeAuto-check passed
  • Hyperframes Media

    chmonitor/chmonitor

    Audio and media assets for HyperFrames compositions, produced by one shared audio engine (scripts/audio.mjs) — multi-provider TTS (HeyGen / ElevenLabs / Kokoro local), background music + sound…

    298 GitHub starsUsed in 1 repo~2.8k tokens
    Media & CreativeAuto-check: notes

More from cosmicstack-labs/mercury-agent-skills

All 12 skills in this repo
  • Before You Build

    cosmicstack-labs/mercury-agent-skills

    Use this before implementing a product, feature, SaaS, AI app, or side project to score product risk and choose the smallest validation step.

    476 GitHub stars~2.2k tokensUpdated 1 mo ago
    Auto-check passed
  • Hyperframes CLI

    cosmicstack-labs/mercury-agent-skills

    HyperFrames CLI dev loop — project scaffolding, validation (lint/inspect), browser preview with live reload, MP4/WebM rendering, and environment troubleshooting (doctor, browser, info, upgrade).

    476 GitHub stars~1.4k tokensUpdated 1 mo ago
    Auto-check passed
  • Agent Handoff Protocols

    cosmicstack-labs/mercury-agent-skills

    Design and implement agent-to-agent handoff protocols for multi-agent systems.

    476 GitHub stars~4.1k tokensUpdated 1 mo ago
    Auto-check passed
  • Agent Health Monitoring

    cosmicstack-labs/mercury-agent-skills

    Monitor AI agent health, detect anomalies, set up alerting, and maintain observability dashboards for production multi-agent systems.

    476 GitHub stars~2.7k tokensUpdated 1 mo ago
    Auto-check passed
  • Agent Task Delegation

    cosmicstack-labs/mercury-agent-skills

    Design and operate task delegation systems for multi-agent fleets.

    476 GitHub stars~3.4k tokensUpdated 1 mo ago
    Auto-check passed
  • Any2pdf

    cosmicstack-labs/mercury-agent-skills

    Convert Markdown to publication-quality PDF with reportlab — CJK/Latin mixed text, themes, cover pages, watermarks, callouts, formulas, and interactive theme selection

    476 GitHub stars~3.3k tokensUpdated 1 mo ago
    Auto-check: notes

Works with

Questions about Hyperframes Media

What does Hyperframes Media do?

Asset preprocessing for HyperFrames compositions — local text-to-speech narration (Kokoro-82M, no API key), audio/video transcription (Whisper), and background removal for transparent overlays…. Hyperframes Media is an agent skill from cosmicstack-labs/mercury-agent-skills. Asset preprocessing for HyperFrames compositions — local text-to-speech narration (Kokoro-82M, no API key), audio/video transcription (Whisper), and background removal for transparent overlays (u2net).

When should I use Hyperframes Media?

Hyperframes Media fits situations like: generating voiceover from text; transcribing speech for captions; removing background from video/images; choosing TTS voices.

How do I install Hyperframes Media in Claude Code?

Run `npx skills add cosmicstack-labs/mercury-agent-skills --skill hyperframes-media -a claude-code`. Or copy the skill folder (categories/development/hyperframes-media in cosmicstack-labs/mercury-agent-skills) into .claude/skills/hyperframes-media in your project. Claude Code loads it when a task matches its description.

How do I install Hyperframes Media in Codex?

Run `npx skills add cosmicstack-labs/mercury-agent-skills --skill hyperframes-media -a codex`. Or copy the skill folder (categories/development/hyperframes-media in cosmicstack-labs/mercury-agent-skills) into .agents/skills/hyperframes-media in your project. Codex loads it when a task matches its description.

Can I use Hyperframes Media in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add cosmicstack-labs/mercury-agent-skills --skill hyperframes-media -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/hyperframes-media, .gemini/skills/hyperframes-media, .github/skills/hyperframes-media and .opencode/skills/hyperframes-media in your project.

What does Hyperframes Media need to run?

Going by SKILL.md and its folder, Hyperframes Media needs the command-line tools its instructions call (npx and pip). Our summary lists: Python 3; Node.js.

Does Hyperframes Media access the network?

SKILL.md contains no URLs. Its commands use npx and pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Hyperframes Media safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Hyperframes Media use?

Hyperframes Media is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Hyperframes Media use?

About 1.7k tokens (SKILL.md is roughly 6.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Hyperframes Media?

Skills that share tags, products or a category with Hyperframes Media: Hyperframes (scott-fryxell/brayness, 124 stars), Hyperframes Media (boraoztunc/skills, 393 stars), Hyperframes CLI (aiskillstore/marketplace, 430 stars) and HyperFrames Media Use (heygen-com/hyperframes, 58k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Hyperframes Media?

cosmicstack-labs (a GitHub organization) maintains it in cosmicstack-labs/mercury-agent-skills, which has 476 GitHub stars. The repository holds 12 skills in this directory. The repository was last updated on August 25, 2026.

Source: cosmicstack-labs/mercury-agent-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.