Agent skill

Audio and SRT to Narrated Video

by sugarforever in sugarforever/01coder-agent-skills

Builds a narration-synced MP4 from a voiceover audio file and its SRT subtitles using HyperFrames, with scenes timed to the subtitle cues; instructions are in Chinese.

MITAuto-check passedMedia & Creative

SKILL.md written in Chinese; this summary is our English description.

Install Audio and SRT to Narrated Video

skills CLI
$ npx skills add sugarforever/01coder-agent-skills --skill producing-video -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install sugarforever/01coder-agent-skills producing-video --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/sugarforever/01coder-agent-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/producing-video .claude/skills/producing-video && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
producing-video
GitHub stars
137
Token cost
~1.9k tokens
SKILL.md length
466 words
Files
2 (incl. scripts)
Skills in repo
21
Repo updated
First seen
Licence
MIT

At a glance

Builds a narration-synced MP4 from a voiceover audio file and its SRT subtitles using HyperFrames, with scenes timed to the subtitle cues; instructions are in Chinese.

  • Works in 6 steps: 选 frame / 品牌 → 起项目 → 解析 SRT → 切场景 → …
  • Turning a recorded narration plus an SRT file into a finished video
  • SKILL.md covers 分工(很重要), 依赖检查(pre-flight), 工作流 and Gotchas(血泪,务必遵守), plus 3 more sections
  • Runs JavaScript scripts from its folder; calls npx, curl and ffmpeg; reaches hyperframes.dev and cdn.jsdelivr.net

What it does

You supply a recorded voiceover and an SRT subtitle file, and the agent turns them into a video whose visuals follow the voice. Audio and SRT are the only source of truth: on-screen content comes from the subtitle text, scene timing comes from the cue timestamps, and the audio goes into the HyperFrames composition as an audio clip so it is mixed in during render, with no separate step of adding sound afterward. The instructions are written in Chinese.

The workflow starts with a pre-flight check using npx hyperframes doctor, which needs Node 22 or newer, FFmpeg and Chrome. The agent then picks a visual style, either a series brand, a HyperFrames frame template or one of eight built-in styles, creates a blank project, and reads the SRT, using the bundled srt-cues.mjs script to list each cue's start time. Cues are grouped into scenes of about 30 seconds, so a 6 to 7 minute piece comes to roughly 12 to 16 scenes.

When your agent uses it

  • Turning a recorded narration plus an SRT file into a finished video
  • Making a daily or series explainer whose visuals follow the voiceover
  • Reusing a series brand's colors and fonts for a new narrated video

Example prompts

  • “Here are audio.mp3 and audio.srt, so make a video from them with HyperFrames.”
  • “Render this narration and subtitle file into an MP4 in the same style as our earlier series covers.”
  • “把这期早读做成视频,音频和字幕都在 ./episode 文件夹里。”

Requirements

  • Node.js 22 or newer
  • FFmpeg
  • Chrome
  • The hyperframes and hyperframes-cli skills
  • A voiceover audio file and an SRT subtitle file

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. 选 frame / 品牌
  2. 起项目
  3. 解析 SRT → 切场景
  4. 编写合成(index.html)
  5. 校对(每次改完都跑)
  6. 渲染 + 验收

What it can do on your machine

Read from SKILL.md and the folder at commit e51fb6e. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (JavaScript), which the agent can run.

    Shell commands in SKILL.md call:

    • npx
    • curl
    • ffmpeg

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • hyperframes.dev
    • cdn.jsdelivr.net

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Audio and SRT to Narrated Video loads about 1.9k tokens when it runs. Until then it costs about 159 tokens; SKILL.md has 466 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~159
When it runs · the whole SKILL.md, loaded when a task matches
~1.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from sugarforever/01coder-agent-skills at commit e51fb6e, republished under its MIT licence (© sugarforever). 466 words, ~1,855 tokens.

Download SKILL.mdSave it as .claude/skills/producing-video/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
producing-video
description
Turn a user-provided voiceover audio file + SRT subtitle into a finished, narration-synced MP4 using HyperFrames (HTML-to-video). The audio + SRT are the source of truth — scenes are timed to the SRT cues, content is read from the SRT, and the audio is muxed in automatically. Use when the user hands over an mp3/wav + srt and wants a video, says "把音频做成视频", "做一期视频", "audio + srt to video", "把这期早读做成视频", "render this narration into a video", or provides a recording + subtitles for an explainer / daily / 解读 / 口播. NOT for generating the voiceover (that's the user's job here) and NOT for slide decks (use a slides skill).

Producing-Video · 音频 + 字幕 → 成片

把"用户已经录好的口播音频 + SRT 字幕"做成一条画面跟着声音走的 MP4。用 HyperFrames(HTML 即视频)出片。

铁律:音频和 SRT 是唯一事实源。 画面的内容来自 SRT,画面的时间轴来自 SRT 的 cue 时间戳,音频作为一个 <audio> clip 直接挂进合成里、渲染时自动合流 —— 没有"先出视频再合音频"这一步。

分工(很重要)

谁做什么
用户写稿 → 录音/合成音频 → 生成 SRT → 把 audio.mp3 + audio.srt 交给你
本 skill(你)选 frame/品牌 → 按 SRT 搭 HyperFrames 合成 → 校对 → 渲染成片

用户不指望你生成配音(那是上游)。如果用户还没有音频、问的是"怎么配音",那不是本 skill —— TTS / 声音克隆是另一条线(见下方"超出范围")。

依赖检查(pre-flight)

bash
npx hyperframes doctor      # 需要 Node ≥ 22 · FFmpeg · Chrome

需要 hyperframes / hyperframes-cli 两个 skill 在场(编写合成 + 跑 CLI)。缺了就让用户 npx hyperframes skills 安装后重来。

确认用户给了两个文件:音频(mp3/wav/m4a)+ SRT。只给音频没给 SRT → 本 skill 需要 SRT 拿时间轴;可让用户补 SRT(很多录音工具/剪辑软件能导出),不要默认去跑 Whisper 转写(用户没给 SRT 往往是有意的,先问)。


工作流

Step 1 · 选 frame / 品牌

画面的风格 = 一个 frame.md / visual-style / 既有系列品牌。三种来源,按情况选:

  1. 延续系列品牌(推荐用于日更/系列)。如果这期属于一个已有系列(如"AI 早读"),去翻该系列的封面/历史成片,沿用同一套 token(颜色、字体、栅格、页眉页脚),让视频和封面一脉相承。
  2. HyperFrames frame.md 模板。用户可能直接点名,例如 creative-mode / biennale-yellow / cobalt-grid。取 token:
    • curl -sSL https://www.hyperframes.dev/design/<slug>.md (站点是 JS 渲染,多半取不到正文);
    • 更靠谱:在本机 open-design 仓库里找 design-templates/*<slug>*/template.json,里面有精确的 palette / typography(hex + 字体名)。
  3. 8 个内置 visual-style(Swiss Pulse / Velvet Standard / Shadow Cut / Maximalist 等)。在 hyperframes skill 的 visual-styles.md 里,直接抄 YAML token。

字体一律本地 woff2(见 Gotcha "字体")。中文必须配 Noto Sans SC(400/500/700/900 视用量);英文 display 按 frame 选(Archivo Black / Manrope / Oswald…);标签数字常用 JetBrains Mono。从 fontsource CDN 下到项目的 fonts/:

bash
curl -sSL -o fonts/<name>.woff2 "https://cdn.jsdelivr.net/fontsource/fonts/<family>@latest/<subset>-<weight>-normal.woff2"
# 中文:subset 用 chinese-simplified;拉丁:latin
Step 2 · 起项目
bash
cd <repo>/studio/videos                       # 仓库约定:成片放这里
npx hyperframes init <YYYYMMDD-slug> --example blank --non-interactive
cd <YYYYMMDD-slug> && mkdir -p fonts audio
cp <user-audio> audio/narration-full.mp3
cp <user-srt>   audio/narration.srt
Step 3 · 解析 SRT → 切场景

读完整 SRT(scripts/srt-cues.mjs 可打印每条 cue 的开始秒数 + 文本,方便规划)。然后:

  1. 按内容把 cue 归成"幕/场景"。一条 SRT cue 通常是一句话;把讲同一件事的几条 cue 合成一个场景(scene)。一支 67 分钟的日更,大约 1216 个场景比较舒服(平均 ~30s/场,画面不至于久不动)。
  2. 场景开始时间 = 它第一条 cue 的开始时间戳(秒)。
  3. 场景时长 = 下一场开始 − 本场开始 + ~0.5s(这 0.5s 重叠让转场 wipe 能盖住上一场,避免穿帮)。最后一场到音频结束。
  4. 场景内的逐步出现(sub-reveal)= 对应子 cue 的开始时间。让标题/要点/数字"在被念到的那一刻"出现 —— 这是同步感的关键。
  5. 画面文案从 SRT 来,且不能和口播打架。可以精炼成海报式短句,但不能说的是 A、画面写 B。顺手核对事实(数字、专有名词、人名)—— 用户很在意准确。

把场景开始时间放进一个 JS 数组 const B = [...],所有 tween 用 B[i-1] + 局部偏移 定位。日后微调时间轴只改数组,不用逐条改 tween。

Step 4 · 编写合成(index.html)
  • 持久 chrome:页眉/页脚/栅格/边框这类每一场都在的元素,放在场景之外(直接挂 #root,不是 clip),整片不动 —— 营造"节目"感。
  • 音频一条连续 clip("音频即时钟"):
    html
    <audio id="vo" src="audio/narration-full.mp3" data-start="0" data-duration="<总时长>" data-track-index="20" data-volume="1"></audio>
    媒体元素不需要 class="clip";给它独立的 data-track-index。
  • 每个场景一个全画布 .scene.clip,各自独立 data-track-index(重叠的 wipe 需要不同 track),z-index 递增(后面的盖前面的)。
  • 转场只用"incoming 场景自身的 clip-path 揭幕"(见 Gotcha「转场」):
    js
    function wipe(sel, at){ tl.fromTo(sel,{clipPath:"inset(0 100% 0 0)"},{clipPath:"inset(0 0% 0 0)",duration:0.5,ease:"power3.inOut"}, at); }
  • 入场动画用 gsap.from(),定位在对应 cue 时间。短(0.3–0.7s),错峰,变化 ease。
  • 复用一套 class 化的 CSS 工具件(kicker / headline / note / stat / chip-xform / cards / numbered-list / accent-mark),15 个场景共享,别每场写一套 id 样式。
  • 品牌纪律:严格按 frame 的 token;强调色当"标点"不当"填充"(深色系:强调色只给小而实的标记;浅色系:强调色做色块、文字用墨色压在色块上)。
Step 5 · 校对(每次改完都跑)
bash
npx hyperframes lint        # 0 error 才继续(var(--x) 字体告警是误报,可忽略)
npx hyperframes validate    # WCAG AA 对比度;改掉过暗的次级灰
npx hyperframes inspect --samples 30   # 版面溢出,带时间戳;场景多就多采样

装饰元素故意出血到画外 → 标 data-layout-ignore。真实溢出 → 改容器/字号/padding。

Show full SKILL.md (213 more words)Show less
Step 6 · 渲染 + 验收
  • 先 standard 跑一版自检(快),用 ffmpeg 抽几帧验证同步:在"你知道这一刻在讲什么"的时间点抽帧,确认画面对得上。
    bash
    ffmpeg -y -ss <秒> -i renders/x.mp4 -frames:v 1 /tmp/f.png   # 然后看图
  • 成片渲染(master):
    bash
    npx hyperframes render --resolution landscape-4k --quality high --output renders/<slug>-4k.mp4
    --resolution landscape-4k 是把同一合成按 2× DPR 真·超采样到 3840×2160(不是放大);4K master 即使观众看 1080p 也更耐平台二压。4K + 长片渲染较久(几分钟到十几分钟),可后台跑。
  • 验收:ffprobe 确认有 video(h264) + audio(aac) 两条轨且时长对得上;抽帧确认每场落在它的 cue 上。

成片留在 studio/videos/<slug>/renders/。不要自动提交(除非用户明确要)。


Gotchas(血泪,务必遵守)

这些是踩过的坑,违反任何一条都会出废片:

  1. 音频即时钟。场景时长从 SRT 量出来,不要凭感觉定时长再硬塞音频。
  2. 有 SRT 就别转写。用户给了 SRT = 精确时间轴免费拿到,不需要 Whisper。(没 SRT 想自己转写前先问用户。)
  3. 转场只用场景自身 clip-path 揭幕。绝不用单独的全屏色块/幕布/刀闸/砸场板去做转场 —— 在渲染引擎里它会"扫进来盖住下一场后卡住不走",整场变成纯色/黑屏。揭幕揭的是 incoming 场景本体。
  4. 不要给 .pad 容器套整体 opacity 的 "pushIn" 包装。容器级 opacity 动画在 seek 渲染里可能留在 0,把整场变黑。用每个元素各自的 gsap.from()。
  5. 配对 tween 不要加 overwrite:"auto"。它会把配对的另一条 tween 杀掉(比如"扫入"在、"扫出"没了)。lint 的 overlapping_gsap_tweens 是无害告警,宁可留着。
  6. 字体必须本地 woff2。Google Fonts <link> 会被 lint 标记、且 sandbox 渲染里不可靠。中文配 Noto Sans SC;中文字在彩色 accent 色块里要给足竖直 padding/line-height(CJK 字形比 em 框高,padding 太紧 inspect 会报 text_box_overflow,给到 ~0.2em 竖直 padding + line-height ~1.12)。
  7. 深色主题别信亮度探测。1×1 平均亮度对深底+稀疏文字永远偏低,会把正常场景误判成黑屏。靠抽帧看图确认。
  8. 浅色主题的对比度误报。validate 在固定几个时间戳采样所有 DOM 文字,包括当时未激活的场景。浅底上"浅色文字"(白字、奶油字)一旦不在自己场景的激活时刻被采到,就报低对比度——这是误报。规避:accent 色块上一律用墨色(深)文字,未激活时墨字压奶油底仍是高对比,零误报。
  9. 确定性。禁止 Date.now() / Math.random()(破坏可复现渲染);要随机用种子化 PRNG。
  10. 每个场景独立 track-index;音频单独高 track-index。装饰出血标 data-layout-ignore。
  11. 画质:原生 1080p 在 Retina 上看会发虚(被放大 + H.264 4:2:0 软化彩色字缘);master 用 --resolution landscape-4k --quality high。

SRT → 场景时间轴(配方)

场景[i].start    = cue[第一条].start                 (秒)
场景[i].duration = 场景[i+1].start − 场景[i].start + 0.5   (末场到音频末尾)
场景内某元素入场 = 它对应 cue 的 start
音频 clip        = data-start=0, data-duration=总时长
root data-duration = 音频末句之后留 ~3s 收尾

JS 里:

js
const B = [0, 38.63, 53.96, /* ...每场 start... */];   // 从 SRT 量
const at = (i, off) => B[i-1] + off;                   // i 是 1-based 场号
function wipe(sel, i){ tl.fromTo(sel, {clipPath:"inset(0 100% 0 0)"}, {clipPath:"inset(0 0% 0 0)", duration:0.5, ease:"power3.inOut"}, B[i-1]); }

超出范围

  • 生成配音 / 声音克隆:本 skill 只吃用户给的音频。要让它"听起来像我"用云端 TTS(MiniMax / ElevenLabs,支持中文 + 克隆),克隆只换"出声那一步",下游时间轴/挂载/渲染不变。HyperFrames 自带的 Kokoro TTS 做不了中文(CLI 传 zh、espeak 要 cmn,且质量差);本机临时方案 macOS say -v Tingting。
  • 幻灯片 deck:要的是横向翻页 deck 而非视频 → 用 slides 类 skill。

输出清单

  • lint 0 error · validate 全过 · inspect 0 issue
  • 抽帧确认每场落在它的 cue 上(尤其数据页/转折页)
  • ffprobe:video + audio 两轨、时长 = 音频时长
  • master 用 4K high(除非用户另说)
  • 成片在 studio/videos/<slug>/renders/,未自动提交

© sugarforever, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (scripts) in skills/producing-video of sugarforever/01coder-agent-skills.

  • SKILL.md
  • scripts/srt-cues.mjs

Open the folder on GitHubat commit e51fb6e

Compare with similar skills

Audio and SRT to Narrated Video next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Audio and SRT to Narrated Video compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Audio and SRT to Narrated Video this skillsugarforever/01coder-agent-skills137—~1.9kAutomated safety check: PassMIT
Manim Video Productionbrowser-use/video-use28k6 repos~3kAutomated safety check: PassMIT
Embedded Video Captionsheygen-com/hyperframes59k3 repos~8.6kAutomated safety check: PassApache-2.0
HyperFrames Video Entry Pointheygen-com/hyperframes59k3 repos~5.2kAutomated safety check: PassApache-2.0
Painted Animationtuzhechen2005/opus-video-skills1371 repos~1.9kAutomated safety check: PassCustom licence
HyperFrames Video Compositionsspinabot/brigade11k—~1.6kAutomated safety check: PassMIT

Similar skills

  • Manim Video Production

    browser-use/video-use

    Produces math and technical explainer videos with Manim Community Edition: concept animations, equation derivations, algorithm walkthroughs and data stories.

    28k GitHub starsUsed in 6 repos~3k tokens
    Media & CreativeAuto-check passed
  • Embedded Video Captions

    heygen-com/hyperframes

    Adds captions to a single-subject talking-head video without editing the footage, from plain subtitles to cinematic text placed behind the speaker.

    59k GitHub starsUsed in 3 repos~8.6k tokens
    Media & CreativeAuto-check passed
  • HyperFrames Video Entry Point

    heygen-com/hyperframes

    Entry point for making, editing and rendering videos from HTML compositions with HyperFrames, routing each request to the right workflow.

    59k GitHub starsUsed in 3 repos~5.2k tokens
    Media & CreativeAuto-check passed
  • Painted Animation

    tuzhechen2005/opus-video-skills

    Make hand-painted watercolour-and-ink cartoon videos (MP4) with code — p5.js + p5.brush rendered frame by frame in headless Chrome, encoded with ffmpeg — starring Clawd or any character.

    137 GitHub starsUsed in 1 repo~1.9k tokens
    Media & CreativeAuto-check passed
  • Writes HTML compositions with a GSAP timeline that the render_video tool turns into deterministic MP4 videos from data, text or layouts.

    11k GitHub stars~1.6k tokensUpdated yesterday
    Media & CreativeAuto-check passed
  • Black White Text Opener

    Pluviobyte/video-production-skills

    Create reusable black-background white-text opening animations for new videos.

    667 GitHub stars~1k tokensUpdated 2 mo ago
    Media & CreativeAuto-check passed

More from sugarforever/01coder-agent-skills

All 21 skills in this repo
  • Typography Cover Designer

    sugarforever/01coder-agent-skills

    Designs typography-driven video covers and thumbnails in HTML/CSS and screenshots them with Chrome DevTools at 16:9, 16:10, 9:16 and 3:4.

    137 GitHub stars~3.3k tokensUpdated 3 mo ago
    Auto-check passed
  • Mining Session Skills

    sugarforever/01coder-agent-skills

    Reviews one finished Claude Code session from its exported markdown and decides whether a skill is worth creating, updating or reusing for similar work.

    137 GitHub stars~1.4k tokensUpdated 3 mo ago
    Auto-check passed
  • Subtitle Correction

    sugarforever/01coder-agent-skills

    Fixes speech-recognition mistakes in .srt subtitle files, in Chinese or English, using terms you supply and leaving every timestamp untouched.

    137 GitHub stars~2.3k tokensUpdated 3 mo ago
    Auto-check passed
  • Claude Session Exporter

    sugarforever/01coder-agent-skills

    Lists Claude Code session transcripts and exports them to organized Markdown with a digest, a clean conversation and linked tool-call details.

    137 GitHub stars~1.7k tokensUpdated 3 mo ago
    Auto-check passed
  • Codex Session Manager

    sugarforever/01coder-agent-skills

    Lists, filters and exports local Codex session transcripts to organized Markdown with a digest, timeline and tool-call details, in full or pick-one mode.

    137 GitHub stars~1.2k tokensUpdated 3 mo ago
    Auto-check passed
  • Diagram to Image Converter

    sugarforever/01coder-agent-skills

    Converts Mermaid diagrams and markdown tables into PNG images through a hosted rendering API, for platforms without rich formatting.

    137 GitHub stars~1.9k tokensUpdated 3 mo ago
    Auto-check passed

Works with

Questions about Audio and SRT to Narrated Video

What does Audio and SRT to Narrated Video do?

Builds a narration-synced MP4 from a voiceover audio file and its SRT subtitles using HyperFrames, with scenes timed to the subtitle cues; instructions are in Chinese. You supply a recorded voiceover and an SRT subtitle file, and the agent turns them into a video whose visuals follow the voice. Audio and SRT are the only source of truth: on-screen content comes from the subtitle text, scene timing comes from the cue timestamps, and the audio goes into the HyperFrames composition as an audio clip so it is mixed in during render, with no separate step of adding sound afterward.

When should I use Audio and SRT to Narrated Video?

Audio and SRT to Narrated Video fits situations like: turning a recorded narration plus an SRT file into a finished video; making a daily or series explainer whose visuals follow the voiceover; reusing a series brand's colors and fonts for a new narrated video.

How do I install Audio and SRT to Narrated Video in Claude Code?

Run `npx skills add sugarforever/01coder-agent-skills --skill producing-video -a claude-code`. Or copy the skill folder (skills/producing-video in sugarforever/01coder-agent-skills) into .claude/skills/producing-video in your project. Claude Code loads it when a task matches its description.

How do I install Audio and SRT to Narrated Video in Codex?

Run `npx skills add sugarforever/01coder-agent-skills --skill producing-video -a codex`. Or copy the skill folder (skills/producing-video in sugarforever/01coder-agent-skills) into .agents/skills/producing-video in your project. Codex loads it when a task matches its description.

Can I use Audio and SRT to Narrated Video in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add sugarforever/01coder-agent-skills --skill producing-video -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/producing-video, .gemini/skills/producing-video, .github/skills/producing-video and .opencode/skills/producing-video in your project.

What does Audio and SRT to Narrated Video need to run?

Going by SKILL.md and its folder, Audio and SRT to Narrated Video needs JavaScript for the scripts in its folder and the command-line tools its instructions call (npx, curl and ffmpeg). Our summary lists: Node.js 22 or newer; FFmpeg; Chrome; The hyperframes and hyperframes-cli skills; A voiceover audio file and an SRT subtitle file.

Does Audio and SRT to Narrated Video access the network?

SKILL.md names 2 domains. In commands or code: hyperframes.dev and cdn.jsdelivr.net; the agent is likely to contact these when it follows the instructions. This is read from the text; nothing was executed.

Is Audio and SRT to Narrated Video safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Audio and SRT to Narrated Video use?

Audio and SRT to Narrated Video is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Audio and SRT to Narrated Video use?

About 1.9k tokens (SKILL.md is roughly 7.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Audio and SRT to Narrated Video?

Skills that share tags, products or a category with Audio and SRT to Narrated Video: Manim Video Production (browser-use/video-use, 28k stars), Embedded Video Captions (heygen-com/hyperframes, 59k stars), HyperFrames Video Entry Point (heygen-com/hyperframes, 59k stars) and Painted Animation (tuzhechen2005/opus-video-skills, 137 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Audio and SRT to Narrated Video?

sugarforever (a GitHub user) maintains it in sugarforever/01coder-agent-skills, which has 137 GitHub stars. The repository holds 21 skills in this directory. The repository was last updated on June 19, 2026.

Source: sugarforever/01coder-agent-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.