Agent skill

Audio-Timed Subtitle Generator

by Pluviobyte in Pluviobyte/rnskill

Generates SRT, VTT and caption JSON for final narration audio or a merged video using Volcengine Doubao ASR word timestamps, with a quality gate before delivery.

Custom licenceAuto-check: notesMedia & Creative

Install Audio-Timed Subtitle Generator

skills CLI
$ npx skills add Pluviobyte/rnskill --skill ra-audio-to-subtitles -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Pluviobyte/rnskill ra-audio-to-subtitles --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Pluviobyte/rnskill.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/ra-audio-to-subtitles .claude/skills/ra-audio-to-subtitles && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
ra-audio-to-subtitles
GitHub stars
1.6k
Token cost
~1.1k tokens
SKILL.md length
430 words
Files
4 (incl. scripts, references)
Skills in repo
45
Repo updated
First seen
Licence
Custom licence

At a glance

Generates SRT, VTT and caption JSON for final narration audio or a merged video using Volcengine Doubao ASR word timestamps, with a quality gate before delivery.

  • Works in 6 steps: Finish and concatenate narration first.… → Run scripts/generate_subtitles.py… → Inspect phrase segmentation. Keep… → …
  • Producing subtitles whose timing must match the final narration audio
  • SKILL.md covers Required workflow, Commands, Output contract and Hard rules
  • Runs Python scripts from its folder; calls python3 and claude; needs VOLCENGINE_API_KEY

What it does

This skill treats the final audio stream as the only source of subtitle timing, never character counts, scene duration, TTS segment length or fixed delays. After narration is finished and concatenated, or a recorded or merged video is final, scripts/generate_subtitles.py is run against that media and the exact narration script, calling Volcengine Doubao ASR for word-level timestamps. The agent then inspects phrase segmentation: English product and model names stay whole, isolated one-word fragments are merged, and Chinese connector words stay with the phrase they introduce.

The output folder holds the raw ASR response, word timestamps with detected gaps, phrase captions for HyperFrames or Remotion, SRT and VTT files for Bilibili, YouTube or editing software, and a caption-qc.json report. Delivery requires that report to show a pass covering alignment, overlap, fragments, connector splits, caption duration and reading speed. If ASR fails or alignment coverage falls below the gate, work stops, and a character-count estimate is allowed only for a clearly labeled scratch preview. A doctor flag runs a health check, and an offline mode reuses a saved ASR response for regression tests.

When your agent uses it

  • Producing subtitles whose timing must match the final narration audio
  • Exporting SRT and VTT files for a finished talking-head video
  • Checking caption quality before the final render or delivery

Example prompts

  • “Generate subtitles from final-voiceover.mp3 using the narration script in narration_segments.json.”
  • “Create SRT and VTT captions for merged.mp4 from the script in handoff.md.”
  • “Run the subtitle health check and tell me whether the ASR setup works.”

Requirements

  • Python 3
  • Volcengine ASR access and credentials

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Finish and concatenate narration first. If the user recorded or merged a
  2. Run scripts/generate_subtitles.py against the final media and the exact
  3. Inspect phrase segmentation. Keep English product/model tokens intact,
  4. Render captions from captions.json; export captions.srt for Bilibili,
  5. Require caption-qc.json to report status: pass before final render or
  6. If ASR fails or alignment coverage is below the gate, stop. A character-

What it can do on your machine

Read from SKILL.md and the folder at commit 83d1783. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python3
    • claude

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • VOLCENGINE_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Audio-Timed Subtitle Generator loads about 1.1k tokens when it runs, and up to ~1.5k if it reads all its reference files. Until then it costs about 95 tokens; SKILL.md has 430 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~95
When it runs · the whole SKILL.md, loaded when a task matches
~1.1k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~1.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteMentions a .env fileSKILL.md:93
    _API_KEY` comes from the workspace root `.env` or the environment.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

Its licence (Custom licence) doesn't allow us to republish the file, so here is its outline and opening line. It has 430 words (~1,082 tokens).

“Use the final audio stream as the only timing source. Never estimate final subtitle timing from character count, scene duration, TTS segment duration, or fixed delays.”

— opening of SKILL.md by Pluviobyte, Custom licence
name
ra-audio-to-subtitles

Read the full SKILL.md on GitHub

Files

SKILL.md and 3 other files (scripts, references) in skills/ra-audio-to-subtitles of Pluviobyte/rnskill.

  • SKILL.md
  • agents/openai.yaml
  • references/artifact-contract.md
  • scripts/generate_subtitles.py

Open the folder on GitHubat commit 83d1783

Compare with similar skills

Audio-Timed Subtitle Generator next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Audio-Timed Subtitle Generator compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Audio-Timed Subtitle Generator this skillPluviobyte/rnskill1.6k—~1.1kAutomated safety check: NotesCustom licence
Video Podcast MakerAgents365-ai/video-podcast-maker1.7k—~4.9kAutomated safety check: PassMIT
Yt Dlp DownloaderMapleShaw/yt-dlp-downloader-skill2091 repos~1.5kAutomated safety check: PassNone
Ffmpeg Skillkajisho5/ffmpeg-skill1.9k—~7.4kAutomated safety check: PassMIT
Video To Subtitle Summaryimlewc/video-to-subtitle-summary-skill218—~4.6kAutomated safety check: NotesMIT
Video To NotesKIRVO-REPORTING/video-to-notes105—~1.5kAutomated safety check: PassMIT

Similar skills

  • Video Podcast Maker

    Agents365-ai/video-podcast-maker

    A skill your agent uses when the user gives a topic and wants an automated topic-driven narrated explainer, podcast, or knowledge-summary video (Bilibili / YouTube / Xiaohongshu / Douyin / WeChat…

    1.7k GitHub stars~4.9k tokensUpdated 9 days ago
    Media & CreativeAuto-check passed
  • Yt Dlp Downloader

    MapleShaw/yt-dlp-downloader-skill

    Download videos from YouTube, Bilibili, Twitter, and thousands of other sites using yt-dlp.

    209 GitHub starsUsed in 1 repo~1.5k tokens
    Media & CreativeAuto-check passed
  • Ffmpeg Skill

    kajisho5/ffmpeg-skill

    Edit video and audio with local FFmpeg from natural-language requests: cut, trim, join, resize/reframe (9:16, 1:1), speed change, captions and subtitles (SRT/ASS, animated, karaoke), logos and text…

    1.9k GitHub stars~7.4k tokensUpdated yesterday
    Media & CreativeAuto-check passed
  • Video To Subtitle Summary

    imlewc/video-to-subtitle-summary-skill

    A skill your agent uses when user provides a short video platform URL or local video/audio file and wants subtitles/AI summary, or when user asks to list their own AI Douyin historical tasks.

    218 GitHub stars~4.6k tokensUpdated 2 mo ago
    Media & CreativeAuto-check: notes
  • Video To Notes

    KIRVO-REPORTING/video-to-notes

    Use immediately for any bare YouTube or YouTube Shorts URL, youtu.be link, Bilibili or b23.tv link, or other video URL; do not ask what the user wants.

    105 GitHub stars~1.5k tokensUpdated 1 mo ago
    Media & CreativeAuto-check passed
  • Video Clip Extractor

    linzzzzzz/openclip

    Processes videos to identify engaging moments, generate transcripts, and create highlight clips with artistic titles and custom cover images.

    569 GitHub stars~2.8k tokensUpdated 1 mo ago
    Media & CreativeAuto-check: warnings

More from Pluviobyte/rnskill

All 45 skills in this repo
  • Motion Replica Builder

    Pluviobyte/rnskill

    Analyzes a reference video's layout, motion and timing over a chosen time range, then builds an original animated replica with new copy and assets plus a final video QC.

    1.6k GitHub stars~1.7k tokensUpdated 20 days ago
    Auto-check passed
  • Niulai Style Image Transfer

    Pluviobyte/rnskill

    Restyles photos, film stills or new scene descriptions in the rough low-budget 3D look observed in the 2026 animated film Niulai, with an iterative scoring loop.

    1.6k GitHub stars~1.3k tokensUpdated 20 days ago
    Auto-check passed
  • Skill Captions

    Pluviobyte/rnskill

    Renders, previews, validates and burns in production subtitles from QC-passed caption timing, with fixed dark or light caption styles scaled to the video size.

    1.6k GitHub stars~1.3k tokensUpdated 20 days ago
    Auto-check passed
  • Creates Chinese social-media video covers in 3:4 and 4:3 from registered styles and shared presenter assets, or swaps the person in a reference cover for a digital-human frame.

    1.6k GitHub stars~1.2k tokensUpdated 20 days ago
    Auto-check passed
  • Generates a HeyGen digital-human video layer using a fixed avatar profile, local narration and an approved circular avatar placement.

    1.6k GitHub stars~2.4k tokensUpdated 20 days ago
    Auto-check passed
  • Produces a local rough cut of a talking-head or narrated screen recording, with transcript correction, an approval gate, pause compression and loudness normalization.

    1.6k GitHub stars~1.8k tokensUpdated 20 days ago
    Auto-check passed

Questions about Audio-Timed Subtitle Generator

What does Audio-Timed Subtitle Generator do?

Generates SRT, VTT and caption JSON for final narration audio or a merged video using Volcengine Doubao ASR word timestamps, with a quality gate before delivery. This skill treats the final audio stream as the only source of subtitle timing, never character counts, scene duration, TTS segment length or fixed delays.py is run against that media and the exact narration script, calling Volcengine Doubao ASR for word-level timestamps.

When should I use Audio-Timed Subtitle Generator?

Audio-Timed Subtitle Generator fits situations like: producing subtitles whose timing must match the final narration audio; exporting SRT and VTT files for a finished talking-head video; checking caption quality before the final render or delivery.

How do I install Audio-Timed Subtitle Generator in Claude Code?

Run `npx skills add Pluviobyte/rnskill --skill ra-audio-to-subtitles -a claude-code`. Or copy the skill folder (skills/ra-audio-to-subtitles in Pluviobyte/rnskill) into .claude/skills/ra-audio-to-subtitles in your project. Claude Code loads it when a task matches its description.

How do I install Audio-Timed Subtitle Generator in Codex?

Run `npx skills add Pluviobyte/rnskill --skill ra-audio-to-subtitles -a codex`. Or copy the skill folder (skills/ra-audio-to-subtitles in Pluviobyte/rnskill) into .agents/skills/ra-audio-to-subtitles in your project. Codex loads it when a task matches its description.

Can I use Audio-Timed Subtitle Generator in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Pluviobyte/rnskill --skill ra-audio-to-subtitles -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ra-audio-to-subtitles, .gemini/skills/ra-audio-to-subtitles, .github/skills/ra-audio-to-subtitles and .opencode/skills/ra-audio-to-subtitles in your project.

What does Audio-Timed Subtitle Generator need to run?

Going by SKILL.md and its folder, Audio-Timed Subtitle Generator needs Python for the scripts in its folder, the command-line tools its instructions call (python3 and claude) and credentials named VOLCENGINE_API_KEY. Our summary lists: Python 3; Volcengine ASR access and credentials.

Does Audio-Timed Subtitle Generator access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Audio-Timed Subtitle Generator safe to install?

Our automated static check of SKILL.md found notes only (mentions a .env file), nothing it rates as a warning. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Audio-Timed Subtitle Generator use?

Audio-Timed Subtitle Generator has a licence file (the repository's licence) that doesn't match a standard licence. Read it on GitHub before reusing the skill.

How many tokens does Audio-Timed Subtitle Generator use?

About 1.1k tokens (SKILL.md is roughly 4.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 465 tokens, read only when the agent opens those files.

What are the alternatives to Audio-Timed Subtitle Generator?

Skills that share tags, products or a category with Audio-Timed Subtitle Generator: Video Podcast Maker (Agents365-ai/video-podcast-maker, 1.7k stars), Yt Dlp Downloader (MapleShaw/yt-dlp-downloader-skill, 209 stars), Ffmpeg Skill (kajisho5/ffmpeg-skill, 1.9k stars) and Video To Subtitle Summary (imlewc/video-to-subtitle-summary-skill, 218 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Audio-Timed Subtitle Generator?

Pluviobyte (a GitHub user) maintains it in Pluviobyte/rnskill, which has 1,639 GitHub stars. The repository holds 45 skills in this directory. The repository was last updated on September 21, 2026.

Source: Pluviobyte/rnskill on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.