HyperFrames Media Use
heygen-com/hyperframes
Finds, generates and edits media for HyperFrames video projects: music, sound effects, images, icons, logos, voiceovers, captions and color grades.
A skill your agent uses when the user has a video + a target-language SRT and wants the video to actually speak that language — generates a time-aligned TTS voice dub.
$ npx skills add jianshuo/claude-skills --skill wjs-dubbing-video -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install jianshuo/claude-skills wjs-dubbing-video --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/jianshuo/claude-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/wjs-dubbing-video .claude/skills/wjs-dubbing-video && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "wjs-dubbing-video" agent skill from https://github.com/jianshuo/claude-skills/tree/main/wjs-dubbing-video into .claude/skills/wjs-dubbing-video/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "wjs-dubbing-video", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/jianshuo/claude-skills/tree/main/wjs-dubbing-videoType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add jianshuo/claude-skills --skill wjs-dubbing-video -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install jianshuo/claude-skills wjs-dubbing-video --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jianshuo/claude-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/wjs-dubbing-video .agents/skills/wjs-dubbing-video && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "wjs-dubbing-video" agent skill from https://github.com/jianshuo/claude-skills/tree/main/wjs-dubbing-video into .agents/skills/wjs-dubbing-video/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "wjs-dubbing-video", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add jianshuo/claude-skills --skill wjs-dubbing-video -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install jianshuo/claude-skills wjs-dubbing-video --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jianshuo/claude-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/wjs-dubbing-video .cursor/skills/wjs-dubbing-video && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "wjs-dubbing-video" agent skill from https://github.com/jianshuo/claude-skills/tree/main/wjs-dubbing-video into .cursor/skills/wjs-dubbing-video/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "wjs-dubbing-video", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/jianshuo/claude-skills.git --path wjs-dubbing-video--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add jianshuo/claude-skills --skill wjs-dubbing-video -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install jianshuo/claude-skills wjs-dubbing-video --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jianshuo/claude-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/wjs-dubbing-video .gemini/skills/wjs-dubbing-video && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "wjs-dubbing-video" agent skill from https://github.com/jianshuo/claude-skills/tree/main/wjs-dubbing-video into .gemini/skills/wjs-dubbing-video/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "wjs-dubbing-video", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install jianshuo/claude-skills wjs-dubbing-videoInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add jianshuo/claude-skills --skill wjs-dubbing-video -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/jianshuo/claude-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/wjs-dubbing-video .github/skills/wjs-dubbing-video && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "wjs-dubbing-video" agent skill from https://github.com/jianshuo/claude-skills/tree/main/wjs-dubbing-video into .github/skills/wjs-dubbing-video/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "wjs-dubbing-video", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add jianshuo/claude-skills --skill wjs-dubbing-video -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install jianshuo/claude-skills wjs-dubbing-video --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jianshuo/claude-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/wjs-dubbing-video .opencode/skills/wjs-dubbing-video && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "wjs-dubbing-video" agent skill from https://github.com/jianshuo/claude-skills/tree/main/wjs-dubbing-video into .opencode/skills/wjs-dubbing-video/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "wjs-dubbing-video", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
wjs-dubbing-videoA skill your agent uses when the user has a video + a target-language SRT and wants the video to actually speak that language — generates a time-aligned TTS voice dub.
Wjs Dubbing Video is an agent skill from jianshuo/claude-skills. Use when the user has a video + a target-language SRT and wants the video to actually speak that language — generates a time-aligned TTS voice dub. Routes by voice ID — Volcano (豆包) TTS for Chinese, edge-tts neural for any language. Defaults to one voice (single-speaker); opt-in multi-speaker via visual diarization. Outputs <langdub.mp4 with the dub audio in place of the original. Final mixing (audio bed + burn-in) is handed off to /wjs-burning-subtitles. Triggers — "配音", "中文配音", "Chinese dub", "voice over this"…
Its SKILL.md is about 5.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files, including scripts (for example `scripts/dub.py` and `scripts/visual_diarize.py`).
It sits in Media & Creative, covering Text to speech and voice and Transcription. The repository describes itself as: 13 Claude Code skills for video production (transcribe / translate / dub / multicam / subtitles / reframe) + WeChat publishing. Compatible with Claude Code, OpenAI Codex CLI… The licence is MIT.
3 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit b2690f5. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 2 files in scripts/ (Python), which the agent can run.
Shell commands in SKILL.md call:
pythonuvuvxFrom the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
openspeech.bytedance.comFrom URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
VOLC_TTS_ACCESS_TOKENFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Wjs Dubbing Video loads about 5.3k tokens when it runs. Until then it costs about 153 tokens; SKILL.md has 2,253 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check noted patterns worth knowing about, such as sudo or a known installer.
entials: most users keep them in `~/code/.env`. Read them at the top of any session via:set -a; source ~/code/.env; set +aAutomated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from jianshuo/claude-skills at commit b2690f5, republished under its MIT licence (© jianshuo). 2,253 words, ~5,287 tokens.
.claude/skills/wjs-dubbing-video/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.Video + target-language SRT → *_<lang>_dub.mp4 with a time-aligned TTS voice. This skill stops at the dub track. Burn-in + audio bed mixing is the next skill (/wjs-burning-subtitles/render.py composites everything in one final encode).
entrevista.zh-CN.srt) and wants the video to speak that language./wjs-transcribing-audio then /wjs-translating-subtitles first./wjs-burning-subtitles.Default: assume one speaker. Use a single voice for the entire dub. This is the right answer for monologues, vlogs, recorded talks, narrator-only clips, and the overwhelming majority of videos people ask about. Don't run diarization, don't tag the SRT with [A]/[B], don't bring up multi-speaker complexity.
Switch to multi-speaker only when the user explicitly says so — phrasings like "two people", "interview", "dialogue", "conversation between", "separate the speakers", "different voice for each", or a direct request to do diarization. When triggered, follow the "Multi-speaker dubbing" section below.
If you're unsure whether a video is one speaker or many, ship the single-voice version first. Adding speaker separation later is cheap (just regenerate the dub); shipping confused multi-speaker output by default wastes the user's time.
scripts/dub.py auto-routes by voice-ID prefix:
| Voice ID pattern | Engine | Auth |
|---|---|---|
zh_..._bigtts | Volcano (字节跳动豆包) TTS | VOLC_TTS_APPID + VOLC_TTS_ACCESS_TOKEN |
zh-CN-...Neural / en-US-...Neural / etc. | edge-tts (Microsoft Edge neural) | none (free) |
For Mandarin, Volcano is markedly more natural than edge-tts, especially for emotional/contemplative content. Use edge-tts when Volcano credentials aren't available or as a debugging fallback.
Endpoint: https://openspeech.bytedance.com/api/v3/tts/unidirectional (used for both TTS 1.0 and 2.0; the Resource-Id header picks the backend).
Headers:
X-Api-App-Id: (env: VOLC_TTS_APPID) # 10-digit speech App ID
X-Api-Access-Key: (env: VOLC_TTS_ACCESS_TOKEN) # 32-char token from speech console
X-Api-Resource-Id: volc.service_type.10029 # see resource ID note below
Content-Type: application/jsonLoading credentials: most users keep them in ~/code/.env. Read them at the top of any session via:
set -a; source ~/code/.env; set +aThe doc lists seed-tts-2.0 as the "TTS 2.0 (recommended)" resource, but a typical TTS-SeedTTS2.0 console instance does not include the popular *_bigtts speaker catalog (爽快斯斯, 高冷御姐, 开朗姐姐, etc.). Trying those speakers against seed-tts-2.0 returns 200 code=55000000 "resource ID is mismatched with speaker related resource". The fix is to use volc.service_type.10029 (the TTS 1.0 V3 endpoint) — the audio quality of the bigtts speakers is identical, and they all work against this resource. The bundled dub.py defaults to volc.service_type.10029; override with VOLC_TTS_RESOURCE env if you have a different instance.
Other 401/403 errors:
401 code=45000010 "load grant: requested grant not found in SaaS storage" — the App ID + key combo is valid against the gateway, but the user has not activated this resource. They must go to 火山引擎 → 语音技术 → 语音合成大模型 → 实例管理 and 开通 the service. No workaround.403 code=45000030 — the speaker isn't included in the user's instance bundle.Despite the doc's casual language, the response is streaming NDJSON, not a single JSON object and not raw audio bytes. Each line is a separate JSON event with a base64-encoded MP3 chunk in data. The terminal event has code: 20000000 (which means OK in this API's success codes — different from code: 0). Concatenate the decoded chunks for the full MP3.
import base64, json, requests
audio = b""
r = requests.post(url, headers=h, json=payload, timeout=60, stream=True)
for line in r.iter_lines():
if not line: continue
evt = json.loads(line)
if evt.get("code") not in (0, None, 20000000):
raise RuntimeError(f"code={evt.get('code')} {evt.get('message')}")
if evt.get("data"):
audio += base64.b64decode(evt["data"])volc.service_type.10029)Full list at volcengine.com/docs/6561/1257544 — but availability depends on your instance bundle. Confirmed-working female voices for the typical SeedTTS-2.0 starter instance:
| Speaker ID | 中文名 | Feel |
|---|---|---|
zh_female_gaolengyujie_moon_bigtts | 高冷御姐 | Best for contemplative/spiritual content. Mature, restrained, calm. |
zh_female_kailangjiejie_moon_bigtts | 开朗姐姐 | Warm older-sister storytelling. |
zh_female_shuangkuaisisi_moon_bigtts | 爽快斯斯 | Versatile, conversational baseline. |
zh_female_linjianvhai_moon_bigtts | 邻家女孩 | Casual, lifestyle-vlog. |
zh_female_yuanqinvyou_moon_bigtts | 元气女友 | Lively, upbeat. |
zh_female_meilinvyou_moon_bigtts | 美丽女友 | Soft, intimate. |
zh_female_shuangkuaisisi_emo_v2_mars_bigtts | 斯斯情感版 | Full emotional range — pair with explicit emotion + scale. |
These voices return 55000000 against the typical instance even though the doc lists them: vv_uranus_bigtts, wenroushunv_moon_bigtts, qingxin_moon_bigtts, yingmaoxiaoyuan_moon_bigtts, tianxinxiaoling_moon_bigtts, shaoergushi_moon_bigtts. Don't promise them without testing.
speech_rate is Volcano's native scale [-50, +100] where the value is a percentage delta (so -8 means 8% slower). The script passes --rate -8% through as -8.
Useful emotion presets:
emotion="calm", emotion_scale=4 — contemplative, default for this skill's spiritual-content niche.emotion="gentle" — softer / more intimate.emotion="neutral" — flat / informational.emotion="sad" — melancholic. Use sparingly.Override dub.py defaults with VOLC_TTS_EMOTION and VOLC_TTS_EMOTION_SCALE env vars without editing code.
No English Volcano voices are wired up in this skill — for English use edge-tts (next section). Volcano does have English speakers (en_male_*_bigtts, en_female_*_bigtts) but they aren't typically included in TTS-SeedTTS-2.0 starter instances. Add them by extending the voice routing in dub.py once verified.
Free, no API key, high-quality but less expressive than Volcano. Install into a project venv — do not call it via uvx once per segment. Each uvx invocation spawns a fresh Python process and the bing endpoint will rate-limit or RST the connection after a handful of rapid hits, breaking mid-render.
uv venv .venv
uv pip install --python .venv/bin/python edge-ttsThen drive it from a single long-lived Python process using edge_tts.Communicate(...) directly, with retry-on-failure logic. The bundled scripts/dub.py does this.
There is no perfect cross-language match — choose gender, age feel, and tone deliberately, then bend with rate/pitch.
Volcano's zh_female_gaolengyujie_moon_bigtts (高冷御姐, calm, speech_rate=-8) is the validated baseline for mature contemplative female speakers — equivalent to or better than any edge-tts option for that profile. See the Volcano speaker table above for the rest.
edge-tts catalog (Chinese):
| Voice | Gender | Default feel |
|---|---|---|
zh-CN-XiaoxiaoNeural | F | Warm, news/novel |
zh-CN-XiaoyiNeural | F | Lively, young |
zh-CN-YunjianNeural | M | Passionate, sports |
zh-CN-YunxiNeural | M | Sunshine, lively |
zh-CN-YunyangNeural | M | Professional newsreader |
zh-HK-HiuMaanNeural | F | Friendly, slightly mature |
zh-TW-HsiaoChenNeural | F | Friendly |
All voices below speak fluent American/British/Australian English; the *Multilingual* ones also handle Spanish names, French/Italian loanwords, etc. without mispronunciation.
| Voice | Gender | Default feel |
|---|---|---|
en-US-AvaMultilingualNeural | F | Best for warm/mature/caring — natural for spiritual or coaching content |
en-US-EmmaMultilingualNeural | F | Cheerful, conversational, younger |
en-US-AndrewMultilingualNeural | M | Warm, confident, sincere |
en-US-BrianMultilingualNeural | M | Approachable, casual |
en-US-AriaNeural | F | Crisp newsreader |
en-US-GuyNeural | M | Steady male newsreader |
en-GB-SoniaNeural | F | British female (RP) |
en-GB-RyanNeural | M | British male (RP) |
en-AU-WilliamMultilingualNeural | M | Australian male |
fr-FR-VivienneMultilingualNeural | F | Mature European female who also reads English |
For matching a mature contemplative Spanish female (this skill's canonical use case), start with en-US-AvaMultilingualNeural at --rate -5% --pitch -3Hz. Do not use the news-style Aria or Guy for spiritual content — they sound clinical.
zh-CN-XiaoxiaoNeural with --rate=-8% --pitch=-10Hz (or Volcano gaolengyujie).zh-CN-YunyangNeural with --rate=-5%. Avoid Yunjian/Yunxi (too energetic).*MultilingualNeural voices.🛑 Checkpoint — sample before full dub. A full-video dub is the most expensive step (TTS API calls + atempo + ffmpeg mux). Before running dub.py over the whole SRT:
Skip the checkpoint only if the user named a specific voice up front AND has already heard a sample of that voice on this video.
The script's scripts/sample_voices.py (if present) is a thin wrapper for exactly this; otherwise drive the same Python loop the dub script uses.
Mandatory smoke test before promising any Volcano voice on a new account: synth one ~5-word cue with that speaker ID first; only quote it to the user if the smoke test returns a non-empty MP3. If the smoke test 401s with code=45000010 ("grant not found"), tell the user they need to 开通 the resource in 火山引擎 console — do not pretend it'll work after a retry.
.venv/bin/python ~/.claude/skills/wjs-dubbing-video/scripts/dub.py [voice] [rate] [pitch]
# Mature Chinese contemplative female (Volcano):
.venv/bin/python ~/.claude/skills/wjs-dubbing-video/scripts/dub.py \
zh_female_gaolengyujie_moon_bigtts -8% +0Hz
# Warm English caring female (edge-tts, multilingual):
.venv/bin/python ~/.claude/skills/wjs-dubbing-video/scripts/dub.py \
en-US-AvaMultilingualNeural -5% -3Hz
# Default Chinese fallback (no Volcano creds needed):
.venv/bin/python ~/.claude/skills/wjs-dubbing-video/scripts/dub.py \
zh-CN-XiaoxiaoNeural -8% -10HzThe script:
*.zh-CN.srt, *.en.srt, etc., or pass --srt).dub_work/seg_NN.mp3.ffprobe.atempo filters to speed it up; if shorter, pads with silence after.*_zh_dub.mp4 / *_en_dub.mp4 keeping the original video stream by -c:v copy.Output: <source-stem>_<lang>_dub.mp4 (e.g., entrevista_zh_dub.mp4). This is the input for the next step — /wjs-burning-subtitles/render.py — which composites the final video.
Mandarin takes 60–80% of the time Spanish does to say the same thing. With strict cue-by-cue timing, that leaves awkward 2–4s silences at the end of most cues. English is closer to ~85% of Spanish. Three levers, in increasing impact:
Slow the native TTS rate. Changing --rate from +0% to -12% to -15% produces clean, natural-sounding slower speech (much better than time-stretching afterward). Try -12% first; -15%/-20% for very contemplative content.
Mild slow-stretch per cue. When a cue's TTS is still shorter than its slot, run atempo between 0.82× and 0.95×. dub.py does this automatically: when slack > 0.5s, it sets atempo = max(0.82, tts_dur / target_dur) and pads the remainder. Below 0.82× the voice starts sounding drugged; above 0.92× the stretch is essentially imperceptible.
Expand the target-language text in the worst cues. When the slot is so long that even 0.82× stretch leaves >2s of silence, the cleanest fix is to lengthen the translation. Add natural Mandarin particles ("嗯,", "其实", "也就是说", "你知道") or unpack a compressed phrase into its full meaning. This changes the on-screen subtitle, so confirm with the user before doing it. Edit the SRT, regenerate just those segments by deleting their dub_work/seg_NN.mp3 and re-running dub.py.
Combine the levers: native rate -12% + stretch-to-fit handles ~80% of cases. Reserve text expansion for the 2–3 worst outliers.
Only invoke this section when the user explicitly says the source has multiple speakers ("interview", "two people", "dialogue", "separate the speakers", "different voice for each", or a direct request to do diarization).
When triggered, generate the dub with a different voice per speaker so the listener can follow who's speaking. Two paths:
scripts/visual_diarize.py watches mouth movement per face per frame and tags each cue with the dominant speaker. Self-contained, no API keys, no audio fingerprinting.
uv pip install --python .venv/bin/python mediapipe opencv-python
.venv/bin/python ~/.claude/skills/wjs-dubbing-video/scripts/visual_diarize.py \
--video input.mp4 --srt input.en.srt \
--out input.en.diarized.srt \
--report diarization_report.json \
--sample-fps 5 --num-speakers 2How it works:
--num-speakers faces per frame, 478 landmarks each.A, B, ... left-to-right.[A]/[B]-prefixed SRT plus a JSON report with per-cue scores and a confidence ratio (winner / runner-up).On first run, downloads the FaceLandmarker model (~3.6 MB) to /tmp/mp_models/face_landmarker.task.
Visual is materially better than guessing from text. In one validation, manual text-based labels split 6/50 between speakers; visual diarization showed the actual split was 29/27 — text-based guessing was wildly wrong because both people take similar-shaped turns. Always prefer visual when the speakers are on camera.
Spot-check low-confidence cues. Any cue in the JSON report with confidence_ratio < 1.5 is borderline — usually overlapping speech or one speaker briefly off-frame. Hand-correct before dubbing.
For very short clips (1–2 minutes), or when speakers are off-camera, or when visual diarization fails:
1
00:00:00,000 --> 00:00:03,400
[A] So what about that AI rewrite thing?
2
00:00:03,400 --> 00:00:08,200
[B] Right — let me explain the workflow.Save as *.tagged.srt. Keep the clean SRT (without tags) for downstream burn-in via /wjs-burning-subtitles.
Pass --voice-map with speaker=voice pairs. The positional voice arg is the default for cues with no tag.
.venv/bin/python ~/.claude/skills/wjs-dubbing-video/scripts/dub.py \
en-US-AndrewMultilingualNeural -3% +0Hz \
--srt input.en.tagged.srt \
--voice-map "A=en-US-BrianMultilingualNeural,B=en-US-AndrewMultilingualNeural"Voice-pairing tips:
en-US- and en-GB- for distinctness.zh_female_gaolengyujie_moon_bigtts (mature) + zh_female_kailangjiejie_moon_bigtts (warm sister).Visual diarization fails when:
For audio-only material (podcasts, voice-overs), fall back to pyannote.audio or whisperx --diarize. This skill does not yet bundle audio-based diarization.
<source-stem>_<lang>_dub.mp4 — video stream-copied from source, audio replaced with the time-aligned dub track. Drop-in input for /wjs-burning-subtitles/render.py.dub_work/seg_NN.mp3 — per-cue TTS clips (kept for resume / per-cue regen)./wjs-burning-subtitles — to mix the original audio as a low-volume bed, burn the SRT, or both. The final encode happens there in one ffmpeg pass (no cascade). Pass --video <source.mp4> --dub <source_lang_dub.mp4> [--srt <srt>] to its render.py.*_<lang>_dub.mp4) is technically a finished video and can ship as-is, but it sounds dubbed (because it is). Mixing the original underneath gives the "professional translation" feel — do that in /wjs-burning-subtitles.uvx edge-tts once per cue. Spawns a Python process each time; bing endpoint rate-limits or RSTs mid-render. Use the persistent library path in dub.py.audio_source without listening. Always sample a 30 s clip before committing.[A]. Wastes time and the dub sounds the same. Default to one voice.code=55000000 against typical SeedTTS-2.0 starter bundles. Always synth a 5-word smoke test before quoting.code=20000000, not code=0. Concatenate base64-decoded data chunks for the full MP3.© jianshuo, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 2 other files (scripts) in wjs-dubbing-video of jianshuo/claude-skills.
Open the folder on GitHubat commit b2690f5
Wjs Dubbing Video next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Wjs Dubbing Video this skilljianshuo/claude-skills | 131 | — | ~5.3k | Automated safety check: Notes | MIT | |
| HyperFrames Media Useheygen-com/hyperframes | 60k | — | ~2.4k | Automated safety check: Pass | Apache-2.0 | |
| Edu Chem Videowy51ai/edulab | 1.4k | — | ~2.1k | Automated safety check: Notes | Apache-2.0 | |
| Edu Math Videowy51ai/edulab | 1.4k | — | ~2.5k | Automated safety check: Notes | Apache-2.0 | |
| Edu Physics Videowy51ai/edulab | 1.4k | — | ~2.3k | Automated safety check: Notes | Apache-2.0 | |
| Elevenlabs Transcribeqdhenry/Claude-Command-Suite | 1.3k | — | ~1.5k | Automated safety check: Notes | None |
heygen-com/hyperframes
Finds, generates and edits media for HyperFrames video projects: music, sound effects, images, icons, logos, voiceovers, captions and color grades.
wy51ai/edulab
A skill your agent uses when asked to make an explainer / walkthrough video (讲解视频、解题视频、例题精讲、微课) for a chemistry problem (化学题: 氧化还原配平 双线桥 电子守恒, 物质的量计算, 化学平衡 三段式 平衡常数 转化率 反应速率, 离子反应, 电化学, 溶液 滴定…
wy51ai/edulab
A skill your agent uses when asked to make an explainer / walkthrough video (讲解视频、解题视频、例题精讲、微课) for a math problem (数学题, geometry, algebra, functions, motion/行程 problems), from a problem screenshot…
wy51ai/edulab
A skill your agent uses when asked to make an explainer / walkthrough video (讲解视频、解题视频、例题精讲、微课) for a physics problem (物理题: mechanics/力学 受力分析 牛顿定律 斜面 传送带 板块 平抛 圆周 能量 动量, optics/光学 折射…
qdhenry/Claude-Command-Suite
Transcribes audio/video files using ElevenLabs Scribe v2 API.
zenstory-ai/video-recap-skills
合成视频解说最终成片:把旁白音频铺到源视频上,按旁白窗口压低原声,生成 SRT / ASS 字幕并可烧录, 最后做响度标准化。作为最终合成阶段使用。输入源视频、ttsmeta.json 与旁白位置; 输出 recap 成片和字幕。触发词:视频合成、混音、字幕、压字幕、assemble video、mux、ducking、subtitles、成片。
jianshuo/claude-skills
A skill your agent uses when the user has a long-form video (interview / lecture / podcast / conversation) and a transcript SRT, and wants to extract 3–6 stand-alone topical short clips from it.
jianshuo/claude-skills
Upload one or many videos to YouTube. An agent skill from jianshuo/claude-skills.
jianshuo/claude-skills
A skill your agent uses when migrating a WordPress site to a Hugo static site on GitHub Pages from a WXR export (.xml) plus the wp-content/uploads folder — preserving /archives/<id/ URLs, localizing…
jianshuo/claude-skills
A skill your agent uses when the user has a video + an SRT and wants the subtitles either burned into the pixels (libass, always-visible) or soft-muxed as a togglable track.
jianshuo/claude-skills
A skill your agent uses when the user complains about spam on his X/Twitter posts — 同城面付 / 寻固炮 / 线下上门 / 免费破处 这类引流号在他推文下刷的 emoji 垃圾回复 — and wants them removed.
jianshuo/claude-skills
A skill your agent uses when the user wants a book turned into YouTube chapter videos — 每章用 VoiceDrop 读书的有声书 mp3 做音轨,配 GPT Image 2 画面和中心思想大字,输出 1920×1080 横屏视频发 YouTube。Triggers — "把这本书做成视频"…
Categories
A skill your agent uses when the user has a video + a target-language SRT and wants the video to actually speak that language — generates a time-aligned TTS voice dub. Wjs Dubbing Video is an agent skill from jianshuo/claude-skills. Use when the user has a video + a target-language SRT and wants the video to actually speak that language — generates a time-aligned TTS voice dub.
Wjs Dubbing Video fits situations like: the user has a video + a target-language SRT and wants the video to actually speak that language — generates a time-aligned TTS voice dub; voice over this; different voice for each speaker.
Run `npx skills add jianshuo/claude-skills --skill wjs-dubbing-video -a claude-code`. Or copy the skill folder (wjs-dubbing-video in jianshuo/claude-skills) into .claude/skills/wjs-dubbing-video in your project. Claude Code loads it when a task matches its description.
Run `npx skills add jianshuo/claude-skills --skill wjs-dubbing-video -a codex`. Or copy the skill folder (wjs-dubbing-video in jianshuo/claude-skills) into .agents/skills/wjs-dubbing-video in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jianshuo/claude-skills --skill wjs-dubbing-video -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/wjs-dubbing-video, .gemini/skills/wjs-dubbing-video, .github/skills/wjs-dubbing-video and .opencode/skills/wjs-dubbing-video in your project.
Going by SKILL.md and its folder, Wjs Dubbing Video needs Python for the scripts in its folder, the command-line tools its instructions call (python, uv and uvx) and credentials named VOLC_TTS_ACCESS_TOKEN. Our summary lists: Python 3; A credential in VOLC_TTS_ACCESS_TOKEN.
SKILL.md names 1 domain. In commands or code: openspeech.bytedance.com; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found notes only (mentions a .env file), nothing it rates as a warning. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Wjs Dubbing Video is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 5.3k tokens (SKILL.md is roughly 21k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Wjs Dubbing Video: HyperFrames Media Use (heygen-com/hyperframes, 60k stars), Edu Chem Video (wy51ai/edulab, 1.4k stars), Edu Math Video (wy51ai/edulab, 1.4k stars) and Edu Physics Video (wy51ai/edulab, 1.4k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
jianshuo (a GitHub user) maintains it in jianshuo/claude-skills, which has 131 GitHub stars. The repository holds 38 skills in this directory. The repository was last updated on August 20, 2026.
Source: jianshuo/claude-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.