Agent skill

Video Translator

by shang-zhu in shang-zhu/violin

Dub a video into another language and generate subtitles using the default Together + Cartesia stack.

MITAuto-check: notesMedia & Creative

Install Video Translator

skills CLI
$ npx skills add shang-zhu/violin --skill video-translator -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install shang-zhu/violin video-translator --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/shang-zhu/violin.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/video-translator .claude/skills/video-translator && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
video-translator
GitHub stars
1.1k
Token cost
~1k tokens
SKILL.md length
451 words
Files
1
Skills in repo
1
Repo updated
First seen
Licence
MIT

At a glance

Dub a video into another language and generate subtitles using the default Together + Cartesia stack.

  • The user wants to translate / dub / voice-over a video file
  • SKILL.md covers Pre-flight, Decisions, Run and Flags, plus 3 more sections
  • Calls uv; needs TOGETHER_API_KEY
  • Generate subtitles for it

What it does

Video Translator is an agent skill from shang-zhu/violin. Dub a video into another language and generate subtitles using the default Together + Cartesia stack. Trigger when the user wants to translate / dub / voice-over a video file, or generate subtitles for it. Handles .mp4 / .mkv / .webm. Installs as the violin CLI (and violin-api for the FastAPI server) via uv tool install. For alternative models (OpenAI / ElevenLabs) or custom configs, point the user to the repo: https://github.com/shang-zhu/violin.

Its SKILL.md is about 1k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Media & Creative, covering Text to speech and voice, Translation and Transcription. It works with ElevenLabs, OpenAI, GitHub and FastAPI. The repository describes itself as: Open-source Video Translation Skill. The licence is MIT.

When your agent uses it

  • The user wants to translate / dub / voice-over a video file
  • Generate subtitles for it

Example prompts

  • “/video-translator”

Requirements

  • A credential in TOGETHER_API_KEY
  • Pre-approved tools (allowed-tools): Bash, Read

What it can do on your machine

Read from SKILL.md and the folder at commit 7a86852. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Bash
    • Read

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • uv

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use uv, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • TOGETHER_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Video Translator loads about 1k tokens when it runs. Until then it costs about 120 tokens; SKILL.md has 451 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~120
When it runs · the whole SKILL.md, loaded when a task matches
~1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteMentions a .env fileSKILL.md:24
    - Inside the Violin repo → populate `.env` (auto-loaded)
  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Bash, Read

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from shang-zhu/violin at commit 7a86852, republished under its MIT licence (© shang-zhu). 451 words, ~1,036 tokens.

Download SKILL.mdSave it as .claude/skills/video-translator/SKILL.md (or your agent's skills folder).
name
video-translator
description
Dub a video into another language and generate subtitles using the default Together + Cartesia stack. Trigger when the user wants to translate / dub / voice-over a video file, or generate subtitles for it. Handles `.mp4` / `.mkv` / `.webm`. Installs as the `violin` CLI (and `violin-api` for the FastAPI server) via `uv tool install`. For alternative models (OpenAI / ElevenLabs) or custom configs, point the user to the repo: https://github.com/shang-zhu/violin.
allowed-tools
Bash, Read

Violin — operating skill

Always uses the default config (Together for translation, cartesia/sonic-3 for TTS). If the user asks for OpenAI, ElevenLabs, or custom configs, stop and point them to the Violin repo — those flows aren't supported through the global CLI.

Pre-flight

Run these silently first. Abort if any fails:

bash
command -v violin                 # 1. CLI on PATH
test -f "<input>"                 # 2. Input exists
printenv TOGETHER_API_KEY         # 3. Key available

If violin is missing: tell the user to uv tool install violin, then violin --install-skill to refresh this skill file. Do not auto-install.

If TOGETHER_API_KEY is missing:

  • Inside the Violin repo → populate .env (auto-loaded)
  • Elsewhere → export TOGETHER_API_KEY=... in ~/.zshrc / ~/.bashrc, then source it

Decisions

  • CLI vs API: single run-and-wait file → CLI (violin). Multi-job / HTTP / web UI → API server (violin-api); print the command, don't auto-start it.
  • Style (--style): default standard. Kids content → kids, formal/lecture → academic, casual → casual, dramatic → storyteller, news → news. Run violin --style list if unsure.
  • Voiceover: keep default (mix dubbed audio over a quiet original). Use --no-voiceover only when the user explicitly says "replace audio entirely".

Run

bash
violin <input> <output> --language <Lang> [flags]

Flags

FlagDefaultWhen to set
--language / -lrequiredTarget language (e.g. Chinese, Spanish, Japanese).
--voice / -vauto (native voice picked by preferences.voice_gender)Only when the user names a specific voice from the catalog (e.g. "warm female narrator"). Otherwise omit and let the default kick in.
--source-languageauto-detectOnly if Whisper mis-detects the source language.
--style / -sstandardSee Decisions above.
--no-subtitlesoffUser says "no SRT" / "video only".
--no-voiceoveroffUser says "replace original audio entirely".
--config / -cconfig/default.yamlDon't use through this skill — repo-only flow.
--timings-outoffOnly when the user wants a per-step timing JSON for debugging / benchmarking.
Show full SKILL.md (193 more words)Show less

Language coverage

33 target languages total. 16 ship with handpicked native-speaker voices: Chinese, Spanish, English, Hindi, Arabic, Portuguese, Russian, Japanese, Turkish, German, Korean, French, Italian, Polish, Dutch, Swedish. The other 17 fall back to the English voice catalog (multilingual under Cartesia Sonic 3) — quality is decent but the voice isn't a native speaker. Mention this caveat only if the user is translating to a fallback language and asks about voice quality.

Report back

  • Output video path + SRT path (printed by the run).
  • Total cost (printed at end — surface, don't hide).
  • If voiceover was on, mention the _original.m4a sidecar.

Don'ts

  • Don't run on multi-GB videos without first quoting the rough cost (audio length × per-provider rates in pipeline/pricing.py).
  • Don't fabricate a "subtitles-only" mode — the CLI requires the full pipeline. If the user only wants SRT, run the full pipeline and hand them just the .srt, warning them of the cost first.
  • Don't try to switch to OpenAI or ElevenLabs from this skill. Point the user to the repo + --config config/other_api.yaml (or their own override).
  • Don't paraphrase the README. For supported languages (33), voice catalog, and full flag docs, point them at README.md or violin --help.

© shang-zhu, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/video-translator of shang-zhu/violin.

Open the folder on GitHubat commit 7a86852

Compare with similar skills

Video Translator next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Video Translator compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Video Translator this skillshang-zhu/violin1.1k—~1kAutomated safety check: NotesMIT
Local AI Useamd/skills408—~5kAutomated safety check: NotesMIT
AI SDK Developmenttrypostit/trypost6921 repos~3.5kAutomated safety check: PassMIT
Speech To Texttadaspetra/loop2962 repos~2kAutomated safety check: PassMIT
Edu Math Videowy51ai/edulab1.4k—~2.5kAutomated safety check: NotesApache-2.0
Hyperframes Mediachmonitor/chmonitor3011 repos~2.8kAutomated safety check: NotesGPL-3.0

Similar skills

  • Local AI Use

    amd/skills

    Makes this agent generate images, transcribe audio, and synthesize speech on the user's own machine through a local Lemonade Server instead of a paid cloud API.

    408 GitHub stars~5k tokensUpdated 2 days ago
    Media & CreativeAuto-check: notes
  • AI SDK Development

    trypostit/trypost

    TRIGGER when working with ai-sdk which is Laravel official first-party AI SDK.

    692 GitHub starsUsed in 1 repo~3.5k tokens
    AI & LLM EngineeringAuto-check passed
  • Speech To Text

    tadaspetra/loop

    Transcribe audio to text using ElevenLabs Scribe v2. An agent skill from tadaspetra/loop.

    296 GitHub starsUsed in 2 repos~2k tokens
    Media & CreativeAuto-check passed
  • Edu Math Video

    wy51ai/edulab

    A skill your agent uses when asked to make an explainer / walkthrough video (讲解视频、解题视频、例题精讲、微课) for a math problem (数学题, geometry, algebra, functions, motion/行程 problems), from a problem screenshot…

    1.4k GitHub stars~2.5k tokensUpdated yesterday
    Media & CreativeAuto-check: notes
  • Hyperframes Media

    chmonitor/chmonitor

    Audio and media assets for HyperFrames compositions, produced by one shared audio engine (scripts/audio.mjs) — multi-provider TTS (HeyGen / ElevenLabs / Kokoro local), background music + sound…

    301 GitHub starsUsed in 1 repo~2.8k tokens
    Media & CreativeAuto-check: notes
  • 9Router Text to Speech

    decolua/9router

    Turns text into spoken audio through a 9Router server, choosing among voices from OpenAI, ElevenLabs, Deepgram, Edge TTS and other providers.

    31k GitHub stars~765 tokensUpdated 3 days ago
    Media & CreativeAuto-check passed

Questions about Video Translator

What does Video Translator do?

Dub a video into another language and generate subtitles using the default Together + Cartesia stack. Video Translator is an agent skill from shang-zhu/violin. Dub a video into another language and generate subtitles using the default Together + Cartesia stack.

When should I use Video Translator?

Video Translator fits situations like: the user wants to translate / dub / voice-over a video file; generate subtitles for it.

How do I install Video Translator in Claude Code?

Run `npx skills add shang-zhu/violin --skill video-translator -a claude-code`. Or copy the skill folder (.claude/skills/video-translator in shang-zhu/violin) into .claude/skills/video-translator in your project. Claude Code loads it when a task matches its description.

How do I install Video Translator in Codex?

Run `npx skills add shang-zhu/violin --skill video-translator -a codex`. Or copy the skill folder (.claude/skills/video-translator in shang-zhu/violin) into .agents/skills/video-translator in your project. Codex loads it when a task matches its description.

Can I use Video Translator in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add shang-zhu/violin --skill video-translator -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/video-translator, .gemini/skills/video-translator, .github/skills/video-translator and .opencode/skills/video-translator in your project.

What does Video Translator need to run?

Going by SKILL.md and its folder, Video Translator needs the command-line tools its instructions call (uv) and credentials named TOGETHER_API_KEY. Our summary lists: A credential in TOGETHER_API_KEY. Its frontmatter pre-approves these tools: Bash, Read.

Does Video Translator access the network?

SKILL.md contains no URLs. Its commands use uv, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Video Translator safe to install?

Our automated static check of SKILL.md found notes only (mentions a .env file; pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Video Translator use?

Video Translator is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Video Translator use?

About 1k tokens (SKILL.md is roughly 4.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Video Translator?

Skills that share tags, products or a category with Video Translator: Local AI Use (amd/skills, 408 stars), AI SDK Development (trypostit/trypost, 692 stars), Speech To Text (tadaspetra/loop, 296 stars) and Edu Math Video (wy51ai/edulab, 1.4k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Video Translator?

shang-zhu (a GitHub user) maintains it in shang-zhu/violin, which has 1,075 GitHub stars. The repository was last updated on September 4, 2026.

Source: shang-zhu/violin on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.