Agent skill

Voiceover

by GTKottman in GTKottman/mortiflix-oss

Narration, sound effects and music beds in the voice the studio set up (ElevenLabs, Qwen3-TTS on this machine's GPU, or the owner's own voice from the recording booth): lines from the approved…

AGPL-3.0Auto-check passedMedia & Creative

Install Voiceover

skills CLI
$ npx skills add GTKottman/mortiflix-oss --skill voiceover -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install GTKottman/mortiflix-oss voiceover --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/GTKottman/mortiflix-oss.git skills-src && mkdir -p .claude/skills && cp -r skills-src/pipelines/_shared/skills/voiceover .claude/skills/voiceover && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
voiceover
GitHub stars
396
Token cost
~1.8k tokens
SKILL.md length
913 words
Files
3
Skills in repo
8
Repo updated
First seen
Licence
AGPL-3.0

At a glance

Narration, sound effects and music beds in the voice the studio set up (ElevenLabs, Qwen3-TTS on this machine's GPU, or the owner's own voice from the recording booth): lines from the approved…

  • Works in 7 steps: Which voice? → Write voice/lines.json → Speak and check → …
  • Any narrated animatic
  • SKILL.md covers 0. Which voice?, 1. Write voice/lines.json, 2. Speak and check and 3. Build, plus 3 more sections
  • Runs JavaScript scripts from its folder; calls node

What it does

Voiceover is an agent skill from GTKottman/mortiflix-oss. Narration, sound effects and music beds in the voice the studio set up (ElevenLabs, Qwen3-TTS on this machine's GPU, or the owner's own voice from the recording booth): lines from the approved script, checked takes, one voice track and a word-timing map the animation follows. Use for any narrated animatic or final, a re-take, or sound design.

Its SKILL.md is about 1.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files.

It sits in Media & Creative, covering Text to speech and voice and Music and audio generation. It works with ElevenLabs and Qwen. The repository describes itself as: A motion design studio on your own machine: Claude makes the video step by step, you approve every stage. Bring your own Claude Code or API key. The licence is AGPL-3.0.

When your agent uses it

  • Any narrated animatic
  • Tasks that involve Text to speech and voice
  • Tasks that involve Music and audio generation

Example prompts

  • “s GPU, or the owner”
  • “/voiceover”

Requirements

  • Node.js

Workflow steps

7 steps, taken from the step headings in SKILL.md.

  1. Which voice?
  2. Write voice/lines.json
  3. Speak and check
  4. Build
  5. Sound effects and music (ElevenLabs, when switched on)
  6. Mix
  7. Record

What it can do on your machine

Read from SKILL.md and the folder at commit aea9533. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (JavaScript), which the agent can run.

    Shell commands in SKILL.md call:

    • node

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Voiceover loads about 1.8k tokens when it runs. Until then it costs about 89 tokens; SKILL.md has 913 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~89
When it runs · the whole SKILL.md, loaded when a task matches
~1.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from GTKottman/mortiflix-oss at commit aea9533, republished under its AGPL-3.0 licence (© GTKottman). 913 words, ~1,835 tokens.

Download SKILL.mdSave it as .claude/skills/voiceover/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
voiceover
description
Narration, sound effects and music beds in the voice the studio set up (ElevenLabs, Qwen3-TTS on this machine's GPU, or the owner's own voice from the recording booth): lines from the approved script, checked takes, one voice track and a word-timing map the animation follows. Use for any narrated animatic or final, a re-take, or sound design.

Voiceover

The approved script is spoken word for word, and the narration is the clock the animation is timed to.

VO=.claude/skills/voiceover/vo.mjs
node $VO check                                   # which engine, voice and model; credits; the rules for this engine
node $VO speak voice/lines.json                  # every line, checked, retaken if words go missing
node $VO speak voice/lines.json --only b03-1     # just the lines you fixed
node $VO build voice/lines.json                  # voice/voice.wav + voice/timing.json

0. Which voice?

node $VO check says what the owner set up in Settings (you don't choose the engine or the voice):

  • elevenlabs: the voice, model and settings are chosen; the key is in your environment (never print it). check shows the plan and credits. On the free plan the audio is non-commercial and must credit ElevenLabs: say so in your submission note.
  • qwen: Qwen3-TTS on this machine's GPU through ComfyUI (free, local). If check says ComfyUI or the suite isn't ready, that's a mfx needs-you.
  • none: no narration. If the brief asked for one, ask at the next gate whether to go on-screen-text-only (the default) or wait while the owner sets a voice up. Never use a robotic system voice.
  • own: the owner narrates in their own voice. Write voice/lines.json exactly as below (short lines; put delivery notes as a [tag] at the start, which the owner sees as direction; skip IPA and v4 tricks, a person reads script). node $VO speak voice/lines.json then lists the lines still to record and prints the mfx needs-you text that sends the owner to the recording booth (the web studio, mortiflix record, or importing files). Send it and stop. When they resume, speak passes and build makes the track, timed per line. If you change a line after it was recorded, speak asks for that line again.
  • The owner's own recording in input/ (one long file, any engine): use it instead (cut it into lines as clips, then build).

Don't change the voice or the model yourself: they're the owner's choice. Suggest a change in the handoff if one would clearly be better.

1. Write voice/lines.json

One entry per sentence or breath group (under ~600 characters; short lines are cheap to retake):

json
[
  { "id": "b01-1", "text": "[warm, unhurried] Every city has a heartbeat.", "script": "Every city has a heartbeat.", "gap_after": 0.5 },
  { "id": "b01-2", "text": "Ours runs on \"/ˈbaɪsɪkəlz/\".", "script": "Ours runs on bicycles." }
]
  • text is what the voice reads; script is the same words spelled normally (needed whenever text has tags or IPA: the check compares against it). gap_after is the pause after the line (default 0.4 s; longer between beats).
  • Ids follow the script's beats (b01-1), letters, digits and dashes.
  • Write numbers, dates and money the way they're said ("twenty twenty-six", "four point five percent").
Directing ElevenLabs Eleven v4 (eleven_v4, the default)
  • Audio tags in square brackets steer delivery: [warm], [measured, curious], [whispers], [sighs], [excited], [quick, light pace]. One at the start of a line; another only where the mood really turns.
  • Describe the voice, not a sound. v4 also makes sound effects, so [rain] or [applause] can come out as a noise. Write [soft, hushed voice], not [quiet room]. The check flags any non-speech sound in a take.
  • Punctuation and capitals: ellipses add pauses and weight, CAPITALS add emphasis. No SSML: v4 ignores <break> and <phoneme> (they can be read aloud).
  • Pronunciation: IPA between slashes inside quotes, "/ˈkoʊmæl/", with the normal spelling in script.
  • v4 has only stability and similarity (set in Settings); there's no style or speed: direct pace with tags and punctuation.
  • Each line is sent with its neighbours' text, so the delivery flows across lines.

Other ElevenLabs models (if the owner chose one): eleven_multilingual_v2 has style and speed settings and takes <break time="1.0s" /> pauses (up to 3 s); eleven_flash_v2 takes <phoneme> tags. Don't use tags v4-style on them.

Show full SKILL.md (363 more words)Show less
Directing Qwen3-TTS (local)
  • Delivery comes from the instruction the owner set (1.7B model); the line is read as plain words. Tags and IPA are dropped before speaking (script is read when present), so spell hard names the way they sound in script.
  • 10 languages (English, Chinese, Japanese, Korean, German, French, Russian, Portuguese, Spanish, Italian).

2. Speak and check

speak makes each line, listens to it with speech to text (ElevenLabs Scribe v2, or Qwen3-ASR locally), and keeps it when at least 92% of the words match the script with no 3-word run missing or added, and no stray sounds. A line that fails is retaken with a new seed (up to --max-takes, default 3); the best take is kept either way. ElevenLabs lines run several at once (the plan's limit minus one); local lines run speak-all, then listen-all, because the two models take turns on the GPU.

  • Results: voice/clips/<id>.(mp3|wav) + <id>.json (engine, take, seed, the transcript, word timings), every take in voice/takes/, and voice/speak-report.json.
  • A line that keeps failing: fix the input, don't re-roll. Split a long sentence, spell out a number, add IPA (ElevenLabs) or a sounds-like spelling (local), then speak --only it.
  • Listen to the first two lines before making the rest: pace, pronunciation, tone. The check catches missing words, not taste.

3. Build

build joins the clips with their gaps into voice/voice.wav (48 kHz mono) and writes voice/timing.json: every line's start and end, and every word's start and end when the check measured them ("timing": "words"), otherwise the line only. Turn the times you animate to into frames (Math.round(seconds * fps)) in video/src/timing.ts, and copy voice.wav into video/public/. Never cut inside a word: cut in the silences.

4. Sound effects and music (ElevenLabs, when switched on)

node .claude/skills/voiceover/sound.mjs sfx "soft glassy whoosh, left to right" --seconds 1.2 --out sfx/whoosh.mp3
node .claude/skills/voiceover/sound.mjs sfx "low city ambience, distant traffic" --seconds 20 --loop --out sfx/city.mp3
node .claude/skills/voiceover/sound.mjs music "warm minimal synth bed, 90 bpm, hopeful" --seconds 45 --out music/bed.mp3

Sound effects: 0.5–30 s, --loop for seamless ambience, --influence 0–1 (how literally it follows the prompt). Music: instrumental unless --vocals, 3 s to 10 min. Note every prompt in assets/SOURCES.md.

5. Mix

Music about 18 dB under the voice while it speaks; it can come up in the gaps. The final mix: -14 LUFS, true peak ≤ -1 dBTP (final-pass skill).

6. Record

Write voice/VOICE.md: engine, voice, model, settings, tags used, pronunciation fixes, rejected takes and why.

© GTKottman, AGPL-3.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files in pipelines/_shared/skills/voiceover of GTKottman/mortiflix-oss.

  • SKILL.md
  • sound.mjs
  • vo.mjs

Open the folder on GitHubat commit aea9533

Compare with similar skills

Voiceover next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Voiceover compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Voiceover this skillGTKottman/mortiflix-oss396—~1.8kAutomated safety check: PassAGPL-3.0
Musictadaspetra/loop2963 repos~827Automated safety check: PassMIT
ElevenLabs Voiceover Generatordigitalsamba/claude-code-video-toolkit2.2k1 repos~2.7kAutomated safety check: NotesMIT
Sound Effectstadaspetra/loop2962 repos~1.1kAutomated safety check: PassMIT
Hyperframes Mediachmonitor/chmonitor2991 repos~2.8kAutomated safety check: NotesGPL-3.0
Z Qwen Audio Studiotjxj/z-skills546—~716Automated safety check: PassNone

Similar skills

  • Music

    tadaspetra/loop

    Generate music using ElevenLabs Music API. An agent skill from tadaspetra/loop.

    296 GitHub starsUsed in 3 repos~827 tokens
    Media & CreativeAuto-check passed
  • ElevenLabs Voiceover Generator

    digitalsamba/claude-code-video-toolkit

    Generates narration, sound effects and cloned voices through the ElevenLabs API, with model and setting choices tuned to the content's style.

    2.2k GitHub starsUsed in 1 repo~2.7k tokens
    Media & CreativeAuto-check: notes
  • Sound Effects

    tadaspetra/loop

    Generate sound effects from text descriptions using ElevenLabs.

    296 GitHub starsUsed in 2 repos~1.1k tokens
    Media & CreativeAuto-check passed
  • Hyperframes Media

    chmonitor/chmonitor

    Audio and media assets for HyperFrames compositions, produced by one shared audio engine (scripts/audio.mjs) — multi-provider TTS (HeyGen / ElevenLabs / Kokoro local), background music + sound…

    299 GitHub starsUsed in 1 repo~2.8k tokens
    Media & CreativeAuto-check: notes
  • Z Qwen Audio Studio

    tjxj/z-skills

    A skill your agent uses when creating complete generated audio with qwen-audio-3.1-tts-next, including podcasts, radio drama, advertisements, multiple speakers, reference voices, ambience, sound…

    546 GitHub stars~716 tokensUpdated 16 days ago
    Media & CreativeAuto-check passed
  • Release Video

    huytieu/COG-second-brain

    Turn a product release (the list of shipped items plus real screen recordings) into a motion recap video and one explained demo per feature, with sound effects tied to on-screen motion and a…

    1.3k GitHub stars~1.7k tokensUpdated 6 days ago
    Media & CreativeAuto-check passed

More from GTKottman/mortiflix-oss

All 8 skills in this repo
  • Blender 3D

    GTKottman/mortiflix-oss

    3D work in the studio's own Blender with its toolkits: MoBlend (MoGraph cloners, effectors, fields, MoText, fracture), Nova FX (particles, fire, sparks, fireworks), Camera (framing, shot presets…

    396 GitHub stars~902 tokensUpdated yesterday
    Auto-check passed
  • Music

    GTKottman/mortiflix-oss

    Scoring a video with original music written in Strudel, after the animatic is approved - a dramatic reading, a spotting map from the locked timing, a blueprint (motif, chart, roles, story sections…

    396 GitHub stars~2.6k tokensUpdated yesterday
    Auto-check passed
  • Remotion Motion

    GTKottman/mortiflix-oss

    Build, preview and render the video in Remotion (React): project setup from the template, stills for style frames, animatic and final renders through the studio's render queue.

    396 GitHub stars~806 tokensUpdated yesterday
    Auto-check passed
  • Transition Board

    GTKottman/mortiflix-oss

    Designing every cut of a video before the animatic - for each change from one style frame to the next, the object or idea that carries it, a transition chosen from the owner's remotion-transitions…

    396 GitHub stars~2.2k tokensUpdated yesterday
    Auto-check passed
  • Assets

    GTKottman/mortiflix-oss

    Finding and downloading stock assets (footage, images, illustrations, 3D models, sound effects, fonts) from the asset sites the owner uses, in the owner's own Chrome with browser-harness.

    396 GitHub stars~601 tokensUpdated yesterday
    Auto-check passed
  • Final Pass

    GTKottman/mortiflix-oss

    Check a rendered video before anyone else sees it: format, black or frozen frames, loudness, true peak, and a frame sheet to look at.

    396 GitHub stars~613 tokensUpdated yesterday
    Auto-check passed

Works with

Questions about Voiceover

What does Voiceover do?

Narration, sound effects and music beds in the voice the studio set up (ElevenLabs, Qwen3-TTS on this machine's GPU, or the owner's own voice from the recording booth): lines from the approved…. Voiceover is an agent skill from GTKottman/mortiflix-oss. Narration, sound effects and music beds in the voice the studio set up (ElevenLabs, Qwen3-TTS on this machine's GPU, or the owner's own voice from the recording booth): lines from the approved script, checked takes, one voice track and a word-timing map the animation follows.

When should I use Voiceover?

Voiceover fits situations like: any narrated animatic; tasks that involve Text to speech and voice; tasks that involve Music and audio generation.

How do I install Voiceover in Claude Code?

Run `npx skills add GTKottman/mortiflix-oss --skill voiceover -a claude-code`. Or copy the skill folder (pipelines/_shared/skills/voiceover in GTKottman/mortiflix-oss) into .claude/skills/voiceover in your project. Claude Code loads it when a task matches its description.

How do I install Voiceover in Codex?

Run `npx skills add GTKottman/mortiflix-oss --skill voiceover -a codex`. Or copy the skill folder (pipelines/_shared/skills/voiceover in GTKottman/mortiflix-oss) into .agents/skills/voiceover in your project. Codex loads it when a task matches its description.

Can I use Voiceover in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GTKottman/mortiflix-oss --skill voiceover -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/voiceover, .gemini/skills/voiceover, .github/skills/voiceover and .opencode/skills/voiceover in your project.

What does Voiceover need to run?

Going by SKILL.md and its folder, Voiceover needs JavaScript for the scripts in its folder and the command-line tools its instructions call (node). Our summary lists: Node.js.

Does Voiceover access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Voiceover safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Voiceover use?

Voiceover is published under the AGPL-3.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Voiceover use?

About 1.8k tokens (SKILL.md is roughly 7.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Voiceover?

Skills that share tags, products or a category with Voiceover: Music (tadaspetra/loop, 296 stars), ElevenLabs Voiceover Generator (digitalsamba/claude-code-video-toolkit, 2.2k stars), Sound Effects (tadaspetra/loop, 296 stars) and Hyperframes Media (chmonitor/chmonitor, 299 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Voiceover?

GTKottman (a GitHub user) maintains it in GTKottman/mortiflix-oss, which has 396 GitHub stars. The repository holds 8 skills in this directory. The repository was last updated on October 7, 2026.

Source: GTKottman/mortiflix-oss on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.