Agent skill

Offline Voice

by notivn in notivn/AIEV

How the on-device narration engine (VieNeu-TTS) and voice cloning work in this system - install tiers, the worker protocol, and the verified traps around torchaudio, voice metadata parsing and…

MITAuto-check passedMedia & Creative

Install Offline Voice

skills CLI
$ npx skills add notivn/AIEV --skill offline-voice -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install notivn/AIEV offline-voice --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/notivn/AIEV.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/offline-voice .claude/skills/offline-voice && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
offline-voice
GitHub stars
126
Token cost
~2k tokens
SKILL.md length
1,037 words
Files
1
Skills in repo
14
Repo updated
First seen
Licence
MIT

At a glance

How the on-device narration engine (VieNeu-TTS) and voice cloning work in this system - install tiers, the worker protocol, and the verified traps around torchaudio, voice metadata parsing and…

  • Works in 2 steps: Speech only: pip install vieneu.… → Cloning: additionally pip install torch…
  • Tasks that involve Text to speech and voice
  • SKILL.md covers Why VieNeu and not F5-TTS /…, Install tiers - keep them…, Known issues and Architecture notes, plus 1 more section
  • Calls pip; needs GEMINI_API_KEY

What it does

Offline Voice is an agent skill from notivn/AIEV. How the on-device narration engine (VieNeu-TTS) and voice cloning work in this system - install tiers, the worker protocol, and the verified traps around torchaudio, voice metadata parsing and duration. Read when working on Text to video narration, the Voices page, the /api/voices or /api/tts endpoints, or when a user reports that offline voices or voice cloning fail to install or sound wrong.

Its SKILL.md is about 2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Media & Creative, covering Text to speech and voice. It works with PyTorch. The repository describes itself as: Automatic AI video editing. Claude directs HyperFrames (HTML + GSAP motion graphics) and Remotion (timeline assembly) to turn raw footage into a finished MP4 - transcript… The licence is MIT.

When your agent uses it

  • Tasks that involve Text to speech and voice

Example prompts

  • “/offline-voice”

Requirements

  • Python 3
  • A credential in GEMINI_API_KEY

Workflow steps

2 steps, taken from the first numbered list in SKILL.md.

  1. Speech only: pip install vieneu. Torch-free (ONNX Runtime), ~30 MB of wheels. This is enough for all 14 preset voices.
  2. Cloning: additionally pip install torch torchaudio, 2-3 GB.

What it can do on your machine

Read from SKILL.md and the folder at commit 1a4c2b0. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • GEMINI_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Offline Voice loads about 2k tokens when it runs. Until then it costs about 103 tokens; SKILL.md has 1,037 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~103
When it runs · the whole SKILL.md, loaded when a task matches
~2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from notivn/AIEV at commit 1a4c2b0, republished under its MIT licence (© notivn). 1,037 words, ~2,015 tokens.

Download SKILL.mdSave it as .claude/skills/offline-voice/SKILL.md (or your agent's skills folder).
name
offline-voice
description
How the on-device narration engine (VieNeu-TTS) and voice cloning work in this system - install tiers, the worker protocol, and the verified traps around torchaudio, voice metadata parsing and duration. Read when working on Text to video narration, the Voices page, the /api/voices or /api/tts endpoints, or when a user reports that offline voices or voice cloning fail to install or sound wrong.

Offline voice - VieNeu-TTS and voice cloning

The system has two narration engines running side by side. They are deliberately not merged.

geminivieneu
Where it runsGoogle APIThe user's own machine
CostPer callFree
Needs networkYesNo
Voices30 preset14 Vietnamese preset + unlimited cloned
Voice cloningNo (API only accepts preset names)Yes
SpeedA few secondsRTF ~1.0 (about real time)
SetupGEMINI_API_KEYpip install vieneu

Keep both. Do not "simplify" by deleting one: a weak machine reads at roughly real time, which is unacceptable for a long script, while Gemini costs money on every preview click.

Why VieNeu and not F5-TTS / XTTS / Chatterbox

Checked against the requirement "Vietnamese + English, cloning, and shippable in an open-source repo":

  • Chatterbox (MIT, excellent English) has no Vietnamese. Disqualified on language.
  • F5-TTS code is MIT but the weights are CC-BY-NC, and every Vietnamese finetune is CC-BY-NC-SA. Disqualified on licence.
  • viXTTS / XTTS-v2 inherit Coqui's CPML non-commercial licence. Disqualified on licence.
  • VietTTS (dangvansam) is Apache-2.0 code but the installer is Linux-only, and its weights are CC-BY-NC too.
  • VieNeu-TTS is Apache 2.0 on the code and on the v3 Turbo weights (pnnbao-ump/VieNeu-TTS-v3-Turbo, trained from scratch on ~10k hours, not a finetune), plus Apache 2.0 on its codec dependency OpenMOSS-Team/MOSS-Audio-Tokenizer-Nano. Vietnamese-first with English code-switching, torch-free on CPU via ONNX, clones from 3-8 seconds of reference audio.

If a future engine is evaluated, check the weights licence separately from the code licence. That is where every other candidate failed.

Licence trap inside the package itself. vieneu/assets/voices.json carries "license": "CC BY-NC 4.0" and "Model and voices are for non-commercial use only". That notice is scoped to that file, which is the v2 preset pack: 6 voices named Binh, Tuyen, Vinh, Doan, Ly, Ngoc, default Binh. We run mode="v3turbo", which reads voices_v3_turbo.json (14 Vietnamese voices, default Phạm Tuyên). The two sets have zero overlap - verified by comparing the key sets - so nothing on our path is NC-licensed. If anyone ever switches the engine mode away from v3turbo, or a preset named Binh/Tuyen/Vinh/Doan/Ly/Ngoc shows up in the picker, the commercial position changes and must be rechecked.

Install tiers - keep them separate

Two independent levels. Never collapse them into one check:

  1. Speech only: pip install vieneu. Torch-free (ONNX Runtime), ~30 MB of wheels. This is enough for all 14 preset voices.
  2. Cloning: additionally pip install torch torchaudio, 2-3 GB.

start/doctor.mjs reports these as vieneu and vieneu-clone, and only asks about torch once vieneu is present. Forcing a 2-3 GB torch download on someone who only wants offline narration is the mistake to avoid.

The first synthesis after a server start downloads ~1 GB of model from HuggingFace and takes 30s; subsequent process starts take ~15s to load. Always surface this in the UI, otherwise the first click looks like a hang.

Known issues

Symptom: Cloning fails with ImportError: TorchCodec is required for load_with_torchcodec, or RuntimeError: Could not load libtorchcodec. Cause: add_voice() reads the reference through torchaudio.load(). From torchaudio 2.9 every backend routes through torchcodec, which needs FFmpeg shared libraries - the usual Windows FFmpeg builds are static, so it can never load. Installing torchcodec does not fix it. Passing backend="soundfile" does not fix it either (verified: all three backends raise the same error). Fix: monkeypatch torchaudio.load to use soundfile before importing vieneu. soundfile is already a vieneu dependency, so this adds nothing to install. Guard it in try/except so that a machine without torch still gets preset voices.

python
import torch, torchaudio, soundfile as sf
def _load(uri, *a, **kw):
    data, sr = sf.read(str(uri), dtype="float32", always_2d=True)
    return torch.from_numpy(data.T).contiguous(), sr
torchaudio.load = _load

Symptom: Voices are filed under the wrong region, or every male voice looks southern. Cause: list_preset_voices() returns labels shaped Name — <Nam|Nữ> · <Bắc|Trung|Nam> · Phong cách <...>. Field 1 is gender, field 2 is region, and the word "Nam" means both "male" and "southern". Fix: parse strictly by position, never by searching for the substring "Nam". Regression cases: Thái Sơn = gender nam / region nam; Mai Anh = gender nu / region bac.

Show full SKILL.md (399 more words)Show less

Symptom: Cloned voices vanish after a restart, or a stray voice appears in the preset list and leaks into vieneu's own error messages. Cause: tts.save_voices() (and add_voice(..., save=True)) rewrites vieneu/assets/voices_v3_turbo.json inside site-packages, permanently adding your test voice as a 15th "preset". This happened for real during development. Fix: never call either. The store lives at assets/voices/ (library.json plus <id>/ref.wav) and voices are re-registered into the worker on demand with add_voice(..., save=False). Registration costs ~3s and is lost whenever the worker restarts, so track what the live process actually has. As defence in depth the worker drops any preset whose description does not parse into gender · region · style, and listLocalVoices() drops presets colliding with a store id - a contaminated install still yields a clean 14. To repair one by hand, delete the offending key from voices_v3_turbo.json or pip install --force-reinstall vieneu.

Symptom: A storytelling voice reads like a newsreader. Cause: infer() has its own style= argument defaulting to "tu_nhien", which is independent of the voice you picked. Fix: look up each preset's own style (get_preset_voice(name) returns {gender, style, description, ...} as real fields) and pass it into infer(). Only region has to be parsed out of the description string.

Symptom: Timeline drifts against the narration. Cause: estimating duration from character count. Fix: every duration is measured with ffprobe, exactly as on the Gemini path. This is not engine-specific advice - see the video-pipeline skill.

Architecture notes

  • apps/server/python/vieneu_worker.py is a long-lived worker speaking JSON-lines on stdin/stdout. stdout is protocol only; all logging goes to stderr. It must sys.stdout.reconfigure(encoding="utf-8") first thing or Vietnamese corrupts on the Windows cp1252 console.
  • ping must answer without loading the model - it backs the availability check.
  • One worker, requests serialised (the model is not re-entrant), 10-minute idle shutdown because it holds ~1 GB of RAM.
  • Output is already 48 kHz mono, which equals the pipeline's OUT_SAMPLE_RATE, so no resampling. It still goes through the same TRIM_FILTER and gap-insertion path as Gemini in synthScript() - two engines with two joining strategies would drift audibly at the seams.
  • The Perth watermark is left at its default (on). It is inaudible and is the responsible default for cloned speech.

Reference audio guidance for users

3-8 seconds, clean speech, no background music, no noise. Longer is not better - the model only uses the head, so anything past ~8s is discarded. Below 3s the embedding is too weak and the clone stops resembling the speaker.

© notivn, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/offline-voice of notivn/AIEV.

Open the folder on GitHubat commit 1a4c2b0

Compare with similar skills

Offline Voice next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Offline Voice compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Offline Voice this skillnotivn/AIEV126—~2kAutomated safety check: PassMIT
Musictadaspetra/loop2962 repos~827Automated safety check: PassMIT
Edu Math Videowy51ai/edulab1.4k—~2.5kAutomated safety check: NotesApache-2.0
Sound Effectstadaspetra/loop2962 repos~1.1kAutomated safety check: PassMIT
Book Video Factorybytec-ai/book-video-factory322—~1.4kAutomated safety check: NotesNone
Podcastzarazhangrui/personalized-podcast437—~2.3kAutomated safety check: NotesNone

Similar skills

  • Music

    tadaspetra/loop

    Generate music using ElevenLabs Music API. An agent skill from tadaspetra/loop.

    296 GitHub starsUsed in 2 repos~827 tokens
    Media & CreativeAuto-check passed
  • Edu Math Video

    wy51ai/edulab

    A skill your agent uses when asked to make an explainer / walkthrough video (讲解视频、解题视频、例题精讲、微课) for a math problem (数学题, geometry, algebra, functions, motion/行程 problems), from a problem screenshot…

    1.4k GitHub stars~2.5k tokensUpdated 10 days ago
    Media & CreativeAuto-check: notes
  • Sound Effects

    tadaspetra/loop

    Generate sound effects from text descriptions using ElevenLabs.

    296 GitHub starsUsed in 2 repos~1.1k tokens
    Media & CreativeAuto-check passed
  • Book Video Factory

    bytec-ai/book-video-factory

    通用的多账号图书短视频生产工作流。用于用户希望建立图书号项目目录、配置账号级片头/声音/BGM/视觉规范,或只提供一本书后依次完成资料研究、口播稿、分镜、图片、配音、字幕、预览与成片导出。适用于新建工作区、批量管理多个账号、继续已有单书任务和检查生产状态;不绑定特定研究、图片、TTS、转录或视频渲染供应商。

    322 GitHub stars~1.4k tokensUpdated 2 mo ago
    Media & CreativeAuto-check: notes
  • Podcast

    zarazhangrui/personalized-podcast

    Generate a podcast episode from content you provide. An agent skill from zarazhangrui/personalized-podcast.

    437 GitHub stars~2.3k tokensUpdated 6 mo ago
    Media & CreativeAuto-check: notes
  • Book Sales Video

    Kianzzz/book-sales-video

    从书名或飞书多维表格中的成稿文案出发,结合微信读书资料与公开点评创作图书带货/书评短视频,并用豆包 TTS、Pexels、Codex 生图和本机 OpenChatCut 完成配音、配图、双语字幕、音效、动效、BGM、可编辑初稿与按需导出。用户提出“根据一本书做带货视频”“读取飞书文案制作图书视频”“写书评口播并自动剪成抖音视频”“仿参考样式做图书推荐短视频”时使用;仅查书、仅写普通书评或无关剪辑…

    217 GitHub stars~3.2k tokensUpdated 1 mo ago
    Media & CreativeAuto-check passed

More from notivn/AIEV

All 14 skills in this repo
  • Background Music

    notivn/AIEV

    Pick background music from the assets/music/ library and configure auto-ducking (music dips automatically under speech) via meta.json audio.music for the Remotion assembly layer.

    126 GitHub stars~1.1k tokensUpdated yesterday
    Auto-check passed
  • Color Grading

    notivn/AIEV

    Color grading video in the AI Edit Video system - delog/tonemap HDR-HLG-log footage, apply the color preset the user approved in the UI, and the visual verification workflow.

    126 GitHub stars~932 tokensUpdated yesterday
    Auto-check passed
  • Build a Vietnamese vertical TikTok explainer in the "MỔ XẺ PAPER AI" (AI paper dissection) format with HyperFrames (HTML/CSS/GSAP → MP4), Noti.vn style.

    126 GitHub stars~4.8k tokensUpdated yesterday
    Auto-check passed
  • Skill Authoring

    notivn/AIEV

    The standard for writing new skills for the AI Edit Video system - file structure, frontmatter, tone of voice, and how to accumulate production lessons into skills.

    126 GitHub stars~925 tokensUpdated yesterday
    Auto-check passed
  • Noti Tiktok Vn

    notivn/AIEV

    Edit a Vietnamese vertical TikTok video (9:16) with HyperFrames following the Noti.vn/GĐT standard - talking-head + kinetic typography + karaoke captions + zoom/punch-in camera + timestamp-synced…

    126 GitHub stars~5.4k tokensUpdated yesterday
    Auto-check passed
  • Build a Vietnamese landscape 16:9 YouTube video (1920×1080) with HyperFrames (HTML/CSS/GSAP → MP4), keeping the Noti.vn/GĐT branding inherited from noti-tiktok-vn.

    126 GitHub stars~5.6k tokensUpdated yesterday
    Auto-check passed

Works with

Questions about Offline Voice

What does Offline Voice do?

How the on-device narration engine (VieNeu-TTS) and voice cloning work in this system - install tiers, the worker protocol, and the verified traps around torchaudio, voice metadata parsing and…. Offline Voice is an agent skill from notivn/AIEV. How the on-device narration engine (VieNeu-TTS) and voice cloning work in this system - install tiers, the worker protocol, and the verified traps around torchaudio, voice metadata parsing and duration.

When should I use Offline Voice?

Offline Voice fits situations like: tasks that involve Text to speech and voice.

How do I install Offline Voice in Claude Code?

Run `npx skills add notivn/AIEV --skill offline-voice -a claude-code`. Or copy the skill folder (.claude/skills/offline-voice in notivn/AIEV) into .claude/skills/offline-voice in your project. Claude Code loads it when a task matches its description.

How do I install Offline Voice in Codex?

Run `npx skills add notivn/AIEV --skill offline-voice -a codex`. Or copy the skill folder (.claude/skills/offline-voice in notivn/AIEV) into .agents/skills/offline-voice in your project. Codex loads it when a task matches its description.

Can I use Offline Voice in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add notivn/AIEV --skill offline-voice -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/offline-voice, .gemini/skills/offline-voice, .github/skills/offline-voice and .opencode/skills/offline-voice in your project.

What does Offline Voice need to run?

Going by SKILL.md and its folder, Offline Voice needs the command-line tools its instructions call (pip) and credentials named GEMINI_API_KEY. Our summary lists: Python 3; A credential in GEMINI_API_KEY.

Does Offline Voice access the network?

SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Offline Voice safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Offline Voice use?

Offline Voice is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Offline Voice use?

About 2k tokens (SKILL.md is roughly 8.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Offline Voice?

Skills that share tags, products or a category with Offline Voice: Music (tadaspetra/loop, 296 stars), Edu Math Video (wy51ai/edulab, 1.4k stars), Sound Effects (tadaspetra/loop, 296 stars) and Book Video Factory (bytec-ai/book-video-factory, 322 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Offline Voice?

notivn (a GitHub organization) maintains it in notivn/AIEV, which has 126 GitHub stars. The repository holds 14 skills in this directory. The repository was last updated on October 8, 2026.

Source: notivn/AIEV on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.