Agent skill

Voice Compose

by Utopai-Research in Utopai-Research/pai-code

Designs and attaches voice samples or final narration/line audio on the filmmaking canvas via the local generatevoice.js CLI.

Custom licenceAuto-check passedMedia & Creative

Install Voice Compose

skills CLI
$ npx skills add Utopai-Research/pai-code --skill voice-compose -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Utopai-Research/pai-code voice-compose --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Utopai-Research/pai-code.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/voice-compose .claude/skills/voice-compose && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
voice-compose
GitHub stars
357
Token cost
~631 tokens
SKILL.md length
239 words
Files
1
Skills in repo
6
Repo updated
First seen
Licence
Custom licence

At a glance

Designs and attaches voice samples or final narration/line audio on the filmmaking canvas via the local generatevoice.js CLI.

  • Works in 2 steps: Character voice sample → Narrator / VO voice sample or final line…
  • Asks to give a character a voice
  • Calls node
  • Preview how a character sounds

What it does

Voice Compose is an agent skill from Utopai-Research/pai-code. Designs and attaches voice samples or final narration/line audio on the filmmaking canvas via the local generatevoice.js CLI. Use before calling generatevoice.js; when the user asks to give a character a voice, preview how a character sounds, create reusable timbre anchors for every speaking character or VO/narration, or create exact narration/VO/final line audio.

Its SKILL.md is about 630 tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Media & Creative, covering Text to speech and voice. The repository describes itself as: Local AI filmmaking studio — skills, canvas, timeline — driven from your coding agent.

When your agent uses it

  • Asks to give a character a voice
  • Preview how a character sounds
  • Create reusable timbre anchors for every speaking character
  • Create exact narration/VO/final line audio

Example prompts

  • “Use the voice-compose skill to design and attaches voice samples or final narration/line audio on the filmmaking canvas via the local…”
  • “/voice-compose”

Workflow steps

2 steps, taken from the step headings in SKILL.md.

  1. Character voice sample
  2. Narrator / VO voice sample or final line audio

What it can do on your machine

Read from SKILL.md and the folder at commit 825bca2. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • node

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Voice Compose loads about 631 tokens when it runs. Until then it costs about 96 tokens; SKILL.md has 239 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~96
When it runs · the whole SKILL.md, loaded when a task matches
~631

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

Its licence (Custom licence) doesn't allow us to republish the file, so here is its outline and opening line. It has 239 words (~631 tokens).

“Default to one short reusable timbre sample per speaking character and one VO/narrator sample when narration exists. video-compose keeps actual shot dialogue/VO in the video prompt. Treat audio_result.data.text as downstream speech only for approved final narration/line reads.”

— opening of SKILL.md by Utopai-Research, Custom licence
name
voice-compose

Read the full SKILL.md on GitHub

Files

Just SKILL.md in skills/voice-compose of Utopai-Research/pai-code.

Open the folder on GitHubat commit 825bca2

Compare with similar skills

Voice Compose next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Voice Compose compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Voice Compose this skillUtopai-Research/pai-code357—~631Automated safety check: PassCustom licence
MoneyPrinterTurbo Video Generatorharry0703/MoneyPrinterTurbo129k—~2.1kAutomated safety check: WarnMIT
HyperFrames Media Useheygen-com/hyperframes60k—~2.4kAutomated safety check: PassApache-2.0
Openspec OnboardSAP/e-mobility-charging-stations-simulator22725 repos~3.5kAutomated safety check: PassMIT
Blog AudioAgriciDaniel/claude-blog2.3k1 repos~2.2kAutomated safety check: NotesMIT
Musictadaspetra/loop2962 repos~827Automated safety check: PassMIT

Similar skills

  • MoneyPrinterTurbo Video Generator

    harry0703/MoneyPrinterTurbo

    Installs and runs MoneyPrinterTurbo to turn a topic or script into a finished short video with voice-over, subtitles, stock footage and music.

    129k GitHub stars~2.1k tokensUpdated yesterday
    Media & CreativeAuto-check: warnings
  • HyperFrames Media Use

    heygen-com/hyperframes

    Finds, generates and edits media for HyperFrames video projects: music, sound effects, images, icons, logos, voiceovers, captions and color grades.

    60k GitHub stars~2.4k tokensUpdated today
    Media & CreativeAuto-check passed
  • Openspec Onboard

    SAP/e-mobility-charging-stations-simulator

    Official

    Guided onboarding for OpenSpec - walk through a complete workflow cycle with narration and real codebase work.

    227 GitHub starsUsed in 25 repos~3.5k tokens
    Media & CreativeAuto-check passed
  • Blog Audio

    AgriciDaniel/claude-blog

    Generate audio narration of blog posts using Google Gemini TTS.

    2.3k GitHub starsUsed in 1 repo~2.2k tokens
    Media & CreativeAuto-check: notes
  • Music

    tadaspetra/loop

    Generate music using ElevenLabs Music API. An agent skill from tadaspetra/loop.

    296 GitHub starsUsed in 2 repos~827 tokens
    Media & CreativeAuto-check passed
  • Create News Video

    hoquanghai/Auto-Create-Video

    Tạo video tin tức ngắn 9:16 (~60s) từ URL bài báo hoặc file .txt tiếng Việt.

    319 GitHub starsUsed in 1 repo~3.7k tokens
    Media & CreativeAuto-check passed

More from Utopai-Research/pai-code

  • Video Compose

    Utopai-Research/pai-code

    Generates and prompts video clips on the filmmaking canvas. An agent skill from Utopai-Research/pai-code.

    357 GitHub stars~4.1k tokensUpdated 11 days ago
    Auto-check passed
  • Groups Compose

    Utopai-Research/pai-code

    Designs and maintains semantic groupings and readable layouts on the filmmaking canvas — scenes, character-reference sets, act beats, and other titled visual frames.

    357 GitHub stars~1.5k tokensUpdated 11 days ago
    Auto-check passed
  • Script Compose

    Utopai-Research/pai-code

    Handles explicit screenplay/story work on the filmmaking canvas.

    357 GitHub stars~2.1k tokensUpdated 11 days ago
    Auto-check passed
  • Story To Video Workflow

    Utopai-Research/pai-code

    Orchestrates story, script, screenplay, concept, product promo, and multi-shot idea work into finished video.

    357 GitHub stars~2k tokensUpdated 11 days ago
    Auto-check passed
  • Image Compose

    Utopai-Research/pai-code

    Generates/edits filmmaking canvas images via generateimage.js and generateimagepro.js.

    357 GitHub stars~3k tokensUpdated 11 days ago
    Auto-check passed

Questions about Voice Compose

What does Voice Compose do?

Designs and attaches voice samples or final narration/line audio on the filmmaking canvas via the local generatevoice.js CLI. Voice Compose is an agent skill from Utopai-Research/pai-code.js CLI.

When should I use Voice Compose?

Voice Compose fits situations like: asks to give a character a voice; preview how a character sounds; create reusable timbre anchors for every speaking character; create exact narration/VO/final line audio.

How do I install Voice Compose in Claude Code?

Run `npx skills add Utopai-Research/pai-code --skill voice-compose -a claude-code`. Or copy the skill folder (skills/voice-compose in Utopai-Research/pai-code) into .claude/skills/voice-compose in your project. Claude Code loads it when a task matches its description.

How do I install Voice Compose in Codex?

Run `npx skills add Utopai-Research/pai-code --skill voice-compose -a codex`. Or copy the skill folder (skills/voice-compose in Utopai-Research/pai-code) into .agents/skills/voice-compose in your project. Codex loads it when a task matches its description.

Can I use Voice Compose in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Utopai-Research/pai-code --skill voice-compose -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/voice-compose, .gemini/skills/voice-compose, .github/skills/voice-compose and .opencode/skills/voice-compose in your project.

What does Voice Compose need to run?

Going by SKILL.md and its folder, Voice Compose needs the command-line tools its instructions call (node).

Does Voice Compose access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Voice Compose safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Voice Compose use?

Voice Compose has a licence file (the repository's licence) that doesn't match a standard licence. Read it on GitHub before reusing the skill.

How many tokens does Voice Compose use?

About 631 tokens (SKILL.md is roughly 2.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Voice Compose?

Skills that share tags, products or a category with Voice Compose: MoneyPrinterTurbo Video Generator (harry0703/MoneyPrinterTurbo, 129k stars), HyperFrames Media Use (heygen-com/hyperframes, 60k stars), Openspec Onboard (SAP/e-mobility-charging-stations-simulator, 227 stars) and Blog Audio (AgriciDaniel/claude-blog, 2.3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Voice Compose?

Utopai-Research (a GitHub organization) maintains it in Utopai-Research/pai-code, which has 357 GitHub stars. The repository holds 6 skills in this directory. The repository was last updated on September 29, 2026.

Source: Utopai-Research/pai-code on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.