Agent skill

Voice

by guaardvark in guaardvark/guaardvark

Narration and text-to-speech on the user's machine through Guaardvark's Audio Foundry (Chatterbox, Kokoro, Piper) and consent-gated voice cloning from a reference clip.

MITAuto-check passedMedia & Creative

Install Voice

skills CLI
$ npx skills add guaardvark/guaardvark --skill voice -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install guaardvark/guaardvark voice --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/guaardvark/guaardvark.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/voice .claude/skills/voice && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
voice
GitHub stars
257
Token cost
~798 tokens
SKILL.md length
289 words
Files
1
Skills in repo
15
Repo updated
First seen
Licence
MIT

At a glance

Narration and text-to-speech on the user's machine through Guaardvark's Audio Foundry (Chatterbox, Kokoro, Piper) and consent-gated voice cloning from a reference clip.

  • Works in 3 steps: The reference must go through the upload… → Generate with "backend": "chatterbox",… → Before uploading, ask whether the voice…
  • The user wants a voiceover
  • SKILL.md covers Which engine, Speak a line or a script, Clone a voice (consent-gated) and Rules
  • Calls curl

What it does

Voice is an agent skill from guaardvark/guaardvark. Narration and text-to-speech on the user's machine through Guaardvark's Audio Foundry (Chatterbox, Kokoro, Piper) and consent-gated voice cloning from a reference clip. Use when the user wants a voiceover, narration of a script, a spoken line, or "make it sound like this voice".

Its SKILL.md is about 800 tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Media & Creative, covering Text to speech and voice. It works with Model Context Protocol. The repository describes itself as: Self-hosted AI studio on your own GPU: chat with your files (RAG), screen and browser agents, coding swarms, LoRA training, an MCP server for Claude Code and Cursor, and local… The licence is MIT.

When your agent uses it

  • The user wants a voiceover
  • Narration of a script
  • Make it sound like this voice

Example prompts

  • “s machine through Guaardvark”
  • “make it sound like this voice”
  • “/voice”

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. The reference must go through the upload route; that is what records consent. Arbitrary
  2. Generate with "backend": "chatterbox", "reference_clip_path": "".
  3. Before uploading, ask whether the voice belongs to the user or someone who consented. Do not

What it can do on your machine

Read from SKILL.md and the folder at commit 2a33110. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • curl

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use curl, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Voice loads about 798 tokens when it runs. Until then it costs about 71 tokens; SKILL.md has 289 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~71
When it runs · the whole SKILL.md, loaded when a task matches
~798

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from guaardvark/guaardvark at commit 2a33110, republished under its MIT licence (© guaardvark). 289 words, ~798 tokens.

Download SKILL.mdSave it as .claude/skills/voice/SKILL.md (or your agent's skills folder).
name
voice
description
Narration and text-to-speech on the user's machine through Guaardvark's Audio Foundry (Chatterbox, Kokoro, Piper) and consent-gated voice cloning from a reference clip. Use when the user wants a voiceover, narration of a script, a spoken line, or "make it sound like this voice".

Voice with Guaardvark

Read setup first. Expressive voices need the audio_foundry plugin running; Piper works without it. B=${GUAARDVARK_URL:-http://localhost:5000}.

Which engine

engineroutewhen
ChatterboxAudio Foundry backend: "chatterbox"expressive, emotion presets, cloning
KokoroAudio Foundry backend: "kokoro"fast, clean, 10+ built-in voices (af_heart default)
Piper/api/voice/text-to-speechoffline fallback, no GPU

GET $B/api/audio-foundry/voices lists what is installed. GET $B/api/voice/voices lists Piper voices.

Speak a line or a script

Over MCP, call generate_speech (text up to 3000 characters, optional voice such as af_heart, engine auto | kokoro | chatterbox); it waits and returns the file with a download link. Cloning is not offered over MCP. Without MCP, or for the Chatterbox knobs, use REST:

bash
curl -s -X POST $B/api/audio-foundry/generate/voice -H 'Content-Type: application/json' -d '{
  "text": "The line to speak.",
  "backend": "auto",            # auto | chatterbox | kokoro
  "voice_id": "af_heart",       # Kokoro voice, or omit
  "emotion": "calm",            # Chatterbox preset, or omit
  "exaggeration": 0.5, "cfg_weight": 0.5, "temperature": 0.8,   # Chatterbox knobs, optional
  "seed": 7, "output_format": "wav", "async": true
}'
  • Short text returns the file directly (path, document_id). With "async": true or long text you get 202 {"job_id"}: poll GET $B/api/audio-foundry/jobs/<job_id> until status is done; the result has path and document_id. Cancel: POST .../jobs/<job_id>/cancel.
  • Multi-section narration with pauses: POST $B/api/voice/narrate {"script": "...", "engine": "kokoro", "voice": "...", "pause_between_sections": 0.6, "output_format": "wav"}.
  • Piper only: POST $B/api/voice/text-to-speech {"text", "voice": "libritts"} returns audio_url.
  1. The reference must go through the upload route; that is what records consent. Arbitrary file paths are refused with 403.
    bash
    curl -s -X POST $B/api/audio-foundry/voice-clips/upload -F file=@/abs/path/ref.wav -F name="Narrator sample"
    The response gives the stored path. GET $B/api/audio-foundry/voice-clips lists clips.
  2. Generate with "backend": "chatterbox", "reference_clip_path": "<that path>".
  3. Before uploading, ask whether the voice belongs to the user or someone who consented. Do not clone a public figure or anyone who has not agreed. Refuse politely if unclear.

Rules

  • 10 to 20 seconds of clean speech is enough for a clone; more is not better.
  • Say which engine ran (the response reports it); auto falls back to Kokoro on a Chatterbox error.
  • Audio files are local under data/outputs/; they also appear in the Audio library page.

© guaardvark, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/voice of guaardvark/guaardvark.

Open the folder on GitHubat commit 2a33110

Compare with similar skills

Voice next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Voice compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Voice this skillguaardvark/guaardvark257—~798Automated safety check: PassMIT
Audio Duckingsonilo-ai/skills115—~1.8kAutomated safety check: NotesMIT
Auto Dubbingsonilo-ai/skills115—~4.5kAutomated safety check: NotesMIT
Proofreadsonilo-ai/skills115—~4.4kAutomated safety check: NotesMIT
Task Recoverysonilo-ai/skills115—~2.1kAutomated safety check: NotesMIT
Generate Narration AudioArcReel/ArcReel5.4k—~524Automated safety check: PassAGPL-3.0

Similar skills

  • Audio Ducking

    sonilo-ai/skills

    Duck a music bed under a voice track using Sonilo — automatically lowers the music wherever the voice speaks and lifts it back in the gaps.

    115 GitHub stars~1.8k tokensUpdated today
    Media & CreativeAuto-check: notes
  • Auto Dubbing

    sonilo-ai/skills

    Dub a video into one or more other languages using Sonilo, translating and re-voicing the speech into a new video per language.

    115 GitHub stars~4.5k tokensUpdated today
    Media & CreativeAuto-check: notes
  • Proofread

    sonilo-ai/skills

    Transcribe a video with Sonilo and translate the transcript into editable .srt files — one per target language, plus the detected source language — so the wording can be read and corrected before…

    115 GitHub stars~4.4k tokensUpdated today
    Media & CreativeAuto-check: notes
  • Task Recovery

    sonilo-ai/skills

    Recover the result of a timed-out Sonilo generation call using its taskid.

    115 GitHub stars~2.1k tokensUpdated today
    Media & CreativeAuto-check: notes
  • 为旁白/解说剧本逐分镜生成旁白配音(TTS)。当 TTS 项目首轮自动剪辑需要补齐缺失配音、用户要求生成或重新生成某个分镜或某集旁白配音,或批量配音中断需要补齐时使用。

    5.4k GitHub stars~524 tokensUpdated today
    Media & CreativeAuto-check passed
  • Scenario Audio

    scenario-labs/skills

    A skill your agent uses when generating or handling audio on Scenario via MCP.

    931 GitHub stars~3k tokensUpdated yesterday
    Media & CreativeAuto-check passed

More from guaardvark/guaardvark

All 15 skills in this repo
  • Cast

    guaardvark/guaardvark

    Build consistent characters, environments and props in Guaardvark's Cast Library and train LoRAs for them locally (reference photos → vision bible → sample plan → approved samples → training).

    257 GitHub stars~673 tokensUpdated today
    Auto-check passed
  • Image

    guaardvark/guaardvark

    Generate or edit images on the user's own GPU through Guaardvark: single images, instruction edits, background cut-outs, inpaint and outpaint, consistent characters from the Cast Library, and batch…

    257 GitHub stars~1.8k tokensUpdated today
    Auto-check passed
  • Music

    guaardvark/guaardvark

    Generate full songs with vocals or instrumentals (ACE-Step) and sound effects or ambience (Stable Audio Open) on the user's GPU through Guaardvark's Audio Foundry.

    257 GitHub stars~710 tokensUpdated today
    Auto-check passed
  • Ops

    guaardvark/guaardvark

    Operate a running Guaardvark: GPU and VRAM state, plugin start/stop, logs, Celery tasks, the Interconnector sync to other machines, overnight RAG autoresearch, and infographics.

    257 GitHub stars~779 tokensUpdated today
    Auto-check passed
  • Setup

    guaardvark/guaardvark

    Connect this agent to a running Guaardvark (self-hosted AI studio) and check what it can do right now.

    257 GitHub stars~1.2k tokensUpdated today
    Auto-check passed
  • Swarm

    guaardvark/guaardvark

    Launch and watch Guaardvark's Swarm Orchestrator: parallel coding agents, each in its own git worktree, working a markdown plan and merging back deterministically.

    257 GitHub stars~746 tokensUpdated today
    Auto-check passed

Questions about Voice

What does Voice do?

Narration and text-to-speech on the user's machine through Guaardvark's Audio Foundry (Chatterbox, Kokoro, Piper) and consent-gated voice cloning from a reference clip. Voice is an agent skill from guaardvark/guaardvark. Narration and text-to-speech on the user's machine through Guaardvark's Audio Foundry (Chatterbox, Kokoro, Piper) and consent-gated voice cloning from a reference clip.

When should I use Voice?

Voice fits situations like: the user wants a voiceover; narration of a script; make it sound like this voice.

How do I install Voice in Claude Code?

Run `npx skills add guaardvark/guaardvark --skill voice -a claude-code`. Or copy the skill folder (.agents/skills/voice in guaardvark/guaardvark) into .claude/skills/voice in your project. Claude Code loads it when a task matches its description.

How do I install Voice in Codex?

Run `npx skills add guaardvark/guaardvark --skill voice -a codex`. Or copy the skill folder (.agents/skills/voice in guaardvark/guaardvark) into .agents/skills/voice in your project. Codex loads it when a task matches its description.

Can I use Voice in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add guaardvark/guaardvark --skill voice -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/voice, .gemini/skills/voice, .github/skills/voice and .opencode/skills/voice in your project.

What does Voice need to run?

Going by SKILL.md and its folder, Voice needs the command-line tools its instructions call (curl).

Does Voice access the network?

SKILL.md contains no URLs. Its commands use curl, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Voice safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Voice use?

Voice is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Voice use?

About 798 tokens (SKILL.md is roughly 3.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Voice?

Skills that share tags, products or a category with Voice: Audio Ducking (sonilo-ai/skills, 115 stars), Auto Dubbing (sonilo-ai/skills, 115 stars), Proofread (sonilo-ai/skills, 115 stars) and Task Recovery (sonilo-ai/skills, 115 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Voice?

guaardvark (a GitHub user) maintains it in guaardvark/guaardvark, which has 257 GitHub stars. The repository holds 15 skills in this directory. The repository was last updated on October 9, 2026.

Source: guaardvark/guaardvark on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.