Vox Director
Alisa0808/vox-director
Turn ONE topic into a finished Vox-style paper-collage explainer / ad video, end to end on the Atlas Cloud API + local ffmpeg — script, collage keyframes, motion, voice-over, music, captions, all…
Convert text, markdown, or a summary produced by another skill into a listenable MP3 using local CPU-only neural text-to-speech.
$ npx skills add github/awesome-copilot --skill speak-summary -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install github/awesome-copilot speak-summary --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/github/awesome-copilot.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/speak-summary .claude/skills/speak-summary && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "speak-summary" agent skill from https://github.com/github/awesome-copilot/tree/main/skills/speak-summary into .claude/skills/speak-summary/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "speak-summary", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/github/awesome-copilot/tree/main/skills/speak-summaryType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add github/awesome-copilot --skill speak-summary -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install github/awesome-copilot speak-summary --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/github/awesome-copilot.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/speak-summary .agents/skills/speak-summary && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "speak-summary" agent skill from https://github.com/github/awesome-copilot/tree/main/skills/speak-summary into .agents/skills/speak-summary/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "speak-summary", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add github/awesome-copilot --skill speak-summary -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install github/awesome-copilot speak-summary --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/github/awesome-copilot.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/speak-summary .cursor/skills/speak-summary && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "speak-summary" agent skill from https://github.com/github/awesome-copilot/tree/main/skills/speak-summary into .cursor/skills/speak-summary/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "speak-summary", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/github/awesome-copilot.git --path skills/speak-summary--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add github/awesome-copilot --skill speak-summary -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install github/awesome-copilot speak-summary --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/github/awesome-copilot.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/speak-summary .gemini/skills/speak-summary && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "speak-summary" agent skill from https://github.com/github/awesome-copilot/tree/main/skills/speak-summary into .gemini/skills/speak-summary/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "speak-summary", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install github/awesome-copilot speak-summaryInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add github/awesome-copilot --skill speak-summary -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/github/awesome-copilot.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/speak-summary .github/skills/speak-summary && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "speak-summary" agent skill from https://github.com/github/awesome-copilot/tree/main/skills/speak-summary into .github/skills/speak-summary/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "speak-summary", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add github/awesome-copilot --skill speak-summary -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install github/awesome-copilot speak-summary --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/github/awesome-copilot.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/speak-summary .opencode/skills/speak-summary && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "speak-summary" agent skill from https://github.com/github/awesome-copilot/tree/main/skills/speak-summary into .opencode/skills/speak-summary/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "speak-summary", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
speak-summaryConvert text, markdown, or a summary produced by another skill into a listenable MP3 using local CPU-only neural text-to-speech.
Speak Summary is an agent skill from github/awesome-copilot, published by the product's own GitHub organization. Convert text, markdown, or a summary produced by another skill into a listenable MP3 using local CPU-only neural text-to-speech. Rewrites written prose for the ear before synthesising. Use when the user asks to "read this out", "turn this into audio", "make an MP3", "I want to listen to this", "podcast version", or wants a spoken digest for a commute or breakfast.
Its SKILL.md is about 1.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including scripts (for example `scripts/tts.sh`).
It sits in Media & Creative, covering Text to speech and voice. It works with Python and FFmpeg. The repository describes itself as: Community-contributed instructions, agents, skills, and configurations to help you make the most of GitHub Copilot. The licence is MIT.
Read from SKILL.md and the folder at commit 727ff2e. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 1 file in scripts/ (Shell), which the agent can run.
Shell commands in SKILL.md call:
brewpipapt-getFrom the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
github.comFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Speak Summary loads about 1.6k tokens when it runs. Until then it costs about 95 tokens; SKILL.md has 857 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from github/awesome-copilot at commit 727ff2e, republished under its MIT licence (© github). 857 words, ~1,591 tokens.
.claude/skills/speak-summary/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.Turn written text into audio someone will actually want to listen to.
This skill is deliberately a terminal step in a chain. Another skill (or you)
produces the text; this one makes it listenable. It pairs naturally with
roundup, daily-prep, meeting-minutes, or any summarisation work.
Everything runs locally on CPU. No text is sent to a cloud speech service, which matters when the content is confidential, and it means the skill works in a headless cloud agent or CI container just as well as on a laptop.
The synthesis engine is Kyutai pocket-tts,
a small neural TTS model designed to run on CPUs.
The bundled script installs it automatically into a cached virtualenv on first use, so usually you need do nothing. To install it explicitly:
pip install pocket-tts # any platform
brew install pocket-tts # macOS, if preferredpocket-tts requires Python >=3.10 and <3.15. The script searches for a
compatible interpreter rather than assuming python3 is one — worth knowing if
you are on a very new Python, where installation would otherwise fail.
You also need an encoder. ffmpeg is strongly preferred (brew install ffmpeg
or apt-get install -y ffmpeg); on macOS the script falls back to the built-in
afconvert and emits .m4a instead of .mp3.
The first run downloads the model (~1GB) from Hugging Face. After that it is fully offline and synthesises roughly 6x faster than real-time.
Do not feed written text straight into the synthesiser. Prose that reads well on screen is tiring to listen to. Rewriting it first is what separates a useful audio digest from an unlistenable one.
Produce a spoken script that:
Write this spoken script to its own .txt file. Keep the original written
version with its links intact — the audio is a companion to it, not a
replacement. The user will want to click through later.
./scripts/tts.sh <input.txt> <output.mp3> [voice.safetensors]The script strips any residual markdown, splits the text on sentence boundaries into ~600 character chunks (quality degrades on long single inputs), synthesises each chunk, and concatenates the result into a mono MP3 at 96kbps — small enough to sync to a phone, good enough for speech.
Environment overrides:
| Variable | Purpose |
|---|---|
SPEAK_TTS_BIN | Path to a specific pocket-tts binary; skips all auto-detection. |
SPEAK_TTS_HOME | Where to create/find the cached virtualenv. Default ~/.cache/speak-summary/venv. |
The default English voice is alba. To use a different one, pocket-tts
supports voice cloning from a short clean audio sample:
pocket-tts export-voice --helpPass the resulting .safetensors file as the third argument to the script.
Only clone a voice you have the rights to use. Do not clone a real person's voice — colleague, customer, or public figure — without their explicit consent.
~/Music/Briefings/ unless the user says otherwise; it is easy to point a phone or podcast app at.<subject>-<YYYY-MM-DD>.mp3.afplay <path> on macOS, ffplay -nodisp -autoexit <path> elsewhere.Aim for 4–6 minutes for a routine digest, which is roughly 600–900 spoken words at a natural pace. If the source would run past about 10 minutes, say so and offer either a tighter edit or a split into multiple files — attention drops off sharply beyond that for informational audio.
The natural pattern is gather → summarise → speak:
roundup → speak-summary — a spoken version of the status briefing.daily-prep → speak-summary — tomorrow's schedule, listened to tonight.meeting-minutes → speak-summary — catch up on a meeting you missed.When invoked as part of a chain, do not re-summarise. The upstream skill owns what to say; this skill owns how it sounds. Take its output, rewrite it for the ear, and synthesise.
To run unattended (a briefing waiting before breakfast), schedule the upstream skill with a workflow and have it finish by calling this one.
Audio cuts off mid-sentence. A chunk exceeded the model's comfortable length. Shorten the sentences in the spoken script.
Words mispronounced. Spell them phonetically in the input — "Kubernetes" as "koo-ber-net-eez". This is a normal part of preparing a spoken script.
First run is slow. That is the one-off model download. Later runs start in about a second.
pocket-tts not found after install. The virtualenv may be stale, or your
python3 may be outside the supported 3.10–3.14 range. Delete
~/.cache/speak-summary/venv and re-run, or point SPEAK_TTS_BIN at a known binary.
© github, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 1 other file (scripts) in skills/speak-summary of github/awesome-copilot.
Open the folder on GitHubat commit 727ff2e
Speak Summary next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Speak Summary this skillgithub/awesome-copilot | 40k | — | ~1.6k | Automated safety check: Pass | MIT | |
| Vox DirectorAlisa0808/vox-director | 2.2k | — | ~5.6k | Automated safety check: Pass | MIT | |
| Whiteboard Videognipbao/codex-whiteboard-video-skill | 323 | — | ~7.2k | Automated safety check: Notes | MIT | |
| Muapi DirectorAnil-matcha/vox-ai-motion-graphics-generator | 240 | — | ~679 | Automated safety check: Pass | None | |
| Webcode Local Windows Tts Installershuyu-labs/WebCode | 278 | — | ~787 | Automated safety check: Pass | Custom licence | |
| Ermdougcalobrisi/erm | 112 | — | ~1.3k | Automated safety check: Notes | MIT |
Alisa0808/vox-director
Turn ONE topic into a finished Vox-style paper-collage explainer / ad video, end to end on the Atlas Cloud API + local ffmpeg — script, collage keyframes, motion, voice-over, music, captions, all…
gnipbao/codex-whiteboard-video-skill
Generate smooth hand-drawn whiteboard and story videos directly inside Codex from text scripts, GPT Image 2 color storyboards, scene plans, SVGs, line art, or local images.
Anil-matcha/vox-ai-motion-graphics-generator
Turn ONE topic, talking-head video, or photo into a finished Vox-style paper-collage explainer / ad video on the MuAPI platform (api.muapi.ai) + local ffmpeg — script, collage keyframes, motion…
shuyu-labs/WebCode
A skill your agent uses when building a local Windows WebCode installer from this repo for machine testing, especially when the package must bundle the Kokoro or sherpa-onnx Reply TTS service, model…
dougcalobrisi/erm
Install and run erm, the local CLI that removes filler words / disfluencies (um, uh, er, erm, ah, hmm, mhm, mm, uh-huh and elongations) from spoken-audio recordings.
dtsola/xiaoyaosearch
Turns a topic into a 4K horizontal video podcast through research, scripting, text-to-speech, Remotion rendering and background music, and can learn styles from references.
github/awesome-copilot
Maps an unfamiliar codebase into seven evidence-backed documents in docs/codebase/, using a scan script and templates, for onboarding or architecture write-ups.
github/awesome-copilot
Designs Azure infrastructure from a natural-language description, or diagrams an existing resource group, then refines the design through conversation and deploys it with Bicep.
github/awesome-copilot
Generates, edits and validates draw.io files with correct mxGraph XML, covering flowcharts, architecture, sequence, ER and UML class diagrams.
github/awesome-copilot
Cleans raw credit data and screens variables before loan modeling, dropping unstable, noisy or redundant features and writing an Excel report of every step.
github/awesome-copilot
Builds a warm, browser-based daily focus board the user updates by talking to their agent, with Eisenhower priorities, a brain-dump box and kind not-today carryover.
github/awesome-copilot
End-to-end skill for building, testing, linting, versioning, and publishing a production-grade Python library to PyPI.
Categories
Convert text, markdown, or a summary produced by another skill into a listenable MP3 using local CPU-only neural text-to-speech. Speak Summary is an agent skill from github/awesome-copilot, published by the product's own GitHub organization. Convert text, markdown, or a summary produced by another skill into a listenable MP3 using local CPU-only neural text-to-speech.
Speak Summary fits situations like: the user asks to read this out; turn this into audio; I want to listen to this; podcast version.
Run `npx skills add github/awesome-copilot --skill speak-summary -a claude-code`. Or copy the skill folder (skills/speak-summary in github/awesome-copilot) into .claude/skills/speak-summary in your project. Claude Code loads it when a task matches its description.
Run `npx skills add github/awesome-copilot --skill speak-summary -a codex`. Or copy the skill folder (skills/speak-summary in github/awesome-copilot) into .agents/skills/speak-summary in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add github/awesome-copilot --skill speak-summary -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/speak-summary, .gemini/skills/speak-summary, .github/skills/speak-summary and .opencode/skills/speak-summary in your project.
Going by SKILL.md and its folder, Speak Summary needs a shell for the scripts in its folder and the command-line tools its instructions call (brew, pip and apt-get). Our summary lists: Python 3; A Bash shell.
SKILL.md names 1 domain. As links in the text: github.com. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Speak Summary is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.6k tokens (SKILL.md is roughly 6.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Speak Summary: Vox Director (Alisa0808/vox-director, 2.2k stars), Whiteboard Video (gnipbao/codex-whiteboard-video-skill, 323 stars), Muapi Director (Anil-matcha/vox-ai-motion-graphics-generator, 240 stars) and Webcode Local Windows Tts Installer (shuyu-labs/WebCode, 278 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
github (a GitHub organization, an official publisher) maintains it in github/awesome-copilot, which has 39,748 GitHub stars. The repository holds 417 skills in this directory. The repository was last updated on October 7, 2026.
Source: github/awesome-copilot on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.