Official agent skill

Speak Summary

by github in github/awesome-copilot

Convert text, markdown, or a summary produced by another skill into a listenable MP3 using local CPU-only neural text-to-speech.

OfficialMITAuto-check passedMedia & Creative

Install Speak Summary

skills CLI
$ npx skills add github/awesome-copilot --skill speak-summary -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install github/awesome-copilot speak-summary --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/github/awesome-copilot.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/speak-summary .claude/skills/speak-summary && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
speak-summary
GitHub stars
40k
Token cost
~1.6k tokens
SKILL.md length
857 words
Files
2 (incl. scripts)
Skills in repo
417
Repo updated
First seen
Licence
MIT

At a glance

Convert text, markdown, or a summary produced by another skill into a listenable MP3 using local CPU-only neural text-to-speech.

  • The user asks to read this out
  • SKILL.md covers Prerequisites, The important step: rewrite…, Synthesise and Voices, plus 4 more sections
  • Runs Shell scripts from its folder; calls brew, pip and apt-get
  • Turn this into audio

What it does

Speak Summary is an agent skill from github/awesome-copilot, published by the product's own GitHub organization. Convert text, markdown, or a summary produced by another skill into a listenable MP3 using local CPU-only neural text-to-speech. Rewrites written prose for the ear before synthesising. Use when the user asks to "read this out", "turn this into audio", "make an MP3", "I want to listen to this", "podcast version", or wants a spoken digest for a commute or breakfast.

Its SKILL.md is about 1.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including scripts (for example `scripts/tts.sh`).

It sits in Media & Creative, covering Text to speech and voice. It works with Python and FFmpeg. The repository describes itself as: Community-contributed instructions, agents, skills, and configurations to help you make the most of GitHub Copilot. The licence is MIT.

When your agent uses it

  • The user asks to read this out
  • Turn this into audio
  • I want to listen to this
  • Podcast version

Example prompts

  • “read this out”
  • “turn this into audio”
  • “make an MP3”
  • “/speak-summary”

Requirements

  • Python 3
  • A Bash shell

What it can do on your machine

Read from SKILL.md and the folder at commit 727ff2e. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Shell), which the agent can run.

    Shell commands in SKILL.md call:

    • brew
    • pip
    • apt-get

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Speak Summary loads about 1.6k tokens when it runs. Until then it costs about 95 tokens; SKILL.md has 857 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~95
When it runs · the whole SKILL.md, loaded when a task matches
~1.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from github/awesome-copilot at commit 727ff2e, republished under its MIT licence (© github). 857 words, ~1,591 tokens.

Download SKILL.mdSave it as .claude/skills/speak-summary/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
speak-summary
description
Convert text, markdown, or a summary produced by another skill into a listenable MP3 using local CPU-only neural text-to-speech. Rewrites written prose for the ear before synthesising. Use when the user asks to "read this out", "turn this into audio", "make an MP3", "I want to listen to this", "podcast version", or wants a spoken digest for a commute or breakfast.

Speak Summary

Turn written text into audio someone will actually want to listen to.

This skill is deliberately a terminal step in a chain. Another skill (or you) produces the text; this one makes it listenable. It pairs naturally with roundup, daily-prep, meeting-minutes, or any summarisation work.

Everything runs locally on CPU. No text is sent to a cloud speech service, which matters when the content is confidential, and it means the skill works in a headless cloud agent or CI container just as well as on a laptop.

Prerequisites

The synthesis engine is Kyutai pocket-tts, a small neural TTS model designed to run on CPUs.

The bundled script installs it automatically into a cached virtualenv on first use, so usually you need do nothing. To install it explicitly:

bash
pip install pocket-tts          # any platform
brew install pocket-tts         # macOS, if preferred

pocket-tts requires Python >=3.10 and <3.15. The script searches for a compatible interpreter rather than assuming python3 is one — worth knowing if you are on a very new Python, where installation would otherwise fail.

You also need an encoder. ffmpeg is strongly preferred (brew install ffmpeg or apt-get install -y ffmpeg); on macOS the script falls back to the built-in afconvert and emits .m4a instead of .mp3.

The first run downloads the model (~1GB) from Hugging Face. After that it is fully offline and synthesises roughly 6x faster than real-time.

The important step: rewrite for the ear

Do not feed written text straight into the synthesiser. Prose that reads well on screen is tiring to listen to. Rewriting it first is what separates a useful audio digest from an unlistenable one.

Produce a spoken script that:

  • Opens with orientation. What this is, what it covers, roughly how long it runs.
  • Replaces bullets with connective prose. "First… The bigger one is… Finally…" — a listener has no visual structure to lean on, so carry it in the language.
  • Expands abbreviations on first use. "PR" becomes "pull request", "CI" becomes "continuous integration". Acronyms that read fine are noise when spoken.
  • Speaks dates and numbers naturally. "the twentieth of August", not "2026-08-20". "About three thousand", not "2,847".
  • Never reads URLs aloud. Say "linked in the written version" instead.
  • Uses short sentences. Split anything past roughly 25 words.
  • Signposts transitions. "Turning to the product side…", "Two things need your attention…".
  • Ends with the actions. Recap what the listener should do, since that is what they need to retain and they cannot scroll back.
  • Drops anything purely visual. Tables, code blocks, and diagrams should be summarised in a sentence or omitted, never read out.

Write this spoken script to its own .txt file. Keep the original written version with its links intact — the audio is a companion to it, not a replacement. The user will want to click through later.

Synthesise

bash
./scripts/tts.sh <input.txt> <output.mp3> [voice.safetensors]

The script strips any residual markdown, splits the text on sentence boundaries into ~600 character chunks (quality degrades on long single inputs), synthesises each chunk, and concatenates the result into a mono MP3 at 96kbps — small enough to sync to a phone, good enough for speech.

Environment overrides:

VariablePurpose
SPEAK_TTS_BINPath to a specific pocket-tts binary; skips all auto-detection.
SPEAK_TTS_HOMEWhere to create/find the cached virtualenv. Default ~/.cache/speak-summary/venv.
Show full SKILL.md (334 more words)Show less

Voices

The default English voice is alba. To use a different one, pocket-tts supports voice cloning from a short clean audio sample:

bash
pocket-tts export-voice --help

Pass the resulting .safetensors file as the third argument to the script.

Only clone a voice you have the rights to use. Do not clone a real person's voice — colleague, customer, or public figure — without their explicit consent.

Output

  • Default to ~/Music/Briefings/ unless the user says otherwise; it is easy to point a phone or podcast app at.
  • Name files <subject>-<YYYY-MM-DD>.mp3.
  • Report the path, duration, and size.
  • Offer to play it: afplay <path> on macOS, ffplay -nodisp -autoexit <path> elsewhere.

Length guidance

Aim for 4–6 minutes for a routine digest, which is roughly 600–900 spoken words at a natural pace. If the source would run past about 10 minutes, say so and offer either a tighter edit or a split into multiple files — attention drops off sharply beyond that for informational audio.

Chaining onto other skills

The natural pattern is gather → summarise → speak:

  • roundup → speak-summary — a spoken version of the status briefing.
  • daily-prep → speak-summary — tomorrow's schedule, listened to tonight.
  • meeting-minutes → speak-summary — catch up on a meeting you missed.

When invoked as part of a chain, do not re-summarise. The upstream skill owns what to say; this skill owns how it sounds. Take its output, rewrite it for the ear, and synthesise.

To run unattended (a briefing waiting before breakfast), schedule the upstream skill with a workflow and have it finish by calling this one.

Troubleshooting

Audio cuts off mid-sentence. A chunk exceeded the model's comfortable length. Shorten the sentences in the spoken script.

Words mispronounced. Spell them phonetically in the input — "Kubernetes" as "koo-ber-net-eez". This is a normal part of preparing a spoken script.

First run is slow. That is the one-off model download. Later runs start in about a second.

pocket-tts not found after install. The virtualenv may be stale, or your python3 may be outside the supported 3.10–3.14 range. Delete ~/.cache/speak-summary/venv and re-run, or point SPEAK_TTS_BIN at a known binary.

© github, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (scripts) in skills/speak-summary of github/awesome-copilot.

  • SKILL.md
  • scripts/tts.sh

Open the folder on GitHubat commit 727ff2e

Compare with similar skills

Speak Summary next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Speak Summary compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Speak Summary this skillgithub/awesome-copilot40k—~1.6kAutomated safety check: PassMIT
Vox DirectorAlisa0808/vox-director2.2k—~5.6kAutomated safety check: PassMIT
Whiteboard Videognipbao/codex-whiteboard-video-skill323—~7.2kAutomated safety check: NotesMIT
Muapi DirectorAnil-matcha/vox-ai-motion-graphics-generator240—~679Automated safety check: PassNone
Webcode Local Windows Tts Installershuyu-labs/WebCode278—~787Automated safety check: PassCustom licence
Ermdougcalobrisi/erm112—~1.3kAutomated safety check: NotesMIT

Similar skills

  • Vox Director

    Alisa0808/vox-director

    Turn ONE topic into a finished Vox-style paper-collage explainer / ad video, end to end on the Atlas Cloud API + local ffmpeg — script, collage keyframes, motion, voice-over, music, captions, all…

    2.2k GitHub stars~5.6k tokensUpdated 2 days ago
    Media & CreativeAuto-check passed
  • Whiteboard Video

    gnipbao/codex-whiteboard-video-skill

    Generate smooth hand-drawn whiteboard and story videos directly inside Codex from text scripts, GPT Image 2 color storyboards, scene plans, SVGs, line art, or local images.

    323 GitHub stars~7.2k tokensUpdated 1 mo ago
    Media & CreativeAuto-check: notes
  • Muapi Director

    Anil-matcha/vox-ai-motion-graphics-generator

    Turn ONE topic, talking-head video, or photo into a finished Vox-style paper-collage explainer / ad video on the MuAPI platform (api.muapi.ai) + local ffmpeg — script, collage keyframes, motion…

    240 GitHub stars~679 tokensUpdated 1 mo ago
    Media & CreativeAuto-check passed
  • A skill your agent uses when building a local Windows WebCode installer from this repo for machine testing, especially when the package must bundle the Kokoro or sherpa-onnx Reply TTS service, model…

    278 GitHub stars~787 tokensUpdated 3 mo ago
    Media & CreativeAuto-check passed
  • Erm

    dougcalobrisi/erm

    Install and run erm, the local CLI that removes filler words / disfluencies (um, uh, er, erm, ah, hmm, mhm, mm, uh-huh and elongations) from spoken-audio recordings.

    112 GitHub stars~1.3k tokensUpdated 1 mo ago
    Media & CreativeAuto-check: notes
  • Video Podcast Maker

    dtsola/xiaoyaosearch

    Turns a topic into a 4K horizontal video podcast through research, scripting, text-to-speech, Remotion rendering and background music, and can learn styles from references.

    1k GitHub stars~3.4k tokensUpdated 5 days ago
    Media & CreativeAuto-check passed

More from github/awesome-copilot

All 417 skills in this repo
  • Acquire Codebase Knowledge

    github/awesome-copilot

    Official

    Maps an unfamiliar codebase into seven evidence-backed documents in docs/codebase/, using a scan script and templates, for onboarding or architecture write-ups.

    40k GitHub starsUsed in 1 repo~2.3k tokens
    Auto-check passed
  • Azure Architecture Autopilot

    github/awesome-copilot

    Official

    Designs Azure infrastructure from a natural-language description, or diagrams an existing resource group, then refines the design through conversation and deploys it with Bicep.

    40k GitHub starsUsed in 1 repo~1.9k tokens
    Auto-check passed
  • Draw.io Diagram Generator

    github/awesome-copilot

    Official

    Generates, edits and validates draw.io files with correct mxGraph XML, covering flowcharts, architecture, sequence, ER and UML class diagrams.

    40k GitHub starsUsed in 1 repo~4.9k tokens
    Auto-check passed
  • Credit Risk Data Cleaning

    github/awesome-copilot

    Official

    Cleans raw credit data and screens variables before loan modeling, dropping unstable, noisy or redundant features and writing an Excel report of every step.

    40k GitHub starsUsed in 1 repo~1.5k tokens
    Auto-check passed
  • Daily Focus Board

    github/awesome-copilot

    Official

    Builds a warm, browser-based daily focus board the user updates by talking to their agent, with Eisenhower priorities, a brain-dump box and kind not-today carryover.

    40k GitHub stars~3k tokensUpdated today
    Auto-check passed
  • Python Pypi Package Builder

    github/awesome-copilot

    Official

    End-to-end skill for building, testing, linting, versioning, and publishing a production-grade Python library to PyPI.

    40k GitHub starsUsed in 1 repo~4.6k tokens
    Auto-check passed

Works with

Questions about Speak Summary

What does Speak Summary do?

Convert text, markdown, or a summary produced by another skill into a listenable MP3 using local CPU-only neural text-to-speech. Speak Summary is an agent skill from github/awesome-copilot, published by the product's own GitHub organization. Convert text, markdown, or a summary produced by another skill into a listenable MP3 using local CPU-only neural text-to-speech.

When should I use Speak Summary?

Speak Summary fits situations like: the user asks to read this out; turn this into audio; I want to listen to this; podcast version.

How do I install Speak Summary in Claude Code?

Run `npx skills add github/awesome-copilot --skill speak-summary -a claude-code`. Or copy the skill folder (skills/speak-summary in github/awesome-copilot) into .claude/skills/speak-summary in your project. Claude Code loads it when a task matches its description.

How do I install Speak Summary in Codex?

Run `npx skills add github/awesome-copilot --skill speak-summary -a codex`. Or copy the skill folder (skills/speak-summary in github/awesome-copilot) into .agents/skills/speak-summary in your project. Codex loads it when a task matches its description.

Can I use Speak Summary in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add github/awesome-copilot --skill speak-summary -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/speak-summary, .gemini/skills/speak-summary, .github/skills/speak-summary and .opencode/skills/speak-summary in your project.

What does Speak Summary need to run?

Going by SKILL.md and its folder, Speak Summary needs a shell for the scripts in its folder and the command-line tools its instructions call (brew, pip and apt-get). Our summary lists: Python 3; A Bash shell.

Does Speak Summary access the network?

SKILL.md names 1 domain. As links in the text: github.com. This is read from the text; nothing was executed.

Is Speak Summary safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Speak Summary use?

Speak Summary is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Speak Summary use?

About 1.6k tokens (SKILL.md is roughly 6.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Speak Summary?

Skills that share tags, products or a category with Speak Summary: Vox Director (Alisa0808/vox-director, 2.2k stars), Whiteboard Video (gnipbao/codex-whiteboard-video-skill, 323 stars), Muapi Director (Anil-matcha/vox-ai-motion-graphics-generator, 240 stars) and Webcode Local Windows Tts Installer (shuyu-labs/WebCode, 278 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Speak Summary?

github (a GitHub organization, an official publisher) maintains it in github/awesome-copilot, which has 39,748 GitHub stars. The repository holds 417 skills in this directory. The repository was last updated on October 7, 2026.

Source: github/awesome-copilot on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.