Agent skill

Horizon Speech

by peters in peters/horizon

Check that a microphone is usable for Horizon dictation and report which layer is at fault — no signal, bad level, or a genuine model/accent limit.

MITAuto-check passedMedia & Creative

Install Horizon Speech

skills CLI
$ npx skills add peters/horizon --skill horizon-speech -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install peters/horizon horizon-speech --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/peters/horizon.git skills-src && mkdir -p .claude/skills && cp -r skills-src/assets/plugins/codex/skills/horizon-speech .claude/skills/horizon-speech && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
horizon-speech
GitHub stars
716
Token cost
~1.8k tokens
SKILL.md length
996 words
Files
2
Skills in repo
6
Repo updated
First seen
Licence
MIT

At a glance

Check that a microphone is usable for Horizon dictation and report which layer is at fault — no signal, bad level, or a genuine model/accent limit.

  • Works in 5 steps: Compare configured device against reality → Record a short sample → Measure the level → …
  • Dictation produces wrong
  • SKILL.md covers Read this first, 1. Compare configured device…, 2. Record a short sample and 3. Measure the level, plus 3 more sections
  • Runs Python scripts from its folder; calls ffmpeg and python3

What it does

Horizon Speech is an agent skill from peters/horizon. Check that a microphone is usable for Horizon dictation and report which layer is at fault — no signal, bad level, or a genuine model/accent limit. Use when dictation produces wrong, unrelated, or empty text.

Its SKILL.md is about 1.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 1 other file (for example `level.py`).

It sits in Media & Creative, covering Transcription. It works with Model Context Protocol. The repository describes itself as: GPU-accelerated terminal board that puts all your sessions on an infinite canvas. The licence is MIT.

When your agent uses it

  • Dictation produces wrong
  • Tasks that involve Transcription

Example prompts

  • “/horizon-speech”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Compare configured device against reality
  2. Record a short sample
  3. Measure the level
  4. Transcribe the same file
  5. Verdict

What it can do on your machine

Read from SKILL.md and the folder at commit e2e6058. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • ffmpeg
    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Horizon Speech loads about 1.8k tokens when it runs. Until then it costs about 56 tokens; SKILL.md has 996 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~56
When it runs · the whole SKILL.md, loaded when a task matches
~1.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from peters/horizon at commit e2e6058, republished under its MIT licence (© peters). 996 words, ~1,754 tokens.

Download SKILL.mdSave it as .claude/skills/horizon-speech/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
horizon-speech
description
Check that a microphone is usable for Horizon dictation and report which layer is at fault — no signal, bad level, or a genuine model/accent limit. Use when dictation produces wrong, unrelated, or empty text.

Horizon speech input check

This skill uses local audio diagnostics. Horizon has no speech MCP tool. Do not treat dictation as an MCP-controlled capability.

Run this when a user says dictation "types the wrong thing", produces text they never said, or produces nothing.

Read this first

Whisper-family models hallucinate fluent, plausible text when fed silence. Given a digitally silent recording, NB-Whisper will confidently emit something like Esther Smith, forfatter — grammatical, idiomatic, and completely unrelated to the user. The output looks like a model, accent, or dialect problem. It is almost always a capture problem.

So: measure the capture level before touching models, languages, or config. Never conclude "the model is bad at your accent" until step 3 shows real signal.

A user cannot tell these apart from the transcript alone. That is the whole reason this skill exists.

1. Compare configured device against reality

Horizon resolves its config by checking, in order, $HOME/.horizon/, then $XDG_CONFIG_HOME/horizon/, then a relative horizon.yaml/horizon.yml in the working directory. Resolution keys off HOME on every platform — Windows included; USERPROFILE is never consulted, so a native Windows launch without HOME set falls back to the relative path. Confirm which file is actually live before trusting it, or you may inspect a config Horizon never loaded.

Read features.speech from that file:

  • input_device: "" means the system default input is used.
  • A non-empty name is matched case-insensitively: exact, then substring, then a normalized identity key.

If the configured name matches no present device, Horizon falls back to the system default and only logs a warning. The user never sees it. A config naming a microphone that is no longer plugged in therefore looks like it works. Check the name against the live device list before anything else.

List input devices:

  • Linux (PipeWire/PulseAudio): wpctl status, or pactl list short sources
  • Linux (ALSA only): arecord -l
  • macOS: system_profiler SPAudioDataType, or ffmpeg -f avfoundation -list_devices true -i ""
  • Windows: ffmpeg -list_devices true -f dshow -i dummy, or Get-PnpDevice -Class AudioEndpoint in PowerShell

Also confirm the device is not a Bluetooth headset in HSP/HFP mode. That profile is narrowband and heavily compressed; it degrades accuracy badly. Prefer a wired or USB microphone, or force the A2DP-era high-quality input if the stack offers one.

2. Record a short sample

Ask the user to speak normally for about six seconds. Record mono at 16 kHz — that is what the models consume.

Record from the device step 1 matched, never the system default. When input_device names a non-default microphone, sampling the default measures a different device than Horizon uses and can yield the exact opposite diagnosis — a silent default while the real microphone works, or the reverse.

  • Linux (ALSA): arecord -D <device> -f S16_LE -r 16000 -c 1 -d 6 sample.wav, where <device> is a name from arecord -L
  • Linux (PipeWire): pw-record --target <node-name-or-id> --rate 16000 --channels 1 sample.wav
  • macOS: ffmpeg -f avfoundation -i ":<device-index>" -ar 16000 -ac 1 -t 6 sample.wav
  • Windows: ffmpeg -f dshow -i audio="<device name>" -ar 16000 -ac 1 -t 6 sample.wav

Only fall back to the default device when input_device is empty, since that is what Horizon itself then uses.

Tell the user when recording starts. If you launch the recorder as a background task, add a visible countdown first — otherwise the window elapses while they are still reading your message, and you will measure an empty room and misdiagnose it as a dead microphone.

Show full SKILL.md (434 more words)Show less

3. Measure the level

level.py sits next to this SKILL.md. Invoke it by its full path — the working directory is normally the user's workspace, not the installed skill directory, so a bare level.py will not be found:

python3 <dir containing this SKILL.md>/level.py sample.wav

On Windows use the standard launcher: py -3 ...\level.py sample.wav.

It needs only the Python 3 standard library, and reports peak, RMS and clipped samples, each as a percentage of full scale so the numbers mean the same thing at every sample width:

ReadingMeaningFix
peak below 0.6%No signal — capture muted, or no deviceUnmute capture; confirm the right device is selected and present
clipped samples > 20Too hot — distortionLower gain; turn microphone boost off first
rms above 25%Hotter than necessaryReduce gain slightly
rms 5–25%Ideal for ASRNothing to do
rms 2–5%UsableA little more gain would help
rms below 2%Too quietRaise gain, or move closer

4. Transcribe the same file

Run the profile's model against sample.wav and compare with what the user actually said. Horizon's models are transcribe.cpp GGUFs; if a transcribe-cli build is available:

transcribe-cli -m <model>.gguf -l <lang> sample.wav

Otherwise have the user dictate the same sentence in Horizon and compare.

5. Verdict

Combine steps 3 and 4 — this is the part the user cannot do alone:

  • No signal + fluent unrelated text → capture is dead. The text is hallucinated. Fix the microphone; the model is fine.
  • Good level + empty transcript → real audio, no speech detected. Usually room noise only, or the user did not speak inside the window.
  • Clipping + roughly correct text → working but distorted. Reduce gain; accuracy will improve, especially on consonants.
  • Good level + mostly correct text with wrong proper nouns or English technical terms → this is the model's genuine limit, not the microphone. Suggest an initial prompt to bias vocabulary, or a profile whose model suits the language better. Only at this point is accent or dialect worth discussing.

Report which of these it is explicitly. "Your microphone is fine, the model mis-heard a term" and "your microphone captured nothing" are opposite fixes and look identical in the transcript.

Adjusting gain

Turn microphone boost off before raising capture volume — boost amplifies the microphone's own noise floor as much as the voice.

  • Linux (PipeWire): wpctl set-volume <source-id> <0.0-1.0>. This maps onto the ALSA hardware control. Boost lives in amixer -c <card>, e.g. amixer -c <card> sset 'Mic Boost' 0.
  • macOS: System Settings → Sound → Input → Input volume.
  • Windows: Settings → System → Sound → Input → device properties. Also disable "Audio Enhancements", which can gate or suppress speech.

Re-run steps 2–3 after each change. Gain that sounds fine to a human can still be clipping.

© peters, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in assets/plugins/codex/skills/horizon-speech of peters/horizon.

  • SKILL.md
  • level.py

Open the folder on GitHubat commit e2e6058

Compare with similar skills

Horizon Speech next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Horizon Speech compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Horizon Speech this skillpeters/horizon716—~1.8kAutomated safety check: PassMIT
Resolve Audiosamuelgursky/davinci-resolve-mcp3.5k—~1.3kAutomated safety check: PassMIT
Premiere Captionsayushozha/AdobePremiereProMCP120—~724Automated safety check: PassMIT
Proofreadsonilo-ai/skills115—~4.4kAutomated safety check: NotesMIT
Transloadit Media Processinggithub/awesome-copilot40k1 repos~1.3kAutomated safety check: PassMIT
Scenario Caption Studioscenario-labs/skills946—~3.4kAutomated safety check: PassMIT

Similar skills

  • Resolve Audio

    samuelgursky/davinci-resolve-mcp

    Audio and Fairlight work in the DaVinci Resolve MCP. An agent skill from samuelgursky/davinci-resolve-mcp.

    3.5k GitHub stars~1.3k tokensUpdated today
    Media & CreativeAuto-check passed
  • Premiere Captions

    ayushozha/AdobePremiereProMCP

    Import, verify, structurally validate, and export timed captions or subtitles in Adobe Premiere Pro.

    120 GitHub stars~724 tokensUpdated 5 days ago
    Media & CreativeAuto-check passed
  • Proofread

    sonilo-ai/skills

    Transcribe a video with Sonilo and translate the transcript into editable .srt files — one per target language, plus the detected source language — so the wording can be read and corrected before…

    115 GitHub stars~4.4k tokensUpdated 2 days ago
    Media & CreativeAuto-check: notes
  • Transloadit Media Processing

    github/awesome-copilot

    Official

    Process media files (video, audio, images, documents) using Transloadit.

    40k GitHub starsUsed in 1 repo~1.3k tokens
    Media & CreativeAuto-check passed
  • Scenario Caption Studio

    scenario-labs/skills

    A skill your agent uses when a video needs its spoken words on screen through Scenario via MCP: burned-in styled captions for a TikTok, Reels, or Shorts cut, ad captions for sound-off feeds, YouTube…

    946 GitHub stars~3.4k tokensUpdated yesterday
    Media & CreativeAuto-check passed
  • Scenario Video Assembly

    scenario-labs/skills

    A skill your agent uses when generated clips must become a finished video on Scenario via MCP: cutting a shot list together, laying a timeline, concatenating with transitions, overlaying a logo or…

    946 GitHub stars~2.3k tokensUpdated yesterday
    Media & CreativeAuto-check passed

More from peters/horizon

  • Horizon Device

    peters/horizon

    Manage Horizon native VNC Device panels and drive isolated local desktops for simulators and native application tests through devicepanel and the horizon-device CLI/MCP.

    716 GitHub stars~4k tokensUpdated today
    Auto-check passed
  • Horizon App Testing

    peters/horizon

    Run or inspect declared iOS and Android native app tests through Horizon devicetestrun and app MCP tools.

    716 GitHub stars~277 tokensUpdated today
    Auto-check passed
  • Horizon Cast

    peters/horizon

    Cast Horizon panels, workspaces, cloud cards, or its main window to Apple TV through the public cast MCP tool.

    716 GitHub stars~227 tokensUpdated today
    Auto-check passed
  • Horizon Browser

    peters/horizon

    Control, inspect, or audit Horizon browser panels through public browser MCP tools.

    716 GitHub stars~5.1k tokensUpdated today
    Auto-check passed
  • Horizon Cloud

    peters/horizon

    Inspect Horizon cloud offers and companions, read or act on the cloud list of your workspace, control explicitly authorized companion workers, use a worker Local Network Bridge, or ask for GitHub…

    716 GitHub stars~408 tokensUpdated today
    Auto-check passed

Questions about Horizon Speech

What does Horizon Speech do?

Check that a microphone is usable for Horizon dictation and report which layer is at fault — no signal, bad level, or a genuine model/accent limit. Horizon Speech is an agent skill from peters/horizon. Check that a microphone is usable for Horizon dictation and report which layer is at fault — no signal, bad level, or a genuine model/accent limit.

When should I use Horizon Speech?

Horizon Speech fits situations like: dictation produces wrong; tasks that involve Transcription.

How do I install Horizon Speech in Claude Code?

Run `npx skills add peters/horizon --skill horizon-speech -a claude-code`. Or copy the skill folder (assets/plugins/codex/skills/horizon-speech in peters/horizon) into .claude/skills/horizon-speech in your project. Claude Code loads it when a task matches its description.

How do I install Horizon Speech in Codex?

Run `npx skills add peters/horizon --skill horizon-speech -a codex`. Or copy the skill folder (assets/plugins/codex/skills/horizon-speech in peters/horizon) into .agents/skills/horizon-speech in your project. Codex loads it when a task matches its description.

Can I use Horizon Speech in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add peters/horizon --skill horizon-speech -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/horizon-speech, .gemini/skills/horizon-speech, .github/skills/horizon-speech and .opencode/skills/horizon-speech in your project.

What does Horizon Speech need to run?

Going by SKILL.md and its folder, Horizon Speech needs Python for the scripts in its folder and the command-line tools its instructions call (ffmpeg and python3). Our summary lists: Python 3.

Does Horizon Speech access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Horizon Speech safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Horizon Speech use?

Horizon Speech is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Horizon Speech use?

About 1.8k tokens (SKILL.md is roughly 7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Horizon Speech?

Skills that share tags, products or a category with Horizon Speech: Resolve Audio (samuelgursky/davinci-resolve-mcp, 3.5k stars), Premiere Captions (ayushozha/AdobePremiereProMCP, 120 stars), Proofread (sonilo-ai/skills, 115 stars) and Transloadit Media Processing (github/awesome-copilot, 40k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Horizon Speech?

peters (a GitHub user) maintains it in peters/horizon, which has 716 GitHub stars. The repository holds 6 skills in this directory. The repository was last updated on October 11, 2026.

Source: peters/horizon on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.