Agent skill

Erm

by dougcalobrisi in dougcalobrisi/erm

Install and run erm, the local CLI that removes filler words / disfluencies (um, uh, er, erm, ah, hmm, mhm, mm, uh-huh and elongations) from spoken-audio recordings.

MITAuto-check: notesMedia & Creative

Install Erm

skills CLI
$ npx skills add dougcalobrisi/erm --skill erm -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install dougcalobrisi/erm erm --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/dougcalobrisi/erm.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/erm .claude/skills/erm && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
erm
GitHub stars
112
Token cost
~1.3k tokens
SKILL.md length
672 words
Files
1
Skills in repo
2
Repo updated
First seen
Licence
MIT

At a glance

Install and run erm, the local CLI that removes filler words / disfluencies (um, uh, er, erm, ah, hmm, mhm, mm, uh-huh and elongations) from spoken-audio recordings.

  • Works in 4 steps: Install / run → Ask before choosing a command → Core workflow (the iterate loop) → …
  • The user wants to install
  • SKILL.md covers Resolving documentation, 1. Install / run, 2. Ask before choosing a command and 3. Core workflow (the iterate…, plus 1 more section
  • Calls uvx, ffmpeg and brew

What it does

Erm is an agent skill from dougcalobrisi/erm. Install and run erm, the local CLI that removes filler words / disfluencies (um, uh, er, erm, ah, hmm, mhm, mm, uh-huh and elongations) from spoken-audio recordings. Use when the user wants to install or set up erm, clean up a recording/podcast/voiceover, strip "ums" and "uhs" from audio, or asks which erm command to run. For fixing imperfect output or adjusting knobs, use the erm-tune skill instead.

Its SKILL.md is about 1.3k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Media & Creative, covering Text to speech and voice and Video production. It works with FFmpeg, Python and NVIDIA AI Platform. The repository describes itself as: Local CLI that strips disfluencies (um, uh, er, erm, ah, hmm, mhm, mm, uh-huh, etc) from recordings of English speech. The licence is MIT.

When your agent uses it

  • The user wants to install
  • Clean up a recording/podcast/voiceover
  • Strip ums and uhs from audio
  • Asks which erm command to run

Example prompts

  • “/erm”

Requirements

  • Python 3
  • Pre-approved tools (allowed-tools): Bash, Read, AskUserQuestion

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Install / run
  2. Ask before choosing a command
  3. Core workflow (the iterate loop)
  4. When results aren't perfect

What it can do on your machine

Read from SKILL.md and the folder at commit af2fe92. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Bash
    • Read
    • AskUserQuestion

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • uvx
    • ffmpeg
    • brew
    • apt
    • choco
    • uv
    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • doug.sh

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Erm loads about 1.3k tokens when it runs. Until then it costs about 102 tokens; SKILL.md has 672 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~102
When it runs · the whole SKILL.md, loaded when a task matches
~1.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Bash, Read, AskUserQuestion

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from dougcalobrisi/erm at commit af2fe92, republished under its MIT licence (© dougcalobrisi). 672 words, ~1,345 tokens.

Download SKILL.mdSave it as .claude/skills/erm/SKILL.md (or your agent's skills folder).
name
erm
description
Install and run erm, the local CLI that removes filler words / disfluencies (um, uh, er, erm, ah, hmm, mhm, mm, uh-huh and elongations) from spoken-audio recordings. Use when the user wants to install or set up erm, clean up a recording/podcast/voiceover, strip "ums" and "uhs" from audio, or asks which erm command to run. For fixing imperfect output or adjusting knobs, use the erm-tune skill instead.
allowed-tools
Bash, Read, AskUserQuestion

erm — install and use

erm strips disfluencies from English speech audio. It transcribes with faster-whisper, runs extra audio-domain detectors for fillers Whisper hides, and splices with ffmpeg (energy-snapped, crossfaded, room-tone-matched).

Resolving documentation

When you need authoritative detail, resolve it in this order (each works in more environments than the last):

  1. erm --help and erm validate --help — definitive flags and defaults; works once installed.
  2. Public docs: https://doug.sh/docs/erm/ — usage, recipes, troubleshooting, etc.
  3. Bundled docs (Claude Code/Cowork plugin only): ${CLAUDE_PLUGIN_ROOT}/docs/*.md and the source of truth for flag defaults, ${CLAUDE_PLUGIN_ROOT}/src/erm/cli.py.

Never guess flag names or defaults — read one of the above.

1. Install / run

erm needs Python 3.11+ and ffmpeg/ffprobe on PATH.

  1. Check ffmpeg: ffmpeg -version. If missing, suggest the OS install (brew install ffmpeg, apt install ffmpeg, choco install ffmpeg).
  2. Resolve a launcher — prefer uv (broadest, no persistent install):
    • Tier 1 — uvx (preferred). If uv --version succeeds, run erm straight from PyPI with uvx erm … — no install step; uv fetches and caches the environment on first run, so later runs are fast. Pin a version with uvx erm@<version> … when needed. Verify: uvx erm --help.
    • Tier 2 — venv fallback (no uv on PATH). Create an isolated env and install from PyPI:
      sh
      python3 -m venv .venv
      source .venv/bin/activate
      pip install erm
      erm --help   # verify

Launcher convention. In the commands throughout this skill, erm means the launcher you resolved above: prefix with uvx under tier 1 (e.g. uvx erm INPUT.wav --dry-run), or use plain erm after activating the venv under tier 2.

Transcription runs on CPU by default (no setup). GPU is optional and needs the CUDA runtime libs; --device auto falls back to CPU. Add the CUDA wheels to the same environment — uvx --with nvidia-cublas-cu12 --with nvidia-cudnn-cu12 erm … under tier 1, or pip install nvidia-cublas-cu12 nvidia-cudnn-cu12 into the venv under tier 2. See the transcription docs page for details.

2. Ask before choosing a command

erm's behavior forks on a couple of choices. Use AskUserQuestion to settle these only when they aren't already clear from the request, then proceed:

  • What kind of audio? podcast/interview · video (caption-timed or A/V sync) · multitrack stem · already-clean studio. This selects the recipe.
  • Render mode? --mode remove (default — excises fillers, timeline shrinks) vs --mode silence (mutes in place, duration preserved — required for video sync and multitrack stems).
  • Video input — audio or picture? For a video file, erm emits the cleaned audio only (.wav) by default (the "pull the audio out" case). Add --video to render the picture too — container inferred from the input, A/V in sync by construction. With --video: --mode silence stream-copies the picture losslessly (caption/lip-sync safe), --video-splice {crossfade,cut} picks the splice style, --vcodec/--crf/--preset tune the re-encode. See the video doc.

If the user already implied the answers (e.g. "clean my podcast"), don't ask — pick the sensible default and say what you chose.

Then read the recipes doc and use the matching copy-paste command.

Show full SKILL.md (192 more words)Show less

3. Core workflow (the iterate loop)

  1. Inspect first: erm INPUT.wav --dry-run — prints/writes the cut-list JSON (*-cuts-*.json); renders nothing. Review what it intends to cut.
  2. Render: erm INPUT.wav — writes INPUT-cleaned-<timestamp>.wav next to the input.
  3. Validate: erm validate INPUT.wav OUTPUT.wav — re-transcribes the output and asserts no fillers survive, plus container/duration sanity. Exit 0 = pass.

Useful flags (confirm with erm --help): -o/--output, --json, --model, --device, --fillers, --video (render the picture from a video input). The full usage doc explains the workflow in depth.

Adjusting the word list. If the user wants to strip an extra word (e.g. "also remove 'basically' / 'like'"), prefer --add-fillers "basically,like" — it keeps the built-in defaults and unions the new words on top. Use --remove-fillers WORD to drop a default that over-matches their voice. Reach for --fillers only to replace the whole set, since it requires re-typing every stem. Custom words match verbatim (no automatic elongation). See the recipes doc → "Custom filler vocabulary".

4. When results aren't perfect

If fillers remain, real words get clipped, splices click/smear, the noise floor pumps, or words run together — hand off to the erm-tune skill, which maps each symptom to the right knob.

© dougcalobrisi, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/erm of dougcalobrisi/erm.

Open the folder on GitHubat commit af2fe92

Compare with similar skills

Erm next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Erm compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Erm this skilldougcalobrisi/erm112—~1.3kAutomated safety check: NotesMIT
Vox DirectorAlisa0808/vox-director2.2k—~5.6kAutomated safety check: PassMIT
Whiteboard Videognipbao/codex-whiteboard-video-skill327—~7.2kAutomated safety check: NotesMIT
Muapi DirectorAnil-matcha/vox-ai-motion-graphics-generator246—~679Automated safety check: PassNone
Webcode Local Windows Tts Installershuyu-labs/WebCode278—~787Automated safety check: PassCustom licence
Video Podcast Makerdtsola/xiaoyaosearch1k—~3.4kAutomated safety check: PassMIT

Similar skills

  • Vox Director

    Alisa0808/vox-director

    Turn ONE topic into a finished Vox-style paper-collage explainer / ad video, end to end on the Atlas Cloud API + local ffmpeg — script, collage keyframes, motion, voice-over, music, captions, all…

    2.2k GitHub stars~5.6k tokensUpdated 5 days ago
    Media & CreativeAuto-check passed
  • Whiteboard Video

    gnipbao/codex-whiteboard-video-skill

    Generate smooth hand-drawn whiteboard and story videos directly inside Codex from text scripts, GPT Image 2 color storyboards, scene plans, SVGs, line art, or local images.

    327 GitHub stars~7.2k tokensUpdated 1 mo ago
    Media & CreativeAuto-check: notes
  • Muapi Director

    Anil-matcha/vox-ai-motion-graphics-generator

    Turn ONE topic, talking-head video, or photo into a finished Vox-style paper-collage explainer / ad video on the MuAPI platform (api.muapi.ai) + local ffmpeg — script, collage keyframes, motion…

    246 GitHub stars~679 tokensUpdated 1 mo ago
    Media & CreativeAuto-check passed
  • A skill your agent uses when building a local Windows WebCode installer from this repo for machine testing, especially when the package must bundle the Kokoro or sherpa-onnx Reply TTS service, model…

    278 GitHub stars~787 tokensUpdated 3 mo ago
    Media & CreativeAuto-check passed
  • Video Podcast Maker

    dtsola/xiaoyaosearch

    Turns a topic into a 4K horizontal video podcast through research, scripting, text-to-speech, Remotion rendering and background music, and can learn styles from references.

    1k GitHub stars~3.4k tokensUpdated 8 days ago
    Media & CreativeAuto-check passed
  • Ffmpeg Skill

    kajisho5/ffmpeg-skill

    Edit video and audio with local FFmpeg from natural-language requests: cut, trim, join, resize/reframe (9:16, 1:1), speed change, captions and subtitles (SRT/ASS, animated, karaoke), logos and text…

    1.9k GitHub stars~7.4k tokensUpdated today
    Media & CreativeAuto-check passed

More from dougcalobrisi/erm

  • Erm Tune

    dougcalobrisi/erm

    Diagnose and tune erm's output quality. An agent skill from dougcalobrisi/erm.

    112 GitHub stars~1k tokensUpdated 1 mo ago
    Auto-check: notes

Questions about Erm

What does Erm do?

Install and run erm, the local CLI that removes filler words / disfluencies (um, uh, er, erm, ah, hmm, mhm, mm, uh-huh and elongations) from spoken-audio recordings. Erm is an agent skill from dougcalobrisi/erm. Install and run erm, the local CLI that removes filler words / disfluencies (um, uh, er, erm, ah, hmm, mhm, mm, uh-huh and elongations) from spoken-audio recordings.

When should I use Erm?

Erm fits situations like: the user wants to install; clean up a recording/podcast/voiceover; strip ums and uhs from audio; asks which erm command to run.

How do I install Erm in Claude Code?

Run `npx skills add dougcalobrisi/erm --skill erm -a claude-code`. Or copy the skill folder (skills/erm in dougcalobrisi/erm) into .claude/skills/erm in your project. Claude Code loads it when a task matches its description.

How do I install Erm in Codex?

Run `npx skills add dougcalobrisi/erm --skill erm -a codex`. Or copy the skill folder (skills/erm in dougcalobrisi/erm) into .agents/skills/erm in your project. Codex loads it when a task matches its description.

Can I use Erm in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add dougcalobrisi/erm --skill erm -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/erm, .gemini/skills/erm, .github/skills/erm and .opencode/skills/erm in your project.

What does Erm need to run?

Going by SKILL.md and its folder, Erm needs the command-line tools its instructions call (uvx, ffmpeg, brew, apt, choco and uv). Our summary lists: Python 3. Its frontmatter pre-approves these tools: Bash, Read, AskUserQuestion.

Does Erm access the network?

SKILL.md names 1 domain. As links in the text: doug.sh. This is read from the text; nothing was executed.

Is Erm safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Erm use?

Erm is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Erm use?

About 1.3k tokens (SKILL.md is roughly 5.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Erm?

Skills that share tags, products or a category with Erm: Vox Director (Alisa0808/vox-director, 2.2k stars), Whiteboard Video (gnipbao/codex-whiteboard-video-skill, 327 stars), Muapi Director (Anil-matcha/vox-ai-motion-graphics-generator, 246 stars) and Webcode Local Windows Tts Installer (shuyu-labs/WebCode, 278 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Erm?

dougcalobrisi (a GitHub user) maintains it in dougcalobrisi/erm, which has 112 GitHub stars. The repository holds 2 skills in this directory. The repository was last updated on August 23, 2026.

Source: dougcalobrisi/erm on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.