Agent skill

Clean Audio

by hassancs91 in hassancs91/claude-youtube-editor

Voice/audio cleanup step of the AI Video Editor pipeline — diagnose a video's background noise, pick the right denoise method, and produce a cleaned master (voice isolated, levels preserved, video…

MITAuto-check passedMedia & Creative

Install Clean Audio

skills CLI
$ npx skills add hassancs91/claude-youtube-editor --skill clean-audio -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install hassancs91/claude-youtube-editor clean-audio --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/hassancs91/claude-youtube-editor.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/clean-audio .claude/skills/clean-audio && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
clean-audio
GitHub stars
328
Token cost
~1.8k tokens
SKILL.md length
847 words
Files
1
Skills in repo
8
Repo updated
First seen
Licence
MIT

At a glance

Voice/audio cleanup step of the AI Video Editor pipeline — diagnose a video's background noise, pick the right denoise method, and produce a cleaned master (voice isolated, levels preserved, video…

  • Works in 7 steps: Diagnose the noise BEFORE choosing a… → Decide the method with the user from the… → A/B on a short sample FIRST (prove… → …
  • The user wants to clean the audio / voice
  • SKILL.md covers The two methods (pick by NOISE…, Inputs (read/measure first,…, Workflow and Decisions to surface to the user, plus 2 more sections
  • Calls ffmpeg and python; needs ELEVENLABS_API_KEY

What it does

Clean Audio is an agent skill from hassancs91/claude-youtube-editor. Voice/audio cleanup step of the AI Video Editor pipeline — diagnose a video's background noise, pick the right denoise method, and produce a cleaned master (voice isolated, levels preserved, video stream copied). Use when the user wants to "clean the audio / voice", "remove background noise", "denoise", "isolate voice", fix outdoor/room/water/hum/hiss noise, run ElevenLabs Voice Isolator or local RNNoise, A/B denoise methods, or produce a cleaned master for a video-N in this repo. Covers diagnosing the noise…

Its SKILL.md is about 1.8k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Media & Creative, covering Text to speech and voice, Video production and Image editing. It works with ElevenLabs and YouTube. The repository describes itself as: Record the talking head, Claude Code does the rest: the cut, the visuals, the voice, the sound effects, the thumbnail, and the YouTube upload. Every screen moment is built as… The licence is MIT.

When your agent uses it

  • The user wants to clean the audio / voice
  • Remove background noise
  • Fix outdoor/room/water/hum/hiss noise
  • Run ElevenLabs Voice Isolator

Example prompts

  • “clean the audio / voice”
  • “remove background noise”
  • “denoise”
  • “/clean-audio”

Requirements

  • Python 3
  • A credential in ELEVENLABS_API_KEY

Workflow steps

7 steps, taken from the first numbered list in SKILL.md.

  1. Diagnose the noise BEFORE choosing a method. Measure and look
  2. Decide the method with the user from the diagnosis (table above). If unsure, A/B both.
  3. A/B on a short sample FIRST (prove before spending / committing): cut a ~15s pause-rich sample,
  4. Clean the full master: python tools/clean_voice.py videos/video-N/reference/.mp4 --method [--model sh]
  5. Levels are preserved by RMS-match, not LUFS (the tool does this). Never match integrated LUFS —
  6. Give the user an in-context preview (optional but recommended): swap the clean audio onto the
  7. On approval, rewire the pipeline: point timeline.json "master" at the clean file so every future

What it can do on your machine

Read from SKILL.md and the folder at commit a6ac742. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • ffmpeg
    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • ELEVENLABS_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Clean Audio loads about 1.8k tokens when it runs. Until then it costs about 205 tokens; SKILL.md has 847 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~205
When it runs · the whole SKILL.md, loaded when a task matches
~1.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from hassancs91/claude-youtube-editor at commit a6ac742, republished under its MIT licence (© hassancs91). 847 words, ~1,759 tokens.

Download SKILL.mdSave it as .claude/skills/clean-audio/SKILL.md (or your agent's skills folder).
name
clean-audio
description
Voice/audio cleanup step of the AI Video Editor pipeline — diagnose a video's background noise, pick the right denoise method, and produce a cleaned master (voice isolated, levels preserved, video stream copied). Use when the user wants to "clean the audio / voice", "remove background noise", "denoise", "isolate voice", fix outdoor/room/water/hum/hiss noise, run ElevenLabs Voice Isolator or local RNNoise, A/B denoise methods, or produce a cleaned master for a video-N in this repo. Covers diagnosing the noise (spectrogram + levels), choosing eleven vs rnnoise by noise type, the sample A/B, tools/clean_voice.py, preserving levels (RMS-match, not LUFS), and rewiring the pipeline to the clean master. Not the SFX/music mix (that is /suggest-sfx + the final-mix step) and not the cut (that is /clean-cut).

clean-audio — voice cleanup

Take a locked master cut and remove its background noise, producing a cleaned master whose voice sounds natural and whose visuals are untouched. Runs early (once the cut is locked) so everything downstream — TSX bake, SFX mix, final assemble — sits on the clean voice. Work with the user; the final loudness/limiting is the final-mix step's job, this step is "denoise only, levels preserved."

The engine is tools/clean_voice.py; this skill is the judgment around it: diagnose → pick method → A/B → clean → rewire.

The two methods (pick by NOISE TYPE — this is the core decision)

MethodWhat it isUse whenCost
--method elevenElevenLabs Voice Isolator (cloud ML voice/noise separation)Dynamic, broadband noise in the voice band — outdoor running water, wind, traffic, crowd, cafe. Local tools CANNOT remove these.1000 credits/min ($1 for a 5.5-min video); needs ELEVENLABS_API_KEY
--method rnnoise --model sh (or cb)Local RNNoise via ffmpeg arnndn (models in tools/models/rnnoise/)Stationary / mild noise (steady hiss, fan, some room tone). Free/offline. Only PARTIALLY removes dynamic noise.free

Proven on video-1 (shot outdoors with a stream): afftdn did ~nothing, RNNoise only partially darkened the water bed, ElevenLabs removed it near-completely (pauses to near-silence, voice + breaths intact). Rule of thumb: stationary noise → try local first; dynamic broadband (water/wind/traffic) → ElevenLabs.

Inputs (read/measure first, every time)

  • The master — videos/video-N/reference/<cut>.mp4 (or the locked cut). Original is NEVER modified; output is a new -clean / -clean-<model> file.
  • The composited preview (if it exists) — videos/video-N/output/video-N-preview.mp4, to make a clean in-context preview by swapping audio (its video is identical — no re-bake needed).
  • videos/video-N/work/timeline.json — its master field; you rewire this to the clean master on approval.

Workflow

  1. Diagnose the noise BEFORE choosing a method. Measure and look:
    • Levels: ffmpeg -i M -vn -af astats (RMS, peak, noise floor) + ebur128 (integrated LUFS, true peak).
    • Find speech-free gaps (grep edited-transcript.json for the biggest inter-word gaps) and measure the pure-noise RMS there vs speech RMS → the real SNR.
    • Spectrogram: ffmpeg -i M -vn -lavfi showspectrumpic=s=1500x600:legend=1:scale=log out.png and LOOK at it. Hum = steady horizontal lines (50/60Hz) → notch. Rumble = low band → high-pass. Broadband bed that fills the voice band and fluctuates = dynamic (water/wind) → ElevenLabs. HF hiss = bright top band.
    • Note if the export is already produced (compressed/normalized/peak-maxed) — it limits what's recoverable.
  2. Decide the method with the user from the diagnosis (table above). If unsure, A/B both.
  3. A/B on a short sample FIRST (prove before spending / committing): cut a ~15s pause-rich sample, run each candidate method, level-match them to each other, and compare — by ear (the real test) AND by spectrogram (pauses going dark = noise removed) and residual level. Let the user pick.
  4. Clean the full master: python tools/clean_voice.py videos/video-N/reference/<cut>.mp4 --method <chosen> [--model sh] → <cut>-clean.mp4 (or -clean-<model>.mp4). Video stream COPIED (fast, non-destructive, keeps 4K60).
  5. Levels are preserved by RMS-match, not LUFS (the tool does this). Never match integrated LUFS — it is gated and inflated by the removed noise, and over-boosts the voice into clipping. The clean file will read a lower integrated LUFS than the noisy original; that is expected (the noise was padding the number), the voice RMS is unchanged. Final loudness to -14 LUFS is the final-mix step's job.
  6. Give the user an in-context preview (optional but recommended): swap the clean audio onto the composited preview — ffmpeg -i preview.mp4 -i <cut>-clean.mp4 -map 0:v -map 1:a -c:v copy -c:a aac -shortest preview-clean.mp4 (video identical, no re-bake). For a full A/B, also export FULL_*.mp3 scrub files.
  7. On approval, rewire the pipeline: point timeline.json "master" at the clean file so every future bake/mix uses the clean voice; re-bake the preview if needed.
Show full SKILL.md (253 more words)Show less

Decisions to surface to the user

  • Method (from the diagnosis) — and A/B if unsure.
  • Dead-silent gaps vs a faint ambience bed. Voice isolation removes ALL background; on an outdoor shot the dead-silent gaps can feel vacuum-sealed. Offer to add back a low-level neutral ambience if wanted.
  • Cost for ElevenLabs (~1000 credits/min) — confirm before running on the full master.

Principles (the house style)

  • Least processing that works. The goal is to remove distraction, not to make the voice sound processed. Prefer the gentlest method that clears the noise; don't over-strip a clean track.
  • Diagnose, then choose. The right tool depends on the noise type — never crank a denoiser blind. Local spectral/RNNoise can't separate dynamic broadband noise; that's ElevenLabs' job.
  • A/B before you commit (and before you spend). Prove on a sample; the user's ears decide.
  • Preserve levels; loudness is the final-mix step's. RMS-match with a peak ceiling, no compression here.
  • Non-destructive. Original master untouched; video stream copied; output is a new file.

Tooling quick reference

  • Clean: python tools/clean_voice.py IN.mp4 [--method eleven|rnnoise] [--model sh|cb] [-o OUT.mp4] [--no-preserve-loudness] [--keep]
  • Diagnose: ffmpeg -i M -vn -af astats -f null - · ffmpeg -i M -vn -lavfi showspectrumpic=... out.png (then Read the png).
  • RNNoise models: tools/models/rnnoise/<model>.rnnn (sh, cb).
  • Scratch samples/spectrograms go in the scratchpad, not the project.

Done = the noise is diagnosed, the method is chosen (A/B'd if needed), the full master is cleaned with levels preserved, the user has approved by ear, and — on approval — timeline.json points at the clean master. Update memory if a noise-type → method lesson emerges.

© hassancs91, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/clean-audio of hassancs91/claude-youtube-editor.

Open the folder on GitHubat commit a6ac742

Compare with similar skills

Clean Audio next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Clean Audio compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Clean Audio this skillhassancs91/claude-youtube-editor328—~1.8kAutomated safety check: PassMIT
Media Genclacky-ai/openclacky1.2k—~7.3kAutomated safety check: PassMIT
Paw Cra Agent Video Producerpawbytes/skill-suites113—~2.3kAutomated safety check: PassMIT
AI Presenter VideoNousResearch/hermes-agent252k—~2.3kAutomated safety check: PassMIT
Super Video MakerBomx/super-video-maker-skill310—~11kAutomated safety check: NotesNone
Hearyourvoicekillernay/HearYourVOICE140—~10kAutomated safety check: NotesMIT

Similar skills

  • Media Gen

    clacky-ai/openclacky

    Generate or edit images, videos, or audio in the current task.

    1.2k GitHub stars~7.3k tokensUpdated today
    Media & CreativeAuto-check passed
  • Paw Cra Agent Video Producer

    pawbytes/skill-suites

    Video production specialist for short-form, long-form, episodic, and motion graphics video.

    113 GitHub stars~2.3k tokensUpdated 7 days ago
    Media & CreativeAuto-check passed
  • AI Presenter Video

    NousResearch/hermes-agent

    Produces a presenter-led video from a topic or script plus one authorized presenter image, with captions, lip-sync checks and acceptance reports.

    252k GitHub stars~2.3k tokensUpdated today
    Media & CreativeAuto-check passed
  • Super Video Maker

    Bomx/super-video-maker-skill

    End-to-end AI video production skill for agentic frameworks.

    310 GitHub stars~11k tokensUpdated 2 mo ago
    Media & CreativeAuto-check: notes
  • Hearyourvoice

    killernay/HearYourVOICE

    The repeatable workflow for short Thai documentary/explainer videos — one topic in, one finished MP4 out.

    140 GitHub stars~10k tokensUpdated 2 mo ago
    Media & CreativeAuto-check: notes
  • AI Video Gen

    aAAaqwq/AGI-Super-Team

    End-to-end AI video generation - create videos from text prompts using image generation, video synthesis, voice-over, and editing.

    105 GitHub starsUsed in 1 repo~819 tokens
    Media & CreativeAuto-check: notes

More from hassancs91/claude-youtube-editor

All 8 skills in this repo
  • Youtube Thumbnail

    hassancs91/claude-youtube-editor

    Dedicated YouTube thumbnail generator — interviews you for exactly the style elements you want (environment, text budget, extras, accent color), then renders high-contrast, vibrant, face-consistent…

    328 GitHub stars~2.1k tokensUpdated 1 mo ago
    Auto-check: notes
  • Brand Setup

    hassancs91/claude-youtube-editor

    Makes this repo's videos look like YOUR channel instead of the house default — interviews you for palette, fonts, wordmark, motion energy, delivery specs and SFX taste, then rewrites brand.md +…

    328 GitHub stars~2.3k tokensUpdated 1 mo ago
    Auto-check passed
  • Clean Cut

    hassancs91/claude-youtube-editor

    Step 1 of the AI Video Editor pipeline — turn raw talking-head footage into a clean master cut.

    328 GitHub stars~4.3k tokensUpdated 1 mo ago
    Auto-check: notes
  • Fake Screencast

    hassancs91/claude-youtube-editor

    Turn static SCREENSHOTS into a simulated screen recording (TSX) — a fake screencast with an animated cursor that eases to targets and clicks, a browser URL bar that updates per page, hard-cut…

    328 GitHub stars~2.5k tokensUpdated 1 mo ago
    Auto-check passed
  • Make Tsx

    hassancs91/claude-youtube-editor

    Step 2 of the AI Video Editor pipeline — build the visual beats (Remotion TSX shots) over a project's master cut and bake a composited preview.

    328 GitHub stars~1.9k tokensUpdated 1 mo ago
    Auto-check: notes
  • Suggest Sfx

    hassancs91/claude-youtube-editor

    Step 4 of the AI Video Editor pipeline — the SFX pass. An agent skill from hassancs91/claude-youtube-editor.

    328 GitHub stars~3k tokensUpdated 1 mo ago
    Auto-check passed

Questions about Clean Audio

What does Clean Audio do?

Voice/audio cleanup step of the AI Video Editor pipeline — diagnose a video's background noise, pick the right denoise method, and produce a cleaned master (voice isolated, levels preserved, video…. Clean Audio is an agent skill from hassancs91/claude-youtube-editor. Voice/audio cleanup step of the AI Video Editor pipeline — diagnose a video's background noise, pick the right denoise method, and produce a cleaned master (voice isolated, levels preserved, video stream copied).

When should I use Clean Audio?

Clean Audio fits situations like: the user wants to clean the audio / voice; remove background noise; fix outdoor/room/water/hum/hiss noise; run ElevenLabs Voice Isolator.

How do I install Clean Audio in Claude Code?

Run `npx skills add hassancs91/claude-youtube-editor --skill clean-audio -a claude-code`. Or copy the skill folder (.claude/skills/clean-audio in hassancs91/claude-youtube-editor) into .claude/skills/clean-audio in your project. Claude Code loads it when a task matches its description.

How do I install Clean Audio in Codex?

Run `npx skills add hassancs91/claude-youtube-editor --skill clean-audio -a codex`. Or copy the skill folder (.claude/skills/clean-audio in hassancs91/claude-youtube-editor) into .agents/skills/clean-audio in your project. Codex loads it when a task matches its description.

Can I use Clean Audio in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add hassancs91/claude-youtube-editor --skill clean-audio -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/clean-audio, .gemini/skills/clean-audio, .github/skills/clean-audio and .opencode/skills/clean-audio in your project.

What does Clean Audio need to run?

Going by SKILL.md and its folder, Clean Audio needs the command-line tools its instructions call (ffmpeg and python) and credentials named ELEVENLABS_API_KEY. Our summary lists: Python 3; A credential in ELEVENLABS_API_KEY.

Does Clean Audio access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Clean Audio safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Clean Audio use?

Clean Audio is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Clean Audio use?

About 1.8k tokens (SKILL.md is roughly 7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Clean Audio?

Skills that share tags, products or a category with Clean Audio: Media Gen (clacky-ai/openclacky, 1.2k stars), Paw Cra Agent Video Producer (pawbytes/skill-suites, 113 stars), AI Presenter Video (NousResearch/hermes-agent, 252k stars) and Super Video Maker (Bomx/super-video-maker-skill, 310 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Clean Audio?

hassancs91 (a GitHub user) maintains it in hassancs91/claude-youtube-editor, which has 328 GitHub stars. The repository holds 8 skills in this directory. The repository was last updated on August 18, 2026.

Source: hassancs91/claude-youtube-editor on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.