Agent skill

Audio Voice Recovery

by pproenca in pproenca/dot-skills

Audio forensics and voice recovery guidelines for CSI-level audio analysis.

MITAuto-check passedMedia & Creative

Install Audio Voice Recovery

skills CLI
$ npx skills add pproenca/dot-skills --skill audio-voice-recovery -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install pproenca/dot-skills audio-voice-recovery --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/pproenca/dot-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/.experimental/audio-voice-recovery .claude/skills/audio-voice-recovery && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
audio-voice-recovery
GitHub stars
215
Token cost
~3.3k tokens
SKILL.md length
1,067 words
Files
55 (incl. scripts, references, assets)
Skills in repo
41
Repo updated
First seen
Licence
MIT

At a glance

Audio forensics and voice recovery guidelines for CSI-level audio analysis.

  • Works in 8 steps: Signal Preservation & Analysis (CRITICAL) → Noise Profiling & Estimation (CRITICAL) → Spectral Processing (HIGH) → …
  • Tasks involving audio enhancement
  • SKILL.md covers When to Apply, Rule Categories by Priority, Quick Reference and Essential Tools, plus 5 more sections
  • Calls python3, brew and pip

What it does

Audio Voice Recovery is an agent skill from pproenca/dot-skills. Audio forensics and voice recovery guidelines for CSI-level audio analysis. This skill should be used when recovering voice from low-quality or low-volume audio, enhancing degraded recordings, performing forensic audio analysis, or transcribing difficult audio. Triggers on tasks involving audio enhancement, noise reduction, voice isolation, forensic authentication, or audio transcription.

Its SKILL.md is about 3.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 57 other files, including scripts, reference files and assets (for example `AGENTS.md`, `assets/templates/_template.md` and `metadata.json`).

It sits in Media & Creative, covering Transcription and Authentication. The repository describes itself as: A collection of AI agent skills following the Agent Skills open format. The licence is MIT.

When your agent uses it

  • Tasks involving audio enhancement
  • Noise reduction
  • Voice isolation
  • Forensic authentication

Example prompts

  • “/audio-voice-recovery”

Requirements

  • Python 3

Workflow steps

8 steps, taken from the step headings in SKILL.md.

  1. Signal Preservation & Analysis (CRITICAL)
  2. Noise Profiling & Estimation (CRITICAL)
  3. Spectral Processing (HIGH)
  4. Voice Isolation & Enhancement (HIGH)
  5. Temporal Processing (MEDIUM-HIGH)
  6. Transcription & Recognition (MEDIUM)
  7. Forensic Authentication (MEDIUM)
  8. Tool Integration & Automation (LOW-MEDIUM)

What it can do on your machine

Read from SKILL.md and the folder at commit cf93c57. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/, which the agent can run.

    Shell commands in SKILL.md call:

    • python3
    • brew
    • pip
    • ffmpeg
    • whisper

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Audio Voice Recovery loads about 3.3k tokens when it runs, and up to ~67k if it reads all its reference files. Until then it costs about 103 tokens; SKILL.md has 1,067 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~103
When it runs · the whole SKILL.md, loaded when a task matches
~3.3k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~67k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from pproenca/dot-skills at commit cf93c57, republished under its MIT licence (© pproenca). 1,067 words, ~3,303 tokens.

Download SKILL.mdSave it as .claude/skills/audio-voice-recovery/SKILL.md (or your agent's skills folder). This skill also uses 54 other files; get the full folder from GitHub.
name
audio-voice-recovery
description
Audio forensics and voice recovery guidelines for CSI-level audio analysis. This skill should be used when recovering voice from low-quality or low-volume audio, enhancing degraded recordings, performing forensic audio analysis, or transcribing difficult audio. Triggers on tasks involving audio enhancement, noise reduction, voice isolation, forensic authentication, or audio transcription.

Forensic Audio Research Audio Voice Recovery Best Practices

Comprehensive audio forensics and voice recovery guide providing CSI-level capabilities for recovering voice from low-quality, low-volume, or damaged audio recordings. Contains 45 rules across 8 categories, prioritized by impact to guide audio enhancement, forensic analysis, and transcription workflows.

When to Apply

Reference these guidelines when:

  • Recovering voice from noisy or low-quality recordings
  • Enhancing audio for transcription or legal evidence
  • Performing forensic audio authentication
  • Analyzing recordings for tampering or splices
  • Building automated audio processing pipelines
  • Transcribing difficult or degraded speech

Rule Categories by Priority

PriorityCategoryImpactPrefixRules
1Signal Preservation & AnalysisCRITICALsignal-5
2Noise Profiling & EstimationCRITICALnoise-5
3Spectral ProcessingHIGHspectral-6
4Voice Isolation & EnhancementHIGHvoice-7
5Temporal ProcessingMEDIUM-HIGHtemporal-5
6Transcription & RecognitionMEDIUMtranscribe-5
7Forensic AuthenticationMEDIUMforensic-5
8Tool Integration & AutomationLOW-MEDIUMtool-7

Quick Reference

1. Signal Preservation & Analysis (CRITICAL)
2. Noise Profiling & Estimation (CRITICAL)
3. Spectral Processing (HIGH)
4. Voice Isolation & Enhancement (HIGH)
5. Temporal Processing (MEDIUM-HIGH)
6. Transcription & Recognition (MEDIUM)
7. Forensic Authentication (MEDIUM)
8. Tool Integration & Automation (LOW-MEDIUM)

Essential Tools

ToolPurposeInstall
FFmpegFormat conversion, filteringbrew install ffmpeg
SoXNoise profiling, effectsbrew install sox
WhisperSpeech transcriptionpip install openai-whisper
librosaPython audio analysispip install librosa
noisereduceML noise reductionpip install noisereduce
AudacityVisual editingbrew install audacity

Use the bundled scripts to generate objective baselines, create a workflow plan, and verify results.

  • scripts/preflight_audio.py - Generate a forensic preflight report (JSON or Markdown).
  • scripts/plan_from_preflight.py - Create a workflow plan template from the preflight report.
  • scripts/compare_audio.py - Compare objective metrics between baseline and processed audio.

Example usage:

bash
# 1) Analyze and capture baseline metrics
python3 skills/.experimental/audio-voice-recovery/scripts/preflight_audio.py evidence.wav --out preflight.json

# 2) Generate a workflow plan template
python3 skills/.experimental/audio-voice-recovery/scripts/plan_from_preflight.py --preflight preflight.json --out plan.md

# 3) Compare baseline vs processed metrics
python3 skills/.experimental/audio-voice-recovery/scripts/compare_audio.py \
  --before evidence.wav \
  --after enhanced.wav \
  --format md \
  --out comparison.md
Show full SKILL.md (495 more words)Show less

Forensic Preflight Workflow (Do This Before Any Changes)

Align preflight with SWGDE Best Practices for the Enhancement of Digital Audio (20-a-001) and SWGDE Best Practices for Forensic Audio (08-a-001). Establish an objective baseline state and plan the workflow so processing does not introduce clipping, artifacts, or false "done" confidence. Use scripts/preflight_audio.py to capture baseline metrics and preserve the report with the case file.

Capture and record before processing:

  • Record evidence identity and integrity: path, filename, file size, SHA-256 checksum, source, format/container, codec
  • Record signal integrity: sample rate, bit depth, channels, duration
  • Measure baseline loudness and levels: LUFS/LKFS, true peak, peak, RMS, dynamic range, DC offset
  • Detect clipping and document clipped-sample percentage, peak headroom, exact time ranges
  • Identify noise profile: stationary vs non-stationary, dominant noise bands, SNR estimate
  • Locate the region of interest (ROI) and document time ranges and changes over time
  • Inspect spectral content and estimate speech-band energy and intelligibility risk
  • Scan for temporal defects: dropouts, discontinuities, splices, drift
  • Evaluate channel correlation and phase anomalies (if stereo)
  • Extract and preserve metadata: timestamps, device/model tags, embedded notes

Procedure:

  1. Prepare a forensic working copy, verify hashes, and preserve the original untouched.
  2. Locate ROI and target signal; document exact time ranges and changes across the recording.
  3. Assess challenges to intelligibility and signal quality; map challenges to mitigation strategies.
  4. Identify required processing and plan a workflow order that avoids unwanted artifacts. Generate a plan draft with scripts/plan_from_preflight.py and complete it with case-specific decisions.
  5. Measure baseline loudness and true peak per ITU-R BS.1770 / EBU R 128 and record peak/RMS/DC offset.
  6. Detect clipping and dropouts; if clipping is present, declip first or pause and document limitations.
  7. Inspect spectral content and noise type; collect representative noise profile segments and estimate SNR.
  8. If stereo, evaluate channel correlation and phase; document anomalies.
  9. Create a baseline listening log (multiple devices) and define success criteria for intelligibility and listenability.

Failure-pattern guardrails:

  • Do not process until every preflight field is captured.
  • Document every process, setting, software version, and time segment to enable repeatability.
  • Compare each processed output to the unprocessed input and assess progress toward intelligibility and listenability.
  • Avoid over-processing; review removed signal (filter residue) to avoid removing target signal components.
  • Keep intermediate files uncompressed and preserve sample rate/bit depth when moving between tools.
  • Perform a final review against the original; if unsatisfactory, revise or stop and report limitations.
  • If the request is not achievable, communicate limitations and do not declare completion.
  • Require objective metrics and A/B listening before declaring completion.
  • Do not rely solely on objective metrics; corroborate with critical listening.
  • Take listening breaks to avoid ear fatigue during extended reviews.

Quick Enhancement Pipeline

bash
# 1. Analyze original (run preflight and capture baseline metrics)
python3 skills/.experimental/audio-voice-recovery/scripts/preflight_audio.py evidence.wav --out preflight.json

# 2. Create working copy with checksum
cp evidence.wav working.wav
sha256sum evidence.wav > evidence.sha256

# 3. Apply enhancement
ffmpeg -i working.wav -af "\
  highpass=f=80,\
  adeclick=w=55:o=75,\
  afftdn=nr=12:nf=-30:nt=w,\
  equalizer=f=2500:t=q:w=1:g=3,\
  loudnorm=I=-16:TP=-1.5:LRA=11\
" enhanced.wav

# 4. Transcribe
whisper enhanced.wav --model large-v3 --language en

# 5. Verify original unchanged
sha256sum -c evidence.sha256

# 6. Verify improvement (objective comparison + A/B listening)
python3 skills/.experimental/audio-voice-recovery/scripts/compare_audio.py \
  --before evidence.wav \
  --after enhanced.wav \
  --format md \
  --out comparison.md

How to Use

Read individual reference files for detailed explanations and code examples:

Reference Files

FileDescription
AGENTS.mdComplete compiled guide with all rules
references/_sections.mdCategory definitions and ordering
assets/templates/_template.mdTemplate for new rules
metadata.jsonVersion and reference information

© pproenca, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 54 other files (scripts, references, assets) in skills/.experimental/audio-voice-recovery of pproenca/dot-skills.

  • SKILL.md
  • AGENTS.md
  • assets/templates/_template.md
  • metadata.json
  • references/_sections.md
  • references/forensic-chain-custody.md
  • references/forensic-enf-analysis.md
  • references/forensic-metadata.md
  • references/forensic-speaker-id.md
  • references/forensic-tampering.md
  • references/noise-adaptive-estimation.md
  • references/noise-avoid-overprocessing.md
  • references/noise-identify-type.md
  • references/noise-profile-silence.md
  • references/noise-snr-assessment.md
  • references/signal-analyze-first.md
  • references/signal-bit-depth.md
  • references/signal-lossless-format.md
  • … and 37 more

Open the folder on GitHubat commit cf93c57

Compare with similar skills

Audio Voice Recovery next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Audio Voice Recovery compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Audio Voice Recovery this skillpproenca/dot-skills215—~3.3kAutomated safety check: PassMIT
Bilibili HubOpenMinis/MinisSkills446—~2.2kAutomated safety check: PassMIT
Gemini Live API Devgoogle-gemini/gemini-skills4.3k—~4.6kAutomated safety check: PassApache-2.0
HyperFrames Media Useheygen-com/hyperframes60k—~2.4kAutomated safety check: PassApache-2.0
Native Subtitle Quote Imagechengyi-ai/native-subtitle-quote-image2.6k—~1.8kAutomated safety check: PassMIT
Edu Chem Videowy51ai/edulab1.4k—~2.1kAutomated safety check: NotesApache-2.0

Similar skills

  • Bilibili Hub

    OpenMinis/MinisSkills

    A skill for reading and writing Bilibili data with Python + UV, using bilibili-api-python + aiohttp.

    446 GitHub stars~2.2k tokensUpdated 2 days ago
    Media & CreativeAuto-check passed
  • Gemini Live API Dev

    google-gemini/gemini-skills

    Official

    A skill your agent uses when building real-time, bidirectional streaming applications with the Gemini Live API, or migrating legacy Live models (2.0/2.5/3.1) to Gemini 3.8 Live.

    4.3k GitHub stars~4.6k tokensUpdated 3 days ago
    Backend & APIsAuto-check passed
  • HyperFrames Media Use

    heygen-com/hyperframes

    Finds, generates and edits media for HyperFrames video projects: music, sound effects, images, icons, logos, voiceovers, captions and color grades.

    60k GitHub stars~2.4k tokensUpdated today
    Media & CreativeAuto-check passed
  • Native Subtitle Quote Image

    chengyi-ai/native-subtitle-quote-image

    将本地视频或用户有权处理的在线视频,经过来源获取、文字稿定位、选题选句、精确取帧、紧凑裁切、拼图和逐张质检,制作成 3:4 或保留画面原比例的视频字幕长图。支持两种明确分开的输出:保留画面内已烧录字幕的原生字幕模式,以及把已审核的时间点与台词绘制到真实视频帧上的脚本字幕模式。用户要求原生字幕截图、字幕帧拼图、YouTube…

    2.6k GitHub stars~1.8k tokensUpdated today
    Media & CreativeAuto-check passed
  • Edu Chem Video

    wy51ai/edulab

    A skill your agent uses when asked to make an explainer / walkthrough video (讲解视频、解题视频、例题精讲、微课) for a chemistry problem (化学题: 氧化还原配平 双线桥 电子守恒, 物质的量计算, 化学平衡 三段式 平衡常数 转化率 反应速率, 离子反应, 电化学, 溶液 滴定…

    1.4k GitHub stars~2.1k tokensUpdated today
    Media & CreativeAuto-check: notes
  • Reconstruct a complete, searchable memory from an untrusted meeting or conversation transcript.

    3k GitHub stars~714 tokensUpdated yesterday
    Media & CreativeAuto-check passed

More from pproenca/dot-skills

All 41 skills in this repo
  • Codemod React Pipeline

    pproenca/dot-skills

    Guided, scripted pipeline for running JSX/TSX/React codemods safely across large legacy codebases.

    215 GitHub stars~1.6k tokensUpdated 1 mo ago
    Auto-check passed
  • Dev Rfc

    pproenca/dot-skills

    Create well-structured RFCs and technical proposals for software projects.

    215 GitHub stars~3.8k tokensUpdated 1 mo ago
    Auto-check passed
  • Dx Harness

    pproenca/dot-skills

    Developer-experience friction auditing and fixing — slow onboarding, repeated manual setup steps, missing bootstrap/reset/seed scripts, undiscoverable conventions.

    215 GitHub stars~1.5k tokensUpdated 1 mo ago
    Auto-check passed
  • Language Spec Author

    pproenca/dot-skills

    Turn a rough idea for a language into a complete, implementable specification — a DSL, query, config/data, template, or protocol language — by interviewing the author dimension by dimension until…

    215 GitHub stars~2.4k tokensUpdated 1 mo ago
    Auto-check passed
  • Python Pep Author

    pproenca/dot-skills

    Drafting Python Enhancement Proposals (PEPs) — proposing a Python language feature, a standard library change, an interoperability standard, or an informational/process document for the Python…

    215 GitHub stars~2.1k tokensUpdated 1 mo ago
    Auto-check passed
  • Designs new features, extensions, or modifications to Uncle Bob's Acceptance Pipeline Specification — new mutation strategies, Gherkin syntax support, report formats, pipeline stages, IR fields, or…

    215 GitHub stars~1.5k tokensUpdated 1 mo ago
    Auto-check passed

Questions about Audio Voice Recovery

What does Audio Voice Recovery do?

Audio forensics and voice recovery guidelines for CSI-level audio analysis. Audio Voice Recovery is an agent skill from pproenca/dot-skills. Audio forensics and voice recovery guidelines for CSI-level audio analysis.

When should I use Audio Voice Recovery?

Audio Voice Recovery fits situations like: tasks involving audio enhancement; noise reduction; voice isolation; forensic authentication.

How do I install Audio Voice Recovery in Claude Code?

Run `npx skills add pproenca/dot-skills --skill audio-voice-recovery -a claude-code`. Or copy the skill folder (skills/.experimental/audio-voice-recovery in pproenca/dot-skills) into .claude/skills/audio-voice-recovery in your project. Claude Code loads it when a task matches its description.

How do I install Audio Voice Recovery in Codex?

Run `npx skills add pproenca/dot-skills --skill audio-voice-recovery -a codex`. Or copy the skill folder (skills/.experimental/audio-voice-recovery in pproenca/dot-skills) into .agents/skills/audio-voice-recovery in your project. Codex loads it when a task matches its description.

Can I use Audio Voice Recovery in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add pproenca/dot-skills --skill audio-voice-recovery -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/audio-voice-recovery, .gemini/skills/audio-voice-recovery, .github/skills/audio-voice-recovery and .opencode/skills/audio-voice-recovery in your project.

What does Audio Voice Recovery need to run?

Going by SKILL.md and its folder, Audio Voice Recovery needs the command-line tools its instructions call (python3, brew, pip, ffmpeg and whisper). Our summary lists: Python 3.

Does Audio Voice Recovery access the network?

SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Audio Voice Recovery safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Audio Voice Recovery use?

Audio Voice Recovery is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Audio Voice Recovery use?

About 3.3k tokens (SKILL.md is roughly 13k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 64k tokens, read only when the agent opens those files.

What are the alternatives to Audio Voice Recovery?

Skills that share tags, products or a category with Audio Voice Recovery: Bilibili Hub (OpenMinis/MinisSkills, 446 stars), Gemini Live API Dev (google-gemini/gemini-skills, 4.3k stars), HyperFrames Media Use (heygen-com/hyperframes, 60k stars) and Native Subtitle Quote Image (chengyi-ai/native-subtitle-quote-image, 2.6k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Audio Voice Recovery?

pproenca (a GitHub user) maintains it in pproenca/dot-skills, which has 215 GitHub stars. The repository holds 41 skills in this directory. The repository was last updated on August 15, 2026.

Source: pproenca/dot-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.