Agent skill

Call Voice Deepfake Liveness Detector

by CALLE-AI in CALLE-AI/awesome-phone-call-agents

Offline experimental scorer of supplied acoustic features for phone-workflow demonstrations.

MITAuto-check passedMedia & Creative

Install Call Voice Deepfake Liveness Detector

skills CLI
$ npx skills add CALLE-AI/awesome-phone-call-agents --skill call-voice-deepfake-liveness-detector -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install CALLE-AI/awesome-phone-call-agents call-voice-deepfake-liveness-detector --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/CALLE-AI/awesome-phone-call-agents.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/call-voice-deepfake-liveness-detector .claude/skills/call-voice-deepfake-liveness-detector && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
call-voice-deepfake-liveness-detector
GitHub stars
107
Token cost
~1.9k tokens
SKILL.md length
842 words
Files
6 (incl. scripts, references)
Skills in repo
101
Repo updated
First seen
Licence
MIT

At a glance

Offline experimental scorer of supplied acoustic features for phone-workflow demonstrations.

  • Works in 5 steps: Spectral Flatness Analysis: Real human… → Phase Noise Floor Inspection: Microphone… → Temporal Periodicity Deviation (Jitter):… → …
  • Tasks that involve Text to speech and voice
  • SKILL.md covers Scientific Foundation, How It Works, Decision Matrix and Key Features, plus 6 more sections
  • Runs Python scripts from its folder

What it does

Call Voice Deepfake Liveness Detector is an agent skill from CALLE-AI/awesome-phone-call-agents. Offline experimental scorer of supplied acoustic features for phone-workflow demonstrations. Returns heuristic review suggestions, not caller authentication or a deployed deepfake detector.

Its SKILL.md is about 1.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 7 other files, including scripts and reference files (for example `references/examples.md`, `references/research-papers.md` and `references/safety.md`).

It sits in Media & Creative, covering Text to speech and voice. The repository describes itself as: Portable phone-call Agent Skills, apps, examples, adapters, and scheduler recipes for AI agents. The licence is MIT.

When your agent uses it

  • Tasks that involve Text to speech and voice

Example prompts

  • “/call-voice-deepfake-liveness-detector”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Spectral Flatness Analysis: Real human speech contains a rich, irregular harmonic structure. TTS engines produce overly smooth…
  2. Phase Noise Floor Inspection: Microphone recordings always contain ambient environmental noise (Gaussian noise floor, ≈−40 to −60 dB)…
  3. Temporal Periodicity Deviation (Jitter): Human vocal production involves micro-variations in pitch period. Synthetic voices exhibit…
  4. Liveness Score Aggregation: The three sub-scores are averaged into a single liveness_score (0.0 = fully synthetic → 1.0 = genuine human)…
  5. Suggested Review Action: Returns an action string only. It does not route, block, escalate, log, or authorize a transaction.

What it can do on your machine

Read from SKILL.md and the folder at commit 38d4118. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 2 files in scripts/ (Python), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Call Voice Deepfake Liveness Detector loads about 1.9k tokens when it runs, and up to ~3.8k if it reads all its reference files. Until then it costs about 57 tokens; SKILL.md has 842 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~57
When it runs · the whole SKILL.md, loaded when a task matches
~1.9k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~3.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from CALLE-AI/awesome-phone-call-agents at commit 38d4118, republished under its MIT licence (© CALLE-AI). 842 words, ~1,901 tokens.

Download SKILL.mdSave it as .claude/skills/call-voice-deepfake-liveness-detector/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.
name
call-voice-deepfake-liveness-detector
description
Offline experimental scorer of supplied acoustic features for phone-workflow demonstrations. Returns heuristic review suggestions, not caller authentication or a deployed deepfake detector.
version
1.0.0

Voice Deepfake Liveness Detector

This skill directly addresses one of the most urgent and growing threats in telephony: AI-powered voice cloning fraud. With tools like ElevenLabs and VALL-E, a bad actor can clone a person's voice from just a few seconds of audio and use it to impersonate family members, executives, or even government officials in phone calls.

Unlike transcript-oriented call-fraud-shield, this prototype accepts three already-extracted numeric acoustic features. It includes no raw-waveform extraction, telephony middleware, call blocking, identity verification, or audit logger. Its synthetic-fixture labels are advisory examples, not proof that a caller is human or synthetic.

Note: This skill and call-fraud-shield are complementary, not competing. Together, they form a two-layer defense: acoustic (this skill) + semantic (call-fraud-shield).

Scientific Foundation

This skill is grounded in the peer-reviewed anti-spoofing literature:

  • ASVspoof 2021 (IEEE/ACM TASLP, 2023): The gold-standard benchmark for countermeasures against TTS and voice-conversion (VC) attacks. Defines the Equal Error Rate (EER) metric and the feature taxonomy adopted by this skill.
  • ADD 2022 Challenge (ICASSP 2022): The first challenge to address Partially Fake Audio Detection (PF) — where only a fragment of the call is synthetic — a critical real-world threat model.
  • FTC Voice Cloning Challenge (April 2024): Official U.S. regulatory acknowledgement of the consumer threat, validating the practical urgency of this skill's design.

How It Works

The helper combines supplied features using illustrative thresholds. A separate host would have to extract and validate these inputs; the following rationale is a design hypothesis:

  1. Spectral Flatness Analysis: Real human speech contains a rich, irregular harmonic structure. TTS engines produce overly smooth spectrograms with unnaturally low spectral flatness variance.
  2. Phase Noise Floor Inspection: Microphone recordings always contain ambient environmental noise (Gaussian noise floor, ≈−40 to −60 dB). AI-generated audio is typically unnaturally clean (≈−85 to −95 dB).
  3. Temporal Periodicity Deviation (Jitter): Human vocal production involves micro-variations in pitch period. Synthetic voices exhibit machine-regular periodicity detectable by statistical analysis.
  4. Liveness Score Aggregation: The three sub-scores are averaged into a single liveness_score (0.0 = fully synthetic → 1.0 = genuine human). Requires ≥ 2 valid features; otherwise returns INDETERMINATE.
  5. Suggested Review Action: Returns an action string only. It does not route, block, escalate, log, or authorize a transaction.

Decision Matrix

liveness_scoreClassificationConfidenceRecommended Action
>= 0.70HUMANHIGHproceed_normally
0.35 – 0.69LIKELY_SYNTHETICMEDIUMinsert_friction_challenge
< 0.35SYNTHETICHIGHblock_and_escalate_to_human
< 2 valid featuresINDETERMINATELOWapply_standard_verification

Key Features

  • Offline scoring: The helper receives numeric features; no extraction latency or end-to-end latency has been established.
  • Fail-open design: Insufficient audio data (packet loss, muted caller) → INDETERMINATE, never a false positive SYNTHETIC block.
  • Complementary to call-fraud-shield: Acoustic layer defense, not a replacement for semantic analysis.
  • Threshold tunable per use-case: High-security financial calls → threshold=0.60; general customer service → threshold=0.30.
  • Partial feature resilience: Operates meaningfully with any 2 of 3 features — robust to selective sensor failure.

Configuration Reference

ParameterDefaultRangeDescription
SYNTHETIC_THRESHOLD0.350.10 – 0.60Score below this → SYNTHETIC. Raise for stricter security.
HUMAN_THRESHOLD0.700.50 – 0.90Score at/above this → HUMAN.
MIN_QUALITY_NOISE_DB-80.0−90 to −60Noise floor below this → suspicious (TTS-like).
JITTER_SYNTHETIC_MAX0.0050.001 – 0.010Pitch jitter below this → TTS-like.
FLATNESS_SYNTHETIC_MAX0.0150.005 – 0.030Spectral flatness variance below this → TTS-like.
segment_duration_ms500250 – 2000Audio window length for feature extraction.
Show full SKILL.md (314 more words)Show less

Expected Outcomes & Metrics

These are unvalidated design targets, not benchmark results of the shipped helper:

MetricTargetNotes
Equal Error Rate (EER)< 5% targetNot validated against ASVspoof by this contribution
False Positive Rate (FPR)< 3%Legitimate callers incorrectly flagged
False Negative Rate (FNR)< 8%Synthetic voices that pass undetected
Latency (per segment)< 20msOn standard cloud compute
Indeterminate Rate< 5%Calls with insufficient audio quality

Use Cases

  • Executive impersonation / CEO fraud prevention: Detect attackers cloning a CFO's voice to authorize wire transfers.
  • Grandparent/family emergency scam defense: Flag synthetic voices claiming to be a grandchild in distress.
  • KYC (Know Your Customer) voice verification: Add a liveness gate to voice-based identity verification flows.
  • Call center integrity monitoring: Continuous background monitoring for high-volume inbound call operations.
  • Partially-fake audio detection: Catch hybrid attacks where only key phrases are synthesized into a real conversation stream (ADD 2022 PF track).

Limitations & Known Constraints

  • Not a forensic tool: The liveness score is a probabilistic risk indicator, not legal-grade proof of synthesis. All SYNTHETIC flags must route to human review for final decision.
  • Codec degradation: Heavy audio compression (G.711, low-bitrate VoIP) can reduce spectral richness and push borderline human voices toward LIKELY_SYNTHETIC. Threshold should be recalibrated for heavily compressed environments.
  • Adversarial robustness: A sufficiently motivated attacker can add artificial noise to a synthetic voice to game the noise floor detector. This is an active research area (ADD 2022 FG track) and future versions will incorporate countermeasure ensembles.
  • Non-English accents & speech impediments: Diverse vocal profiles must be included in threshold validation datasets to prevent demographic bias.

Integration

Proposed integration only: a host could pass the advisory result to an agent after implementing feature extraction and its own verification controls. The diagram is not a shipped middleware hook:

[Inbound Audio Stream]
        |
[call-voice-deepfake-liveness-detector]  <-- this skill
        |
  liveness_score=0.28, flag=SYNTHETIC, confidence=HIGH
        |
[Conversational Agent]  <-- adapts: inserts friction challenge
        |
[call-fraud-shield]  <-- semantic/transcript layer

References

See references/research-papers.md for research inspiration, not validation of this heuristic. See references/safety.md for safety, bias, and compliance guidelines. See references/examples.md for end-to-end scenario walkthroughs.

© CALLE-AI, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 5 other files (scripts, references) in skills/call-voice-deepfake-liveness-detector of CALLE-AI/awesome-phone-call-agents.

  • SKILL.md
  • references/examples.md
  • references/research-papers.md
  • references/safety.md
  • scripts/deepfake_detector.py
  • scripts/test_deepfake_detector.py

Open the folder on GitHubat commit 38d4118

Compare with similar skills

Call Voice Deepfake Liveness Detector next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Call Voice Deepfake Liveness Detector compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Call Voice Deepfake Liveness Detector this skillCALLE-AI/awesome-phone-call-agents107—~1.9kAutomated safety check: PassMIT
MoneyPrinterTurbo Video Generatorharry0703/MoneyPrinterTurbo130k—~2.1kAutomated safety check: WarnMIT
HyperFrames Media Useheygen-com/hyperframes60k—~2.4kAutomated safety check: PassApache-2.0
Openspec OnboardSAP/e-mobility-charging-stations-simulator22725 repos~3.5kAutomated safety check: PassMIT
Blog AudioAgriciDaniel/claude-blog2.3k1 repos~2.2kAutomated safety check: NotesMIT
Musictadaspetra/loop2962 repos~827Automated safety check: PassMIT

Similar skills

  • MoneyPrinterTurbo Video Generator

    harry0703/MoneyPrinterTurbo

    Installs and runs MoneyPrinterTurbo to turn a topic or script into a finished short video with voice-over, subtitles, stock footage and music.

    130k GitHub stars~2.1k tokensUpdated yesterday
    Media & CreativeAuto-check: warnings
  • HyperFrames Media Use

    heygen-com/hyperframes

    Finds, generates and edits media for HyperFrames video projects: music, sound effects, images, icons, logos, voiceovers, captions and color grades.

    60k GitHub stars~2.4k tokensUpdated today
    Media & CreativeAuto-check passed
  • Openspec Onboard

    SAP/e-mobility-charging-stations-simulator

    Official

    Guided onboarding for OpenSpec - walk through a complete workflow cycle with narration and real codebase work.

    227 GitHub starsUsed in 25 repos~3.5k tokens
    Media & CreativeAuto-check passed
  • Blog Audio

    AgriciDaniel/claude-blog

    Generate audio narration of blog posts using Google Gemini TTS.

    2.3k GitHub starsUsed in 1 repo~2.2k tokens
    Media & CreativeAuto-check: notes
  • Music

    tadaspetra/loop

    Generate music using ElevenLabs Music API. An agent skill from tadaspetra/loop.

    296 GitHub starsUsed in 2 repos~827 tokens
    Media & CreativeAuto-check passed
  • Create News Video

    hoquanghai/Auto-Create-Video

    Tạo video tin tức ngắn 9:16 (~60s) từ URL bài báo hoặc file .txt tiếng Việt.

    319 GitHub starsUsed in 1 repo~3.7k tokens
    Media & CreativeAuto-check passed

More from CALLE-AI/awesome-phone-call-agents

All 101 skills in this repo
  • Accessible Outing Verifier

    CALLE-AI/awesome-phone-call-agents

    Demonstrates advisory accessibility-planning checks with offline fixtures and a proposed bounded CALL-E workflow; use for exploring unknown or qualified venue claims without making calls.

    107 GitHub stars~1.4k tokensUpdated yesterday
    Auto-check passed
  • Ground Truth Gate

    CALLE-AI/awesome-phone-call-agents

    A skill your agent uses when an agent holds some evidence for a physical-world claim but the evidence is broader, narrower, or older than the exact question asked, and it must first decide whether a…

    107 GitHub stars~3.3k tokensUpdated yesterday
    Auto-check passed
  • Is It Accessible

    CALLE-AI/awesome-phone-call-agents

    Call a venue and ask the accessibility questions that matter to one specific person — step-free entry, hearing loop, guide dogs, quiet hours, changing places — then return a per-need verdict backed…

    107 GitHub stars~4.3k tokensUpdated yesterday
    Auto-check passed
  • Landmark Navigation Assist

    CALLE-AI/awesome-phone-call-agents

    Turns a pre-written, building-level location config into a CALL-E outbound phone-call task that guides a delivery driver through the last few hundred metres to a specific building using landmarks…

    107 GitHub stars~1.6k tokensUpdated yesterday
    Auto-check passed
  • Research Gap Call Verifier

    CALLE-AI/awesome-phone-call-agents

    Turn cited business research into a bounded, approval-gated phone-call plan that asks only unresolved factual questions, then reconcile CALL-E-compatible results without treating voicemail, refusal…

    107 GitHub stars~1.7k tokensUpdated yesterday
    Auto-check passed
  • Structured Outcome Followup Call

    CALLE-AI/awesome-phone-call-agents

    Place a goal-driven CALL-E call that collects specific structured answers, score those answers against a deterministic rubric you supply, and conditionally trigger a follow-up action — all runnable…

    107 GitHub stars~1.4k tokensUpdated yesterday
    Auto-check passed

Questions about Call Voice Deepfake Liveness Detector

What does Call Voice Deepfake Liveness Detector do?

Offline experimental scorer of supplied acoustic features for phone-workflow demonstrations. Call Voice Deepfake Liveness Detector is an agent skill from CALLE-AI/awesome-phone-call-agents. Offline experimental scorer of supplied acoustic features for phone-workflow demonstrations.

When should I use Call Voice Deepfake Liveness Detector?

Call Voice Deepfake Liveness Detector fits situations like: tasks that involve Text to speech and voice.

How do I install Call Voice Deepfake Liveness Detector in Claude Code?

Run `npx skills add CALLE-AI/awesome-phone-call-agents --skill call-voice-deepfake-liveness-detector -a claude-code`. Or copy the skill folder (skills/call-voice-deepfake-liveness-detector in CALLE-AI/awesome-phone-call-agents) into .claude/skills/call-voice-deepfake-liveness-detector in your project. Claude Code loads it when a task matches its description.

How do I install Call Voice Deepfake Liveness Detector in Codex?

Run `npx skills add CALLE-AI/awesome-phone-call-agents --skill call-voice-deepfake-liveness-detector -a codex`. Or copy the skill folder (skills/call-voice-deepfake-liveness-detector in CALLE-AI/awesome-phone-call-agents) into .agents/skills/call-voice-deepfake-liveness-detector in your project. Codex loads it when a task matches its description.

Can I use Call Voice Deepfake Liveness Detector in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add CALLE-AI/awesome-phone-call-agents --skill call-voice-deepfake-liveness-detector -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/call-voice-deepfake-liveness-detector, .gemini/skills/call-voice-deepfake-liveness-detector, .github/skills/call-voice-deepfake-liveness-detector and .opencode/skills/call-voice-deepfake-liveness-detector in your project.

What does Call Voice Deepfake Liveness Detector need to run?

Going by SKILL.md and its folder, Call Voice Deepfake Liveness Detector needs Python for the scripts in its folder. Our summary lists: Python 3.

Does Call Voice Deepfake Liveness Detector access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Call Voice Deepfake Liveness Detector safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Call Voice Deepfake Liveness Detector use?

Call Voice Deepfake Liveness Detector is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Call Voice Deepfake Liveness Detector use?

About 1.9k tokens (SKILL.md is roughly 7.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.9k tokens, read only when the agent opens those files.

What are the alternatives to Call Voice Deepfake Liveness Detector?

Skills that share tags, products or a category with Call Voice Deepfake Liveness Detector: MoneyPrinterTurbo Video Generator (harry0703/MoneyPrinterTurbo, 130k stars), HyperFrames Media Use (heygen-com/hyperframes, 60k stars), Openspec Onboard (SAP/e-mobility-charging-stations-simulator, 227 stars) and Blog Audio (AgriciDaniel/claude-blog, 2.3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Call Voice Deepfake Liveness Detector?

CALLE-AI (a GitHub organization) maintains it in CALLE-AI/awesome-phone-call-agents, which has 107 GitHub stars. The repository holds 101 skills in this directory. The repository was last updated on October 10, 2026.

Source: CALLE-AI/awesome-phone-call-agents on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.