Agent skill

Call Post Summary Faithfulness Auditor

by CALLE-AI in CALLE-AI/awesome-phone-call-agents

Offline experimental CALL-E helper that decomposes the agent postsummary into atomic claims and anchors each against the transcript, flagging unsupported and outcome-contradicted statements before…

MITAuto-check passed

Install Call Post Summary Faithfulness Auditor

skills CLI
$ npx skills add CALLE-AI/awesome-phone-call-agents --skill call-post-summary-faithfulness-auditor -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install CALLE-AI/awesome-phone-call-agents call-post-summary-faithfulness-auditor --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/CALLE-AI/awesome-phone-call-agents.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/call-post-summary-faithfulness-auditor .claude/skills/call-post-summary-faithfulness-auditor && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
call-post-summary-faithfulness-auditor
GitHub stars
107
Token cost
~1.9k tokens
SKILL.md length
951 words
Files
11 (incl. scripts, references)
Skills in repo
101
Repo updated
First seen
Licence
MIT

At a glance

Offline experimental CALL-E helper that decomposes the agent postsummary into atomic claims and anchors each against the transcript, flagging unsupported and outcome-contradicted statements before…

  • Works in 3 steps: Mask first. Any 7+-digit run (separators… → Decompose. Each summary sentence becomes… → Anchor. Each claim is folded (case,…
  • SKILL.md covers When To Use, When Not To Use, Verdicts and How It Works, plus 4 more sections
  • Runs Python scripts from its folder; calls python3

What it does

Call Post Summary Faithfulness Auditor is an agent skill from CALLE-AI/awesome-phone-call-agents. Offline experimental CALL-E helper that decomposes the agent postsummary into atomic claims and anchors each against the transcript, flagging unsupported and outcome-contradicted statements before automation acts on them, plus a faithful-summary goal template. It is not semantic entailment, proof of lying, or authorization to act.

Its SKILL.md is about 1.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 12 other files, including scripts and reference files (for example `references/example-call-result-contradicted.json`, `references/example-call-result-flat.json` and `references/example-call-result-no-checkable.json`).

The repository describes itself as: Portable phone-call Agent Skills, apps, examples, adapters, and scheduler recipes for AI agents. The licence is MIT.

Example prompts

  • “/call-post-summary-faithfulness-auditor”

Requirements

  • Python 3

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. Mask first. Any 7+-digit run (separators included) is masked in
  2. Decompose. Each summary sentence becomes atomic claims by kind -
  3. Anchor. Each claim is folded (case, meridiem to 24h, month names,

What it can do on your machine

Read from SKILL.md and the folder at commit 38d4118. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 2 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Call Post Summary Faithfulness Auditor loads about 1.9k tokens when it runs, and up to ~6k if it reads all its reference files. Until then it costs about 93 tokens; SKILL.md has 951 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~93
When it runs · the whole SKILL.md, loaded when a task matches
~1.9k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from CALLE-AI/awesome-phone-call-agents at commit 38d4118, republished under its MIT licence (© CALLE-AI). 951 words, ~1,887 tokens.

Download SKILL.mdSave it as .claude/skills/call-post-summary-faithfulness-auditor/SKILL.md (or your agent's skills folder). This skill also uses 10 other files; get the full folder from GitHub.
name
call-post-summary-faithfulness-auditor
description
Offline experimental CALL-E helper that decomposes the agent post_summary into atomic claims and anchors each against the transcript, flagging unsupported and outcome-contradicted statements before automation acts on them, plus a faithful-summary goal template. It is not semantic entailment, proof of lying, or authorization to act.
license
MIT

call-post-summary-faithfulness-auditor

The post_summary is the only artifact most automations ever parse. Audit it, claim by claim.

Every CALL-E application downstream of a call reads the agent-written post_summary blindly: outcome tokens route workflows, amounts and dates feed writebacks, action promises become follow-up tasks. If the agent hallucinates in that free-text field - a dollar figure nobody spoke, a "confirmed" that was really a voicemail - the hallucination silently becomes a wrong action. No existing sibling skill audits the summary itself: call-review gates schema fields, verity gates a single task_completed claim, and the PR-gated null-result helpers gate missing extraction. This skill catches the opposite failure - fabricated content that made it into the record.

When To Use

  • before any workflow acts on the post_summary - writebacks, outcome parsing, CRM updates, follow-up scheduling
  • after any call whose result feeds automation, as a gate between get_call_run and the action layer
  • when an outcome token appears in the summary with no matching conversation anywhere in the transcript

When Not To Use

  • when semantic nuance matters; anchoring is lexical, so a faithful paraphrase can read as UNSUPPORTED - human review decides, always
  • to prove the agent intended to deceive; this is claim anchoring, not a lie detector or an intent test
  • on masked spans; phone digits are masked before analysis by design and are unverifiable by the same token

Verdicts

VerdictMeaningSuggested routing
FAITHFULevery checkable claim anchored to a transcript turnproceed, keep the card as evidence
UNSUPPORTED_CLAIMSat least one claim value not found verbatim in any turnhold the writeback; verify each flagged claim against the call record
CONTRADICTED_CLAIMSan outcome claim conflicts with late callee speech under the polarity ruleblock automation; a human must read the call
NO_CHECKABLE_CLAIMSnothing machine-checkable: reason is summary_missing (empty field) or opinion_only (no anchored values)do not infer an outcome from the summary; use the structured outcome fields

How It Works

Three deterministic stages, offline, no LLM:

  1. Mask first. Any 7+-digit run (separators included) is masked in the summary and every turn before any claim work, keeping the last two characters. Phone numbers can never anchor and never leak into cards.
  2. Decompose. Each summary sentence becomes atomic claims by kind - outcome (confirm/cancel/decline/reschedule/... with a polarity), numeric ($X, party of N, unit-suffixed counts), date_time (month-day in either order, numeric MM/DD, weekday, and clock times in 12h or 24h form, all canonicalized to month-day and 24h strings), action (will send/email/call back/...), plus explicitly non_checkable notes for spelled numbers, masked-only sentences, and pure-opinion sentences.
  3. Anchor. Each claim is folded (case, meridiem to 24h, month names, day order, digit commas and decimals) and searched across all turns with boundary guards so substrings cannot false-anchor ("oct 1" does not match "oct 14", "14:00" does not match "2:14:00"). Outcome claims additionally check polarity against callee negatives in the final third of the call; CONTRADICTED claims record the firing turn in contradicted_by_turn, and UNSUPPORTED claims whose kind has no lexical presence anywhere in the transcript carry reason: kind_absent_from_transcript. Numerics never contradict - two different values can legitimately coexist in a call, so a value either anchors or reads UNSUPPORTED.

The card reports per-claim grades, per-grade counts, an overall verdict, a coverage_gaps list (the transcript contains an outcome word the summary never mentioned), and a fixed honesty disclaimer.

Craft: the faithful-summary goal

bash
python3 scripts/post_summary_faithfulness_auditor.py craft \
  --task "Confirm the reservation" \
  --facts "party of 4; Wednesday October 14; 2 p.m.; $20 deposit" \
  --outcome-token CONFIRMED

Emits a plan_call goal whose SUMMARY DISCIPLINE block constrains the agent to restate only what was spoken aloud, repeat numbers as digit words, never introduce an unspoken value, and state the outcome word exactly once - so the summary this skill later audits starts faithful.

Show full SKILL.md (361 more words)Show less

Limitations

  • spelled numbers ("four", "twelve") are flagged but not machine-checkable
  • paraphrase misses: "half past two" will not anchor a "2 p.m." claim
  • UNSUPPORTED is not proof of falsehood; the value may exist in audio nuance the transcript renders differently
  • date formats are month-day (either order), numeric MM/DD, weekday, and clock forms (12h and 24h); numeric dates are read as US month/day order, ISO dates and relative dates ("next Friday") are out of scope
  • contraction and negation edge cases ("can't" vs "cannot") may slip past the polarity regexes in rare phrasings
  • heuristic anchor: a positive outcome claim with no literal outcome verb in the transcript may read SUPPORTED when a LATE callee turn carries an unambiguous affirmative ack ("yes", "yep", "perfect", "lock it in", "we'll take it", ...) with no negative token in that turn - it mirrors how real callees confirm, but it is a heuristic, not entailment
  • heuristic action-ack relaxation: an action claim can grade SUPPORTED from an acknowledgment token ("okay", "ok", "sure", "will do", ...) found in ANY turn, not only late callee turns - a deliberate relaxation that keeps polite callees from sinking faithful summaries, but it is a heuristic, not entailment

Testing

bash
python3 scripts/test_post_summary_faithfulness_auditor.py

Covers masking, decomposition, anchoring, polarity contradiction, verdict routing, the CLI surface, and craft output against the bundled fixtures.

Scientific Foundation

ResearchRelevance
Long-form factuality in large language models (Wei, Yang, Song, Lu, Hu, Huang, Tran, Peng, Liu, Huang, Du, Le. NeurIPS 2024, arXiv 2403.18802)Claim-decomposition plus per-fact verification (SAFE): long free text is graded by splitting it into atomic facts and checking each independently - exactly this skill's stage model
MiniCheck: Efficient Fact-Checking of LLMs on Grounding Documents (Tang, Laban, Durrett. EMNLP 2024, arXiv 2404.10774)Grounding-document fact-checking as a defined task and the LLM-AggreFact benchmark - summary claims checked against a grounding source before the text is used
SummaC: Re-Visiting NLI-based Models for Inconsistency Detection in Summarization (Laban, Schnabel, Bennett, Hearst. TACL vol. 10, 2022, arXiv 2111.09525)Summary-vs-source consistency as a formal summarization requirement - the post_summary is a summary and the transcript is its source

This skill is the deterministic, offline, lexical-anchoring variant of that methodology - not a learned model. Every output is labeled with a fixed disclaimer stating exactly that.

© CALLE-AI, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 10 other files (scripts, references) in skills/call-post-summary-faithfulness-auditor of CALLE-AI/awesome-phone-call-agents.

  • SKILL.md
  • references/example-call-result-contradicted.json
  • references/example-call-result-flat.json
  • references/example-call-result-no-checkable.json
  • references/example-call-result-unsupported.json
  • references/example-call-result.json
  • references/example-goal.txt
  • references/examples.md
  • references/safety.md
  • scripts/post_summary_faithfulness_auditor.py
  • scripts/test_post_summary_faithfulness_auditor.py

Open the folder on GitHubat commit 38d4118

Compare with similar skills

Call Post Summary Faithfulness Auditor next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Call Post Summary Faithfulness Auditor compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Call Post Summary Faithfulness Auditor this skillCALLE-AI/awesome-phone-call-agents107—~1.9kAutomated safety check: PassMIT
Speak Summarygithub/awesome-copilot40k—~1.6kAutomated safety check: PassMIT
Experimental Designaiming-lab/AutoResearchClaw15k—~286Automated safety check: PassMIT
Pua Offlinetanweai/pua20k—~143Automated safety check: PassMIT
Indexing Issue Auditorsickn33/agentic-awesome-skills47k1 repos~1.5kAutomated safety check: PassMIT
Offline Fallbackthedaviddias/Front-End-Checklist74k—~488Automated safety check: PassMIT

Similar skills

  • Speak Summary

    github/awesome-copilot

    Official

    Convert text, markdown, or a summary produced by another skill into a listenable MP3 using local CPU-only neural text-to-speech.

    40k GitHub stars~1.6k tokensUpdated 2 days ago
    Media & CreativeAuto-check passed
  • Experimental Design

    aiming-lab/AutoResearchClaw

    Best practices for designing reproducible ML experiments. An agent skill from aiming-lab/AutoResearchClaw.

    15k GitHub stars~286 tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed
  • Pua Offline

    tanweai/pua

    PUA offline alias for Codex. An agent skill from tanweai/pua.

    20k GitHub stars~143 tokensUpdated 1 mo ago
    Auto-check passed
  • Indexing Issue Auditor

    sickn33/agentic-awesome-skills

    High-level technical SEO and site architecture auditor. An agent skill from sickn33/agentic-awesome-skills.

    47k GitHub starsUsed in 1 repo~1.5k tokens
    Marketing & SEOAuto-check passed
  • Offline Fallback

    thedaviddias/Front-End-Checklist

    A skill your agent uses when adding PWA capabilities, implementing a service worker, or improving the experience for users on unreliable network connections.

    74k GitHub stars~488 tokensUpdated 4 days ago
    Auto-check passed
  • Cost Summary

    ruvnet/ruflo

    Single-shot programmatic dump of all cost data — total spend, per-tier, top session, budget status, federation aggregate.

    74k GitHub stars~594 tokensUpdated yesterday
    Auto-check: notes

More from CALLE-AI/awesome-phone-call-agents

All 101 skills in this repo
  • Accessible Outing Verifier

    CALLE-AI/awesome-phone-call-agents

    Demonstrates advisory accessibility-planning checks with offline fixtures and a proposed bounded CALL-E workflow; use for exploring unknown or qualified venue claims without making calls.

    107 GitHub stars~1.4k tokensUpdated yesterday
    Auto-check passed
  • Ground Truth Gate

    CALLE-AI/awesome-phone-call-agents

    A skill your agent uses when an agent holds some evidence for a physical-world claim but the evidence is broader, narrower, or older than the exact question asked, and it must first decide whether a…

    107 GitHub stars~3.3k tokensUpdated yesterday
    Auto-check passed
  • Is It Accessible

    CALLE-AI/awesome-phone-call-agents

    Call a venue and ask the accessibility questions that matter to one specific person — step-free entry, hearing loop, guide dogs, quiet hours, changing places — then return a per-need verdict backed…

    107 GitHub stars~4.3k tokensUpdated yesterday
    Auto-check passed
  • Landmark Navigation Assist

    CALLE-AI/awesome-phone-call-agents

    Turns a pre-written, building-level location config into a CALL-E outbound phone-call task that guides a delivery driver through the last few hundred metres to a specific building using landmarks…

    107 GitHub stars~1.6k tokensUpdated yesterday
    Auto-check passed
  • Research Gap Call Verifier

    CALLE-AI/awesome-phone-call-agents

    Turn cited business research into a bounded, approval-gated phone-call plan that asks only unresolved factual questions, then reconcile CALL-E-compatible results without treating voicemail, refusal…

    107 GitHub stars~1.7k tokensUpdated yesterday
    Auto-check passed
  • Structured Outcome Followup Call

    CALLE-AI/awesome-phone-call-agents

    Place a goal-driven CALL-E call that collects specific structured answers, score those answers against a deterministic rubric you supply, and conditionally trigger a follow-up action — all runnable…

    107 GitHub stars~1.4k tokensUpdated yesterday
    Auto-check passed

Questions about Call Post Summary Faithfulness Auditor

What does Call Post Summary Faithfulness Auditor do?

Offline experimental CALL-E helper that decomposes the agent postsummary into atomic claims and anchors each against the transcript, flagging unsupported and outcome-contradicted statements before…. Call Post Summary Faithfulness Auditor is an agent skill from CALLE-AI/awesome-phone-call-agents. Offline experimental CALL-E helper that decomposes the agent postsummary into atomic claims and anchors each against the transcript, flagging unsupported and outcome-contradicted statements before automation acts on them, plus a faithful-summary goal template.

How do I install Call Post Summary Faithfulness Auditor in Claude Code?

Run `npx skills add CALLE-AI/awesome-phone-call-agents --skill call-post-summary-faithfulness-auditor -a claude-code`. Or copy the skill folder (skills/call-post-summary-faithfulness-auditor in CALLE-AI/awesome-phone-call-agents) into .claude/skills/call-post-summary-faithfulness-auditor in your project. Claude Code loads it when a task matches its description.

How do I install Call Post Summary Faithfulness Auditor in Codex?

Run `npx skills add CALLE-AI/awesome-phone-call-agents --skill call-post-summary-faithfulness-auditor -a codex`. Or copy the skill folder (skills/call-post-summary-faithfulness-auditor in CALLE-AI/awesome-phone-call-agents) into .agents/skills/call-post-summary-faithfulness-auditor in your project. Codex loads it when a task matches its description.

Can I use Call Post Summary Faithfulness Auditor in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add CALLE-AI/awesome-phone-call-agents --skill call-post-summary-faithfulness-auditor -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/call-post-summary-faithfulness-auditor, .gemini/skills/call-post-summary-faithfulness-auditor, .github/skills/call-post-summary-faithfulness-auditor and .opencode/skills/call-post-summary-faithfulness-auditor in your project.

What does Call Post Summary Faithfulness Auditor need to run?

Going by SKILL.md and its folder, Call Post Summary Faithfulness Auditor needs Python for the scripts in its folder and the command-line tools its instructions call (python3). Our summary lists: Python 3.

Does Call Post Summary Faithfulness Auditor access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Call Post Summary Faithfulness Auditor safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Call Post Summary Faithfulness Auditor use?

Call Post Summary Faithfulness Auditor is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Call Post Summary Faithfulness Auditor use?

About 1.9k tokens (SKILL.md is roughly 7.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 4.1k tokens, read only when the agent opens those files.

What are the alternatives to Call Post Summary Faithfulness Auditor?

Skills that share tags, products or a category with Call Post Summary Faithfulness Auditor: Speak Summary (github/awesome-copilot, 40k stars), Experimental Design (aiming-lab/AutoResearchClaw, 15k stars), Pua Offline (tanweai/pua, 20k stars) and Indexing Issue Auditor (sickn33/agentic-awesome-skills, 47k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Call Post Summary Faithfulness Auditor?

CALLE-AI (a GitHub organization) maintains it in CALLE-AI/awesome-phone-call-agents, which has 107 GitHub stars. The repository holds 101 skills in this directory. The repository was last updated on October 10, 2026.

Source: CALLE-AI/awesome-phone-call-agents on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.