Agent skill

Citation Audit

by Mathews-Tom in Mathews-Tom/armory

Verify that every citation in a manuscript is real, correctly attributed, and accurately described.

MITAuto-check passedResearch & Science

Install Citation Audit

skills CLI
$ npx skills add Mathews-Tom/armory --skill citation-audit -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Mathews-Tom/armory citation-audit --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Mathews-Tom/armory.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/citation-audit .claude/skills/citation-audit && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
citation-audit
GitHub stars
329
Token cost
~2.5k tokens
SKILL.md length
1,200 words
Files
2
Skills in repo
80
Repo updated
First seen
Licence
MIT

At a glance

Verify that every citation in a manuscript is real, correctly attributed, and accurately described.

  • Works in 4 steps: Extract citation contexts → Verify bib entry metadata → Verify claim accuracy → …
  • : check my citations
  • SKILL.md covers Purpose, Why This Exists, Inputs and Execution, plus 5 more sections
  • Reaches arxiv.org

What it does

Citation Audit is an agent skill from Mathews-Tom/armory. Verify that every citation in a manuscript is real, correctly attributed, and accurately described. Detects ghost papers, wrong arXiv IDs, inverted claims, and dead links by fetching each cited work. Optional fix mode applies bib metadata corrections and surfaces prose rewrites for claim errors. Triggers on: "check my citations", "verify references", "citation audit", "are my references real", "check bib", "reference check", "bib audit", "citation verification". Companion to manuscript-review (Pass 5 hygiene)…

Its SKILL.md is about 2.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files (for example `evals/cases.yaml`).

It sits in Research & Science, covering Citation management and Academic paper search. It works with arXiv. The repository describes itself as: Curated, production-grade skills for AI coding agents. Battle-tested workflows for developers who use AI seriously. The licence is MIT.

When your agent uses it

  • : check my citations
  • Verify references
  • Are my references real
  • Reference check

Example prompts

  • “check my citations”
  • “verify references”
  • “citation audit”
  • “/citation-audit”

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Extract citation contexts
  2. Verify bib entry metadata
  3. Verify claim accuracy
  4. Check for missing citations

What it can do on your machine

Read from SKILL.md and the folder at commit 4594fb7. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • arxiv.org

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Citation Audit loads about 2.5k tokens when it runs. Until then it costs about 141 tokens; SKILL.md has 1,200 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~141
When it runs · the whole SKILL.md, loaded when a task matches
~2.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Mathews-Tom/armory at commit 4594fb7, republished under its MIT licence (© Mathews-Tom). 1,200 words, ~2,464 tokens.

Download SKILL.mdSave it as .claude/skills/citation-audit/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
citation-audit
description
Verify that every citation in a manuscript is real, correctly attributed, and accurately described. Detects ghost papers, wrong arXiv IDs, inverted claims, and dead links by fetching each cited work. Optional fix mode applies bib metadata corrections and surfaces prose rewrites for claim errors. Triggers on: "check my citations", "verify references", "citation audit", "are my references real", "check bib", "reference check", "bib audit", "citation verification". Companion to manuscript-review (Pass 5 hygiene); this skill audits factual truth.
metadata.version
1.0.0
metadata.complements
manuscript-review, research-critique, literature-review
metadata.category
review
metadata.tags
citations, bibliography, fact-check, hallucination, references
metadata.difficulty
advanced
metadata.phase
review

Citation Audit Skill

Purpose

Verify every citation in a manuscript against its actual source. LLMs hallucinate citations, invent arXiv IDs, misattribute findings, and confuse authors. This skill catches all of that by fetching and reading each cited work.

Why This Exists

LLMs are unreliable with citations in three distinct ways:

  1. Ghost papers — The paper does not exist. Title, authors, or venue are fabricated.
  2. Wrong metadata — The paper exists but the bib entry has the wrong arXiv ID, wrong authors, wrong year, or wrong venue.
  3. Inverted claims — The paper exists and the bib is correct, but the manuscript mischaracterizes what the paper says.

All three are invisible to structural audits (cross-reference checks, compilation tests). They require reading the actual cited work.

Inputs

  • The manuscript .tex file(s)
  • The .bib file
  • Web access (to fetch papers from arXiv, conference sites, URLs)

Execution

Phase 1: Extract citation contexts

For each \citep{}, \citet{}, \cite{} in the manuscript:

  1. Record the bib key
  2. Record the surrounding sentence or paragraph (the claim context)
  3. Classify the claim type:
    • FACTUAL: "X et al. found Y" / "X et al. measured Y"
    • METHODOLOGICAL: "We follow X" / "We use the benchmark from X"
    • POSITIONAL: "Unlike X, we..." / "X does not measure..."
    • PARENTHETICAL: "(X, 2024)" — no specific claim, just a reference
  4. For FACTUAL and POSITIONAL claims, extract the specific assertion the manuscript makes about the cited work
Phase 2: Verify bib entry metadata

For each bib entry, verify against the actual source:

For arXiv papers (eprint field present):

  1. Fetch https://arxiv.org/abs/{eprint_id}
  2. Compare: title, authors, year
  3. If the fetched paper has a DIFFERENT title/authors than the bib entry, this is a WRONG ID or GHOST PAPER

For conference/journal papers (booktitle or journal field):

  1. Search for the paper by title + author on the web
  2. Verify: venue, year, author list
  3. If the paper cannot be found at the stated venue, flag as UNVERIFIABLE or GHOST PAPER

For web resources (howpublished with URL):

  1. Fetch the URL
  2. Verify it loads and the content matches the described resource
  3. If the URL is dead or redirects to unrelated content, flag as DEAD LINK

For each entry, check:

  • Paper exists (reachable via arXiv, DOI, URL, or web search)
  • Title matches (exact or near-exact)
  • Authors match (at least first author correct)
  • Year matches
  • Venue matches (if applicable)
  • Entry type appropriate (@inproceedings for conferences, @article for journals, @misc for preprints/blogs)
Phase 3: Verify claim accuracy

For each FACTUAL or POSITIONAL claim:

  1. Read the cited paper (abstract + relevant sections at minimum)
  2. Compare the manuscript's claim against what the paper actually says
  3. Classify:
    • ACCURATE — The claim faithfully represents the cited work
    • INACCURATE — The claim mischaracterizes the cited work
    • INVERTED — The claim says the opposite of what the paper found
    • OVERCLAIMED — The claim is stronger than what the paper supports
    • UNDERCLAIMED — The cited work supports a stronger claim than stated
    • UNVERIFIABLE — Cannot access the paper to verify

For INACCURATE and INVERTED findings, provide:

  • What the manuscript claims
  • What the cited paper actually says
  • The specific section/page of the cited paper that contradicts the claim
  • A suggested correction
Phase 4: Check for missing citations

Scan the manuscript for:

  1. Claims that cite no source but should (empirical claims without attribution)
  2. Tools, benchmarks, or datasets mentioned by name without citation
  3. Methods described as "standard" or "well-known" that have a canonical citation

Output Format

Per-citation report
### [bib_key] — [VERDICT]

**Bib entry:** [title] by [authors] ([year])
**Actual paper:** [actual title] by [actual authors] ([actual year])
**Metadata match:** title [✓/✗] | authors [✓/✗] | year [✓/✗] | venue [✓/✗]

**Claim in manuscript (line N):** "[exact text]"
**What the paper actually says:** "[summary of actual finding]"
**Claim accuracy:** [ACCURATE / INACCURATE / INVERTED / OVERCLAIMED / UNDERCLAIMED]

**Fix required:** [description of what needs to change, or "None"]
Summary table
| Bib Key | Exists | Metadata | Claim | Verdict |
|---------|--------|----------|-------|---------|
| key1    | ✓      | ✓        | ✓     | PASS    |
| key2    | ✓      | ✗        | ✗     | FAIL    |
| key3    | ✗      | —        | —     | GHOST   |
Verdict categories
  • PASS — Paper exists, metadata correct, claims accurate
  • METADATA — Paper exists, bib entry has errors (wrong ID, wrong authors, wrong year)
  • CLAIM — Paper exists, metadata correct, but manuscript mischaracterizes it
  • GHOST — Paper does not exist as described
  • DEAD — URL/link is broken
  • UNVERIFIABLE — Cannot access the paper to verify

Severity

  • CRITICAL: GHOST papers, INVERTED claims
  • HIGH: Wrong arXiv IDs, wrong authors, INACCURATE claims
  • MEDIUM: Wrong year, wrong venue, OVERCLAIMED
  • LOW: Missing citations, incomplete bib entries, UNDERCLAIMED

Important notes

  • NEVER trust your own knowledge of papers. ALWAYS fetch and verify. Your training data contains hallucinated citations. The only way to verify is to read the actual source.
  • For arXiv papers, always fetch the abstract page to confirm the paper exists and matches.
  • For conference papers, search DBLP, ACM DL, or the conference site.
  • WebFetch and WebSearch are your primary tools. Do not skip verification because a citation "looks right."
  • Blog posts and documentation URLs change. Always check that the URL still works and points to the described content.
  • When a bib entry has both an eprint (arXiv ID) and a booktitle (venue), verify both independently.
Show full SKILL.md (448 more words)Show less

Phase 5: Fix (when invoked with "fix" or "on")

When the user invokes with an argument containing "fix" or "on", execute Phases 1–4 as above, then apply fixes for every non-PASS citation.

What to auto-fix (no user confirmation needed)

These are mechanical corrections with a single correct answer:

METADATA errors (paper exists, bib entry wrong):

  • Wrong arXiv ID → replace eprint with the correct ID
  • Wrong authors → replace with authors from the actual paper
  • Wrong year → replace with year from the actual paper
  • Wrong title → replace with title from the actual paper
  • Wrong venue → replace with venue from the actual paper
  • Wrong entry type → change @misc/@inproceedings as appropriate

DEAD links:

  • URL redirects → update howpublished URL to the final destination
  • URL 404 but resource found at different URL → update URL
  • URL 404 and resource gone → flag as HUMAN-REQUIRED

Minor author corrections:

  • Misspelled author names → fix spelling
  • Missing authors from author list → add them
  • Collective author name where individual names are available → replace (keep collective name as a note if it is how the group identifies)
What requires HUMAN-REQUIRED decision

Present these and wait for the user:

GHOST papers:

  • Paper does not exist at all → present options: (a) Replace with a real paper that makes the same point (b) Remove the citation and adjust the prose (c) The user knows the paper exists and provides the correct reference

INVERTED or INACCURATE claims:

  • The manuscript says X about a paper that actually says Y → present:
    • What the manuscript claims
    • What the paper actually says
    • A suggested rewrite of the prose that accurately represents the paper
    • Whether the paper still supports the manuscript's argument (and how)
    • Let the user decide the final wording

Dead URLs with no replacement found:

  • Blog post / resource deleted with no archive or alternative
Fix procedure
  1. Apply all auto-fixes to the .bib file
  2. For each HUMAN-REQUIRED item, present the options clearly
  3. After user decisions, apply prose changes to the .tex file
  4. Verify: re-read the .bib and .tex to confirm all fixes applied
  5. Update the audit report: mark each finding as [FIXED], [RESOLVED], or [DEFERRED]
Safety rules
  • NEVER invent a replacement citation. If a ghost paper needs replacing, search for real papers that make the cited point. Present candidates to the user with abstracts. Let the user choose.
  • NEVER change the manuscript's argument. If an inverted claim needs fixing, present the rewrite as a suggestion, not an edit.
  • NEVER remove a citation without user confirmation, even if it is a ghost paper. The user may know something you do not.
  • When fixing URLs, always verify the new URL loads and contains the expected content before writing it.

Save report as

[name]-citation-audit.md in the manuscript directory.

© Mathews-Tom, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in skills/citation-audit of Mathews-Tom/armory.

  • SKILL.md
  • evals/cases.yaml

Open the folder on GitHubat commit 4594fb7

Compare with similar skills

Citation Audit next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Citation Audit compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Citation Audit this skillMathews-Tom/armory329—~2.5kAutomated safety check: PassMIT
Literature Reviewneflibata-feng/MyArxiv-Agent12620 repos~5.9kAutomated safety check: NotesMIT
Openalex Databaseneflibata-feng/MyArxiv-Agent12612 repos~3kAutomated safety check: PassCustom licence
Citation ManagementK-Dense-AI/claude-scientific-writer2.4k2 repos~3.9kAutomated safety check: NotesMIT
Citation Managementneflibata-feng/MyArxiv-Agent12619 repos~8.1kAutomated safety check: NotesMIT
Systematic Literature Review Builderbytedance/deer-flow84k2 repos~4.3kAutomated safety check: PassMIT

Similar skills

  • Literature Review

    neflibata-feng/MyArxiv-Agent

    Conduct comprehensive, systematic literature reviews using multiple academic databases (PubMed, arXiv, bioRxiv, Semantic Scholar, etc.).

    126 GitHub starsUsed in 20 repos~5.9k tokens
    Research & ScienceAuto-check: notes
  • Openalex Database

    neflibata-feng/MyArxiv-Agent

    Query and analyze scholarly literature using the OpenAlex database.

    126 GitHub starsUsed in 12 repos~3k tokens
    Research & ScienceAuto-check passed
  • Citation Management

    K-Dense-AI/claude-scientific-writer

    Finds papers in OpenAlex, PubMed and Google Scholar, turns DOIs, PMIDs and arXiv IDs into clean BibTeX, and validates citations for a manuscript or thesis.

    2.4k GitHub starsUsed in 2 repos~3.9k tokens
    Research & ScienceAuto-check: notes
  • Citation Management

    neflibata-feng/MyArxiv-Agent

    Comprehensive citation management for academic research. An agent skill from neflibata-feng/MyArxiv-Agent.

    126 GitHub starsUsed in 19 repos~8.1k tokens
    Research & ScienceAuto-check: notes
  • Searches arXiv across many papers on one topic, extracts each paper's methodology and findings in parallel, and synthesizes a cited literature review.

    84k GitHub starsUsed in 2 repos~4.3k tokens
    Research & ScienceAuto-check passed
  • Paper2code

    PrathamLearnsToCode/paper2code

    Converts an arxiv paper into a minimal, citation-anchored Python implementation.

    1.5k GitHub stars~1.3k tokensUpdated 6 mo ago
    Research & ScienceAuto-check passed

More from Mathews-Tom/armory

All 80 skills in this repo
  • Architecture Reviewer

    Mathews-Tom/armory

    Architecture reviews across 7 dimensions (structural, scalability, enterprise readiness, performance, security, ops, data) with scored reports.

    329 GitHub stars~4.6k tokensUpdated 4 days ago
    Auto-check passed
  • Concept To Image

    Mathews-Tom/armory

    Turn concepts into static HTML visuals exported as PNG or SVG files via HTML/CSS/SVG.

    329 GitHub stars~2.6k tokensUpdated 4 days ago
    Auto-check passed
  • Watch

    Mathews-Tom/armory

    A skill your agent uses when analyzing an existing video URL or local recording: "watch this video", "analyze youtube video", "summarize this video", "youtube transcript", "find this moment", "what…

    329 GitHub stars~2.8k tokensUpdated 4 days ago
    Auto-check passed
  • Code Refiner

    Mathews-Tom/armory

    Deep code simplification and refactoring preserving behavior across Python, Go, TypeScript, Rust.

    329 GitHub stars~3.1k tokensUpdated 4 days ago
    Auto-check passed
  • Concept To Video

    Mathews-Tom/armory

    Turn concepts into animated explainer videos using Manim (Python) with MP4/GIF output, audio overlay, multi-scene composition.

    329 GitHub stars~4.9k tokensUpdated 4 days ago
    Auto-check passed
  • Decision Map

    Mathews-Tom/armory

    Maps the unresolved architecture, policy, and scope decisions that must be answered before planning can start: one durable decision ticket per question on the issue tracker, typed and blocker-linked…

    329 GitHub stars~2.7k tokensUpdated 4 days ago
    Auto-check passed

Works with

Questions about Citation Audit

What does Citation Audit do?

Verify that every citation in a manuscript is real, correctly attributed, and accurately described. Citation Audit is an agent skill from Mathews-Tom/armory. Verify that every citation in a manuscript is real, correctly attributed, and accurately described.

When should I use Citation Audit?

Citation Audit fits situations like: : check my citations; verify references; are my references real; reference check.

How do I install Citation Audit in Claude Code?

Run `npx skills add Mathews-Tom/armory --skill citation-audit -a claude-code`. Or copy the skill folder (skills/citation-audit in Mathews-Tom/armory) into .claude/skills/citation-audit in your project. Claude Code loads it when a task matches its description.

How do I install Citation Audit in Codex?

Run `npx skills add Mathews-Tom/armory --skill citation-audit -a codex`. Or copy the skill folder (skills/citation-audit in Mathews-Tom/armory) into .agents/skills/citation-audit in your project. Codex loads it when a task matches its description.

Can I use Citation Audit in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Mathews-Tom/armory --skill citation-audit -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/citation-audit, .gemini/skills/citation-audit, .github/skills/citation-audit and .opencode/skills/citation-audit in your project.

What does Citation Audit need to run?

SKILL.md names no scripts, command-line tools or credentials: Citation Audit is instructions for the agent only.

Does Citation Audit access the network?

SKILL.md names 1 domain. In commands or code: arxiv.org; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is Citation Audit safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Citation Audit use?

Citation Audit is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Citation Audit use?

About 2.5k tokens (SKILL.md is roughly 9.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Citation Audit?

Skills that share tags, products or a category with Citation Audit: Literature Review (neflibata-feng/MyArxiv-Agent, 126 stars), Openalex Database (neflibata-feng/MyArxiv-Agent, 126 stars), Citation Management (K-Dense-AI/claude-scientific-writer, 2.4k stars) and Citation Management (neflibata-feng/MyArxiv-Agent, 126 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Citation Audit?

Mathews-Tom (a GitHub user) maintains it in Mathews-Tom/armory, which has 329 GitHub stars. The repository holds 80 skills in this directory. The repository was last updated on October 6, 2026.

Source: Mathews-Tom/armory on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.