Turns DOIs, PMIDs and arXiv IDs into clean BibTeX, searches Google Scholar and PubMed, and checks and deduplicates a reference list.

MITAuto-check: notesResearch & Science

Install Citation Management

skills CLI
$ npx skills add foryourhealth111-pixel/Vibe-Skills --skill citation-management -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install foryourhealth111-pixel/Vibe-Skills citation-management --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/foryourhealth111-pixel/Vibe-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/bundled/skills/citation-management .claude/skills/citation-management && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
citation-management
GitHub stars
3.6k
Token cost
~7.6k tokens
SKILL.md length
1,972 words
Files
14 (incl. scripts, references, assets)
Skills in repo
81
Repo updated
First seen
Licence
MIT

At a glance

Turns DOIs, PMIDs and arXiv IDs into clean BibTeX, searches Google Scholar and PubMed, and checks and deduplicates a reference list.

  • Works in 5 steps: Citation Source Lookup → Metadata Extraction → BibTeX Formatting → …
  • Converting a list of DOIs or PMIDs into BibTeX entries for a manuscript
  • SKILL.md covers Overview, When to Use This Skill, Core Workflow and Search Strategies, plus 4 more sections
  • Runs Python scripts from its folder; calls python and pip; reaches nature.com

What it does

This skill organizes citation work around six bundled Python scripts: `search_google_scholar.py` and `search_pubmed.py` find source records, `doi_to_bibtex.py` and `extract_metadata.py` turn identifiers into complete metadata, `format_bibtex.py` cleans BibTeX files, and `validate_citations.py` checks entries against the actual publication. Reference notes cover search syntax, metadata extraction, BibTeX formatting and validation.

Typical jobs include converting DOIs, PMIDs or arXiv IDs to BibTeX, finding duplicate references, checking citation counts for known papers, keeping formatting consistent and assembling a bibliography for a manuscript or thesis. A BibTeX template and a citation checklist ship in the assets folder. Google Scholar and PubMed lookups need network access, and the agent is allowed to read, write, edit and run shell commands.

When your agent uses it

  • Converting a list of DOIs or PMIDs into BibTeX entries for a manuscript
  • Checking an existing .bib file for wrong metadata or duplicate references
  • Looking up a known paper on Google Scholar or PubMed and capturing its record
  • Building a consistent bibliography for a thesis

Example prompts

  • “Convert these five DOIs into BibTeX and add them to refs.bib.”
  • “Validate every entry in thesis/references.bib and list the ones that do not match their publications.”
  • “Find duplicate citations in my bibliography and merge them.”
  • “Look up the original Transformer paper on arXiv and give me a properly formatted BibTeX entry.”

Requirements

  • Python 3 to run the bundled scripts
  • Network access for Google Scholar and PubMed lookups
  • Pre-approved tools (allowed-tools): Read, Write, Edit, Bash

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Citation Source Lookup
  2. Metadata Extraction
  3. BibTeX Formatting
  4. Citation Validation
  5. Integration with Writing Workflow

What it can do on your machine

Read from SKILL.md and the folder at commit ddcaa2a. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Write
    • Edit
    • Bash

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 6 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python
    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • nature.com

    Also links to:

    • meshb.nlm.nih.gov
    • pubmed.ncbi.nlm.nih.gov
    • bibtex.org
    • scholar.google.com
    • api.crossref.org
    • ncbi.nlm.nih.gov
    • arxiv.org
    • api.datacite.org
    • doi.org
    • overleaf.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Citation Management loads about 7.6k tokens when it runs, and up to ~30k if it reads all its reference files. Until then it costs about 99 tokens; SKILL.md has 1,972 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~99
When it runs · the whole SKILL.md, loaded when a task matches
~7.6k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~30k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Read, Write, Edit, Bash

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from foryourhealth111-pixel/Vibe-Skills at commit ddcaa2a, republished under its MIT licence (© foryourhealth111-pixel). 1,972 words, ~7,593 tokens.

Download SKILL.mdSave it as .claude/skills/citation-management/SKILL.md (or your agent's skills folder). This skill also uses 13 other files; get the full folder from GitHub.
name
citation-management
description
Comprehensive citation management for academic research. Resolve bibliographic identifiers, extract accurate metadata, validate citations, deduplicate references, and generate properly formatted BibTeX entries. This skill should be used when you need to verify citation information, convert DOIs/PMIDs/arXiv IDs to BibTeX, or ensure reference accuracy in scientific writing.
allowed-tools
Read, Write, Edit, Bash
license
MIT License
metadata.skill-author
K-Dense Inc.

Citation Management

Overview

Manage citations systematically throughout the research and writing process. This skill provides tools and strategies for extracting accurate metadata from identifiers and bibliographic sources, validating citation information, cleaning duplicate references, and generating properly formatted BibTeX entries.

Critical for maintaining citation accuracy, avoiding reference errors, and ensuring reproducible research.

When to Use This Skill

Use this skill when:

  • Resolving known or candidate papers from identifiers, titles, Google Scholar, or PubMed records
  • Converting DOIs, PMIDs, or arXiv IDs to properly formatted BibTeX
  • Extracting complete metadata for citations (authors, title, journal, year, etc.)
  • Validating existing citations for accuracy
  • Cleaning and formatting BibTeX files
  • Checking citation counts for known papers in a specific field
  • Verifying that citation information matches the actual publication
  • Building a bibliography for a manuscript or thesis
  • Checking for duplicate citations
  • Ensuring consistent citation formatting

Core Workflow

Citation management follows a systematic process:

Phase 1: Citation Source Lookup

Goal: Locate exact source records for known references or tightly scoped citation candidates.

Google Scholar provides the most comprehensive coverage across disciplines.

Basic Search:

bash
# Search for papers on a topic
python scripts/search_google_scholar.py "CRISPR gene editing" \
  --limit 50 \
  --output results.json

# Search with year filter
python scripts/search_google_scholar.py "machine learning protein folding" \
  --year-start 2020 \
  --year-end 2024 \
  --limit 100 \
  --output ml_proteins.json

Advanced Search Strategies (see references/google_scholar_search.md):

  • Use quotation marks for exact phrases: "deep learning"
  • Search by author: author:LeCun
  • Search in title: intitle:"neural networks"
  • Exclude terms: machine learning -survey
  • Find highly cited papers using sort options
  • Filter by date ranges to get recent work

Best Practices:

  • Use specific, targeted search terms
  • Include key technical terms and acronyms
  • Filter by recent years for fast-moving fields
  • Check "Cited by" to find seminal papers
  • Export top results for further analysis

PubMed specializes in biomedical and life sciences literature (35+ million citations).

Basic Search:

bash
# Search PubMed
python scripts/search_pubmed.py "Alzheimer's disease treatment" \
  --limit 100 \
  --output alzheimers.json

# Search with MeSH terms and filters
python scripts/search_pubmed.py \
  --query '"Alzheimer Disease"[MeSH] AND "Drug Therapy"[MeSH]' \
  --date-start 2020 \
  --date-end 2024 \
  --publication-types "Clinical Trial,Review" \
  --output alzheimers_trials.json

Advanced PubMed Queries (see references/pubmed_search.md):

  • Use MeSH terms: "Diabetes Mellitus"[MeSH]
  • Field tags: "cancer"[Title], "Smith J"[Author]
  • Boolean operators: AND, OR, NOT
  • Date filters: 2020:2024[Publication Date]
  • Publication types: "Review"[Publication Type]
  • Combine with E-utilities API for automation

Best Practices:

  • Use MeSH Browser to find correct controlled vocabulary
  • Construct complex queries in PubMed Advanced Search Builder first
  • Include multiple synonyms with OR
  • Retrieve PMIDs for easy metadata extraction
  • Export to JSON or directly to BibTeX
Phase 2: Metadata Extraction

Goal: Convert paper identifiers (DOI, PMID, arXiv ID) to complete, accurate metadata.

Quick DOI to BibTeX Conversion

For single DOIs, use the quick conversion tool:

bash
# Convert single DOI
python scripts/doi_to_bibtex.py 10.1038/s41586-021-03819-2

# Convert multiple DOIs from a file
python scripts/doi_to_bibtex.py --input dois.txt --output references.bib

# Different output formats
python scripts/doi_to_bibtex.py 10.1038/nature12345 --format json
Comprehensive Metadata Extraction

For DOIs, PMIDs, arXiv IDs, or URLs:

bash
# Extract from DOI
python scripts/extract_metadata.py --doi 10.1038/s41586-021-03819-2

# Extract from PMID
python scripts/extract_metadata.py --pmid 34265844

# Extract from arXiv ID
python scripts/extract_metadata.py --arxiv 2103.14030

# Extract from URL
python scripts/extract_metadata.py --url "https://www.nature.com/articles/s41586-021-03819-2"

# Batch extraction from file (mixed identifiers)
python scripts/extract_metadata.py --input identifiers.txt --output citations.bib

Metadata Sources (see references/metadata_extraction.md):

  1. CrossRef API: Primary source for DOIs

    • Comprehensive metadata for journal articles
    • Publisher-provided information
    • Includes authors, title, journal, volume, pages, dates
    • Free, no API key required
  2. PubMed E-utilities: Biomedical literature

    • Official NCBI metadata
    • Includes MeSH terms, abstracts
    • PMID and PMCID identifiers
    • Free, API key recommended for high volume
  3. arXiv API: Preprints in physics, math, CS, q-bio

    • Complete metadata for preprints
    • Version tracking
    • Author affiliations
    • Free, open access
  4. DataCite API: Research datasets, software, other resources

    • Metadata for non-traditional scholarly outputs
    • DOIs for datasets and code
    • Free access

What Gets Extracted:

  • Required fields: author, title, year
  • Journal articles: journal, volume, number, pages, DOI
  • Books: publisher, ISBN, edition
  • Conference papers: booktitle, conference location, pages
  • Preprints: repository (arXiv, bioRxiv), preprint ID
  • Additional: abstract, keywords, URL
Phase 3: BibTeX Formatting

Goal: Generate clean, properly formatted BibTeX entries.

Understanding BibTeX Entry Types

See references/bibtex_formatting.md for complete guide.

Common Entry Types:

  • @article: Journal articles (most common)
  • @book: Books
  • @inproceedings: Conference papers
  • @incollection: Book chapters
  • @phdthesis: Dissertations
  • @misc: Preprints, software, datasets

Required Fields by Type:

bibtex
@article{citationkey,
  author  = {Last1, First1 and Last2, First2},
  title   = {Article Title},
  journal = {Journal Name},
  year    = {2024},
  volume  = {10},
  number  = {3},
  pages   = {123--145},
  doi     = {10.1234/example}
}

@inproceedings{citationkey,
  author    = {Last, First},
  title     = {Paper Title},
  booktitle = {Conference Name},
  year      = {2024},
  pages     = {1--10}
}

@book{citationkey,
  author    = {Last, First},
  title     = {Book Title},
  publisher = {Publisher Name},
  year      = {2024}
}
Formatting and Cleaning

Use the formatter to standardize BibTeX files:

bash
# Format and clean BibTeX file
python scripts/format_bibtex.py references.bib \
  --output formatted_references.bib

# Sort entries by citation key
python scripts/format_bibtex.py references.bib \
  --sort key \
  --output sorted_references.bib

# Sort by year (newest first)
python scripts/format_bibtex.py references.bib \
  --sort year \
  --descending \
  --output sorted_references.bib

# Remove duplicates
python scripts/format_bibtex.py references.bib \
  --deduplicate \
  --output clean_references.bib

# Validate and report issues
python scripts/format_bibtex.py references.bib \
  --validate \
  --report validation_report.txt

Formatting Operations:

  • Standardize field order
  • Consistent indentation and spacing
  • Proper capitalization in titles (protected with {})
  • Standardized author name format
  • Consistent citation key format
  • Remove unnecessary fields
  • Fix common errors (missing commas, braces)
Phase 4: Citation Validation

Goal: Verify all citations are accurate and complete.

Comprehensive Validation
bash
# Validate BibTeX file
python scripts/validate_citations.py references.bib

# Validate and fix common issues
python scripts/validate_citations.py references.bib \
  --auto-fix \
  --output validated_references.bib

# Generate detailed validation report
python scripts/validate_citations.py references.bib \
  --report validation_report.json \
  --verbose

Validation Checks (see references/citation_validation.md):

  1. DOI Verification:

    • DOI resolves correctly via doi.org
    • Metadata matches between BibTeX and CrossRef
    • No broken or invalid DOIs
  2. Required Fields:

    • All required fields present for entry type
    • No empty or missing critical information
    • Author names properly formatted
  3. Data Consistency:

    • Year is valid (4 digits, reasonable range)
    • Volume/number are numeric
    • Pages formatted correctly (e.g., 123--145)
    • URLs are accessible
  4. Duplicate Detection:

    • Same DOI used multiple times
    • Similar titles (possible duplicates)
    • Same author/year/title combinations
  5. Format Compliance:

    • Valid BibTeX syntax
    • Proper bracing and quoting
    • Citation keys are unique
    • Special characters handled correctly

Validation Output:

json
{
  "total_entries": 150,
  "valid_entries": 145,
  "errors": [
    {
      "citation_key": "Smith2023",
      "error_type": "missing_field",
      "field": "journal",
      "severity": "high"
    },
    {
      "citation_key": "Jones2022",
      "error_type": "invalid_doi",
      "doi": "10.1234/broken",
      "severity": "high"
    }
  ],
  "warnings": [
    {
      "citation_key": "Brown2021",
      "warning_type": "possible_duplicate",
      "duplicate_of": "Brown2021a",
      "severity": "medium"
    }
  ]
}
Phase 5: Integration with Writing Workflow
Building References for Manuscripts

Complete workflow for creating a bibliography:

bash
# 1. Search for papers on your topic
python scripts/search_pubmed.py \
  '"CRISPR-Cas Systems"[MeSH] AND "Gene Editing"[MeSH]' \
  --date-start 2020 \
  --limit 200 \
  --output crispr_papers.json

# 2. Extract DOIs from search results and convert to BibTeX
python scripts/extract_metadata.py \
  --input crispr_papers.json \
  --output crispr_refs.bib

# 3. Add specific papers by DOI
python scripts/doi_to_bibtex.py 10.1038/nature12345 >> crispr_refs.bib
python scripts/doi_to_bibtex.py 10.1126/science.abcd1234 >> crispr_refs.bib

# 4. Format and clean the BibTeX file
python scripts/format_bibtex.py crispr_refs.bib \
  --deduplicate \
  --sort year \
  --descending \
  --output references.bib

# 5. Validate all citations
python scripts/validate_citations.py references.bib \
  --auto-fix \
  --report validation.json \
  --output final_references.bib

# 6. Review validation report and fix any remaining issues
cat validation.json

# 7. Use in your LaTeX document
# \bibliography{final_references}
Citation Validation In Review Documents

When a review document already exists, use this skill for the technical citation layer:

  1. Extract all cited identifiers and bibliography records
  2. Validate metadata against DOI, PMID, arXiv, CrossRef, or PubMed sources
  3. Normalize BibTeX keys and citation style fields
  4. Verify the final bibliography before manuscript build or submission
bash
# After completing literature review
# Verify all citations in the review document
python scripts/validate_citations.py my_review_references.bib --report review_validation.json

# Format for specific citation style if needed
python scripts/format_bibtex.py my_review_references.bib \
  --style nature \
  --output formatted_refs.bib

Search Strategies

Google Scholar Best Practices

Finding Seminal and High-Impact Papers (CRITICAL):

Always prioritize papers based on citation count, venue quality, and author reputation:

Citation Count Thresholds:

Paper AgeCitationsClassification
0-3 years20+Noteworthy
0-3 years100+Highly Influential
3-7 years100+Significant
3-7 years500+Landmark Paper
7+ years500+Seminal Work
7+ years1000+Foundational

Venue Quality Tiers:

  • Tier 1 (Prefer): Nature, Science, Cell, NEJM, Lancet, JAMA, PNAS
  • Tier 2 (High Priority): Impact Factor >10, top conferences (NeurIPS, ICML, ICLR)
  • Tier 3 (Good): Specialized journals (IF 5-10)
  • Tier 4 (Sparingly): Lower-impact peer-reviewed venues

Author Reputation Indicators:

  • Senior researchers with h-index >40
  • Multiple publications in Tier-1 venues
  • Leadership at recognized institutions
  • Awards and editorial positions

Search Strategies for High-Impact Papers:

  • Sort by citation count (most cited first)
  • Look for review articles from Tier-1 journals for overview
  • Check "Cited by" for impact assessment and recent follow-up work
  • Use citation alerts for tracking new citations to key papers
  • Filter by top venues using source:Nature or source:Science
  • Search for papers by known field leaders using author:LastName

Advanced Operators (full list in references/google_scholar_search.md):

"exact phrase"           # Exact phrase matching
author:lastname          # Search by author
intitle:keyword          # Search in title only
source:journal           # Search specific journal
-exclude                 # Exclude terms
OR                       # Alternative terms
2020..2024              # Year range

Example Searches:

# Find recent reviews on a topic
"CRISPR" intitle:review 2023..2024

# Find papers by specific author on topic
author:Church "synthetic biology"

# Find highly cited foundational work
"deep learning" 2012..2015 sort:citations

# Exclude surveys and focus on methods
"protein folding" -survey -review intitle:method
PubMed Best Practices

Using MeSH Terms: MeSH (Medical Subject Headings) provides controlled vocabulary for precise searching.

  1. Find MeSH terms at https://meshb.nlm.nih.gov/search
  2. Use in queries: "Diabetes Mellitus, Type 2"[MeSH]
  3. Combine with keywords for comprehensive coverage

Field Tags:

[Title]              # Search in title only
[Title/Abstract]     # Search in title or abstract
[Author]             # Search by author name
[Journal]            # Search specific journal
[Publication Date]   # Date range
[Publication Type]   # Article type
[MeSH]              # MeSH term

Building Complex Queries:

bash
# Clinical trials on diabetes treatment published recently
"Diabetes Mellitus, Type 2"[MeSH] AND "Drug Therapy"[MeSH] 
AND "Clinical Trial"[Publication Type] AND 2020:2024[Publication Date]

# Reviews on CRISPR in specific journal
"CRISPR-Cas Systems"[MeSH] AND "Nature"[Journal] AND "Review"[Publication Type]

# Specific author's recent work
"Smith AB"[Author] AND cancer[Title/Abstract] AND 2022:2024[Publication Date]

E-utilities for Automation: The scripts use NCBI E-utilities API for programmatic access:

  • ESearch: Search and retrieve PMIDs
  • EFetch: Retrieve full metadata
  • ESummary: Get summary information
  • ELink: Find related articles

See references/pubmed_search.md for complete API documentation.

Tools and Scripts

search_google_scholar.py

Search Google Scholar and export results.

Features:

  • Automated searching with rate limiting
  • Pagination support
  • Year range filtering
  • Export to JSON or BibTeX
  • Citation count information

Usage:

bash
# Basic search
python scripts/search_google_scholar.py "quantum computing"

# Advanced search with filters
python scripts/search_google_scholar.py "quantum computing" \
  --year-start 2020 \
  --year-end 2024 \
  --limit 100 \
  --sort-by citations \
  --output quantum_papers.json

# Export directly to BibTeX
python scripts/search_google_scholar.py "machine learning" \
  --limit 50 \
  --format bibtex \
  --output ml_papers.bib
search_pubmed.py

Search PubMed using E-utilities API.

Features:

  • Complex query support (MeSH, field tags, Boolean)
  • Date range filtering
  • Publication type filtering
  • Batch retrieval with metadata
  • Export to JSON or BibTeX

Usage:

bash
# Simple keyword search
python scripts/search_pubmed.py "CRISPR gene editing"

# Complex query with filters
python scripts/search_pubmed.py \
  --query '"CRISPR-Cas Systems"[MeSH] AND "therapeutic"[Title/Abstract]' \
  --date-start 2020-01-01 \
  --date-end 2024-12-31 \
  --publication-types "Clinical Trial,Review" \
  --limit 200 \
  --output crispr_therapeutic.json

# Export to BibTeX
python scripts/search_pubmed.py "Alzheimer's disease" \
  --limit 100 \
  --format bibtex \
  --output alzheimers.bib
extract_metadata.py

Extract complete metadata from paper identifiers.

Features:

  • Supports DOI, PMID, arXiv ID, URL
  • Queries CrossRef, PubMed, arXiv APIs
  • Handles multiple identifier types
  • Batch processing
  • Multiple output formats

Usage:

bash
# Single DOI
python scripts/extract_metadata.py --doi 10.1038/s41586-021-03819-2

# Single PMID
python scripts/extract_metadata.py --pmid 34265844

# Single arXiv ID
python scripts/extract_metadata.py --arxiv 2103.14030

# From URL
python scripts/extract_metadata.py \
  --url "https://www.nature.com/articles/s41586-021-03819-2"

# Batch processing (file with one identifier per line)
python scripts/extract_metadata.py \
  --input paper_ids.txt \
  --output references.bib

# Different output formats
python scripts/extract_metadata.py \
  --doi 10.1038/nature12345 \
  --format json  # or bibtex, yaml
validate_citations.py

Validate BibTeX entries for accuracy and completeness.

Features:

  • DOI verification via doi.org and CrossRef
  • Required field checking
  • Duplicate detection
  • Format validation
  • Auto-fix common issues
  • Detailed reporting

Usage:

bash
# Basic validation
python scripts/validate_citations.py references.bib

# With auto-fix
python scripts/validate_citations.py references.bib \
  --auto-fix \
  --output fixed_references.bib

# Detailed validation report
python scripts/validate_citations.py references.bib \
  --report validation_report.json \
  --verbose

# Only check DOIs
python scripts/validate_citations.py references.bib \
  --check-dois-only
format_bibtex.py

Format and clean BibTeX files.

Features:

  • Standardize formatting
  • Sort entries (by key, year, author)
  • Remove duplicates
  • Validate syntax
  • Fix common errors
  • Enforce citation key conventions

Usage:

bash
# Basic formatting
python scripts/format_bibtex.py references.bib

# Sort by year (newest first)
python scripts/format_bibtex.py references.bib \
  --sort year \
  --descending \
  --output sorted_refs.bib

# Remove duplicates
python scripts/format_bibtex.py references.bib \
  --deduplicate \
  --output clean_refs.bib

# Complete cleanup
python scripts/format_bibtex.py references.bib \
  --deduplicate \
  --sort year \
  --validate \
  --auto-fix \
  --output final_refs.bib
Show full SKILL.md (790 more words)Show less
doi_to_bibtex.py

Quick DOI to BibTeX conversion.

Features:

  • Fast single DOI conversion
  • Batch processing
  • Multiple output formats
  • Clipboard support

Usage:

bash
# Single DOI
python scripts/doi_to_bibtex.py 10.1038/s41586-021-03819-2

# Multiple DOIs
python scripts/doi_to_bibtex.py \
  10.1038/nature12345 \
  10.1126/science.abc1234 \
  10.1016/j.cell.2023.01.001

# From file (one DOI per line)
python scripts/doi_to_bibtex.py --input dois.txt --output references.bib

# Copy to clipboard
python scripts/doi_to_bibtex.py 10.1038/nature12345 --clipboard

Best Practices

Search Strategy
  1. Start broad, then narrow:

    • Begin with general terms to understand the field
    • Refine with specific keywords and filters
    • Use synonyms and related terms
  2. Use multiple sources:

    • Google Scholar for comprehensive coverage
    • PubMed for biomedical focus
    • arXiv for preprints
    • Combine results for completeness
  3. Leverage citations:

    • Check "Cited by" for seminal papers
    • Review references from key papers
    • Use citation networks to discover related work
  4. Document your searches:

    • Save search queries and dates
    • Record number of results
    • Note any filters or restrictions applied
Metadata Extraction
  1. Always use DOIs when available:

    • Most reliable identifier
    • Permanent link to the publication
    • Best metadata source via CrossRef
  2. Verify extracted metadata:

    • Check author names are correct
    • Verify journal/conference names
    • Confirm publication year
    • Validate page numbers and volume
  3. Handle edge cases:

    • Preprints: Include repository and ID
    • Preprints later published: Use published version
    • Conference papers: Include conference name and location
    • Book chapters: Include book title and editors
  4. Maintain consistency:

    • Use consistent author name format
    • Standardize journal abbreviations
    • Use same DOI format (URL preferred)
BibTeX Quality
  1. Follow conventions:

    • Use meaningful citation keys (FirstAuthor2024keyword)
    • Protect capitalization in titles with {}
    • Use -- for page ranges (not single dash)
    • Include DOI field for all modern publications
  2. Keep it clean:

    • Remove unnecessary fields
    • No redundant information
    • Consistent formatting
    • Validate syntax regularly
  3. Organize systematically:

    • Sort by year or topic
    • Group related papers
    • Use separate files for different projects
    • Merge carefully to avoid duplicates
Validation
  1. Validate early and often:

    • Check citations when adding them
    • Validate complete bibliography before submission
    • Re-validate after any manual edits
  2. Fix issues promptly:

    • Broken DOIs: Find correct identifier
    • Missing fields: Extract from original source
    • Duplicates: Choose best version, remove others
    • Format errors: Use auto-fix when safe
  3. Manual review for critical citations:

    • Verify key papers cited correctly
    • Check author names match publication
    • Confirm page numbers and volume
    • Ensure URLs are current

Common Pitfalls to Avoid

  1. Single source bias: Only using Google Scholar or PubMed

    • Solution: Search multiple databases for comprehensive coverage
  2. Accepting metadata blindly: Not verifying extracted information

    • Solution: Spot-check extracted metadata against original sources
  3. Ignoring DOI errors: Broken or incorrect DOIs in bibliography

    • Solution: Run validation before final submission
  4. Inconsistent formatting: Mixed citation key styles, formatting

    • Solution: Use format_bibtex.py to standardize
  5. Duplicate entries: Same paper cited multiple times with different keys

    • Solution: Use duplicate detection in validation
  6. Missing required fields: Incomplete BibTeX entries

    • Solution: Validate and ensure all required fields present
  7. Outdated preprints: Citing preprint when published version exists

    • Solution: Check if preprints have been published, update to journal version
  8. Special character issues: Broken LaTeX compilation due to characters

    • Solution: Use proper escaping or Unicode in BibTeX
  9. No validation before submission: Submitting with citation errors

    • Solution: Always run validation as final check
  10. Manual BibTeX entry: Typing entries by hand

    • Solution: Always extract from metadata sources using scripts

Example Workflows

Example 1: Building a Bibliography for a Paper
bash
# Step 1: Find key papers on your topic
python scripts/search_google_scholar.py "transformer neural networks" \
  --year-start 2017 \
  --limit 50 \
  --output transformers_gs.json

python scripts/search_pubmed.py "deep learning medical imaging" \
  --date-start 2020 \
  --limit 50 \
  --output medical_dl_pm.json

# Step 2: Extract metadata from search results
python scripts/extract_metadata.py \
  --input transformers_gs.json \
  --output transformers.bib

python scripts/extract_metadata.py \
  --input medical_dl_pm.json \
  --output medical.bib

# Step 3: Add specific papers you already know
python scripts/doi_to_bibtex.py 10.1038/s41586-021-03819-2 >> specific.bib
python scripts/doi_to_bibtex.py 10.1126/science.aam9317 >> specific.bib

# Step 4: Combine all BibTeX files
cat transformers.bib medical.bib specific.bib > combined.bib

# Step 5: Format and deduplicate
python scripts/format_bibtex.py combined.bib \
  --deduplicate \
  --sort year \
  --descending \
  --output formatted.bib

# Step 6: Validate
python scripts/validate_citations.py formatted.bib \
  --auto-fix \
  --report validation.json \
  --output final_references.bib

# Step 7: Review any issues
cat validation.json | grep -A 3 '"errors"'

# Step 8: Use in LaTeX
# \bibliography{final_references}
Example 2: Converting a List of DOIs
bash
# You have a text file with DOIs (one per line)
# dois.txt contains:
# 10.1038/s41586-021-03819-2
# 10.1126/science.aam9317
# 10.1016/j.cell.2023.01.001

# Convert all to BibTeX
python scripts/doi_to_bibtex.py --input dois.txt --output references.bib

# Validate the result
python scripts/validate_citations.py references.bib --verbose
Example 3: Cleaning an Existing BibTeX File
bash
# You have a messy BibTeX file from various sources
# Clean it up systematically

# Step 1: Format and standardize
python scripts/format_bibtex.py messy_references.bib \
  --output step1_formatted.bib

# Step 2: Remove duplicates
python scripts/format_bibtex.py step1_formatted.bib \
  --deduplicate \
  --output step2_deduplicated.bib

# Step 3: Validate and auto-fix
python scripts/validate_citations.py step2_deduplicated.bib \
  --auto-fix \
  --output step3_validated.bib

# Step 4: Sort by year
python scripts/format_bibtex.py step3_validated.bib \
  --sort year \
  --descending \
  --output clean_references.bib

# Step 5: Final validation report
python scripts/validate_citations.py clean_references.bib \
  --report final_validation.json \
  --verbose

# Review report
cat final_validation.json
Example 4: Finding and Citing Seminal Papers
bash
# Find highly cited papers on a topic
python scripts/search_google_scholar.py "AlphaFold protein structure" \
  --year-start 2020 \
  --year-end 2024 \
  --sort-by citations \
  --limit 20 \
  --output alphafold_seminal.json

# Extract the top 10 by citation count
# (script will have included citation counts in JSON)

# Convert to BibTeX
python scripts/extract_metadata.py \
  --input alphafold_seminal.json \
  --output alphafold_refs.bib

# The BibTeX file now contains the most influential papers

Citation Outputs

Citation management outputs are reusable in writing and submission workflows:

  • Export validated BibTeX for LaTeX manuscripts
  • Verify citations match publication standards
  • Format references according to journal requirements
  • Keep bibliography files deduplicated and reproducible
  • Generate properly formatted references
  • Validate citations meet venue requirements

Resources

Bundled Resources

References (in references/):

  • google_scholar_search.md: Complete Google Scholar search guide
  • pubmed_search.md: PubMed and E-utilities API documentation
  • metadata_extraction.md: Metadata sources and field requirements
  • citation_validation.md: Validation criteria and quality checks
  • bibtex_formatting.md: BibTeX entry types and formatting rules

Scripts (in scripts/):

  • search_google_scholar.py: Google Scholar search automation
  • search_pubmed.py: PubMed E-utilities API client
  • extract_metadata.py: Universal metadata extractor
  • validate_citations.py: Citation validation and verification
  • format_bibtex.py: BibTeX formatter and cleaner
  • doi_to_bibtex.py: Quick DOI to BibTeX converter

Assets (in assets/):

  • bibtex_template.bib: Example BibTeX entries for all types
  • citation_checklist.md: Quality assurance checklist
External Resources

Search Engines:

Metadata APIs:

Tools and Validators:

Citation Styles:

Dependencies

Required Python Packages
bash
# Core dependencies
pip install requests  # HTTP requests for APIs
pip install bibtexparser  # BibTeX parsing and formatting
pip install biopython  # PubMed E-utilities access

# Optional (for Google Scholar)
pip install scholarly  # Google Scholar API wrapper
# or
pip install selenium  # For more robust Scholar scraping
Optional Tools
bash
# For advanced validation
pip install crossref-commons  # Enhanced CrossRef API access
pip install pylatexenc  # LaTeX special character handling

Summary

The citation-management skill provides:

  1. Comprehensive search capabilities for Google Scholar and PubMed
  2. Automated metadata extraction from DOI, PMID, arXiv ID, URLs
  3. Citation validation with DOI verification and completeness checking
  4. BibTeX formatting with standardization and cleaning tools
  5. Quality assurance through validation and reporting
  6. Integration with scientific writing workflow
  7. Reproducibility through documented search and extraction methods

Use this skill to maintain accurate, complete citations throughout your research and ensure publication-ready bibliographies.

© foryourhealth111-pixel, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 13 other files (scripts, references, assets) in bundled/skills/citation-management of foryourhealth111-pixel/Vibe-Skills.

  • SKILL.md
  • assets/bibtex_template.bib
  • assets/citation_checklist.md
  • references/bibtex_formatting.md
  • references/citation_validation.md
  • references/google_scholar_search.md
  • references/metadata_extraction.md
  • references/pubmed_search.md
  • scripts/doi_to_bibtex.py
  • scripts/extract_metadata.py
  • scripts/format_bibtex.py
  • scripts/search_google_scholar.py
  • scripts/search_pubmed.py
  • scripts/validate_citations.py

Open the folder on GitHubat commit ddcaa2a

Compare with similar skills

Citation Management next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Citation Management compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Citation Management this skillforyourhealth111-pixel/Vibe-Skills3.6k—~7.6kAutomated safety check: NotesMIT
Citation ManagementK-Dense-AI/claude-scientific-writer2.4k2 repos~3.9kAutomated safety check: NotesMIT
Citation Managementneflibata-feng/MyArxiv-Agent12619 repos~8.1kAutomated safety check: NotesMIT
Arxiv MCP Serverblazickjp/arxiv-mcp-server3.2k—~353Automated safety check: PassApache-2.0
Hugging Face Paper Publisherhuggingface/skills11k4 repos~4.2kAutomated safety check: PassApache-2.0
Arxiv Paper Writeryunshenwuchuxun/latex-paper-skills267—~3.2kAutomated safety check: PassMIT

Similar skills

  • Citation Management

    K-Dense-AI/claude-scientific-writer

    Finds papers in OpenAlex, PubMed and Google Scholar, turns DOIs, PMIDs and arXiv IDs into clean BibTeX, and validates citations for a manuscript or thesis.

    2.4k GitHub starsUsed in 2 repos~3.9k tokens
    Research & ScienceAuto-check: notes
  • Citation Management

    neflibata-feng/MyArxiv-Agent

    Comprehensive citation management for academic research. An agent skill from neflibata-feng/MyArxiv-Agent.

    126 GitHub starsUsed in 19 repos~8.1k tokens
    Research & ScienceAuto-check: notes
  • Arxiv MCP Server

    blazickjp/arxiv-mcp-server

    A skill your agent uses when finding, comparing, reading, or monitoring arXiv papers, including requests for abstracts, citation graphs, original LaTeX, section-level technical details, or…

    3.2k GitHub stars~353 tokensUpdated 3 days ago
    Research & ScienceAuto-check passed
  • Official

    Indexes research papers on the Hugging Face Hub from arXiv, links them to models and datasets, claims authorship and generates markdown research articles from templates.

    11k GitHub starsUsed in 4 repos~4.2k tokens
    Research & ScienceAuto-check passed
  • Arxiv Paper Writer

    yunshenwuchuxun/latex-paper-skills

    Writes ML/AI review and survey papers for arXiv using the IEEEtran LaTeX template with verified BibTeX citations.

    267 GitHub stars~3.2k tokensUpdated 6 mo ago
    Research & ScienceAuto-check passed
  • Nature Academic Search

    jing1312/nature-figure-skill

    Multi-source literature search, citation verification, MeSH search strategy, citation file management (.nbib/.ris/.bib conversion), and reference management (BibTeX, related articles, ID conversion)…

    171 GitHub stars~1.3k tokensUpdated 1 mo ago
    Research & ScienceAuto-check: notes

More from foryourhealth111-pixel/Vibe-Skills

All 81 skills in this repo
  • Market Research Reports

    foryourhealth111-pixel/Vibe-Skills

    Produces long consulting-style market research and industry reports covering market sizing, competitive landscape, market entry and investment theses.

    3.6k GitHub stars~2.5k tokensUpdated 1 mo ago
    Auto-check: notes
  • Academic Venue Templates

    foryourhealth111-pixel/Vibe-Skills

    Supplies venue-specific LaTeX templates and formatting rules for journals, conferences and posters, and checks a manuscript against page limits and submission requirements.

    3.6k GitHub stars~3.9k tokensUpdated 1 mo ago
    Auto-check: notes
  • Digital Brain

    foryourhealth111-pixel/Vibe-Skills

    This skill should be used when the user asks to "write a post", "check my voice", "look up contact", "prepare for meeting", "weekly review", "track goals", or mentions personal brand, content…

    3.6k GitHub stars~1.7k tokensUpdated 1 mo ago
    Auto-check passed
  • Smart File Writer

    foryourhealth111-pixel/Vibe-Skills

    Diagnoses why a file write failed (permissions, disk space, path length, locks, read-only mounts) before retrying, instead of repeating the same call blindly.

    3.6k GitHub stars~2.6k tokensUpdated 1 mo ago
    Auto-check passed
  • Automated Video Studio

    foryourhealth111-pixel/Vibe-Skills

    Turns footage, audio and a storyboard plan into a finished short video with FFmpeg jump-cuts, subtitle burn-in and a final polish pass.

    3.6k GitHub stars~838 tokensUpdated 1 mo ago
    Auto-check passed
  • Creating Data Visualizations

    foryourhealth111-pixel/Vibe-Skills

    Create analytical charts and plots from existing data. An agent skill from foryourhealth111-pixel/Vibe-Skills.

    3.6k GitHub stars~403 tokensUpdated 1 mo ago
    Auto-check passed

Questions about Citation Management

What does Citation Management do?

Turns DOIs, PMIDs and arXiv IDs into clean BibTeX, searches Google Scholar and PubMed, and checks and deduplicates a reference list. py` checks entries against the actual publication. Reference notes cover search syntax, metadata extraction, BibTeX formatting and validation.

When should I use Citation Management?

Citation Management fits situations like: converting a list of DOIs or PMIDs into BibTeX entries for a manuscript; checking an existing .bib file for wrong metadata or duplicate references; looking up a known paper on Google Scholar or PubMed and capturing its record; building a consistent bibliography for a thesis.

How do I install Citation Management in Claude Code?

Run `npx skills add foryourhealth111-pixel/Vibe-Skills --skill citation-management -a claude-code`. Or copy the skill folder (bundled/skills/citation-management in foryourhealth111-pixel/Vibe-Skills) into .claude/skills/citation-management in your project. Claude Code loads it when a task matches its description.

How do I install Citation Management in Codex?

Run `npx skills add foryourhealth111-pixel/Vibe-Skills --skill citation-management -a codex`. Or copy the skill folder (bundled/skills/citation-management in foryourhealth111-pixel/Vibe-Skills) into .agents/skills/citation-management in your project. Codex loads it when a task matches its description.

Can I use Citation Management in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add foryourhealth111-pixel/Vibe-Skills --skill citation-management -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/citation-management, .gemini/skills/citation-management, .github/skills/citation-management and .opencode/skills/citation-management in your project.

What does Citation Management need to run?

Going by SKILL.md and its folder, Citation Management needs Python for the scripts in its folder and the command-line tools its instructions call (python and pip). Our summary lists: Python 3 to run the bundled scripts; Network access for Google Scholar and PubMed lookups. Its frontmatter pre-approves these tools: Read, Write, Edit, Bash.

Does Citation Management access the network?

SKILL.md names 11 domains. In commands or code: nature.com; the agent is likely to contact it when it follows the instructions. As links in the text: meshb.nlm.nih.gov, pubmed.ncbi.nlm.nih.gov, bibtex.org, scholar.google.com, api.crossref.org, ncbi.nlm.nih.gov, arxiv.org, api.datacite.org, doi.org and overleaf.com. This is read from the text; nothing was executed.

Is Citation Management safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Citation Management use?

Citation Management is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Citation Management use?

About 7.6k tokens (SKILL.md is roughly 30k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 22k tokens, read only when the agent opens those files.

What are the alternatives to Citation Management?

Skills that share tags, products or a category with Citation Management: Citation Management (K-Dense-AI/claude-scientific-writer, 2.4k stars), Citation Management (neflibata-feng/MyArxiv-Agent, 126 stars), Arxiv MCP Server (blazickjp/arxiv-mcp-server, 3.2k stars) and Hugging Face Paper Publisher (huggingface/skills, 11k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Citation Management?

foryourhealth111-pixel (a GitHub user) maintains it in foryourhealth111-pixel/Vibe-Skills, which has 3,627 GitHub stars. The repository holds 81 skills in this directory. The repository was last updated on August 31, 2026.

Source: foryourhealth111-pixel/Vibe-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.