Agent skill

Literature Statistics

by aipoch in aipoch/medical-research-skills

Generate statistics for publication-year and journal distributions from local references or PDFs; use when you need standardized Year/Journal tables and a summary without any network access.

MITAuto-check passedDocuments & Office

Install Literature Statistics

skills CLI
$ npx skills add aipoch/medical-research-skills --skill literature-statistics -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install aipoch/medical-research-skills literature-statistics --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/aipoch/medical-research-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/scientific-skills/Other/literature-statistics .claude/skills/literature-statistics && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
literature-statistics
GitHub stars
1.9k
Token cost
~1.9k tokens
SKILL.md length
846 words
Files
5 (incl. scripts, references)
Skills in repo
578
Repo updated
First seen
Licence
MIT

At a glance

Generate statistics for publication-year and journal distributions from local references or PDFs; use when you need standardized Year/Journal tables and a summary without any network access.

  • Works in 3 steps: Process a local PDF directory → Process a local reference file (example… → Expected output format (Markdown)
  • You need standardized Year/Journal tables and a summary without any network access
  • SKILL.md covers When to Use, Key Features, Dependencies and Example Usage, plus 8 more sections
  • Runs Python scripts from its folder; calls python and pip

What it does

Literature Statistics is an agent skill from aipoch/medical-research-skills. Generate statistics for publication-year and journal distributions from local references or PDFs; use when you need standardized Year/Journal tables and a summary without any network access.

Its SKILL.md is about 1.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files, including scripts and reference files (for example `literature-statistics_audit_result_v2.json`, `references/examples.md` and `scripts/process_pdfs.py`).

It sits in Documents & Office, covering Statistics, Citation management and PDF. It works with LaTeX. The repository describes itself as: Hundreds of agent skills for medical research, including protocol design, data analysis, evidence insights, and academic writing. The licence is MIT.

When your agent uses it

  • You need standardized Year/Journal tables and a summary without any network access
  • Tasks that involve Statistics
  • Tasks that involve Citation management

Example prompts

  • “/literature-statistics”

Requirements

  • Python 3

Workflow steps

3 steps, taken from the step headings in SKILL.md.

  1. Process a local PDF directory
  2. Process a local reference file (example pattern)
  3. Expected output format (Markdown)

What it can do on your machine

Read from SKILL.md and the folder at commit 686e09d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 2 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python
    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Literature Statistics loads about 1.9k tokens when it runs, and up to ~2.3k if it reads all its reference files. Until then it costs about 53 tokens; SKILL.md has 846 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~53
When it runs · the whole SKILL.md, loaded when a task matches
~1.9k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~2.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from aipoch/medical-research-skills at commit 686e09d, republished under its MIT licence (© aipoch). 846 words, ~1,884 tokens.

Download SKILL.mdSave it as .claude/skills/literature-statistics/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
literature-statistics
description
Generate statistics for publication-year and journal distributions from local references or PDFs; use when you need standardized Year/Journal tables and a summary without any network access.
license
MIT
author
AIPOCH

Source: https://github.com/aipoch/medical-research-skills

When to Use

  • You have a batch of references and need a publication year distribution table (counts and percentages).
  • You need a journal distribution table (Top N optional) for a literature review or report appendix.
  • Your input is pasted citations (BibTeX/RIS/EndNote/plain text/mixed) and you want quick aggregation.
  • Your input is local reference files (.bib/.ris/.txt/.csv) and you want consistent, standardized output.
  • You have a local PDF folder and want to extract year/journal signals (best-effort) and summarize them.

Key Features

  • Supports multiple input types: pasted text, local reference files, and local PDF directories (via script).
  • Extracts Year and Journal using format-specific parsing rules (BibTeX/RIS/plain text/PDF).
  • Produces two standardized tables:
    • Year distribution: year, count, percent
    • Journal distribution: journal title, count, percent
  • Provides a summary including totals and unknown-field counts (unknown year / unknown journal).
  • Conservative extraction: does not guess when metadata is unclear; ambiguous items are counted as unknown.
  • Local-only operation: no network calls, no external APIs, no credential usage.

Dependencies

  • Python 3.9+
  • Python packages (pinned by your project file):
    • pip install -r scripts/requirements.txt

Example Usage

1) Process a local PDF directory
bash
python scripts/process_pdfs.py --input-dir "./pdfs" --output "./literature_stats.md"
2) Process a local reference file (example pattern)

If your repository provides a CLI entry or script for reference files, run it similarly to the PDF script. For example:

bash
python scripts/process_references.py --input "./refs/library.bib" --output "./literature_stats.md"
3) Expected output format (Markdown)
md
## Summary
- Total processed: 120
- Unknown year: 7
- Unknown journal: 15

## Year Distribution
| Year | Count | Percent |
|------|-------|---------|
| 2023 | 18    | 15.0%   |
| 2022 | 22    | 18.3%   |
| ...  | ...   | ...     |

## Journal Distribution
| Journal | Count | Percent |
|---------|-------|---------|
| Journal of X | 9 | 7.5% |
| ...         | ... | ... |

For additional examples, see: references/examples.md.

Implementation Details

Processing Pipeline
  1. Detect input type: pasted text / file path / PDF directory.
  2. Read content from pasted text or local files.
  3. Split into individual citations using format cues:
    • BibTeX entries
    • RIS records
    • blank-line separation for plain text/mixed inputs
  4. Extract year and journal using the parsing rules below.
  5. Normalize journal names using the normalization rules below.
  6. Aggregate counts and compute percentages.
  7. Output:
    • Table 1: Year distribution
    • Table 2: Journal distribution
    • Summary: totals + unknown counts
  8. For PDF directories, use:
    bash
    python scripts/process_pdfs.py --input-dir "<pdf_dir>" --output "<output_md>"
Parsing Rules
BibTeX
  • Year: year field
  • Journal: journal field
RIS
  • Year: PY or Y1 (use the first 4-digit year)
  • Journal: first non-empty value among JO / JF / T2
Plain Text / Mixed Citations
  • Year: first 4-digit year in the range 1900-2099 found near the end of the citation
  • Journal: infer only when patterns are unambiguous (e.g., Journal Name. 2022; or Journal Name, 2022); otherwise set to unknown
PDF Directory (Script-Based)
  • Year: prefer PDF metadata; otherwise use the first 4-digit year found on the first page
  • Journal: prefer PDF metadata; otherwise scan first-page lines containing keywords such as:
    • Journal, Proceedings, Transactions If unclear, set to unknown.
Journal Normalization Rules
  • Trim leading/trailing whitespace.
  • Collapse multiple spaces into a single space.
  • Remove trailing periods and commas.
  • If casing is inconsistent, convert to Title Case; otherwise keep original casing.
  • Do not expand abbreviations or infer aliases.
Failure Handling and Safety Constraints
  • Do not guess missing/unclear year or journal values.
  • Count ambiguous entries as unknown and report the totals in the summary.
  • No network access; no external APIs; no credentials.
  • Do not read files outside the user-provided paths.
Sorting and Reporting Requirements
  • Tables are sorted by:
    1. count descending
    2. then by name ascending (year or journal title)
  • Always report:
    • total processed count
    • unknown year count
    • unknown journal count
Show full SKILL.md (330 more words)Show less

When Not to Use

  • Do not use this skill when the required source data, identifiers, files, or credentials are missing.
  • Do not use this skill when the user asks for fabricated results, unsupported claims, or out-of-scope conclusions.
  • Do not use this skill when a simpler direct answer is more appropriate than the documented workflow.

Required Inputs

  • A clearly specified task goal aligned with the documented scope.
  • All required files, identifiers, parameters, or environment variables before execution.
  • Any domain constraints, formatting requirements, and expected output destination if applicable.
  1. Validate the request against the skill boundary and confirm all required inputs are present.
  2. Select the documented execution path and prefer the simplest supported command or procedure.
  3. Produce the expected output using the documented file format, schema, or narrative structure.
  4. Run a final validation pass for completeness, consistency, and safety before returning the result.

Output Contract

  • Return a structured deliverable that is directly usable without reformatting.
  • If a file is produced, prefer a deterministic output name such as literature_statistics_result.md unless the skill documentation defines a better convention.
  • Include a short validation summary describing what was checked, what assumptions were made, and any remaining limitations.

Validation and Safety Rules

  • Validate required inputs before execution and stop early when mandatory fields or files are missing.
  • Do not fabricate measurements, references, findings, or conclusions that are not supported by the provided source material.
  • Emit a clear warning when credentials, privacy constraints, safety boundaries, or unsupported requests affect the result.
  • Keep the output safe, reproducible, and within the documented scope at all times.

Failure Handling

  • If validation fails, explain the exact missing field, file, or parameter and show the minimum fix required.
  • If an external dependency or script fails, surface the command path, likely cause, and the next recovery step.
  • If partial output is returned, label it clearly and identify which checks could not be completed.

Quick Validation

Run this minimal verification path before full execution when possible:

bash
python scripts/process_pdfs.py --help

Expected output format:

text
Result file: literature_statistics_result.md
Validation summary: PASS/FAIL with brief notes
Assumptions: explicit list if any

© aipoch, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files (scripts, references) in scientific-skills/Other/literature-statistics of aipoch/medical-research-skills.

  • SKILL.md
  • literature-statistics_audit_result_v2.json
  • references/examples.md
  • scripts/process_pdfs.py
  • scripts/requirements.txt

Open the folder on GitHubat commit 686e09d

Compare with similar skills

Literature Statistics next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Literature Statistics compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Literature Statistics this skillaipoch/medical-research-skills1.9k—~1.9kAutomated safety check: PassMIT
Evidence Ledgerwanshuiyin/Anti-Autoresearch160—~11kAutomated safety check: NotesMIT
Paper Summarybrycewang-stanford/Auto-Empirical-Research-Skills4.6k—~3kAutomated safety check: PassCustom licence
Paper Auditbahayonghang/academic-writing-skills500—~5.1kAutomated safety check: PassNone
Paper Reading ZhMrGeDiao/paper-reading-zh103—~1.3kAutomated safety check: PassMIT
Pdf2texCalix-L/awesome-latex-skills181—~1.4kAutomated safety check: PassMIT

Similar skills

  • Evidence Ledger

    wanshuiyin/Anti-Autoresearch

    Build the deterministic evidence ledger (artifactmanifest.json + claims.json) that every other Anti-Autoresearch auditor reads.

    160 GitHub stars~11k tokensUpdated 4 days ago
    Documents & OfficeAuto-check: notes
  • Paper Summary

    brycewang-stanford/Auto-Empirical-Research-Skills

    Writes expert academic paper summaries for social science research, particularly political science and applied statistics.

    4.6k GitHub stars~3k tokensUpdated 5 days ago
    Research & ScienceAuto-check passed
  • Paper Audit

    bahayonghang/academic-writing-skills

    Reviewer-style audit and submission gate for academic papers in .tex, .typ, or .pdf.

    500 GitHub stars~5.1k tokensUpdated 12 days ago
    Documents & OfficeAuto-check passed
  • Paper Reading Zh

    MrGeDiao/paper-reading-zh

    中文论文深读、总结/TL;DR、工程拆解、按论文实现/复现、比较与证据审计。用户给出论文 PDF、链接、标题、摘要、正文或图表,或明确要读论文时使用。只有论文锚点且无附言时澄清阅读目标;只有阅读意图时问哪篇。已有论文时,“看看这篇”直接深读。不用于仅翻译、解释单个术语、生成 BibTeX、找或下载论文。

    103 GitHub stars~1.3k tokensUpdated 7 days ago
    Research & ScienceAuto-check passed
  • Pdf2tex

    Calix-L/awesome-latex-skills

    Reconstruct editable LaTeX from PDF content using page-aware extraction and visual comparison.

    181 GitHub stars~1.4k tokensUpdated 3 days ago
    Documents & OfficeAuto-check passed
  • 01 Paper Review

    agentscope-ai/OpenJudge

    Review academic papers for correctness, quality, and novelty using OpenJudge's multi-stage pipeline.

    871 GitHub stars~2.4k tokensUpdated 29 days ago
    Research & ScienceAuto-check passed

More from aipoch/medical-research-skills

All 578 skills in this repo
  • Academic Poster Generator

    aipoch/medical-research-skills

    Complete workflow for generating academic research posters from PDF literature; use when you need to extract paper content from PDFs and produce a LaTeX-based poster…

    1.9k GitHub stars~2.2k tokensUpdated 23 days ago
    Auto-check passed
  • Diagnostic Study Quality Assessment Quadas

    aipoch/medical-research-skills

    Analyzes clinical diagnostic accuracy studies for bias using the QUADAS-2 tool.

    1.9k GitHub stars~1.4k tokensUpdated 23 days ago
    Auto-check passed
  • Exploratory Data Analysis

    aipoch/medical-research-skills

    Perform comprehensive exploratory data analysis on scientific data files across 200+ file formats.

    1.9k GitHub stars~3.7k tokensUpdated 23 days ago
    Auto-check passed
  • Iso Certification

    aipoch/medical-research-skills

    A toolkit for preparing ISO 13485:2016 certification documentation for medical device QMS.

    1.9k GitHub stars~1.8k tokensUpdated 23 days ago
    Auto-check passed
  • Journal Skills

    aipoch/medical-research-skills

    Recommends target journals for manuscript submission by analyzing the paper topic/abstract and the journal distribution of similar PubMed literature; use when users ask for journal…

    1.9k GitHub stars~1.7k tokensUpdated 23 days ago
    Auto-check passed
  • Latex Posters

    aipoch/medical-research-skills

    Creates academic-poster writing packages for LaTeX using beamerposter, tikzposter, or baposter.

    1.9k GitHub stars~1.3k tokensUpdated 23 days ago
    Auto-check passed

Works with

Questions about Literature Statistics

What does Literature Statistics do?

Generate statistics for publication-year and journal distributions from local references or PDFs; use when you need standardized Year/Journal tables and a summary without any network access. Literature Statistics is an agent skill from aipoch/medical-research-skills. Generate statistics for publication-year and journal distributions from local references or PDFs; use when you need standardized Year/Journal tables and a summary without any network access.

When should I use Literature Statistics?

Literature Statistics fits situations like: you need standardized Year/Journal tables and a summary without any network access; tasks that involve Statistics; tasks that involve Citation management.

How do I install Literature Statistics in Claude Code?

Run `npx skills add aipoch/medical-research-skills --skill literature-statistics -a claude-code`. Or copy the skill folder (scientific-skills/Other/literature-statistics in aipoch/medical-research-skills) into .claude/skills/literature-statistics in your project. Claude Code loads it when a task matches its description.

How do I install Literature Statistics in Codex?

Run `npx skills add aipoch/medical-research-skills --skill literature-statistics -a codex`. Or copy the skill folder (scientific-skills/Other/literature-statistics in aipoch/medical-research-skills) into .agents/skills/literature-statistics in your project. Codex loads it when a task matches its description.

Can I use Literature Statistics in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add aipoch/medical-research-skills --skill literature-statistics -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/literature-statistics, .gemini/skills/literature-statistics, .github/skills/literature-statistics and .opencode/skills/literature-statistics in your project.

What does Literature Statistics need to run?

Going by SKILL.md and its folder, Literature Statistics needs Python for the scripts in its folder and the command-line tools its instructions call (python and pip). Our summary lists: Python 3.

Does Literature Statistics access the network?

SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Literature Statistics safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Literature Statistics use?

Literature Statistics is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Literature Statistics use?

About 1.9k tokens (SKILL.md is roughly 7.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 457 tokens, read only when the agent opens those files.

What are the alternatives to Literature Statistics?

Skills that share tags, products or a category with Literature Statistics: Evidence Ledger (wanshuiyin/Anti-Autoresearch, 160 stars), Paper Summary (brycewang-stanford/Auto-Empirical-Research-Skills, 4.6k stars), Paper Audit (bahayonghang/academic-writing-skills, 500 stars) and Paper Reading Zh (MrGeDiao/paper-reading-zh, 103 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Literature Statistics?

aipoch (a GitHub organization) maintains it in aipoch/medical-research-skills, which has 1,937 GitHub stars. The repository holds 578 skills in this directory. The repository was last updated on September 17, 2026.

Source: aipoch/medical-research-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.