Agent skill

Paper Research on arXiv

by XiaomiMiMo in XiaomiMiMo/MiMo-Code

Searches arXiv, fetches metadata, generates BibTeX, downloads PDFs and finds citations and related papers using a bundled Python script.

MITAuto-check passedResearch & Science

Install Paper Research on arXiv

skills CLI
$ npx skills add XiaomiMiMo/MiMo-Code --skill arxiv -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install XiaomiMiMo/MiMo-Code arxiv --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/XiaomiMiMo/MiMo-Code.git skills-src && mkdir -p .claude/skills && cp -r skills-src/packages/cli/src/skill/builtin/.bundle/arxiv .claude/skills/arxiv && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
arxiv
GitHub stars
14k
Used in
1 other repo
Token cost
~1.5k tokens
SKILL.md length
546 words
Files
2 (incl. scripts)
Skills in repo
22
Repo updated
First seen
Licence
MIT

At a glance

Searches arXiv, fetches metadata, generates BibTeX, downloads PDFs and finds citations and related papers using a bundled Python script.

  • Works in 5 steps: search "topic" --sort date --max 15 —… → search "topic" --max 15 — seminal work… → Cross-check impact: python… → …
  • Finding recent papers on a topic or from a given author
  • SKILL.md covers Quick Start: the arxiv.py script, Reading Paper Content, Recommended Research Workflows and Raw API Reference (when the…, plus 1 more section
  • Runs Python scripts from its folder; calls python and curl; reaches arxiv.org and api.semanticscholar.org

What it does

The skill bundles scripts/arxiv.py, which uses the free arXiv and Semantic Scholar APIs with no keys and only the Python standard library. Subcommands cover search with author, category and sort filters, full metadata and abstracts for one or several IDs, BibTeX output, PDF download to a chosen folder, the latest submissions in a category, who cites a paper, what it cites, and similar papers. Adding the json flag gives machine-readable output.

The script handles Atom XML parsing, versioned IDs and withdrawn-paper detection. To read a paper, the agent fetches the abstract page, the HTML version or the PDF; if only a PDF is available it downloads it and hands it to a PDF skill. Suggested workflows include a literature review that combines recent and relevance-sorted searches, checks citation counts and collects BibTeX, a deep dive on one paper through its references, citations and similar works, and staying current in a field.

When your agent uses it

  • Finding recent papers on a topic or from a given author
  • Generating BibTeX for papers you plan to cite
  • Checking who cites a paper and what it builds on
  • Listing the latest submissions in a category such as cs.CL
  • Downloading a paper's PDF for local processing

Example prompts

  • “Search arXiv for recent GRPO reinforcement learning papers, newest first.”
  • “Give me BibTeX for 1706.03762.”
  • “What is new in cs.CL this week?”
  • “Who cites 2601.02780, and what does it reference?”

Requirements

  • Python 3, standard library only
  • Network access to the arXiv and Semantic Scholar APIs

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. search "topic" --sort date --max 15 — recent work
  2. search "topic" --max 15 — seminal work (relevance-sorted)
  3. Cross-check impact: python scripts/arxiv.py get ID then cites ID --max 5 for citation counts
  4. Read the top candidates via webfetch on the abs/html URLs
  5. bibtex ID1,ID2,... for the papers you keep

What it can do on your machine

Read from SKILL.md and the folder at commit 6babeb0. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python
    • curl

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • arxiv.org
    • api.semanticscholar.org
    • export.arxiv.org

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Paper Research on arXiv loads about 1.5k tokens when it runs. Until then it costs about 167 tokens; SKILL.md has 546 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~167
When it runs · the whole SKILL.md, loaded when a task matches
~1.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from XiaomiMiMo/MiMo-Code at commit 6babeb0, republished under its MIT licence (© XiaomiMiMo). 546 words, ~1,503 tokens.

Download SKILL.mdSave it as .claude/skills/arxiv/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
arxiv
description
Use this skill whenever the user wants to find, read, cite, track, download, or analyze academic papers on arXiv. That includes: searching papers by topic, author, category, or arXiv ID; fetching abstracts or full metadata; generating BibTeX citations; downloading PDFs; listing the latest submissions in a field (e.g. cs.AI daily digest); checking a paper's citation impact; finding who cites a paper, what it references, or related-paper recommendations. Trigger on mentions of 'arXiv', an arXiv ID (e.g. 2601.02780 or hep-th/0601001), an arxiv.org URL, 'paper search', 'literature review', 'find papers about X', 'cite this paper', or 'what's new in cs.LG'.
version
2.0.0
license
MIT
platforms
linux, macos, windows

arXiv Research

Search, read, cite, and analyze academic papers using the free arXiv API and Semantic Scholar API. No API keys, no dependencies — the bundled script uses only the Python stdlib.

Quick Start: the arxiv.py script

Prefer scripts/arxiv.py over raw curl — it handles Atom XML parsing, versioned IDs, withdrawn-paper detection, and produces clean readable (or --json) output.

GoalCommand
Search paperspython scripts/arxiv.py search "GRPO reinforcement learning" --max 10 --sort date
Filtered searchpython scripts/arxiv.py search "attention" --author vaswani --category cs.CL
Full metadata + abstractpython scripts/arxiv.py get 2601.02780,1706.03762
BibTeX citationpython scripts/arxiv.py bibtex 2601.02780
Download PDFpython scripts/arxiv.py download 2601.02780 --dest ./papers
Latest in a categorypython scripts/arxiv.py new cs.CL --max 10
Who cites this paperpython scripts/arxiv.py cites 2601.02780 --max 20
What this paper citespython scripts/arxiv.py refs 2601.02780
Related-paper recommendationspython scripts/arxiv.py similar 2601.02780
Machine-readable outputappend --json to any command

Common flags: --max N (result count), --sort relevance|date|updated, --start N (pagination offset, search only), --json.

Reading Paper Content

After finding a paper, read it with the webfetch tool:

  • Abstract page (fast, metadata + abstract): https://arxiv.org/abs/2601.02780
  • Full paper as HTML (best for reading, when available): https://arxiv.org/html/2601.02780
  • PDF: https://arxiv.org/pdf/2601.02780

If HTML is unavailable and the PDF must be processed locally, download it first, then use a PDF-processing skill.

Literature review on a topic

  1. search "topic" --sort date --max 15 — recent work
  2. search "topic" --max 15 — seminal work (relevance-sorted)
  3. Cross-check impact: python scripts/arxiv.py get ID then cites ID --max 5 for citation counts
  4. Read the top candidates via webfetch on the abs/html URLs
  5. bibtex ID1,ID2,... for the papers you keep

Deep-dive a single paper

  1. get ID — full abstract, versions, journal ref, DOI
  2. refs ID — what it builds on
  3. cites ID — follow-up work (sorted by citation count)
  4. similar ID — related papers you might have missed
  5. webfetch the HTML/PDF for full text

Stay current in a field

  • new cs.AI --max 20 — latest submissions in a category
  • Category taxonomy: https://arxiv.org/category_taxonomy — common ones: cs.AI, cs.CL (NLP), cs.CV, cs.LG, cs.CR, stat.ML, math.OC
Show full SKILL.md (207 more words)Show less

Raw API Reference (when the script isn't enough)

The script covers most needs; use the raw APIs for advanced queries.

arXiv API (Atom XML)
bash
curl -s "https://export.arxiv.org/api/query?search_query=all:transformer&max_results=5"

Field prefixes: all: (everything), ti: (title), au: (author), abs: (abstract), cat: (category), co: (comment, e.g. co:accepted+NeurIPS).

Boolean syntax (URL-encode spaces as +):

all:GPT+OR+all:BERT              # OR
all:language+model+ANDNOT+all:vision   # AND NOT
ti:"chain+of+thought"            # exact phrase
au:hinton+AND+cat:cs.LG          # combined

Parameters: sortBy (relevance|lastUpdatedDate|submittedDate), sortOrder, start, max_results, id_list (comma-separated IDs).

Semantic Scholar API (JSON)

arXiv has no citation data — Semantic Scholar fills that gap (free, ~1 req/sec unauthenticated).

bash
# Paper details with citation counts
curl -s "https://api.semanticscholar.org/graph/v1/paper/arXiv:2601.02780?fields=title,citationCount,influentialCitationCount,tldr"

# Author profile
curl -s "https://api.semanticscholar.org/graph/v1/author/search?query=Yann+LeCun&fields=name,hIndex,citationCount,paperCount"

# Keyword search returning JSON (alternative to arXiv search)
curl -s "https://api.semanticscholar.org/graph/v1/paper/search?query=GRPO&limit=5&fields=title,year,citationCount,externalIds"

Useful fields: title, authors, year, abstract, tldr (AI summary), citationCount, influentialCitationCount, isOpenAccess, openAccessPdf, fieldsOfStudy, publicationVenue, externalIds (arXiv ID, DOI).

Important Details

Rate limits — arXiv: ~1 request / 3 seconds; Semantic Scholar: ~1 request / second. Space out consecutive calls; the script exits with a clear message on HTTP 429.

ID formats — new style 2402.03300, old style hep-th/0601001. Both work everywhere.

Versioning — arxiv.org/abs/2601.02780 resolves to the latest version; ...v2 is a specific immutable version. bibtex and download preserve the version suffix so citations match the content you actually read (later versions can change substantially).

Withdrawn papers — the script flags entries whose abstract indicates withdrawal/retraction with [WITHDRAWN]. Don't cite these without noting the status.

Listings caveat — new CATEGORY sorts by submission date via the search API; the official "new today" listing at https://arxiv.org/list/cs.AI/new may group slightly differently (cross-lists, replacements).

© XiaomiMiMo, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (scripts) in packages/cli/src/skill/builtin/.bundle/arxiv of XiaomiMiMo/MiMo-Code.

  • SKILL.md
  • scripts/arxiv.py

Open the folder on GitHubat commit 6babeb0

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in XiaomiMiMo/MiMo-Code, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Paper Research on arXiv next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Paper Research on arXiv compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Paper Research on arXiv this skillXiaomiMiMo/MiMo-Code14k1 repos~1.5kAutomated safety check: PassMIT
Literature ReviewK-Dense-AI/scientific-agent-skills48k1 repos~3.2kAutomated safety check: NotesMIT
Literature Reviewneflibata-feng/MyArxiv-Agent12621 repos~5.9kAutomated safety check: NotesMIT
Literature Review AgentAr9av/PaperOrchestra6772 repos~5.2kAutomated safety check: PassCustom licence
Paper AutoratersAr9av/PaperOrchestra6772 repos~1.6kAutomated safety check: PassCustom licence
Arxiv MCP Serverblazickjp/arxiv-mcp-server3.2k—~353Automated safety check: PassApache-2.0

Similar skills

  • Literature Review

    K-Dense-AI/scientific-agent-skills

    Runs systematic, scoping or narrative literature reviews across PubMed, arXiv, bioRxiv and Semantic Scholar, with citation checks and Markdown or PDF output.

    48k GitHub starsUsed in 1 repo~3.2k tokens
    Research & ScienceAuto-check: notes
  • Literature Review

    neflibata-feng/MyArxiv-Agent

    Conduct comprehensive, systematic literature reviews using multiple academic databases (PubMed, arXiv, bioRxiv, Semantic Scholar, etc.).

    126 GitHub starsUsed in 21 repos~5.9k tokens
    Research & ScienceAuto-check: notes
  • Literature Review Agent

    Ar9av/PaperOrchestra

    Step 3 of the PaperOrchestra pipeline (arXiv:2604.05018). An agent skill from Ar9av/PaperOrchestra.

    677 GitHub starsUsed in 2 repos~5.2k tokens
    Research & ScienceAuto-check passed
  • Paper Autoraters

    Ar9av/PaperOrchestra

    Run the four paper-quality autoraters from PaperOrchestra (arXiv:2604.05018, App.

    677 GitHub starsUsed in 2 repos~1.6k tokens
    Research & ScienceAuto-check passed
  • Arxiv MCP Server

    blazickjp/arxiv-mcp-server

    A skill your agent uses when finding, comparing, reading, or monitoring arXiv papers, including requests for abstracts, citation graphs, original LaTeX, section-level technical details, or…

    3.2k GitHub stars~353 tokensUpdated 2 days ago
    Research & ScienceAuto-check passed
  • PaperSeek Literature Search

    MingfengHong/paperseek

    Routes literature searches through the PaperSeek launcher, picks suitable scholarly sources, parses JSON output and keeps API keys out of the chat.

    200 GitHub stars~1.6k tokensUpdated 3 days ago
    Research & ScienceAuto-check passed

More from XiaomiMiMo/MiMo-Code

All 22 skills in this repo
  • Agent Skill Creator

    XiaomiMiMo/MiMo-Code

    Interactive guide for creating, reviewing and fixing agent skills (SKILL.md folders), covering structure, frontmatter rules, trigger phrases and validation before sharing.

    14k GitHub stars~1.9k tokensUpdated 5 days ago
    Auto-check passed
  • DOCX Toolkit

    XiaomiMiMo/MiMo-Code

    Produces, edits and reads Microsoft Word files through python-docx and lxml, with a decision table for picking the lightest workflow for a given task.

    14k GitHub stars~2.4k tokensUpdated 5 days ago
    Auto-check passed
  • Drive MiMo Code

    XiaomiMiMo/MiMo-Code

    Lets one MiMoCode process drive another, headless with JSON events or interactively through tmux, to test behavior and visual regressions with parseable evidence.

    14k GitHub stars~3.9k tokensUpdated 5 days ago
    Auto-check passed
  • PDF Toolkit

    XiaomiMiMo/MiMo-Code

    Reads, transforms, composes and fills PDFs with Python scripts for extraction, merging, watermarking, encryption, OCR and form filling.

    14k GitHub stars~1.7k tokensUpdated 5 days ago
    Auto-check passed
  • XLSX Spreadsheet Toolkit

    XiaomiMiMo/MiMo-Code

    Builds, edits, cleans, recalculates and reads Excel workbooks and CSV files with openpyxl and pandas, plus LibreOffice for recalculation and PDF export.

    14k GitHub stars~2.9k tokensUpdated 5 days ago
    Auto-check passed
  • Claude Code Delegation

    XiaomiMiMo/MiMo-Code

    Hands coding work to the Claude Code CLI from the terminal in print, interactive tmux or background mode, only when you explicitly ask for Claude Code.

    14k GitHub stars~1.3k tokensUpdated 5 days ago
    Auto-check passed

Questions about Paper Research on arXiv

What does Paper Research on arXiv do?

Searches arXiv, fetches metadata, generates BibTeX, downloads PDFs and finds citations and related papers using a bundled Python script. py, which uses the free arXiv and Semantic Scholar APIs with no keys and only the Python standard library. Subcommands cover search with author, category and sort filters, full metadata and abstracts for one or several IDs, BibTeX output, PDF download to a chosen folder, the latest submissions in a category, who cites a paper, what it cites, and similar papers.

When should I use Paper Research on arXiv?

Paper Research on arXiv fits situations like: finding recent papers on a topic or from a given author; generating BibTeX for papers you plan to cite; checking who cites a paper and what it builds on; listing the latest submissions in a category such as cs.CL.

How do I install Paper Research on arXiv in Claude Code?

Run `npx skills add XiaomiMiMo/MiMo-Code --skill arxiv -a claude-code`. Or copy the skill folder (packages/cli/src/skill/builtin/.bundle/arxiv in XiaomiMiMo/MiMo-Code) into .claude/skills/arxiv in your project. Claude Code loads it when a task matches its description.

How do I install Paper Research on arXiv in Codex?

Run `npx skills add XiaomiMiMo/MiMo-Code --skill arxiv -a codex`. Or copy the skill folder (packages/cli/src/skill/builtin/.bundle/arxiv in XiaomiMiMo/MiMo-Code) into .agents/skills/arxiv in your project. Codex loads it when a task matches its description.

Can I use Paper Research on arXiv in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add XiaomiMiMo/MiMo-Code --skill arxiv -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/arxiv, .gemini/skills/arxiv, .github/skills/arxiv and .opencode/skills/arxiv in your project.

What does Paper Research on arXiv need to run?

Going by SKILL.md and its folder, Paper Research on arXiv needs Python for the scripts in its folder and the command-line tools its instructions call (python and curl). Our summary lists: Python 3, standard library only; Network access to the arXiv and Semantic Scholar APIs.

Does Paper Research on arXiv access the network?

SKILL.md names 3 domains. In commands or code: arxiv.org, api.semanticscholar.org and export.arxiv.org; the agent is likely to contact these when it follows the instructions. This is read from the text; nothing was executed.

Is Paper Research on arXiv safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Paper Research on arXiv use?

Paper Research on arXiv is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Paper Research on arXiv use?

About 1.5k tokens (SKILL.md is roughly 6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Paper Research on arXiv?

Skills that share tags, products or a category with Paper Research on arXiv: Literature Review (K-Dense-AI/scientific-agent-skills, 48k stars), Literature Review (neflibata-feng/MyArxiv-Agent, 126 stars), Literature Review Agent (Ar9av/PaperOrchestra, 677 stars) and Paper Autoraters (Ar9av/PaperOrchestra, 677 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Paper Research on arXiv?

XiaomiMiMo (a GitHub organization) maintains it in XiaomiMiMo/MiMo-Code, which has 13,611 GitHub stars. The repository holds 22 skills in this directory. The repository was last updated on October 3, 2026.

Source: XiaomiMiMo/MiMo-Code on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.