Agent skill

Arxiv

by AlexAI-MCP in AlexAI-MCP/hermes-CCC

Search and retrieve academic papers from arXiv using their free REST API.

MITAuto-check passedResearch & Science

Install Arxiv

skills CLI
$ npx skills add AlexAI-MCP/hermes-CCC --skill arxiv -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install AlexAI-MCP/hermes-CCC arxiv --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/AlexAI-MCP/hermes-CCC.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/arxiv .claude/skills/arxiv && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
arxiv
GitHub stars
135
Token cost
~2k tokens
SKILL.md length
602 words
Files
1
Skills in repo
44
Repo updated
First seen
Licence
MIT

At a glance

Search and retrieve academic papers from arXiv using their free REST API.

  • Works in 6 steps: Search broadly with all: and a category. → Narrow with ti: if the result set is… → Pull the top 5 to 20 entries. → …
  • Tasks that involve Academic paper search
  • SKILL.md covers Purpose, Base Endpoint, Core Record Shape and Useful Categories, plus 20 more sections
  • Calls curl; reaches export.arxiv.org and arxiv.org

What it does

Arxiv is an agent skill from AlexAI-MCP/hermes-CCC. Search and retrieve academic papers from arXiv using their free REST API. No API key needed.

Its SKILL.md is about 2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Research & Science, covering Academic paper search. It works with arXiv. The repository describes itself as: Hermes Agent ported to Claude Code Channel — 46 native skills, no OAuth, no external process. The licence is MIT.

When your agent uses it

  • Tasks that involve Academic paper search

Example prompts

  • “/arxiv”

Requirements

  • Python 3

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Search broadly with all: and a category.
  2. Narrow with ti: if the result set is noisy.
  3. Pull the top 5 to 20 entries.
  4. Parse title, summary, and categories.
  5. Keep the paper id for citation and later retrieval.
  6. Download only shortlisted PDFs.

What it can do on your machine

Read from SKILL.md and the folder at commit 8107e89. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • curl

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • export.arxiv.org
    • arxiv.org
    • w3.org

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Arxiv loads about 2k tokens when it runs. Until then it costs about 25 tokens; SKILL.md has 602 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~25
When it runs · the whole SKILL.md, loaded when a task matches
~2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from AlexAI-MCP/hermes-CCC at commit 8107e89, republished under its MIT licence (© AlexAI-MCP). 602 words, ~2,034 tokens.

Download SKILL.mdSave it as .claude/skills/arxiv/SKILL.md (or your agent's skills folder).
name
arxiv
description
Search and retrieve academic papers from arXiv using their free REST API. No API key needed.
version
1.0.0
author
hermes-CCC (ported from Hermes Agent by NousResearch)
license
MIT

arXiv

Purpose

  • Use this skill to search, inspect, and download academic papers from arXiv.
  • Prefer it for literature review, paper triage, and reproducible paper retrieval.
  • arXiv access is free and does not require an API key.
  • The primary API is Atom XML over HTTP.

Base Endpoint

  • Base URL: https://export.arxiv.org/api/query
  • Response format: Atom XML feed
  • Transport: HTTP GET
  • Authentication: none

Core Record Shape

  • id: canonical entry URL for the paper
  • title: paper title
  • summary: abstract text
  • authors: ordered author list
  • published: original publication timestamp
  • categories: arXiv subject tags
  • pdf_url: direct PDF URL when present

Useful Categories

  • cs.AI
  • cs.LG
  • cs.CL
  • stat.ML
  • cs.CV

Search Syntax

  • Search terms are passed in search_query=...
  • Field prefixes narrow the query:
  • ti: title search
  • au: author search
  • abs: abstract search
  • all: broad metadata search
  • Combine terms with AND, OR, and ANDNOT
  • URL-encode spaces as + or %20

Common Query Patterns

  • all:reasoning+AND+cat:cs.AI
  • ti:transformer+AND+cat:cs.CL
  • au:Goodfellow+AND+cat:cs.LG
  • abs:diffusion+AND+cat:cs.CV
  • all:reinforcement+learning+AND+cat:stat.ML

CLI Subcommands

  • /arxiv search
  • /arxiv get
  • /arxiv recent
  • /arxiv download

Subcommand Intent

  • /arxiv search: run a query and print matching entries
  • /arxiv get: fetch one known arXiv identifier and display parsed metadata
  • /arxiv recent: list recent submissions in a category or topic
  • /arxiv download: save the PDF locally from a known identifier or PDF URL

Search Example With curl

bash
curl -s "https://export.arxiv.org/api/query?search_query=all:large+language+models+AND+cat:cs.CL&start=0&max_results=5&sortBy=relevance&sortOrder=descending"

Recent Example With curl

bash
curl -s "https://export.arxiv.org/api/query?search_query=cat:cs.LG&start=0&max_results=10&sortBy=submittedDate&sortOrder=descending"

PDF Download Example With curl

bash
curl -L "https://arxiv.org/pdf/2401.01234.pdf" -o 2401.01234.pdf

Fetch A Single Entry By Identifier

bash
curl -s "https://export.arxiv.org/api/query?id_list=2401.01234"

Rate Limit Guidance

  • Keep requests at or below 3 req/sec
  • Sleep between loops when paginating large result sets
  • Cache parsed results locally if you will revisit them
  • Avoid hammering the endpoint with concurrent workers

Python Parsing Pattern

python
import xml.etree.ElementTree as ET
from urllib.request import urlopen

API_URL = "https://export.arxiv.org/api/query?search_query=all:reasoning+AND+cat:cs.AI&start=0&max_results=3"
NS = {"atom": "http://www.w3.org/2005/Atom"}

with urlopen(API_URL) as response:
    xml_bytes = response.read()

root = ET.fromstring(xml_bytes)

for entry in root.findall("atom:entry", NS):
    paper_id = entry.findtext("atom:id", default="", namespaces=NS)
    title = entry.findtext("atom:title", default="", namespaces=NS).strip()
    summary = entry.findtext("atom:summary", default="", namespaces=NS).strip()
    published = entry.findtext("atom:published", default="", namespaces=NS)
    authors = [
        author.findtext("atom:name", default="", namespaces=NS)
        for author in entry.findall("atom:author", NS)
    ]
    categories = [node.attrib.get("term", "") for node in entry.findall("atom:category", NS)]
    pdf_url = ""
    for link in entry.findall("atom:link", NS):
        if link.attrib.get("title") == "pdf":
            pdf_url = link.attrib.get("href", "")
            break
    print("id:", paper_id)
    print("title:", title)
    print("published:", published)
    print("authors:", ", ".join(authors))
    print("categories:", ", ".join(categories))
    print("pdf_url:", pdf_url)
    print("summary:", summary[:240], "...")
    print("-" * 60)

Minimal Search Helper

python
import xml.etree.ElementTree as ET
from urllib.parse import quote_plus
from urllib.request import urlopen

def search_arxiv(query: str, max_results: int = 5) -> list[dict]:
    encoded = quote_plus(query)
    url = (
        "https://export.arxiv.org/api/query"
        f"?search_query={encoded}&start=0&max_results={max_results}"
    )
    ns = {"atom": "http://www.w3.org/2005/Atom"}
    with urlopen(url) as response:
        root = ET.fromstring(response.read())
    rows = []
    for entry in root.findall("atom:entry", ns):
        authors = [
            node.findtext("atom:name", default="", namespaces=ns)
            for node in entry.findall("atom:author", ns)
        ]
        categories = [node.attrib.get("term", "") for node in entry.findall("atom:category", ns)]
        rows.append(
            {
                "id": entry.findtext("atom:id", default="", namespaces=ns),
                "title": entry.findtext("atom:title", default="", namespaces=ns).strip(),
                "summary": entry.findtext("atom:summary", default="", namespaces=ns).strip(),
                "authors": authors,
                "published": entry.findtext("atom:published", default="", namespaces=ns),
                "categories": categories,
            }
        )
    return rows

for row in search_arxiv("all:multimodal AND cat:cs.CV", max_results=3):
    print(row["title"])

Example AI And ML Queries

  • all:chain-of-thought AND cat:cs.AI
  • all:instruction tuning AND cat:cs.CL
  • all:reward modeling AND cat:cs.LG
  • all:vision transformer AND cat:cs.CV
  • all:mixture of experts AND cat:stat.ML
  • ti:retrieval augmented generation
  • abs:alignment AND cat:cs.AI

Practical Retrieval Workflow

  1. Search broadly with all: and a category.
  2. Narrow with ti: if the result set is noisy.
  3. Pull the top 5 to 20 entries.
  4. Parse title, summary, and categories.
  5. Keep the paper id for citation and later retrieval.
  6. Download only shortlisted PDFs.

Sorting And Pagination

  • start controls the starting offset
  • max_results controls page size
  • sortBy=relevance is useful for topic search
  • sortBy=submittedDate is useful for recent monitoring
  • sortOrder=descending is typical for recent feeds
Show full SKILL.md (231 more words)Show less

ID And PDF Notes

  • arXiv IDs may appear as modern identifiers like 2401.01234
  • older identifiers can include subject prefixes
  • PDF links usually resolve as https://arxiv.org/pdf/<id>.pdf
  • the canonical entry page remains useful for metadata stability

Bulk Access Notes

  • Use the API for ordinary search and retrieval tasks.
  • For large-scale corpus access, arXiv also publishes bulk data and S3-style access paths in some workflows.
  • Bulk S3 access is appropriate for offline indexing, not for interactive one-off queries.
  • If you need many thousands of records, prefer bulk snapshots over high-frequency API pagination.

Failure Handling

  • If the feed is empty, print the final URL and query string first.
  • If XML parsing fails, save the raw response before retrying.
  • If pdf_url is missing, synthesize it from the parsed identifier.
  • If you get throttled, back off and reduce concurrency.

Good Defaults

  • Start with max_results=5
  • Use sortBy=relevance for topic search
  • Use sortBy=submittedDate for /arxiv recent
  • Keep summaries trimmed in terminal output
  • Persist parsed metadata as JSON if you will cite or compare papers later

Output Contract

  • Always print id
  • Always print title
  • Always print summary
  • Always print authors
  • Always print published
  • Always print categories
  • Always print pdf_url if available

When Not To Use This Skill

  • Do not use this as your only citation source for final camera-ready metadata.
  • Do not assume arXiv category tags equal venue topics.
  • Do not use the interactive API for massive mirror-scale ingestion jobs.

© AlexAI-MCP, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/arxiv of AlexAI-MCP/hermes-CCC.

Open the folder on GitHubat commit 8107e89

Compare with similar skills

Arxiv next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Arxiv compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Arxiv this skillAlexAI-MCP/hermes-CCC135—~2kAutomated safety check: PassMIT
Read arXiv Paperkarpathy/nanochat59k1 repos~494Automated safety check: PassMIT
Literature Reviewneflibata-feng/MyArxiv-Agent12620 repos~5.9kAutomated safety check: NotesMIT
Openalex Databaseneflibata-feng/MyArxiv-Agent12612 repos~3kAutomated safety check: PassCustom licence
Citation ManagementK-Dense-AI/claude-scientific-writer2.4k2 repos~3.9kAutomated safety check: NotesMIT
Citation Managementneflibata-feng/MyArxiv-Agent12619 repos~8.1kAutomated safety check: NotesMIT

Similar skills

  • Read arXiv Paper

    karpathy/nanochat

    Fetches the TeX source of an arXiv paper from its URL, reads it and writes a markdown summary tied to the nanochat project.

    59k GitHub starsUsed in 1 repo~494 tokens
    Research & ScienceAuto-check passed
  • Literature Review

    neflibata-feng/MyArxiv-Agent

    Conduct comprehensive, systematic literature reviews using multiple academic databases (PubMed, arXiv, bioRxiv, Semantic Scholar, etc.).

    126 GitHub starsUsed in 20 repos~5.9k tokens
    Research & ScienceAuto-check: notes
  • Openalex Database

    neflibata-feng/MyArxiv-Agent

    Query and analyze scholarly literature using the OpenAlex database.

    126 GitHub starsUsed in 12 repos~3k tokens
    Research & ScienceAuto-check passed
  • Citation Management

    K-Dense-AI/claude-scientific-writer

    Finds papers in OpenAlex, PubMed and Google Scholar, turns DOIs, PMIDs and arXiv IDs into clean BibTeX, and validates citations for a manuscript or thesis.

    2.4k GitHub starsUsed in 2 repos~3.9k tokens
    Research & ScienceAuto-check: notes
  • Citation Management

    neflibata-feng/MyArxiv-Agent

    Comprehensive citation management for academic research. An agent skill from neflibata-feng/MyArxiv-Agent.

    126 GitHub starsUsed in 19 repos~8.1k tokens
    Research & ScienceAuto-check: notes
  • Searches arXiv across many papers on one topic, extracts each paper's methodology and findings in parallel, and synthesizes a cited literature review.

    84k GitHub starsUsed in 2 repos~4.3k tokens
    Research & ScienceAuto-check passed

More from AlexAI-MCP/hermes-CCC

All 44 skills in this repo
  • GitHub Code Review

    AlexAI-MCP/hermes-CCC

    Review GitHub pull requests with a findings-first engineering mindset.

    135 GitHub stars~1.3k tokensUpdated 6 mo ago
    Auto-check passed
  • GitHub PR Workflow

    AlexAI-MCP/hermes-CCC

    Run a disciplined GitHub pull request workflow from branch creation through merge.

    135 GitHub stars~1.4k tokensUpdated 6 mo ago
    Auto-check passed
  • Hermes Memory

    AlexAI-MCP/hermes-CCC

    Manage durable project memory for Claude Code. An agent skill from AlexAI-MCP/hermes-CCC.

    135 GitHub stars~1.7k tokensUpdated 6 mo ago
    Auto-check passed
  • Hermes Route

    AlexAI-MCP/hermes-CCC

    Route Claude Code work by complexity, risk, and tool needs. An agent skill from AlexAI-MCP/hermes-CCC.

    135 GitHub stars~1.9k tokensUpdated 6 mo ago
    Auto-check passed
  • Hermes Skill

    AlexAI-MCP/hermes-CCC

    Create, improve, inventory, and audit Claude Code skills. An agent skill from AlexAI-MCP/hermes-CCC.

    135 GitHub stars~1.7k tokensUpdated 6 mo ago
    Auto-check passed
  • Hermes Traj

    AlexAI-MCP/hermes-CCC

    Capture Claude Code interaction trajectories in training-friendly formats.

    135 GitHub stars~1.6k tokensUpdated 6 mo ago
    Auto-check passed

Works with

Questions about Arxiv

What does Arxiv do?

Search and retrieve academic papers from arXiv using their free REST API. Arxiv is an agent skill from AlexAI-MCP/hermes-CCC. Search and retrieve academic papers from arXiv using their free REST API.

When should I use Arxiv?

Arxiv fits situations like: tasks that involve Academic paper search.

How do I install Arxiv in Claude Code?

Run `npx skills add AlexAI-MCP/hermes-CCC --skill arxiv -a claude-code`. Or copy the skill folder (skills/arxiv in AlexAI-MCP/hermes-CCC) into .claude/skills/arxiv in your project. Claude Code loads it when a task matches its description.

How do I install Arxiv in Codex?

Run `npx skills add AlexAI-MCP/hermes-CCC --skill arxiv -a codex`. Or copy the skill folder (skills/arxiv in AlexAI-MCP/hermes-CCC) into .agents/skills/arxiv in your project. Codex loads it when a task matches its description.

Can I use Arxiv in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add AlexAI-MCP/hermes-CCC --skill arxiv -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/arxiv, .gemini/skills/arxiv, .github/skills/arxiv and .opencode/skills/arxiv in your project.

What does Arxiv need to run?

Going by SKILL.md and its folder, Arxiv needs the command-line tools its instructions call (curl). Our summary lists: Python 3.

Does Arxiv access the network?

SKILL.md names 3 domains. In commands or code: export.arxiv.org, arxiv.org and w3.org; the agent is likely to contact these when it follows the instructions. This is read from the text; nothing was executed.

Is Arxiv safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Arxiv use?

Arxiv is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Arxiv use?

About 2k tokens (SKILL.md is roughly 8.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Arxiv?

Skills that share tags, products or a category with Arxiv: Read arXiv Paper (karpathy/nanochat, 59k stars), Literature Review (neflibata-feng/MyArxiv-Agent, 126 stars), Openalex Database (neflibata-feng/MyArxiv-Agent, 126 stars) and Citation Management (K-Dense-AI/claude-scientific-writer, 2.4k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Arxiv?

AlexAI-MCP (a GitHub user) maintains it in AlexAI-MCP/hermes-CCC, which has 135 GitHub stars. The repository holds 44 skills in this directory. The repository was last updated on April 8, 2026.

Source: AlexAI-MCP/hermes-CCC on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.