Agent skill

Arxiv Latex Source

by wentorai in wentorai/research-plugins

Download and parse LaTeX source files from arXiv preprints. An agent skill from wentorai/research-plugins.

MITAuto-check passedResearch & Science

Install Arxiv Latex Source

skills CLI
$ npx skills add wentorai/research-plugins --skill arxiv-latex-source -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install wentorai/research-plugins arxiv-latex-source --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/wentorai/research-plugins.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/literature/fulltext/arxiv-latex-source .claude/skills/arxiv-latex-source && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
arxiv-latex-source
GitHub stars
298
Used in
1 other repo
Token cost
~2k tokens
SKILL.md length
487 words
Files
1
Skills in repo
405
Repo updated
First seen
Licence
MIT

At a glance

Download and parse LaTeX source files from arXiv preprints. An agent skill from wentorai/research-plugins.

  • Works in 3 steps: Look for \documentclass in .tex files —… → Check for a README.txt that may specify… → If multiple .tex files contain…
  • Tasks that involve Academic paper search
  • SKILL.md covers Overview, Authentication, Core Endpoints and LaTeX Source Parsing Guide, plus 4 more sections
  • Calls curl; reaches arxiv.org and export.arxiv.org

What it does

Arxiv Latex Source is an agent skill from wentorai/research-plugins. Download and parse LaTeX source files from arXiv preprints

Its SKILL.md is about 2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Research & Science, covering Academic paper search and LaTeX. It works with LaTeX and arXiv. The repository describes itself as: 350+ academic research skills, MCP configs, and plugins for Research-Claw and AI agents. The licence is MIT.

When your agent uses it

  • Tasks that involve Academic paper search
  • Tasks that involve LaTeX

Example prompts

  • “/arxiv-latex-source”

Requirements

  • Python 3

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. Look for \documentclass in .tex files — this marks the root document
  2. Check for a README.txt that may specify the main file
  3. If multiple .tex files contain \documentclass, prefer the one with \begin{document}

What it can do on your machine

Read from SKILL.md and the folder at commit bf44b3c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • curl

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • arxiv.org
    • export.arxiv.org

    Also links to:

    • info.arxiv.org

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Arxiv Latex Source loads about 2k tokens when it runs. Until then it costs about 19 tokens; SKILL.md has 487 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~19
When it runs · the whole SKILL.md, loaded when a task matches
~2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from wentorai/research-plugins at commit bf44b3c, republished under its MIT licence (© wentorai). 487 words, ~1,995 tokens.

Download SKILL.mdSave it as .claude/skills/arxiv-latex-source/SKILL.md (or your agent's skills folder).
name
arxiv-latex-source
description
Download and parse LaTeX source files from arXiv preprints

arXiv LaTeX Source Access Guide

Overview

arXiv stores the original LaTeX source files for the vast majority of its 2.4 million+ preprints. Accessing LaTeX source provides major advantages over PDF parsing: exact mathematical notation as written by the author, structured sections and labels, machine-readable bibliography entries, and intact figure captions, table data, and cross-references.

For formula extraction, citation graph construction, section-level text analysis, or training data curation for scientific language models, LaTeX source is the gold standard. PDF parsing introduces OCR errors in equations, loses structural hierarchy, and mangles complex tables.

The e-print endpoint serves source bundles as gzip-compressed tarballs (.tar.gz) containing .tex files, figures, .bib/.bbl bibliography files, style files, and supplementary materials. No authentication is required.

Authentication

No authentication or API key is required. The e-print endpoint is publicly accessible. However, arXiv asks that automated tools set a descriptive User-Agent header and comply with rate limits.

Core Endpoints

Download LaTeX Source
  • URL: GET https://arxiv.org/e-print/{arxiv_id}

  • Response: application/gzip — a .tar.gz archive containing the source files

  • Parameters:

    ParamTypeRequiredDescription
    arxiv_idstringYesarXiv identifier, e.g. 2301.00001 or 2301.00001v2 for a specific version
  • Example:

    bash
    # Download source archive (response: 200, application/gzip, ~1.3 MB)
    curl -sL -o source.tar.gz "https://arxiv.org/e-print/2301.00001"
    
    # List archive contents
    tar tz -f source.tar.gz | head -10
    # ACM-Reference-Format.bbx
    # ACM-Reference-Format.bst
    # Image_1.jpg
    # README.txt
    # acmart.cls
  • Content-Disposition header: attachment; filename="arXiv-2301.00001v1.tar.gz"

  • ETag: SHA-256 hash provided for caching: sha256:f1ffe8ec...

Format Detection

The endpoint almost always returns a gzip-compressed tar archive. Rare cases (very old or single-file submissions) may return a single gzip-compressed .tex file without tar wrapper. Always verify format before extracting:

bash
curl -sL "https://arxiv.org/e-print/{arxiv_id}" -o source.gz
file source.gz  # "gzip compressed data, was 'XXXX.tar', ..."
Metadata API (Companion)

Pair source downloads with the arXiv Atom API for structured metadata:

  • URL: GET https://export.arxiv.org/api/query?id_list={arxiv_id}
  • Response: Atom XML with <title>, <author>, <summary>, <category>, <published>
  • Example: curl -s "https://export.arxiv.org/api/query?id_list=2301.00001"

LaTeX Source Parsing Guide

Show full SKILL.md (226 more words)Show less
Locating the Main .tex File

A source archive typically contains multiple files. To find the main document:

  1. Look for \documentclass in .tex files — this marks the root document
  2. Check for a README.txt that may specify the main file
  3. If multiple .tex files contain \documentclass, prefer the one with \begin{document}
python
import tarfile, re

def find_main_tex(tar_path):
    with tarfile.open(tar_path, 'r:gz') as tar:
        tex_files = [m for m in tar.getmembers() if m.name.endswith('.tex')]
        for member in tex_files:
            content = tar.extractfile(member).read().decode('utf-8', errors='ignore')
            if r'\documentclass' in content and r'\begin{document}' in content:
                return member.name, content
    return None, None
Extracting Sections

LaTeX sections follow a predictable hierarchy:

python
import re

def extract_sections(tex_content):
    pattern = r'\\(section|subsection|subsubsection)\{([^}]+)\}'
    sections = re.findall(pattern, tex_content)
    return [(level, title) for level, title in sections]
    # [('section', 'Introduction'), ('section', 'Related Work'), ...]
Extracting Equations
python
def extract_equations(tex_content):
    patterns = [
        r'\\\[(.+?)\\\]',
        r'\\begin\{equation\}(.+?)\\end\{equation\}',
        r'\\begin\{align\*?\}(.+?)\\end\{align\*?\}',
    ]
    equations = []
    for pat in patterns:
        equations.extend(re.findall(pat, tex_content, re.DOTALL))
    return equations
Extracting Bibliography

Parse .bib files (BibTeX entries) or .bbl files (compiled \bibitem commands):

python
def extract_bibliography(tar_path):
    refs = []
    with tarfile.open(tar_path, 'r:gz') as tar:
        for member in tar.getmembers():
            if member.name.endswith('.bib'):
                content = tar.extractfile(member).read().decode('utf-8', errors='ignore')
                refs.extend(re.findall(r'@\w+\{([^,]+),(.+?)\n\}', content, re.DOTALL))
            elif member.name.endswith('.bbl'):
                content = tar.extractfile(member).read().decode('utf-8', errors='ignore')
                refs.extend(re.findall(r'\\bibitem.*?\{(.+?)\}', content))
    return refs

Rate Limits

  • Maximum: 4 requests per second for automated access
  • Recommended: 1 request/second with delays between sequential downloads
  • Bulk access: For 1000+ papers, use the arXiv S3 bulk data mirror instead
  • HTTP 429: Rate limit exceeded; implement exponential backoff
  • User-Agent: Required — set a descriptive string: MyTool/1.0 (mailto:user@university.edu)
  • Persistent abuse may result in IP-level blocks

Academic Use Cases

  • Formula extraction for ML training — Build equation datasets with ground-truth LaTeX notation, free of OCR noise from PDF parsing
  • Citation network analysis — Parse .bib/.bbl files for exact reference keys to construct citation graphs
  • Section-level text analysis — Extract specific sections (e.g., all "Related Work" across a subfield) for systematic reviews
  • Reproducibility auditing — Examine algorithm environments, hyperparameter tables, and methodology sections
  • Cross-paper notation alignment — Compare and normalize equation environments across papers in a subfield

Complete Python Example

python
import requests, tarfile, io, re, time, gzip

def download_arxiv_source(arxiv_id, delay=1.0):
    """Download and extract all .tex files from an arXiv paper's source."""
    url = f"https://arxiv.org/e-print/{arxiv_id}"
    headers = {"User-Agent": "ResearchTool/1.0 (mailto:user@example.com)"}
    resp = requests.get(url, headers=headers)
    resp.raise_for_status()
    time.sleep(delay)

    buf = io.BytesIO(resp.content)
    try:
        with tarfile.open(fileobj=buf, mode='r:gz') as tar:
            return {m.name: tar.extractfile(m).read().decode('utf-8', errors='ignore')
                    for m in tar.getmembers() if m.name.endswith('.tex') and m.isfile()}
    except tarfile.ReadError:
        buf.seek(0)
        return {"main.tex": gzip.decompress(buf.read()).decode('utf-8', errors='ignore')}

# Usage
sources = download_arxiv_source("2301.00001")
for fname, content in sources.items():
    if r'\documentclass' in content:
        sections = re.findall(r'\\section\{([^}]+)\}', content)
        equations = re.findall(r'\\begin\{equation\}(.+?)\\end\{equation\}', content, re.DOTALL)
        print(f"{fname}: {len(sections)} sections, {len(equations)} equations")

References

© wentorai, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/literature/fulltext/arxiv-latex-source of wentorai/research-plugins.

Open the folder on GitHubat commit bf44b3c

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in wentorai/research-plugins, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Arxiv Latex Source next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Arxiv Latex Source compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Arxiv Latex Source this skillwentorai/research-plugins2981 repos~2kAutomated safety check: PassMIT
Paper OrchestraAr9av/PaperOrchestra6791 repos~3.5kAutomated safety check: PassCustom licence
Section Writing AgentAr9av/PaperOrchestra6791 repos~3.3kAutomated safety check: PassCustom licence
SearchMuuuun/luxas1.2k—~2.3kAutomated safety check: WarnMIT
Oleafly Pre SubmissionOleafly/Oleafly212—~2.9kAutomated safety check: PassMIT
Anmath Submissionfranklee16/academic-research-skills2231 repos~1.1kAutomated safety check: PassNone

Similar skills

  • Paper Orchestra

    Ar9av/PaperOrchestra

    Orchestrate the full PaperOrchestra (Song et al., 2026, arXiv:2604.05018) five-agent pipeline to turn unstructured research materials (idea, experimental log, LaTeX template, conference guidelines…

    679 GitHub starsUsed in 1 repo~3.5k tokens
    Research & ScienceAuto-check passed
  • Section Writing Agent

    Ar9av/PaperOrchestra

    Step 4 of the PaperOrchestra pipeline (arXiv:2604.05018). An agent skill from Ar9av/PaperOrchestra.

    679 GitHub starsUsed in 1 repo~3.3k tokens
    Research & ScienceAuto-check passed
  • Search

    Muuuun/luxas

    Unified academic paper search, citation chains, paper download (arXiv LaTeX/PDF, Sci-Hub), figure extraction from papers, LaTeX source reading, BibTeX fetching, web search, and browser automation…

    1.2k GitHub stars~2.3k tokensUpdated 1 mo ago
    Research & ScienceAuto-check: warnings
  • Oleafly Pre Submission

    Oleafly/Oleafly

    Run a pre-flight pass over the project before uploading to arXiv or a venue, and write a pass or fail checklist.

    212 GitHub stars~2.9k tokensUpdated yesterday
    Research & ScienceAuto-check passed
  • Anmath Submission

    franklee16/academic-research-skills

    A skill your agent uses when running the final pre-submission preflight for an Annals of Mathematics manuscript — AMS-LaTeX compile, theorem environments, abstract and MSC, references, arXiv…

    223 GitHub starsUsed in 1 repo~1.1k tokens
    Research & ScienceAuto-check passed
  • Arxiv Package

    Mathews-Tom/armory

    Package a TeX/LaTeX project into a clean tarball or zip for arXiv upload: file selection, build-artifact exclusion, 00README.XXX generation, ancillary file organization, archive validation.

    329 GitHub stars~1.6k tokensUpdated 5 days ago
    Research & ScienceAuto-check: notes

More from wentorai/research-plugins

All 405 skills in this repo
  • Abstract Writing Guide

    wentorai/research-plugins

    Craft structured research abstracts that maximize clarity and journal acceptance

    298 GitHub starsUsed in 1 repo~1.7k tokens
    Auto-check passed
  • Academic Citation Manager

    wentorai/research-plugins

    Manage academic citations across BibTeX, APA, MLA, and Chicago formats

    298 GitHub starsUsed in 1 repo~2.7k tokens
    Auto-check passed
  • Academic Paper Summarizer

    wentorai/research-plugins

    Summarize academic papers with structured extraction of key elements

    298 GitHub starsUsed in 1 repo~1.4k tokens
    Auto-check passed
  • Academic Study Methods

    wentorai/research-plugins

    Evidence-based study techniques for academic learning and retention

    298 GitHub starsUsed in 1 repo~1.8k tokens
    Auto-check passed
  • Academic Tone Guide

    wentorai/research-plugins

    Adjust writing tone and register for academic audiences and venues

    298 GitHub starsUsed in 1 repo~1.9k tokens
    Auto-check passed
  • Academic Translation Guide

    wentorai/research-plugins

    Academic translation, post-editing, and Chinglish correction guide

    298 GitHub starsUsed in 1 repo~1.6k tokens
    Auto-check passed

Works with

Questions about Arxiv Latex Source

What does Arxiv Latex Source do?

Download and parse LaTeX source files from arXiv preprints. An agent skill from wentorai/research-plugins. Arxiv Latex Source is an agent skill from wentorai/research-plugins.

When should I use Arxiv Latex Source?

Arxiv Latex Source fits situations like: tasks that involve Academic paper search; tasks that involve LaTeX.

How do I install Arxiv Latex Source in Claude Code?

Run `npx skills add wentorai/research-plugins --skill arxiv-latex-source -a claude-code`. Or copy the skill folder (skills/literature/fulltext/arxiv-latex-source in wentorai/research-plugins) into .claude/skills/arxiv-latex-source in your project. Claude Code loads it when a task matches its description.

How do I install Arxiv Latex Source in Codex?

Run `npx skills add wentorai/research-plugins --skill arxiv-latex-source -a codex`. Or copy the skill folder (skills/literature/fulltext/arxiv-latex-source in wentorai/research-plugins) into .agents/skills/arxiv-latex-source in your project. Codex loads it when a task matches its description.

Can I use Arxiv Latex Source in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add wentorai/research-plugins --skill arxiv-latex-source -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/arxiv-latex-source, .gemini/skills/arxiv-latex-source, .github/skills/arxiv-latex-source and .opencode/skills/arxiv-latex-source in your project.

What does Arxiv Latex Source need to run?

Going by SKILL.md and its folder, Arxiv Latex Source needs the command-line tools its instructions call (curl). Our summary lists: Python 3.

Does Arxiv Latex Source access the network?

SKILL.md names 3 domains. In commands or code: arxiv.org and export.arxiv.org; the agent is likely to contact these when it follows the instructions. As links in the text: info.arxiv.org. This is read from the text; nothing was executed.

Is Arxiv Latex Source safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Arxiv Latex Source use?

Arxiv Latex Source is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Arxiv Latex Source use?

About 2k tokens (SKILL.md is roughly 8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Arxiv Latex Source?

Skills that share tags, products or a category with Arxiv Latex Source: Paper Orchestra (Ar9av/PaperOrchestra, 679 stars), Section Writing Agent (Ar9av/PaperOrchestra, 679 stars), Search (Muuuun/luxas, 1.2k stars) and Oleafly Pre Submission (Oleafly/Oleafly, 212 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Arxiv Latex Source?

wentorai (a GitHub user) maintains it in wentorai/research-plugins, which has 298 GitHub stars. The repository holds 405 skills in this directory. The repository was last updated on June 19, 2026.

Source: wentorai/research-plugins on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.