Agent skill

Pdf2tex

by Calix-L in Calix-L/awesome-latex-skills

Reconstruct editable LaTeX from PDF content using page-aware extraction and visual comparison.

MITAuto-check passedDocuments & Office

Install Pdf2tex

skills CLI
$ npx skills add Calix-L/awesome-latex-skills --skill pdf2tex -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Calix-L/awesome-latex-skills pdf2tex --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Calix-L/awesome-latex-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/pdf2tex .claude/skills/pdf2tex && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
pdf2tex
GitHub stars
181
Token cost
~1.4k tokens
SKILL.md length
665 words
Files
10 (incl. scripts, references, assets)
Skills in repo
5
Repo updated
First seen
Licence
MIT

At a glance

Reconstruct editable LaTeX from PDF content using page-aware extraction and visual comparison.

  • Tasks that involve LaTeX
  • SKILL.md covers Establish the reconstruction…, Extract evidence, Reconstruct without inventing… and Build, compare, and deliver
  • Runs Python scripts from its folder; calls python
  • Tasks that involve PDF

What it does

Pdf2tex is an agent skill from Calix-L/awesome-latex-skills. Reconstruct editable LaTeX from PDF content using page-aware extraction and visual comparison. Preserve source evidence, flag uncertain math/tables/citations, and distinguish text extraction, OCR, reconstruction, and verified compilation.

Its SKILL.md is about 1.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 13 other files, including scripts, reference files and assets (for example `agents/config.yaml`, `agents/openai.yaml` and `references/math-reconstruction.md`).

It sits in Documents & Office, covering LaTeX, PDF and Citation management. It works with LaTeX. The repository describes itself as: LaTeX manuscript workflows with five agent skills, project diagnostics, reproducible builds and offline change review. The licence is MIT.

When your agent uses it

  • Tasks that involve LaTeX
  • Tasks that involve PDF
  • Tasks that involve Citation management

Example prompts

  • “/pdf2tex”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit 7c29d3c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Pdf2tex loads about 1.4k tokens when it runs, and up to ~7.9k if it reads all its reference files. Until then it costs about 62 tokens; SKILL.md has 665 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~62
When it runs · the whole SKILL.md, loaded when a task matches
~1.4k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~7.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from Calix-L/awesome-latex-skills at commit 7c29d3c, republished under its MIT licence (© Calix-L). 665 words, ~1,378 tokens.

Download SKILL.mdSave it as .claude/skills/pdf2tex/SKILL.md (or your agent's skills folder). This skill also uses 9 other files; get the full folder from GitHub.
name
pdf2tex
description
Reconstruct editable LaTeX from PDF content using page-aware extraction and visual comparison. Preserve source evidence, flag uncertain math/tables/citations, and distinguish text extraction, OCR, reconstruction, and verified compilation.
metadata.version
1.25.0

Establish the reconstruction target

Inspect the PDF, requested pages, available tools, and desired output. Determine whether the user wants content recovery or close visual reconstruction. Preserve the original PDF and write new artifacts to a separate destination.

A PDF may expose text, font names, coordinates, images, and metadata. It does not reliably encode its original document class, packages, macros, bibliography database, comments, or source-file boundaries. Font/creator metadata is evidence for a candidate setup, not proof of the original engine or class.

Extract evidence

When PyMuPDF is available, use the bundled helper from this skill's own directory. The following paths are relative to the repository root; for an installed skill, substitute its actual location. Dependency installation is separate from extraction.

sh
python -m pip install -r pdf2tex/requirements.txt
python pdf2tex/scripts/extract_pdf.py paper.pdf --output extraction --pages 1-3,5 --images --render

The helper creates a new directory with report.html, text.txt, layout.json, and optional embedded images and whole-page PNG previews when requested. It records page numbers, raw text spans/font/position data, metadata, selected-page coverage, and warnings. It refuses existing output directories and refuses publication if the input fingerprint changes during extraction. Open report.html for offline page/text review; keep the entire directory together when sharing. It performs no OCR or conversion.

Read PDF extraction guide for API details, alternative readers, columns, fonts, and OCR. Sorted text is not guaranteed reading order; inspect page layouts and use coordinates. Images can be repeated or carry separate soft masks. Vector figures and composite panels often need a page crop or another export workflow. Use optional --render previews to inspect selected pages, including vector/composite figures; these are visual evidence, not OCR or segmented assets. --dpi accepts 72–300 with a per-page pixel limit.

Use optional --chars when inspecting scripts or small notation. It adds character origins/bounding boxes while retaining span text. Page geometry and rotation matrices help relate unrotated text coordinates to rendered previews; positions are evidence for candidate readings, not an automatic math parser.

A page without text may be blank, graphical, or scanned. Check it visually before choosing OCR. OCR requires separate tools and cannot establish the correctness of equations or tables. Retain page provenance and flag OCR-derived uncertainty. For a password-protected PDF, use an authorized readable copy.

Show full SKILL.md (315 more words)Show less

Reconstruct without inventing content

Use structure detection to interpret blocks, math reconstruction for notation, and table reconstruction for cells and merged regions. These heuristics need comparison with the rendered original.

  • Select an available class and engine suitable for the target; state inferred choices. Use a supplied official author kit when exact publication layout is required.
  • Preserve selected-page coverage, section order, prose, equations, table values, captions, footnotes, and references. Escape LaTeX-special characters in prose without indiscriminately escaping math or generated commands.
  • Associate citation markers with bibliography entries only when the mapping is supported. Keep unmatched markers and uncertainty visible; do not invent bibliographic metadata or silently assign the nearest reference.
  • Preserve ambiguous glyphs, merged table cells, missing images, and illegible content as source evidence with % [UNCERTAIN: ...] or a visible placeholder. A comment alone must not hide missing content from the generated document.
  • Remove headers/footers or join hyphenated lines only after checking that they are layout artifacts. Preserve meaningful hyphens and repeated scientific text.
  • Do not guess original macros or file splitting. A self-contained source is a useful default, not a claim that it matches the original organization.

Build, compare, and deliver

Use the selected engine and actual bibliography backend, with additional passes for cross-references. latex-rescue can help when available. If tools or assets are unavailable, preserve the source and report compilation as unverified.

Compare the rendered reconstruction with the selected original pages: completeness, reading order, math symbols, tables, figure appearances, captions, and citations. Matching page counts does not establish fidelity. Check merged cells and OCR math manually, and distinguish a visual approximation from content verification.

Deliver the new source/assets, input version and selected pages, extraction and OCR methods actually used, inferred class/engine, build and visual-check results, and uncertainty locations. Separate recovered content from placeholders. Do not promise exact original source, perfect reconstruction, or immediate compilation. Use latex-polish or latex-fmt only for a further requested editing task.

© Calix-L, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 9 other files (scripts, references, assets) in pdf2tex of Calix-L/awesome-latex-skills.

  • SKILL.md
  • agents/config.yaml
  • agents/openai.yaml
  • assets/review.css
  • references/math-reconstruction.md
  • references/pdf-extraction-guide.md
  • references/structure-detection.md
  • references/table-reconstruction.md
  • requirements.txt
  • scripts/extract_pdf.py

Open the folder on GitHubat commit 7c29d3c

Compare with similar skills

Pdf2tex next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Pdf2tex compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Pdf2tex this skillCalix-L/awesome-latex-skills181—~1.4kAutomated safety check: PassMIT
Evidence Ledgerwanshuiyin/Anti-Autoresearch161—~11kAutomated safety check: NotesMIT
Paper CovertGRIND-Lab-Core/night_owl_research_agent106—~2.1kAutomated safety check: NotesNone
Kimi PDFthvroyal/kimi-skills238—~1.9kAutomated safety check: PassNone
Paper Auditbahayonghang/academic-writing-skills498—~5.1kAutomated safety check: PassNone
Thu ThesisLeoYeAI/openclaw-master-skills2.2k—~4kAutomated safety check: PassMIT

Similar skills

  • Evidence Ledger

    wanshuiyin/Anti-Autoresearch

    Build the deterministic evidence ledger (artifactmanifest.json + claims.json) that every other Anti-Autoresearch auditor reads.

    161 GitHub stars~11k tokensUpdated 2 days ago
    Documents & OfficeAuto-check: notes
  • Paper Covert

    GRIND-Lab-Core/night_owl_research_agent

    Converts the final Markdown manuscript from paper-draft / paper-review-loop into a submission package for the target venue — modular LaTeX (one file per section), compiled PDF, and Word .docx.

    106 GitHub stars~2.1k tokensUpdated 5 mo ago
    Documents & OfficeAuto-check: notes
  • Kimi PDF

    thvroyal/kimi-skills

    Professional PDF solution. An agent skill from thvroyal/kimi-skills.

    238 GitHub stars~1.9k tokensUpdated 6 mo ago
    Documents & OfficeAuto-check passed
  • Paper Audit

    bahayonghang/academic-writing-skills

    Reviewer-style audit and submission gate for academic papers in .tex, .typ, or .pdf.

    498 GitHub stars~5.1k tokensUpdated 11 days ago
    Documents & OfficeAuto-check passed
  • Thu Thesis

    LeoYeAI/openclaw-master-skills

    清华大学毕业论文 Word → PDF 一键格式规范化工具。输入任意 Word (.docx) 格式的清华毕业论文,自动转换为符合清华 thuthesis 官方 LaTeX 模板规范的高质量 PDF。适用于所有清华学位论文(MBA/学硕/专硕),一条命令搞定。功能:自动提取章节结构、中英文摘要、参考文献(自动生成 BibTeX)、图片(含…

    2.2k GitHub stars~4k tokensUpdated 2 mo ago
    Documents & OfficeAuto-check passed
  • 01 Paper Review

    agentscope-ai/OpenJudge

    Review academic papers for correctness, quality, and novelty using OpenJudge's multi-stage pipeline.

    870 GitHub stars~2.4k tokensUpdated 28 days ago
    Research & ScienceAuto-check passed

More from Calix-L/awesome-latex-skills

  • Latex Rescue

    Calix-L/awesome-latex-skills

    Diagnose and repair LaTeX build failures in local projects or supplied logs.

    181 GitHub stars~1.4k tokensUpdated yesterday
    Auto-check passed
  • Latex Fmt

    Calix-L/awesome-latex-skills

    Reformat LaTeX papers for a specified venue, year, track, and submission stage.

    181 GitHub stars~1.4k tokensUpdated yesterday
    Auto-check passed
  • Latex Polish

    Calix-L/awesome-latex-skills

    Edit academic prose in LaTeX manuscripts or supplied text at light, moderate, or strict intensity.

    181 GitHub stars~981 tokensUpdated yesterday
    Auto-check passed
  • Paper Read

    Calix-L/awesome-latex-skills

    Summarize and critically analyze academic papers supplied as PDFs, text, arXiv links, or journal URLs.

    181 GitHub stars~1.1k tokensUpdated yesterday
    Auto-check passed

Works with

Questions about Pdf2tex

What does Pdf2tex do?

Reconstruct editable LaTeX from PDF content using page-aware extraction and visual comparison. Pdf2tex is an agent skill from Calix-L/awesome-latex-skills. Reconstruct editable LaTeX from PDF content using page-aware extraction and visual comparison.

When should I use Pdf2tex?

Pdf2tex fits situations like: tasks that involve LaTeX; tasks that involve PDF; tasks that involve Citation management.

How do I install Pdf2tex in Claude Code?

Run `npx skills add Calix-L/awesome-latex-skills --skill pdf2tex -a claude-code`. Or copy the skill folder (pdf2tex in Calix-L/awesome-latex-skills) into .claude/skills/pdf2tex in your project. Claude Code loads it when a task matches its description.

How do I install Pdf2tex in Codex?

Run `npx skills add Calix-L/awesome-latex-skills --skill pdf2tex -a codex`. Or copy the skill folder (pdf2tex in Calix-L/awesome-latex-skills) into .agents/skills/pdf2tex in your project. Codex loads it when a task matches its description.

Can I use Pdf2tex in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Calix-L/awesome-latex-skills --skill pdf2tex -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/pdf2tex, .gemini/skills/pdf2tex, .github/skills/pdf2tex and .opencode/skills/pdf2tex in your project.

What does Pdf2tex need to run?

Going by SKILL.md and its folder, Pdf2tex needs Python for the scripts in its folder and the command-line tools its instructions call (python). Our summary lists: Python 3.

Does Pdf2tex access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Pdf2tex safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Pdf2tex use?

Pdf2tex is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Pdf2tex use?

About 1.4k tokens (SKILL.md is roughly 5.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 6.5k tokens, read only when the agent opens those files.

What are the alternatives to Pdf2tex?

Skills that share tags, products or a category with Pdf2tex: Evidence Ledger (wanshuiyin/Anti-Autoresearch, 161 stars), Paper Covert (GRIND-Lab-Core/night_owl_research_agent, 106 stars), Kimi PDF (thvroyal/kimi-skills, 238 stars) and Paper Audit (bahayonghang/academic-writing-skills, 498 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Pdf2tex?

Calix-L (a GitHub user) maintains it in Calix-L/awesome-latex-skills, which has 181 GitHub stars. The repository holds 5 skills in this directory. The repository was last updated on October 7, 2026.

Source: Calix-L/awesome-latex-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.