Agent skill

Latex OCR Guide

by wentorai in wentorai/research-plugins

Extract and convert mathematical formulas from images and PDFs to LaTeX code

MITAuto-check passedDocuments & Office

Install Latex OCR Guide

skills CLI
$ npx skills add wentorai/research-plugins --skill latex-ocr-guide -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install wentorai/research-plugins latex-ocr-guide --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/wentorai/research-plugins.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/tools/ocr-translate/latex-ocr-guide .claude/skills/latex-ocr-guide && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
latex-ocr-guide
GitHub stars
298
Used in
1 other repo
Token cost
~1.6k tokens
SKILL.md length
193 words
Files
1
Skills in repo
405
Repo updated
First seen
Licence
MIT

At a glance

Extract and convert mathematical formulas from images and PDFs to LaTeX code

  • Tasks that involve LaTeX
  • SKILL.md covers Tool Landscape, Batch Processing Workflow, Using Mathpix API and Verification and Correction
  • Calls pip; reaches github.com and api.mathpix.com

What it does

Latex OCR Guide is an agent skill from wentorai/research-plugins. Extract and convert mathematical formulas from images and PDFs to LaTeX code

Its SKILL.md is about 1.6k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Documents & Office, covering LaTeX. It works with LaTeX. The repository describes itself as: 350+ academic research skills, MCP configs, and plugins for Research-Claw and AI agents. The licence is MIT.

When your agent uses it

  • Tasks that involve LaTeX

Example prompts

  • “/latex-ocr-guide”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit bf44b3c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • github.com
    • api.mathpix.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Latex OCR Guide loads about 1.6k tokens when it runs. Until then it costs about 23 tokens; SKILL.md has 193 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~23
When it runs · the whole SKILL.md, loaded when a task matches
~1.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from wentorai/research-plugins at commit bf44b3c, republished under its MIT licence (© wentorai). 193 words, ~1,571 tokens.

Download SKILL.mdSave it as .claude/skills/latex-ocr-guide/SKILL.md (or your agent's skills folder).
name
latex-ocr-guide
description
Extract and convert mathematical formulas from images and PDFs to LaTeX code

LaTeX OCR Guide

A skill for extracting mathematical formulas from images, PDFs, and handwritten notes and converting them to LaTeX code. Covers tool selection, batch processing workflows, and quality verification techniques.

Tool Landscape

Available Math OCR Tools
ToolTypeAccuracyBest ForLicense
MathpixCloud APIVery highAll math, diagramsCommercial ($)
LaTeX-OCR (Lukas Blecher)Local modelHighPrinted formulasMIT
Pix2TexLocal modelHighSingle equationsMIT
Nougat (Meta)Local modelHighFull papers with mathMIT
InftyReaderDesktopHighPrinted math, JapaneseCommercial
img2latexLocal modelModerateSimple equationsMIT
Quick Start with LaTeX-OCR
bash
# Install the open-source LaTeX-OCR package
pip install "pix2tex[gui]"

# Or install from GitHub for latest version
pip install git+https://github.com/lukas-blecher/LaTeX-OCR.git
python
from pix2tex.cli import LatexOCR
from PIL import Image

def recognize_formula(image_path: str) -> str:
    """
    Convert a formula image to LaTeX code.

    Args:
        image_path: Path to image containing a mathematical formula
    Returns:
        LaTeX string representation of the formula
    """
    model = LatexOCR()
    img = Image.open(image_path)
    latex_code = model(img)
    return latex_code

# Single image
result = recognize_formula('formula.png')
print(result)
# Output: E = mc^{2}

Batch Processing Workflow

Processing Multiple Formulas from a PDF
python
import fitz  # PyMuPDF
from PIL import Image
import io

def extract_formulas_from_pdf(pdf_path: str, output_dir: str,
                                min_height: int = 30) -> list[dict]:
    """
    Extract formula regions from a PDF and convert to LaTeX.

    Args:
        pdf_path: Path to the PDF file
        output_dir: Directory to save extracted formula images
        min_height: Minimum height (px) to consider as formula region
    """
    doc = fitz.open(pdf_path)
    model = LatexOCR()
    results = []

    for page_num in range(len(doc)):
        page = doc[page_num]
        # Extract images from page
        image_list = page.get_images(full=True)

        for img_idx, img_info in enumerate(image_list):
            xref = img_info[0]
            pix = fitz.Pixmap(doc, xref)

            if pix.height >= min_height:
                img_data = pix.tobytes("png")
                img = Image.open(io.BytesIO(img_data))

                try:
                    latex = model(img)
                    results.append({
                        'page': page_num + 1,
                        'image_index': img_idx,
                        'latex': latex,
                        'confidence': 'high' if len(latex) > 3 else 'low'
                    })
                except Exception as e:
                    results.append({
                        'page': page_num + 1,
                        'image_index': img_idx,
                        'latex': None,
                        'error': str(e)
                    })

    return results
Processing Handwritten Notes

For handwritten mathematics, preprocessing improves accuracy significantly:

python
import cv2
import numpy as np

def preprocess_handwritten(image_path: str) -> Image.Image:
    """
    Preprocess a handwritten formula image for better OCR accuracy.
    """
    img = cv2.imread(image_path, cv2.IMREAD_GRAYSCALE)

    # 1. Denoise
    img = cv2.fastNlMeansDenoising(img, h=10)

    # 2. Adaptive thresholding for varying illumination
    img = cv2.adaptiveThreshold(
        img, 255, cv2.ADAPTIVE_THRESH_GAUSSIAN_C,
        cv2.THRESH_BINARY, 15, 8
    )

    # 3. Dilation to connect broken strokes
    kernel = np.ones((2, 2), np.uint8)
    img = cv2.dilate(img, kernel, iterations=1)

    # 4. Crop to content with padding
    coords = cv2.findNonZero(255 - img)
    x, y, w, h = cv2.boundingRect(coords)
    pad = 20
    img = img[max(0, y-pad):y+h+pad, max(0, x-pad):x+w+pad]

    return Image.fromarray(img)

Using Mathpix API

Pricing note: Mathpix is a paid service (starting at $5/month). For free open-source alternatives, use pix2tex/LaTeX-OCR or Nougat (Meta), both MIT-licensed and capable of running locally.

For production-quality results, the Mathpix API provides the highest accuracy:

python
import requests
import base64

def mathpix_ocr(image_path: str, app_id: str, app_key: str) -> dict:
    """
    Use Mathpix API for high-accuracy math OCR.
    """
    with open(image_path, 'rb') as f:
        image_data = base64.b64encode(f.read()).decode()

    response = requests.post(
        'https://api.mathpix.com/v3/text',
        headers={
            'app_id': app_id,
            'app_key': app_key,
            'Content-type': 'application/json'
        },
        json={
            'src': f'data:image/png;base64,{image_data}',
            'formats': ['latex_styled', 'text'],
            'data_options': {'include_asciimath': True}
        }
    )
    return response.json()

Verification and Correction

Always verify OCR output by rendering the LaTeX:

python
import matplotlib.pyplot as plt

def verify_latex(latex_string: str, output_path: str = 'verify.png'):
    """Render LaTeX formula and save as image for visual verification."""
    fig, ax = plt.subplots(figsize=(8, 2))
    ax.text(0.5, 0.5, f'${latex_string}$', fontsize=20,
            ha='center', va='center', transform=ax.transAxes)
    ax.axis('off')
    fig.savefig(output_path, dpi=150, bbox_inches='tight')
    plt.close()
    print(f"Verification image saved to {output_path}")

Common OCR errors to watch for: confusing l with 1, O with 0, missing superscripts/subscripts, incorrect fraction nesting, and misrecognized Greek letters. Always proofread critical equations before submission.

© wentorai, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/tools/ocr-translate/latex-ocr-guide of wentorai/research-plugins.

Open the folder on GitHubat commit bf44b3c

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in wentorai/research-plugins, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Latex OCR Guide next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Latex OCR Guide compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Latex OCR Guide this skillwentorai/research-plugins2981 repos~1.6kAutomated safety check: PassMIT
Research Writingalfonso0512/research-writing-skill4901 repos~818Automated safety check: PassMIT
Paper WritingMLNLP-World/Paper-Writing-Tips4.7k—~630Automated safety check: PassNone
Evomath TaoEvoScientist/EvoSkills4782 repos~3.8kAutomated safety check: PassApache-2.0
PaperjurySpark-To-Paper-Skills/paperjury1.2k—~5.3kAutomated safety check: PassMIT
Thesis Defense PPTX Builderzouchenzhen/thesis-defense-pptx-skill266—~2.4kAutomated safety check: PassApache-2.0

Similar skills

  • Research Writing

    alfonso0512/research-writing-skill

    科研论文写作助手,提供 30 个 Prompt 模板覆盖论文写作全流程. An agent skill from alfonso0512/research-writing-skill.

    490 GitHub starsUsed in 1 repo~818 tokens
    Documents & OfficeAuto-check passed
  • Paper Writing

    MLNLP-World/Paper-Writing-Tips

    学术论文写作检查与优化助手。基于 MLNLP-World 社区整理的论文写作技巧,帮助检查和优化学术论文。Use when: (1) 检查论文 LaTeX 格式和排版, (2) 优化公式符号使用, (3) 改进图表设计, (4) 润色英文学术表达, (5) 检查参考文献格式, (6) 投稿前终稿检查, (7) 用户询问论文写作技巧或规范。

    4.7k GitHub stars~630 tokensUpdated 14 days ago
    Documents & OfficeAuto-check passed
  • Evomath Tao

    EvoScientist/EvoSkills

    A skill your agent uses whenever the user submits a non-trivial mathematical claim that needs a rigorous proof or audit.

    478 GitHub starsUsed in 2 repos~3.8k tokens
    Documents & OfficeAuto-check passed
  • Paperjury

    Spark-To-Paper-Skills/paperjury

    Three modes for CS-conference papers (CVPR/ICCV/ECCV vision, ACL/EMNLP/NAACL NLP, ICLR/NeurIPS/ICML/AAAI ML).

    1.2k GitHub stars~5.3k tokensUpdated 1 mo ago
    Documents & OfficeAuto-check passed
  • Thesis Defense PPTX Builder

    zouchenzhen/thesis-defense-pptx-skill

    Builds an editable thesis defense PowerPoint from a thesis PDF or LaTeX project while preserving a supplied university or lab template, then runs a visual quality check.

    266 GitHub stars~2.4k tokensUpdated 4 mo ago
    Documents & OfficeAuto-check passed
  • Mathmodel Skill

    handsomeZR-netizen/mathmodel-skill

    CUMCM 国赛、MCM/ICM 美赛与电工杯数学建模竞赛的端到端协作工作流。Use when a user explicitly works on one of these modeling contests or asks to run/review a modeling-competition paper from problem selection through modeling…

    292 GitHub stars~2.5k tokensUpdated 14 days ago
    Documents & OfficeAuto-check passed

More from wentorai/research-plugins

All 405 skills in this repo
  • Abstract Writing Guide

    wentorai/research-plugins

    Craft structured research abstracts that maximize clarity and journal acceptance

    298 GitHub starsUsed in 1 repo~1.7k tokens
    Auto-check passed
  • Academic Citation Manager

    wentorai/research-plugins

    Manage academic citations across BibTeX, APA, MLA, and Chicago formats

    298 GitHub starsUsed in 1 repo~2.7k tokens
    Auto-check passed
  • Academic Paper Summarizer

    wentorai/research-plugins

    Summarize academic papers with structured extraction of key elements

    298 GitHub starsUsed in 1 repo~1.4k tokens
    Auto-check passed
  • Academic Study Methods

    wentorai/research-plugins

    Evidence-based study techniques for academic learning and retention

    298 GitHub starsUsed in 1 repo~1.8k tokens
    Auto-check passed
  • Academic Tone Guide

    wentorai/research-plugins

    Adjust writing tone and register for academic audiences and venues

    298 GitHub starsUsed in 1 repo~1.9k tokens
    Auto-check passed
  • Academic Translation Guide

    wentorai/research-plugins

    Academic translation, post-editing, and Chinglish correction guide

    298 GitHub starsUsed in 1 repo~1.6k tokens
    Auto-check passed

Works with

Questions about Latex OCR Guide

What does Latex OCR Guide do?

Extract and convert mathematical formulas from images and PDFs to LaTeX code. Latex OCR Guide is an agent skill from wentorai/research-plugins.

When should I use Latex OCR Guide?

Latex OCR Guide fits situations like: tasks that involve LaTeX.

How do I install Latex OCR Guide in Claude Code?

Run `npx skills add wentorai/research-plugins --skill latex-ocr-guide -a claude-code`. Or copy the skill folder (skills/tools/ocr-translate/latex-ocr-guide in wentorai/research-plugins) into .claude/skills/latex-ocr-guide in your project. Claude Code loads it when a task matches its description.

How do I install Latex OCR Guide in Codex?

Run `npx skills add wentorai/research-plugins --skill latex-ocr-guide -a codex`. Or copy the skill folder (skills/tools/ocr-translate/latex-ocr-guide in wentorai/research-plugins) into .agents/skills/latex-ocr-guide in your project. Codex loads it when a task matches its description.

Can I use Latex OCR Guide in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add wentorai/research-plugins --skill latex-ocr-guide -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/latex-ocr-guide, .gemini/skills/latex-ocr-guide, .github/skills/latex-ocr-guide and .opencode/skills/latex-ocr-guide in your project.

What does Latex OCR Guide need to run?

Going by SKILL.md and its folder, Latex OCR Guide needs the command-line tools its instructions call (pip). Our summary lists: Python 3.

Does Latex OCR Guide access the network?

SKILL.md names 2 domains. In commands or code: github.com and api.mathpix.com; the agent is likely to contact these when it follows the instructions. This is read from the text; nothing was executed.

Is Latex OCR Guide safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Latex OCR Guide use?

Latex OCR Guide is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Latex OCR Guide use?

About 1.6k tokens (SKILL.md is roughly 6.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Latex OCR Guide?

Skills that share tags, products or a category with Latex OCR Guide: Research Writing (alfonso0512/research-writing-skill, 490 stars), Paper Writing (MLNLP-World/Paper-Writing-Tips, 4.7k stars), Evomath Tao (EvoScientist/EvoSkills, 478 stars) and Paperjury (Spark-To-Paper-Skills/paperjury, 1.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Latex OCR Guide?

wentorai (a GitHub user) maintains it in wentorai/research-plugins, which has 298 GitHub stars. The repository holds 405 skills in this directory. The repository was last updated on June 19, 2026.

Source: wentorai/research-plugins on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.