Agent skill

Liteparse

by sammcj in sammcj/agentic-coding

Provides fast document to markdown extraction. An agent skill from sammcj/agentic-coding.

Apache-2.0Auto-check passedDocuments & Office

Install Liteparse

skills CLI
$ npx skills add sammcj/agentic-coding --skill liteparse -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install sammcj/agentic-coding liteparse --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/sammcj/agentic-coding.git skills-src && mkdir -p .claude/skills && cp -r skills-src/Skills/liteparse .claude/skills/liteparse && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
liteparse
GitHub stars
162
Token cost
~1.1k tokens
SKILL.md length
322 words
Files
1
Skills in repo
64
Repo updated
First seen
Licence
Apache-2.0

At a glance

Provides fast document to markdown extraction. An agent skill from sammcj/agentic-coding.

  • Works in 5 steps: Use via npx, or install LiteParse → Produce the CLI Command or Script → Key Options Reference → …
  • The user asks to parse
  • SKILL.md covers Step 0 - Use via npx, or…, Step 1 - Produce the CLI…, Step 3 - Key Options Reference and Step 4 - Using a Config File, plus 2 more sections
  • Calls npx and pnpm

What it does

Liteparse is an agent skill from sammcj/agentic-coding. Provides fast document to markdown extraction. Use this skill when the user asks to parse, perform multi-format document conversion or spatially extract text from an unstructured file (PDF, DOCX, PPTX, XLSX, images, etc.) locally.

Its SKILL.md is about 1.1k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Documents & Office, covering Document parsing, PowerPoint presentations and Word documents. It works with Microsoft PowerPoint, Microsoft Word and Microsoft Excel. The repository describes itself as: Agentic Coding Rules, Templates etc... The licence is Apache-2.0.

When your agent uses it

  • The user asks to parse
  • Perform multi-format document conversion
  • Spatially extract text from an unstructured file (PDF

Example prompts

  • “Use the liteparse skill to provide fast document to markdown extraction. An agent skill from sammcj/agentic-coding”
  • “/liteparse”

Requirements

  • Node.js

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Use via npx, or install LiteParse
  2. Produce the CLI Command or Script
  3. Key Options Reference
  4. Using a Config File
  5. HTTP OCR Server API (Advanced)

What it can do on your machine

Read from SKILL.md and the folder at commit 2f25ced. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • npx
    • pnpm

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use npx and pnpm, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Liteparse loads about 1.1k tokens when it runs. Until then it costs about 60 tokens; SKILL.md has 322 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~60
When it runs · the whole SKILL.md, loaded when a task matches
~1.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from sammcj/agentic-coding at commit 2f25ced, republished under its Apache-2.0 licence (© sammcj). 322 words, ~1,114 tokens.

Download SKILL.mdSave it as .claude/skills/liteparse/SKILL.md (or your agent's skills folder).
name
liteparse
description
Provides fast document to markdown extraction. Use this skill when the user asks to parse, perform multi-format document conversion or spatially extract text from an unstructured file (PDF, DOCX, PPTX, XLSX, images, etc.) locally.

LiteParse Skill

Parse unstructured documents (PDF, DOCX, PPTX, XLSX, images, and more) locally with LiteParse: fast, lightweight, no cloud dependencies or LLM required.

Step 0 - Use via npx, or install LiteParse

NOTE: Rather than installing liteparse globally, you can instead run it directly with npx, substituting lit <args> with npx -y @llamaindex/liteparse <args> in the commands below.

bash
npx -y @llamaindex/liteparse

Otherwise if installing globally, use pnpm install -g @llamaindex/liteparse and then run lit <args>.


Step 1 - Produce the CLI Command or Script

Parse a Single File
bash
# Basic text extraction
lit parse document.pdf

# JSON output saved to a file
lit parse document.pdf --format json -o output.json

# Specific page range
lit parse document.pdf --target-pages "1-5,10,15-20"

# Disable OCR (faster, text-only PDFs)
lit parse document.pdf --no-ocr

# Use an external HTTP OCR server for higher accuracy
lit parse document.pdf --ocr-server-url http://localhost:8828/ocr

# Higher DPI for better quality
lit parse document.pdf --dpi 300
Batch Parse a Directory
bash
lit batch-parse ./input-directory ./output-directory

# Only process PDFs, recursively
lit batch-parse ./input ./output --extension .pdf --recursive
Generate Page Screenshots

Screenshots are useful for LLM agents that need to see visual layout.

bash
# All pages
lit screenshot document.pdf -o ./screenshots

# Specific pages
lit screenshot document.pdf --pages "1,3,5" -o ./screenshots

# High-DPI PNG
lit screenshot document.pdf --dpi 300 --format png -o ./screenshots

# Page range
lit screenshot document.pdf --pages "1-10" -o ./screenshots

Step 3 - Key Options Reference

OCR Options
OptionDescription
(default)Tesseract.js - zero setup, built-in
--ocr-language fraSet OCR language (ISO code)
--ocr-server-url <url>Use external HTTP OCR server (EasyOCR, PaddleOCR, custom)
--no-ocrDisable OCR entirely
Output Options
OptionDescription
--format jsonStructured JSON with bounding boxes
--format textPlain text (default)
-o <file>Save output to file
Performance / Quality Options
OptionDescription
--dpi <n>Rendering DPI (default: 150; use 300 for high quality)
--max-pages <n>Limit pages parsed
--target-pages <pages>Parse specific pages (e.g. "1-5,10")
--no-precise-bboxDisable precise bounding boxes (faster)
--skip-diagonal-textIgnore rotated/diagonal text
--preserve-small-textKeep very small text that would otherwise be dropped

Step 4 - Using a Config File

For repeated use with consistent options, generate a liteparse.config.json:

json
{
  "ocrLanguage": "en",
  "ocrEnabled": true,
  "maxPages": 1000,
  "dpi": 150,
  "outputFormat": "json",
  "preciseBoundingBox": true,
  "skipDiagonalText": false,
  "preserveVerySmallText": false
}

For an HTTP OCR server:

json
{
  "ocrServerUrl": "http://localhost:8828/ocr",
  "ocrLanguage": "en",
  "outputFormat": "json"
}

Use with:

bash
lit parse document.pdf --config liteparse.config.json

Step 5 - HTTP OCR Server API (Advanced)

If the user wants to plug in a custom OCR backend, the server must implement:

  • Endpoint: POST /ocr
  • Accepts: file (multipart) and language (string) parameters
  • Returns:
json
{
  "results": [
    { "text": "Hello", "bbox": [x1, y1, x2, y2], "confidence": 0.98 }
  ]
}

Ready-to-use wrappers exist for EasyOCR and PaddleOCR in the LiteParse repo.


Supported Input Formats

CategoryFormats
PDF.pdf
Word.doc, .docx, .docm, .odt, .rtf
PowerPoint.ppt, .pptx, .pptm, .odp
Spreadsheets.xls, .xlsx, .xlsm, .ods, .csv, .tsv
Images.jpg, .jpeg, .png, .gif, .bmp, .tiff, .webp, .svg

Office documents require LibreOffice; images require ImageMagick. LiteParse auto-converts these formats to PDF before parsing.

© sammcj, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in Skills/liteparse of sammcj/agentic-coding.

Open the folder on GitHubat commit 2f25ced

Compare with similar skills

Liteparse next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Liteparse compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Liteparse this skillsammcj/agentic-coding162—~1.1kAutomated safety check: PassApache-2.0
MarkitdownImCa0/just-laws78214 repos~3.2kAutomated safety check: NotesMIT
Markitdownjimmc414/Kosmos5952 repos~1.7kAutomated safety check: PassNone
Liteparsebastani-inc/atomic850—~1.4kAutomated safety check: PassMIT
Nutrient Document Processingaffaan-m/ECC276k4 repos~1.5kAutomated safety check: PassMIT
Markdown Converterintellectronica/agent-skills2954 repos~492Automated safety check: PassCC0-1.0

Similar skills

  • Markitdown

    ImCa0/just-laws

    Convert files and office documents to Markdown. An agent skill from ImCa0/just-laws.

    782 GitHub starsUsed in 14 repos~3.2k tokens
    Documents & OfficeAuto-check: notes
  • Markitdown

    jimmc414/Kosmos

    Convert various file formats (PDF, Office documents, images, audio, web content, structured data) to Markdown optimized for LLM processing.

    595 GitHub starsUsed in 2 repos~1.7k tokens
    Documents & OfficeAuto-check passed
  • Liteparse

    bastani-inc/atomic

    A skill your agent uses whenever a task involves a document file (PDF, DOCX, PPTX, XLSX, or image) and you need to read it or pull text, tables, or specific values out of it — to answer a question…

    850 GitHub stars~1.4k tokensUpdated today
    Documents & OfficeAuto-check passed
  • Process, convert, OCR, extract, redact, sign, and fill documents using the Nutrient DWS API.

    276k GitHub starsUsed in 4 repos~1.5k tokens
    Documents & OfficeAuto-check passed
  • Markdown Converter

    intellectronica/agent-skills

    Convert documents and files to Markdown using markitdown. An agent skill from intellectronica/agent-skills.

    295 GitHub starsUsed in 4 repos~492 tokens
    Documents & OfficeAuto-check passed
  • Document Converter

    wentorai/Research-Claw

    Convert Office documents (PPTX, DOCX, XLSX, PDF, HTML, CSV, JSON, XML, images) to Markdown using Microsoft MarkItDown.

    858 GitHub stars~1.3k tokensUpdated 1 mo ago
    Documents & OfficeAuto-check passed

More from sammcj/agentic-coding

All 64 skills in this repo
  • Yue2 Music

    sammcj/agentic-coding

    A skill your agent uses when generating songs with YuE2, covering a recording via SheetSage2 audio-to-ABC, editing a score or lyrics with melody preservation, or building a reproducible listening…

    162 GitHub stars~2.3k tokensUpdated today
    Auto-check passed
  • Bento Slides

    sammcj/agentic-coding

    A skill your agent uses when creating or editing Bento (.bento.html) slide decks, including any request for a single-file HTML slide deck.

    162 GitHub stars~2.9k tokensUpdated today
    Auto-check passed
  • Idrive Backup

    sammcj/agentic-coding

    A skill your agent uses whenever the user wants you to manage, discuss or diagnose iDrive Backup configuration on macOS

    162 GitHub stars~1.7k tokensUpdated today
    Auto-check: notes
  • Piper Tts Training

    sammcj/agentic-coding

    Train custom TTS voices for Piper (ONNX format) using fine-tuning or from-scratch approaches.

    162 GitHub stars~1.4k tokensUpdated today
    Auto-check passed
  • PPTX To Md

    sammcj/agentic-coding

    Convert a PPTX slide deck into per-slide markdown that preserves both the verbatim text and the meaning of embedded screenshots, diagrams and charts in their original layout positions.

    162 GitHub stars~1.8k tokensUpdated today
    Auto-check passed
  • Skill Creator Primer

    sammcj/agentic-coding

    You MUST load this skill before the skill-creator skill AND before making ANY change to, or conducting a review of ANY Agent Skill.

    162 GitHub stars~9.8k tokensUpdated today
    Auto-check passed

Questions about Liteparse

What does Liteparse do?

Provides fast document to markdown extraction. An agent skill from sammcj/agentic-coding. Liteparse is an agent skill from sammcj/agentic-coding. Provides fast document to markdown extraction.

When should I use Liteparse?

Liteparse fits situations like: the user asks to parse; perform multi-format document conversion; spatially extract text from an unstructured file (PDF.

How do I install Liteparse in Claude Code?

Run `npx skills add sammcj/agentic-coding --skill liteparse -a claude-code`. Or copy the skill folder (Skills/liteparse in sammcj/agentic-coding) into .claude/skills/liteparse in your project. Claude Code loads it when a task matches its description.

How do I install Liteparse in Codex?

Run `npx skills add sammcj/agentic-coding --skill liteparse -a codex`. Or copy the skill folder (Skills/liteparse in sammcj/agentic-coding) into .agents/skills/liteparse in your project. Codex loads it when a task matches its description.

Can I use Liteparse in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add sammcj/agentic-coding --skill liteparse -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/liteparse, .gemini/skills/liteparse, .github/skills/liteparse and .opencode/skills/liteparse in your project.

What does Liteparse need to run?

Going by SKILL.md and its folder, Liteparse needs the command-line tools its instructions call (npx and pnpm). Our summary lists: Node.js.

Does Liteparse access the network?

SKILL.md contains no URLs. Its commands use npx, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Liteparse safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Liteparse use?

Liteparse is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Liteparse use?

About 1.1k tokens (SKILL.md is roughly 4.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Liteparse?

Skills that share tags, products or a category with Liteparse: Markitdown (ImCa0/just-laws, 782 stars), Markitdown (jimmc414/Kosmos, 595 stars), Liteparse (bastani-inc/atomic, 850 stars) and Nutrient Document Processing (affaan-m/ECC, 276k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Liteparse?

sammcj (a GitHub user) maintains it in sammcj/agentic-coding, which has 162 GitHub stars. The repository holds 64 skills in this directory. The repository was last updated on October 9, 2026.

Source: sammcj/agentic-coding on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.