Agent skill

Liteparse

by bastani-inc in bastani-inc/atomic

A skill your agent uses whenever a task involves a document file (PDF, DOCX, PPTX, XLSX, or image) and you need to read it or pull text, tables, or specific values out of it — to answer a question…

MITAuto-check passedDocuments & Office

Install Liteparse

skills CLI
$ npx skills add bastani-inc/atomic --skill liteparse -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install bastani-inc/atomic liteparse --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/bastani-inc/atomic.git skills-src && mkdir -p .claude/skills && cp -r skills-src/packages/subagents/skills/liteparse .claude/skills/liteparse && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
liteparse
GitHub stars
855
Token cost
~1.4k tokens
SKILL.md length
696 words
Files
2 (incl. scripts)
Skills in repo
15
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses whenever a task involves a document file (PDF, DOCX, PPTX, XLSX, or image) and you need to read it or pull text, tables, or specific values out of it — to answer a question…

  • A task involves a document file (PDF
  • SKILL.md covers The golden rule: parse ONCE to…, Search discipline — minimize…, Ranked search when keywords… and Born-digital vs scanned, plus 4 more sections
  • Runs Python scripts from its folder; calls npm
  • Image) and you need to read it

What it does

Liteparse is an agent skill from bastani-inc/atomic. Use this skill whenever a task involves a document file (PDF, DOCX, PPTX, XLSX, or image) and you need to read it or pull text, tables, or specific values out of it — to answer a question about its contents, look up a figure, or extract data. Provides fast, local, model-free extraction via the lit CLI with disciplined, low-cost search patterns.

Its SKILL.md is about 1.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including scripts (for example `scripts/search.py`). Compatibility notes: Requires Node 18+ and @llamaindex/liteparse (npm i -g @llamaindex/liteparse, verify lit --version). LibreOffice for Office files; ImageMagick for images. The…

It sits in Documents & Office, covering Document parsing, Word documents and PowerPoint presentations. It works with Microsoft Excel, Microsoft PowerPoint and Microsoft Word. The repository describes itself as: The verifiable coding agent runtime. Define your coding agent's process in natural language with stages, checks, and approval gates instead of hoping it follows your… The licence is MIT.

When your agent uses it

  • A task involves a document file (PDF
  • Image) and you need to read it
  • Specific values out of it — to answer a question about its contents
  • Look up a figure

Example prompts

  • “/liteparse”

Requirements

  • Python 3
  • Node.js
  • Compatibility (from SKILL.md): Requires Node 18+ and `@llamaindex/liteparse` (`npm i -g @llamaindex/liteparse`, verify `lit --version`). LibreOffice for Office files; ImageMagick for images. The bundled search.py helper needs `uv`.

What it can do on your machine

Read from SKILL.md and the folder at commit 23dbfd7. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • npm

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use npm, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Requires Node 18+ and `@llamaindex/liteparse` (`npm i -g @llamaindex/liteparse`, verify `lit --version`). LibreOffice for Office files; ImageMagick for images. The bundled search.py helper needs `uv`.

    From compatibility in the SKILL.md frontmatter.

Context cost

Liteparse loads about 1.4k tokens when it runs. Until then it costs about 90 tokens; SKILL.md has 696 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~90
When it runs · the whole SKILL.md, loaded when a task matches
~1.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from bastani-inc/atomic at commit 23dbfd7, republished under its MIT licence (© bastani-inc). 696 words, ~1,447 tokens.

Download SKILL.mdSave it as .claude/skills/liteparse/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
liteparse
description
Use this skill whenever a task involves a document file (PDF, DOCX, PPTX, XLSX, or image) and you need to read it or pull text, tables, or specific values out of it — to answer a question about its contents, look up a figure, or extract data. Provides fast, local, model-free extraction via the `lit` CLI with disciplined, low-cost search patterns.
compatibility
Requires Node 18+ and `@llamaindex/liteparse` (`npm i -g @llamaindex/liteparse`, verify `lit --version`). LibreOffice for Office files; ImageMagick for images. The bundled search.py helper needs `uv`.
license
MIT
metadata.author
LlamaIndex
metadata.version
1.0.1

LiteParse

Extract text from documents locally with the lit CLI — a fast, model-free parser. This skill is about using it cheaply: each lit parse re-runs full extraction, and every line you dump into the conversation is paid for on every subsequent turn. The patterns below come from analyzing real agent traces where the same PDF was parsed up to 9 times and single image reads cost 140k+ characters of context. Don't repeat those mistakes.

The golden rule: parse ONCE to a file, then search the file

lit parse re-extracts the whole document every time you call it. Re-parsing per search is the #1 waste seen in traces. Parse a document exactly once, to a temp file, then run all your searches against that file:

bash
# ONE TIME, per document. --no-ocr for born-digital PDFs (almost all reports) — much faster.
lit parse "/abs/path/doc.pdf" --format text --no-ocr -o /tmp/doc.txt && wc -l /tmp/doc.txt

Then search the file with cheap shell tools — never re-run lit parse to search again.

Search discipline — minimize ROUND-TRIPS, then keep results small

Every Bash call is a full model round-trip (latency + re-read of context). The biggest waste after parsing is a serial loop: grep → look → grep again → sed to read the window → grep again. In traces this doubled the turn count versus just reading the doc. Two rules fix it:

1. Get context in the SAME command — don't grep then sed. Use grep -C so the surrounding lines come back with the hit. This removes the follow-up sed turn for the common case:

bash
grep -n -i -C4 "total assets" /tmp/doc.txt | head -40      # location AND its window, one turn

Only fall back to sed -n 'A,Bp' when you already know the exact line and need a wider window than -C gave you.

2. Batch independent lookups into ONE command. When a question needs several distinct facts (e.g. emissions and revenue), don't spend one turn per term. Probe them together with labels:

bash
for q in "carbon intensity" "scope 1" "total revenue"; do \
  echo "=== $q ==="; grep -n -i -C3 "$q" /tmp/doc.txt | head -25; done

Then keep results small:

  • Always bound output with head and use -n for line numbers.
  • Don't fan out blindly. Aim to resolve a question in ≤3 search commands. If two targeted greps don't pin it down, switch to search.py (below) — don't keep firing keyword variations one per turn.
  • Prefer Bash grep/sed on the saved file over the Read and Grep tools — fewer round-trips and you control output size precisely.

Ranked search when keywords are uncertain (bundled helper)

When two targeted greps haven't pinned the answer, stop greping — don't iterate keyword variants one turn at a time. Run the bundled BM25 ranker ONCE to surface the most relevant line-windows in a single command:

bash
scripts/search.py /tmp/doc.txt -q "materiality assessment priority topics" -k 8 -e 5

-k = number of matches, -e = lines of context around each (so the window comes back inline — no follow-up sed turn). It returns ranked windows with line numbers. Use a rich natural-language query (several synonyms in one string), not a single keyword. This replaces a long chain of speculative greps.

Show full SKILL.md (255 more words)Show less

Born-digital vs scanned

  • Born-digital PDF (real text layer — nearly all corporate/finance/ESG reports): always pass --no-ocr. It's much faster and the text is identical. Leaving OCR on wastes time.
  • Scanned PDF / image: drop --no-ocr. If the value is missing or digits look wrong, read the page visually (see below) rather than trusting OCR.

Reading a page visually — last resort, ONE screenshot, modest DPI

Screenshots are the most expensive thing you can put in context: a single high-DPI page PNG ran ~140k characters in one trace, and agents often rendered the same page twice (default + hi-res).

Only screenshot when text/tables genuinely can't answer the question (dense multi-column tables, figures, charts). Then:

  • Render one page at a time with --target-pages "N" (note: it's --target-pages, NOT --pages).
  • Use modest DPI (~150–200). Do not start at 300+; do not re-render the same page at higher DPI unless the text is actually illegible.
bash
lit screenshot "/abs/path/doc.pdf" --target-pages "13" --dpi 150 -o /tmp/shots/   # then Read the PNG

Many questions about the same document

Parsing once to a file already covers this: keep the /tmp/doc.txt and reuse it across every question instead of re-parsing.

Don't waste turns on preamble

Skip lit --version, ls -la, and lit … --help unless something actually failed. Go straight to the parse. Core flags you need:

--format text|json · --no-ocr · --target-pages "1-5,10" · --dpi <n> (default 150) · --ocr-language <iso>. Use --format json only when you need bounding boxes/layout — it's much larger; still search it, never load it whole.

Setup

PDFs work out of the box. If lit is missing: npm i -g @llamaindex/liteparse. Office docs need LibreOffice; images need ImageMagick (both auto-converted to PDF).

© bastani-inc, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (scripts) in packages/subagents/skills/liteparse of bastani-inc/atomic.

  • SKILL.md
  • scripts/search.py

Open the folder on GitHubat commit 23dbfd7

Compare with similar skills

Liteparse next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Liteparse compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Liteparse this skillbastani-inc/atomic855—~1.4kAutomated safety check: PassMIT
MarkitdownImCa0/just-laws78114 repos~3.2kAutomated safety check: NotesMIT
Markitdownjimmc414/Kosmos5952 repos~1.7kAutomated safety check: PassNone
Nutrient Document Processingaffaan-m/ECC276k4 repos~1.5kAutomated safety check: PassMIT
Markdown Converterintellectronica/agent-skills2954 repos~492Automated safety check: PassCC0-1.0
Document Converterwentorai/Research-Claw858—~1.3kAutomated safety check: PassCustom licence

Similar skills

  • Markitdown

    ImCa0/just-laws

    Convert files and office documents to Markdown. An agent skill from ImCa0/just-laws.

    781 GitHub starsUsed in 14 repos~3.2k tokens
    Documents & OfficeAuto-check: notes
  • Markitdown

    jimmc414/Kosmos

    Convert various file formats (PDF, Office documents, images, audio, web content, structured data) to Markdown optimized for LLM processing.

    595 GitHub starsUsed in 2 repos~1.7k tokens
    Documents & OfficeAuto-check passed
  • Process, convert, OCR, extract, redact, sign, and fill documents using the Nutrient DWS API.

    276k GitHub starsUsed in 4 repos~1.5k tokens
    Documents & OfficeAuto-check passed
  • Markdown Converter

    intellectronica/agent-skills

    Convert documents and files to Markdown using markitdown. An agent skill from intellectronica/agent-skills.

    295 GitHub starsUsed in 4 repos~492 tokens
    Documents & OfficeAuto-check passed
  • Document Converter

    wentorai/Research-Claw

    Convert Office documents (PPTX, DOCX, XLSX, PDF, HTML, CSV, JSON, XML, images) to Markdown using Microsoft MarkItDown.

    858 GitHub stars~1.3k tokensUpdated 1 mo ago
    Documents & OfficeAuto-check passed
  • Learn

    iurykrieger/claude-bedrock

    Ingests an external data source into the Second Brain. An agent skill from iurykrieger/claude-bedrock.

    105 GitHub stars~6.5k tokensUpdated 5 mo ago
    Documents & OfficeAuto-check: notes

More from bastani-inc/atomic

All 15 skills in this repo
  • How

    bastani-inc/atomic

    A skill your agent uses for "how does X work", code walkthroughs before changing something, and placement / ownership / layering questions ("where should this live", "which package owns this", "is…

    855 GitHub starsUsed in 3 repos~2.3k tokens
    Auto-check passed
  • Bun

    bastani-inc/atomic

    A skill your agent uses when building, testing, and deploying JavaScript/TypeScript applications.

    855 GitHub stars~2.8k tokensUpdated today
    Auto-check passed
  • Feedback

    bastani-inc/atomic

    Draft, revise and post privacy-scrubbed Atomic bug reports or enhancements through ordinary conversation.

    855 GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Prek

    bastani-inc/atomic

    A skill your agent uses when setting up or running hooks with prek in any repository.

    855 GitHub stars~2.1k tokensUpdated today
    Auto-check passed
  • TDD

    bastani-inc/atomic

    Test-driven development with red-green-refactor loop, plus a value gate and audit workflow for tests.

    855 GitHub stars~1.4k tokensUpdated today
    Auto-check passed
  • Tmux

    bastani-inc/atomic

    Control tmux-compatible sessions/windows/panes for interactive CLIs: list, capture output, send keys, paste text, monitor prompts.

    855 GitHub stars~1.6k tokensUpdated today
    Auto-check: notes

Questions about Liteparse

What does Liteparse do?

A skill your agent uses whenever a task involves a document file (PDF, DOCX, PPTX, XLSX, or image) and you need to read it or pull text, tables, or specific values out of it — to answer a question…. Liteparse is an agent skill from bastani-inc/atomic. Use this skill whenever a task involves a document file (PDF, DOCX, PPTX, XLSX, or image) and you need to read it or pull text, tables, or specific values out of it — to answer a question about its contents, look up a figure, or extract data.

When should I use Liteparse?

Liteparse fits situations like: A task involves a document file (PDF; image) and you need to read it; specific values out of it — to answer a question about its contents; look up a figure.

How do I install Liteparse in Claude Code?

Run `npx skills add bastani-inc/atomic --skill liteparse -a claude-code`. Or copy the skill folder (packages/subagents/skills/liteparse in bastani-inc/atomic) into .claude/skills/liteparse in your project. Claude Code loads it when a task matches its description.

How do I install Liteparse in Codex?

Run `npx skills add bastani-inc/atomic --skill liteparse -a codex`. Or copy the skill folder (packages/subagents/skills/liteparse in bastani-inc/atomic) into .agents/skills/liteparse in your project. Codex loads it when a task matches its description.

Can I use Liteparse in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add bastani-inc/atomic --skill liteparse -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/liteparse, .gemini/skills/liteparse, .github/skills/liteparse and .opencode/skills/liteparse in your project.

What does Liteparse need to run?

Going by SKILL.md and its folder, Liteparse needs Python for the scripts in its folder and the command-line tools its instructions call (npm). Our summary lists: Python 3; Node.js. Compatibility (from SKILL.md): Requires Node 18+ and `@llamaindex/liteparse` (`npm i -g @llamaindex/liteparse`, verify `lit --version`). LibreOffice for Office files; ImageMagick for images. The bundled search.py helper needs `uv`..

Does Liteparse access the network?

SKILL.md contains no URLs. Its commands use npm, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Liteparse safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Liteparse use?

Liteparse is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Liteparse use?

About 1.4k tokens (SKILL.md is roughly 5.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Liteparse?

Skills that share tags, products or a category with Liteparse: Markitdown (ImCa0/just-laws, 781 stars), Markitdown (jimmc414/Kosmos, 595 stars), Nutrient Document Processing (affaan-m/ECC, 276k stars) and Markdown Converter (intellectronica/agent-skills, 295 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Liteparse?

bastani-inc (a GitHub organization) maintains it in bastani-inc/atomic, which has 855 GitHub stars. The repository holds 15 skills in this directory. The repository was last updated on October 10, 2026.

Source: bastani-inc/atomic on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.