Markitdown
ImCa0/just-laws
Convert files and office documents to Markdown. An agent skill from ImCa0/just-laws.
A skill your agent uses whenever a task involves a document file (PDF, DOCX, PPTX, XLSX, or image) and you need to read it or pull text, tables, or specific values out of it — to answer a question…
$ npx skills add bastani-inc/atomic --skill liteparse -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install bastani-inc/atomic liteparse --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/bastani-inc/atomic.git skills-src && mkdir -p .claude/skills && cp -r skills-src/packages/subagents/skills/liteparse .claude/skills/liteparse && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "liteparse" agent skill from https://github.com/bastani-inc/atomic/tree/main/packages/subagents/skills/liteparse into .claude/skills/liteparse/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "liteparse", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/bastani-inc/atomic/tree/main/packages/subagents/skills/liteparseType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add bastani-inc/atomic --skill liteparse -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install bastani-inc/atomic liteparse --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/bastani-inc/atomic.git skills-src && mkdir -p .agents/skills && cp -r skills-src/packages/subagents/skills/liteparse .agents/skills/liteparse && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "liteparse" agent skill from https://github.com/bastani-inc/atomic/tree/main/packages/subagents/skills/liteparse into .agents/skills/liteparse/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "liteparse", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add bastani-inc/atomic --skill liteparse -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install bastani-inc/atomic liteparse --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/bastani-inc/atomic.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/packages/subagents/skills/liteparse .cursor/skills/liteparse && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "liteparse" agent skill from https://github.com/bastani-inc/atomic/tree/main/packages/subagents/skills/liteparse into .cursor/skills/liteparse/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "liteparse", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/bastani-inc/atomic.git --path packages/subagents/skills/liteparse--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add bastani-inc/atomic --skill liteparse -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install bastani-inc/atomic liteparse --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/bastani-inc/atomic.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/packages/subagents/skills/liteparse .gemini/skills/liteparse && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "liteparse" agent skill from https://github.com/bastani-inc/atomic/tree/main/packages/subagents/skills/liteparse into .gemini/skills/liteparse/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "liteparse", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install bastani-inc/atomic liteparseInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add bastani-inc/atomic --skill liteparse -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/bastani-inc/atomic.git skills-src && mkdir -p .github/skills && cp -r skills-src/packages/subagents/skills/liteparse .github/skills/liteparse && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "liteparse" agent skill from https://github.com/bastani-inc/atomic/tree/main/packages/subagents/skills/liteparse into .github/skills/liteparse/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "liteparse", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add bastani-inc/atomic --skill liteparse -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install bastani-inc/atomic liteparse --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/bastani-inc/atomic.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/packages/subagents/skills/liteparse .opencode/skills/liteparse && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "liteparse" agent skill from https://github.com/bastani-inc/atomic/tree/main/packages/subagents/skills/liteparse into .opencode/skills/liteparse/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "liteparse", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
liteparseA skill your agent uses whenever a task involves a document file (PDF, DOCX, PPTX, XLSX, or image) and you need to read it or pull text, tables, or specific values out of it — to answer a question…
Liteparse is an agent skill from bastani-inc/atomic. Use this skill whenever a task involves a document file (PDF, DOCX, PPTX, XLSX, or image) and you need to read it or pull text, tables, or specific values out of it — to answer a question about its contents, look up a figure, or extract data. Provides fast, local, model-free extraction via the lit CLI with disciplined, low-cost search patterns.
Its SKILL.md is about 1.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including scripts (for example `scripts/search.py`). Compatibility notes: Requires Node 18+ and @llamaindex/liteparse (npm i -g @llamaindex/liteparse, verify lit --version). LibreOffice for Office files; ImageMagick for images. The…
It sits in Documents & Office, covering Document parsing, Word documents and PowerPoint presentations. It works with Microsoft Excel, Microsoft PowerPoint and Microsoft Word. The repository describes itself as: The verifiable coding agent runtime. Define your coding agent's process in natural language with stages, checks, and approval gates instead of hoping it follows your… The licence is MIT.
Read from SKILL.md and the folder at commit 23dbfd7. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 1 file in scripts/ (Python), which the agent can run.
Shell commands in SKILL.md call:
npmFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use npm, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Requires Node 18+ and `@llamaindex/liteparse` (`npm i -g @llamaindex/liteparse`, verify `lit --version`). LibreOffice for Office files; ImageMagick for images. The bundled search.py helper needs `uv`.
From compatibility in the SKILL.md frontmatter.
Liteparse loads about 1.4k tokens when it runs. Until then it costs about 90 tokens; SKILL.md has 696 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from bastani-inc/atomic at commit 23dbfd7, republished under its MIT licence (© bastani-inc). 696 words, ~1,447 tokens.
.claude/skills/liteparse/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.Extract text from documents locally with the lit CLI — a fast, model-free parser. This skill is
about using it cheaply: each lit parse re-runs full extraction, and every line you dump into
the conversation is paid for on every subsequent turn. The patterns below come from analyzing real
agent traces where the same PDF was parsed up to 9 times and single image reads cost
140k+ characters of context. Don't repeat those mistakes.
lit parse re-extracts the whole document every time you call it. Re-parsing per search is the #1
waste seen in traces. Parse a document exactly once, to a temp file, then run all your searches
against that file:
# ONE TIME, per document. --no-ocr for born-digital PDFs (almost all reports) — much faster.
lit parse "/abs/path/doc.pdf" --format text --no-ocr -o /tmp/doc.txt && wc -l /tmp/doc.txtThen search the file with cheap shell tools — never re-run lit parse to search again.
Every Bash call is a full model round-trip (latency + re-read of context). The biggest waste after
parsing is a serial loop: grep → look → grep again → sed to read the window → grep again. In
traces this doubled the turn count versus just reading the doc. Two rules fix it:
1. Get context in the SAME command — don't grep then sed. Use grep -C so the surrounding
lines come back with the hit. This removes the follow-up sed turn for the common case:
grep -n -i -C4 "total assets" /tmp/doc.txt | head -40 # location AND its window, one turnOnly fall back to sed -n 'A,Bp' when you already know the exact line and need a wider window
than -C gave you.
2. Batch independent lookups into ONE command. When a question needs several distinct facts (e.g. emissions and revenue), don't spend one turn per term. Probe them together with labels:
for q in "carbon intensity" "scope 1" "total revenue"; do \
echo "=== $q ==="; grep -n -i -C3 "$q" /tmp/doc.txt | head -25; doneThen keep results small:
head and use -n for line numbers.search.py (below) — don't keep firing keyword variations one per turn.grep/sed on the saved file over the Read and Grep tools — fewer round-trips and
you control output size precisely.When two targeted greps haven't pinned the answer, stop greping — don't iterate keyword variants one turn at a time. Run the bundled BM25 ranker ONCE to surface the most relevant line-windows in a single command:
scripts/search.py /tmp/doc.txt -q "materiality assessment priority topics" -k 8 -e 5-k = number of matches, -e = lines of context around each (so the window comes back inline — no
follow-up sed turn). It returns ranked windows with line numbers. Use a rich natural-language query
(several synonyms in one string), not a single keyword. This replaces a long chain of speculative greps.
--no-ocr. It's much faster and the text is identical. Leaving OCR on wastes time.--no-ocr. If the value is missing or digits look wrong, read the
page visually (see below) rather than trusting OCR.Screenshots are the most expensive thing you can put in context: a single high-DPI page PNG ran ~140k characters in one trace, and agents often rendered the same page twice (default + hi-res).
Only screenshot when text/tables genuinely can't answer the question (dense multi-column tables, figures, charts). Then:
--target-pages "N" (note: it's --target-pages, NOT --pages).lit screenshot "/abs/path/doc.pdf" --target-pages "13" --dpi 150 -o /tmp/shots/ # then Read the PNGParsing once to a file already covers this: keep the /tmp/doc.txt and reuse it across every
question instead of re-parsing.
Skip lit --version, ls -la, and lit … --help unless something actually failed. Go straight to
the parse. Core flags you need:
--format text|json · --no-ocr · --target-pages "1-5,10" · --dpi <n> (default 150) ·
--ocr-language <iso>. Use --format json only when you need bounding boxes/layout — it's much
larger; still search it, never load it whole.
PDFs work out of the box. If lit is missing: npm i -g @llamaindex/liteparse. Office docs need
LibreOffice; images need ImageMagick (both auto-converted to PDF).
© bastani-inc, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 1 other file (scripts) in packages/subagents/skills/liteparse of bastani-inc/atomic.
Open the folder on GitHubat commit 23dbfd7
Liteparse next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Liteparse this skillbastani-inc/atomic | 855 | — | ~1.4k | Automated safety check: Pass | MIT | |
| MarkitdownImCa0/just-laws | 781 | 14 repos | ~3.2k | Automated safety check: Notes | MIT | |
| Markitdownjimmc414/Kosmos | 595 | 2 repos | ~1.7k | Automated safety check: Pass | None | |
| Nutrient Document Processingaffaan-m/ECC | 276k | 4 repos | ~1.5k | Automated safety check: Pass | MIT | |
| Markdown Converterintellectronica/agent-skills | 295 | 4 repos | ~492 | Automated safety check: Pass | CC0-1.0 | |
| Document Converterwentorai/Research-Claw | 858 | — | ~1.3k | Automated safety check: Pass | Custom licence |
ImCa0/just-laws
Convert files and office documents to Markdown. An agent skill from ImCa0/just-laws.
jimmc414/Kosmos
Convert various file formats (PDF, Office documents, images, audio, web content, structured data) to Markdown optimized for LLM processing.
affaan-m/ECC
Process, convert, OCR, extract, redact, sign, and fill documents using the Nutrient DWS API.
intellectronica/agent-skills
Convert documents and files to Markdown using markitdown. An agent skill from intellectronica/agent-skills.
wentorai/Research-Claw
Convert Office documents (PPTX, DOCX, XLSX, PDF, HTML, CSV, JSON, XML, images) to Markdown using Microsoft MarkItDown.
iurykrieger/claude-bedrock
Ingests an external data source into the Second Brain. An agent skill from iurykrieger/claude-bedrock.
bastani-inc/atomic
A skill your agent uses for "how does X work", code walkthroughs before changing something, and placement / ownership / layering questions ("where should this live", "which package owns this", "is…
bastani-inc/atomic
A skill your agent uses when building, testing, and deploying JavaScript/TypeScript applications.
bastani-inc/atomic
Draft, revise and post privacy-scrubbed Atomic bug reports or enhancements through ordinary conversation.
bastani-inc/atomic
A skill your agent uses when setting up or running hooks with prek in any repository.
bastani-inc/atomic
Test-driven development with red-green-refactor loop, plus a value gate and audit workflow for tests.
bastani-inc/atomic
Control tmux-compatible sessions/windows/panes for interactive CLIs: list, capture output, send keys, paste text, monitor prompts.
Categories
A skill your agent uses whenever a task involves a document file (PDF, DOCX, PPTX, XLSX, or image) and you need to read it or pull text, tables, or specific values out of it — to answer a question…. Liteparse is an agent skill from bastani-inc/atomic. Use this skill whenever a task involves a document file (PDF, DOCX, PPTX, XLSX, or image) and you need to read it or pull text, tables, or specific values out of it — to answer a question about its contents, look up a figure, or extract data.
Liteparse fits situations like: A task involves a document file (PDF; image) and you need to read it; specific values out of it — to answer a question about its contents; look up a figure.
Run `npx skills add bastani-inc/atomic --skill liteparse -a claude-code`. Or copy the skill folder (packages/subagents/skills/liteparse in bastani-inc/atomic) into .claude/skills/liteparse in your project. Claude Code loads it when a task matches its description.
Run `npx skills add bastani-inc/atomic --skill liteparse -a codex`. Or copy the skill folder (packages/subagents/skills/liteparse in bastani-inc/atomic) into .agents/skills/liteparse in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add bastani-inc/atomic --skill liteparse -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/liteparse, .gemini/skills/liteparse, .github/skills/liteparse and .opencode/skills/liteparse in your project.
Going by SKILL.md and its folder, Liteparse needs Python for the scripts in its folder and the command-line tools its instructions call (npm). Our summary lists: Python 3; Node.js. Compatibility (from SKILL.md): Requires Node 18+ and `@llamaindex/liteparse` (`npm i -g @llamaindex/liteparse`, verify `lit --version`). LibreOffice for Office files; ImageMagick for images. The bundled search.py helper needs `uv`..
SKILL.md contains no URLs. Its commands use npm, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Liteparse is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.4k tokens (SKILL.md is roughly 5.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Liteparse: Markitdown (ImCa0/just-laws, 781 stars), Markitdown (jimmc414/Kosmos, 595 stars), Nutrient Document Processing (affaan-m/ECC, 276k stars) and Markdown Converter (intellectronica/agent-skills, 295 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
bastani-inc (a GitHub organization) maintains it in bastani-inc/atomic, which has 855 GitHub stars. The repository holds 15 skills in this directory. The repository was last updated on October 10, 2026.
Source: bastani-inc/atomic on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.