Markitdown
ImCa0/just-laws
Convert files and office documents to Markdown. An agent skill from ImCa0/just-laws.
Convert born-digital, scanned, or mixed PDFs into auditable Markdown while preserving reading order, equations, source-page anchors, and information-bearing images as adjacent non-original text…
$ npx skills add libnyx/LT2MD --skill lt2md -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install libnyx/LT2MD lt2md --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
Claude Code skills documentation · loads skills from .claude/skills/
Install the "lt2md" agent skill from https://github.com/libnyx/LT2MD/tree/master into .claude/skills/lt2md/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "lt2md", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add libnyx/LT2MD --skill lt2md -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install libnyx/LT2MD lt2md --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "lt2md" agent skill from https://github.com/libnyx/LT2MD/tree/master into .agents/skills/lt2md/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "lt2md", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add libnyx/LT2MD --skill lt2md -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install libnyx/LT2MD lt2md --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "lt2md" agent skill from https://github.com/libnyx/LT2MD/tree/master into .cursor/skills/lt2md/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "lt2md", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add libnyx/LT2MD --skill lt2md -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install libnyx/LT2MD lt2md --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "lt2md" agent skill from https://github.com/libnyx/LT2MD/tree/master into .gemini/skills/lt2md/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "lt2md", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install libnyx/LT2MD lt2mdInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add libnyx/LT2MD --skill lt2md -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "lt2md" agent skill from https://github.com/libnyx/LT2MD/tree/master into .github/skills/lt2md/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "lt2md", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add libnyx/LT2MD --skill lt2md -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install libnyx/LT2MD lt2md --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "lt2md" agent skill from https://github.com/libnyx/LT2MD/tree/master into .opencode/skills/lt2md/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "lt2md", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
lt2mdConvert born-digital, scanned, or mixed PDFs into auditable Markdown while preserving reading order, equations, source-page anchors, and information-bearing images as adjacent non-original text…
Lt2md is an agent skill from libnyx/LT2MD. Convert born-digital, scanned, or mixed PDFs into auditable Markdown while preserving reading order, equations, source-page anchors, and information-bearing images as adjacent non-original text descriptions. Use this skill whenever a user asks to transcribe, OCR, understand, or convert a PDF into Markdown, especially for scanned PDFs, image-heavy pages, formulas, multi-column layouts, page or section ranges, or token-efficient reuse. LT2MD (Long Transcribe to Markdown) is a workflow contract, not a replacement…
Its SKILL.md is about 4.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 39 other files, including scripts and reference files (for example `CHANGELOG.md`, `CONTRIBUTING.md` and `README.md`).
It sits in Documents & Office, covering PDF, Transcription and Markdown. The repository describes itself as: AI-agent skill producing reusable Markdown from PDFs. It turns flowcharts, diagrams, and charts into text beside each caption instead of empty links. It checks an earlier… The licence is AGPL-3.0.
5 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit c26a18e. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 1 file in scripts/, which the agent can run.
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Lt2md loads about 4.4k tokens when it runs, and up to ~26k if it reads all its reference files. Until then it costs about 140 tokens; SKILL.md has 2,345 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from libnyx/LT2MD at commit c26a18e, republished under its AGPL-3.0 licence (© libnyx). 2,345 words, ~4,417 tokens.
.claude/skills/lt2md/SKILL.md (or your agent's skills folder). This skill also uses 35 other files; get the full folder from GitHub.LT2MD turns observable PDF content into Markdown that an agent or a person can audit later. It is designed for born-digital, scanned, and mixed PDFs. The goal is not merely to obtain text: preserve reading order, formulas, figure meaning, scope boundaries, and a path back to the source page.
The PDF remains the only authority for content. OCR, extracted text, model guesses, and formatting preferences are candidates or transformations, never evidence that can overrule the rendered page.
scripts/manage_job.py.scripts/validate_markdown.py, scripts/audit_markdown.py, and scripts/manage_job.py verify as separate final gates. Do not place OCR, model calls, or PDF interpretation inside the static tools.SOURCE HTML comment on its own line before every complete paragraph, display equation, figure block, table or example block. Do not insert an anchor inside a word, sentence, inline formula, display-math block, table row, caption or image description. A cross-page block uses one physical-page range before the merged block.转录注 with the exact page and ambiguity. Never silently normalize an uncertain value into a familiar one.manage_job.py set-book-page-offset <job> --offset <N>; otherwise retain unmapped rather than guessing. Once recorded, that mapping is source evidence: the batch scaffold's source_print_pages and every SOURCE BOOK_PAGE must follow it, and manager review/checkpoint/finalization rejects contradictions.manage_job.py batch-plan <job> --json after initialization, then visually lower any recommendation that contains formulas, tables, multi-column order, dense figures, poor legibility, or a cross-page semantic block. The raster-only plan is a conservative starting point, not visual proof. Prefer complete paragraphs, sections, or examples as cut points; keep a sentence crossing a page boundary with one transcriber.manage_job.py source-inventory-template, fill only source objects and evidence, then freeze it with manage_job.py seal-source-inventory. Only after that seal may the transcriber use manage_job.py batch-template --author-id <transcriber> to create a fresh, non-overwriting batch-scoped candidate. This order is a hard gate: candidate block IDs, review decisions and candidate text must not be retrofitted into the source inventory. Never copy an unreviewed full-document V1 draft into the batch candidate and mistake a whole-document audit failure for a batch transcription attempt. Separate body text, equations, figures, captions, examples, headers, footers, and scan noise. Preserve literal Markdown backslashes while writing formulas: an escape-interpreting string layer must not turn a formula command into TAB, FF, or another C0 control byte. Merge only print line breaks and cross-page continuation; do not insert a page boundary inside a word, sentence, or LaTeX expression. An existing Markdown draft is an untrusted candidate, not evidence: visually re-check every retained block. If an inventory item has no source-grounded candidate block, leave the batch blocked; do not omit it merely because the candidate lacks an anchor. If a block is left unchanged, preserve page-specific review evidence; if the page cannot be read, stop there rather than calling the unchanged draft complete.manage_job.py reviewer-handoff. The manager, not reviewer-supplied JSON, owns the reviewer actor ID, local security-principal record, candidate digest, and sealed-inventory binding. The default policy is an auditable process handoff: it does not prove subjective independence merely because labels differ. An optional init --review-identity-policy os-security-principal-v1 also requires the reviewer process to use a different local OS security principal from the candidate and source-inventory authoring processes; it still cannot prove distinct people or model contexts. The handoff reviewer re-reads the rendered source and completes mappings against the already sealed source-only inventory. The reviewer may add candidate mappings, dispositions and risk closures, but may not rewrite sealed source facts. The coordinator changes content only after confirming the source. An omitted footnote, caption, heading, or cross-page continuation remains blocking even when static Markdown checks pass.checkpoint-review JSON manifest through manage_job.py checkpoint before starting later pages. Use manage_job.py review-template only after the sealed-inventory-backed candidate passes the static contract and its exact-byte reviewer handoff is recorded; it produces a blocked identity/hash scaffold and does not replace source review. The manager rejects a missing handoff, a stale candidate digest, a forged reviewer label/principal, an indented-code pseudo-anchor, a review block spanning multiple SOURCE blocks, or a structural modification hidden by whitespace normalization. A failed static check, audit, source-inventory mapping, or independent review is a stop condition: repair the same batch or leave it explicitly incomplete; never treat a failure report as permission to continue. If a source object visibly continues to the next physical page before any independent review, do not accept the short batch or anchor a fragment. Use manage_job.py extend-unclosed-source-inventory only to preserve its sealed source facts and exact unclosed candidate while expanding the same-start range to at most six pages; then re-inventory every page, create a fresh candidate, and complete the normal independent review. This extension is blocked evidence, never acceptance, and cannot change a checkpointed range. When a reviewer supplies source-grounded omissions, misreads, ordering defects, or wrong-page anchors, return only that batch and the exact evidence to the transcriber, then obtain a new independent reread—never relabel the old review as accepted. If that review proves the source-only inventory facts themselves are incomplete or wrong, do not mutate the old seal: use manage_job.py source-inventory-revision-template with that independent blocked review, reread and seal the new source-only inventory, then create a fresh replacement candidate for the same range. The manager freezes a SHA-named copy of the blocked candidate and review; the new candidate receipt must bind the new active seal, and verification checks both the forward and backward revision chain. It rejects self-review, stale candidate replay, altered lineage, and unchanged source facts; a blocked review can never become acceptance. For jobs created before frozen-candidate evidence existed, use the strict manage_job.py backfill-revision-evidence <job> --pages <range> migration only when the preserved bytes, hashes, receipt, old seal and blocked review agree exactly. Every 16 accepted physical pages or 4 accepted batches, whichever comes first, actually reread task brief, render manifest, progress, frozen evidence and risk queue, then record the receipt with manage_job.py reread. This longer cadence supplements, rather than replaces, the per-batch evidence checkpoint; a cache-only or partial job must not pass manage_job.py verify.FORMAT_OK or strict FORMAT_CHANGE records. The reviewer must not read the source PDF or change content.--before <snapshot>. All content projections and source-anchor values must remain unchanged.final-review-template, obtain an independent source-page acceptance, then use manage_job.py finalize; verify with manage_job.py verify --require-finalization. If interruption leaves a prepared promotion, do not start another finalization: run manage_job.py recover-finalization <job> and re-verify. Report the final Markdown path, converted range, cache/job identity, validation result and any 转录注. Do not deliver internal candidate drafts unless the user asks for them.Follow the four roles in workflow.md:
转录注.LT2MD itself is a text workflow and does not require a GPU. Its local page renderer uses CPython 3.10+, pypdfium2 and Pillow; these are PDF/image dependencies, not OCR engines. It does not require Tesseract, PaddleOCR, Poppler or another dedicated OCR executable. The actual scan transcription and image-description quality still depends on a host agent with usable visual reading capability. CPU-only execution is allowed, but it may be slower and a host without a usable visual backend cannot promise accurate scan transcription or image descriptions. Do not describe LT2MD as an unconditional guarantee that every computer can complete every PDF.
The final file must be UTF-8 Markdown with the YAML fields, block-level page anchors, LaTeX delimiters, figure-caption/description adjacency, example structure, uncertainty notes, and strict range termination required by markdown-contract.md. It should be self-contained and not depend on external image files unless the user explicitly requests image assets. Cache PNGs remain in the local workspace by default.
The static validator, risk audit and job verifier are separate format/state gates. None proves that the text, equations, or image descriptions are factually correct; that requires the source-grounded visual reviews above.
Before sealing each source inventory, use cached predecessor/successor pages as independent semantic-boundary evidence. They are context only, never automatic output pages: a successor continuation requires bounded expansion while still unsealed; a predecessor continuation blocks the job rather than rewriting any accepted checkpoint. See references/workflow.md and references/job-state.md.
© libnyx, AGPL-3.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 35 other files (scripts, references) in the repository root of libnyx/LT2MD.
Open the folder on GitHubat commit c26a18e
Lt2md next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Lt2md this skilllibnyx/LT2MD | 109 | — | ~4.4k | Automated safety check: Pass | AGPL-3.0 | |
| MarkitdownImCa0/just-laws | 781 | 14 repos | ~3.2k | Automated safety check: Notes | MIT | |
| Huashu Markdown Publishing Pipelinealchaincyf/huashu-md-html | 908 | — | ~4.8k | Automated safety check: Pass | MIT | |
| Markitdownjimmc414/Kosmos | 594 | 2 repos | ~1.7k | Automated safety check: Pass | None | |
| MineruNebutra/MinerU-Skill | 122 | — | ~504 | Automated safety check: Pass | MIT | |
| Doc To Markdowndaymade/claude-code-skills | 1.4k | — | ~2.5k | Automated safety check: Pass | MIT |
ImCa0/just-laws
Convert files and office documents to Markdown. An agent skill from ImCa0/just-laws.
alchaincyf/huashu-md-html
Converts files and web pages into clean Markdown, then turns Markdown into polished HTML, Word, PDF and EPUB using four templates.
jimmc414/Kosmos
Convert various file formats (PDF, Office documents, images, audio, web content, structured data) to Markdown optimized for LLM processing.
Nebutra/MinerU-Skill
An AI-Native skill for parsing PDF / Office / image files into Markdown with MinerU — a fast, zero-config document parser for AI agents.
daymade/claude-code-skills
Converts DOCX/PDF/PPTX and saved HTML/HTM to high-quality Markdown with automatic post-processing.
intellectronica/agent-skills
Convert documents and files to Markdown using markitdown. An agent skill from intellectronica/agent-skills.
Categories
Convert born-digital, scanned, or mixed PDFs into auditable Markdown while preserving reading order, equations, source-page anchors, and information-bearing images as adjacent non-original text…. Lt2md is an agent skill from libnyx/LT2MD. Convert born-digital, scanned, or mixed PDFs into auditable Markdown while preserving reading order, equations, source-page anchors, and information-bearing images as adjacent non-original text descriptions.
Lt2md fits situations like: A user asks to transcribe; convert a PDF into Markdown; especially for scanned PDFs; image-heavy pages.
Run `npx skills add libnyx/LT2MD --skill lt2md -a claude-code`. Or copy the skill folder (the libnyx/LT2MD repository) into .claude/skills/lt2md in your project. Claude Code loads it when a task matches its description.
Run `npx skills add libnyx/LT2MD --skill lt2md -a codex`. Or copy the skill folder (the libnyx/LT2MD repository) into .agents/skills/lt2md in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add libnyx/LT2MD --skill lt2md -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/lt2md, .gemini/skills/lt2md, .github/skills/lt2md and .opencode/skills/lt2md in your project.
SKILL.md names no scripts, command-line tools or credentials: Lt2md is instructions for the agent only.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Lt2md is published under the AGPL-3.0 licence (from the LICENSE file in the skill folder). It allows redistribution, so the full SKILL.md is shown on this page.
About 4.4k tokens (SKILL.md is roughly 18k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 22k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Lt2md: Markitdown (ImCa0/just-laws, 781 stars), Huashu Markdown Publishing Pipeline (alchaincyf/huashu-md-html, 908 stars), Markitdown (jimmc414/Kosmos, 594 stars) and Mineru (Nebutra/MinerU-Skill, 122 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
libnyx (a GitHub user) maintains it in libnyx/LT2MD, which has 109 GitHub stars. The repository was last updated on August 23, 2026.
Source: libnyx/LT2MD on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.