Mineru
Nebutra/MinerU-Skill
An AI-Native skill for parsing PDF / Office / image files into Markdown with MinerU — a fast, zero-config document parser for AI agents.
Extract single figures from PDF textbooks on-demand with built-in QC verification.
$ npx skills add drpwchen/textbook-to-note --skill figure-remap -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install drpwchen/textbook-to-note figure-remap --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/drpwchen/textbook-to-note.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/figure-remap .claude/skills/figure-remap && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "figure-remap" agent skill from https://github.com/drpwchen/textbook-to-note/tree/main/skills/figure-remap into .claude/skills/figure-remap/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "figure-remap", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/drpwchen/textbook-to-note/tree/main/skills/figure-remapType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add drpwchen/textbook-to-note --skill figure-remap -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install drpwchen/textbook-to-note figure-remap --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/drpwchen/textbook-to-note.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/figure-remap .agents/skills/figure-remap && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "figure-remap" agent skill from https://github.com/drpwchen/textbook-to-note/tree/main/skills/figure-remap into .agents/skills/figure-remap/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "figure-remap", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add drpwchen/textbook-to-note --skill figure-remap -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install drpwchen/textbook-to-note figure-remap --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/drpwchen/textbook-to-note.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/figure-remap .cursor/skills/figure-remap && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "figure-remap" agent skill from https://github.com/drpwchen/textbook-to-note/tree/main/skills/figure-remap into .cursor/skills/figure-remap/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "figure-remap", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/drpwchen/textbook-to-note.git --path skills/figure-remap--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add drpwchen/textbook-to-note --skill figure-remap -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install drpwchen/textbook-to-note figure-remap --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/drpwchen/textbook-to-note.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/figure-remap .gemini/skills/figure-remap && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "figure-remap" agent skill from https://github.com/drpwchen/textbook-to-note/tree/main/skills/figure-remap into .gemini/skills/figure-remap/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "figure-remap", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install drpwchen/textbook-to-note figure-remapInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add drpwchen/textbook-to-note --skill figure-remap -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/drpwchen/textbook-to-note.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/figure-remap .github/skills/figure-remap && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "figure-remap" agent skill from https://github.com/drpwchen/textbook-to-note/tree/main/skills/figure-remap into .github/skills/figure-remap/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "figure-remap", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add drpwchen/textbook-to-note --skill figure-remap -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install drpwchen/textbook-to-note figure-remap --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/drpwchen/textbook-to-note.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/figure-remap .opencode/skills/figure-remap && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "figure-remap" agent skill from https://github.com/drpwchen/textbook-to-note/tree/main/skills/figure-remap into .opencode/skills/figure-remap/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "figure-remap", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
figure-remapExtract single figures from PDF textbooks on-demand with built-in QC verification.
Figure Remap is an agent skill from drpwchen/textbook-to-note. Extract single figures from PDF textbooks on-demand with built-in QC verification. Primary use: when writing/supplementing a note that needs an embedded figure (anatomy, classification, algorithm, imaging). Each call checks an existing fast-path crop, re-extracts from PDF if missing/wrong, retries with local-vision guidance, and escalates to a frontier-model vision read only if all else fails. Also contains legacy batch tools for whole-book re-extraction. Trigger: "extract figure", "figure for note", or…
Its SKILL.md is about 6.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Documents & Office, covering PDF. It works with Obsidian. The repository describes itself as: Turn your own PDF textbooks into an AI-searchable knowledge base and structured, fully-cited notes — figures included. Local-first, token-frugal. The licence is MIT.
3 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 65e7690. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
pythonFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Figure Remap loads about 6.4k tokens when it runs. Until then it costs about 149 tokens; SKILL.md has 3,050 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from drpwchen/textbook-to-note at commit 65e7690, republished under its MIT licence (© drpwchen). 3,050 words, ~6,441 tokens.
.claude/skills/figure-remap/SKILL.md (or your agent's skills folder).This skill calls scripts in your clone of the textbook-to-note repo. At install time, replace
{REPO}below with the absolute path of the clone.
figure-remap is a multi-backend deterministic document-grounding system, not a "crop utility." The system has six explicit layers:
┌─ Contract ──────────────────────────────────────────┐
│ figure_remap.extract() → {status, match_quality, │
│ hard_fail, file, fig_id, reason} │
├─ Policy ─────────────────────────────────────────────┤
│ strict = deterministic + ambiguity-intolerant │
│ L1 / L2 / L3 hard-fail hierarchy │
├─ Junk pre-gate ─────────────────────────────────────┤
│ pregate.verdict() on a QC-passed crop: │
│ chapter-banner geometry / blank → kill, no retry │
├─ Backend selection ─────────────────────────────────┤
│ Capability-based: _select_backend() │
│ "geometric" | "caption_anchor" | [future] │
├─ Backends ──────────────────────────────────────────┤
│ geometric_match_bbox() — born-digital rasters │
│ caption_anchor_bbox() — scanned / hybrid scans │
│ [future] layout-aware / table-aware │
├─ Analysis & cache ──────────────────────────────────┤
│ PageDerived (per-page, sha1-keyed, policy-versioned)│
│ _render_cache/<sha1>_pg<N>_<dpi>dpi.png │
├─ Debug artifacts ───────────────────────────────────┤
│ <out>.fail.json on every scanned-mode hard_fail │
└─────────────────────────────────────────────────────┘Future backends plug in at the Backend layer with their own capability check and a function returning bbox + ambiguity flags — nothing above changes.
==Architectural freeze rule==: do not add per-book heuristics or
if scanned/elif born-digital branches into the gate function. New behavior
goes into a new backend plus a new capability check in _select_backend.
_select_backend() routes pages by structural capability, not by document
type:
| Category | Signature | Backend |
|---|---|---|
| Pure born-digital | ≥2 separable rasters, no full-page background, native fonts | geometric (0.95 confidence) |
| Born-digital text page | 1 small raster + many text blocks + embedded fonts | geometric (0.75 confidence) |
| Hybrid scan (near-full-page raster + embedded figure rasters + OCR overlay) | e.g. a scanned reference book with an OCR text layer laid over the scan | caption_anchor (0.92 confidence) |
| Pure scanned | Near-full-page raster only, no embedded figures, sparse OCR text | caption_anchor (0.9 confidence) |
The hybrid-scan category is the one that surprises naive backend selection: the PDF library reports multiple "assignable" rasters (the embedded figure photos), but the surrounding OCR-overlay text contaminates the geometric backend's text-bleed QC check. Detecting the page-background raster and routing to caption-anchor is the fix.
Module: figure_scanned.py. Algorithm:
Fig X-Y caption text
bounding boxes (with an inline-reference filter — must start at a line
beginning, not mid-sentence after a word).l/I → 1, S → 5, O → 0) only inside the numeric
portion of the fig_id.L1_no_caption_match hard_failL1_target_matches_multiple_captions hard_fail.fail.json debug artifact with policy
version, captions, columns, backend reasoning.Known Phase-1 limitations (deferred):
4-28 vs 4-2B) currently fails
L1_no_caption_match for OCR-mangled targets.fail.json schemaEvery scanned-mode hard_fail writes a sibling <out>.fail.json debug
artifact, with a versioned schema. Stable fields:
| Field | Type | Meaning |
|---|---|---|
schema_version | string | bumps only on breaking schema changes |
policy_version | string | policy that produced this fail |
fig_id_target | string | the normalized fig_id the caller asked for |
fail_reason | string | reason code (e.g., L1_no_caption_match) |
page_idx | int | 0-indexed page |
pdf_sha1 | string | first 12 hex chars of PDF hash |
backend_selected | string | "geometric" or "caption_anchor" |
backend_confidence | float | 0..1 |
backend_reasons | list[string] | one-line tokens explaining the choice |
captions | list[object] | each: {raw, aliases, ambiguity, bbox} |
columns | list[[x0,x1]] | column boundaries in PDF points |
direction | string | null | "above" | "below" | null |
render_path | string | path to the cached rendered page image |
Downstream consumers must check schema_version and refuse to parse
unrecognized majors.
extract-page for scanned booksWhen a workflow needs all figures on one scanned page, use the batch CLI
subcommand instead of N independent extract calls:
python {REPO}/figures/figure_remap.py extract-page \
--book "BookID_PartI" \
--pdf "<path>.pdf" --page N --out-dir "attachments/topic/" \
--fig-ids "4-1A,4-1B,4-1C,4-1D" \
--name-template "Fig_{fig_id}_BookShort.png"==This is NOT a contract addition== — extract-page is a render-cache-aware
loop that calls extract() once per fig_id. Each emitted JSON line is a full
single-figure contract dict; workflows can substitute N independent extract
calls without behavior change — only speed differs.
There is ONE sanctioned entrypoint for every figure request —
figure_remap.py extract. Workflows MUST go through it and MUST NOT call
the internal gate script directly. The entrypoint is the abstraction
boundary: it returns a stable minimal contract and hides the implementation
(matching method, strict mode, fallback policy, QC ladder), so logging stays
uniform and policy is enforced in one place.
python {REPO}/figures/figure_remap.py extract \
--book "{BookID}" \
--fig-id "X.Y" \
--caption "{caption text from textbook md}" \
--out "attachments/Fig_X-Y_{BookShort}_{topic}.png" \
--pdf "{path/to/textbook.pdf}" \
--page N \
[--source {auto|pdf|existing}] # default auto (see "Source modes" below)
[--existing "{book_dir}/figures/Fig_X-Y.{png,jpeg}"] \
[--no-strict] # opt OUT of deterministic mode into the fallback ladder (discouraged)--source selects where the figure comes from. Default is auto — callers
don't need to think about a pre-extracted fast path; the entrypoint resolves
it from the book convention <OUTPUT_DIR>/<book>/figures/Fig_<id>.*
automatically.
--source | Behaviour | When to use |
|---|---|---|
auto (default) | (1) Try book-convention pre-extract; QC-gate it. (2) On miss/QC fail → fall through to strict geometric match + ±1 page sweep. | Normal note-writing. |
pdf | Skip the existing fast path entirely; ignore even an explicit --existing. Always re-extract from PDF. | Validating a code change, suspect a pre-extract is contaminated. |
existing | Cache-only: hard_fail if the convention path is missing or QC-fails. No PDF fallback. | Trust-the-batch mode after offline validation, or when the PDF is unavailable. |
Contract returned (and ONLY this — never the gate's internal shape):
{ "status": "pass|fail|escalate", "match_quality": "exact|uncertain|failed",
"hard_fail": true|false, "file": "<path|null>", "fig_id": "<normalized>", "reason": "<text>",
"qc_degraded": true|false, "qc_skipped": ["<check name>", "..."] }Those eight keys, no more and no fewer — _validate() in
figures/figure_remap.py raises on any extra or missing key.
qc_degraded / qc_skipped: one of the two BLOCKING checks could not run
(e.g. the source page render was unavailable), so a pass here is weaker
than a fully-gated one. Treat a degraded pass as "embed, but eyeball it". A
skipped advisory check never sets this — it was never gating anything.match_quality: exact = deterministic geometric match · uncertain = a
fallback crop (only with --no-strict) · failed = no crop. ==Branch on
this, never on the engine method.== The precise method is logged only in a
local QC log, deliberately kept out of the contract.status:pass (exit 0) → file is a QC-passed crop; embed it.
==Embed the file string verbatim — never the --out you asked for, never
a name rebuilt from the Fig_{id}_{Book} convention.== file is the
authoritative path and it legitimately differs from --out: the auto fast
path returns the book's pre-extracted image, so a template saying .png can
come back .jpeg. A rebuilt name is derived from the real one, so it reads
correct in review and only a machine catches it. Prove it:python {REPO}/figures/figure_embed_lint.py --notes-dir {NOTES_DIR} "<note.md>"status:fail (exit 1) → deterministic miss / hard_fail. ==NOT a wrong
figure — a correct refusal.== Fix --page, escalate to vision, or leave a
<!-- TODO -->.status:fail with reason starting pregate= (exit 1, hard_fail:false)
→ different animal: extraction and QC both succeeded, then the junk pre-gate
(figures/pregate.py) judged the crop to be a chapter-title banner or a
blank. ==Skip this fig_id entirely — do not retry, do not fix --page, do
not leave a retry TODO.== Retrying only re-crops the same banner. Applies to
born-digital geometric crops; the scanned caption_anchor path and the
existing fast path return no bbox, so the pre-gate abstains there.status:escalate (exit 2) → non-strict only; read the page render, then
re-call with --bbox.Default is strict (deterministic). --no-strict re-enables the
legacy-compatibility fallback ladder, which can yield a plausible-but-wrong
crop — avoid for note-writing.
⚠️ A pass guarantees the right raster for the fig_id, not that the figure depicts what you assume. Still read the real caption text before trusting its content.
figure-remap policy (current)
--no-strict.figure_remap.py extract ONLY; the internal gate script is
private.match_quality, not the engine method.{REPO}/figures/test_contract.py) runs after any gate
edit.extract auto-sweeps ±1 neighbor pages on strict hard_fail before
returning fail. Still deterministic (geometric match only, no vision
model); just widens the search window by 2 pages. See "Page resilience"
below.(A)(B)… or panel X), strict mode is overridden to relaxed
up front — see "Multi-panel figures" below. This is policy-level routing
(a capability mismatch, not a silent fallback): strict geometric match's
"caption owns exactly one raster" assumption is invalid for composite
layouts, so retrying strict can only waste attempts.Composite figures with sub-panels (multi-location diagrams, staged classification figures, surgical step sequences) break a core assumption of strict mode and require ==different handling at three levels==: routing, semantic completeness, and uncertainty reporting.
Strict geometric matching assumes ==one caption owns one raster==.
Multi-panel captions reference sub-regions ((A)/(B)/(C)/(D)/(E)) of a
composite layout — those panels are not independently owned rasters. This
is not an implementation bug; it's a representation mismatch.
Detection triggers on ≥2 distinct (letter) or panel <letter> references
(a single (A) is allowed through — common in non-panel prose). When
triggered, extract overrides strict=False and records
policy=multipanel_caption_relaxed in the contract's reason field.
==Multi-panel figures extracted via raster/xref cropping often lose vector overlays== — panel letter labels, arrows, dotted guides, annotation lines. The crop is technically correct (right page region, right raster) but ==semantically incomplete==: it can be hard to map a sub-region back to its caption text.
Caller doctrine: when embedding a multi-panel figure in a note, write the surrounding description as self-contained — map each panel letter to its content in the prose, so readers can interpret the figure even when overlay labels were lost in extraction.
match_quality: uncertain today collapses three distinct failure modes:
identity uncertainty (vision-model fallback selected the raster, not
deterministically verified), semantic-completeness uncertainty (overlay
labels lost), and geometry uncertainty (crop bbox loose). Caller doctrine:
when match_quality:uncertain is returned, prose-annotate the embed (e.g.
"vision-extracted, recommend visual confirmation") rather than relying on
machine-readable fields.
figure_remap.extract() delegates to the internal gate implementation,
which handles the ladder + QC in one call (first success wins):
0. geometric_match (primary, deterministic, no vision model) —
direction-agnostic caption↔image matching: auto-detects whether captions
sit above/below the figure, assigns each embedded raster to its nearest
caption, crops the raster(s) owned by the target fig_id (union for
multi-panel).
--existing)2 → read the page render, estimate bbox,
re-call the gate with --bboxQC on each candidate: two checks, both pure computation, both block (whitespace fill ≥80%, text-bleed <50 prose chars). No model participates in QC — a model-backed check gave the same crop different verdicts on repeat runs (issue #16), so nothing a vision model says can gate, and the QC chain no longer calls one at all. The only vision call left in the gate is the bbox suggestion in rung 3, which the strict default never reaches; its output is a proposal that must still pass the computed checks, and it is pinned to greedy decoding (temperature 0, top_k 1, fixed seed) so it does not drift between runs. Whether a QC-passed crop is worth embedding is judged downstream by the workflow's frontier-model classification step.
Exit codes: 0 ok (saved to --out) · 2 escalate to frontier vision ·
other → log + skip.
The best-effort ladder is heuristic routing: when geometric matching
misses, it falls through to raw-xref pick / local-vision-guided crop, which
can silently produce a plausible-but-wrong crop (e.g. grabbing the
adjacent figure's raster on a multi-figure page, or any raster when the
caller passes a wrong --page). The entrypoint therefore defaults to the
deterministic contract; the fallback ladder is opt-in via --no-strict.
Behavior under strict mode:
{"status":"fail", "match_quality":"failed", "hard_fail":true, "file":null, "reason":...} with exit 1. It never substitutes a different raster.
(Exit 2 is status:escalate, which strict mode never returns.)match_method (geometric_match | FAIL)
and the per-book QC log records a method counter — but that key is
deliberately not on the contract above; branch on match_quality.When strict HARD_FAILs, the caller decides: fix --page, or escalate to
vision (read the page render → --bbox), or leave a <!-- TODO -->. A
HARD_FAIL is the correct outcome — it converts a silent mislabel into an
explicit miss.
Caption/figure cross-page offsets are common: a caption may be on the page the markdown lists while the raster sits on N±1 (frequent in multi-column layouts where a figure spans a column break). A single-page strict match can't recover this on its own.
figure_remap.py extract handles this transparently:
--page N (strict geometric_match)N-1, then N+1 (still strict, still
geometric_match only)reason annotates which neighbor matchedstatus:fail with reason listing the
neighbor attempts==This is NOT a fallback ladder.== Every attempt is still deterministic geometric matching; no vision model, no raw-xref pick. The contract is unchanged.
--existing is intentionally skipped for neighbor attempts: the fast-path
image is keyed to the original page assertion, so reusing it on a different
page would short-circuit the real geometric retry.
When a true hard_fail returns after the sweep, the caller's next move is:
--page== — the caption may live on a different page than the
markdown saysextract against the best alternative--no-strict) — last resort<!-- TODO --> per the convention
belowfig_id is dash-variant agnostic: normalization folds hyphen, en-dash,
em-dash, figure-dash, horizontal-bar and period to a canonical dotted id, so
64-5 / 64–5 / Fig 64.5 / 30 — 44 all compare equal.
Caption format varies per book — calibrate first to set the correct
--fig-id:
(?:FIGURE|Fig\.?)\s*\S+) to see how captions are writtenfigures/figures_manifest.json's fig_id field if presentfigures/Fig_* filenames==The calibration output is a list to copy from, not a rule to apply.==
--book, --fig-id and --caption are exact-match strings: paste --caption
verbatim from the converted markdown (don't retype it, don't tidy the
spacing, don't translate it) and take --book verbatim from the directory name
under your corpus root. If the fig_id you need isn't in the manifest or the grep
output, ==say so and stop — do not construct one from the pattern==; a
plausible-but-wrong fig_id is what makes geometric matching claim the
neighbouring figure.
Common variants:
FIGURE / Figure / Fig. / Fig / FIG.- / . / en-dash / em-dash5.42a / 5-42A / 5.42 (a) / 5.42, ASome books have only an empty figures/ — the fast path is skipped, and the
gate goes straight to re-extract. For these, supply --pdf + --page
directly; do not pass --existing.
Every figure passes through QC before being saved:
If QC fails on the existing fast-path crop → automatic re-extract. If QC fails on re-extract → local-vision-guided retry with bbox. If QC fails after retries → escalate (exit 2) for a frontier-model read.
When a figure can't be extracted now (PDF missing, caption ambiguous), leave:
<!-- TODO: insert Fig X.Y from {Book} ({reason}) — render PDF page N + crop -->==TODO is the LAST step, not the first.== Before writing this comment the caller MUST have:
extract)--no-strict for a frontier-vision escalationDo not imply a batch run will "fill this in later" — each agent extracts its own figures on demand.
==A failed extract gets a TODO comment, never an embed.== Writing
![[Fig_X-Y_Book.png]] for a figure that never passed QC produces a note that
looks complete and renders a broken link — exactly what figure_embed_lint.py
exists to catch. Run the lint on every note you touched before reporting done.
The original batch workflow is preserved for cases where the core extractor is materially changed and you want to validate the fix across the whole corpus:
python {REPO}/figures/batch_remap.py --list
python {REPO}/figures/batch_remap.py --book NAME --dry
python {REPO}/figures/batch_remap.py --book NAME --applyThis runs caption-based extraction on every chapter, computes QC metrics
(green/yellow/red), and visual-checks a few random samples per book via a
local vision model. See {REPO}/figures/CALIBRATION.md for accumulated
per-book quirks.
Do not run batch as part of normal note-writing — too slow, and the on-demand gate is more reliable for single figures because it can fall back to local-vision-guided bbox retry, which batch mode doesn't try.
Historically some setups also had a separate legacy caption-based batch extractor (
extract_figures.py). It is not shipped in this repo —{REPO}/figures/batch_remap.pydegrades gracefully without it (falls back to the gate's own extraction path per figure).
Everything referenced above lives under {REPO}/figures/:
figure_remap.py — sanctioned entrypoint (extract, extract-page)figure_qc_gate.py — primary internal implementation: single-figure
on-demand extraction + QC gatefigure_scanned.py — scanned/hybrid caption_anchor backendqc_metrics.py — batch QC metric computationvisual_check.py — legacy: local-vision-model sample verification, used
only by the batch tools (never by the gate or the note-writing workflow)visual_check_batch.py — legacy: batch-mode visual-check runnerbatch_remap.py — legacy: whole-book batch re-extractiontest_contract.py — regression test suite for the entrypoint contractCALIBRATION.md — per-book quirks accumulated from batch runs (still
useful for setting --fig-id format expectations)--out
writes a fresh copy to your notes' attachments folder; the original
figures/ (if present) is left untouched.<!-- TODO: ... -->
comment in the note rather than silently skipping.© drpwchen, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in skills/figure-remap of drpwchen/textbook-to-note.
Open the folder on GitHubat commit 65e7690
Figure Remap next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Figure Remap this skilldrpwchen/textbook-to-note | 105 | — | ~6.4k | Automated safety check: Pass | MIT | |
| MineruNebutra/MinerU-Skill | 122 | — | ~504 | Automated safety check: Pass | MIT | |
| Summarizereysu/ai-life-skills | 269 | — | ~7.7k | Automated safety check: Pass | MIT | |
| Study Vaultten-builder/ten-builder | 111 | — | ~718 | Automated safety check: Pass | Custom licence | |
| Doc Importerwpsnote/wpsnote-skills | 179 | — | ~2.7k | Automated safety check: Pass | MIT | |
| PDFaxoviq-ai/synthadoc | 1.6k | — | ~293 | Automated safety check: Pass | AGPL-3.0-or-later |
Nebutra/MinerU-Skill
An AI-Native skill for parsing PDF / Office / image files into Markdown with MinerU — a fast, zero-config document parser for AI agents.
reysu/ai-life-skills
Summarize any content (YouTube video, article, whitepaper/PDF, podcast episode, book chapter, etc.) into a rich Obsidian note with section-by-section breakdowns, wikilinks to all technical concepts…
ten-builder/ten-builder
PDF/문서를 Obsidian StudyVault(학습 노트)로 변환합니다. An agent skill from ten-builder/ten-builder.
wpsnote/wpsnote-skills
将本地文档批量导入到 WPS 笔记。支持扫描 Obsidian Vault、思源笔记、微信公众号存档、 下载目录或任意用户指定目录,自动识别 HTML、Markdown、PDF、DOCX、PPTX、XLSX 等格式, 转换为 WPS 笔记可读内容并保留图片和富文本格式。
axoviq-ai/synthadoc
Extract text from PDF documents
tjxj/z-skills
将 Markdown、Obsidian 笔记、技术文章、白皮书或长文排成中文 PDF。用户说“转成 PDF”“Markdown 转 PDF”“md 转 pdf”“排版成电子书”“生成指定风格 PDF”“一次生成多种风格”“学术风/书籍风/报告风”,或要编译 my-girlfriend-jingtian-latex 项目时,都应使用本…
drpwchen/textbook-to-note
Convert PDF/EPUB textbooks to searchable markdown files for an AI agent's own reference.
Works with
Categories
Extract single figures from PDF textbooks on-demand with built-in QC verification. Figure Remap is an agent skill from drpwchen/textbook-to-note. Extract single figures from PDF textbooks on-demand with built-in QC verification.
Figure Remap fits situations like: tasks that involve PDF.
Run `npx skills add drpwchen/textbook-to-note --skill figure-remap -a claude-code`. Or copy the skill folder (skills/figure-remap in drpwchen/textbook-to-note) into .claude/skills/figure-remap in your project. Claude Code loads it when a task matches its description.
Run `npx skills add drpwchen/textbook-to-note --skill figure-remap -a codex`. Or copy the skill folder (skills/figure-remap in drpwchen/textbook-to-note) into .agents/skills/figure-remap in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add drpwchen/textbook-to-note --skill figure-remap -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/figure-remap, .gemini/skills/figure-remap, .github/skills/figure-remap and .opencode/skills/figure-remap in your project.
Going by SKILL.md and its folder, Figure Remap needs the command-line tools its instructions call (python). Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Figure Remap is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 6.4k tokens (SKILL.md is roughly 26k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Figure Remap: Mineru (Nebutra/MinerU-Skill, 122 stars), Summarize (reysu/ai-life-skills, 269 stars), Study Vault (ten-builder/ten-builder, 111 stars) and Doc Importer (wpsnote/wpsnote-skills, 179 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
drpwchen (a GitHub user) maintains it in drpwchen/textbook-to-note, which has 105 GitHub stars. The repository holds 2 skills in this directory. The repository was last updated on August 26, 2026.
Source: drpwchen/textbook-to-note on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.