Agent skill

Figure Remap

by drpwchen in drpwchen/textbook-to-note

Extract single figures from PDF textbooks on-demand with built-in QC verification.

MITAuto-check passedDocuments & Office

Install Figure Remap

skills CLI
$ npx skills add drpwchen/textbook-to-note --skill figure-remap -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install drpwchen/textbook-to-note figure-remap --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/drpwchen/textbook-to-note.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/figure-remap .claude/skills/figure-remap && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
figure-remap
GitHub stars
105
Token cost
~6.4k tokens
SKILL.md length
3,050 words
Files
1
Skills in repo
2
Repo updated
First seen
Licence
MIT

At a glance

Extract single figures from PDF textbooks on-demand with built-in QC verification.

  • Works in 3 steps: Routing (automated) → Vector overlay loss (caller awareness) → Uncertainty taxonomy (observation)
  • Tasks that involve PDF
  • SKILL.md covers Architecture (read this before…, Backends — when each is used, Scanned/hybrid backend —… and fail.json schema, plus 14 more sections
  • Calls python

What it does

Figure Remap is an agent skill from drpwchen/textbook-to-note. Extract single figures from PDF textbooks on-demand with built-in QC verification. Primary use: when writing/supplementing a note that needs an embedded figure (anatomy, classification, algorithm, imaging). Each call checks an existing fast-path crop, re-extracts from PDF if missing/wrong, retries with local-vision guidance, and escalates to a frontier-model vision read only if all else fails. Also contains legacy batch tools for whole-book re-extraction. Trigger: "extract figure", "figure for note", or…

Its SKILL.md is about 6.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Documents & Office, covering PDF. It works with Obsidian. The repository describes itself as: Turn your own PDF textbooks into an AI-searchable knowledge base and structured, fully-cited notes — figures included. Local-first, token-frugal. The licence is MIT.

When your agent uses it

  • Tasks that involve PDF

Example prompts

  • “extract figure”
  • “figure for note”
  • “/figure-remap”

Requirements

  • Python 3

Workflow steps

3 steps, taken from the step headings in SKILL.md.

  1. Routing (automated)
  2. Vector overlay loss (caller awareness)
  3. Uncertainty taxonomy (observation)

What it can do on your machine

Read from SKILL.md and the folder at commit 65e7690. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Figure Remap loads about 6.4k tokens when it runs. Until then it costs about 149 tokens; SKILL.md has 3,050 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~149
When it runs · the whole SKILL.md, loaded when a task matches
~6.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from drpwchen/textbook-to-note at commit 65e7690, republished under its MIT licence (© drpwchen). 3,050 words, ~6,441 tokens.

Download SKILL.mdSave it as .claude/skills/figure-remap/SKILL.md (or your agent's skills folder).
name
figure-remap
description
Extract single figures from PDF textbooks on-demand with built-in QC verification. Primary use: when writing/supplementing a note that needs an embedded figure (anatomy, classification, algorithm, imaging). Each call checks an existing fast-path crop, re-extracts from PDF if missing/wrong, retries with local-vision guidance, and escalates to a frontier-model vision read only if all else fails. Also contains legacy batch tools for whole-book re-extraction. Trigger: "extract figure", "figure for note", or proactively when a note's TODO/REF references a figure not yet embedded.

figure-remap — deterministic document-grounding system

This skill calls scripts in your clone of the textbook-to-note repo. At install time, replace {REPO} below with the absolute path of the clone.

Architecture (read this before implementation details)

figure-remap is a multi-backend deterministic document-grounding system, not a "crop utility." The system has six explicit layers:

┌─ Contract ──────────────────────────────────────────┐
│  figure_remap.extract() → {status, match_quality,   │
│  hard_fail, file, fig_id, reason}                   │
├─ Policy ─────────────────────────────────────────────┤
│  strict = deterministic + ambiguity-intolerant      │
│  L1 / L2 / L3 hard-fail hierarchy                   │
├─ Junk pre-gate ─────────────────────────────────────┤
│  pregate.verdict() on a QC-passed crop:             │
│  chapter-banner geometry / blank → kill, no retry   │
├─ Backend selection ─────────────────────────────────┤
│  Capability-based: _select_backend()                │
│    "geometric"  | "caption_anchor" | [future]       │
├─ Backends ──────────────────────────────────────────┤
│  geometric_match_bbox()  — born-digital rasters     │
│  caption_anchor_bbox()   — scanned / hybrid scans   │
│  [future] layout-aware / table-aware                │
├─ Analysis & cache ──────────────────────────────────┤
│  PageDerived (per-page, sha1-keyed, policy-versioned)│
│  _render_cache/<sha1>_pg<N>_<dpi>dpi.png            │
├─ Debug artifacts ───────────────────────────────────┤
│  <out>.fail.json on every scanned-mode hard_fail    │
└─────────────────────────────────────────────────────┘

Future backends plug in at the Backend layer with their own capability check and a function returning bbox + ambiguity flags — nothing above changes.

==Architectural freeze rule==: do not add per-book heuristics or if scanned/elif born-digital branches into the gate function. New behavior goes into a new backend plus a new capability check in _select_backend.

Backends — when each is used

_select_backend() routes pages by structural capability, not by document type:

CategorySignatureBackend
Pure born-digital≥2 separable rasters, no full-page background, native fontsgeometric (0.95 confidence)
Born-digital text page1 small raster + many text blocks + embedded fontsgeometric (0.75 confidence)
Hybrid scan (near-full-page raster + embedded figure rasters + OCR overlay)e.g. a scanned reference book with an OCR text layer laid over the scancaption_anchor (0.92 confidence)
Pure scannedNear-full-page raster only, no embedded figures, sparse OCR textcaption_anchor (0.9 confidence)

The hybrid-scan category is the one that surprises naive backend selection: the PDF library reports multiple "assignable" rasters (the embedded figure photos), but the surrounding OCR-overlay text contaminates the geometric backend's text-bleed QC check. Detecting the page-background raster and routing to caption-anchor is the fix.

Scanned/hybrid backend — caption_anchor

Module: figure_scanned.py. Algorithm:

  1. Render the page at 200 dpi to a cache file (keyed by PDF hash + page + dpi).
  2. Walk the page's text-with-position data for Fig X-Y caption text bounding boxes (with an inline-reference filter — must start at a line beginning, not mid-sentence after a word).
  3. Apply Class A OCR normalization (silent, 1-to-1): common OCR confusions (l/I → 1, S → 5, O → 0) only inside the numeric portion of the fig_id.
  4. Build a caption-match object per caption with an alias set.
  5. Look up the target fig_id across all caption alias sets:
    • 0 matches → L1_no_caption_match hard_fail
    • ≥2 matches → L1_target_matches_multiple_captions hard_fail
  6. Infer page-global caption direction (caption-below is the layout default).
  7. Compute the figure region: the area opposite the caption, bounded by the previous/next caption in the same column.
  8. Crop the rendered page; run an ink-density trim to tighten the bbox.
  9. Scanned QC: whitespace-fill check only (text-bleed is meaningless on scanned PDFs — the OCR overlay covers the whole page).
  10. On hard_fail: write a sibling .fail.json debug artifact with policy version, captions, columns, backend reasoning.

Known Phase-1 limitations (deferred):

  • Class B OCR ambiguity (e.g. 4-28 vs 4-2B) currently fails L1_no_caption_match for OCR-mangled targets.
  • L2 ownership-overlap guard, L3 direction-quality guard not yet built.
  • Direction inference is page-global, not per-column.

fail.json schema

Every scanned-mode hard_fail writes a sibling <out>.fail.json debug artifact, with a versioned schema. Stable fields:

FieldTypeMeaning
schema_versionstringbumps only on breaking schema changes
policy_versionstringpolicy that produced this fail
fig_id_targetstringthe normalized fig_id the caller asked for
fail_reasonstringreason code (e.g., L1_no_caption_match)
page_idxint0-indexed page
pdf_sha1stringfirst 12 hex chars of PDF hash
backend_selectedstring"geometric" or "caption_anchor"
backend_confidencefloat0..1
backend_reasonslist[string]one-line tokens explaining the choice
captionslist[object]each: {raw, aliases, ambiguity, bbox}
columnslist[[x0,x1]]column boundaries in PDF points
directionstring | null"above" | "below" | null
render_pathstringpath to the cached rendered page image

Downstream consumers must check schema_version and refuse to parse unrecognized majors.

Performance helper — extract-page for scanned books

When a workflow needs all figures on one scanned page, use the batch CLI subcommand instead of N independent extract calls:

bash
python {REPO}/figures/figure_remap.py extract-page \
  --book "BookID_PartI" \
  --pdf "<path>.pdf" --page N --out-dir "attachments/topic/" \
  --fig-ids "4-1A,4-1B,4-1C,4-1D" \
  --name-template "Fig_{fig_id}_BookShort.png"

==This is NOT a contract addition== — extract-page is a render-cache-aware loop that calls extract() once per fig_id. Each emitted JSON line is a full single-figure contract dict; workflows can substitute N independent extract calls without behavior change — only speed differs.

Entrypoint contract (call THIS, not the gate)

There is ONE sanctioned entrypoint for every figure request — figure_remap.py extract. Workflows MUST go through it and MUST NOT call the internal gate script directly. The entrypoint is the abstraction boundary: it returns a stable minimal contract and hides the implementation (matching method, strict mode, fallback policy, QC ladder), so logging stays uniform and policy is enforced in one place.

bash
python {REPO}/figures/figure_remap.py extract \
  --book "{BookID}" \
  --fig-id "X.Y" \
  --caption "{caption text from textbook md}" \
  --out "attachments/Fig_X-Y_{BookShort}_{topic}.png" \
  --pdf "{path/to/textbook.pdf}" \
  --page N \
  [--source {auto|pdf|existing}]   # default auto (see "Source modes" below)
  [--existing "{book_dir}/figures/Fig_X-Y.{png,jpeg}"] \
  [--no-strict]      # opt OUT of deterministic mode into the fallback ladder (discouraged)
Source modes

--source selects where the figure comes from. Default is auto — callers don't need to think about a pre-extracted fast path; the entrypoint resolves it from the book convention <OUTPUT_DIR>/<book>/figures/Fig_<id>.* automatically.

--sourceBehaviourWhen to use
auto (default)(1) Try book-convention pre-extract; QC-gate it. (2) On miss/QC fail → fall through to strict geometric match + ±1 page sweep.Normal note-writing.
pdfSkip the existing fast path entirely; ignore even an explicit --existing. Always re-extract from PDF.Validating a code change, suspect a pre-extract is contaminated.
existingCache-only: hard_fail if the convention path is missing or QC-fails. No PDF fallback.Trust-the-batch mode after offline validation, or when the PDF is unavailable.

Contract returned (and ONLY this — never the gate's internal shape):

json
{ "status": "pass|fail|escalate", "match_quality": "exact|uncertain|failed",
  "hard_fail": true|false, "file": "<path|null>", "fig_id": "<normalized>", "reason": "<text>",
  "qc_degraded": true|false, "qc_skipped": ["<check name>", "..."] }

Those eight keys, no more and no fewer — _validate() in figures/figure_remap.py raises on any extra or missing key.

  • qc_degraded / qc_skipped: one of the two BLOCKING checks could not run (e.g. the source page render was unavailable), so a pass here is weaker than a fully-gated one. Treat a degraded pass as "embed, but eyeball it". A skipped advisory check never sets this — it was never gating anything.
  • match_quality: exact = deterministic geometric match · uncertain = a fallback crop (only with --no-strict) · failed = no crop. ==Branch on this, never on the engine method.== The precise method is logged only in a local QC log, deliberately kept out of the contract.
  • status:pass (exit 0) → file is a QC-passed crop; embed it. ==Embed the file string verbatim — never the --out you asked for, never a name rebuilt from the Fig_{id}_{Book} convention.== file is the authoritative path and it legitimately differs from --out: the auto fast path returns the book's pre-extracted image, so a template saying .png can come back .jpeg. A rebuilt name is derived from the real one, so it reads correct in review and only a machine catches it. Prove it:
    bash
    python {REPO}/figures/figure_embed_lint.py --notes-dir {NOTES_DIR} "<note.md>"
    exit 0 clean · 1 = MISSING (a guessed filename, or an embed written for an extract that actually failed) or CASE MISMATCH (opens on Windows/macOS, breaks on git and Linux) · 3 = notes dir not found = unverifiable, not a pass.
  • status:fail (exit 1) → deterministic miss / hard_fail. ==NOT a wrong figure — a correct refusal.== Fix --page, escalate to vision, or leave a <!-- TODO -->.
  • status:fail with reason starting pregate= (exit 1, hard_fail:false) → different animal: extraction and QC both succeeded, then the junk pre-gate (figures/pregate.py) judged the crop to be a chapter-title banner or a blank. ==Skip this fig_id entirely — do not retry, do not fix --page, do not leave a retry TODO.== Retrying only re-crops the same banner. Applies to born-digital geometric crops; the scanned caption_anchor path and the existing fast path return no bbox, so the pre-gate abstains there.
  • status:escalate (exit 2) → non-strict only; read the page render, then re-call with --bbox.

Default is strict (deterministic). --no-strict re-enables the legacy-compatibility fallback ladder, which can yield a plausible-but-wrong crop — avoid for note-writing.

⚠️ A pass guarantees the right raster for the fig_id, not that the figure depicts what you assume. Still read the real caption text before trusting its content.

Policy version

figure-remap policy (current)

  • Deterministic mode (geometric match) is the DEFAULT and MUST NOT be bypassed without an explicit --no-strict.
  • The fallback ladder (raw-xref pick / local-vision guidance / size-pick) is legacy compatibility only — never the default path.
  • Workflows call figure_remap.py extract ONLY; the internal gate script is private.
  • The contract exposes match_quality, not the engine method.
  • ==Future changes MUST NOT re-introduce silent fallback (default→ heuristic)== without bumping the policy version. A bug-fix patch must not quietly flip the default. Enforcement: the gate's CLI is guarded, and a regression test suite ({REPO}/figures/test_contract.py) runs after any gate edit.
  • extract auto-sweeps ±1 neighbor pages on strict hard_fail before returning fail. Still deterministic (geometric match only, no vision model); just widens the search window by 2 pages. See "Page resilience" below.
  • Multi-panel caption auto-relax: when a caption contains ≥2 distinct panel references ((A)(B)… or panel X), strict mode is overridden to relaxed up front — see "Multi-panel figures" below. This is policy-level routing (a capability mismatch, not a silent fallback): strict geometric match's "caption owns exactly one raster" assumption is invalid for composite layouts, so retrying strict can only waste attempts.

Multi-panel figures — capability mismatch

Composite figures with sub-panels (multi-location diagrams, staged classification figures, surgical step sequences) break a core assumption of strict mode and require ==different handling at three levels==: routing, semantic completeness, and uncertainty reporting.

1. Routing (automated)

Strict geometric matching assumes ==one caption owns one raster==. Multi-panel captions reference sub-regions ((A)/(B)/(C)/(D)/(E)) of a composite layout — those panels are not independently owned rasters. This is not an implementation bug; it's a representation mismatch.

Detection triggers on ≥2 distinct (letter) or panel <letter> references (a single (A) is allowed through — common in non-panel prose). When triggered, extract overrides strict=False and records policy=multipanel_caption_relaxed in the contract's reason field.

2. Vector overlay loss (caller awareness)

==Multi-panel figures extracted via raster/xref cropping often lose vector overlays== — panel letter labels, arrows, dotted guides, annotation lines. The crop is technically correct (right page region, right raster) but ==semantically incomplete==: it can be hard to map a sub-region back to its caption text.

Caller doctrine: when embedding a multi-panel figure in a note, write the surrounding description as self-contained — map each panel letter to its content in the prose, so readers can interpret the figure even when overlay labels were lost in extraction.

3. Uncertainty taxonomy (observation)

match_quality: uncertain today collapses three distinct failure modes: identity uncertainty (vision-model fallback selected the raster, not deterministically verified), semantic-completeness uncertainty (overlay labels lost), and geometry uncertainty (crop bbox loose). Caller doctrine: when match_quality:uncertain is returned, prose-annotate the embed (e.g. "vision-extracted, recommend visual confirmation") rather than relying on machine-readable fields.

Implementation: internal gate script (do not call from workflows)

figure_remap.extract() delegates to the internal gate implementation, which handles the ladder + QC in one call (first success wins): 0. geometric_match (primary, deterministic, no vision model) — direction-agnostic caption↔image matching: auto-detects whether captions sit above/below the figure, assigns each embedded raster to its nearest caption, crops the raster(s) owned by the target fig_id (union for multi-panel).

  1. Existing file — reuse a prior crop if it passes QC (if --existing)
  2. Raw page candidates — largest embedded rasters, decoration-filtered
  3. Local-vision-guided retry (×2) — a small local vision model suggests a bbox (weak; last resort before escalation)
  4. Escalate — exit code 2 → read the page render, estimate bbox, re-call the gate with --bbox

QC on each candidate: two checks, both pure computation, both block (whitespace fill ≥80%, text-bleed <50 prose chars). No model participates in QC — a model-backed check gave the same crop different verdicts on repeat runs (issue #16), so nothing a vision model says can gate, and the QC chain no longer calls one at all. The only vision call left in the gate is the bbox suggestion in rung 3, which the strict default never reaches; its output is a proposal that must still pass the computed checks, and it is pinned to greedy decoding (temperature 0, top_k 1, fixed seed) so it does not drift between runs. Whether a QC-passed crop is worth embedding is judged downstream by the workflow's frontier-model classification step.

Exit codes: 0 ok (saved to --out) · 2 escalate to frontier vision · other → log + skip.

Show full SKILL.md (1,156 more words)Show less

Deterministic mode — no-fallback strict mode

The best-effort ladder is heuristic routing: when geometric matching misses, it falls through to raw-xref pick / local-vision-guided crop, which can silently produce a plausible-but-wrong crop (e.g. grabbing the adjacent figure's raster on a multi-figure page, or any raster when the caller passes a wrong --page). The entrypoint therefore defaults to the deterministic contract; the fallback ladder is opt-in via --no-strict.

Behavior under strict mode:

  • geometric_match is the ONLY extractor on the execution path. Raw-xref pick, vision-guided bbox, existing-file reuse, and size-based picks are skipped entirely.
  • If geometric_match cannot deterministically own a raster for the normalized fig_id → HARD_FAIL, surfaced on the entrypoint contract as {"status":"fail", "match_quality":"failed", "hard_fail":true, "file":null, "reason":...} with exit 1. It never substitutes a different raster. (Exit 2 is status:escalate, which strict mode never returns.)
  • Internally the gate tracks a match_method (geometric_match | FAIL) and the per-book QC log records a method counter — but that key is deliberately not on the contract above; branch on match_quality.

When strict HARD_FAILs, the caller decides: fix --page, or escalate to vision (read the page render → --bbox), or leave a <!-- TODO -->. A HARD_FAIL is the correct outcome — it converts a silent mislabel into an explicit miss.

Page resilience: ±1 neighbor sweep

Caption/figure cross-page offsets are common: a caption may be on the page the markdown lists while the raster sits on N±1 (frequent in multi-column layouts where a figure spans a column break). A single-page strict match can't recover this on its own.

figure_remap.py extract handles this transparently:

  1. Try --page N (strict geometric_match)
  2. If hard_fail: silently retry N-1, then N+1 (still strict, still geometric_match only)
  3. First pass wins; reason annotates which neighbor matched
  4. If all three fail → return status:fail with reason listing the neighbor attempts

==This is NOT a fallback ladder.== Every attempt is still deterministic geometric matching; no vision model, no raw-xref pick. The contract is unchanged.

--existing is intentionally skipped for neighbor attempts: the fast-path image is keyed to the original page assertion, so reusing it on a different page would short-circuit the real geometric retry.

When a true hard_fail returns after the sweep, the caller's next move is:

  1. ==Re-check --page== — the caption may live on a different page than the markdown says
  2. ==Cross-book search== (workflow-level, not script-level): search the same concept across your other priority reference books and re-issue extract against the best alternative
  3. ==Vision escalate== (--no-strict) — last resort
  4. Only if all the above fail → leave <!-- TODO --> per the convention below

fig_id is dash-variant agnostic: normalization folds hyphen, en-dash, em-dash, figure-dash, horizontal-bar and period to a canonical dotted id, so 64-5 / 64–5 / Fig 64.5 / 30 — 44 all compare equal.

Per-book calibration (do this BEFORE calling extract)

Caption format varies per book — calibrate first to set the correct --fig-id:

  • Grep one chapter's markdown for a figure-caption pattern (e.g. (?:FIGURE|Fig\.?)\s*\S+) to see how captions are written
  • Inspect figures/figures_manifest.json's fig_id field if present
  • Inspect existing figures/Fig_* filenames

==The calibration output is a list to copy from, not a rule to apply.== --book, --fig-id and --caption are exact-match strings: paste --caption verbatim from the converted markdown (don't retype it, don't tidy the spacing, don't translate it) and take --book verbatim from the directory name under your corpus root. If the fig_id you need isn't in the manifest or the grep output, ==say so and stop — do not construct one from the pattern==; a plausible-but-wrong fig_id is what makes geometric matching claim the neighbouring figure.

Common variants:

  • Caption marker: FIGURE / Figure / Fig. / Fig / FIG.
  • Separator: - / . / en-dash / em-dash
  • Sub-letter: 5.42a / 5-42A / 5.42 (a) / 5.42, A
  • Page offset: PDF page ≠ printed page (varies per book — check front matter)

Layout B books (no pre-extracted figures)

Some books have only an empty figures/ — the fast path is skipped, and the gate goes straight to re-extract. For these, supply --pdf + --page directly; do not pass --existing.

QC criteria (built into the gate)

Every figure passes through QC before being saved:

  • Content type matches caption modality (diagram / imaging / photo / chart)
  • Multi-part figures (A/B/C) include all panels
  • No bleed-in from adjacent figures
  • Text labels readable (not a blurred thumbnail)

If QC fails on the existing fast-path crop → automatic re-extract. If QC fails on re-extract → local-vision-guided retry with bbox. If QC fails after retries → escalate (exit 2) for a frontier-model read.

TODO comment convention in notes

When a figure can't be extracted now (PDF missing, caption ambiguous), leave:

markdown
<!-- TODO: insert Fig X.Y from {Book} ({reason}) — render PDF page N + crop -->

==TODO is the LAST step, not the first.== Before writing this comment the caller MUST have:

  1. Confirmed the page number against the textbook markdown (grep the caption text)
  2. Let the ±1 neighbor sweep run (automatic in extract)
  3. Tried at least one cross-book alternative for the same concept — the same figure/photo/classification often appears in more than one reference book
  4. Optionally used --no-strict for a frontier-vision escalation

Do not imply a batch run will "fill this in later" — each agent extracts its own figures on demand.

==A failed extract gets a TODO comment, never an embed.== Writing ![[Fig_X-Y_Book.png]] for a figure that never passed QC produces a note that looks complete and renders a broken link — exactly what figure_embed_lint.py exists to catch. Run the lint on every note you touched before reporting done.

Legacy: batch re-extraction (rarely used now)

The original batch workflow is preserved for cases where the core extractor is materially changed and you want to validate the fix across the whole corpus:

bash
python {REPO}/figures/batch_remap.py --list
python {REPO}/figures/batch_remap.py --book NAME --dry
python {REPO}/figures/batch_remap.py --book NAME --apply

This runs caption-based extraction on every chapter, computes QC metrics (green/yellow/red), and visual-checks a few random samples per book via a local vision model. See {REPO}/figures/CALIBRATION.md for accumulated per-book quirks.

Do not run batch as part of normal note-writing — too slow, and the on-demand gate is more reliable for single figures because it can fall back to local-vision-guided bbox retry, which batch mode doesn't try.

Historically some setups also had a separate legacy caption-based batch extractor (extract_figures.py). It is not shipped in this repo — {REPO}/figures/batch_remap.py degrades gracefully without it (falls back to the gate's own extraction path per figure).

Files

Everything referenced above lives under {REPO}/figures/:

  • figure_remap.py — sanctioned entrypoint (extract, extract-page)
  • figure_qc_gate.py — primary internal implementation: single-figure on-demand extraction + QC gate
  • figure_scanned.py — scanned/hybrid caption_anchor backend
  • qc_metrics.py — batch QC metric computation
  • visual_check.py — legacy: local-vision-model sample verification, used only by the batch tools (never by the gate or the note-writing workflow)
  • visual_check_batch.py — legacy: batch-mode visual-check runner
  • batch_remap.py — legacy: whole-book batch re-extraction
  • test_contract.py — regression test suite for the entrypoint contract
  • CALIBRATION.md — per-book quirks accumulated from batch runs (still useful for setting --fig-id format expectations)

Key constraints

  • 0 hosted-LLM tokens for QC: only a local vision model via a local inference server. Frontier-model vision is used only on escalation (exit code 2).
  • No destructive edits: source PDFs are the source of truth. --out writes a fresh copy to your notes' attachments folder; the original figures/ (if present) is left untouched.
  • Idempotent: re-running extract on the same figure produces the same output.
  • TODO trail: any failed extraction leaves a <!-- TODO: ... --> comment in the note rather than silently skipping.

© drpwchen, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/figure-remap of drpwchen/textbook-to-note.

Open the folder on GitHubat commit 65e7690

Compare with similar skills

Figure Remap next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Figure Remap compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Figure Remap this skilldrpwchen/textbook-to-note105—~6.4kAutomated safety check: PassMIT
MineruNebutra/MinerU-Skill122—~504Automated safety check: PassMIT
Summarizereysu/ai-life-skills269—~7.7kAutomated safety check: PassMIT
Study Vaultten-builder/ten-builder111—~718Automated safety check: PassCustom licence
Doc Importerwpsnote/wpsnote-skills179—~2.7kAutomated safety check: PassMIT
PDFaxoviq-ai/synthadoc1.6k—~293Automated safety check: PassAGPL-3.0-or-later

Similar skills

  • Mineru

    Nebutra/MinerU-Skill

    An AI-Native skill for parsing PDF / Office / image files into Markdown with MinerU — a fast, zero-config document parser for AI agents.

    122 GitHub stars~504 tokensUpdated 15 days ago
    Documents & OfficeAuto-check passed
  • Summarize

    reysu/ai-life-skills

    Summarize any content (YouTube video, article, whitepaper/PDF, podcast episode, book chapter, etc.) into a rich Obsidian note with section-by-section breakdowns, wikilinks to all technical concepts…

    269 GitHub stars~7.7k tokensUpdated 1 mo ago
    Documents & OfficeAuto-check passed
  • Study Vault

    ten-builder/ten-builder

    PDF/문서를 Obsidian StudyVault(학습 노트)로 변환합니다. An agent skill from ten-builder/ten-builder.

    111 GitHub stars~718 tokensUpdated 3 mo ago
    Documents & OfficeAuto-check passed
  • Doc Importer

    wpsnote/wpsnote-skills

    将本地文档批量导入到 WPS 笔记。支持扫描 Obsidian Vault、思源笔记、微信公众号存档、 下载目录或任意用户指定目录,自动识别 HTML、Markdown、PDF、DOCX、PPTX、XLSX 等格式, 转换为 WPS 笔记可读内容并保留图片和富文本格式。

    179 GitHub stars~2.7k tokensUpdated 4 mo ago
    Documents & OfficeAuto-check passed
  • PDF

    axoviq-ai/synthadoc

    Extract text from PDF documents

    1.6k GitHub stars~293 tokensUpdated today
    Documents & OfficeAuto-check passed
  • Z Md To PDF

    tjxj/z-skills

    将 Markdown、Obsidian 笔记、技术文章、白皮书或长文排成中文 PDF。用户说“转成 PDF”“Markdown 转 PDF”“md 转 pdf”“排版成电子书”“生成指定风格 PDF”“一次生成多种风格”“学术风/书籍风/报告风”,或要编译 my-girlfriend-jingtian-latex 项目时,都应使用本…

    547 GitHub stars~711 tokensUpdated 17 days ago
    Documents & OfficeAuto-check passed

More from drpwchen/textbook-to-note

  • Textbook To Md

    drpwchen/textbook-to-note

    Convert PDF/EPUB textbooks to searchable markdown files for an AI agent's own reference.

    105 GitHub stars~3.5k tokensUpdated 1 mo ago
    Auto-check passed

Works with

Questions about Figure Remap

What does Figure Remap do?

Extract single figures from PDF textbooks on-demand with built-in QC verification. Figure Remap is an agent skill from drpwchen/textbook-to-note. Extract single figures from PDF textbooks on-demand with built-in QC verification.

When should I use Figure Remap?

Figure Remap fits situations like: tasks that involve PDF.

How do I install Figure Remap in Claude Code?

Run `npx skills add drpwchen/textbook-to-note --skill figure-remap -a claude-code`. Or copy the skill folder (skills/figure-remap in drpwchen/textbook-to-note) into .claude/skills/figure-remap in your project. Claude Code loads it when a task matches its description.

How do I install Figure Remap in Codex?

Run `npx skills add drpwchen/textbook-to-note --skill figure-remap -a codex`. Or copy the skill folder (skills/figure-remap in drpwchen/textbook-to-note) into .agents/skills/figure-remap in your project. Codex loads it when a task matches its description.

Can I use Figure Remap in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add drpwchen/textbook-to-note --skill figure-remap -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/figure-remap, .gemini/skills/figure-remap, .github/skills/figure-remap and .opencode/skills/figure-remap in your project.

What does Figure Remap need to run?

Going by SKILL.md and its folder, Figure Remap needs the command-line tools its instructions call (python). Our summary lists: Python 3.

Does Figure Remap access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Figure Remap safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Figure Remap use?

Figure Remap is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Figure Remap use?

About 6.4k tokens (SKILL.md is roughly 26k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Figure Remap?

Skills that share tags, products or a category with Figure Remap: Mineru (Nebutra/MinerU-Skill, 122 stars), Summarize (reysu/ai-life-skills, 269 stars), Study Vault (ten-builder/ten-builder, 111 stars) and Doc Importer (wpsnote/wpsnote-skills, 179 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Figure Remap?

drpwchen (a GitHub user) maintains it in drpwchen/textbook-to-note, which has 105 GitHub stars. The repository holds 2 skills in this directory. The repository was last updated on August 26, 2026.

Source: drpwchen/textbook-to-note on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.