Agent skill

Mapping Documents

by oaustegard in oaustegard/claude-skills

Generate navigable semantic maps from PDF documents. An agent skill from oaustegard/claude-skills.

MITAuto-check passedAgent Workflows

Install Mapping Documents

skills CLI
$ npx skills add oaustegard/claude-skills --skill mapping-documents -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install oaustegard/claude-skills mapping-documents --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/oaustegard/claude-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/mapping-documents .claude/skills/mapping-documents && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
mapping-documents
GitHub stars
150
Token cost
~1.4k tokens
SKILL.md length
402 words
Files
4 (incl. scripts)
Skills in repo
69
Repo updated
First seen
Licence
MIT

At a glance

Generate navigable semantic maps from PDF documents. An agent skill from oaustegard/claude-skills.

  • Works in 6 steps: Read _USAGE.md block in CLAUDE.md for… → Read top-level TOC in _MAP.md for… → Drill into relevant sections for typed… → …
  • Analyzing papers
  • SKILL.md covers Installation, Generate Maps, Output Artifacts and After Generating: Wire It Up, plus 4 more sections
  • Runs Python scripts from its folder; calls python, python3 and pip; needs ANTHROPIC_API_KEY and API_KEY

What it does

Mapping Documents is an agent skill from oaustegard/claude-skills. Generate navigable semantic maps from PDF documents. Extracts section structure via font analysis, then runs LLM extraction per section for claims, symbols, and dependencies — all page-anchored. Produces MAP.md (progressive disclosure), .symbols.json (definition index), .anchors.json (claim references), and a USAGE.md snippet for CLAUDE.md. Use when analyzing papers, specs, or legal docs; when asked to "map this document", "index this PDF", "what does this paper say"; or when a coding agent needs grounded…

Its SKILL.md is about 1.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files, including scripts (for example `CHANGELOG.md`, `README.md` and `scripts/docmap.py`).

It sits in Agent Workflows, covering Agent instruction files and PDF. The repository describes itself as: My collection of Claude skills. The licence is MIT.

When your agent uses it

  • Analyzing papers
  • Asked to map this document
  • What does this paper say
  • A coding agent needs grounded reference material from a PDF source

Example prompts

  • “map this document”
  • “index this PDF”
  • “what does this paper say”
  • “/mapping-documents”

Requirements

  • Python 3
  • A credential in ANTHROPIC_API_KEY
  • A credential in API_KEY

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Read _USAGE.md block in CLAUDE.md for orientation
  2. Read top-level TOC in _MAP.md for structure and section summaries
  3. Drill into relevant sections for typed claims and symbol definitions
  4. Query .symbols.json for "where is X defined?" lookups
  5. Query .anchors.json for claim filtering by type or section
  6. Read the raw PDF only when exact wording or figures are needed

What it can do on your machine

Read from SKILL.md and the folder at commit 6fc82b8. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python
    • python3
    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • ANTHROPIC_API_KEY
    • API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Mapping Documents loads about 1.4k tokens when it runs. Until then it costs about 155 tokens; SKILL.md has 402 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~155
When it runs · the whole SKILL.md, loaded when a task matches
~1.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from oaustegard/claude-skills at commit 6fc82b8, republished under its MIT licence (© oaustegard). 402 words, ~1,442 tokens.

Download SKILL.mdSave it as .claude/skills/mapping-documents/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
mapping-documents
description
Generate navigable semantic maps from PDF documents. Extracts section structure via font analysis, then runs LLM extraction per section for claims, symbols, and dependencies — all page-anchored. Produces _MAP.md (progressive disclosure), .symbols.json (definition index), .anchors.json (claim references), and a _USAGE.md snippet for CLAUDE.md. Use when analyzing papers, specs, or legal docs; when asked to "map this document", "index this PDF", "what does this paper say"; or when a coding agent needs grounded reference material from a PDF source. Analogous to tree-sitting but for prose documents.
metadata.version
0.2.0

Mapping Documents

Generate _MAP.md files providing hierarchical document structure with semantic annotations. Maps show section summaries, typed claims (result/definition/method/caveat/open-question), symbol definitions, and cross-section dependencies — all anchored to page numbers.

The structural analog to tree-sitting: tree-sitter parses code via grammar, docmap parses documents via font analysis + LLM extraction.

Installation

bash
pip install pdfplumber anthropic --break-system-packages -q

Generate Maps

bash
# Full run (structure + semantic extraction via Claude API)
python /mnt/skills/user/mapping-documents/scripts/docmap.py paper.pdf \
  --out docs/ --genre paper --workers 4

# Structure only (no API calls, no cost)
python /mnt/skills/user/mapping-documents/scripts/docmap.py paper.pdf \
  --out docs/ --structure-only

API key resolution: --api-key flag > ANTHROPIC_API_KEY env > API_KEY env.

Output Artifacts

Four files, forming a three-layer progressive-disclosure stack:

CLAUDE.md / project instructions     ← curated invariants (you write this)
    ↕ (_USAGE.md bridges the gap)
_MAP.md + JSON indexes               ← navigable document map (docmap generates)
    ↕
raw PDF                              ← the source document
FilePurposeWhen to read
{stem}_USAGE.mdSnippet for pasting into CLAUDE.md / AGENTS.md / project knowledge. Describes the reading order and JSON query patterns.Once, at setup
{stem}_MAP.mdSection map: TOC with summaries, typed claims, defined symbols, dependencies. All page-anchored.Any question about what the document says
{stem}.symbols.jsonFlat symbol index: where defined, where used, what it means."Where is X defined?"
{stem}.anchors.jsonEvery claim: section ID, type, text, page number."What caveats exist?" / "What does §3 claim?"

After Generating: Wire It Up

Generating the map is step 1. Step 2 is telling the agent the map exists.

For a code repo (CLAUDE.md / AGENTS.md):

bash
# Paste the generated usage snippet into your agent instructions
cat docs/paper_USAGE.md >> CLAUDE.md

For Claude.ai project knowledge: Upload _MAP.md as a project knowledge file, or paste the _USAGE.md content into project instructions.

The _USAGE.md snippet includes copy-pasteable query commands for the JSON indexes. Replace QUERY and SECTION_ID placeholders with actual values.

Show full SKILL.md (194 more words)Show less

Navigate Via Maps

After generating and wiring up, use the map for navigation — read _MAP.md, not the raw PDF.

Workflow:

  1. Read _USAGE.md block in CLAUDE.md for orientation
  2. Read top-level TOC in _MAP.md for structure and section summaries
  3. Drill into relevant sections for typed claims and symbol definitions
  4. Query .symbols.json for "where is X defined?" lookups
  5. Query .anchors.json for claim filtering by type or section
  6. Read the raw PDF only when exact wording or figures are needed

Querying the JSON indexes:

bash
# Symbol lookup
python3 -c "import json; [print(f'§{s[\"defined_in\"]} p.{s[\"defined_at_page\"]}') \
  for s in json.load(open('docs/paper.symbols.json')) if 'edl' in s['symbol']]"

# All caveats in the document
python3 -c "import json; [print(f'p.{c[\"page\"]} {c[\"text\"]}') \
  for c in json.load(open('docs/paper.anchors.json')) if c['type'] == 'caveat']"

# All claims in a section
python3 -c "import json; [print(f'[{c[\"type\"]}] {c[\"text\"]}') \
  for c in json.load(open('docs/paper.anchors.json')) if c['section'] == '4.3']"

Genre Support

Genre controls the claim taxonomy used in semantic extraction.

GenreClaim typesBest for
paper (default)definition, result, method, claim, caveat, open-questionAcademic papers, arXiv preprints
specrequirement, definition, constraint, example, noteRFCs, API specs, technical standards
legaldefinition, obligation, right, exception, condition, referenceContracts, policy documents, regulations

Limitations (v0.1.x)

  • PDF-only. No DOCX, HTML, or plain text input yet.
  • Single-column layout assumed. Two-column papers may mis-order text within sections.
  • No caching. Re-running re-extracts everything.
  • No citation cross-referencing.
  • Genre must be specified manually.
  • Semantic extraction can hallucinate. Every claim is page-anchored, but the page number comes from the LLM. Verify critical claims against the source.

CLI Reference

python docmap.py paper.pdf [options]

Options:
  --genre {paper,spec,legal}   Claim taxonomy (default: paper)
  --structure-only             Skip LLM pass (free, fast)
  --out DIR                    Output directory (default: .)
  --api-key KEY                Anthropic API key
  --model MODEL                Model (default: claude-sonnet-5-5)
  --workers N                  Parallel workers (default: 4)
  --no-usage-snippet           Skip _USAGE.md generation
  -v                           Verbose structural parsing

© oaustegard, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files (scripts) in mapping-documents of oaustegard/claude-skills.

  • SKILL.md
  • CHANGELOG.md
  • README.md
  • scripts/docmap.py

Open the folder on GitHubat commit 6fc82b8

Compare with similar skills

Mapping Documents next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Mapping Documents compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Mapping Documents this skilloaustegard/claude-skills150—~1.4kAutomated safety check: PassMIT
Create SkillPINA-org/PINA797—~1.8kAutomated safety check: PassMIT
Writing Docsc15t/c15t1.9k—~872Automated safety check: PassApache-2.0
Project Context Auditorjsmastery-pro/skills1.4k—~3.2kAutomated safety check: NotesMIT
Demo City ReviewmikeOnBreeze/cc-crossbeam293—~1.8kAutomated safety check: PassMIT
Cv BuilderLeoYeAI/openclaw-master-skills2.2k—~2.7kAutomated safety check: PassMIT

Similar skills

  • Create Skill

    PINA-org/PINA

    Create, modify, and improve PINA skills. An agent skill from PINA-org/PINA.

    797 GitHub stars~1.8k tokensUpdated yesterday
    Agent WorkflowsAuto-check passed
  • Writing Docs

    c15t/c15t

    Author or edit c15t documentation in docs//.mdx — the source for both the c15t.com site and the docs bundled into published packages.

    1.9k GitHub stars~872 tokensUpdated today
    Agent WorkflowsAuto-check passed
  • Project Context Auditor

    jsmastery-pro/skills

    Bootstraps a project's tool-agnostic AGENTS.md files for a greenfield project, an undocumented codebase or one area, adding only what is missing and never overwriting curated content.

    1.4k GitHub stars~3.2k tokensUpdated 1 mo ago
    Agent WorkflowsAuto-check: notes
  • Demo City Review

    mikeOnBreeze/cc-crossbeam

    Local demo: City-side ADU plan review. An agent skill from mikeOnBreeze/cc-crossbeam.

    293 GitHub stars~1.8k tokensUpdated 7 mo ago
    Agent WorkflowsAuto-check passed
  • Cv Builder

    LeoYeAI/openclaw-master-skills

    Creates professional CVs with permanent URLs at talent.de. An agent skill from LeoYeAI/openclaw-master-skills.

    2.2k GitHub stars~2.7k tokensUpdated 2 mo ago
    Agent WorkflowsAuto-check passed
  • Qiaomu Campus Resume

    aiskillstore/marketplace

    为大学生、应届生和实习求职者从 0 创建、定制或美化可投递简历:支持借鉴 Grill Me 的一问一答访谈,从零收集经历并生成 PDF;读取岗位 JD 与旧简历,建立要求—证据映射并针对性改写;在不改变事实的前提下优化已有简历内容结构或只美化排版;按用户明确要求一次生成六种不同风格的 ATS 友好 PDF 简历。适用于“采访我做简历”“没有简历从零问答”“根据 JD…

    430 GitHub stars~1.6k tokensUpdated today
    Agent WorkflowsAuto-check passed

More from oaustegard/claude-skills

All 69 skills in this repo
  • Bluesky Zeitgeist Sampler

    oaustegard/claude-skills

    Deprecated sampler that captures short windows of the Bluesky firehose, clusters trending terms and builds an HTML report; replaced by the browsing-bluesky skill.

    150 GitHub starsUsed in 1 repo~1.4k tokens
    Auto-check passed
  • Vega-Lite Interactive Charts

    oaustegard/claude-skills

    Builds interactive Vega-Lite charts from uploaded data: analyzes the fields, picks five to ten fitting chart types, and produces a React artifact with the data embedded inline.

    150 GitHub stars~2.1k tokensUpdated today
    Auto-check passed
  • Single-File HTML Composer

    oaustegard/claude-skills

    Builds self-contained single-file HTML pages such as reports, decks, postmortems, flowcharts and prototypes from a small spec using a bundled Python composer and templates.

    150 GitHub stars~3.2k tokensUpdated today
    Auto-check passed
  • Declauding

    oaustegard/claude-skills

    Rewrites model-sounding prose into plain technical writing and checks that every claim survives, for PR text, docs, commit messages and similar drafts.

    150 GitHub stars~5.2k tokensUpdated today
    Auto-check passed
  • Forecasting Reverso

    oaustegard/claude-skills

    Zero-shot univariate time series forecasting using the Reverso foundation model (NumPy/Numba CPU-only inference).

    150 GitHub starsUsed in 1 repo~1.5k tokens
    Auto-check passed
  • Preact Developer

    oaustegard/claude-skills

    Guides building standards-based Preact apps with native-first choices, HTM syntax, import maps and vendored ESM, from single-file demos to larger builds.

    150 GitHub stars~4.6k tokensUpdated today
    Auto-check passed

Questions about Mapping Documents

What does Mapping Documents do?

Generate navigable semantic maps from PDF documents. An agent skill from oaustegard/claude-skills. Mapping Documents is an agent skill from oaustegard/claude-skills. Generate navigable semantic maps from PDF documents.

When should I use Mapping Documents?

Mapping Documents fits situations like: analyzing papers; asked to map this document; what does this paper say; A coding agent needs grounded reference material from a PDF source.

How do I install Mapping Documents in Claude Code?

Run `npx skills add oaustegard/claude-skills --skill mapping-documents -a claude-code`. Or copy the skill folder (mapping-documents in oaustegard/claude-skills) into .claude/skills/mapping-documents in your project. Claude Code loads it when a task matches its description.

How do I install Mapping Documents in Codex?

Run `npx skills add oaustegard/claude-skills --skill mapping-documents -a codex`. Or copy the skill folder (mapping-documents in oaustegard/claude-skills) into .agents/skills/mapping-documents in your project. Codex loads it when a task matches its description.

Can I use Mapping Documents in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add oaustegard/claude-skills --skill mapping-documents -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/mapping-documents, .gemini/skills/mapping-documents, .github/skills/mapping-documents and .opencode/skills/mapping-documents in your project.

What does Mapping Documents need to run?

Going by SKILL.md and its folder, Mapping Documents needs Python for the scripts in its folder, the command-line tools its instructions call (python, python3 and pip) and credentials named ANTHROPIC_API_KEY and API_KEY. Our summary lists: Python 3; A credential in ANTHROPIC_API_KEY; A credential in API_KEY.

Does Mapping Documents access the network?

SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Mapping Documents safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Mapping Documents use?

Mapping Documents is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Mapping Documents use?

About 1.4k tokens (SKILL.md is roughly 5.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Mapping Documents?

Skills that share tags, products or a category with Mapping Documents: Create Skill (PINA-org/PINA, 797 stars), Writing Docs (c15t/c15t, 1.9k stars), Project Context Auditor (jsmastery-pro/skills, 1.4k stars) and Demo City Review (mikeOnBreeze/cc-crossbeam, 293 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Mapping Documents?

oaustegard (a GitHub user) maintains it in oaustegard/claude-skills, which has 150 GitHub stars. The repository holds 69 skills in this directory. The repository was last updated on October 8, 2026.

Source: oaustegard/claude-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.