Agent skill

Clinical Document Ingestion

by maziyarpanahi in maziyarpanahi/openmed

Converts scanned faxes, images, CSV/TSV exports and C-CDA XML into clean text on-device, ready for OpenMed de-identification and named-entity recognition.

Apache-2.0Auto-check passedDocuments & Office

Install Clinical Document Ingestion

skills CLI
$ npx skills add maziyarpanahi/openmed --skill ingesting-clinical-documents -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install maziyarpanahi/openmed ingesting-clinical-documents --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/maziyarpanahi/openmed.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/ingesting-clinical-documents .claude/skills/ingesting-clinical-documents && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
ingesting-clinical-documents
GitHub stars
5.5k
Token cost
~2k tokens
SKILL.md length
612 words
Files
2 (incl. references)
Skills in repo
74
Repo updated
First seen
Licence
Apache-2.0

At a glance

Converts scanned faxes, images, CSV/TSV exports and C-CDA XML into clean text on-device, ready for OpenMed de-identification and named-entity recognition.

  • Extracting text from scanned or photographed clinical notes
  • SKILL.md covers When to use, What is supported today, Install and Quick start: OCR an image,…, plus 6 more sections
  • Calls pip
  • Handling CSV or TSV patient data exports before de-identification

What it does

The skill is the intake stage of the OpenMed pipeline, turning clinical documents that are not plain text into a normalized ExtractedDocument of clean text with character-offset spans back to their source location. Everything runs on-device through local OCR backends, so no document leaves the machine. It applies to scanned or photographed notes needing OCR, CSV or TSV patient exports needing column-aware handling, and C-CDA XML needing flattening, feeding into the deidentifying-clinical-text and extracting-clinical-entities skills afterward.

redact_document dispatches by file extension today: image formats go through the ocr() function or an OcrEngine, CSV and TSV go through column-aware tabular redaction, and detected C-CDA XML goes through a standard-library CDA adapter. PDF and DOCX have no live handler yet and raise an UnsupportedDocumentError, so PDFs must first be converted to page images or text by another tool.

Installation adds the multimodal extra for the document intake contract and image dependencies, plus the ocr-paddle extra for the PaddleOCR engine, while the Tesseract engine needs the system binary installed separately. The quick-start shows calling ocr() directly for a two-step intake-then-deidentify flow, where the engine parameter can be left as auto-select or pinned to tesseract or paddleocr, and OcrResult exposes per-word bounding boxes and confidence.

When your agent uses it

  • Extracting text from scanned or photographed clinical notes
  • Handling CSV or TSV patient data exports before de-identification
  • Flattening a C-CDA XML document into text
  • Building an on-device intake pipeline before OpenMed de-identification

Example prompts

  • “OCR this scanned fax of a clinical note and prepare it for de-identification.”
  • “Redact the patient CSV export in ./exports/patients.csv using the tabular handler.”
  • “Flatten this C-CDA XML file into text for the NER pipeline.”

Requirements

  • Python with openmed[multimodal]
  • Tesseract or PaddleOCR, installed via openmed[ocr-paddle] or the system package

What it can do on your machine

Read from SKILL.md and the folder at commit 9dca507. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • github.com
    • hhs.gov
    • hl7.org

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Clinical Document Ingestion loads about 2k tokens when it runs, and up to ~3.6k if it reads all its reference files. Until then it costs about 162 tokens; SKILL.md has 612 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~162
When it runs · the whole SKILL.md, loaded when a task matches
~2k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~3.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from maziyarpanahi/openmed at commit 9dca507, republished under its Apache-2.0 licence (© maziyarpanahi). 612 words, ~2,005 tokens.

Download SKILL.mdSave it as .claude/skills/ingesting-clinical-documents/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
ingesting-clinical-documents
description
Turn scanned faxes, images, and CSV/CDA exports into clean text ready for OpenMed de-identification and NER, fully on-device. Use when the user has clinical documents (image scans, photographed/faxed notes, tabular CSV/TSV exports, C-CDA XML) and needs OCR or structured intake before openmed.deidentify and openmed.analyze_text, asks about openmed.multimodal, OCR engines (Tesseract / PaddleOCR), tabular redaction, or layout and reading order. Covers the verified ocr() and redact_document() entry points and the ExtractedDocument contract. Pairs before deidentifying-clinical-text and extracting-clinical-entities.
license
Apache-2.0
metadata.project
OpenMed
metadata.category
imaging-ocr
metadata.pairs
before
metadata.version
1.0

Ingesting Clinical Documents

Clinical text often arrives as scanned faxes, photographed notes, CSV exports, or C-CDA XML — not plain text. openmed.multimodal converts these into a normalized ExtractedDocument (clean text + character-offset → source-location spans) so you can run de-identification and NER. It runs on-device: OCR backends are local, no document leaves the machine.

When to use

  • You have images / scanned faxes of clinical notes and need text out (OCR).
  • You have CSV/TSV patient exports that need column-aware handling.
  • You have C-CDA XML to flatten into text.
  • You are building the intake stage that feeds openmed.deidentify and openmed.analyze_text.

This is the first stage. After intake, hand off to deidentifying-clinical-text then extracting-clinical-entities.

What is supported today

redact_document dispatches by file extension. Live handlers:

InputExtensionsPath
Images / scans.png .jpg .jpeg .tif .tiff .bmp .gif .webpOCR (ocr() / image handler)
Tables.csv .tsvcolumn-aware tabular redaction
C-CDA.xml (detected as CDA)stdlib CDA adapter

PDF and DOCX have no live handler yet — redact_document("x.pdf") raises UnsupportedDocumentError. Convert PDFs to page images first (or to text with your own tool) and feed the images through OCR. See references/multimodal-ingest.md for the full contract, engines, and the tabular pipeline.

Install

bash
pip install "openmed[multimodal]"      # document intake contract + image deps
pip install "openmed[ocr-paddle]"      # add the PaddleOCR engine
# Tesseract engine also needs the system binary, e.g.:  brew install tesseract

Quick start: OCR an image, then de-identify

The clean two-step intake path. ocr() lives in the submodule (it is intentionally not re-exported from openmed.multimodal):

python
from openmed.multimodal.ocr import ocr
import openmed

# 1) OCR a scanned/faxed note -> OcrResult -> ExtractedDocument -> plain text
result = ocr("fax_page.png", engine=None)   # None = auto-select an installed engine
doc    = result.to_document()                # ExtractedDocument
text   = doc.text                            # clean text for downstream OpenMed

# 2) De-identify, then run NER (privacy-first order)
deid = openmed.deidentify(text, method="mask", policy="hipaa_safe_harbor")
ner  = openmed.analyze_text(deid.deidentified_text, output_format="dict")

for ent in ner.entities:
    print(ent.label, ent.text, ent.confidence)

engine may be None (auto-select), "tesseract", "paddleocr", or an OcrEngine instance. OcrResult exposes .text and per-word boxes via .words (each OcrWord has text, bbox, confidence, page).

One-step intake + redaction with redact_document

For images, CSV/TSV, and CDA, redact_document performs intake and de-identification in a single, format-aware call, returning an already-redacted ExtractedDocument:

python
from openmed.multimodal import redact_document

# Image scan: OCR + redact in one call
doc = redact_document("fax_page.png")
print(doc.text)        # redacted text
print(doc.spans[:3])   # SourceSpan offsets -> page / bbox in the original scan

# CSV export: per-column classification (direct id / quasi-id / safe) + redaction
table_doc = redact_document("patients.csv")
print(table_doc.text)

Use redact_document when you want OpenMed to own intake and redaction (especially for tables, where redaction is column-scoped, not free-text NER). Use the ocr() → to_document() → deidentify path when you want to control the de-identification method, policy, or mapping yourself.

Tabular CSV/TSV redaction

CSV columns get classified before any cell is touched, so a free-text NER pass is not run blindly over structured data:

python
from openmed.multimodal import read_table, redact_table

view = read_table("patients.csv")            # TableView with column decisions
for col in view.columns:
    print(col.name, "->", col.assigned_class, col.action, col.canonical_label)

redacted = redact_table("patients.csv", keep_year=True)
print(redacted.text)            # redacted CSV
for entry in redacted.manifest: # PHI-SAFE audit: counts/actions per column, no raw values
    print(entry)

redact_table(...) returns a RedactedTable with .text, .headers, .rows, .columns, and a PHI-safe .manifest (no raw cell values). See references/multimodal-ingest.md for column classes and actions.

Show full SKILL.md (259 more words)Show less

Preserve layout / reading order and map back to the source

Every ExtractedDocument keeps character offset → source location. After detecting PHI on doc.text, project a span's offset back to its page and bounding box:

python
from openmed.multimodal.ocr import ocr
import openmed

doc  = ocr("fax_page.png").to_document()
deid = openmed.deidentify(doc.text, method="mask")

for ent in deid.pii_entities:
    loc = doc.location_at(ent.start)   # SourceSpan or None
    if loc is not None:
        print(ent.label, "page", loc.page, "bbox", loc.bbox)

This lets you redact pixels on the original scan, not just the extracted text.

Hand-off to / from OpenMed

  • To deidentifying-clinical-text: pass doc.text to openmed.deidentify(...) with a policy profile; this is the required next stage for PHI.
  • To extracting-clinical-entities: run openmed.analyze_text on the redacted text, not raw OCR output.
  • From file conversion (out-of-process): for PDFs/DOCX, render to page images with your own tool, then OCR those images through this skill.

Edge cases & gotchas

  • ocr() is imported from the submodule: from openmed.multimodal.ocr import ocr. It is deliberately not re-exported from openmed.multimodal.
  • No PDF/DOCX handler yet: redact_document raises UnsupportedDocumentError for them. Rasterize to images first.
  • OCR needs a backend: install [ocr-paddle] for PaddleOCR, or the system Tesseract binary for pytesseract. Missing backends raise MissingDependencyError with an install hint.
  • OCR is noisy: misreads lower downstream recall. Prefer higher-DPI scans; inspect OcrWord.confidence to flag low-quality pages.
  • Tables are not free text: redact_table redacts per column classification — don't run whole-table NER and expect structured columns to be handled correctly.
  • No raw PHI in artifacts: the table manifest and any logs record counts/actions/labels, never raw values. Keep OCR intermediates on-device and out of logs.
  • Local-first: OCR engines run locally; do not send scans to a cloud OCR API in a PHI workflow.

Standards & references

© maziyarpanahi, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (references) in skills/ingesting-clinical-documents of maziyarpanahi/openmed.

  • SKILL.md
  • references/multimodal-ingest.md

Open the folder on GitHubat commit 9dca507

Compare with similar skills

Clinical Document Ingestion next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Clinical Document Ingestion compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Clinical Document Ingestion this skillmaziyarpanahi/openmed5.5k—~2kAutomated safety check: PassApache-2.0
Docling Document Conversiondocling-project/docling69k—~1.1kAutomated safety check: PassMIT
PDF Processinganthropics/skills180k47 repos~2kAutomated safety check: PassProprietary
PDF Processing GuideshareAI-lab/learn-claude-code78k4 repos~646Automated safety check: PassMIT
Excel Spreadsheet Creation and Editinganthropics/skills180k4 repos~2.1kAutomated safety check: PassProprietary
XLSXrvdbreemen/OTGW-firmware20735 repos~2.9kAutomated safety check: PassProprietary

Similar skills

  • Docling Document Conversion

    docling-project/docling

    Converts PDFs, Office files, HTML, images and other documents into a unified DoclingDocument with Markdown or JSON output, through the docling CLI, Python SDK or a remote service.

    69k GitHub stars~1.1k tokensUpdated today
    Documents & OfficeAuto-check passed
  • PDF Processing

    anthropics/skills

    Official

    Handles everyday PDF jobs in Python and on the command line: extract text and tables, merge, split, rotate, watermark, fill forms, encrypt and OCR.

    180k GitHub starsUsed in 47 repos~2k tokens
    Documents & OfficeAuto-check passed
  • PDF Processing Guide

    shareAI-lab/learn-claude-code

    Gives the agent command-line and Python recipes for reading, creating, merging and splitting PDF files, plus tips for large and scanned documents.

    78k GitHub starsUsed in 4 repos~646 tokens
    Documents & OfficeAuto-check passed
  • Official

    Creates, edits and analyzes spreadsheets (.xlsx, .xlsm, .csv, .tsv) with openpyxl and pandas, writing live formulas and recalculating to confirm zero formula errors.

    180k GitHub starsUsed in 4 repos~2.1k tokens
    Documents & OfficeAuto-check passed
  • XLSX

    rvdbreemen/OTGW-firmware

    Use this skill any time a spreadsheet file is the primary input or output.

    207 GitHub starsUsed in 35 repos~2.9k tokens
    Documents & OfficeAuto-check passed
  • MinerU Document Reader

    opendatalab/MinerU

    Reads, OCRs, searches and cites local documents through the mineru CLI, covering PDF, images, Office files, EPUB, HTML and CSV.

    81k GitHub stars~9.4k tokensUpdated yesterday
    Documents & OfficeAuto-check: warnings

More from maziyarpanahi/openmed

All 74 skills in this repo
  • Checks OpenMed de-identified clinical text against the 18 HIPAA Safe Harbor identifier categories and reports gaps and residual re-identification risk.

    5.5k GitHub stars~1.7k tokensUpdated yesterday
    Auto-check passed
  • OpenMed Model Card Writer

    maziyarpanahi/openmed

    Fills in a model card for an OpenMed clinical NER or de-identification model from its evaluation reports: intended use, metrics, subgroups and limitations.

    5.5k GitHub stars~1.8k tokensUpdated yesterday
    Auto-check passed
  • Walks a data pipeline against the HIPAA Privacy and Security Rule checklist and produces a gap report before it processes patient data.

    5.5k GitHub stars~2k tokensUpdated yesterday
    Auto-check passed
  • ICD-10 Coding Assistant

    maziyarpanahi/openmed

    Suggests candidate ICD-10-CM diagnosis and ICD-10-PCS procedure codes for clinical text extracted by OpenMed, with rationale for a certified coder to review.

    5.5k GitHub stars~2k tokensUpdated yesterday
    Auto-check passed
  • OpenMed ETL to OMOP CDM

    maziyarpanahi/openmed

    Maps OpenMed-extracted, terminology-coded conditions, drugs and measurements into OMOP CDM v5.4 tables for OHDSI and ATLAS analytics.

    5.5k GitHub stars~1.9k tokensUpdated yesterday
    Auto-check passed
  • Extracting SDOH and Z-Codes

    maziyarpanahi/openmed

    Finds social risks such as housing instability or food insecurity in clinical notes and proposes matching ICD-10-CM Z-codes for a coder to confirm.

    5.5k GitHub stars~1.9k tokensUpdated yesterday
    Auto-check passed

Works with

Questions about Clinical Document Ingestion

What does Clinical Document Ingestion do?

Converts scanned faxes, images, CSV/TSV exports and C-CDA XML into clean text on-device, ready for OpenMed de-identification and named-entity recognition. The skill is the intake stage of the OpenMed pipeline, turning clinical documents that are not plain text into a normalized ExtractedDocument of clean text with character-offset spans back to their source location. Everything runs on-device through local OCR backends, so no document leaves the machine.

When should I use Clinical Document Ingestion?

Clinical Document Ingestion fits situations like: extracting text from scanned or photographed clinical notes; handling CSV or TSV patient data exports before de-identification; flattening a C-CDA XML document into text; building an on-device intake pipeline before OpenMed de-identification.

How do I install Clinical Document Ingestion in Claude Code?

Run `npx skills add maziyarpanahi/openmed --skill ingesting-clinical-documents -a claude-code`. Or copy the skill folder (skills/ingesting-clinical-documents in maziyarpanahi/openmed) into .claude/skills/ingesting-clinical-documents in your project. Claude Code loads it when a task matches its description.

How do I install Clinical Document Ingestion in Codex?

Run `npx skills add maziyarpanahi/openmed --skill ingesting-clinical-documents -a codex`. Or copy the skill folder (skills/ingesting-clinical-documents in maziyarpanahi/openmed) into .agents/skills/ingesting-clinical-documents in your project. Codex loads it when a task matches its description.

Can I use Clinical Document Ingestion in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add maziyarpanahi/openmed --skill ingesting-clinical-documents -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ingesting-clinical-documents, .gemini/skills/ingesting-clinical-documents, .github/skills/ingesting-clinical-documents and .opencode/skills/ingesting-clinical-documents in your project.

What does Clinical Document Ingestion need to run?

Going by SKILL.md and its folder, Clinical Document Ingestion needs the command-line tools its instructions call (pip). Our summary lists: Python with openmed[multimodal]; Tesseract or PaddleOCR, installed via openmed[ocr-paddle] or the system package.

Does Clinical Document Ingestion access the network?

SKILL.md names 3 domains. As links in the text: github.com, hhs.gov and hl7.org. This is read from the text; nothing was executed.

Is Clinical Document Ingestion safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Clinical Document Ingestion use?

Clinical Document Ingestion is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Clinical Document Ingestion use?

About 2k tokens (SKILL.md is roughly 8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.6k tokens, read only when the agent opens those files.

What are the alternatives to Clinical Document Ingestion?

Skills that share tags, products or a category with Clinical Document Ingestion: Docling Document Conversion (docling-project/docling, 69k stars), PDF Processing (anthropics/skills, 180k stars), PDF Processing Guide (shareAI-lab/learn-claude-code, 78k stars) and Excel Spreadsheet Creation and Editing (anthropics/skills, 180k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Clinical Document Ingestion?

maziyarpanahi (a GitHub user) maintains it in maziyarpanahi/openmed, which has 5,500 GitHub stars. The repository holds 74 skills in this directory. The repository was last updated on October 9, 2026.

Source: maziyarpanahi/openmed on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.