Agent skill

Batch Processing Clinical Text

by maziyarpanahi in maziyarpanahi/openmed

Run large-scale batch NER, PII extraction, or de-identification over many clinical notes on-device with OpenMed, with sharding, checkpointing, resumability, and append-only JSONL output.

Apache-2.0Auto-check passedDatabases

Install Batch Processing Clinical Text

skills CLI
$ npx skills add maziyarpanahi/openmed --skill batch-processing-clinical-text -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install maziyarpanahi/openmed batch-processing-clinical-text --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/maziyarpanahi/openmed.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/batch-processing-clinical-text .claude/skills/batch-processing-clinical-text && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
batch-processing-clinical-text
GitHub stars
5.5k
Token cost
~2.2k tokens
SKILL.md length
599 words
Files
1
Skills in repo
74
Repo updated
First seen
Licence
Apache-2.0

At a glance

Run large-scale batch NER, PII extraction, or de-identification over many clinical notes on-device with OpenMed, with sharding, checkpointing, resumability, and append-only JSONL output.

  • Works in 7 steps: De-identify upstream if the corpus has… → Pick the operation + model — one… → Shard the corpus into independent… → …
  • The user needs to process a corpus
  • SKILL.md covers When to use this skill, Quick start, Choosing the operation and Streaming + PHI-safe progress, plus 6 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Batch Processing Clinical Text is an agent skill from maziyarpanahi/openmed. Run large-scale batch NER, PII extraction, or de-identification over many clinical notes on-device with OpenMed, with sharding, checkpointing, resumability, and append-only JSONL output. Use when the user needs to process a corpus or folder of notes, de-identify a dataset, run NER over thousands of documents, build a resumable batch pipeline, or stream results to JSONL without holding everything in memory. Covers processbatch / BatchProcessor / BatchItem / BatchResult, the operation= selector (analyzetext |…

Its SKILL.md is about 2.2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Databases, covering Data pipelines and ETL and Database administration. The repository describes itself as: Local-first healthcare AI: clinical NER & HIPAA PII de-identification that runs 100% on-device. 2,200+ medical models, 21 languages, Apple MLX + Python, no cloud, no patient data…. The licence is Apache-2.0.

When your agent uses it

  • The user needs to process a corpus
  • Folder of notes
  • De-identify a dataset
  • Run NER over thousands of documents

Example prompts

  • “/batch-processing-clinical-text”

Requirements

  • Python 3

Workflow steps

7 steps, taken from the first numbered list in SKILL.md.

  1. De-identify upstream if the corpus has PHI, or key everything by stable
  2. Pick the operation + model — one BatchProcessor per
  3. Shard the corpus into independent partitions writing to separate JSONL
  4. Stream with iter_process and append each BatchItemResult to JSONL
  5. Track progress with the PHI-safe on_progress callback (counts/timing
  6. Re-run to resume: skip ids already present in the output; with
  7. Reconcile via result.get_failed_results() / a failures pass, then hand

What it can do on your machine

Read from SKILL.md and the folder at commit 252806a. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • jsonlines.org
    • hhs.gov

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Batch Processing Clinical Text loads about 2.2k tokens when it runs. Until then it costs about 183 tokens; SKILL.md has 599 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~183
When it runs · the whole SKILL.md, loaded when a task matches
~2.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from maziyarpanahi/openmed at commit 252806a, republished under its Apache-2.0 licence (© maziyarpanahi). 599 words, ~2,224 tokens.

Download SKILL.mdSave it as .claude/skills/batch-processing-clinical-text/SKILL.md (or your agent's skills folder).
name
batch-processing-clinical-text
description
Run large-scale batch NER, PII extraction, or de-identification over many clinical notes on-device with OpenMed, with sharding, checkpointing, resumability, and append-only JSONL output. Use when the user needs to process a corpus or folder of notes, de-identify a dataset, run NER over thousands of documents, build a resumable batch pipeline, or stream results to JSONL without holding everything in memory. Covers process_batch / BatchProcessor / BatchItem / BatchResult, the operation= selector (analyze_text | extract_pii | deidentify), iter_process streaming, the PHI-safe on_progress callback, chunking long documents, and no-PHI logging. Produces a resumable batch runner over an OpenMed model.
license
Apache-2.0
metadata.project
OpenMed
metadata.category
deployment-ops
metadata.pairs
adjacent
metadata.version
1.0

Batch processing clinical text

openmed.processing runs OpenMed over many documents efficiently, with progress tracking, per-item error isolation, and streaming. It runs fully on-device: the corpus, the model, and the output never leave the host. This skill shows a resumable runner — sharded, checkpointed, append-only JSONL — that you can restart without reprocessing.

When to use this skill

For corpora, folders, or datasets — anything beyond a handful of notes. For a single note, just call openmed.analyze_text / deidentify directly (extracting-clinical-entities, deidentifying-clinical-text). For an always-on HTTP service, see serving-openmed-rest-api.

Quick start

python
from openmed import process_batch

texts = ["Patient has type 2 diabetes.", "No acute distress. BP 120/80."]
result = process_batch(texts, model_name="disease_detection_superclinical")

print(result.summary())          # PHI-safe counts + timing
print(result.successful_items, "/", result.total_items)
for item in result.get_successful_results():
    print(item.id, item.result.to_dict()["entities"])   # spans only; avoid raw text in logs

process_batch(...) is a thin wrapper over BatchProcessor. Real signatures (openmed/processing/batch.py):

  • process_batch(texts, model_name="disease_detection_superclinical", ids=None, config=None, progress_callback=None, on_progress=None, **kwargs) -> BatchResult
  • BatchProcessor(model_name=..., operation="analyze_text", batch_size=8, continue_on_error=True, **analyze_kwargs) with operation ∈ {"analyze_text", "extract_pii", "deidentify"}.
  • BatchItem(id, text, source=None, metadata=None)
  • BatchResult — .items, .total_items, .successful_items, .failed_items, .success_rate, .average_processing_time, .summary(), .to_dict(), .get_successful_results(), .get_failed_results().
  • BatchItemResult — .id, .result (a PredictionResult/DeidentificationResult), .error, .processing_time, .source, .success, .to_dict().

Choosing the operation

python
from openmed import BatchProcessor

# NER (default)
ner = BatchProcessor(model_name="disease_detection_superclinical")          # operation="analyze_text"
# Detect PHI spans
pii = BatchProcessor(operation="extract_pii", model_name="OpenMed/OpenMed-PII-SuperClinical-Small-44M-v1")
# De-identify (rewrites text)
deid = BatchProcessor(operation="deidentify", method="mask", confidence_threshold=0.7)

BatchProcessor reuses one model loader across items (and a cached privacy filter for PII ops), so a single processor over many texts is far cheaper than many one-off calls.

Streaming + PHI-safe progress

python
from openmed.processing import BatchProcessor, BatchProgress

proc = BatchProcessor(operation="deidentify", method="mask", batch_size=16)

def on_progress(p: BatchProgress) -> None:   # frozen record: completed/total/current_index/elapsed
    if p.completed % 500 == 0:
        print(f"{p.completed}/{p.total} ({p.elapsed:.1f}s)")   # NO PHI

for item in proc.iter_process(texts, ids=doc_ids, on_progress=on_progress):
    write_jsonl(item)   # one result at a time — constant memory for huge corpora

on_progress receives only counts/timing (never text), so it is safe to log. iter_process yields BatchItemResults without holding the whole corpus.

Workflow

  1. De-identify upstream if the corpus has PHI, or key everything by stable internal ids so logs/output never carry identifiers.
  2. Pick the operation + model — one BatchProcessor per operation (analyze_text / extract_pii / deidentify); reuse it across all items so the loader is shared.
  3. Shard the corpus into independent partitions writing to separate JSONL files; run shards as separate processes for parallelism.
  4. Stream with iter_process and append each BatchItemResult to JSONL (flush per item) — that append-only file is your checkpoint.
  5. Track progress with the PHI-safe on_progress callback (counts/timing only).
  6. Re-run to resume: skip ids already present in the output; with continue_on_error=True, failures are recorded, not raised.
  7. Reconcile via result.get_failed_results() / a failures pass, then hand JSONL to the downstream consumer.

A resumable batch runner

python
import json
from pathlib import Path
from openmed.processing import BatchProcessor

def run_resumable(texts, ids, out_path: Path, *, operation="deidentify", **kw):
    out_path.parent.mkdir(parents=True, exist_ok=True)
    # 1) checkpoint = ids already written (append-only JSONL is the source of truth)
    done = set()
    if out_path.exists():
        with out_path.open() as f:
            done = {json.loads(line)["id"] for line in f if line.strip()}
    todo = [(i, t) for i, t in zip(ids, texts) if i not in done]
    if not todo:
        return
    pending_ids, pending_texts = zip(*todo)

    proc = BatchProcessor(operation=operation, continue_on_error=True, **kw)
    # 2) append each result as it completes -> safe to kill/restart anytime
    with out_path.open("a") as f:
        for item in proc.iter_process(list(pending_texts), ids=list(pending_ids)):
            record = {"id": item.id, "ok": item.success,
                      "result": item.result.to_dict() if item.success else None,
                      "error": item.error}            # item.error is a message, keep PHI out
            f.write(json.dumps(record) + "\n")
            f.flush()

Re-running skips finished ids (continue_on_error=True keeps one bad doc from killing the run; failures are recorded, not raised). Shard a corpus by writing to out/shard_000.jsonl, shard_001.jsonl, … and run shards in separate processes.

Show full SKILL.md (246 more words)Show less

Chunking long documents

BatchProcessor does not split single documents. For notes longer than the model's max sequence length, pre-split into sentences/windows (openmed.processing.sentences) — or for analyze_text, rely on built-in sentence handling — then re-stitch entities by adding the chunk offset back to each entity's start/end so spans point into the original document.

Hand-off to / from OpenMed

  • Per-item engine: each operation calls the same openmed.analyze_text / extract_pii / deidentify you'd call directly — same results, batched.
  • Downstream: JSONL feeds building-patient-timelines, etl-to-omop-cdm, and exporting-to-fhir. De-id JSONL feeds evaluating-with-leakage-gates.
  • Service vs batch: for request/response use serving-openmed-rest-api; for corpora use this batch path.

Edge cases & gotchas

  • No PHI in logs/JSONL keys. Use stable ids (BatchItem.id), offsets, and labels. BatchItemResult.error is a message — keep raw text out of inputs to exceptions you log.
  • continue_on_error=True is the default and recommended for corpora; check result.get_failed_results() afterward. Set False only when one failure should abort everything.
  • batch_size is a throughput dial, not correctness — tune to memory/CPU. Larger isn't always faster on CPU.
  • Append-only is the checkpoint. Don't buffer results in memory and write at the end; you lose progress on a crash. Append + flush per item.
  • process_files/process_directory read files for you and set BatchItem.source; unreadable files become failed items (not crashes) under continue_on_error.
  • Mixed languages: pass lang= per run for PII ops; don't run an English PII model across other languages (deidentifying-multilingual-text).

Standards & references

© maziyarpanahi, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/batch-processing-clinical-text of maziyarpanahi/openmed.

Open the folder on GitHubat commit 252806a

Compare with similar skills

Batch Processing Clinical Text next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Batch Processing Clinical Text compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Batch Processing Clinical Text this skillmaziyarpanahi/openmed5.5k—~2.2kAutomated safety check: PassApache-2.0
Replication Driven Researchbrycewang-stanford/Auto-Empirical-Research-Skills4.5k—~1.7kAutomated safety check: PassCustom licence
Ddia Systemswondelai/skills2.4k—~4.2kAutomated safety check: PassMIT
DB SculptorEliasOulkadi/shokunin114—~3.1kAutomated safety check: NotesMIT
Postgresmagnus919/agent-skills111—~4kAutomated safety check: PassMIT
Clickhouse IohellangleZ/burn-in-cceverywhere-ralph11214 repos~2.5kAutomated safety check: PassNone

Similar skills

  • Replication Driven Research

    brycewang-stanford/Auto-Empirical-Research-Skills

    A skill your agent uses when starting empirical analysis, creating a data pipeline, generating results, or when data or model specifications change.

    4.5k GitHub stars~1.7k tokensUpdated 2 days ago
    Testing & QAAuto-check passed
  • Ddia Systems

    wondelai/skills

    Design data systems by understanding storage engines, replication, partitioning, transactions, and consistency models.

    2.4k GitHub stars~4.2k tokensUpdated 26 days ago
    DatabasesAuto-check passed
  • DB Sculptor

    EliasOulkadi/shokunin

    Design database schemas with Prisma/Drizzle, PostgreSQL index strategy (B-tree, GIN, GiST, BRIN, Hash), query optimization (EXPLAIN ANALYZE), migration safety (expand/contract, zero-downtime), and…

    114 GitHub stars~3.1k tokensUpdated 2 days ago
    DatabasesAuto-check: notes
  • Postgres

    magnus919/agent-skills

    Operate PostgreSQL instances safely: configuration review, index and query-plan analysis, vacuum and bloat management, WAL archiving and point-in-time recovery, replication and failover, extensions…

    111 GitHub stars~4k tokensUpdated today
    DatabasesAuto-check passed
  • Clickhouse Io

    hellangleZ/burn-in-cceverywhere-ralph

    ClickHouse database patterns, query optimization, analytics, and data engineering best practices for high-performance analytical workloads.

    112 GitHub starsUsed in 14 repos~2.5k tokens
    DatabasesAuto-check passed
  • Analyzing Data

    astronomer/agents

    Queries the data warehouse with SQL and answers business questions about data.

    450 GitHub stars~1.3k tokensUpdated yesterday
    DatabasesAuto-check passed

More from maziyarpanahi/openmed

All 74 skills in this repo
  • Checks OpenMed de-identified clinical text against the 18 HIPAA Safe Harbor identifier categories and reports gaps and residual re-identification risk.

    5.5k GitHub stars~1.7k tokensUpdated yesterday
    Auto-check passed
  • OpenMed Model Card Writer

    maziyarpanahi/openmed

    Fills in a model card for an OpenMed clinical NER or de-identification model from its evaluation reports: intended use, metrics, subgroups and limitations.

    5.5k GitHub stars~1.8k tokensUpdated yesterday
    Auto-check passed
  • Walks a data pipeline against the HIPAA Privacy and Security Rule checklist and produces a gap report before it processes patient data.

    5.5k GitHub stars~2k tokensUpdated yesterday
    Auto-check passed
  • ICD-10 Coding Assistant

    maziyarpanahi/openmed

    Suggests candidate ICD-10-CM diagnosis and ICD-10-PCS procedure codes for clinical text extracted by OpenMed, with rationale for a certified coder to review.

    5.5k GitHub stars~2k tokensUpdated yesterday
    Auto-check passed
  • OpenMed ETL to OMOP CDM

    maziyarpanahi/openmed

    Maps OpenMed-extracted, terminology-coded conditions, drugs and measurements into OMOP CDM v5.4 tables for OHDSI and ATLAS analytics.

    5.5k GitHub stars~1.9k tokensUpdated yesterday
    Auto-check passed
  • Extracting SDOH and Z-Codes

    maziyarpanahi/openmed

    Finds social risks such as housing instability or food insecurity in clinical notes and proposes matching ICD-10-CM Z-codes for a coder to confirm.

    5.5k GitHub stars~1.9k tokensUpdated yesterday
    Auto-check passed

Questions about Batch Processing Clinical Text

What does Batch Processing Clinical Text do?

Run large-scale batch NER, PII extraction, or de-identification over many clinical notes on-device with OpenMed, with sharding, checkpointing, resumability, and append-only JSONL output. Batch Processing Clinical Text is an agent skill from maziyarpanahi/openmed. Run large-scale batch NER, PII extraction, or de-identification over many clinical notes on-device with OpenMed, with sharding, checkpointing, resumability, and append-only JSONL output.

When should I use Batch Processing Clinical Text?

Batch Processing Clinical Text fits situations like: the user needs to process a corpus; folder of notes; de-identify a dataset; run NER over thousands of documents.

How do I install Batch Processing Clinical Text in Claude Code?

Run `npx skills add maziyarpanahi/openmed --skill batch-processing-clinical-text -a claude-code`. Or copy the skill folder (skills/batch-processing-clinical-text in maziyarpanahi/openmed) into .claude/skills/batch-processing-clinical-text in your project. Claude Code loads it when a task matches its description.

How do I install Batch Processing Clinical Text in Codex?

Run `npx skills add maziyarpanahi/openmed --skill batch-processing-clinical-text -a codex`. Or copy the skill folder (skills/batch-processing-clinical-text in maziyarpanahi/openmed) into .agents/skills/batch-processing-clinical-text in your project. Codex loads it when a task matches its description.

Can I use Batch Processing Clinical Text in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add maziyarpanahi/openmed --skill batch-processing-clinical-text -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/batch-processing-clinical-text, .gemini/skills/batch-processing-clinical-text, .github/skills/batch-processing-clinical-text and .opencode/skills/batch-processing-clinical-text in your project.

What does Batch Processing Clinical Text need to run?

SKILL.md names no scripts, command-line tools or credentials: Batch Processing Clinical Text is instructions for the agent only. Our summary lists: Python 3.

Does Batch Processing Clinical Text access the network?

SKILL.md names 2 domains. As links in the text: jsonlines.org and hhs.gov. This is read from the text; nothing was executed.

Is Batch Processing Clinical Text safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Batch Processing Clinical Text use?

Batch Processing Clinical Text is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Batch Processing Clinical Text use?

About 2.2k tokens (SKILL.md is roughly 8.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Batch Processing Clinical Text?

Skills that share tags, products or a category with Batch Processing Clinical Text: Replication Driven Research (brycewang-stanford/Auto-Empirical-Research-Skills, 4.5k stars), Ddia Systems (wondelai/skills, 2.4k stars), DB Sculptor (EliasOulkadi/shokunin, 114 stars) and Postgres (magnus919/agent-skills, 111 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Batch Processing Clinical Text?

maziyarpanahi (a GitHub user) maintains it in maziyarpanahi/openmed, which has 5,452 GitHub stars. The repository holds 74 skills in this directory. The repository was last updated on October 6, 2026.

Source: maziyarpanahi/openmed on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.