Agent skill

Extracting Pii Entities

by maziyarpanahi in maziyarpanahi/openmed

Detect PHI/PII spans in clinical text with OpenMed's extractpii without altering the text.

Apache-2.0Auto-check passed

Install Extracting Pii Entities

skills CLI
$ npx skills add maziyarpanahi/openmed --skill extracting-pii-entities -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install maziyarpanahi/openmed extracting-pii-entities --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/maziyarpanahi/openmed.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/extracting-pii-entities .claude/skills/extracting-pii-entities && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
extracting-pii-entities
GitHub stars
5.5k
Token cost
~1.7k tokens
SKILL.md length
502 words
Files
1
Skills in repo
74
Repo updated
First seen
Licence
Apache-2.0

At a glance

Detect PHI/PII spans in clinical text with OpenMed's extractpii without altering the text.

  • The user wants to find names
  • SKILL.md covers When to use, extract_pii vs deidentify, Install and Quick start, plus 6 more sections
  • Calls pip
  • Other identifiers and get their offsets and labels (not redact them)

What it does

Extracting Pii Entities is an agent skill from maziyarpanahi/openmed. Detect PHI/PII spans in clinical text with OpenMed's extractpii without altering the text. Use when the user wants to find names, dates, MRNs, phone numbers, addresses, SSNs, or other identifiers and get their offsets and labels (not redact them), inspect what would be removed before de-identifying, route spans to a custom redactor, normalize labels to a canonical taxonomy, or filter by confidence and language. Covers extractpii, the PIIEntity fields, CANONICALLABELS / normalizelabel, and how it differs from…

Its SKILL.md is about 1.7k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

The repository describes itself as: Local-first healthcare AI: clinical NER and HIPAA PII de-identification on hardware you control. 2,200+ medical models, 35 model-backed PII languages, and Python, MLX, Android… The licence is Apache-2.0.

When your agent uses it

  • The user wants to find names
  • Other identifiers and get their offsets and labels (not redact them)
  • Inspect what would be removed before de-identifying
  • Route spans to a custom redactor

Example prompts

  • “/extracting-pii-entities”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit 34d7b8c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • hhs.gov
    • huggingface.co

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Extracting Pii Entities loads about 1.7k tokens when it runs. Until then it costs about 155 tokens; SKILL.md has 502 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~155
When it runs · the whole SKILL.md, loaded when a task matches
~1.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from maziyarpanahi/openmed at commit 34d7b8c, republished under its Apache-2.0 licence (© maziyarpanahi). 502 words, ~1,745 tokens.

Download SKILL.mdSave it as .claude/skills/extracting-pii-entities/SKILL.md (or your agent's skills folder).
name
extracting-pii-entities
description
Detect PHI/PII spans in clinical text with OpenMed's extract_pii without altering the text. Use when the user wants to find names, dates, MRNs, phone numbers, addresses, SSNs, or other identifiers and get their offsets and labels (not redact them), inspect what would be removed before de-identifying, route spans to a custom redactor, normalize labels to a canonical taxonomy, or filter by confidence and language. Covers extract_pii, the PIIEntity fields, CANONICAL_LABELS / normalize_label, and how it differs from deidentify. Pairs before reidentifying-text and deidentifying-clinical-text.
license
Apache-2.0
metadata.project
OpenMed
metadata.category
openmed-core
metadata.pairs
adjacent
metadata.version
1.0

Extracting PII Entities

openmed.extract_pii finds PHI/PII spans and returns them without changing the text. Use it when you need to see the identifiers — to audit, route to a custom redactor, or decide a policy — rather than produce redacted output. It runs on-device.

When to use

  • You want the spans and labels of identifiers, with the original text intact.
  • You need a preview of what deidentify would act on before committing.
  • You are feeding detected spans into a downstream redactor (your own, Presidio, or deidentify).
  • You want to normalize model labels to a stable canonical taxonomy.

If you instead want redacted/masked output directly, use deidentifying-clinical-text (openmed.deidentify). If you need reversible masking, see reidentifying-text.

extract_pii vs deidentify

extract_piideidentify
Changes the text?NoYes (mask/remove/replace/hash/shift)
ReturnsPredictionResult (spans)DeidentificationResult (redacted text)
Default threshold0.50.7 (safety-biased)
Use fordetection, audit, routingproducing safe output

Install

bash
pip install "openmed[hf]"

Quick start

python
import openmed

note = "Patient John Doe (MRN 00481726), DOB 1970-01-15, phone 617-555-0142."

result = openmed.extract_pii(note, confidence_threshold=0.5)

for ent in result.entities:
    print(f"{ent.label:10} {ent.text!r:18} {ent.confidence:.2f} [{ent.start}:{ent.end}]")

extract_pii(...) returns a PredictionResult. Its .entities are PIIEntity objects (synthetic example fields shown):

text
ent.text            # the identifier surface string, e.g. "617-555-0142"
ent.label           # detected label, e.g. "PHONE"
ent.confidence      # model score in [0, 1]   (NOTE: .confidence, not .score)
ent.start / ent.end # character offsets into the original note
ent.canonical_label # label mapped to OpenMed's canonical taxonomy (if set)
ent.entity_type     # same as label

The text is unchanged — result.text is your original input.

Signature & key parameters

python
openmed.extract_pii(
    text,
    model_name="OpenMed/OpenMed-PII-SuperClinical-Small-44M-v1",  # default EN model
    confidence_threshold=0.5,     # raise for precision, lower for recall
    use_smart_merging=True,       # merge fragmented spans into whole units
    lang="en",                    # en es pt fr de it nl hi te ar tr ja
    loader=None,                  # reuse a ModelLoader across calls
)
  • use_smart_merging=True (default) reassembles fragmented predictions into complete units (a full phone number, a full date) — keep it on.
  • lang selects the language-appropriate default model and regex patterns. Pass the right language; do not run the English model on non-English text. Use openmed.get_default_pii_model(lang) to confirm coverage.

Normalize labels to the canonical taxonomy

Different models may emit slightly different label spellings. Normalize them to OpenMed's canonical set so downstream logic is stable:

python
import openmed
from openmed import CANONICAL_LABELS, normalize_label

result = openmed.extract_pii("Email jane.roe@example.org; SSN 123-45-6789.")

for ent in result.entities:
    canon = ent.canonical_label or normalize_label(ent.label)
    assert canon in CANONICAL_LABELS or canon == "OTHER"
    print(ent.text, "->", canon)

CANONICAL_LABELS is a frozenset of UPPER_SNAKE_CASE labels (e.g. PERSON, DATE, PHONE, EMAIL, SSN, ID_NUM, LOCATION). normalize_label(label) accepts messy inputs ("FIRSTNAME", "first_name", "B-EMAIL") and maps unknown labels to OTHER rather than raising.

Feed spans to a downstream redactor

extract_pii gives you offsets; you decide the action. A simple offset-based redactor (replace highest-offset first so positions stay valid):

python
import openmed

note = "Patient John Doe, MRN 00481726, seen 2024-03-02."
result = openmed.extract_pii(note, confidence_threshold=0.6)

redacted = note
for ent in sorted(result.entities, key=lambda e: e.start, reverse=True):
    redacted = redacted[:ent.start] + f"[{ent.label}]" + redacted[ent.end:]

print(redacted)   # Patient [PERSON], MRN [ID_NUM], seen [DATE].

For production redaction, masking strategies, and policy profiles, hand the work to openmed.deidentify instead of hand-rolling — it adds a safety sweep and date-shifting (see deidentifying-clinical-text).

Show full SKILL.md (181 more words)Show less

Hand-off to / from OpenMed

  • To deidentifying-clinical-text: once you have reviewed the spans, call openmed.deidentify(note, method="mask", policy="hipaa_safe_harbor") to produce safe output — it re-detects with a higher default threshold for safety.
  • To reidentifying-text: if you need reversibility, use openmed.deidentify(..., keep_mapping=True) and store the mapping securely.
  • To Presidio / custom anonymizers: extract_pii spans (label, start, end) translate cleanly into other recognizers' result formats; OpenMed also ships an Anonymizer (openmed.Anonymizer) for richer surrogate generation.

Edge cases & gotchas

  • Attribute is .confidence, not .score. PIIEntity extends EntityPrediction.
  • Detection is not redaction. extract_pii never changes text — if a caller expected redacted output, they want deidentify.
  • Threshold trade-off: 0.5 favors recall (good for finding PHI to review). For removing PHI, prefer deidentify's safety-biased 0.7 default.
  • Language matters: wrong lang silently lowers recall. Verify with get_default_pii_model(lang).
  • No raw PHI in logs/audit. Record offsets, labels, and hashes — never the identifier text. Use synthetic data in examples and tests.
  • Local-first. No cloud calls in PHI workflows; models run on-device after a one-time download.

Standards & references

© maziyarpanahi, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/extracting-pii-entities of maziyarpanahi/openmed.

Open the folder on GitHubat commit 34d7b8c

Compare with similar skills

Extracting Pii Entities next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Extracting Pii Entities compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Extracting Pii Entities this skillmaziyarpanahi/openmed5.5k—~1.7kAutomated safety check: PassApache-2.0
Pii Detectruvnet/ruflo74k—~350Automated safety check: NotesMIT
Detecting Model Extraction Attacksmukul975/Anthropic-Cybersecurity-Skills34k—~2.9kAutomated safety check: PassApache-2.0
Entity ExtractionNxcoreAI/EverRoom3k—~158Automated safety check: PassCustom licence
Outcome Extraction For Clinical Trialsaipoch/medical-research-skills1.9k—~1.4kAutomated safety check: PassMIT
Baseline Extraction For Clinical Trialsaipoch/medical-research-skills1.9k—~1.2kAutomated safety check: PassMIT

Similar skills

  • Pii Detect

    ruvnet/ruflo

    Detect and flag personally identifiable information (PII) in text, code, and configurations.

    74k GitHub stars~350 tokensUpdated yesterday
    Auto-check: notes
  • Detecting Model Extraction Attacks

    mukul975/Anthropic-Cybersecurity-Skills

    Detect MITRE ATLAS AML.T0024 attacks (model stealing, inversion, membership inference) performed via inference-API abuse, by monitoring per-principal query volume/distribution, rate-limiting and…

    34k GitHub stars~2.9k tokensUpdated 1 mo ago
    SecurityAuto-check passed
  • Entity Extraction

    NxcoreAI/EverRoom

    Extract searchable entities and a concise summary from a document for EverRoom knowledge routing.

    3k GitHub stars~158 tokensUpdated 2 days ago
    Auto-check passed
  • Outcome Extraction For Clinical Trials

    aipoch/medical-research-skills

    Clinical research outcome extraction for meta-analysis. An agent skill from aipoch/medical-research-skills.

    1.9k GitHub stars~1.4k tokensUpdated 24 days ago
    Research & ScienceAuto-check passed
  • Baseline Extraction For Clinical Trials

    aipoch/medical-research-skills

    Extracts clinical trial baseline data (study, region, participants, etc.) from article text or PMID.

    1.9k GitHub stars~1.2k tokensUpdated 24 days ago
    Research & ScienceAuto-check passed
  • Clinical Reports

    davila7/claude-code-templates

    Write comprehensive clinical reports including case reports (CARE guidelines), diagnostic reports (radiology/pathology/lab), clinical trial reports (ICH-E3, SAE, CSR), and patient documentation…

    33k GitHub starsUsed in 11 repos~9.9k tokens
    Legal & ComplianceAuto-check: notes

More from maziyarpanahi/openmed

All 74 skills in this repo
  • Checks OpenMed de-identified clinical text against the 18 HIPAA Safe Harbor identifier categories and reports gaps and residual re-identification risk.

    5.5k GitHub stars~1.7k tokensUpdated today
    Auto-check passed
  • OpenMed Model Card Writer

    maziyarpanahi/openmed

    Fills in a model card for an OpenMed clinical NER or de-identification model from its evaluation reports: intended use, metrics, subgroups and limitations.

    5.5k GitHub stars~1.8k tokensUpdated today
    Auto-check passed
  • Walks a data pipeline against the HIPAA Privacy and Security Rule checklist and produces a gap report before it processes patient data.

    5.5k GitHub stars~2k tokensUpdated today
    Auto-check passed
  • ICD-10 Coding Assistant

    maziyarpanahi/openmed

    Suggests candidate ICD-10-CM diagnosis and ICD-10-PCS procedure codes for clinical text extracted by OpenMed, with rationale for a certified coder to review.

    5.5k GitHub stars~2k tokensUpdated today
    Auto-check passed
  • OpenMed ETL to OMOP CDM

    maziyarpanahi/openmed

    Maps OpenMed-extracted, terminology-coded conditions, drugs and measurements into OMOP CDM v5.4 tables for OHDSI and ATLAS analytics.

    5.5k GitHub stars~1.9k tokensUpdated today
    Auto-check passed
  • Extracting SDOH and Z-Codes

    maziyarpanahi/openmed

    Finds social risks such as housing instability or food insecurity in clinical notes and proposes matching ICD-10-CM Z-codes for a coder to confirm.

    5.5k GitHub stars~1.9k tokensUpdated today
    Auto-check passed

Questions about Extracting Pii Entities

What does Extracting Pii Entities do?

Detect PHI/PII spans in clinical text with OpenMed's extractpii without altering the text. Extracting Pii Entities is an agent skill from maziyarpanahi/openmed. Detect PHI/PII spans in clinical text with OpenMed's extractpii without altering the text.

When should I use Extracting Pii Entities?

Extracting Pii Entities fits situations like: the user wants to find names; other identifiers and get their offsets and labels (not redact them); inspect what would be removed before de-identifying; route spans to a custom redactor.

How do I install Extracting Pii Entities in Claude Code?

Run `npx skills add maziyarpanahi/openmed --skill extracting-pii-entities -a claude-code`. Or copy the skill folder (skills/extracting-pii-entities in maziyarpanahi/openmed) into .claude/skills/extracting-pii-entities in your project. Claude Code loads it when a task matches its description.

How do I install Extracting Pii Entities in Codex?

Run `npx skills add maziyarpanahi/openmed --skill extracting-pii-entities -a codex`. Or copy the skill folder (skills/extracting-pii-entities in maziyarpanahi/openmed) into .agents/skills/extracting-pii-entities in your project. Codex loads it when a task matches its description.

Can I use Extracting Pii Entities in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add maziyarpanahi/openmed --skill extracting-pii-entities -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/extracting-pii-entities, .gemini/skills/extracting-pii-entities, .github/skills/extracting-pii-entities and .opencode/skills/extracting-pii-entities in your project.

What does Extracting Pii Entities need to run?

Going by SKILL.md and its folder, Extracting Pii Entities needs the command-line tools its instructions call (pip). Our summary lists: Python 3.

Does Extracting Pii Entities access the network?

SKILL.md names 2 domains. As links in the text: hhs.gov and huggingface.co. This is read from the text; nothing was executed.

Is Extracting Pii Entities safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Extracting Pii Entities use?

Extracting Pii Entities is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Extracting Pii Entities use?

About 1.7k tokens (SKILL.md is roughly 7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Extracting Pii Entities?

Skills that share tags, products or a category with Extracting Pii Entities: Pii Detect (ruvnet/ruflo, 74k stars), Detecting Model Extraction Attacks (mukul975/Anthropic-Cybersecurity-Skills, 34k stars), Entity Extraction (NxcoreAI/EverRoom, 3k stars) and Outcome Extraction For Clinical Trials (aipoch/medical-research-skills, 1.9k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Extracting Pii Entities?

maziyarpanahi (a GitHub user) maintains it in maziyarpanahi/openmed, which has 5,506 GitHub stars. The repository holds 74 skills in this directory. The repository was last updated on October 11, 2026.

Source: maziyarpanahi/openmed on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.