Agent skill

Deidentifying Multilingual Text

by maziyarpanahi in maziyarpanahi/openmed

De-identify non-English clinical text on-device with OpenMed by passing lang= and locale= to deidentify().

Apache-2.0Auto-check passedDatabases

Install Deidentifying Multilingual Text

skills CLI
$ npx skills add maziyarpanahi/openmed --skill deidentifying-multilingual-text -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install maziyarpanahi/openmed deidentifying-multilingual-text --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/maziyarpanahi/openmed.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/deidentifying-multilingual-text .claude/skills/deidentifying-multilingual-text && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
deidentifying-multilingual-text
GitHub stars
5.5k
Token cost
~1.7k tokens
SKILL.md length
551 words
Files
1
Skills in repo
74
Repo updated
First seen
Licence
Apache-2.0

At a glance

De-identify non-English clinical text on-device with OpenMed by passing lang= and locale= to deidentify().

  • Works in 5 steps: Confirm the language is supported by… → Pass lang= to deidentify/extract_pii.… → Set locale= for surrogates when… → …
  • The user has Spanish
  • SKILL.md covers When to use this skill, Discover supported languages…, Quick start (Spanish) and Workflow, plus 4 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Deidentifying Multilingual Text is an agent skill from maziyarpanahi/openmed. De-identify non-English clinical text on-device with OpenMed by passing lang= and locale= to deidentify(). Use when the user has Spanish, German, French, Italian, Portuguese, Dutch, Hindi, Telugu, Arabic, Japanese, or Turkish medical notes, needs locale-aware fake surrogates, must handle language-specific national IDs (DNI, NIR, Steuer-ID, codice fiscale, BSN, CPF, TCKN, Aadhaar), or asks which languages OpenMed PII supports. Covers SUPPORTEDLANGUAGES, getpiimodelsbylanguage, getpatternsforlanguage, LANGTOLOCALE…

Its SKILL.md is about 1.7k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Databases, covering Database schema design and Internationalization. The repository describes itself as: Local-first healthcare AI: clinical NER and HIPAA PII de-identification on hardware you control. 2,200+ medical models, 35 model-backed PII languages, and Python, MLX, Android… The licence is Apache-2.0.

When your agent uses it

  • The user has Spanish
  • Turkish medical notes
  • Needs locale-aware fake surrogates
  • Must handle language-specific national IDs (DNI

Example prompts

  • “/deidentifying-multilingual-text”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Confirm the language is supported by checking SUPPORTED_LANGUAGES.
  2. Pass lang= to deidentify/extract_pii. This selects the
  3. Set locale= for surrogates when method="replace". If omitted, the
  4. Let accent normalization happen. For models trained on accent-free text
  5. Keep surrogates stable across a document with consistent=True, seed=....

What it can do on your machine

Read from SKILL.md and the folder at commit 9dca507. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • hhs.gov

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Deidentifying Multilingual Text loads about 1.7k tokens when it runs. Until then it costs about 168 tokens; SKILL.md has 551 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~168
When it runs · the whole SKILL.md, loaded when a task matches
~1.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from maziyarpanahi/openmed at commit 9dca507, republished under its Apache-2.0 licence (© maziyarpanahi). 551 words, ~1,695 tokens.

Download SKILL.mdSave it as .claude/skills/deidentifying-multilingual-text/SKILL.md (or your agent's skills folder).
name
deidentifying-multilingual-text
description
De-identify non-English clinical text on-device with OpenMed by passing lang= and locale= to deidentify(). Use when the user has Spanish, German, French, Italian, Portuguese, Dutch, Hindi, Telugu, Arabic, Japanese, or Turkish medical notes, needs locale-aware fake surrogates, must handle language-specific national IDs (DNI, NIR, Steuer-ID, codice fiscale, BSN, CPF, TCKN, Aadhaar), or asks which languages OpenMed PII supports. Covers SUPPORTED_LANGUAGES, get_pii_models_by_language, get_patterns_for_language, LANG_TO_LOCALE, and accent normalization. Pairs with OpenMed deidentifying-clinical-text and generating-synthetic-surrogates.
license
Apache-2.0
metadata.project
OpenMed
metadata.category
de-identification
metadata.pairs
adjacent
metadata.version
1.0

De-identifying multilingual text

OpenMed de-identifies clinical text in many languages, each with a dedicated PII model, language-specific regex patterns (national IDs, phone formats), and a locale-aware surrogate generator. Pass lang= to deidentify / extract_pii and the right model, patterns, and fake-data tables are selected automatically. Everything runs on-device.

When to use this skill

Use it whenever the source text is not English, or when surrogates must look native to the locale (a German note should get German-looking fake names and a valid-format Steuer-ID surrogate, not a US SSN).

Discover supported languages at runtime — don't hardcode

python
import openmed
from openmed.core.pii_i18n import SUPPORTED_LANGUAGES, get_patterns_for_language

print(sorted(SUPPORTED_LANGUAGES))   # query it; the set is the source of truth
# Language-appropriate default model for a code:
models = openmed.get_pii_models_by_language("es")
# Language-specific regex patterns (national IDs, phones, etc.):
patterns = get_patterns_for_language("de")

The set currently spans English plus European, South Asian, Middle Eastern, and East Asian languages — but always read SUPPORTED_LANGUAGES rather than trusting a number, since it changes as models ship. MCP exposes the same list via openmed_list_pii_languages.

Quick start (Spanish)

python
import openmed

nota = (
    "El paciente Carlos Hernández (DNI 12345678Z), nacido el 11/04/1979, "
    "vive en Calle Mayor 5, Madrid. Teléfono 612 345 678."
)

result = openmed.deidentify(
    nota,
    lang="es",                # selects the Spanish PII model + ES patterns
    method="replace",         # locale-native fake values
    locale="es_ES",           # Faker locale (defaults from lang via LANG_TO_LOCALE)
)
print(result.deidentified_text)
# El paciente [surrogate name] (DNI [surrogate]), nacido el [date], ...

For German, just switch the code:

python
befund = "Patientin Anna Müller, geb. 11.04.1979, Steuer-ID 12 345 678 901."
result = openmed.deidentify(befund, lang="de", method="replace")

Workflow

  1. Confirm the language is supported by checking SUPPORTED_LANGUAGES.
  2. Pass lang= to deidentify/extract_pii. This selects the language-specific model (via get_pii_models_by_language) and the regex pattern set (via get_patterns_for_language) for national IDs and formats.
  3. Set locale= for surrogates when method="replace". If omitted, the locale is derived from lang through LANG_TO_LOCALE (e.g. pt→pt_PT). Override for regional variants (pt_BR, en_GB, Gulf/Levant Arabic).
  4. Let accent normalization happen. For models trained on accent-free text (Spanish), deidentify auto-strips diacritics before inference and maps spans back to the original accented text. You normally do not set normalize_accents yourself.
  5. Keep surrogates stable across a document with consistent=True, seed=....

Language-specific national IDs

The pattern sets encode and the validators check real national identifier formats and checksums, so structured IDs are caught even when the model is unsure. Examples available in openmed.core.pii_i18n:

LanguageIdentifierValidator
FrenchNIR / INSEEvalidate_french_nir
GermanSteuer-IDvalidate_german_steuer_id
ItalianCodice Fiscalevalidate_italian_codice_fiscale
SpanishDNI / NIEvalidate_spanish_dni, validate_spanish_nie
DutchBSNvalidate_dutch_bsn
HindiAadhaarvalidate_aadhaar
PortugueseCPF / CNPJvalidate_portuguese_cpf, validate_portuguese_cnpj
TurkishTCKNvalidate_turkish_tckn

These map to OpenMed CANONICAL_LABELS (ID_NUM, SSN) and are redacted by the same policy actions as any other identifier.

Show full SKILL.md (223 more words)Show less

Hand-off to / from OpenMed

  • Core de-id: deidentifying-clinical-text — methods, thresholds, keep_mapping, policies (all accept lang/locale).
  • Surrogates: generating-synthetic-surrogates — locale-native fakes and custom providers per language.
  • Policies: configuring-privacy-policies — policy= works with any lang.
  • Audit: auditing-deidentification-runs records the model and language in the no-PHI report.
  • Other surfaces: MCP openmed_deidentify / openmed_list_pii_languages; REST POST /pii/deidentify (both take a language parameter).

Edge cases & gotchas

  • Never run the English model on other languages. Recall collapses. Always pass lang=; the default model is English only.
  • lang ≠ locale. lang picks the detection model and patterns; locale shapes the replacement fakes. Set both when surrogate realism matters (e.g. lang="pt", locale="pt_BR").
  • Some locales are approximations. Faker has no Telugu locale, so OpenMed maps te→en_IN and warns once. Override locale= if you need closer regional surrogates.
  • Date order varies. Day-first languages (fr, de, it, es, nl, pt, …) parse 11/04/1979 as 11 April; the date logic is language-aware — keep lang set when shifting dates.
  • Mixed-language notes (e.g. English headers in a Spanish chart) may lower recall; verify residual risk with audit=True and consider a second pass.
  • No raw PHI in logs regardless of language — offsets, labels, hashes only.

Standards & references

  • HIPAA de-identification, 45 CFR 164.514(b) (US) and GDPR / national DPAs for EU subjects: https://www.hhs.gov/hipaa/for-professionals/privacy/special-topics/de-identification/index.html
  • National ID format references are encoded in OpenMed validators (no external registry bundled).
  • OpenMed source: openmed/core/pii_i18n.py (SUPPORTED_LANGUAGES, get_patterns_for_language, validators), openmed/core/model_registry.py (get_pii_models_by_language), openmed/core/anonymizer/locales.py (LANG_TO_LOCALE).

© maziyarpanahi, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/deidentifying-multilingual-text of maziyarpanahi/openmed.

Open the folder on GitHubat commit 9dca507

Compare with similar skills

Deidentifying Multilingual Text next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Deidentifying Multilingual Text compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Deidentifying Multilingual Text this skillmaziyarpanahi/openmed5.5k—~1.7kAutomated safety check: PassApache-2.0
Sanity Best Practicesrobotostudio/turbo-start-sanity183—~940Automated safety check: PassMIT
Mas SkillsAUTO-MAS-Project/AUTO-MAS711—~3.6kAutomated safety check: PassAGPL-3.0
SQL Optimization Patternsynulihao/AgentSkillOS61811 repos~3.3kAutomated safety check: PassNone
Datamodellmnimbalyst/nimbalyst1.9k—~713Automated safety check: PassMIT
Add Mpk Taskmirage-project/mirage2.5k—~4.5kAutomated safety check: PassApache-2.0

Similar skills

  • Sanity Best Practices

    robotostudio/turbo-start-sanity

    Sanity development best practices for schema design, GROQ queries, TypeGen, Visual Editing, images, Portable Text, Studio structure, localization, migrations, Sanity Functions, Blueprints, and…

    183 GitHub stars~940 tokensUpdated 4 days ago
    Frontend & DesignAuto-check passed
  • Mas Skills

    AUTO-MAS-Project/AUTO-MAS

    A skill your agent uses when a task needs AUTO-MAS engineering conventions across frontend, UI, code style, schema naming, module boundaries, function design, API contracts, data modeling, script…

    711 GitHub stars~3.6k tokensUpdated today
    DatabasesAuto-check passed
  • SQL Optimization Patterns

    ynulihao/AgentSkillOS

    Master SQL query optimization, indexing strategies, and EXPLAIN analysis to dramatically improve database performance and eliminate slow queries.

    618 GitHub starsUsed in 11 repos~3.3k tokens
    DatabasesAuto-check passed
  • Datamodellm

    nimbalyst/nimbalyst

    Create visual data models for database schemas using Nimbalyst's DataModelLM editor.

    1.9k GitHub stars~713 tokensUpdated yesterday
    DatabasesAuto-check passed
  • Add Mpk Task

    mirage-project/mirage

    Step-by-step guide for adding a new task implementation to Mirage Persistent Kernel (MPK).

    2.5k GitHub stars~4.5k tokensUpdated 2 days ago
    DatabasesAuto-check passed
  • B200 Flash Attention4 Planner

    mirage-project/mirage

    A skill your agent uses when the user wants to design or extend a FlashAttention-style forward kernel on B200/Blackwell, involving the two MMAs QKᵀ and PV, online softmax, S/P/O in TMEM, warp roles…

    2.5k GitHub stars~1.9k tokensUpdated 2 days ago
    DatabasesAuto-check passed

More from maziyarpanahi/openmed

All 74 skills in this repo
  • Checks OpenMed de-identified clinical text against the 18 HIPAA Safe Harbor identifier categories and reports gaps and residual re-identification risk.

    5.5k GitHub stars~1.7k tokensUpdated today
    Auto-check passed
  • OpenMed Model Card Writer

    maziyarpanahi/openmed

    Fills in a model card for an OpenMed clinical NER or de-identification model from its evaluation reports: intended use, metrics, subgroups and limitations.

    5.5k GitHub stars~1.8k tokensUpdated today
    Auto-check passed
  • Walks a data pipeline against the HIPAA Privacy and Security Rule checklist and produces a gap report before it processes patient data.

    5.5k GitHub stars~2k tokensUpdated today
    Auto-check passed
  • ICD-10 Coding Assistant

    maziyarpanahi/openmed

    Suggests candidate ICD-10-CM diagnosis and ICD-10-PCS procedure codes for clinical text extracted by OpenMed, with rationale for a certified coder to review.

    5.5k GitHub stars~2k tokensUpdated today
    Auto-check passed
  • OpenMed ETL to OMOP CDM

    maziyarpanahi/openmed

    Maps OpenMed-extracted, terminology-coded conditions, drugs and measurements into OMOP CDM v5.4 tables for OHDSI and ATLAS analytics.

    5.5k GitHub stars~1.9k tokensUpdated today
    Auto-check passed
  • Extracting SDOH and Z-Codes

    maziyarpanahi/openmed

    Finds social risks such as housing instability or food insecurity in clinical notes and proposes matching ICD-10-CM Z-codes for a coder to confirm.

    5.5k GitHub stars~1.9k tokensUpdated today
    Auto-check passed

Questions about Deidentifying Multilingual Text

What does Deidentifying Multilingual Text do?

De-identify non-English clinical text on-device with OpenMed by passing lang= and locale= to deidentify(). Deidentifying Multilingual Text is an agent skill from maziyarpanahi/openmed. De-identify non-English clinical text on-device with OpenMed by passing lang= and locale= to deidentify().

When should I use Deidentifying Multilingual Text?

Deidentifying Multilingual Text fits situations like: the user has Spanish; turkish medical notes; needs locale-aware fake surrogates; must handle language-specific national IDs (DNI.

How do I install Deidentifying Multilingual Text in Claude Code?

Run `npx skills add maziyarpanahi/openmed --skill deidentifying-multilingual-text -a claude-code`. Or copy the skill folder (skills/deidentifying-multilingual-text in maziyarpanahi/openmed) into .claude/skills/deidentifying-multilingual-text in your project. Claude Code loads it when a task matches its description.

How do I install Deidentifying Multilingual Text in Codex?

Run `npx skills add maziyarpanahi/openmed --skill deidentifying-multilingual-text -a codex`. Or copy the skill folder (skills/deidentifying-multilingual-text in maziyarpanahi/openmed) into .agents/skills/deidentifying-multilingual-text in your project. Codex loads it when a task matches its description.

Can I use Deidentifying Multilingual Text in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add maziyarpanahi/openmed --skill deidentifying-multilingual-text -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/deidentifying-multilingual-text, .gemini/skills/deidentifying-multilingual-text, .github/skills/deidentifying-multilingual-text and .opencode/skills/deidentifying-multilingual-text in your project.

What does Deidentifying Multilingual Text need to run?

SKILL.md names no scripts, command-line tools or credentials: Deidentifying Multilingual Text is instructions for the agent only. Our summary lists: Python 3.

Does Deidentifying Multilingual Text access the network?

SKILL.md names 1 domain. As links in the text: hhs.gov. This is read from the text; nothing was executed.

Is Deidentifying Multilingual Text safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Deidentifying Multilingual Text use?

Deidentifying Multilingual Text is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Deidentifying Multilingual Text use?

About 1.7k tokens (SKILL.md is roughly 6.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Deidentifying Multilingual Text?

Skills that share tags, products or a category with Deidentifying Multilingual Text: Sanity Best Practices (robotostudio/turbo-start-sanity, 183 stars), Mas Skills (AUTO-MAS-Project/AUTO-MAS, 711 stars), SQL Optimization Patterns (ynulihao/AgentSkillOS, 618 stars) and Datamodellm (nimbalyst/nimbalyst, 1.9k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Deidentifying Multilingual Text?

maziyarpanahi (a GitHub user) maintains it in maziyarpanahi/openmed, which has 5,500 GitHub stars. The repository holds 74 skills in this directory. The repository was last updated on October 9, 2026.

Source: maziyarpanahi/openmed on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.