Agent skill

Reviewing Reidentification Risk

by maziyarpanahi in maziyarpanahi/openmed

Run expert-determination-style quasi-identifier risk scoring (k-anonymity, l-diversity) plus OpenMed's empirical re-identification attack on a de-identified dataset, then document residual risk in a…

Apache-2.0Auto-check passedLegal & Compliance

Install Reviewing Reidentification Risk

skills CLI
$ npx skills add maziyarpanahi/openmed --skill reviewing-reidentification-risk -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install maziyarpanahi/openmed reviewing-reidentification-risk --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/maziyarpanahi/openmed.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/reviewing-reidentification-risk .claude/skills/reviewing-reidentification-risk && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
reviewing-reidentification-risk
GitHub stars
5.5k
Token cost
~1.9k tokens
SKILL.md length
683 words
Files
1
Skills in repo
74
Repo updated
First seen
Licence
Apache-2.0

At a glance

Run expert-determination-style quasi-identifier risk scoring (k-anonymity, l-diversity) plus OpenMed's empirical re-identification attack on a de-identified dataset, then document residual risk in a…

  • Works in 6 steps: Enumerate quasi-identifiers (QIs). List… → Compute k-anonymity. For each… → Check l-diversity on sensitive… → …
  • The user needs HIPAA Expert Determination (45 CFR 164.514(b)
  • SKILL.md covers When to use, Quick start, Workflow and Hand-off to / from OpenMed, plus 2 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Reviewing Reidentification Risk is an agent skill from maziyarpanahi/openmed. Run expert-determination-style quasi-identifier risk scoring (k-anonymity, l-diversity) plus OpenMed's empirical re-identification attack on a de-identified dataset, then document residual risk in a defensible memo. Use when the user needs HIPAA Expert Determination (45 CFR 164.514(b)(1)) support, asks whether a dataset is safe to release, worries about singling-out via age/ZIP/dates, or wants a statistical "very small risk" determination. Covers identifying quasi-identifiers, computing k-anonymity / l-diversity…

Its SKILL.md is about 1.9k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Legal & Compliance, covering Healthcare and finance regulation. The repository describes itself as: Local-first healthcare AI: clinical NER and HIPAA PII de-identification on hardware you control. 2,200+ medical models, 35 model-backed PII languages, and Python, MLX, Android… The licence is Apache-2.0.

When your agent uses it

  • The user needs HIPAA Expert Determination (45 CFR 164.514(b)
  • Asks whether a dataset is safe to release
  • Worries about singling-out via age/ZIP/dates
  • Wants a statistical very small risk determination

Example prompts

  • “very small risk”
  • “/reviewing-reidentification-risk”

Requirements

  • Python 3

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Enumerate quasi-identifiers (QIs). List every field an outsider could
  2. Compute k-anonymity. For each equivalence class (records sharing the same
  3. Check l-diversity on sensitive attributes. k-anonymity hides which
  4. Run the empirical attack. run_reid_attack / run_reid_benchmark model an
  5. Generalize or suppress, then re-score. For singletons / low-k classes,
  6. Write the determination memo. Record the QIs considered, methods applied,

What it can do on your machine

Read from SKILL.md and the folder at commit 34d7b8c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • hhs.gov
    • dataprivacylab.org
    • dl.acm.org
    • csrc.nist.gov

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Reviewing Reidentification Risk loads about 1.9k tokens when it runs. Until then it costs about 186 tokens; SKILL.md has 683 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~186
When it runs · the whole SKILL.md, loaded when a task matches
~1.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from maziyarpanahi/openmed at commit 34d7b8c, republished under its Apache-2.0 licence (© maziyarpanahi). 683 words, ~1,910 tokens.

Download SKILL.mdSave it as .claude/skills/reviewing-reidentification-risk/SKILL.md (or your agent's skills folder).
name
reviewing-reidentification-risk
description
Run expert-determination-style quasi-identifier risk scoring (k-anonymity, l-diversity) plus OpenMed's empirical re-identification attack on a de-identified dataset, then document residual risk in a defensible memo. Use when the user needs HIPAA Expert Determination (45 CFR 164.514(b)(1)) support, asks whether a dataset is safe to release, worries about singling-out via age/ZIP/dates, or wants a statistical "very small risk" determination. Covers identifying quasi-identifiers, computing k-anonymity / l-diversity, running openmed.eval.attacks.reid (run_reid_attack / run_reid_benchmark) as the adversarial attack, and writing the risk memo. Pairs after deidentifying-clinical-text and auditing-deid-leakage.
license
Apache-2.0
metadata.project
OpenMed
metadata.category
de-identification
metadata.pairs
after
metadata.version
1.0

Reviewing re-identification risk

Removing direct identifiers is not enough. A record stripped of name, SSN, and MRN can still be singled out by a combination of quasi-identifiers — age, ZIP/region, admission date, sex, rare diagnosis. The HIPAA Expert Determination pathway (45 CFR 164.514(b)(1)) requires a qualified person to apply statistical methods and document that the risk of re-identification is "very small." This skill produces that evidence: quasi-identifier risk metrics (k-anonymity, l-diversity) plus OpenMed's empirical re-identification attack, written up as a residual-risk memo.

When to use

  • After direct-identifier removal passes auditing-deid-leakage (no leaks) and you must decide whether the dataset is releasable.
  • The user invokes Expert Determination, asks for a re-identification risk score, k-anonymity, l-diversity, or a "very small risk" determination memo.
  • You need an adversarial linkage attack — modeling an attacker with auxiliary data — not just a structural metric.

Quick start

python
from openmed.eval.attacks.reid import run_reid_attack, run_reid_benchmark

# Synthetic de-identified records; each row is the released, de-id'd data.
deidentified = [
    {"record_id": "r1", "text": "[NAME], 47F, ZIP 021xx, admitted 2024-03."},
    {"record_id": "r2", "text": "[NAME], 47F, ZIP 021xx, admitted 2024-03."},
    {"record_id": "r3", "text": "[NAME], 88M, ZIP 597xx, admitted 2024-03."},  # singleton
]
# Auxiliary = what an attacker might already hold (e.g. a voter list).
auxiliary = [{"record_id": "v9", "text": "88M ZIP 597xx"}]

result = run_reid_attack(
    fixtures=[],                          # bring your own records below
    deidentified_records=deidentified,
    auxiliary_records=auxiliary,
)
metric = result.to_metric()
print(metric["aux_linkage_rate"],        # empirical linkage success
      metric["k_min"],                   # smallest equivalence-class size
      metric["singleton_count"],         # k=1 records (uniquely identifiable)
      metric["quasi_identifier_count"])

k_min is the population k-anonymity floor across the dataset; a k_min of 1 means at least one record is unique on its quasi-identifiers and is the highest re-identification risk. aux_linkage_rate is the empirical attack: how often the adversary's auxiliary data successfully links back to a released record.

To run against the bundled golden suite and emit a leaderboard-style report:

python
report = run_reid_benchmark(
    suite="golden",
    deidentified_records=deidentified,
    auxiliary_records=auxiliary,
    output_markdown="reid_risk.md",
)

Workflow

  1. Enumerate quasi-identifiers (QIs). List every field an outsider could plausibly know and cross-reference: age/DOB, ZIP/region, dates of service, sex, race, rare conditions, provider. Direct identifiers should already be gone (verified by auditing-deid-leakage); QIs are what's left to worry about.
  2. Compute k-anonymity. For each equivalence class (records sharing the same QI combination), the class size is k. run_reid_attack returns k_min and the list of singleton_records (k=1). A common Expert Determination target is k ≥ a documented threshold (e.g. k ≥ 5 or k ≥ 11) for every record.
  3. Check l-diversity on sensitive attributes. k-anonymity hides which record, but if every record in a class shares the same sensitive value (e.g. all HIV-positive), the attribute leaks anyway. Require ≥ l distinct sensitive values per class; flag homogeneous classes.
  4. Run the empirical attack. run_reid_attack / run_reid_benchmark model an adversary with auxiliary_records and measure actual linkage success (aux_linkage_rate), residual leakage (leakage_rate), surrogate-consistency leaks, and date-shift-inversion leaks. Structural metrics bound risk; the attack demonstrates it.
  5. Generalize or suppress, then re-score. For singletons / low-k classes, coarsen QIs (age → age band, ZIP5 → ZIP3, exact date → month/quarter) or suppress the record, then re-run until k_min and linkage rate meet your documented threshold.
  6. Write the determination memo. Record the QIs considered, methods applied, k_min, l-diversity, the attack's aux_linkage_rate, the assumptions about attacker capability, and the conclusion that residual risk is "very small." Cite the metrics — never paste raw records into the memo.
Show full SKILL.md (250 more words)Show less

Hand-off to / from OpenMed

  • From auditing-deid-leakage: only score QI risk once direct-identifier leakage is zero. A leak short-circuits the whole determination.
  • OpenMed calls: from openmed.eval.attacks.reid import run_reid_attack, run_reid_benchmark, generate_reid_leaderboard. The attack delegates to openmed.risk.risk_report for k-anonymity / linkage internals.
  • To evaluating-with-leakage-gates: register reid_leakage_rate as a gate in the eval harness so re-identification risk regressions fail CI.
  • From pseudonymizing-for-gdpr: pseudonymized output is still re-identifiable via QIs — run this attack before claiming a dataset is low-risk or anonymized.

Edge cases & gotchas

  • Expert Determination is a human judgment. OpenMed produces the statistics; a qualified expert signs the determination. The tool supports the memo, it is not the memo.
  • Auxiliary data assumptions drive the result. Linkage rate is only as meaningful as the auxiliary_records you model. Document the assumed attacker (motivated insider vs. public voter list) — different aux sets, different risk.
  • Singletons are the headline. A single k=1 record can sink a release; check singleton_count and singleton_records first.
  • Date-shift can be inverted. Preserving intervals across a date shift lets an attacker re-anchor the timeline; the attack flags date_shift_inversion_rate. Watch it when de-id used method="shift_dates".
  • No raw records in artifacts. Reports carry counts, rates, and offsets. Keep the underlying dataset out of the memo and out of logs.
  • Local-first. Run the attack on-device; never ship candidate-release data to a third party to "test" re-identifiability.

Standards & references

© maziyarpanahi, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/reviewing-reidentification-risk of maziyarpanahi/openmed.

Open the folder on GitHubat commit 34d7b8c

Compare with similar skills

Reviewing Reidentification Risk next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Reviewing Reidentification Risk compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Reviewing Reidentification Risk this skillmaziyarpanahi/openmed5.5k—~1.9kAutomated safety check: PassApache-2.0
Hipaa ComplianceSushegaad/Claude-Skills-Governance-Risk-and-Compliance9461 repos~2.3kAutomated safety check: PassMIT
ISO Standards Readiness EvidenceK-Dense-AI/scientific-agent-skills48k1 repos~4.6kAutomated safety check: NotesMIT
Fda Consultant Specialistdavila7/claude-code-templates33k1 repos~2.7kAutomated safety check: PassMIT
Grc Knowledgemlunato47/claude-grc-plugin184—~6.1kAutomated safety check: PassMIT
Clinical Reportsdavila7/claude-code-templates33k11 repos~9.9kAutomated safety check: NotesMIT

Similar skills

  • Hipaa Compliance

    Sushegaad/Claude-Skills-Governance-Risk-and-Compliance

    Expert HIPAA compliance assistant for healthcare and software contexts.

    946 GitHub starsUsed in 1 repo~2.3k tokens
    Legal & ComplianceAuto-check passed
  • ISO Standards Readiness Evidence

    K-Dense-AI/scientific-agent-skills

    Organizes scope, controlled documents, risk files and traceability into draft evidence for human review against ISO 13485, 14971, 17025 and 15189.

    48k GitHub starsUsed in 1 repo~4.6k tokens
    Legal & ComplianceAuto-check: notes
  • Fda Consultant Specialist

    davila7/claude-code-templates

    Senior FDA consultant and specialist for medical device companies including HIPAA compliance and requirement management.

    33k GitHub starsUsed in 1 repo~2.7k tokens
    Legal & ComplianceAuto-check passed
  • Grc Knowledge

    mlunato47/claude-grc-plugin

    Senior GRC analyst expertise across 18 compliance frameworks — NIST 800-53, FedRAMP (Rev5 + 20x/CR26, KSIs, VDR/VER, Certification Classes A–D), DoD/DoW Impact Levels (IL2–IL6, DISA Cloud SRG), ITAR…

    184 GitHub stars~6.1k tokensUpdated 4 days ago
    Legal & ComplianceAuto-check passed
  • Clinical Reports

    davila7/claude-code-templates

    Write comprehensive clinical reports including case reports (CARE guidelines), diagnostic reports (radiology/pathology/lab), clinical trial reports (ICH-E3, SAE, CSR), and patient documentation…

    33k GitHub starsUsed in 11 repos~9.9k tokens
    Legal & ComplianceAuto-check: notes
  • Audit PCI DSS penetration-test methodology, scope, internal and external reports, segmentation results, tester independence, remediation, retesting, retention, and multi-tenant support evidence.

    135 GitHub stars~1k tokensUpdated 1 mo ago
    Legal & ComplianceAuto-check passed

More from maziyarpanahi/openmed

All 74 skills in this repo
  • Checks OpenMed de-identified clinical text against the 18 HIPAA Safe Harbor identifier categories and reports gaps and residual re-identification risk.

    5.5k GitHub stars~1.7k tokensUpdated today
    Auto-check passed
  • OpenMed Model Card Writer

    maziyarpanahi/openmed

    Fills in a model card for an OpenMed clinical NER or de-identification model from its evaluation reports: intended use, metrics, subgroups and limitations.

    5.5k GitHub stars~1.8k tokensUpdated today
    Auto-check passed
  • Walks a data pipeline against the HIPAA Privacy and Security Rule checklist and produces a gap report before it processes patient data.

    5.5k GitHub stars~2k tokensUpdated today
    Auto-check passed
  • ICD-10 Coding Assistant

    maziyarpanahi/openmed

    Suggests candidate ICD-10-CM diagnosis and ICD-10-PCS procedure codes for clinical text extracted by OpenMed, with rationale for a certified coder to review.

    5.5k GitHub stars~2k tokensUpdated today
    Auto-check passed
  • OpenMed ETL to OMOP CDM

    maziyarpanahi/openmed

    Maps OpenMed-extracted, terminology-coded conditions, drugs and measurements into OMOP CDM v5.4 tables for OHDSI and ATLAS analytics.

    5.5k GitHub stars~1.9k tokensUpdated today
    Auto-check passed
  • Extracting SDOH and Z-Codes

    maziyarpanahi/openmed

    Finds social risks such as housing instability or food insecurity in clinical notes and proposes matching ICD-10-CM Z-codes for a coder to confirm.

    5.5k GitHub stars~1.9k tokensUpdated today
    Auto-check passed

Questions about Reviewing Reidentification Risk

What does Reviewing Reidentification Risk do?

Run expert-determination-style quasi-identifier risk scoring (k-anonymity, l-diversity) plus OpenMed's empirical re-identification attack on a de-identified dataset, then document residual risk in a…. Reviewing Reidentification Risk is an agent skill from maziyarpanahi/openmed. Run expert-determination-style quasi-identifier risk scoring (k-anonymity, l-diversity) plus OpenMed's empirical re-identification attack on a de-identified dataset, then document residual risk in a defensible memo.

When should I use Reviewing Reidentification Risk?

Reviewing Reidentification Risk fits situations like: the user needs HIPAA Expert Determination (45 CFR 164.514(b); asks whether a dataset is safe to release; worries about singling-out via age/ZIP/dates; wants a statistical very small risk determination.

How do I install Reviewing Reidentification Risk in Claude Code?

Run `npx skills add maziyarpanahi/openmed --skill reviewing-reidentification-risk -a claude-code`. Or copy the skill folder (skills/reviewing-reidentification-risk in maziyarpanahi/openmed) into .claude/skills/reviewing-reidentification-risk in your project. Claude Code loads it when a task matches its description.

How do I install Reviewing Reidentification Risk in Codex?

Run `npx skills add maziyarpanahi/openmed --skill reviewing-reidentification-risk -a codex`. Or copy the skill folder (skills/reviewing-reidentification-risk in maziyarpanahi/openmed) into .agents/skills/reviewing-reidentification-risk in your project. Codex loads it when a task matches its description.

Can I use Reviewing Reidentification Risk in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add maziyarpanahi/openmed --skill reviewing-reidentification-risk -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/reviewing-reidentification-risk, .gemini/skills/reviewing-reidentification-risk, .github/skills/reviewing-reidentification-risk and .opencode/skills/reviewing-reidentification-risk in your project.

What does Reviewing Reidentification Risk need to run?

SKILL.md names no scripts, command-line tools or credentials: Reviewing Reidentification Risk is instructions for the agent only. Our summary lists: Python 3.

Does Reviewing Reidentification Risk access the network?

SKILL.md names 4 domains. As links in the text: hhs.gov, dataprivacylab.org, dl.acm.org and csrc.nist.gov. This is read from the text; nothing was executed.

Is Reviewing Reidentification Risk safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Reviewing Reidentification Risk use?

Reviewing Reidentification Risk is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Reviewing Reidentification Risk use?

About 1.9k tokens (SKILL.md is roughly 7.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Reviewing Reidentification Risk?

Skills that share tags, products or a category with Reviewing Reidentification Risk: Hipaa Compliance (Sushegaad/Claude-Skills-Governance-Risk-and-Compliance, 946 stars), ISO Standards Readiness Evidence (K-Dense-AI/scientific-agent-skills, 48k stars), Fda Consultant Specialist (davila7/claude-code-templates, 33k stars) and Grc Knowledge (mlunato47/claude-grc-plugin, 184 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Reviewing Reidentification Risk?

maziyarpanahi (a GitHub user) maintains it in maziyarpanahi/openmed, which has 5,506 GitHub stars. The repository holds 74 skills in this directory. The repository was last updated on October 11, 2026.

Source: maziyarpanahi/openmed on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.