Agent skill

Auditing Subgroup Fairness

by maziyarpanahi in maziyarpanahi/openmed

Audit an OpenMed NER or de-identification model for performance disparities across demographic subgroups (sex, age band, race/ethnicity when available) using openmed.eval.fairnessreport.

Apache-2.0Auto-check passed

Install Auditing Subgroup Fairness

skills CLI
$ npx skills add maziyarpanahi/openmed --skill auditing-subgroup-fairness -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install maziyarpanahi/openmed auditing-subgroup-fairness --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/maziyarpanahi/openmed.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/auditing-subgroup-fairness .claude/skills/auditing-subgroup-fairness && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
auditing-subgroup-fairness
GitHub stars
5.5k
Token cost
~1.5k tokens
SKILL.md length
540 words
Files
1
Skills in repo
74
Repo updated
First seen
Licence
Apache-2.0

At a glance

Audit an OpenMed NER or de-identification model for performance disparities across demographic subgroups (sex, age band, race/ethnicity when available) using openmed.eval.fairnessreport.

  • Works in 6 steps: Tag the gold corpus by group. Add a… → Run fairness_report on the model + suite. → Read leakage first, recall second. For… → …
  • The user wants per-subgroup recall and leakage
  • SKILL.md covers When to use this skill, What it measures, Quick start and Workflow, plus 3 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Auditing Subgroup Fairness is an agent skill from maziyarpanahi/openmed. Audit an OpenMed NER or de-identification model for performance disparities across demographic subgroups (sex, age band, race/ethnicity when available) using openmed.eval.fairnessreport. Use when the user wants per-subgroup recall and leakage, wants to check whether de-identification under-protects a group, wants to surface a documentation gap where subgroup data is missing, or needs equalized-odds-style disparity numbers for a clinical model. Trigger on "fairness", "subgroup", "bias audit", "disparity"…

Its SKILL.md is about 1.5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

The repository describes itself as: Local-first healthcare AI: clinical NER and HIPAA PII de-identification on hardware you control. 2,200+ medical models, 35 model-backed PII languages, and Python, MLX, Android… The licence is Apache-2.0.

When your agent uses it

  • The user wants per-subgroup recall and leakage
  • Wants to check whether de-identification under-protects a group
  • Wants to surface a documentation gap where subgroup data is missing
  • Needs equalized-odds-style disparity numbers for a clinical model

Example prompts

  • “fairness”
  • “subgroup”
  • “bias audit”
  • “/auditing-subgroup-fairness”

Requirements

  • Python 3

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Tag the gold corpus by group. Add a synthetic group to each PHI span's
  2. Run fairness_report on the model + suite.
  3. Read leakage first, recall second. For de-id, a group with higher leakage
  4. Compute the disparity (leakage_disparity) and locate worst_group.
  5. Document the gap. If race/ethnicity surrogates are absent, report that the
  6. Feed it forward. Put per-group numbers and the gap into the model card and

What it can do on your machine

Read from SKILL.md and the folder at commit 34d7b8c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • datadiversity.org
    • arxiv.org
    • doi.org

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Auditing Subgroup Fairness loads about 1.5k tokens when it runs. Until then it costs about 161 tokens; SKILL.md has 540 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~161
When it runs · the whole SKILL.md, loaded when a task matches
~1.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from maziyarpanahi/openmed at commit 34d7b8c, republished under its Apache-2.0 licence (© maziyarpanahi). 540 words, ~1,471 tokens.

Download SKILL.mdSave it as .claude/skills/auditing-subgroup-fairness/SKILL.md (or your agent's skills folder).
name
auditing-subgroup-fairness
description
Audit an OpenMed NER or de-identification model for performance disparities across demographic subgroups (sex, age band, race/ethnicity when available) using openmed.eval.fairness_report. Use when the user wants per-subgroup recall and leakage, wants to check whether de-identification under-protects a group, wants to surface a documentation gap where subgroup data is missing, or needs equalized-odds-style disparity numbers for a clinical model. Trigger on "fairness", "subgroup", "bias audit", "disparity", "equalized odds", "under-protected group", "per-group recall", or "STANDING Together" for an OpenMed model.
license
Apache-2.0
metadata.project
OpenMed
metadata.category
evaluation-quality
metadata.pairs
adjacent
metadata.version
1.0

Auditing Subgroup Fairness

An aggregate pass can hide a group the model fails. For de-identification that failure has a name: under-protection — PHI that leaks more often for one demographic group than another. openmed.eval.fairness_report slices leakage and recall by gold-span group so disparities surface before deployment, not after a breach.

When to use this skill

  • You want per-subgroup recall and leakage for a de-id or NER model.
  • You suspect (or must rule out) that one group is under-protected.
  • You need disparity numbers for a clinical AI governance review.
  • You need to document which subgroups you couldn't evaluate (the data gap).

What it measures

For each surrogate group fairness_report returns:

  • leakage_rate — fraction of that group's gold PHI characters left exposed (the de-id harm metric).
  • recall — fraction of that group's gold spans detected.
  • leakage_disparity — max - min leakage across groups (the gap to close).
  • worst_group / worst_group_leakage — the most-failed group.

Group membership comes from a group tag in each gold span's metadata (keys group, demographic_group, or surrogate_group); ungrouped spans fall into unspecified.

Quick start

python
from openmed.eval import fairness_report

# Gold fixtures must tag spans with a surrogate group, e.g.
#   {"start": 4, "end": 12, "label": "PERSON", "metadata": {"group": "female"}}
fair = fairness_report(
    "OpenMed/Privacy-PII-Detection",
    "golden",                 # named suite, or pass a list of fixtures
    device="cpu",
)

print("leakage disparity:", fair.leakage_disparity)
print("worst group      :", fair.worst_group, fair.worst_group_leakage)
for group, m in sorted(fair.per_group.items()):
    print(f"  {group:14s} recall={m.recall:.3f}  leakage={m.leakage_rate:.4f}")

# Under-protection alarm: any group leaking more than the rest.
LEAKAGE_GAP_LIMIT = 0.0       # leakage-first: ideally zero leakage everywhere
assert fair.leakage_disparity <= LEAKAGE_GAP_LIMIT or fair.worst_group_leakage == 0

FairnessReport.to_dict() is JSON-ready and PHI-free — drop it straight into a model card.

Workflow

  1. Tag the gold corpus by group. Add a synthetic group to each PHI span's metadata (sex, age band, race/ethnicity surrogate). Use synthetic surrogates, not real protected attributes (see building-gold-corpus).
  2. Run fairness_report on the model + suite.
  3. Read leakage first, recall second. For de-id, a group with higher leakage is under-protected — that is the headline finding.
  4. Compute the disparity (leakage_disparity) and locate worst_group. Equalized-odds framing: equal true-positive (recall) and equal leakage across groups.
  5. Document the gap. If race/ethnicity surrogates are absent, report that the audit could not cover them — most clinical NLP studies omit race entirely, so silence is the default failure mode, not equity.
  6. Feed it forward. Put per-group numbers and the gap into the model card and the governance review.
Show full SKILL.md (232 more words)Show less

Hand-off to / from OpenMed

  • From building-gold-corpus: supplies group-tagged synthetic fixtures.
  • From evaluating-with-leakage-gates: an aggregate RELEASABLE decision should be paired with this audit — overall pass, subgroup fail is exactly the trap this catches.
  • To authoring-model-cards: FairnessReport.to_dict() fills the quantitative-analysis / subgroup section.
  • Pairs with benchmarking-clinical-ner: same run, different slice (label vs group).

Edge cases & gotchas

  • Under-protection is the de-id harm; lead with leakage. A group with equal recall but higher leakage is still failed.
  • The race documentation gap is the norm. Most clinical NLP corpora don't record race/ethnicity, so most fairness audits silently can't measure it. Report the absence explicitly — don't let missing data read as parity.
  • unspecified is not a real group. A pile of spans in unspecified means your gold isn't tagged; fix the corpus before trusting the disparity.
  • Small groups give noisy rates. Report span counts (span_count, total_chars) alongside rates; a 1-of-2 leak isn't a 50% population rate.
  • Synthetic surrogates only. Never store real protected attributes in eval fixtures; use fabricated group labels for slicing.
  • Disparity ≈ 0 with high leakage everywhere is not "fair". Equal failure is still failure — check absolute leakage, not just the gap.

Standards & references

© maziyarpanahi, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/auditing-subgroup-fairness of maziyarpanahi/openmed.

Open the folder on GitHubat commit 34d7b8c

Compare with similar skills

Auditing Subgroup Fairness next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Auditing Subgroup Fairness compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Auditing Subgroup Fairness this skillmaziyarpanahi/openmed5.5k—~1.5kAutomated safety check: PassApache-2.0
Production Auditaffaan-m/ECC277k1 repos~1.9kAutomated safety check: PassMIT
Aims Auditalirezarezvani/claude-skills28k—~1.3kAutomated safety check: PassMIT
Geo Auditsickn33/agentic-awesome-skills47k1 repos~3.4kAutomated safety check: NotesMIT
OmniRoute Audit and Policy CLIdiegosouzapw/OmniRoute75k—~733Automated safety check: PassMIT
Audit Preparationsickn33/agentic-awesome-skills47k1 repos~5.3kAutomated safety check: PassMIT

Similar skills

  • Production Audit

    affaan-m/ECC

    Local-evidence production readiness audit for shipped apps, pre-launch reviews, post-merge checks, and "what breaks in prod?" questions without sending repo data to an external audit service.

    277k GitHub starsUsed in 1 repo~1.9k tokens
    Product & Project ManagementAuto-check passed
  • Aims Audit

    alirezarezvani/claude-skills

    /cs:aims-audit <scope — ISO/IEC 42001 AIMS internal-audit 6-question forcing interrogation.

    28k GitHub stars~1.3k tokensUpdated 1 mo ago
    Legal & ComplianceAuto-check passed
  • Geo Audit

    sickn33/agentic-awesome-skills

    Full website GEO+SEO audit with parallel subagent delegation.

    47k GitHub starsUsed in 1 repo~3.4k tokens
    Marketing & SEOAuto-check: notes
  • OmniRoute Audit and Policy CLI

    diegosouzapw/OmniRoute

    Command reference for omniroute's audit, logs, policy and telemetry commands: search and export audit trails, manage access policies and review request history for compliance work.

    75k GitHub stars~733 tokensUpdated today
    SecurityAuto-check passed
  • Audit Preparation

    sickn33/agentic-awesome-skills

    Audit preparation register: required document, period covered, request and receipt dates, preparer and reviewer, auditor queries and adjustments.

    47k GitHub starsUsed in 1 repo~5.3k tokens
    Legal & ComplianceAuto-check passed
  • Qms Audit Expert

    davila7/claude-code-templates

    Senior QMS Audit Expert for internal and external quality management system auditing.

    33k GitHub starsUsed in 1 repo~2.7k tokens
    Legal & ComplianceAuto-check passed

More from maziyarpanahi/openmed

All 74 skills in this repo
  • Checks OpenMed de-identified clinical text against the 18 HIPAA Safe Harbor identifier categories and reports gaps and residual re-identification risk.

    5.5k GitHub stars~1.7k tokensUpdated today
    Auto-check passed
  • OpenMed Model Card Writer

    maziyarpanahi/openmed

    Fills in a model card for an OpenMed clinical NER or de-identification model from its evaluation reports: intended use, metrics, subgroups and limitations.

    5.5k GitHub stars~1.8k tokensUpdated today
    Auto-check passed
  • Walks a data pipeline against the HIPAA Privacy and Security Rule checklist and produces a gap report before it processes patient data.

    5.5k GitHub stars~2k tokensUpdated today
    Auto-check passed
  • ICD-10 Coding Assistant

    maziyarpanahi/openmed

    Suggests candidate ICD-10-CM diagnosis and ICD-10-PCS procedure codes for clinical text extracted by OpenMed, with rationale for a certified coder to review.

    5.5k GitHub stars~2k tokensUpdated today
    Auto-check passed
  • OpenMed ETL to OMOP CDM

    maziyarpanahi/openmed

    Maps OpenMed-extracted, terminology-coded conditions, drugs and measurements into OMOP CDM v5.4 tables for OHDSI and ATLAS analytics.

    5.5k GitHub stars~1.9k tokensUpdated today
    Auto-check passed
  • Extracting SDOH and Z-Codes

    maziyarpanahi/openmed

    Finds social risks such as housing instability or food insecurity in clinical notes and proposes matching ICD-10-CM Z-codes for a coder to confirm.

    5.5k GitHub stars~1.9k tokensUpdated today
    Auto-check passed

Questions about Auditing Subgroup Fairness

What does Auditing Subgroup Fairness do?

Audit an OpenMed NER or de-identification model for performance disparities across demographic subgroups (sex, age band, race/ethnicity when available) using openmed.eval.fairnessreport. Auditing Subgroup Fairness is an agent skill from maziyarpanahi/openmed.fairnessreport.

When should I use Auditing Subgroup Fairness?

Auditing Subgroup Fairness fits situations like: the user wants per-subgroup recall and leakage; wants to check whether de-identification under-protects a group; wants to surface a documentation gap where subgroup data is missing; needs equalized-odds-style disparity numbers for a clinical model.

How do I install Auditing Subgroup Fairness in Claude Code?

Run `npx skills add maziyarpanahi/openmed --skill auditing-subgroup-fairness -a claude-code`. Or copy the skill folder (skills/auditing-subgroup-fairness in maziyarpanahi/openmed) into .claude/skills/auditing-subgroup-fairness in your project. Claude Code loads it when a task matches its description.

How do I install Auditing Subgroup Fairness in Codex?

Run `npx skills add maziyarpanahi/openmed --skill auditing-subgroup-fairness -a codex`. Or copy the skill folder (skills/auditing-subgroup-fairness in maziyarpanahi/openmed) into .agents/skills/auditing-subgroup-fairness in your project. Codex loads it when a task matches its description.

Can I use Auditing Subgroup Fairness in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add maziyarpanahi/openmed --skill auditing-subgroup-fairness -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/auditing-subgroup-fairness, .gemini/skills/auditing-subgroup-fairness, .github/skills/auditing-subgroup-fairness and .opencode/skills/auditing-subgroup-fairness in your project.

What does Auditing Subgroup Fairness need to run?

SKILL.md names no scripts, command-line tools or credentials: Auditing Subgroup Fairness is instructions for the agent only. Our summary lists: Python 3.

Does Auditing Subgroup Fairness access the network?

SKILL.md names 3 domains. As links in the text: datadiversity.org, arxiv.org and doi.org. This is read from the text; nothing was executed.

Is Auditing Subgroup Fairness safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Auditing Subgroup Fairness use?

Auditing Subgroup Fairness is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Auditing Subgroup Fairness use?

About 1.5k tokens (SKILL.md is roughly 5.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Auditing Subgroup Fairness?

Skills that share tags, products or a category with Auditing Subgroup Fairness: Production Audit (affaan-m/ECC, 277k stars), Aims Audit (alirezarezvani/claude-skills, 28k stars), Geo Audit (sickn33/agentic-awesome-skills, 47k stars) and OmniRoute Audit and Policy CLI (diegosouzapw/OmniRoute, 75k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Auditing Subgroup Fairness?

maziyarpanahi (a GitHub user) maintains it in maziyarpanahi/openmed, which has 5,506 GitHub stars. The repository holds 74 skills in this directory. The repository was last updated on October 11, 2026.

Source: maziyarpanahi/openmed on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.