Agent skill

Benchmarking Clinical Ner

by maziyarpanahi in maziyarpanahi/openmed

Score an OpenMed clinical or biomedical NER model against a user-supplied gold corpus with entity-level precision, recall, and F1, then break errors down per label.

Apache-2.0Auto-check passedAI & LLM Engineering

Install Benchmarking Clinical Ner

skills CLI
$ npx skills add maziyarpanahi/openmed --skill benchmarking-clinical-ner -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install maziyarpanahi/openmed benchmarking-clinical-ner --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/maziyarpanahi/openmed.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/benchmarking-clinical-ner .claude/skills/benchmarking-clinical-ner && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
benchmarking-clinical-ner
GitHub stars
5.5k
Token cost
~1.7k tokens
SKILL.md length
577 words
Files
1
Skills in repo
74
Repo updated
First seen
Licence
Apache-2.0

At a glance

Score an OpenMed clinical or biomedical NER model against a user-supplied gold corpus with entity-level precision, recall, and F1, then break errors down per label.

  • Works in 6 steps: Align the corpus to OpenMed fixtures.… → Normalize labels to OpenMed's canonical… → Run run_suite / run_benchmark to get a… → …
  • The user wants a seqeval-style scorecard
  • SKILL.md covers When to use this skill, Match modes, Quick start and Workflow, plus 3 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Benchmarking Clinical Ner is an agent skill from maziyarpanahi/openmed. Score an OpenMed clinical or biomedical NER model against a user-supplied gold corpus with entity-level precision, recall, and F1, then break errors down per label. Use when the user wants a seqeval-style scorecard, strict vs partial (relaxed) span matching, a per-label confusion matrix, false-negative / false-positive examples, or to debug why a model misses entities. Trigger on "evaluate NER", "entity-level F1", "seqeval", "precision recall F1", "confusion matrix", "error analysis", "strict vs partial match"…

Its SKILL.md is about 1.7k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering Natural language processing. The repository describes itself as: Local-first healthcare AI: clinical NER and HIPAA PII de-identification on hardware you control. 2,200+ medical models, 35 model-backed PII languages, and Python, MLX, Android… The licence is Apache-2.0.

When your agent uses it

  • The user wants a seqeval-style scorecard
  • Strict vs partial (relaxed) span matching
  • A per-label confusion matrix
  • False-negative / false-positive examples

Example prompts

  • “evaluate NER”
  • “entity-level F1”
  • “seqeval”
  • “/benchmarking-clinical-ner”

Requirements

  • Python 3

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Align the corpus to OpenMed fixtures. Convert CoNLL/BIO or BRAT
  2. Normalize labels to OpenMed's canonical set so DRUG/MEDICATION variants
  3. Run run_suite / run_benchmark to get a BenchmarkReport.
  4. Read both F1s. A large strict↓ / relaxed↑ gap means boundary errors, not
  5. Run error_report for the per-label confusion matrix and capped
  6. Triage per label. Fix the worst-recall label first; in clinical NER a few

What it can do on your machine

Read from SKILL.md and the folder at commit 34d7b8c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • github.com
    • davidsbatista.net
    • aclanthology.org
    • brat.nlplab.org

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Benchmarking Clinical Ner loads about 1.7k tokens when it runs. Until then it costs about 166 tokens; SKILL.md has 577 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~166
When it runs · the whole SKILL.md, loaded when a task matches
~1.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from maziyarpanahi/openmed at commit 34d7b8c, republished under its Apache-2.0 licence (© maziyarpanahi). 577 words, ~1,666 tokens.

Download SKILL.mdSave it as .claude/skills/benchmarking-clinical-ner/SKILL.md (or your agent's skills folder).
name
benchmarking-clinical-ner
description
Score an OpenMed clinical or biomedical NER model against a user-supplied gold corpus with entity-level precision, recall, and F1, then break errors down per label. Use when the user wants a seqeval-style scorecard, strict vs partial (relaxed) span matching, a per-label confusion matrix, false-negative / false-positive examples, or to debug why a model misses entities. Trigger on "evaluate NER", "entity-level F1", "seqeval", "precision recall F1", "confusion matrix", "error analysis", "strict vs partial match", or "score against gold" in an OpenMed context. The gold corpus is user-supplied; OpenMed bundles no i2b2/n2c2/MIMIC data.
license
Apache-2.0
metadata.project
OpenMed
metadata.category
evaluation-quality
metadata.pairs
adjacent
metadata.version
1.0

Benchmarking Clinical NER

This skill produces an honest entity-level scorecard for an OpenMed NER model: precision / recall / F1 plus a per-label error breakdown. It scores spans, not tokens, because clinical entities are multi-token ("type 2 diabetes mellitus") and token-level accuracy hides boundary errors. Reported numbers are entity-level in the seqeval tradition (CoNLL-2000 / SemEval-2013 families).

When to use this skill

  • You have a gold-annotated clinical corpus and an OpenMed NER model to score.
  • You want strict (exact-boundary) and partial (relaxed-overlap) span F1.
  • You need per-label numbers, not one aggregate — DRUG recall ≠ DISEASE recall.
  • You need to explain the errors: what was missed, what was spurious, what was mislabeled.

For PHI de-id specifically, gate on leakage with evaluating-with-leakage-gates instead of (or in addition to) F1.

Match modes

ModeCounts a hit when…Use for
Strict / exactpredicted span boundaries and label match gold exactlyrelease scoring, boundary-sensitive tasks
Partial / relaxedpredicted span overlaps gold with the right labelrecall-oriented triage, tokenizer-mismatch tolerance

OpenMed exposes both: compute_exact_span_f1 (strict) and compute_relaxed_span_f1 (partial), with the full bundle in compute_metrics_bundle.

Quick start

Run a model over a user-supplied gold fixtures file and print a scorecard:

python
from openmed.eval import run_suite, error_report

# Fixtures: JSON list of {"id", "text", "gold_spans": [{start, end, label}, ...]}
report = run_suite(
    "eval/gold/clinical_ner.json",        # YOUR gold corpus, not bundled
    suite="golden",
    model_name="OpenMed/Disease-Detection",
    device="cpu",
)

m = report.metrics
print("exact F1 :", m["exact_span_f1"]["f1"])      # strict
print("relaxed F1:", m["relaxed_span_f1"]["f1"])    # partial
print("recall by label:", m["recall_slices"]["by_label"])

# Per-label confusion matrix + capped, no-PHI error examples.
errors = error_report(
    "OpenMed/Disease-Detection",
    "eval/gold/clinical_ner.json",
    suite_name="clinical_ner",
    example_cap=5,
)
print(errors.to_markdown())                 # confusion matrix + FN/FP tables
errors.write_json("eval/out/error_analysis.json")

Need just the metrics on spans you already have? Call the metric functions directly:

python
from openmed.eval import compute_exact_span_f1, compute_relaxed_span_f1

strict = compute_exact_span_f1(gold_spans, predicted_spans)
partial = compute_relaxed_span_f1(gold_spans, predicted_spans)

Workflow

  1. Align the corpus to OpenMed fixtures. Convert CoNLL/BIO or BRAT standoff into the fixture shape: text + gold_spans of {start, end, label} character offsets. (CoNLL → offsets; BRAT .ann is already character offsets.)
  2. Normalize labels to OpenMed's canonical set so DRUG/MEDICATION variants don't count as label confusion. Mislabeled-but-overlapping spans show up in the confusion matrix, not as misses.
  3. Run run_suite / run_benchmark to get a BenchmarkReport.
  4. Read both F1s. A large strict↓ / relaxed↑ gap means boundary errors, not detection failures — often tokenizer or whitespace issues.
  5. Run error_report for the per-label confusion matrix and capped examples. MISSED = false negatives (recall problem); SPURIOUS = false positives (precision problem); off-diagonal = label confusion.
  6. Triage per label. Fix the worst-recall label first; in clinical NER a few labels usually dominate the error budget.
Show full SKILL.md (243 more words)Show less

Hand-off to / from OpenMed

  • From extracting-clinical-entities (openmed.analyze_text): the model and predictions you score here come from the NER pipeline.
  • To evaluating-with-leakage-gates: for de-id models, F1 is necessary but not sufficient — pass the same fixtures through the release gates.
  • To authoring-model-cards: drop error_report confusion matrices and per-label F1 straight into the model card's quantitative-analysis section.
  • Pairs with building-gold-corpus (supplies the fixtures) and auditing-subgroup-fairness (slices the same run by demographic group).

Edge cases & gotchas

  • Token F1 lies; report span F1. Always use the span metrics (compute_exact_span_f1 / compute_relaxed_span_f1), not token accuracy.
  • Overlapping/nested gold spans need a documented matching rule. OpenMed's matcher picks the best single overlapping prediction per gold span; nested schemes (e.g. DISEASE inside ANATOMY) should be flattened or scored per layer.
  • Class imbalance hides failures. A macro view per label surfaces a rare-but- critical entity (e.g. ALLERGY) that micro-F1 buries.
  • Error examples are no-PHI by design. ErrorSpanExample stores offsets, context windows, and sha256: text hashes — never plaintext. Keep it that way.
  • Gold quality caps your ceiling. If inter-annotator agreement is low, a "low-F1" model may be right and the gold wrong. Spot-check disagreements before blaming the model.
  • No restricted corpora in the repo. i2b2/n2c2/MIMIC are DUA-gated: load them from the user's licensed copy at eval time; never commit them.

Standards & references

© maziyarpanahi, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/benchmarking-clinical-ner of maziyarpanahi/openmed.

Open the folder on GitHubat commit 34d7b8c

Compare with similar skills

Benchmarking Clinical Ner next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Benchmarking Clinical Ner compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Benchmarking Clinical Ner this skillmaziyarpanahi/openmed5.5k—~1.7kAutomated safety check: PassApache-2.0
Hugging Face TokenizersOrchestra-Research/AI-Research-SKILLs13k6 repos~3.4kAutomated safety check: PassMIT
Gptqmodel Tokenizer NormalizationModelCloud/GPTQModel1.3k—~1.1kAutomated safety check: PassCustom licence
Andrej KarpathyK-Dense-AI/mimeo282—~1.9kAutomated safety check: PassMIT
Comparetaishi-i/awesome-japanese-nlp-resources1k—~4.1kAutomated safety check: NotesCC0-1.0
Researchtaishi-i/awesome-japanese-nlp-resources1k—~3.5kAutomated safety check: NotesCC0-1.0

Similar skills

  • Hugging Face Tokenizers

    Orchestra-Research/AI-Research-SKILLs

    Shows how to load, train and use fast Hugging Face tokenizers, with BPE, WordPiece and Unigram models, padding, truncation and alignment tracking.

    13k GitHub starsUsed in 6 repos~3.4k tokens
    AI & LLM EngineeringAuto-check passed
  • Diagnose and correct GPT-QModel tokenizer initialization, tokenization normalization, special-token handling, prompt rendering, and chat-template problems.

    1.3k GitHub stars~1.1k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Andrej Karpathy

    K-Dense-AI/mimeo

    Applies the mental models and frameworks of Andrej Karpathy (deep learning, former Director of AI at Tesla, founding member of OpenAI, Eureka Labs).

    282 GitHub stars~1.9k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Compare

    taishi-i/awesome-japanese-nlp-resources

    Compare several Japanese NLP libraries, models, or datasets for a keyword (a specific tool name, or a function/task like '形態素解析') across a handful of criteria chosen for that comparison, rendered as…

    1k GitHub stars~4.1k tokensUpdated 4 days ago
    AI & LLM EngineeringAuto-check: notes
  • Research

    taishi-i/awesome-japanese-nlp-resources

    Analyze current trends and challenges in Japanese NLP for a topic.

    1k GitHub stars~3.5k tokensUpdated 4 days ago
    AI & LLM EngineeringAuto-check: notes
  • Sentence Transformers Embeddings

    Orchestra-Research/AI-Research-SKILLs

    Generates text embeddings locally with the sentence-transformers library for RAG, semantic search, clustering and similarity, with model picks for general, multilingual and legal text.

    13k GitHub starsUsed in 2 repos~1.6k tokens
    AI & LLM EngineeringAuto-check passed

More from maziyarpanahi/openmed

All 74 skills in this repo
  • Checks OpenMed de-identified clinical text against the 18 HIPAA Safe Harbor identifier categories and reports gaps and residual re-identification risk.

    5.5k GitHub stars~1.7k tokensUpdated today
    Auto-check passed
  • OpenMed Model Card Writer

    maziyarpanahi/openmed

    Fills in a model card for an OpenMed clinical NER or de-identification model from its evaluation reports: intended use, metrics, subgroups and limitations.

    5.5k GitHub stars~1.8k tokensUpdated today
    Auto-check passed
  • Walks a data pipeline against the HIPAA Privacy and Security Rule checklist and produces a gap report before it processes patient data.

    5.5k GitHub stars~2k tokensUpdated today
    Auto-check passed
  • ICD-10 Coding Assistant

    maziyarpanahi/openmed

    Suggests candidate ICD-10-CM diagnosis and ICD-10-PCS procedure codes for clinical text extracted by OpenMed, with rationale for a certified coder to review.

    5.5k GitHub stars~2k tokensUpdated today
    Auto-check passed
  • OpenMed ETL to OMOP CDM

    maziyarpanahi/openmed

    Maps OpenMed-extracted, terminology-coded conditions, drugs and measurements into OMOP CDM v5.4 tables for OHDSI and ATLAS analytics.

    5.5k GitHub stars~1.9k tokensUpdated today
    Auto-check passed
  • Extracting SDOH and Z-Codes

    maziyarpanahi/openmed

    Finds social risks such as housing instability or food insecurity in clinical notes and proposes matching ICD-10-CM Z-codes for a coder to confirm.

    5.5k GitHub stars~1.9k tokensUpdated today
    Auto-check passed

Questions about Benchmarking Clinical Ner

What does Benchmarking Clinical Ner do?

Score an OpenMed clinical or biomedical NER model against a user-supplied gold corpus with entity-level precision, recall, and F1, then break errors down per label. Benchmarking Clinical Ner is an agent skill from maziyarpanahi/openmed. Score an OpenMed clinical or biomedical NER model against a user-supplied gold corpus with entity-level precision, recall, and F1, then break errors down per label.

When should I use Benchmarking Clinical Ner?

Benchmarking Clinical Ner fits situations like: the user wants a seqeval-style scorecard; strict vs partial (relaxed) span matching; A per-label confusion matrix; false-negative / false-positive examples.

How do I install Benchmarking Clinical Ner in Claude Code?

Run `npx skills add maziyarpanahi/openmed --skill benchmarking-clinical-ner -a claude-code`. Or copy the skill folder (skills/benchmarking-clinical-ner in maziyarpanahi/openmed) into .claude/skills/benchmarking-clinical-ner in your project. Claude Code loads it when a task matches its description.

How do I install Benchmarking Clinical Ner in Codex?

Run `npx skills add maziyarpanahi/openmed --skill benchmarking-clinical-ner -a codex`. Or copy the skill folder (skills/benchmarking-clinical-ner in maziyarpanahi/openmed) into .agents/skills/benchmarking-clinical-ner in your project. Codex loads it when a task matches its description.

Can I use Benchmarking Clinical Ner in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add maziyarpanahi/openmed --skill benchmarking-clinical-ner -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/benchmarking-clinical-ner, .gemini/skills/benchmarking-clinical-ner, .github/skills/benchmarking-clinical-ner and .opencode/skills/benchmarking-clinical-ner in your project.

What does Benchmarking Clinical Ner need to run?

SKILL.md names no scripts, command-line tools or credentials: Benchmarking Clinical Ner is instructions for the agent only. Our summary lists: Python 3.

Does Benchmarking Clinical Ner access the network?

SKILL.md names 4 domains. As links in the text: github.com, davidsbatista.net, aclanthology.org and brat.nlplab.org. This is read from the text; nothing was executed.

Is Benchmarking Clinical Ner safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Benchmarking Clinical Ner use?

Benchmarking Clinical Ner is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Benchmarking Clinical Ner use?

About 1.7k tokens (SKILL.md is roughly 6.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Benchmarking Clinical Ner?

Skills that share tags, products or a category with Benchmarking Clinical Ner: Hugging Face Tokenizers (Orchestra-Research/AI-Research-SKILLs, 13k stars), Gptqmodel Tokenizer Normalization (ModelCloud/GPTQModel, 1.3k stars), Andrej Karpathy (K-Dense-AI/mimeo, 282 stars) and Compare (taishi-i/awesome-japanese-nlp-resources, 1k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Benchmarking Clinical Ner?

maziyarpanahi (a GitHub user) maintains it in maziyarpanahi/openmed, which has 5,506 GitHub stars. The repository holds 74 skills in this directory. The repository was last updated on October 11, 2026.

Source: maziyarpanahi/openmed on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.