Agent skill

Building Gold Corpus

by maziyarpanahi in maziyarpanahi/openmed

Scaffold a synthetic gold-standard annotation project for evaluating OpenMed NER and de-identification models — label schema, annotation guidelines, BRAT or Label Studio config, and disjoint…

Apache-2.0Auto-check passedAI & LLM Engineering

Install Building Gold Corpus

skills CLI
$ npx skills add maziyarpanahi/openmed --skill building-gold-corpus -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install maziyarpanahi/openmed building-gold-corpus --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/maziyarpanahi/openmed.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/building-gold-corpus .claude/skills/building-gold-corpus && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
building-gold-corpus
GitHub stars
5.5k
Token cost
~1.7k tokens
SKILL.md length
596 words
Files
1
Skills in repo
74
Repo updated
First seen
Licence
Apache-2.0

At a glance

Scaffold a synthetic gold-standard annotation project for evaluating OpenMed NER and de-identification models — label schema, annotation guidelines, BRAT or Label Studio config, and disjoint…

  • Works in 7 steps: Define the label schema. Reuse OpenMed… → Write annotation guidelines. The manual… → Generate synthetic source text. Compose… → …
  • The user wants to create eval fixtures
  • SKILL.md covers When to use this skill, The OpenMed fixture shape…, Quick start — scaffold the… and Workflow, plus 3 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Building Gold Corpus is an agent skill from maziyarpanahi/openmed. Scaffold a synthetic gold-standard annotation project for evaluating OpenMed NER and de-identification models — label schema, annotation guidelines, BRAT or Label Studio config, and disjoint train/dev/test splits. Use when the user wants to create eval fixtures, set up annotation, define a label set, write guidelines, configure an annotation tool, or build a held-out gold set for the OpenMed eval harness. Trigger on "gold corpus", "annotation project", "label schema", "annotation guidelines", "BRAT", "Label…

Its SKILL.md is about 1.7k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering LLM evaluation. The repository describes itself as: Local-first healthcare AI: clinical NER and HIPAA PII de-identification on hardware you control. 2,200+ medical models, 35 model-backed PII languages, and Python, MLX, Android… The licence is Apache-2.0.

When your agent uses it

  • The user wants to create eval fixtures
  • Set up annotation
  • Define a label set
  • Write guidelines

Example prompts

  • “gold corpus”
  • “annotation project”
  • “label schema”
  • “/building-gold-corpus”

Requirements

  • Python 3

Workflow steps

7 steps, taken from the first numbered list in SKILL.md.

  1. Define the label schema. Reuse OpenMed canonical labels (PERSON, DATE,
  2. Write annotation guidelines. The manual is the contract: span boundaries,
  3. Generate synthetic source text. Compose realistic clinical narratives
  4. Configure the tool. BRAT uses annotation.conf (entity types) producing
  5. Double-annotate and measure agreement. Have ≥2 annotators on an overlap
  6. Split with discipline. Partition by document/patient, not by sentence,
  7. Validate and commit. Run load_fixtures; confirm spans align and ids are

What it can do on your machine

Read from SKILL.md and the folder at commit 34d7b8c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are json and python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • brat.nlplab.org
    • labelstud.io
    • i2b2.org
    • n2c2.dbmi.hms.harvard.edu
    • aclanthology.org

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Building Gold Corpus loads about 1.7k tokens when it runs. Until then it costs about 176 tokens; SKILL.md has 596 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~176
When it runs · the whole SKILL.md, loaded when a task matches
~1.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from maziyarpanahi/openmed at commit 34d7b8c, republished under its Apache-2.0 licence (© maziyarpanahi). 596 words, ~1,689 tokens.

Download SKILL.mdSave it as .claude/skills/building-gold-corpus/SKILL.md (or your agent's skills folder).
name
building-gold-corpus
description
Scaffold a synthetic gold-standard annotation project for evaluating OpenMed NER and de-identification models — label schema, annotation guidelines, BRAT or Label Studio config, and disjoint train/dev/test splits. Use when the user wants to create eval fixtures, set up annotation, define a label set, write guidelines, configure an annotation tool, or build a held-out gold set for the OpenMed eval harness. Trigger on "gold corpus", "annotation project", "label schema", "annotation guidelines", "BRAT", "Label Studio", "train dev test split", or "build eval fixtures" for OpenMed. Committed gold must be synthetic; licensed (i2b2/n2c2/MIMIC) data is eval-only and never committed.
license
Apache-2.0
metadata.project
OpenMed
metadata.category
evaluation-quality
metadata.pairs
adjacent
metadata.version
1.0

Building a Gold Corpus

You can't evaluate what you can't measure against. This skill scaffolds a gold-standard annotation project whose output drops straight into the OpenMed eval harness as fixtures. The hard rule: anything committed to the repo is synthetic. Licensed clinical corpora (i2b2, n2c2, MIMIC) are DUA-gated — use them at eval time from the user's own copy, never check them in.

When to use this skill

  • You need eval fixtures for benchmarking-clinical-ner or evaluating-with-leakage-gates and have none.
  • You're standing up an annotation effort: schema, guidelines, tool config.
  • You need disciplined train/dev/test splits with no leakage between them.
  • You want a small synthetic golden set you can commit and gate on in CI.

The OpenMed fixture shape (your target output)

Annotations must serialize to character-offset spans the harness understands:

json
{
  "fixtures": [
    {
      "id": "synthetic-0001",
      "language": "en",
      "text": "Ms. Jane Roe (MRN 0000000) seen 2099-01-02 for type 2 diabetes.",
      "gold_spans": [
        {"start": 4,  "end": 12, "label": "PERSON"},
        {"start": 18, "end": 25, "label": "ID_NUM"},
        {"start": 32, "end": 42, "label": "DATE"},
        {"start": 47, "end": 62, "label": "DISEASE"}
      ]
    }
  ]
}

openmed.eval.harness.load_fixtures accepts a top-level list or a {"fixtures": [...]} mapping. Offsets are character indices into text; labels are OpenMed-canonical.

Quick start — scaffold the project

eval/
  gold/
    guidelines.md            # annotation manual + edge-case decisions
    label_schema.json        # canonical labels + definitions + examples
    synthetic/               # COMMITTED synthetic fixtures (CI-gateable)
      train.json
      dev.json
      test.json
  external/                  # GITIGNORED: licensed DUA corpora, eval-only
    .gitignore               # *  (never commit i2b2/n2c2/MIMIC)

Verify your synthetic fixtures load and validate spans before you trust them:

python
from openmed.eval.harness import load_fixtures

fixtures = load_fixtures("eval/gold/synthetic/test.json")
print(len(fixtures), "fixtures;", sum(len(f.gold_spans) for f in fixtures), "spans")
# load_fixtures normalizes spans against source text and rejects duplicate ids.

Workflow

  1. Define the label schema. Reuse OpenMed canonical labels (PERSON, DATE, ID_NUM, EMAIL, PHONE, DISEASE, DRUG, ...). Each label gets a one-line definition, in/out examples, and a boundary rule (include titles? trailing punctuation?).
  2. Write annotation guidelines. The manual is the contract: span boundaries, nested/overlapping policy, ambiguous cases, and a decision log appended as real cases force calls. Vague guidelines → low agreement → unusable gold.
  3. Generate synthetic source text. Compose realistic clinical narratives with fabricated identifiers (Faker-style names, impossible dates like 2099-, all-zero MRNs). Never paste real notes into committed data.
  4. Configure the tool. BRAT uses annotation.conf (entity types) producing .ann standoff; Label Studio uses a labeling-config XML producing JSON. Map either back to the fixture shape above.
  5. Double-annotate and measure agreement. Have ≥2 annotators on an overlap set; compute span-level inter-annotator agreement (F1 or Cohen's κ). Adjudicate disagreements and fold the resolutions into the decision log.
  6. Split with discipline. Partition by document/patient, not by sentence, so no patient appears in two splits. Freeze test; never tune on it.
  7. Validate and commit. Run load_fixtures; confirm spans align and ids are unique. Commit only the synthetic splits.
Show full SKILL.md (242 more words)Show less

Hand-off to / from OpenMed

  • To benchmarking-clinical-ner: dev/test fixtures feed run_suite and error_report for the NER scorecard.
  • To evaluating-with-leakage-gates and gating-deid-leakage: the synthetic held-out set is exactly what the release gates and the CI gate run against.
  • To building-with-openmed: synthetic notes can be generated by running surrogate replacement through openmed.deidentify(method="replace").
  • Pairs with auditing-subgroup-fairness: tag each gold span with a group in metadata so fairness_report can slice by demographic surrogate.

Edge cases & gotchas

  • Committed = synthetic. No exceptions. Real PHI in the repo is a breach even if the repo is private. Generate identifiers; don't transcribe them.
  • DUA data is eval-only. Load i2b2/n2c2/MIMIC from eval/external/ (gitignored) at runtime under the user's license; results may be reported, data never shared.
  • Split by patient, not by line. Sentence-level splitting leaks a patient's style/identifiers across train and test and inflates scores.
  • Offsets must be character indices into this text. Re-tokenization or whitespace edits silently shift offsets; re-validate with load_fixtures.
  • Label the fairness surrogate, not real demographics. Put a synthetic group tag in span metadata; don't store real protected attributes.
  • Decision log is the gold's source of truth. Without it, two re-annotations disagree and your "ceiling" F1 is noise.

Standards & references

© maziyarpanahi, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/building-gold-corpus of maziyarpanahi/openmed.

Open the folder on GitHubat commit 34d7b8c

Compare with similar skills

Building Gold Corpus next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Building Gold Corpus compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Building Gold Corpus this skillmaziyarpanahi/openmed5.5k—~1.7kAutomated safety check: PassApache-2.0
LLM Benchmarking with lm-evaluation-harnessOrchestra-Research/AI-Research-SKILLs13k8 repos~3kAutomated safety check: PassMIT
Hugging Face Local Model Evalshuggingface/skills11k2 repos~1.6kAutomated safety check: PassApache-2.0
Looperksimback/looper710—~2.7kAutomated safety check: NotesMIT
Agent Eval Engineeringlangchain-ai/langchain-skills1.3k—~4kAutomated safety check: PassMIT
Quality FlywheelGoogleCloudPlatform/vertex-ai-samples792—~2kAutomated safety check: PassApache-2.0

Similar skills

  • LLM Benchmarking with lm-evaluation-harness

    Orchestra-Research/AI-Research-SKILLs

    Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.

    13k GitHub starsUsed in 8 repos~3k tokens
    AI & LLM EngineeringAuto-check passed
  • Official

    Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.

    11k GitHub starsUsed in 2 repos~1.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Looper

    ksimback/looper

    Scaffold a well-designed agent loop with best-practice coaching and a cross-model review council.

    710 GitHub stars~2.7k tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check: notes
  • Agent Eval Engineering

    langchain-ai/langchain-skills

    Official

    Builds agent evaluations in stages: inspect the repository and traces, agree a Task Spec with you, then build, audit and run a Harbor task with an independent verifier.

    1.3k GitHub stars~4k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • Quality Flywheel

    GoogleCloudPlatform/vertex-ai-samples

    Evaluate and improve GenAI models and agents using the Google GenAI Evaluation SDK.

    792 GitHub stars~2k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Eval Harness

    cloudnative-co/claude-code-starter-kit

    Formal evaluation framework for Claude Code sessions implementing eval-driven development (EDD) principles.

    153 GitHub starsUsed in 9 repos~1.3k tokens
    AI & LLM EngineeringAuto-check passed

More from maziyarpanahi/openmed

All 74 skills in this repo
  • Checks OpenMed de-identified clinical text against the 18 HIPAA Safe Harbor identifier categories and reports gaps and residual re-identification risk.

    5.5k GitHub stars~1.7k tokensUpdated today
    Auto-check passed
  • OpenMed Model Card Writer

    maziyarpanahi/openmed

    Fills in a model card for an OpenMed clinical NER or de-identification model from its evaluation reports: intended use, metrics, subgroups and limitations.

    5.5k GitHub stars~1.8k tokensUpdated today
    Auto-check passed
  • Walks a data pipeline against the HIPAA Privacy and Security Rule checklist and produces a gap report before it processes patient data.

    5.5k GitHub stars~2k tokensUpdated today
    Auto-check passed
  • ICD-10 Coding Assistant

    maziyarpanahi/openmed

    Suggests candidate ICD-10-CM diagnosis and ICD-10-PCS procedure codes for clinical text extracted by OpenMed, with rationale for a certified coder to review.

    5.5k GitHub stars~2k tokensUpdated today
    Auto-check passed
  • OpenMed ETL to OMOP CDM

    maziyarpanahi/openmed

    Maps OpenMed-extracted, terminology-coded conditions, drugs and measurements into OMOP CDM v5.4 tables for OHDSI and ATLAS analytics.

    5.5k GitHub stars~1.9k tokensUpdated today
    Auto-check passed
  • Extracting SDOH and Z-Codes

    maziyarpanahi/openmed

    Finds social risks such as housing instability or food insecurity in clinical notes and proposes matching ICD-10-CM Z-codes for a coder to confirm.

    5.5k GitHub stars~1.9k tokensUpdated today
    Auto-check passed

Questions about Building Gold Corpus

What does Building Gold Corpus do?

Scaffold a synthetic gold-standard annotation project for evaluating OpenMed NER and de-identification models — label schema, annotation guidelines, BRAT or Label Studio config, and disjoint…. Building Gold Corpus is an agent skill from maziyarpanahi/openmed. Scaffold a synthetic gold-standard annotation project for evaluating OpenMed NER and de-identification models — label schema, annotation guidelines, BRAT or Label Studio config, and disjoint train/dev/test splits.

When should I use Building Gold Corpus?

Building Gold Corpus fits situations like: the user wants to create eval fixtures; set up annotation; define a label set; write guidelines.

How do I install Building Gold Corpus in Claude Code?

Run `npx skills add maziyarpanahi/openmed --skill building-gold-corpus -a claude-code`. Or copy the skill folder (skills/building-gold-corpus in maziyarpanahi/openmed) into .claude/skills/building-gold-corpus in your project. Claude Code loads it when a task matches its description.

How do I install Building Gold Corpus in Codex?

Run `npx skills add maziyarpanahi/openmed --skill building-gold-corpus -a codex`. Or copy the skill folder (skills/building-gold-corpus in maziyarpanahi/openmed) into .agents/skills/building-gold-corpus in your project. Codex loads it when a task matches its description.

Can I use Building Gold Corpus in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add maziyarpanahi/openmed --skill building-gold-corpus -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/building-gold-corpus, .gemini/skills/building-gold-corpus, .github/skills/building-gold-corpus and .opencode/skills/building-gold-corpus in your project.

What does Building Gold Corpus need to run?

SKILL.md names no scripts, command-line tools or credentials: Building Gold Corpus is instructions for the agent only. Our summary lists: Python 3.

Does Building Gold Corpus access the network?

SKILL.md names 5 domains. As links in the text: brat.nlplab.org, labelstud.io, i2b2.org, n2c2.dbmi.hms.harvard.edu and aclanthology.org. This is read from the text; nothing was executed.

Is Building Gold Corpus safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Building Gold Corpus use?

Building Gold Corpus is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Building Gold Corpus use?

About 1.7k tokens (SKILL.md is roughly 6.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Building Gold Corpus?

Skills that share tags, products or a category with Building Gold Corpus: LLM Benchmarking with lm-evaluation-harness (Orchestra-Research/AI-Research-SKILLs, 13k stars), Hugging Face Local Model Evals (huggingface/skills, 11k stars), Looper (ksimback/looper, 710 stars) and Agent Eval Engineering (langchain-ai/langchain-skills, 1.3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Building Gold Corpus?

maziyarpanahi (a GitHub user) maintains it in maziyarpanahi/openmed, which has 5,506 GitHub stars. The repository holds 74 skills in this directory. The repository was last updated on October 11, 2026.

Source: maziyarpanahi/openmed on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.