Agent skill

OpenMed ETL to OMOP CDM

by maziyarpanahi in maziyarpanahi/openmed

Maps OpenMed-extracted, terminology-coded conditions, drugs and measurements into OMOP CDM v5.4 tables for OHDSI and ATLAS analytics.

Apache-2.0Auto-check passedData & Analytics

Install OpenMed ETL to OMOP CDM

skills CLI
$ npx skills add maziyarpanahi/openmed --skill etl-to-omop-cdm -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install maziyarpanahi/openmed etl-to-omop-cdm --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/maziyarpanahi/openmed.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/etl-to-omop-cdm .claude/skills/etl-to-omop-cdm && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
etl-to-omop-cdm
GitHub stars
5.5k
Token cost
~1.9k tokens
SKILL.md length
709 words
Files
2 (incl. references)
Skills in repo
74
Repo updated
First seen
Licence
Apache-2.0

At a glance

Maps OpenMed-extracted, terminology-coded conditions, drugs and measurements into OMOP CDM v5.4 tables for OHDSI and ATLAS analytics.

  • Works in 2 steps: *_source_concept_id — the OHDSI CONCEPT… → *_concept_id — the standard concept,…
  • Loading NLP-derived clinical facts into an OMOP database
  • SKILL.md covers When to use this skill, Quick start, The source → standard pattern… and Domain → table → fields, plus 5 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

The skill is a deterministic downstream step. After OpenMed's `analyze_text` has extracted entities and you have linked them to source vocabularies (ICD-10-CM or SNOMED for conditions, RxNorm for drugs, LOINC for labs), it turns them into rows for `condition_occurrence`, `drug_exposure` and `measurement`. It is not a clinical NER or code-linking skill and assumes both steps are done.

Every event row carries a source concept id for the original code and a standard concept id found by following the 'Maps to' relationship in `CONCEPT_RELATIONSHIP`, standardizing conditions to SNOMED, drugs to RxNorm and measurements to LOINC, with 0 when no mapping exists. Rows also need type concepts that mark them as NLP-derived, and a reference file lists the CDM fields per table. You supply the OHDSI vocabulary bundle from Athena and do the lookups under your own license, since OpenMed ships no UMLS, SNOMED, RxNorm or LOINC content.

When your agent uses it

  • Loading NLP-derived clinical facts into an OMOP database
  • Building an OHDSI ETL from clinical notes
  • Standardizing note-derived findings to OMOP standard concepts

Example prompts

  • “Turn the coded OpenMed output for this note into condition_occurrence and drug_exposure rows.”
  • “Map our RxNorm-linked medications to standard concepts and fill drug_exposure.”
  • “Which required OMOP fields am I missing for measurement rows built from lab text?”

Requirements

  • OpenMed `analyze_text` output already linked to SNOMED, RxNorm or LOINC
  • An OHDSI vocabulary bundle (Athena) with CONCEPT and CONCEPT_RELATIONSHIP

Workflow steps

2 steps, taken from the first numbered list in SKILL.md.

  1. *_source_concept_id — the OHDSI CONCEPT for your original code (e.g. the
  2. *_concept_id — the standard concept, obtained by following

What it can do on your machine

Read from SKILL.md and the folder at commit 34d7b8c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • ohdsi.github.io
    • athena.ohdsi.org

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

OpenMed ETL to OMOP CDM loads about 1.9k tokens when it runs, and up to ~3.1k if it reads all its reference files. Until then it costs about 187 tokens; SKILL.md has 709 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~187
When it runs · the whole SKILL.md, loaded when a task matches
~1.9k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~3.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from maziyarpanahi/openmed at commit 34d7b8c, republished under its Apache-2.0 licence (© maziyarpanahi). 709 words, ~1,895 tokens.

Download SKILL.mdSave it as .claude/skills/etl-to-omop-cdm/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
etl-to-omop-cdm
description
Map OpenMed-extracted, terminology-coded conditions, drugs, and measurements into OMOP CDM v5.4 clinical tables (condition_occurrence, drug_exposure, measurement) for OHDSI/ATLAS analytics. Use when the user wants to load NLP-derived facts into an OMOP database, build an OHDSI ETL from clinical notes, populate condition_occurrence or drug_exposure from text, or standardize note-derived findings to OMOP standard concepts. Covers the source-to-standard concept mapping pattern, required vs optional CDM fields, type concepts for NLP-derived rows, and the user-supplied OHDSI vocabulary (CONCEPT/CONCEPT_RELATIONSHIP). Consumes coded OpenMed analyze_text output (after SNOMED/RxNorm/LOINC linking) and produces OMOP-conformant rows.
license
Apache-2.0
metadata.project
OpenMed
metadata.category
analytics-reporting
metadata.pairs
after
metadata.version
1.0

ETL to OMOP CDM

The OMOP Common Data Model (CDM) is the OHDSI standard for observational health data. This skill maps OpenMed-derived clinical facts — entities from analyze_text that you have already linked to a source terminology — into the OMOP clinical event tables condition_occurrence, drug_exposure, and measurement. The NLP runs on-device; OMOP loading is a downstream, deterministic transform.

When to use this skill

After you have (a) extracted entities with OpenMed and (b) coded them to a source vocabulary (ICD-10-CM / SNOMED for conditions, RxNorm for drugs, LOINC for labs — see the linking skills). Use this skill to turn those coded facts into OMOP rows. It is not a clinical NER skill and not a code-linking skill; it assumes both are done.

Quick start

python
import openmed

note = "Assessment: type 2 diabetes mellitus. Started metformin 500 mg PO BID. HbA1c 8.2%."
result = openmed.analyze_text(note, output_format="dict")
# result["entities"] -> [{text,label,confidence,start,end}, ...]

# You then code each entity to a SOURCE concept using the OHDSI vocabulary you
# downloaded (see linking-umls-concepts / normalizing-rxnorm / mapping-loinc),
# and map SOURCE -> STANDARD via CONCEPT_RELATIONSHIP ('Maps to').
fact = {
    "person_id": 1001,
    "domain": "Condition",
    "source_code": "E11.9",          # ICD-10-CM, from your coding step
    "source_vocabulary": "ICD10CM",
    "source_concept_id": 45533010,    # OHDSI CONCEPT for E11.9 (lookup)
    "standard_concept_id": 201826,    # 'Maps to' -> SNOMED 'Type 2 diabetes mellitus'
    "start_date": "2024-03-12",       # from building-patient-timelines
    "char_span": (fact_start, fact_end),
}

OpenMed never ships UMLS/SNOMED/RxNorm/LOINC content. You supply the OHDSI vocabulary bundle (Athena download) and do the lookups under your own license. OpenMed provides the spans and labels.

The source → standard pattern (the heart of OMOP)

Every clinical event row carries two concept ids:

  1. *_source_concept_id — the OHDSI CONCEPT for your original code (e.g. the ICD-10-CM or RxNorm code your linking step produced).
  2. *_concept_id — the standard concept, obtained by following CONCEPT_RELATIONSHIP.relationship_id = 'Maps to' from the source concept. Conditions standardize to SNOMED, drugs to RxNorm, measurements to LOINC. If no mapping exists, set the standard id to 0.

Domain → table → fields

See references/omop_cdm_v5_4_fields.md for the full per-table field list. Core mapping by OpenMed entity domain:

OpenMed entity domainOMOP tableStandard vocabKey date / value fields
Disease / Conditioncondition_occurrenceSNOMEDcondition_start_date, optional condition_end_date
Drug / Medicationdrug_exposureRxNormdrug_exposure_start_date, drug_exposure_end_date, quantity, sig
Lab / MeasurementmeasurementLOINCmeasurement_date, value_as_number, unit_concept_id, value_as_concept_id

Type concepts: mark rows as NLP-derived

Every event row needs a *_type_concept_id recording provenance. For facts derived from clinical text, OHDSI uses the type concept 32831 "EHR episode record" / "Note" family — specifically prefer a "...from note" / "NLP"-flavored standard type concept from the Type Concept vocabulary in your bundle. Do not invent ids; resolve the type concept against the vocabulary you loaded so cohort builders can filter NLP-derived rows.

Workflow

  1. Extract & code. analyze_text → entities; link each to a source code (linking skills). Resolve source_concept_id and the 'Maps to' standard concept from your Athena vocabulary.
  2. Resolve dates. Attach start (and end, where known) dates per building-patient-timelines. OMOP date fields are DATE; keep the matching *_datetime only if you truly have a time.
  3. Assign person_id. Join to your person table by an internal key — not by any PHI string. De-identify upstream.
  4. Build rows with required keys, the source+standard concept pair, the NLP type concept, and a unique surrogate *_occurrence_id / *_exposure_id / measurement_id.
  5. Stage *_source_value (the raw surface string, after de-id) for QA traceability — but never put raw PHI there.
  6. Conform & validate. Run OHDSI Achilles/DataQualityDashboard on the loaded CDM before analytics.
Show full SKILL.md (247 more words)Show less

Hand-off to / from OpenMed

  • From OpenMed: analyze_text entities (offsets + labels), deidentify upstream, and the per-domain linking skills (linking-umls-concepts, normalizing-rxnorm, mapping-loinc, mapping-to-snomed, coding-icd10).
  • To OHDSI: loaded condition_occurrence / drug_exposure / measurement rows are consumed by ATLAS, Achilles, and cohort definitions — and by computing-ecqms for measure denominators/numerators.

Edge cases & gotchas

  • No license bundling. OpenMed does not include SNOMED/RxNorm/LOINC/UMLS. Download the OHDSI vocabularies from Athena and run lookups under your own agreement. This is a hard rule.
  • Unmapped → standard concept 0. When 'Maps to' yields nothing, set *_concept_id = 0 and keep the source ids. Never fabricate a standard id.
  • Domain routing follows the standard concept's domain, not the source code's apparent type. An ICD-10 code can map to a SNOMED concept whose domain_id is Observation or Measurement — load it into the table the standard concept dictates.
  • NLP rows are lower-assurance. Tag them with the NLP/from note type concept and carry confidence (e.g. in a companion table) so analysts can threshold. Don't silently mix them with structured EHR rows.
  • Dates are required and must be valid. Undated note facts can't populate a *_start_date; route them to your "needs review" staging, not into the CDM with a placeholder date.
  • measurement units and values. Parse value_as_number + unit (mapped to a unit_concept_id); for qualitative results use value_as_concept_id.

Standards & references

© maziyarpanahi, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (references) in skills/etl-to-omop-cdm of maziyarpanahi/openmed.

  • SKILL.md
  • references/omop_cdm_v5_4_fields.md

Open the folder on GitHubat commit 34d7b8c

Compare with similar skills

OpenMed ETL to OMOP CDM next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

OpenMed ETL to OMOP CDM compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
OpenMed ETL to OMOP CDM this skillmaziyarpanahi/openmed5.5k—~1.9kAutomated safety check: PassApache-2.0
Authoritative Data Harvesteryushui2022/MathModel-Skill4541 repos~1.1kAutomated safety check: PassMIT
openFDA Regulatory Data Queriesdavila7/claude-code-templates33k11 repos~3.6kAutomated safety check: PassMIT
Experiment AgentImbad0202/experiment-agent199—~3.1kAutomated safety check: PassCC-BY-NC-4.0
RoundingRConsortium/pharma-skills120—~3.8kAutomated safety check: PassMIT
Multimodal Media Feature ExtractionTyrealQ/q-skills108—~2kAutomated safety check: NotesMIT

Similar skills

  • Authoritative Data Harvester

    yushui2022/MathModel-Skill

    Finds authoritative public data sources for modeling tasks, prefers official APIs and bulk downloads, and outputs a reproducible fetch and cleaning plan with citations.

    454 GitHub starsUsed in 1 repo~1.1k tokens
    Data & AnalyticsAuto-check passed
  • openFDA Regulatory Data Queries

    davila7/claude-code-templates

    Queries the openFDA API from Python for drug, device, food and veterinary data: adverse events, recalls, labels, approvals, NDC and UNII lookups.

    33k GitHub starsUsed in 11 repos~3.6k tokens
    Data & AnalyticsAuto-check passed
  • Experiment Agent

    Imbad0202/experiment-agent

    Experiment executor and monitor for academic research. An agent skill from Imbad0202/experiment-agent.

    199 GitHub stars~3.1k tokensUpdated 16 days ago
    Data & AnalyticsAuto-check passed
  • Rounding

    RConsortium/pharma-skills

    Audit R code that prepares CSR/TLF statistics for SAS-compatible rounding compliance (ties away from zero, round-once-at-display, fixed trailing-zero precision).

    120 GitHub stars~3.8k tokensUpdated 7 days ago
    Data & AnalyticsAuto-check passed
  • Extracts pixel, video-frame, speech, music and visual-semantic features from image, video and audio files for research datasets, using local tools or the Gemini API.

    108 GitHub stars~2k tokensUpdated 17 days ago
    Data & AnalyticsAuto-check: notes
  • PyHealth Clinical ML Toolkit

    davila7/claude-code-templates

    Builds machine learning pipelines on clinical data with PyHealth: EHR datasets, prediction tasks, medical code mapping, healthcare models and evaluation.

    33k GitHub starsUsed in 11 repos~4.4k tokens
    Research & ScienceAuto-check passed

More from maziyarpanahi/openmed

All 74 skills in this repo
  • Checks OpenMed de-identified clinical text against the 18 HIPAA Safe Harbor identifier categories and reports gaps and residual re-identification risk.

    5.5k GitHub stars~1.7k tokensUpdated today
    Auto-check passed
  • OpenMed Model Card Writer

    maziyarpanahi/openmed

    Fills in a model card for an OpenMed clinical NER or de-identification model from its evaluation reports: intended use, metrics, subgroups and limitations.

    5.5k GitHub stars~1.8k tokensUpdated today
    Auto-check passed
  • Walks a data pipeline against the HIPAA Privacy and Security Rule checklist and produces a gap report before it processes patient data.

    5.5k GitHub stars~2k tokensUpdated today
    Auto-check passed
  • ICD-10 Coding Assistant

    maziyarpanahi/openmed

    Suggests candidate ICD-10-CM diagnosis and ICD-10-PCS procedure codes for clinical text extracted by OpenMed, with rationale for a certified coder to review.

    5.5k GitHub stars~2k tokensUpdated today
    Auto-check passed
  • Extracting SDOH and Z-Codes

    maziyarpanahi/openmed

    Finds social risks such as housing instability or food insecurity in clinical notes and proposes matching ICD-10-CM Z-codes for a coder to confirm.

    5.5k GitHub stars~1.9k tokensUpdated today
    Auto-check passed
  • Clinical Document Ingestion

    maziyarpanahi/openmed

    Converts scanned faxes, images, CSV/TSV exports and C-CDA XML into clean text on-device, ready for OpenMed de-identification and named-entity recognition.

    5.5k GitHub stars~2k tokensUpdated today
    Auto-check passed

Questions about OpenMed ETL to OMOP CDM

What does OpenMed ETL to OMOP CDM do?

Maps OpenMed-extracted, terminology-coded conditions, drugs and measurements into OMOP CDM v5.4 tables for OHDSI and ATLAS analytics. The skill is a deterministic downstream step. After OpenMed's `analyze_text` has extracted entities and you have linked them to source vocabularies (ICD-10-CM or SNOMED for conditions, RxNorm for drugs, LOINC for labs), it turns them into rows for `condition_occurrence`, `drug_exposure` and `measurement`.

When should I use OpenMed ETL to OMOP CDM?

OpenMed ETL to OMOP CDM fits situations like: loading NLP-derived clinical facts into an OMOP database; building an OHDSI ETL from clinical notes; standardizing note-derived findings to OMOP standard concepts.

How do I install OpenMed ETL to OMOP CDM in Claude Code?

Run `npx skills add maziyarpanahi/openmed --skill etl-to-omop-cdm -a claude-code`. Or copy the skill folder (skills/etl-to-omop-cdm in maziyarpanahi/openmed) into .claude/skills/etl-to-omop-cdm in your project. Claude Code loads it when a task matches its description.

How do I install OpenMed ETL to OMOP CDM in Codex?

Run `npx skills add maziyarpanahi/openmed --skill etl-to-omop-cdm -a codex`. Or copy the skill folder (skills/etl-to-omop-cdm in maziyarpanahi/openmed) into .agents/skills/etl-to-omop-cdm in your project. Codex loads it when a task matches its description.

Can I use OpenMed ETL to OMOP CDM in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add maziyarpanahi/openmed --skill etl-to-omop-cdm -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/etl-to-omop-cdm, .gemini/skills/etl-to-omop-cdm, .github/skills/etl-to-omop-cdm and .opencode/skills/etl-to-omop-cdm in your project.

What does OpenMed ETL to OMOP CDM need to run?

SKILL.md names no scripts, command-line tools or credentials: OpenMed ETL to OMOP CDM is instructions for the agent only. Our summary lists: OpenMed `analyze_text` output already linked to SNOMED, RxNorm or LOINC; An OHDSI vocabulary bundle (Athena) with CONCEPT and CONCEPT_RELATIONSHIP.

Does OpenMed ETL to OMOP CDM access the network?

SKILL.md names 2 domains. As links in the text: ohdsi.github.io and athena.ohdsi.org. This is read from the text; nothing was executed.

Is OpenMed ETL to OMOP CDM safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does OpenMed ETL to OMOP CDM use?

OpenMed ETL to OMOP CDM is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does OpenMed ETL to OMOP CDM use?

About 1.9k tokens (SKILL.md is roughly 7.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.2k tokens, read only when the agent opens those files.

What are the alternatives to OpenMed ETL to OMOP CDM?

Skills that share tags, products or a category with OpenMed ETL to OMOP CDM: Authoritative Data Harvester (yushui2022/MathModel-Skill, 454 stars), openFDA Regulatory Data Queries (davila7/claude-code-templates, 33k stars), Experiment Agent (Imbad0202/experiment-agent, 199 stars) and Rounding (RConsortium/pharma-skills, 120 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains OpenMed ETL to OMOP CDM?

maziyarpanahi (a GitHub user) maintains it in maziyarpanahi/openmed, which has 5,506 GitHub stars. The repository holds 74 skills in this directory. The repository was last updated on October 11, 2026.

Source: maziyarpanahi/openmed on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.