Agent skill

Defining Cohort Phenotypes

by maziyarpanahi in maziyarpanahi/openmed

Authors computable phenotype and cohort definitions in the OHDSI ATLAS / CIRCE style over the OMOP CDM, combining standard concept sets with NLP-derived features that OpenMed extracts.

Apache-2.0Auto-check passedAI & LLM Engineering

Install Defining Cohort Phenotypes

skills CLI
$ npx skills add maziyarpanahi/openmed --skill defining-cohort-phenotypes -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install maziyarpanahi/openmed defining-cohort-phenotypes --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/maziyarpanahi/openmed.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/defining-cohort-phenotypes .claude/skills/defining-cohort-phenotypes && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
defining-cohort-phenotypes
GitHub stars
5.5k
Token cost
~1.9k tokens
SKILL.md length
676 words
Files
1
Skills in repo
74
Repo updated
First seen
Licence
Apache-2.0

At a glance

Authors computable phenotype and cohort definitions in the OHDSI ATLAS / CIRCE style over the OMOP CDM, combining standard concept sets with NLP-derived features that OpenMed extracts.

  • Works in 6 steps: Start from a library definition if one… → Build concept sets from standard OMOP… → Identify text-only criteria the codes… → …
  • The user wants to define a patient cohort
  • SKILL.md covers When to use, Anatomy of a CIRCE cohort…, Augmenting with OpenMed NLP… and Workflow, plus 3 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Defining Cohort Phenotypes is an agent skill from maziyarpanahi/openmed. Authors computable phenotype and cohort definitions in the OHDSI ATLAS / CIRCE style over the OMOP CDM, combining standard concept sets with NLP-derived features that OpenMed extracts. Use when the user wants to define a patient cohort, write a computable phenotype, reuse PheKB or OHDSI Phenotype Library logic, build concept sets, or augment code-based criteria with text features. Trigger keywords: phenotype, cohort definition, OHDSI, ATLAS, CIRCE, OMOP CDM, concept set, PheKB, Phenotype Library, eMERGE…

Its SKILL.md is about 1.9k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering Natural language processing. The repository describes itself as: Local-first healthcare AI: clinical NER and HIPAA PII de-identification on hardware you control. 2,200+ medical models, 35 model-backed PII languages, and Python, MLX, Android… The licence is Apache-2.0.

When your agent uses it

  • The user wants to define a patient cohort
  • Write a computable phenotype
  • OHDSI Phenotype Library logic
  • Build concept sets

Example prompts

  • “Use the defining-cohort-phenotypes skill to author computable phenotype and cohort definitions in the OHDSI ATLAS / CIRCE style over the OMOP CDM…”
  • “/defining-cohort-phenotypes”

Requirements

  • Python 3

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Start from a library definition if one exists (OHDSI Phenotype Library /
  2. Build concept sets from standard OMOP concepts; set includeDescendants
  3. Identify text-only criteria the codes miss; extract them with
  4. Assemble the CIRCE JSON (concept sets + expression) — in ATLAS or directly.
  5. Validate against OMOP CDM: generate SQL, run on a (synthetic/de-identified)
  6. Document human-readable logic alongside the JSON for portability.

What it can do on your machine

Read from SKILL.md and the folder at commit 34d7b8c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are jsonc and python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • github.com
    • ohdsi.github.io
    • phekb.org

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Defining Cohort Phenotypes loads about 1.9k tokens when it runs. Until then it costs about 203 tokens; SKILL.md has 676 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~203
When it runs · the whole SKILL.md, loaded when a task matches
~1.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from maziyarpanahi/openmed at commit 34d7b8c, republished under its Apache-2.0 licence (© maziyarpanahi). 676 words, ~1,928 tokens.

Download SKILL.mdSave it as .claude/skills/defining-cohort-phenotypes/SKILL.md (or your agent's skills folder).
name
defining-cohort-phenotypes
description
Authors computable phenotype and cohort definitions in the OHDSI ATLAS / CIRCE style over the OMOP CDM, combining standard concept sets with NLP-derived features that OpenMed extracts. Use when the user wants to define a patient cohort, write a computable phenotype, reuse PheKB or OHDSI Phenotype Library logic, build concept sets, or augment code-based criteria with text features. Trigger keywords: phenotype, cohort definition, OHDSI, ATLAS, CIRCE, OMOP CDM, concept set, PheKB, Phenotype Library, eMERGE, computable phenotype. Pairs adjacent to OpenMed: NLP features from openmed.analyze_text augment code-based phenotypes for entities that are poorly captured by structured codes. OMOP CDM and OHDSI tools are open source; restricted vocabularies (SNOMED, CPT) are user-supplied.
license
Apache-2.0
metadata.project
OpenMed
metadata.category
research-genomics
metadata.pairs
adjacent
metadata.version
1.0

Defining cohort phenotypes (OHDSI / OMOP CDM)

A computable phenotype is a portable, executable definition of "which patients have condition X" — concept sets plus inclusion logic that runs against any OMOP CDM-compliant database. In the OHDSI stack, ATLAS authors these visually, CIRCE serializes them to a standardized JSON representation, and that JSON compiles to database-specific SQL. This skill helps you author such definitions and augment them with NLP features that OpenMed extracts from clinical text — exactly the signals that structured codes miss.

OMOP CDM, ATLAS, CIRCE, and the OHDSI Phenotype Library are open source. The vocabulary content you reference (SNOMED CT, CPT4, ICD) is user-supplied — do not bundle restricted terminologies; load them into your own OMOP vocabulary tables with your own licenses.

When to use

  • You need a reproducible cohort definition for analytics or research.
  • You want to reuse an existing PheKB or OHDSI Phenotype Library definition and adapt it.
  • A phenotype depends on facts that live only in free text (e.g. smoking status, symptom severity, social context) and code-based logic alone is weak.

For terminology grounding of individual entities, see coding-icd10, normalizing-rxnorm, mapping-loinc; this skill is about composing them into a cohort.

Anatomy of a CIRCE cohort definition

A CIRCE cohort definition JSON has two parts: ConceptSets (the code lists) and an expression (entry event + inclusion rules). Shape (abridged):

jsonc
{
  "ConceptSets": [{
    "id": 0, "name": "Type 2 diabetes",
    "expression": { "items": [{
      "concept": { "CONCEPT_ID": 201826,           // OMOP standard concept
                   "CONCEPT_CODE": "44054006",      // SNOMED (user vocab)
                   "VOCABULARY_ID": "SNOMED" },
      "includeDescendants": true                    // pull the hierarchy
    }] }
  }],
  "PrimaryCriteria": {                              // entry event
    "CriteriaList": [{ "ConditionOccurrence": { "CodesetId": 0 } }],
    "ObservationWindow": { "PriorDays": 0, "PostDays": 0 },
    "PrimaryCriteriaLimit": { "Type": "First" }
  },
  "InclusionRules": [{
    "name": "Adult at index",
    "expression": { "Type": "ALL", "CriteriaList": [{
      "Criteria": { "ConditionEra": { "AgeAtStart": { "Value": 18, "Op": "gte" } } }
    }] }
  }]
}

You author this in ATLAS (recommended) or by hand. The OHDSI Phenotype Library ships hundreds of vetted definitions as exactly this JSON; reuse before you write.

Augmenting with OpenMed NLP features

Code-based phenotypes are blind to facts that only appear in notes. The pattern is materialize an NLP feature as OMOP rows, then reference it like any concept set.

python
import openmed

# 1) Extract the text feature OpenMed is good at (e.g. tobacco use, symptom)
note = "Patient is a current smoker, ~1 pack/day, with worsening dyspnea."
res = openmed.analyze_text(note, model_name="disease_detection_superclinical",
                           output_format="dict")

# 2) Write a derived OBSERVATION (or a custom cohort attribute) per patient,
#    mapping each extracted entity to a standard concept (grounded out-of-process).
#    e.g. Observation: "Current smoker" -> a SNOMED concept in your vocab.

# 3) Reference that concept in a CIRCE ConceptSet, so the phenotype combines
#    structured codes AND the NLP-derived flag in one inclusion rule.

This mirrors how eMERGE and PheKB phenotypes mix structured codes with NLP: the NLP step contributes high-recall flags for concepts that ICD/CPT capture poorly, and CIRCE composes them with the rest of the logic.

Workflow

  1. Start from a library definition if one exists (OHDSI Phenotype Library / PheKB) and adapt; otherwise design entry event + inclusion rules.
  2. Build concept sets from standard OMOP concepts; set includeDescendants to capture hierarchies. Vocabulary content comes from your own licensed tables.
  3. Identify text-only criteria the codes miss; extract them with openmed.analyze_text and materialize as OMOP rows / cohort attributes.
  4. Assemble the CIRCE JSON (concept sets + expression) — in ATLAS or directly.
  5. Validate against OMOP CDM: generate SQL, run on a (synthetic/de-identified) database, review cohort counts; iterate with PheValuator-style checks.
  6. Document human-readable logic alongside the JSON for portability.
Show full SKILL.md (263 more words)Show less

Hand-off to / from OpenMed

  • OpenMed → phenotype features. openmed.analyze_text over notes yields Disease, Pharmaceutical, Genomics, Oncology, and social/behavioral spans. Ground each to a standard concept (coding-icd10, normalizing-rxnorm, mapping-loinc, or your SNOMED map) and write it into OMOP so CIRCE can reference it.
  • Phenotype → OpenMed scope. A cohort definition tells you which notes to process: run OpenMed only on the cohort's documents to extract the features the phenotype needs, keeping compute and PHI exposure minimal.
  • Run locally on de-identified or synthetic OMOP data. De-identify notes with openmed.deidentify before they enter any shared analytics environment.

Edge cases & gotchas

  • Standard vs source concepts. OMOP maps source codes (ICD-10-CM) to standard concepts (usually SNOMED). Build concept sets on standard concepts and let the source-to-standard map do the translation, or you will miss rows.
  • Descendants matter. Forgetting includeDescendants silently drops the hierarchy (e.g. all diabetes subtypes). Forgetting nothing can over-capture — review the resolved concept list.
  • NLP feature provenance. Tag NLP-derived OMOP rows distinctly (e.g. a type_concept indicating "derived from NLP") so analysts know the signal is probabilistic, not adjudicated.
  • Vocabulary licensing. SNOMED CT, CPT4, and similar require their own licenses and are not redistributed here — load them into your OMOP vocab.
  • Portability ≠ equivalence. The same JSON runs everywhere, but data capture differs by site; validate cohort counts per source before trusting them.
  • Not clinical advice. Phenotype membership supports research/analytics; it is not a diagnosis.

Standards & references

© maziyarpanahi, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/defining-cohort-phenotypes of maziyarpanahi/openmed.

Open the folder on GitHubat commit 34d7b8c

Compare with similar skills

Defining Cohort Phenotypes next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Defining Cohort Phenotypes compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Defining Cohort Phenotypes this skillmaziyarpanahi/openmed5.5k—~1.9kAutomated safety check: PassApache-2.0
Hugging Face TokenizersOrchestra-Research/AI-Research-SKILLs13k6 repos~3.4kAutomated safety check: PassMIT
Gptqmodel Tokenizer NormalizationModelCloud/GPTQModel1.3k—~1.1kAutomated safety check: PassCustom licence
Andrej KarpathyK-Dense-AI/mimeo282—~1.9kAutomated safety check: PassMIT
Comparetaishi-i/awesome-japanese-nlp-resources1k—~4.1kAutomated safety check: NotesCC0-1.0
Researchtaishi-i/awesome-japanese-nlp-resources1k—~3.5kAutomated safety check: NotesCC0-1.0

Similar skills

  • Hugging Face Tokenizers

    Orchestra-Research/AI-Research-SKILLs

    Shows how to load, train and use fast Hugging Face tokenizers, with BPE, WordPiece and Unigram models, padding, truncation and alignment tracking.

    13k GitHub starsUsed in 6 repos~3.4k tokens
    AI & LLM EngineeringAuto-check passed
  • Diagnose and correct GPT-QModel tokenizer initialization, tokenization normalization, special-token handling, prompt rendering, and chat-template problems.

    1.3k GitHub stars~1.1k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Andrej Karpathy

    K-Dense-AI/mimeo

    Applies the mental models and frameworks of Andrej Karpathy (deep learning, former Director of AI at Tesla, founding member of OpenAI, Eureka Labs).

    282 GitHub stars~1.9k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Compare

    taishi-i/awesome-japanese-nlp-resources

    Compare several Japanese NLP libraries, models, or datasets for a keyword (a specific tool name, or a function/task like '形態素解析') across a handful of criteria chosen for that comparison, rendered as…

    1k GitHub stars~4.1k tokensUpdated 5 days ago
    AI & LLM EngineeringAuto-check: notes
  • Research

    taishi-i/awesome-japanese-nlp-resources

    Analyze current trends and challenges in Japanese NLP for a topic.

    1k GitHub stars~3.5k tokensUpdated 5 days ago
    AI & LLM EngineeringAuto-check: notes
  • Sentence Transformers Embeddings

    Orchestra-Research/AI-Research-SKILLs

    Generates text embeddings locally with the sentence-transformers library for RAG, semantic search, clustering and similarity, with model picks for general, multilingual and legal text.

    13k GitHub starsUsed in 2 repos~1.6k tokens
    AI & LLM EngineeringAuto-check passed

More from maziyarpanahi/openmed

All 74 skills in this repo
  • Checks OpenMed de-identified clinical text against the 18 HIPAA Safe Harbor identifier categories and reports gaps and residual re-identification risk.

    5.5k GitHub stars~1.7k tokensUpdated today
    Auto-check passed
  • OpenMed Model Card Writer

    maziyarpanahi/openmed

    Fills in a model card for an OpenMed clinical NER or de-identification model from its evaluation reports: intended use, metrics, subgroups and limitations.

    5.5k GitHub stars~1.8k tokensUpdated today
    Auto-check passed
  • Walks a data pipeline against the HIPAA Privacy and Security Rule checklist and produces a gap report before it processes patient data.

    5.5k GitHub stars~2k tokensUpdated today
    Auto-check passed
  • ICD-10 Coding Assistant

    maziyarpanahi/openmed

    Suggests candidate ICD-10-CM diagnosis and ICD-10-PCS procedure codes for clinical text extracted by OpenMed, with rationale for a certified coder to review.

    5.5k GitHub stars~2k tokensUpdated today
    Auto-check passed
  • OpenMed ETL to OMOP CDM

    maziyarpanahi/openmed

    Maps OpenMed-extracted, terminology-coded conditions, drugs and measurements into OMOP CDM v5.4 tables for OHDSI and ATLAS analytics.

    5.5k GitHub stars~1.9k tokensUpdated today
    Auto-check passed
  • Extracting SDOH and Z-Codes

    maziyarpanahi/openmed

    Finds social risks such as housing instability or food insecurity in clinical notes and proposes matching ICD-10-CM Z-codes for a coder to confirm.

    5.5k GitHub stars~1.9k tokensUpdated today
    Auto-check passed

Questions about Defining Cohort Phenotypes

What does Defining Cohort Phenotypes do?

Authors computable phenotype and cohort definitions in the OHDSI ATLAS / CIRCE style over the OMOP CDM, combining standard concept sets with NLP-derived features that OpenMed extracts. Defining Cohort Phenotypes is an agent skill from maziyarpanahi/openmed. Authors computable phenotype and cohort definitions in the OHDSI ATLAS / CIRCE style over the OMOP CDM, combining standard concept sets with NLP-derived features that OpenMed extracts.

When should I use Defining Cohort Phenotypes?

Defining Cohort Phenotypes fits situations like: the user wants to define a patient cohort; write a computable phenotype; OHDSI Phenotype Library logic; build concept sets.

How do I install Defining Cohort Phenotypes in Claude Code?

Run `npx skills add maziyarpanahi/openmed --skill defining-cohort-phenotypes -a claude-code`. Or copy the skill folder (skills/defining-cohort-phenotypes in maziyarpanahi/openmed) into .claude/skills/defining-cohort-phenotypes in your project. Claude Code loads it when a task matches its description.

How do I install Defining Cohort Phenotypes in Codex?

Run `npx skills add maziyarpanahi/openmed --skill defining-cohort-phenotypes -a codex`. Or copy the skill folder (skills/defining-cohort-phenotypes in maziyarpanahi/openmed) into .agents/skills/defining-cohort-phenotypes in your project. Codex loads it when a task matches its description.

Can I use Defining Cohort Phenotypes in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add maziyarpanahi/openmed --skill defining-cohort-phenotypes -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/defining-cohort-phenotypes, .gemini/skills/defining-cohort-phenotypes, .github/skills/defining-cohort-phenotypes and .opencode/skills/defining-cohort-phenotypes in your project.

What does Defining Cohort Phenotypes need to run?

SKILL.md names no scripts, command-line tools or credentials: Defining Cohort Phenotypes is instructions for the agent only. Our summary lists: Python 3.

Does Defining Cohort Phenotypes access the network?

SKILL.md names 3 domains. As links in the text: github.com, ohdsi.github.io and phekb.org. This is read from the text; nothing was executed.

Is Defining Cohort Phenotypes safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Defining Cohort Phenotypes use?

Defining Cohort Phenotypes is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Defining Cohort Phenotypes use?

About 1.9k tokens (SKILL.md is roughly 7.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Defining Cohort Phenotypes?

Skills that share tags, products or a category with Defining Cohort Phenotypes: Hugging Face Tokenizers (Orchestra-Research/AI-Research-SKILLs, 13k stars), Gptqmodel Tokenizer Normalization (ModelCloud/GPTQModel, 1.3k stars), Andrej Karpathy (K-Dense-AI/mimeo, 282 stars) and Compare (taishi-i/awesome-japanese-nlp-resources, 1k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Defining Cohort Phenotypes?

maziyarpanahi (a GitHub user) maintains it in maziyarpanahi/openmed, which has 5,506 GitHub stars. The repository holds 74 skills in this directory. The repository was last updated on October 11, 2026.

Source: maziyarpanahi/openmed on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.