Agent skill

Generating Synthea Data

by maziyarpanahi in maziyarpanahi/openmed

Generates synthetic but realistic patient records (FHIR R4 bundles, C-CDA documents, CSV) with MITRE Synthea for development, CI fixtures, demos, and leakage-gate test sets — zero real PHI.

Apache-2.0Auto-check passedTesting & QA

Install Generating Synthea Data

skills CLI
$ npx skills add maziyarpanahi/openmed --skill generating-synthea-data -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install maziyarpanahi/openmed generating-synthea-data --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/maziyarpanahi/openmed.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/generating-synthea-data .claude/skills/generating-synthea-data && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
generating-synthea-data
GitHub stars
5.5k
Token cost
~1.6k tokens
SKILL.md length
603 words
Files
1
Skills in repo
74
Repo updated
First seen
Licence
Apache-2.0

At a glance

Generates synthetic but realistic patient records (FHIR R4 bundles, C-CDA documents, CSV) with MITRE Synthea for development, CI fixtures, demos, and leakage-gate test sets — zero real PHI.

  • Works in 6 steps: Choose scale & geography. -p N sets… → Pick formats. Enable FHIR… → Pin a seed (-s) for reproducible… → …
  • Shareable test data for an OpenMed pipeline
  • SKILL.md covers When to use, Quick start, Workflow and Hand-off to / from OpenMed, plus 2 more sections
  • Calls git; reaches github.com

What it does

Generating Synthea Data is an agent skill from maziyarpanahi/openmed. Generates synthetic but realistic patient records (FHIR R4 bundles, C-CDA documents, CSV) with MITRE Synthea for development, CI fixtures, demos, and leakage-gate test sets — zero real PHI. Use when you need safe, shareable test data for an OpenMed pipeline, reproducible fixtures for tests, or a held-out set for de-identification leakage gates, instead of touching real clinical data. Synthea output feeds the FHIR/C-CDA ingestion skills and openmed.eval. Trigger keywords: Synthea, synthetic data, fake patients…

Its SKILL.md is about 1.6k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Testing & QA, covering Clinical and healthcare research and Test data and fixtures. The repository describes itself as: Local-first healthcare AI: clinical NER & HIPAA PII de-identification that runs 100% on-device. 2,200+ medical models, 21 languages, Apple MLX + Python, no cloud, no patient data…. The licence is Apache-2.0.

When your agent uses it

  • Shareable test data for an OpenMed pipeline
  • Reproducible fixtures for tests
  • A held-out set for de-identification leakage gates
  • Instead of touching real clinical data

Example prompts

  • “Use the generating-synthea-data skill to generate synthetic but realistic patient records (FHIR R4 bundles, C-CDA documents, CSV) with MITRE Synthea…”
  • “/generating-synthea-data”

Requirements

  • Python 3

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Choose scale & geography. -p N sets population; the state/location
  2. Pick formats. Enable FHIR (exporter.fhir.export), C-CDA
  3. Pin a seed (-s) for reproducible fixtures so test assertions are stable.
  4. Select modules (optional). Synthea ships disease modules
  5. Use as a leakage-gate corpus. Synthea emits known fake names, MRNs,
  6. Commit fixtures under your test tree (e.g. tests/fixtures/synthea/) —

What it can do on your machine

Read from SKILL.md and the folder at commit 6b1bb2c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • git

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • github.com

    Also links to:

    • hl7.org
    • doi.org

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Generating Synthea Data loads about 1.6k tokens when it runs. Until then it costs about 153 tokens; SKILL.md has 603 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~153
When it runs · the whole SKILL.md, loaded when a task matches
~1.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from maziyarpanahi/openmed at commit 6b1bb2c, republished under its Apache-2.0 licence (© maziyarpanahi). 603 words, ~1,610 tokens.

Download SKILL.mdSave it as .claude/skills/generating-synthea-data/SKILL.md (or your agent's skills folder).
name
generating-synthea-data
description
Generates synthetic but realistic patient records (FHIR R4 bundles, C-CDA documents, CSV) with MITRE Synthea for development, CI fixtures, demos, and leakage-gate test sets — zero real PHI. Use when you need safe, shareable test data for an OpenMed pipeline, reproducible fixtures for tests, or a held-out set for de-identification leakage gates, instead of touching real clinical data. Synthea output feeds the FHIR/C-CDA ingestion skills and openmed.eval. Trigger keywords: Synthea, synthetic data, fake patients, test fixtures, demo data, FHIR bundle generator, synthetic EHR, no PHI.
license
Apache-2.0
metadata.project
OpenMed
metadata.category
data-ingestion
metadata.pairs
adjacent
metadata.version
1.0

Generating Synthetic Patient Data with Synthea

You cannot develop, test, or demo a clinical NLP pipeline on real PHI without a mountain of governance — and you shouldn't have to. Synthea (MITRE's Synthetic Patient Population Simulator) generates statistically realistic, fully synthetic patients: complete longitudinal records as FHIR R4 bundles, C-CDA documents, and flat CSV, with zero real-PHI risk. Use it for OpenMed dev fixtures, CI, demos, and — importantly — as held-out test sets for de-identification leakage gates, where you need known-synthetic "PHI" to measure recall.

When to use

  • Building or demoing an OpenMed ingestion pipeline (FHIR, C-CDA) and need shareable input that is safe to commit and pass around.
  • Creating deterministic CI fixtures so tests don't depend on protected data.
  • Producing a leakage-gate test corpus: synthetic notes with known fake identifiers, so you can score whether openmed.deidentify removed them all.
  • Teaching/onboarding without a data-use agreement.

Quick start

Synthea is a Java tool. Generate a small population in multiple formats:

bash
# Requires Java 11+. Clone and build once.
git clone https://github.com/synthetichealth/synthea && cd synthea
./gradlew build -x test

# Generate 50 patients in Massachusetts as FHIR R4 + C-CDA + CSV.
./run_synthea -p 50 Massachusetts \
  --exporter.fhir.export=true \
  --exporter.ccda.export=true \
  --exporter.csv.export=true \
  --exporter.baseDirectory=./output

# Reproducible runs: fix the seed so fixtures are stable across CI.
./run_synthea -s 12345 -p 20 --exporter.baseDirectory=./fixtures

Output lands under output/fhir/, output/ccda/, and output/csv/. Feed the FHIR bundles to parsing-... skills, or hand narrative straight to OpenMed:

python
import json, openmed

bundle = json.load(open("output/fhir/Patient_xyz.json"))
for entry in bundle.get("entry", []):
    res = entry.get("resource", {})
    div = (res.get("text") or {}).get("div", "")     # narrative XHTML
    if div.strip():
        deid = openmed.deidentify(div, method="replace", policy="hipaa_safe_harbor")
        result = openmed.analyze_text(deid.text, output_format="dict")

Synthea data is synthetic, so de-identifying it is exercising the pipeline, not a privacy requirement — which is exactly what makes it a great test bed.

Workflow

  1. Choose scale & geography. -p N sets population; the state/location argument shapes demographics and addresses. Start small (10–50) for fixtures.
  2. Pick formats. Enable FHIR (exporter.fhir.export), C-CDA (exporter.ccda.export), and/or CSV per your ingestion path. FHIR R4 is the default and pairs with fetching-fhir-resources; C-CDA pairs with parsing-ccda-documents.
  3. Pin a seed (-s) for reproducible fixtures so test assertions are stable.
  4. Select modules (optional). Synthea ships disease modules (-m "diabetes*" to filter); choose modules matching the entities your OpenMed pipeline targets.
  5. Use as a leakage-gate corpus. Synthea emits known fake names, MRNs, addresses, and dates — inject/collect these as ground-truth PHI spans and score openmed.deidentify recall with openmed.eval (evaluating-with-leakage-gates). Because the "PHI" is synthetic and known, you can measure misses without exposing anyone.
  6. Commit fixtures under your test tree (e.g. tests/fixtures/synthea/) — it is safe to version-control synthetic output.
Show full SKILL.md (261 more words)Show less

Hand-off to / from OpenMed

  • To OpenMed (as input): Synthea FHIR/C-CDA narrative → openmed.deidentify → openmed.analyze_text, via the fetching-fhir-resources and parsing-ccda-documents skills.
  • To OpenMed eval: use Synthea's known synthetic identifiers as ground truth for openmed.eval de-identification leakage gates — the daily-release thesis gates on leakage, not F1 alone, and synthetic data lets you build that test set without governance overhead.
  • Adjacent, not in-pipeline: Synthea is a source of safe data; it does not call OpenMed and OpenMed does not call it. Keep it in dev/CI, never as a production data source.

Edge cases & gotchas

  • Synthetic ≠ statistically perfect. Synthea reproduces realistic disease progression and demographics but is not a substitute for real-world distribution validation; never report clinical model accuracy only on synthetic data.
  • Narrative is templated. FHIR text.div narrative is generated from templates, so it is more regular than dictated notes. For NER robustness, supplement with varied real (de-identified) text where governance allows.
  • Determinism needs the seed. Without -s, every run differs — CI fixtures will churn. Always pin the seed for committed fixtures.
  • Version drift. Synthea modules and FHIR profile output change across releases; pin the Synthea version (git tag) alongside your fixtures.
  • Large populations are heavy. -p 100000 produces gigabytes; size to need.
  • Licensing. Synthea and its generated output are permissively licensed (Apache-2.0), so output is safe to redistribute — unlike MIMIC/i2b2/n2c2, which require data-use agreements and must stay user-supplied.

Standards & references

© maziyarpanahi, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/generating-synthea-data of maziyarpanahi/openmed.

Open the folder on GitHubat commit 6b1bb2c

Compare with similar skills

Generating Synthea Data next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Generating Synthea Data compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Generating Synthea Data this skillmaziyarpanahi/openmed5.5k—~1.6kAutomated safety check: PassApache-2.0
Synthetic Dataflonat/flonat-research146—~2.5kAutomated safety check: PassMIT
Phi Prompt Guardaipoch/medical-research-skills2k—~4.2kAutomated safety check: PassMIT
Fs Fixtureprivatenumber/fs-fixture100—~1.2kAutomated safety check: PassMIT
Dev Tenant APInightscout/nocturne139—~1.4kAutomated safety check: PassNone
Rsibench Data Factoryevolvent-ai/RSIBench-Data171—~640Automated safety check: NotesNone

Similar skills

  • Synthetic Data

    flonat/flonat-research

    Generate structurally realistic synthetic datasets for pipeline prototyping, test coverage, or prospective power analysis.

    146 GitHub stars~2.5k tokensUpdated 10 days ago
    Testing & QAAuto-check passed
  • Phi Prompt Guard

    aipoch/medical-research-skills

    Runtime, prompt-time behavioral guardrail that helps reduce PHI exposure in LLM-assisted workflows by detecting PHI-bearing prompts, avoiding unsafe tool actions that would pull more PHI in, and…

    2k GitHub stars~4.2k tokensUpdated 22 days ago
    AI & LLM EngineeringAuto-check passed
  • Fs Fixture

    privatenumber/fs-fixture

    Create disposable file system test fixtures from objects, templates, or empty directories with automatic cleanup.

    100 GitHub stars~1.2k tokensUpdated 1 mo ago
    Testing & QAAuto-check passed
  • Dev Tenant API

    nightscout/nocturne

    Interact with Nocturne's local dev-only API: seed a loginable tenant preloaded with realistic sample data, obtain a browser session (loginLink) or bearer token headlessly, export/re-seed the dev…

    139 GitHub stars~1.4k tokensUpdated today
    Testing & QAAuto-check passed
  • Rsibench Data Factory

    evolvent-ai/RSIBench-Data

    Use inside RSIBench-Data when testing whether an automation agent can improve a target model on a configured benchmark through synthetic Tinker SFT data, Tinker sampling, and E2B-based Harbor…

    171 GitHub stars~640 tokensUpdated 1 mo ago
    Testing & QAAuto-check: notes
  • Eval Design

    agentscope-ai/OpenJudge

    A skill your agent uses when the user needs to design evaluation datasets, create test cases, stratify samples, generate adversarial examples, extract eval dimensions from traces/specs, or build a…

    870 GitHub stars~2.8k tokensUpdated 28 days ago
    Testing & QAAuto-check: warnings

More from maziyarpanahi/openmed

All 74 skills in this repo
  • Checks OpenMed de-identified clinical text against the 18 HIPAA Safe Harbor identifier categories and reports gaps and residual re-identification risk.

    5.5k GitHub stars~1.7k tokensUpdated today
    Auto-check passed
  • OpenMed Model Card Writer

    maziyarpanahi/openmed

    Fills in a model card for an OpenMed clinical NER or de-identification model from its evaluation reports: intended use, metrics, subgroups and limitations.

    5.5k GitHub stars~1.8k tokensUpdated today
    Auto-check passed
  • Walks a data pipeline against the HIPAA Privacy and Security Rule checklist and produces a gap report before it processes patient data.

    5.5k GitHub stars~2k tokensUpdated today
    Auto-check passed
  • ICD-10 Coding Assistant

    maziyarpanahi/openmed

    Suggests candidate ICD-10-CM diagnosis and ICD-10-PCS procedure codes for clinical text extracted by OpenMed, with rationale for a certified coder to review.

    5.5k GitHub stars~2k tokensUpdated today
    Auto-check passed
  • OpenMed ETL to OMOP CDM

    maziyarpanahi/openmed

    Maps OpenMed-extracted, terminology-coded conditions, drugs and measurements into OMOP CDM v5.4 tables for OHDSI and ATLAS analytics.

    5.5k GitHub stars~1.9k tokensUpdated today
    Auto-check passed
  • Extracting SDOH and Z-Codes

    maziyarpanahi/openmed

    Finds social risks such as housing instability or food insecurity in clinical notes and proposes matching ICD-10-CM Z-codes for a coder to confirm.

    5.5k GitHub stars~1.9k tokensUpdated today
    Auto-check passed

Questions about Generating Synthea Data

What does Generating Synthea Data do?

Generates synthetic but realistic patient records (FHIR R4 bundles, C-CDA documents, CSV) with MITRE Synthea for development, CI fixtures, demos, and leakage-gate test sets — zero real PHI. Generating Synthea Data is an agent skill from maziyarpanahi/openmed. Generates synthetic but realistic patient records (FHIR R4 bundles, C-CDA documents, CSV) with MITRE Synthea for development, CI fixtures, demos, and leakage-gate test sets — zero real PHI.

When should I use Generating Synthea Data?

Generating Synthea Data fits situations like: shareable test data for an OpenMed pipeline; reproducible fixtures for tests; A held-out set for de-identification leakage gates; instead of touching real clinical data.

How do I install Generating Synthea Data in Claude Code?

Run `npx skills add maziyarpanahi/openmed --skill generating-synthea-data -a claude-code`. Or copy the skill folder (skills/generating-synthea-data in maziyarpanahi/openmed) into .claude/skills/generating-synthea-data in your project. Claude Code loads it when a task matches its description.

How do I install Generating Synthea Data in Codex?

Run `npx skills add maziyarpanahi/openmed --skill generating-synthea-data -a codex`. Or copy the skill folder (skills/generating-synthea-data in maziyarpanahi/openmed) into .agents/skills/generating-synthea-data in your project. Codex loads it when a task matches its description.

Can I use Generating Synthea Data in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add maziyarpanahi/openmed --skill generating-synthea-data -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/generating-synthea-data, .gemini/skills/generating-synthea-data, .github/skills/generating-synthea-data and .opencode/skills/generating-synthea-data in your project.

What does Generating Synthea Data need to run?

Going by SKILL.md and its folder, Generating Synthea Data needs the command-line tools its instructions call (git). Our summary lists: Python 3.

Does Generating Synthea Data access the network?

SKILL.md names 3 domains. In commands or code: github.com; the agent is likely to contact it when it follows the instructions. As links in the text: hl7.org and doi.org. This is read from the text; nothing was executed.

Is Generating Synthea Data safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Generating Synthea Data use?

Generating Synthea Data is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Generating Synthea Data use?

About 1.6k tokens (SKILL.md is roughly 6.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Generating Synthea Data?

Skills that share tags, products or a category with Generating Synthea Data: Synthetic Data (flonat/flonat-research, 146 stars), Phi Prompt Guard (aipoch/medical-research-skills, 2k stars), Fs Fixture (privatenumber/fs-fixture, 100 stars) and Dev Tenant API (nightscout/nocturne, 139 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Generating Synthea Data?

maziyarpanahi (a GitHub user) maintains it in maziyarpanahi/openmed, which has 5,493 GitHub stars. The repository holds 74 skills in this directory. The repository was last updated on October 9, 2026.

Source: maziyarpanahi/openmed on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.