Generating Synthea Data
maziyarpanahi/openmed
Generates synthetic but realistic patient records (FHIR R4 bundles, C-CDA documents, CSV) with MITRE Synthea for development, CI fixtures, demos, and leakage-gate test sets — zero real PHI.
Generate structurally realistic synthetic datasets for pipeline prototyping, test coverage, or prospective power analysis.
$ npx skills add flonat/flonat-research --skill synthetic-data -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install flonat/flonat-research synthetic-data --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/flonat/flonat-research.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/synthetic-data .claude/skills/synthetic-data && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "synthetic-data" agent skill from https://github.com/flonat/flonat-research/tree/main/skills/synthetic-data into .claude/skills/synthetic-data/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "synthetic-data", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/flonat/flonat-research/tree/main/skills/synthetic-dataType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add flonat/flonat-research --skill synthetic-data -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install flonat/flonat-research synthetic-data --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/flonat/flonat-research.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/synthetic-data .agents/skills/synthetic-data && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "synthetic-data" agent skill from https://github.com/flonat/flonat-research/tree/main/skills/synthetic-data into .agents/skills/synthetic-data/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "synthetic-data", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add flonat/flonat-research --skill synthetic-data -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install flonat/flonat-research synthetic-data --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/flonat/flonat-research.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/synthetic-data .cursor/skills/synthetic-data && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "synthetic-data" agent skill from https://github.com/flonat/flonat-research/tree/main/skills/synthetic-data into .cursor/skills/synthetic-data/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "synthetic-data", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/flonat/flonat-research.git --path skills/synthetic-data--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add flonat/flonat-research --skill synthetic-data -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install flonat/flonat-research synthetic-data --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/flonat/flonat-research.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/synthetic-data .gemini/skills/synthetic-data && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "synthetic-data" agent skill from https://github.com/flonat/flonat-research/tree/main/skills/synthetic-data into .gemini/skills/synthetic-data/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "synthetic-data", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install flonat/flonat-research synthetic-dataInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add flonat/flonat-research --skill synthetic-data -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/flonat/flonat-research.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/synthetic-data .github/skills/synthetic-data && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "synthetic-data" agent skill from https://github.com/flonat/flonat-research/tree/main/skills/synthetic-data into .github/skills/synthetic-data/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "synthetic-data", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add flonat/flonat-research --skill synthetic-data -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install flonat/flonat-research synthetic-data --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/flonat/flonat-research.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/synthetic-data .opencode/skills/synthetic-data && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "synthetic-data" agent skill from https://github.com/flonat/flonat-research/tree/main/skills/synthetic-data into .opencode/skills/synthetic-data/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "synthetic-data", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
synthetic-dataGenerate structurally realistic synthetic datasets for pipeline prototyping, test coverage, or prospective power analysis.
Synthetic Data is an agent skill from flonat/flonat-research. Generate structurally realistic synthetic datasets for pipeline prototyping, test coverage, or prospective power analysis. Use when code or design must be exercised before real data are available or accessible. Never substitute synthetic records for governed raw data.
Its SKILL.md is about 2.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files, including reference files (for example `references/calibration-targets.md`, `references/dgp-recipes-experimental.md` and `references/dgp-recipes-observational.md`).
It sits in Testing & QA, covering Test data and fixtures, Experimental design and Prototyping. The repository describes itself as: Shareable Claude Code + Codex infrastructure for PhD researchers — skills, agents, hooks, and rules for academic workflows. The licence is MIT.
6 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit da27600. It shows what the files ask for, not the result of running them.
Pre-approves these tools, so the agent can use them without asking each time:
Bash(uv*Rscript*R*mkdir*ls*)ReadWriteEditGlobGrep…and 1 more on the same allowed-tools line.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are markdown).
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Synthetic Data loads about 2.5k tokens when it runs, and up to ~10k if it reads all its reference files. Until then it costs about 71 tokens; SKILL.md has 1,059 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from flonat/flonat-research at commit da27600, republished under its MIT licence (© flonat). 1,059 words, ~2,547 tokens.
.claude/skills/synthetic-data/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.Generate structurally realistic synthetic datasets for pilot testing, power analysis, and method development.
| Mode | What it produces | Entry point |
|---|---|---|
| From design | Synthetic data matching an existing experiment design document | "Generate test data for my experiment" |
| From schema | Synthetic data from a user-described structure | "Create a dataset with these variables" |
| Calibrated | Synthetic data calibrated to published summary statistics | "Make fake data matching these descriptives" |
Default: From schema. If an experiment design document exists in docs/, auto-select From design. If user provides published statistics, auto-select Calibrated.
experiment-design Power mode)experiment-designdata-analysiscausal-designDetect from context or ask:
| Signal | Mode |
|---|---|
docs/experiment-design.md exists | From design |
| User describes variables, types, relationships | From schema |
| User provides means, SDs, correlations from a paper | Calibrated |
| Ambiguous | Ask |
Gather the following (adapt questions to mode):
| Parameter | Question | Default |
|---|---|---|
| Variables | What variables do you need? | — |
| Types | Continuous, binary, ordinal, categorical? | Infer from name |
| Sample size | How many observations? | 500 |
| Treatment | Is there a treatment variable? How many arms? | — |
| Effect size | Expected treatment effect (Cohen's d, OR, etc.)? | 0.3 (small-medium) |
| Correlations | Which variables should be correlated? How strongly? | — |
| Clustering | Are observations nested (e.g., students in classrooms)? | No |
| Panel structure | Multiple time periods? How many? | Cross-section |
| Missing data | Should the data include realistic missingness? | No |
| Language | R or Python? | Detect from project or ask |
For From design mode, extract most parameters from the design document automatically and confirm with the user.
For Calibrated mode, require the user to provide published summary statistics. Read references/calibration-targets.md for the calibration procedure.
Read references/dgp-recipes.md for code patterns matching the requested design.
Read shared/multi-language-conventions.md for language-specific code style.
Generate a self-contained script that:
.rds/.parquet)For large Monte Carlo sweeps (10k+ simulations, multi-condition grids, or long-running bootstrap): run on [HPC cluster] instead of locally. Drop the generation script into hpc/ with submit.sbatch (compute partition) or sweep.sbatch (array over seeds/conditions) — templates at Task Management templates/slurm/, guide at docs/guides/hpc.md. All SLURM templates log git-sha.txt to OUT_DIR so synthetic datasets remain traceable to the DGP code version.
Run the script and verify:
Output routing:
| File | Location |
|---|---|
| Generated dataset | data/synthetic/{name}.csv |
| Generation script | code/generate_synthetic_{name}.R (or .py) |
| Data dictionary | data/synthetic/{name}_dictionary.md |
NEVER write to data/raw/. The data-sensitivity rule applies -- synthetic data is not raw data and must be clearly separated.
Create data/synthetic/ if it does not exist.
Produce a markdown data dictionary alongside the dataset:
# Data Dictionary: {name}
Generated: {date}
Script: `code/generate_synthetic_{name}.R`
Seed: {seed}
N: {sample_size}
| Variable | Type | Description | Distribution | Parameters |
|----------|------|-------------|-------------- |------------|
| id | integer | Unique identifier | Sequential | 1 to N |
| treatment | binary | Treatment assignment | Bernoulli | p = 0.5 |
| outcome | continuous | Primary outcome | Normal | mu = 0, sigma = 1 (control) |
| ... | ... | ... | ... | ... |
## Treatment Effects
- True ATE: {value}
- True CATE by subgroup: {if applicable}
## Missing Data
- Mechanism: {MCAR/MAR/MNAR}
- Rate: {percentage} of {variable}references/calibration-targets.md for the full procedure.docs/experiment-design.md, docs/pre-analysis-plan.md, log/plans/experiment-designThe design document produced by experiment-design Design mode contains everything needed:
| User says | Interpret as |
|---|---|
| "A standard survey dataset" | Demographics + Likert DVs + treatment + attention checks |
| "Panel data" | Units x time with unit and time fixed effects |
| "Something like [paper name]" | Switch to Calibrated mode, use paper's descriptives |
references/calibration-targets.md for the calibration procedure| Resource | When read |
|---|---|
references/dgp-recipes.md | All modes (code patterns for each design type) |
references/calibration-targets.md | Calibrated mode (matching procedure) |
shared/multi-language-conventions.md | All modes (code style) |
experiment-design skill | From design mode (reads its design documents) |
data-analysis skill | Consumes synthetic data for pipeline testing |
data-sensitivity rule | Never write to data/raw/ |
design-before-results rule | Synthetic data supports locking the design before real data |
© flonat, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 5 other files (references) in skills/synthetic-data of flonat/flonat-research.
Open the folder on GitHubat commit da27600
Synthetic Data next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Synthetic Data this skillflonat/flonat-research | 145 | — | ~2.5k | Automated safety check: Pass | MIT | |
| Generating Synthea Datamaziyarpanahi/openmed | 5.5k | — | ~1.6k | Automated safety check: Pass | Apache-2.0 | |
| Slate Ar Qualityudecode/plate | 17k | — | ~474 | Automated safety check: Pass | Custom licence | |
| Pytest Patternscohen-liel/hivemind | 110 | — | ~806 | Automated safety check: Pass | Apache-2.0 | |
| Requirementsrizsotto/Bear | 6.5k | — | ~2k | Automated safety check: Pass | GPL-3.0 | |
| Crap Analysisardalis/RiverBooks | 135 | 2 repos | ~3.4k | Automated safety check: Pass | None |
maziyarpanahi/openmed
Generates synthetic but realistic patient records (FHIR R4 bundles, C-CDA documents, CSV) with MITRE Synthea for development, CI fixtures, demos, and leakage-gate test sets — zero real PHI.
udecode/plate
Slate v2 quality-gap Autoresearch shortcut. An agent skill from udecode/plate.
cohen-liel/hivemind
pytest best practices for writing comprehensive test suites.
rizsotto/Bear
Write, modify, or review a requirement file under docs/requirements -- pick the single owning file, keep the text contract-only, name IDs so they need no explanation, and verify cross-references and…
ardalis/RiverBooks
Analyze code coverage and CRAP (Change Risk Anti-Patterns) scores to identify high-risk code.
privatenumber/fs-fixture
Create disposable file system test fixtures from objects, templates, or empty directories with automatic cleanup.
flonat/flonat-research
Create a large-format academic poster in LaTeX using beamerposter, tikzposter, or baposter.
flonat/flonat-research
Create, revise, and evaluate reusable AI workflow skills, including trigger-quality tests.
flonat/flonat-research
Create, read, edit, or convert Microsoft Word documents while preserving professional document structure.
flonat/flonat-research
Read, create, combine, split, rotate, OCR, watermark, secure, or extract content from PDF files.
flonat/flonat-research
Create or migrate project-level agents, repeatable project workflows, and planning state from one client-neutral contract, then render repository-scoped adapters for both Claude Code and Codex.
flonat/flonat-research
Deliver a fast pre-commit safety scan: file size, anonymity (author / affiliation strings in tex/bib), hardcoded secrets, and invisible-Unicode carriers.
Categories
Generate structurally realistic synthetic datasets for pipeline prototyping, test coverage, or prospective power analysis. Synthetic Data is an agent skill from flonat/flonat-research. Generate structurally realistic synthetic datasets for pipeline prototyping, test coverage, or prospective power analysis.
Synthetic Data fits situations like: design must be exercised before real data are available; tasks that involve Test data and fixtures; tasks that involve Experimental design.
Run `npx skills add flonat/flonat-research --skill synthetic-data -a claude-code`. Or copy the skill folder (skills/synthetic-data in flonat/flonat-research) into .claude/skills/synthetic-data in your project. Claude Code loads it when a task matches its description.
Run `npx skills add flonat/flonat-research --skill synthetic-data -a codex`. Or copy the skill folder (skills/synthetic-data in flonat/flonat-research) into .agents/skills/synthetic-data in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add flonat/flonat-research --skill synthetic-data -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/synthetic-data, .gemini/skills/synthetic-data, .github/skills/synthetic-data and .opencode/skills/synthetic-data in your project.
SKILL.md names no scripts, command-line tools or credentials: Synthetic Data is instructions for the agent only. Its frontmatter pre-approves these tools: Bash(uv*, Rscript*, R*, mkdir*, ls*), Read, Write, Edit, Glob, Grep, AskUserQuestion.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Synthetic Data is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.5k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 7.8k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Synthetic Data: Generating Synthea Data (maziyarpanahi/openmed, 5.5k stars), Slate Ar Quality (udecode/plate, 17k stars), Pytest Patterns (cohen-liel/hivemind, 110 stars) and Requirements (rizsotto/Bear, 6.5k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
flonat (a GitHub user) maintains it in flonat/flonat-research, which has 145 GitHub stars. The repository holds 83 skills in this directory. The repository was last updated on September 29, 2026.
Source: flonat/flonat-research on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.