Agent skill

Synthetic Data

by flonat in flonat/flonat-research

Generate structurally realistic synthetic datasets for pipeline prototyping, test coverage, or prospective power analysis.

MITAuto-check passedTesting & QA

Install Synthetic Data

skills CLI
$ npx skills add flonat/flonat-research --skill synthetic-data -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install flonat/flonat-research synthetic-data --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/flonat/flonat-research.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/synthetic-data .claude/skills/synthetic-data && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
synthetic-data
GitHub stars
145
Token cost
~2.5k tokens
SKILL.md length
1,059 words
Files
6 (incl. references)
Skills in repo
83
Repo updated
First seen
Licence
MIT

At a glance

Generate structurally realistic synthetic datasets for pipeline prototyping, test coverage, or prospective power analysis.

  • Works in 6 steps: Detect Mode → Interview for Data Structure → Generate Script → …
  • Design must be exercised before real data are available
  • SKILL.md covers Modes, When to Use, When NOT to Use and Workflow, plus 5 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Synthetic Data is an agent skill from flonat/flonat-research. Generate structurally realistic synthetic datasets for pipeline prototyping, test coverage, or prospective power analysis. Use when code or design must be exercised before real data are available or accessible. Never substitute synthetic records for governed raw data.

Its SKILL.md is about 2.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files, including reference files (for example `references/calibration-targets.md`, `references/dgp-recipes-experimental.md` and `references/dgp-recipes-observational.md`).

It sits in Testing & QA, covering Test data and fixtures, Experimental design and Prototyping. The repository describes itself as: Shareable Claude Code + Codex infrastructure for PhD researchers — skills, agents, hooks, and rules for academic workflows. The licence is MIT.

When your agent uses it

  • Design must be exercised before real data are available
  • Tasks that involve Test data and fixtures
  • Tasks that involve Experimental design

Example prompts

  • “/synthetic-data”

Requirements

  • Pre-approved tools (allowed-tools): Bash(uv*, Rscript*, R*, mkdir*, ls*), Read, Write, Edit, Glob, Grep, AskUserQuestion

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Detect Mode
  2. Interview for Data Structure
  3. Generate Script
  4. Execute
  5. Save Output
  6. Data Dictionary

What it can do on your machine

Read from SKILL.md and the folder at commit da27600. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Bash(uv*
    • Rscript*
    • R*
    • mkdir*
    • ls*)
    • Read
    • Write
    • Edit
    • Glob
    • Grep

    …and 1 more on the same allowed-tools line.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are markdown).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Synthetic Data loads about 2.5k tokens when it runs, and up to ~10k if it reads all its reference files. Until then it costs about 71 tokens; SKILL.md has 1,059 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~71
When it runs · the whole SKILL.md, loaded when a task matches
~2.5k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~10k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from flonat/flonat-research at commit da27600, republished under its MIT licence (© flonat). 1,059 words, ~2,547 tokens.

Download SKILL.mdSave it as .claude/skills/synthetic-data/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.
name
synthetic-data
description
Generate structurally realistic synthetic datasets for pipeline prototyping, test coverage, or prospective power analysis. Use when code or design must be exercised before real data are available or accessible. Never substitute synthetic records for governed raw data.
allowed-tools
Bash(uv*, Rscript*, R*, mkdir*, ls*), Read, Write, Edit, Glob, Grep, AskUserQuestion
argument-hint
[--mode from-design|from-schema|calibrated]

Synthetic Data Generation

Generate structurally realistic synthetic datasets for pilot testing, power analysis, and method development.

Modes

ModeWhat it producesEntry point
From designSynthetic data matching an existing experiment design document"Generate test data for my experiment"
From schemaSynthetic data from a user-described structure"Create a dataset with these variables"
CalibratedSynthetic data calibrated to published summary statistics"Make fake data matching these descriptives"

Default: From schema. If an experiment design document exists in docs/, auto-select From design. If user provides published statistics, auto-select Calibrated.

When to Use

  • Testing analysis code before real data collection
  • Power analysis via simulation (complements experiment-design Power mode)
  • Method development and debugging estimation pipelines
  • Generating pilot data for grant proposals or ethics applications
  • Teaching demonstrations with realistic-looking data

When NOT to Use

  • Designing the experiment itself --> experiment-design
  • Running analysis on real data --> data-analysis
  • Auditing identification strategy --> causal-design

Workflow

Step 1: Detect Mode

Detect from context or ask:

SignalMode
docs/experiment-design.md existsFrom design
User describes variables, types, relationshipsFrom schema
User provides means, SDs, correlations from a paperCalibrated
AmbiguousAsk
Step 2: Interview for Data Structure

Gather the following (adapt questions to mode):

ParameterQuestionDefault
VariablesWhat variables do you need?—
TypesContinuous, binary, ordinal, categorical?Infer from name
Sample sizeHow many observations?500
TreatmentIs there a treatment variable? How many arms?—
Effect sizeExpected treatment effect (Cohen's d, OR, etc.)?0.3 (small-medium)
CorrelationsWhich variables should be correlated? How strongly?—
ClusteringAre observations nested (e.g., students in classrooms)?No
Panel structureMultiple time periods? How many?Cross-section
Missing dataShould the data include realistic missingness?No
LanguageR or Python?Detect from project or ask

For From design mode, extract most parameters from the design document automatically and confirm with the user.

For Calibrated mode, require the user to provide published summary statistics. Read references/calibration-targets.md for the calibration procedure.

Step 3: Generate Script

Read references/dgp-recipes.md for code patterns matching the requested design.

Read shared/multi-language-conventions.md for language-specific code style.

Generate a self-contained script that:

  1. Sets a random seed (document the seed value)
  2. Defines the data generating process with clear comments
  3. Generates the dataset
  4. Adds realistic noise and distributional features
  5. Introduces missing data patterns if requested
  6. Saves the dataset as CSV (and optionally .rds/.parquet)
  7. Prints a summary of the generated data
Step 4: Execute

For large Monte Carlo sweeps (10k+ simulations, multi-condition grids, or long-running bootstrap): run on [HPC cluster] instead of locally. Drop the generation script into hpc/ with submit.sbatch (compute partition) or sweep.sbatch (array over seeds/conditions) — templates at Task Management templates/slurm/, guide at docs/guides/hpc.md. All SLURM templates log git-sha.txt to OUT_DIR so synthetic datasets remain traceable to the DGP code version.

Run the script and verify:

  • Dataset has the expected dimensions
  • Variable types are correct
  • Treatment/control groups are balanced (if applicable)
  • Summary statistics are plausible
  • No degenerate columns (all zeros, all missing, zero variance)
Step 5: Save Output

Output routing:

FileLocation
Generated datasetdata/synthetic/{name}.csv
Generation scriptcode/generate_synthetic_{name}.R (or .py)
Data dictionarydata/synthetic/{name}_dictionary.md

NEVER write to data/raw/. The data-sensitivity rule applies -- synthetic data is not raw data and must be clearly separated.

Create data/synthetic/ if it does not exist.

Step 6: Data Dictionary

Produce a markdown data dictionary alongside the dataset:

markdown
# Data Dictionary: {name}

Generated: {date}
Script: `code/generate_synthetic_{name}.R`
Seed: {seed}
N: {sample_size}

| Variable | Type | Description | Distribution | Parameters |
|----------|------|-------------|-------------- |------------|
| id | integer | Unique identifier | Sequential | 1 to N |
| treatment | binary | Treatment assignment | Bernoulli | p = 0.5 |
| outcome | continuous | Primary outcome | Normal | mu = 0, sigma = 1 (control) |
| ... | ... | ... | ... | ... |

## Treatment Effects
- True ATE: {value}
- True CATE by subgroup: {if applicable}

## Missing Data
- Mechanism: {MCAR/MAR/MNAR}
- Rate: {percentage} of {variable}

Key Principles

Reproducibility
  • Always set random seeds. Document the seed in the script header and the data dictionary.
  • Script must be fully self-contained -- running it again with the same seed produces identical output.
Realism
  • Match variable types to realistic distributions. Not everything is normal:
    • Income, reaction times --> log-normal or gamma
    • Likert scales --> ordinal with floor/ceiling effects
    • Count data --> Poisson or negative binomial
    • Proportions --> beta
    • Binary outcomes --> Bernoulli with realistic base rates
  • Include realistic noise. Treatment effects should be embedded in noisy data, not clean signal.
  • Add heterogeneity. Real data has heterogeneous treatment effects -- consider adding subgroup variation.
Show full SKILL.md (421 more words)Show less
Missing Data
  • If the real data will have missingness, simulate it:
    • MCAR: Random deletion at a fixed rate
    • MAR: Missingness depends on observed variables (e.g., higher income --> less missing)
    • MNAR: Missingness depends on the missing value itself (e.g., depressed people less likely to respond)
  • Document the missingness mechanism in the data dictionary.
Calibration (Calibrated Mode)
  • Read references/calibration-targets.md for the full procedure.
  • Require published summary statistics as input (means, SDs, correlations, effect sizes).
  • Validate that the synthetic data matches the targets within tolerance.
  • Report discrepancies between target and achieved statistics.

Mode: From Design

Workflow
  1. Locate design document -- check docs/experiment-design.md, docs/pre-analysis-plan.md, log/plans/
  2. Extract parameters:
    • Treatment arms and assignment mechanism
    • Primary and secondary outcomes (types, expected distributions)
    • Covariates and their roles
    • Sample size and any clustering/stratification
    • Expected effect sizes from the power analysis
  3. Confirm with user -- present extracted parameters and ask for adjustments
  4. Generate -- follow Steps 3-6 above
Integration with experiment-design

The design document produced by experiment-design Design mode contains everything needed:

  • Hypotheses with expected signs --> treatment effect directions
  • Conditions and randomization --> treatment assignment DGP
  • Outcome measures and scales --> variable types and distributions
  • Power analysis results --> sample size and effect size

Mode: From Schema

Workflow
  1. Interview -- ask user to describe each variable (or provide a sketch)
  2. Infer structure -- detect relationships from the description:
    • "X causes Y" --> include causal pathway
    • "A and B are correlated" --> induce correlation via shared latent factor
    • "C moderates the effect" --> include interaction term
  3. Generate -- follow Steps 3-6 above
Common Requests
User saysInterpret as
"A standard survey dataset"Demographics + Likert DVs + treatment + attention checks
"Panel data"Units x time with unit and time fixed effects
"Something like [paper name]"Switch to Calibrated mode, use paper's descriptives

Mode: Calibrated

Workflow
  1. Collect targets -- user provides published summary statistics:
    • Means and standard deviations
    • Correlation matrix (or key correlations)
    • Effect sizes from regression tables
    • Sample size
    • Distribution shapes (skew, kurtosis) if reported
  2. Read references/calibration-targets.md for the calibration procedure
  3. Generate with matching -- use the Cholesky decomposition or copula method to match the target correlation structure
  4. Validate -- compare synthetic descriptives to targets, report match quality
  5. Save -- follow Steps 5-6 above, include calibration targets in the data dictionary

Cross-References

ResourceWhen read
references/dgp-recipes.mdAll modes (code patterns for each design type)
references/calibration-targets.mdCalibrated mode (matching procedure)
shared/multi-language-conventions.mdAll modes (code style)
experiment-design skillFrom design mode (reads its design documents)
data-analysis skillConsumes synthetic data for pipeline testing
data-sensitivity ruleNever write to data/raw/
design-before-results ruleSynthetic data supports locking the design before real data

© flonat, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 5 other files (references) in skills/synthetic-data of flonat/flonat-research.

  • SKILL.md
  • references/calibration-targets.md
  • references/dgp-recipes-experimental.md
  • references/dgp-recipes-observational.md
  • references/dgp-recipes-survey-mediation.md
  • references/dgp-recipes.md

Open the folder on GitHubat commit da27600

Compare with similar skills

Synthetic Data next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Synthetic Data compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Synthetic Data this skillflonat/flonat-research145—~2.5kAutomated safety check: PassMIT
Generating Synthea Datamaziyarpanahi/openmed5.5k—~1.6kAutomated safety check: PassApache-2.0
Slate Ar Qualityudecode/plate17k—~474Automated safety check: PassCustom licence
Pytest Patternscohen-liel/hivemind110—~806Automated safety check: PassApache-2.0
Requirementsrizsotto/Bear6.5k—~2kAutomated safety check: PassGPL-3.0
Crap Analysisardalis/RiverBooks1352 repos~3.4kAutomated safety check: PassNone

Similar skills

  • Generating Synthea Data

    maziyarpanahi/openmed

    Generates synthetic but realistic patient records (FHIR R4 bundles, C-CDA documents, CSV) with MITRE Synthea for development, CI fixtures, demos, and leakage-gate test sets — zero real PHI.

    5.5k GitHub stars~1.6k tokensUpdated yesterday
    Testing & QAAuto-check passed
  • Slate Ar Quality

    udecode/plate

    Slate v2 quality-gap Autoresearch shortcut. An agent skill from udecode/plate.

    17k GitHub stars~474 tokensUpdated yesterday
    Testing & QAAuto-check passed
  • Pytest Patterns

    cohen-liel/hivemind

    pytest best practices for writing comprehensive test suites.

    110 GitHub stars~806 tokensUpdated 5 mo ago
    Testing & QAAuto-check passed
  • Requirements

    rizsotto/Bear

    Write, modify, or review a requirement file under docs/requirements -- pick the single owning file, keep the text contract-only, name IDs so they need no explanation, and verify cross-references and…

    6.5k GitHub stars~2k tokensUpdated yesterday
    Testing & QAAuto-check passed
  • Crap Analysis

    ardalis/RiverBooks

    Analyze code coverage and CRAP (Change Risk Anti-Patterns) scores to identify high-risk code.

    135 GitHub starsUsed in 2 repos~3.4k tokens
    Testing & QAAuto-check passed
  • Fs Fixture

    privatenumber/fs-fixture

    Create disposable file system test fixtures from objects, templates, or empty directories with automatic cleanup.

    100 GitHub stars~1.2k tokensUpdated 1 mo ago
    Testing & QAAuto-check passed

More from flonat/flonat-research

All 83 skills in this repo
  • Latex Posters

    flonat/flonat-research

    Create a large-format academic poster in LaTeX using beamerposter, tikzposter, or baposter.

    145 GitHub stars~1.5k tokensUpdated 9 days ago
    Auto-check: notes
  • Skill Creator

    flonat/flonat-research

    Create, revise, and evaluate reusable AI workflow skills, including trigger-quality tests.

    145 GitHub stars~4.4k tokensUpdated 9 days ago
    Auto-check passed
  • DOCX

    flonat/flonat-research

    Create, read, edit, or convert Microsoft Word documents while preserving professional document structure.

    145 GitHub stars~1.2k tokensUpdated 9 days ago
    Auto-check passed
  • PDF

    flonat/flonat-research

    Read, create, combine, split, rotate, OCR, watermark, secure, or extract content from PDF files.

    145 GitHub stars~488 tokensUpdated 9 days ago
    Auto-check passed
  • Init Project Orchestration

    flonat/flonat-research

    Create or migrate project-level agents, repeatable project workflows, and planning state from one client-neutral contract, then render repository-scoped adapters for both Claude Code and Codex.

    145 GitHub stars~1.6k tokensUpdated 9 days ago
    Auto-check passed
  • Pre Commit Audit

    flonat/flonat-research

    Deliver a fast pre-commit safety scan: file size, anonymity (author / affiliation strings in tex/bib), hardcoded secrets, and invisible-Unicode carriers.

    145 GitHub stars~2.8k tokensUpdated 9 days ago
    Auto-check: notes

Questions about Synthetic Data

What does Synthetic Data do?

Generate structurally realistic synthetic datasets for pipeline prototyping, test coverage, or prospective power analysis. Synthetic Data is an agent skill from flonat/flonat-research. Generate structurally realistic synthetic datasets for pipeline prototyping, test coverage, or prospective power analysis.

When should I use Synthetic Data?

Synthetic Data fits situations like: design must be exercised before real data are available; tasks that involve Test data and fixtures; tasks that involve Experimental design.

How do I install Synthetic Data in Claude Code?

Run `npx skills add flonat/flonat-research --skill synthetic-data -a claude-code`. Or copy the skill folder (skills/synthetic-data in flonat/flonat-research) into .claude/skills/synthetic-data in your project. Claude Code loads it when a task matches its description.

How do I install Synthetic Data in Codex?

Run `npx skills add flonat/flonat-research --skill synthetic-data -a codex`. Or copy the skill folder (skills/synthetic-data in flonat/flonat-research) into .agents/skills/synthetic-data in your project. Codex loads it when a task matches its description.

Can I use Synthetic Data in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add flonat/flonat-research --skill synthetic-data -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/synthetic-data, .gemini/skills/synthetic-data, .github/skills/synthetic-data and .opencode/skills/synthetic-data in your project.

What does Synthetic Data need to run?

SKILL.md names no scripts, command-line tools or credentials: Synthetic Data is instructions for the agent only. Its frontmatter pre-approves these tools: Bash(uv*, Rscript*, R*, mkdir*, ls*), Read, Write, Edit, Glob, Grep, AskUserQuestion.

Does Synthetic Data access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Synthetic Data safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Synthetic Data use?

Synthetic Data is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Synthetic Data use?

About 2.5k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 7.8k tokens, read only when the agent opens those files.

What are the alternatives to Synthetic Data?

Skills that share tags, products or a category with Synthetic Data: Generating Synthea Data (maziyarpanahi/openmed, 5.5k stars), Slate Ar Quality (udecode/plate, 17k stars), Pytest Patterns (cohen-liel/hivemind, 110 stars) and Requirements (rizsotto/Bear, 6.5k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Synthetic Data?

flonat (a GitHub user) maintains it in flonat/flonat-research, which has 145 GitHub stars. The repository holds 83 skills in this directory. The repository was last updated on September 29, 2026.

Source: flonat/flonat-research on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.