Official agent skill

Databricks Synthetic Data Gen

by databricks in databricks/databricks-agent-skills

Generate realistic synthetic data using Spark + Faker (strongly recommended).

OfficialCustom licenceAuto-check passedTesting & QA

Install Databricks Synthetic Data Gen

skills CLI
$ npx skills add databricks/databricks-agent-skills --skill databricks-synthetic-data-gen -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install databricks/databricks-agent-skills databricks-synthetic-data-gen --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/databricks/databricks-agent-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/databricks-synthetic-data-gen .claude/skills/databricks-synthetic-data-gen && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
databricks-synthetic-data-gen
GitHub stars
345
Token cost
~3.2k tokens
SKILL.md length
1,190 words
Files
7 (incl. scripts, references, assets)
Skills in repo
32
Repo updated
First seen
Licence
Custom licence

At a glance

Generate realistic synthetic data using Spark + Faker (strongly recommended).

  • Works in 3 steps: Gather Requirements → Present Plan with Story → Ask About Data Features
  • User mentions synthetic data
  • SKILL.md covers Data Must Tell a Business Story, References, Critical Rules and Generation Planning Workflow, plus 6 more sections
  • Runs Python scripts from its folder; calls databricks and uv

What it does

Databricks Synthetic Data Gen is an agent skill from databricks/databricks-agent-skills, published by the product's own GitHub organization. Generate realistic synthetic data using Spark + Faker (strongly recommended). Supports serverless execution, multiple output formats (Parquet/JSON/CSV/Delta), and scales from thousands to millions of rows. For small datasets (<10K rows), can optionally generate locally and upload to volumes. Use when user mentions 'synthetic data', 'test data', 'generate data', 'demo dataset', 'Faker', or 'sample data'.

Its SKILL.md is about 3.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 10 other files, including scripts, reference files and assets (for example `agents/openai.yaml`, `references/1-data-patterns.md` and `references/2-troubleshooting.md`). Compatibility notes: Requires databricks CLI (= v1.0.0)

It sits in Testing & QA, covering Test data and fixtures. It works with Databricks. The repository describes itself as: Databricks AI Tools: skills and plugins for building on Databricks with Claude Code, Cursor, Codex, GitHub Copilot, and other AI coding agents.

When your agent uses it

  • User mentions synthetic data
  • Tasks that involve Test data and fixtures

Example prompts

  • “synthetic data”
  • “test data”
  • “generate data”
  • “/databricks-synthetic-data-gen”

Requirements

  • Python 3
  • Compatibility (from SKILL.md): Requires databricks CLI (>= v1.0.0)

Workflow steps

3 steps, taken from the step headings in SKILL.md.

  1. Gather Requirements
  2. Present Plan with Story
  3. Ask About Data Features

What it can do on your machine

Read from SKILL.md and the folder at commit f4fcec5. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • databricks
    • uv

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use uv, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Requires databricks CLI (>= v1.0.0)

    From compatibility in the SKILL.md frontmatter.

Context cost

Databricks Synthetic Data Gen loads about 3.2k tokens when it runs, and up to ~7.1k if it reads all its reference files. Until then it costs about 109 tokens; SKILL.md has 1,190 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~109
When it runs · the whole SKILL.md, loaded when a task matches
~3.2k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~7.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

Its licence (Custom licence) doesn't allow us to republish the file, so here is its outline and opening line. It has 1,190 words (~3,202 tokens).

“Generate realistic, story-driven synthetic data for Databricks using Spark + Faker + Pandas UDFs (strongly recommended).”

— opening of SKILL.md by databricks, Custom licence
name
databricks-synthetic-data-gen
compatibility
Requires databricks CLI (>= v1.0.0)
metadata.version
0.1.0
parent
databricks-core

Read the full SKILL.md on GitHub

Files

SKILL.md and 6 other files (scripts, references, assets) in skills/databricks-synthetic-data-gen of databricks/databricks-agent-skills.

  • SKILL.md
  • agents/openai.yaml
  • assets/databricks.png
  • assets/databricks.svg
  • references/1-data-patterns.md
  • references/2-troubleshooting.md
  • scripts/generate_synthetic_data.py

Open the folder on GitHubat commit f4fcec5

Compare with similar skills

Databricks Synthetic Data Gen next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Databricks Synthetic Data Gen compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Databricks Synthetic Data Gen this skilldatabricks/databricks-agent-skills345—~3.2kAutomated safety check: PassCustom licence
Azure Data Manager For AgriMicrosoftDocs/Agent-Skills775—~1.6kAutomated safety check: PassCC-BY-4.0
Azure Data FactoryMicrosoftDocs/Agent-Skills7751 repos~16kAutomated safety check: PassCC-BY-4.0
Azure Synapse AnalyticsMicrosoftDocs/Agent-Skills7751 repos~13kAutomated safety check: PassCC-BY-4.0
Azure DatabricksMicrosoftDocs/Agent-Skills7751 repos~14kAutomated safety check: PassCC-BY-4.0
Azure Energy Data ServicesMicrosoftDocs/Agent-Skills775—~2.5kAutomated safety check: PassCC-BY-4.0

Similar skills

  • Azure Data Manager For Agri

    MicrosoftDocs/Agent-Skills

    Official

    Expert knowledge for Azure Data Manager for Agriculture development including limits & quotas, security, configuration, and integrations & coding patterns.

    775 GitHub stars~1.6k tokensUpdated 2 days ago
    Testing & QAAuto-check passed
  • Azure Data Factory

    MicrosoftDocs/Agent-Skills

    Official

    Expert knowledge for Azure Data Factory development including troubleshooting, best practices, decision making, architecture & design patterns, limits & quotas, security, configuration, integrations…

    775 GitHub starsUsed in 1 repo~16k tokens
    DevOps & CloudAuto-check passed
  • Azure Synapse Analytics

    MicrosoftDocs/Agent-Skills

    Official

    Expert knowledge for Azure Synapse Analytics development including troubleshooting, best practices, decision making, architecture & design patterns, limits & quotas, security, configuration…

    775 GitHub starsUsed in 1 repo~13k tokens
    DevelopmentAuto-check passed
  • Azure Databricks

    MicrosoftDocs/Agent-Skills

    Official

    Expert knowledge for Azure Databricks development including troubleshooting, best practices, decision making, architecture & design patterns, limits & quotas, security, configuration, integrations &…

    775 GitHub starsUsed in 1 repo~14k tokens
    DevelopmentAuto-check passed
  • Azure Energy Data Services

    MicrosoftDocs/Agent-Skills

    Official

    Expert knowledge for Azure Energy Data Services development including troubleshooting, decision making, architecture & design patterns, security, configuration, integrations & coding patterns, and…

    775 GitHub stars~2.5k tokensUpdated 2 days ago
    DevelopmentAuto-check passed
  • Jest Testing Patterns

    ChrisWiles/claude-code-showcase

    Jest patterns for React Native style tests: TDD discipline, mock factory functions, module and GraphQL hook mocking, custom render helpers and anti-patterns to avoid.

    6.1k GitHub starsUsed in 7 repos~1.5k tokens
    Testing & QAAuto-check passed

More from databricks/databricks-agent-skills

All 32 skills in this repo
  • Databricks Dbsql

    databricks/databricks-agent-skills

    Official

    Databricks SQL (DBSQL) advanced features and SQL warehouse capabilities.

    345 GitHub starsUsed in 1 repo~2.8k tokens
    Auto-check passed
  • Databricks Model Serving

    databricks/databricks-agent-skills

    Official

    Databricks Model Serving endpoint lifecycle and ops. An agent skill from databricks/databricks-agent-skills.

    345 GitHub starsUsed in 1 repo~3.3k tokens
    Auto-check passed
  • Databricks Python SDK

    databricks/databricks-agent-skills

    Official

    Databricks development guidance including Python SDK, Databricks Connect, CLI, and REST API.

    345 GitHub starsUsed in 1 repo~4.6k tokens
    Auto-check passed
  • Databricks Spark Structured Streaming

    databricks/databricks-agent-skills

    Official

    Comprehensive guide to Spark Structured Streaming for production workloads.

    345 GitHub starsUsed in 1 repo~956 tokens
    Auto-check passed
  • Databricks App Design

    databricks/databricks-agent-skills

    Official

    Design the UX of custom-code Databricks Apps (AppKit/React) data screens — KPI/overview pages, reports, charts, tables, and Genie/chat data assistants — mapped to concrete AppKit components.

    345 GitHub stars~1.7k tokensUpdated yesterday
    Auto-check passed
  • Databricks Apps Python

    databricks/databricks-agent-skills

    Official

    Python backend for Databricks Apps — FastAPI (default), Flask, Dash, Streamlit, Gradio, Reflex.

    345 GitHub stars~3k tokensUpdated yesterday
    Auto-check passed

Works with

Categories

Questions about Databricks Synthetic Data Gen

What does Databricks Synthetic Data Gen do?

Generate realistic synthetic data using Spark + Faker (strongly recommended). Databricks Synthetic Data Gen is an agent skill from databricks/databricks-agent-skills, published by the product's own GitHub organization. Generate realistic synthetic data using Spark + Faker (strongly recommended).

When should I use Databricks Synthetic Data Gen?

Databricks Synthetic Data Gen fits situations like: user mentions synthetic data; tasks that involve Test data and fixtures.

How do I install Databricks Synthetic Data Gen in Claude Code?

Run `npx skills add databricks/databricks-agent-skills --skill databricks-synthetic-data-gen -a claude-code`. Or copy the skill folder (skills/databricks-synthetic-data-gen in databricks/databricks-agent-skills) into .claude/skills/databricks-synthetic-data-gen in your project. Claude Code loads it when a task matches its description.

How do I install Databricks Synthetic Data Gen in Codex?

Run `npx skills add databricks/databricks-agent-skills --skill databricks-synthetic-data-gen -a codex`. Or copy the skill folder (skills/databricks-synthetic-data-gen in databricks/databricks-agent-skills) into .agents/skills/databricks-synthetic-data-gen in your project. Codex loads it when a task matches its description.

Can I use Databricks Synthetic Data Gen in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add databricks/databricks-agent-skills --skill databricks-synthetic-data-gen -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/databricks-synthetic-data-gen, .gemini/skills/databricks-synthetic-data-gen, .github/skills/databricks-synthetic-data-gen and .opencode/skills/databricks-synthetic-data-gen in your project.

What does Databricks Synthetic Data Gen need to run?

Going by SKILL.md and its folder, Databricks Synthetic Data Gen needs Python for the scripts in its folder and the command-line tools its instructions call (databricks and uv). Our summary lists: Python 3. Compatibility (from SKILL.md): Requires databricks CLI (>= v1.0.0).

Does Databricks Synthetic Data Gen access the network?

SKILL.md contains no URLs. Its commands use uv, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Databricks Synthetic Data Gen safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Databricks Synthetic Data Gen use?

Databricks Synthetic Data Gen has a licence file (the repository's licence) that doesn't match a standard licence. Read it on GitHub before reusing the skill.

How many tokens does Databricks Synthetic Data Gen use?

About 3.2k tokens (SKILL.md is roughly 13k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3.9k tokens, read only when the agent opens those files.

What are the alternatives to Databricks Synthetic Data Gen?

Skills that share tags, products or a category with Databricks Synthetic Data Gen: Azure Data Manager For Agri (MicrosoftDocs/Agent-Skills, 775 stars), Azure Data Factory (MicrosoftDocs/Agent-Skills, 775 stars), Azure Synapse Analytics (MicrosoftDocs/Agent-Skills, 775 stars) and Azure Databricks (MicrosoftDocs/Agent-Skills, 775 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Databricks Synthetic Data Gen?

databricks (a GitHub organization, an official publisher) maintains it in databricks/databricks-agent-skills, which has 345 GitHub stars. The repository holds 32 skills in this directory. The repository was last updated on October 6, 2026.

Source: databricks/databricks-agent-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.