Official agent skill

Spark Python Data Source

by databricks in databricks/databricks-agent-skills

Build custom Python data sources for Apache Spark using the PySpark DataSource API — batch and streaming readers/writers for external systems.

OfficialCustom licenceAuto-check passedData & Analytics

Install Spark Python Data Source

skills CLI
$ npx skills add databricks/databricks-agent-skills --skill spark-python-data-source -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install databricks/databricks-agent-skills spark-python-data-source --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/databricks/databricks-agent-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/experimental/spark-python-data-source .claude/skills/spark-python-data-source && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
spark-python-data-source
GitHub stars
345
Token cost
~2.1k tokens
SKILL.md length
607 words
Files
12 (incl. references, assets)
Skills in repo
32
Repo updated
First seen
Licence
Custom licence

At a glance

Build custom Python data sources for Apache Spark using the PySpark DataSource API — batch and streaming readers/writers for external systems.

  • Works in 4 steps: DataSource class — entry point that… → Base Reader/Writer classes — shared… → Batch classes — inherit from base +… → …
  • Someone wants to connect Spark to an external system (database
  • SKILL.md covers Instructions, Example Prompts, Related and References
  • Calls uv

What it does

Spark Python Data Source is an agent skill from databricks/databricks-agent-skills, published by the product's own GitHub organization. Build custom Python data sources for Apache Spark using the PySpark DataSource API — batch and streaming readers/writers for external systems. Use this skill whenever someone wants to connect Spark to an external system (database, API, message queue, custom protocol), build a Spark connector or plugin in Python, implement a DataSourceReader or DataSourceWriter, pull data from or push data to a system via Spark, or work with the PySpark DataSource API in any way. Even if they just say "read from X in Spark" or…

Its SKILL.md is about 2.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 14 other files, including reference files and assets (for example `agents/openai.yaml`, `references/authentication-patterns.md` and `references/error-handling.md`). Compatibility notes: Requires databricks CLI (= v1.0.0)

It sits in Data & Analytics, covering Event-driven systems and DataFrames. It works with Python and Apache Spark. The repository describes itself as: Databricks AI Tools: skills and plugins for building on Databricks with Claude Code, Cursor, Codex, GitHub Copilot, and other AI coding agents.

When your agent uses it

  • Someone wants to connect Spark to an external system (database
  • Custom protocol)
  • Build a Spark connector
  • Plugin in Python

Example prompts

  • “read from X in Spark”
  • “write DataFrame to Y”
  • “/spark-python-data-source”

Requirements

  • Python 3
  • Compatibility (from SKILL.md): Requires databricks CLI (>= v1.0.0)

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. DataSource class — entry point that returns readers/writers
  2. Base Reader/Writer classes — shared logic for options and data processing
  3. Batch classes — inherit from base + DataSourceReader/DataSourceWriter
  4. Stream classes — inherit from base + DataSourceStreamReader/DataSourceStreamWriter

What it can do on your machine

Read from SKILL.md and the folder at commit f4fcec5. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • uv

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • github.com
    • docs.databricks.com
    • spark.apache.org

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Requires databricks CLI (>= v1.0.0)

    From compatibility in the SKILL.md frontmatter.

Context cost

Spark Python Data Source loads about 2.1k tokens when it runs, and up to ~24k if it reads all its reference files. Until then it costs about 154 tokens; SKILL.md has 607 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~154
When it runs · the whole SKILL.md, loaded when a task matches
~2.1k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~24k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

Its licence (Custom licence) doesn't allow us to republish the file, so here is its outline and opening line. It has 607 words (~2,068 tokens).

“Build custom Python data sources for Apache Spark 4.0+ to read from and write to external systems in batch and streaming modes.”

— opening of SKILL.md by databricks, Custom licence
name
spark-python-data-source
compatibility
Requires databricks CLI (>= v1.0.0)
metadata.version
0.1.0

Read the full SKILL.md on GitHub

Files

SKILL.md and 11 other files (references, assets) in experimental/spark-python-data-source of databricks/databricks-agent-skills.

  • SKILL.md
  • agents/openai.yaml
  • assets/databricks.png
  • assets/databricks.svg
  • references/authentication-patterns.md
  • references/error-handling.md
  • references/implementation-template.md
  • references/partitioning-patterns.md
  • references/production-patterns.md
  • references/streaming-patterns.md
  • references/testing-patterns.md
  • references/type-conversion.md

Open the folder on GitHubat commit f4fcec5

Compare with similar skills

Spark Python Data Source next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Spark Python Data Source compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Spark Python Data Source this skilldatabricks/databricks-agent-skills345—~2.1kAutomated safety check: PassCustom licence
Flowfile Frame And CodegenEdwardvaneechoud/Flowfile370—~12kAutomated safety check: PassMIT
Datafusion Pythonapache/datafusion-python606—~7.8kAutomated safety check: PassApache-2.0
Neo4j Spark Skillneo4j-contrib/neo4j-skills114—~4.1kAutomated safety check: NotesMIT
Spark EngineerFerroxLabs/wayland608—~4.2kAutomated safety check: PassApache-2.0
Polar Python SDKpolarsource/polar10k—~1.8kAutomated safety check: PassApache-2.0

Similar skills

  • Flowfile Frame And Codegen

    Edwardvaneechoud/Flowfile

    Deep dive into flowfileframe — the Polars-LazyFrame-shaped Python API that builds an in-process flowfilecore FlowGraph as a side effect of every method call — covering the FlowFrame/Expr internals…

    370 GitHub stars~12k tokensUpdated yesterday
    Data & AnalyticsAuto-check passed
  • Datafusion Python

    apache/datafusion-python

    A skill your agent uses when the user is writing datafusion-python (Apache DataFusion Python bindings) DataFrame or SQL code.

    606 GitHub stars~7.8k tokensUpdated today
    Data & AnalyticsAuto-check passed
  • Neo4j Spark Skill

    neo4j-contrib/neo4j-skills

    A skill your agent uses when reading from or writing to Neo4j with Apache Spark or Databricks using the Neo4j Connector for Apache Spark 6.0 (org.neo4j.connectors:spark) or 5.x…

    114 GitHub stars~4.1k tokensUpdated today
    Data & AnalyticsAuto-check: notes
  • Spark Engineer

    FerroxLabs/wayland

    Apache Spark expertise covering RDD vs DataFrame vs Dataset APIs, partitioning strategies, shuffle optimization, broadcast joins, caching, Spark SQL, structured streaming, UDFs, cluster sizing…

    608 GitHub stars~4.2k tokensUpdated yesterday
    Data & AnalyticsAuto-check passed
  • Polar Python SDK

    polarsource/polar

    Integrate Polar billing in server-side Python applications using the versioned Polar and PolarAsync clients.

    10k GitHub stars~1.8k tokensUpdated today
    Data & AnalyticsAuto-check passed
  • Speed up slow Apache Spark jobs by tuning partitions, shuffles, data skew, caching and executor memory, with do and don't rules for PySpark code.

    40k GitHub starsUsed in 8 repos~789 tokens
    Data & AnalyticsAuto-check passed

More from databricks/databricks-agent-skills

All 32 skills in this repo
  • Databricks Dbsql

    databricks/databricks-agent-skills

    Official

    Databricks SQL (DBSQL) advanced features and SQL warehouse capabilities.

    345 GitHub starsUsed in 1 repo~2.8k tokens
    Auto-check passed
  • Databricks Synthetic Data Gen

    databricks/databricks-agent-skills

    Official

    Generate realistic synthetic data using Spark + Faker (strongly recommended).

    345 GitHub stars~3.2k tokensUpdated yesterday
    Auto-check passed
  • Databricks Model Serving

    databricks/databricks-agent-skills

    Official

    Databricks Model Serving endpoint lifecycle and ops. An agent skill from databricks/databricks-agent-skills.

    345 GitHub starsUsed in 1 repo~3.3k tokens
    Auto-check passed
  • Databricks Python SDK

    databricks/databricks-agent-skills

    Official

    Databricks development guidance including Python SDK, Databricks Connect, CLI, and REST API.

    345 GitHub starsUsed in 1 repo~4.6k tokens
    Auto-check passed
  • Databricks Spark Structured Streaming

    databricks/databricks-agent-skills

    Official

    Comprehensive guide to Spark Structured Streaming for production workloads.

    345 GitHub starsUsed in 1 repo~956 tokens
    Auto-check passed
  • Databricks App Design

    databricks/databricks-agent-skills

    Official

    Design the UX of custom-code Databricks Apps (AppKit/React) data screens — KPI/overview pages, reports, charts, tables, and Genie/chat data assistants — mapped to concrete AppKit components.

    345 GitHub stars~1.7k tokensUpdated yesterday
    Auto-check passed

Questions about Spark Python Data Source

What does Spark Python Data Source do?

Build custom Python data sources for Apache Spark using the PySpark DataSource API — batch and streaming readers/writers for external systems. Spark Python Data Source is an agent skill from databricks/databricks-agent-skills, published by the product's own GitHub organization. Build custom Python data sources for Apache Spark using the PySpark DataSource API — batch and streaming readers/writers for external systems.

When should I use Spark Python Data Source?

Spark Python Data Source fits situations like: someone wants to connect Spark to an external system (database; custom protocol); build a Spark connector; plugin in Python.

How do I install Spark Python Data Source in Claude Code?

Run `npx skills add databricks/databricks-agent-skills --skill spark-python-data-source -a claude-code`. Or copy the skill folder (experimental/spark-python-data-source in databricks/databricks-agent-skills) into .claude/skills/spark-python-data-source in your project. Claude Code loads it when a task matches its description.

How do I install Spark Python Data Source in Codex?

Run `npx skills add databricks/databricks-agent-skills --skill spark-python-data-source -a codex`. Or copy the skill folder (experimental/spark-python-data-source in databricks/databricks-agent-skills) into .agents/skills/spark-python-data-source in your project. Codex loads it when a task matches its description.

Can I use Spark Python Data Source in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add databricks/databricks-agent-skills --skill spark-python-data-source -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/spark-python-data-source, .gemini/skills/spark-python-data-source, .github/skills/spark-python-data-source and .opencode/skills/spark-python-data-source in your project.

What does Spark Python Data Source need to run?

Going by SKILL.md and its folder, Spark Python Data Source needs the command-line tools its instructions call (uv). Our summary lists: Python 3. Compatibility (from SKILL.md): Requires databricks CLI (>= v1.0.0).

Does Spark Python Data Source access the network?

SKILL.md names 3 domains. As links in the text: github.com, docs.databricks.com and spark.apache.org. This is read from the text; nothing was executed.

Is Spark Python Data Source safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Spark Python Data Source use?

Spark Python Data Source has a licence file (the repository's licence) that doesn't match a standard licence. Read it on GitHub before reusing the skill.

How many tokens does Spark Python Data Source use?

About 2.1k tokens (SKILL.md is roughly 8.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 21k tokens, read only when the agent opens those files.

What are the alternatives to Spark Python Data Source?

Skills that share tags, products or a category with Spark Python Data Source: Flowfile Frame And Codegen (Edwardvaneechoud/Flowfile, 370 stars), Datafusion Python (apache/datafusion-python, 606 stars), Neo4j Spark Skill (neo4j-contrib/neo4j-skills, 114 stars) and Spark Engineer (FerroxLabs/wayland, 608 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Spark Python Data Source?

databricks (a GitHub organization, an official publisher) maintains it in databricks/databricks-agent-skills, which has 345 GitHub stars. The repository holds 32 skills in this directory. The repository was last updated on October 6, 2026.

Source: databricks/databricks-agent-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.