Spark Python Data Source is an agent skill from databricks/databricks-agent-skills, published by the product's own GitHub organization. Build custom Python data sources for Apache Spark using the PySpark DataSource API — batch and streaming readers/writers for external systems. Use this skill whenever someone wants to connect Spark to an external system (database, API, message queue, custom protocol), build a Spark connector or plugin in Python, implement a DataSourceReader or DataSourceWriter, pull data from or push data to a system via Spark, or work with the PySpark DataSource API in any way. Even if they just say "read from X in Spark" or…
Its SKILL.md is about 2.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 14 other files, including reference files and assets (for example `agents/openai.yaml`, `references/authentication-patterns.md` and `references/error-handling.md`). Compatibility notes: Requires databricks CLI (= v1.0.0)
It sits in Data & Analytics, covering Event-driven systems and DataFrames. It works with Python and Apache Spark. The repository describes itself as: Databricks AI Tools: skills and plugins for building on Databricks with Claude Code, Cursor, Codex, GitHub Copilot, and other AI coding agents.