Official agent skill

Databricks Spark Structured Streaming

by databricks in databricks/databricks-agent-skills

Comprehensive guide to Spark Structured Streaming for production workloads.

OfficialCustom licenceAuto-check passedBackend & APIs

Install Databricks Spark Structured Streaming

skills CLI
$ npx skills add databricks/databricks-agent-skills --skill databricks-spark-structured-streaming -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install databricks/databricks-agent-skills databricks-spark-structured-streaming --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/databricks/databricks-agent-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/databricks-spark-structured-streaming .claude/skills/databricks-spark-structured-streaming && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
databricks-spark-structured-streaming
GitHub stars
345
Used in
1 other repo
Token cost
~956 tokens
SKILL.md length
202 words
Files
15 (incl. references, assets)
Skills in repo
32
Repo updated
First seen
Licence
Custom licence

At a glance

Comprehensive guide to Spark Structured Streaming for production workloads.

  • Building streaming pipelines
  • SKILL.md covers Quick Start, Core Patterns, Configuration and Best Practices, plus 1 more section
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • Working with Kafka ingestion

What it does

Databricks Spark Structured Streaming is an agent skill from databricks/databricks-agent-skills, published by the product's own GitHub organization. Comprehensive guide to Spark Structured Streaming for production workloads. Use when building streaming pipelines, working with Kafka ingestion, implementing Real-Time Mode (RTM), configuring triggers (processingTime, availableNow), handling stateful operations with watermarks, optimizing checkpoints, performing stream-stream or stream-static joins, writing to multiple sinks, or tuning streaming cost and performance.

Its SKILL.md is about 960 tokens, which your agent loads only when the skill is triggered. The skill folder holds 17 other files, including reference files and assets (for example `agents/openai.yaml`, `references/checkpoint-best-practices.md` and `references/kafka-streaming.md`). Compatibility notes: Requires databricks CLI (= v1.0.0)

It sits in Backend & APIs, covering Event-driven systems. It works with Databricks and Apache Kafka. The repository describes itself as: Databricks AI Tools: skills and plugins for building on Databricks with Claude Code, Cursor, Codex, GitHub Copilot, and other AI coding agents.

When your agent uses it

  • Building streaming pipelines
  • Working with Kafka ingestion
  • Implementing Real-Time Mode (RTM)
  • Configuring triggers (processingTime

Example prompts

  • “/databricks-spark-structured-streaming”

Requirements

  • Python 3
  • Compatibility (from SKILL.md): Requires databricks CLI (>= v1.0.0)

What it can do on your machine

Read from SKILL.md and the folder at commit f4fcec5. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Requires databricks CLI (>= v1.0.0)

    From compatibility in the SKILL.md frontmatter.

Context cost

Databricks Spark Structured Streaming loads about 956 tokens when it runs, and up to ~40k if it reads all its reference files. Until then it costs about 115 tokens; SKILL.md has 202 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~115
When it runs · the whole SKILL.md, loaded when a task matches
~956
With references · SKILL.md plus every file in references/, read only if the agent opens them
~40k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

Its licence (Custom licence) doesn't allow us to republish the file, so here is its outline and opening line. It has 202 words (~956 tokens).

“Production-ready streaming pipelines with Spark Structured Streaming. This skill provides navigation to detailed patterns and best practices.”

— opening of SKILL.md by databricks, Custom licence
name
databricks-spark-structured-streaming
compatibility
Requires databricks CLI (>= v1.0.0)
metadata.version
0.1.0
parent
databricks-core

Read the full SKILL.md on GitHub

Files

SKILL.md and 14 other files (references, assets) in skills/databricks-spark-structured-streaming of databricks/databricks-agent-skills.

  • SKILL.md
  • agents/openai.yaml
  • assets/databricks.png
  • assets/databricks.svg
  • references/checkpoint-best-practices.md
  • references/kafka-streaming.md
  • references/lakebase-sink-python.md
  • references/merge-operations.md
  • references/multi-sink-writes.md
  • references/real-time-mode.md
  • references/stateful-operations.md
  • references/stream-static-joins.md
  • references/stream-stream-joins.md
  • references/streaming-best-practices.md
  • references/trigger-and-cost-optimization.md

Open the folder on GitHubat commit f4fcec5

Used in 1 other repository

We found 5 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in databricks/databricks-agent-skills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Databricks Spark Structured Streaming next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Databricks Spark Structured Streaming compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Databricks Spark Structured Streaming this skilldatabricks/databricks-agent-skills3451 repos~956Automated safety check: PassCustom licence
Windmill Trigger Type Checklistwindmill-labs/windmill18k—~4.7kAutomated safety check: PassCustom licence
FoundatioFoundatioFx/Foundatio2.1k—~3.9kAutomated safety check: PassApache-2.0
Opensource Guide Coachcalf-ai/calfkit-sdk1491 repos~2.1kAutomated safety check: PassApache-2.0
Create Environmentgodatadriven/whirl205—~1.9kAutomated safety check: PassApache-2.0
Monstermq Graphql Configvogler75/monster-mq143—~2.3kAutomated safety check: PassGPL-3.0

Similar skills

  • Windmill Trigger Type Checklist

    windmill-labs/windmill

    Checklist of every backend, frontend, CLI and capture change needed to add a new TriggerCrud-based trigger type, such as Azure, GCP or Kafka, to Windmill.

    18k GitHub stars~4.7k tokensUpdated today
    Backend & APIsAuto-check passed
  • Foundatio

    FoundatioFx/Foundatio

    A skill your agent uses when working with Foundatio infrastructure abstractions for .NET -- caching, queuing, messaging, file storage, distributed locking, or background jobs.

    2.1k GitHub stars~3.9k tokensUpdated today
    Backend & APIsAuto-check passed
  • Opensource Guide Coach

    calf-ai/calfkit-sdk

    A skill your agent uses when a user wants guidance on starting, contributing to, growing, governing, funding, securing, or sustaining an open source project, or asks about contributor onboarding…

    149 GitHub starsUsed in 1 repo~2.1k tokens
    Backend & APIsAuto-check passed
  • Create Environment

    godatadriven/whirl

    Create a new Whirl environment in the envs/ directory. An agent skill from godatadriven/whirl.

    205 GitHub stars~1.9k tokensUpdated 6 days ago
    Backend & APIsAuto-check passed
  • Monstermq Graphql Config

    vogler75/monster-mq

    Guide for configuring, managing, and mutating MonsterMQ settings, devices, flows, AI agents, users, loggers, archive groups, topic schemas, and publishing messages via the GraphQL API.

    143 GitHub stars~2.3k tokensUpdated yesterday
    Backend & APIsAuto-check passed
  • Event Store Design

    wshobson/agents

    Designs event stores for event-sourced systems: requirements, a comparison of EventStoreDB, PostgreSQL, Kafka, DynamoDB and Marten, and stream and versioning practices.

    40k GitHub starsUsed in 9 repos~828 tokens
    Backend & APIsAuto-check passed

More from databricks/databricks-agent-skills

All 32 skills in this repo
  • Databricks Dbsql

    databricks/databricks-agent-skills

    Official

    Databricks SQL (DBSQL) advanced features and SQL warehouse capabilities.

    345 GitHub starsUsed in 1 repo~2.8k tokens
    Auto-check passed
  • Databricks Synthetic Data Gen

    databricks/databricks-agent-skills

    Official

    Generate realistic synthetic data using Spark + Faker (strongly recommended).

    345 GitHub stars~3.2k tokensUpdated today
    Auto-check passed
  • Databricks Model Serving

    databricks/databricks-agent-skills

    Official

    Databricks Model Serving endpoint lifecycle and ops. An agent skill from databricks/databricks-agent-skills.

    345 GitHub starsUsed in 1 repo~3.3k tokens
    Auto-check passed
  • Databricks Python SDK

    databricks/databricks-agent-skills

    Official

    Databricks development guidance including Python SDK, Databricks Connect, CLI, and REST API.

    345 GitHub starsUsed in 1 repo~4.6k tokens
    Auto-check passed
  • Databricks App Design

    databricks/databricks-agent-skills

    Official

    Design the UX of custom-code Databricks Apps (AppKit/React) data screens — KPI/overview pages, reports, charts, tables, and Genie/chat data assistants — mapped to concrete AppKit components.

    345 GitHub stars~1.7k tokensUpdated today
    Auto-check passed
  • Databricks Apps Python

    databricks/databricks-agent-skills

    Official

    Python backend for Databricks Apps — FastAPI (default), Flask, Dash, Streamlit, Gradio, Reflex.

    345 GitHub stars~3k tokensUpdated today
    Auto-check passed

Categories

Questions about Databricks Spark Structured Streaming

What does Databricks Spark Structured Streaming do?

Comprehensive guide to Spark Structured Streaming for production workloads. Databricks Spark Structured Streaming is an agent skill from databricks/databricks-agent-skills, published by the product's own GitHub organization. Comprehensive guide to Spark Structured Streaming for production workloads.

When should I use Databricks Spark Structured Streaming?

Databricks Spark Structured Streaming fits situations like: building streaming pipelines; working with Kafka ingestion; implementing Real-Time Mode (RTM); configuring triggers (processingTime.

How do I install Databricks Spark Structured Streaming in Claude Code?

Run `npx skills add databricks/databricks-agent-skills --skill databricks-spark-structured-streaming -a claude-code`. Or copy the skill folder (skills/databricks-spark-structured-streaming in databricks/databricks-agent-skills) into .claude/skills/databricks-spark-structured-streaming in your project. Claude Code loads it when a task matches its description.

How do I install Databricks Spark Structured Streaming in Codex?

Run `npx skills add databricks/databricks-agent-skills --skill databricks-spark-structured-streaming -a codex`. Or copy the skill folder (skills/databricks-spark-structured-streaming in databricks/databricks-agent-skills) into .agents/skills/databricks-spark-structured-streaming in your project. Codex loads it when a task matches its description.

Can I use Databricks Spark Structured Streaming in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add databricks/databricks-agent-skills --skill databricks-spark-structured-streaming -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/databricks-spark-structured-streaming, .gemini/skills/databricks-spark-structured-streaming, .github/skills/databricks-spark-structured-streaming and .opencode/skills/databricks-spark-structured-streaming in your project.

What does Databricks Spark Structured Streaming need to run?

SKILL.md names no scripts, command-line tools or credentials: Databricks Spark Structured Streaming is instructions for the agent only. Our summary lists: Python 3. Compatibility (from SKILL.md): Requires databricks CLI (>= v1.0.0).

Does Databricks Spark Structured Streaming access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Databricks Spark Structured Streaming safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Databricks Spark Structured Streaming use?

Databricks Spark Structured Streaming has a licence file (the repository's licence) that doesn't match a standard licence. Read it on GitHub before reusing the skill.

How many tokens does Databricks Spark Structured Streaming use?

About 956 tokens (SKILL.md is roughly 3.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 39k tokens, read only when the agent opens those files.

What are the alternatives to Databricks Spark Structured Streaming?

Skills that share tags, products or a category with Databricks Spark Structured Streaming: Windmill Trigger Type Checklist (windmill-labs/windmill, 18k stars), Foundatio (FoundatioFx/Foundatio, 2.1k stars), Opensource Guide Coach (calf-ai/calfkit-sdk, 149 stars) and Create Environment (godatadriven/whirl, 205 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Databricks Spark Structured Streaming?

databricks (a GitHub organization, an official publisher) maintains it in databricks/databricks-agent-skills, which has 345 GitHub stars. The repository holds 32 skills in this directory. The repository was last updated on October 8, 2026.

Source: databricks/databricks-agent-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.