Agent skill

Databricks Streaming Guardian

by jeremylongshore in jeremylongshore/tons-of-skills-marketplace

Guard production Databricks data pipelines — Delta Lake, Liquid Clustering, Structured Streaming, Auto Loader, and DLT — against the twelve foot-guns that fire at scale: OPTIMIZE/auto-compaction…

MITAuto-check passedDevelopment

Install Databricks Streaming Guardian

skills CLI
$ npx skills add jeremylongshore/tons-of-skills-marketplace --skill databricks-streaming-guardian -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install jeremylongshore/tons-of-skills-marketplace databricks-streaming-guardian --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/jeremylongshore/tons-of-skills-marketplace.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/.curated/databricks-streaming-guardian .claude/skills/databricks-streaming-guardian && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
databricks-streaming-guardian
GitHub stars
2.8k
Token cost
~3.5k tokens
SKILL.md length
1,366 words
Files
14 (incl. scripts, references)
Skills in repo
3,342
Repo updated
First seen
Licence
MIT

At a glance

Guard production Databricks data pipelines — Delta Lake, Liquid Clustering, Structured Streaming, Auto Loader, and DLT — against the twelve foot-guns that fire at scale: OPTIMIZE/auto-compaction…

  • Works in 6 steps: Before a destructive op (the hook does… → A Delta write conflict (D01, D02) → A broken streaming source (D03, D04, D12) → …
  • A Delta MERGE/OPTIMIZE fails with a concurrency exception
  • SKILL.md covers Overview, Prerequisites, Instructions and Output, plus 3 more sections
  • Runs Python and Shell scripts from its folder; calls bash, python3 and git

What it does

Databricks Streaming Guardian is an agent skill from jeremylongshore/tons-of-skills-marketplace. Guard production Databricks data pipelines — Delta Lake, Liquid Clustering, Structured Streaming, Auto Loader, and DLT — against the twelve foot-guns that fire at scale: OPTIMIZE/auto-compaction conflicts, Liquid-Clustering merge conflicts, VACUUM breaking a streaming checkpoint, RocksDB OOM, Auto Loader schema-evolution stops, and DLT refresh data loss. Includes a PreToolUse hook that blocks DROP/CREATE-OR-REPLACE/VACUUM on a table with active streaming consumers. Use when a Delta MERGE/OPTIMIZE fails with a…

Its SKILL.md is about 3.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 18 other files, including scripts and reference files (for example `agents/merge-rewriter.md`, `docs/ADR.md` and `docs/ONE-PAGER.md`). Compatibility notes: Designed for Claude Code

It sits in Development, covering Database administration, Git workflow and Data pipelines and ETL. It works with Databricks. The repository describes itself as: Model-agnostic agent-skills platform with a harness-free canonical layer, verified adapters, and the ccpi package manager. Explore at tonsofskills.com. The licence is MIT.

When your agent uses it

  • A Delta MERGE/OPTIMIZE fails with a concurrency exception
  • A stream breaks after VACUUM
  • A table replace
  • An Auto Loader stream stops on a new column

Example prompts

  • “ConcurrentAppendException”
  • “ConcurrentDeleteDeleteException”
  • “DELTAFILENOTFOUND”
  • “/databricks-streaming-guardian”

Requirements

  • Python 3
  • A Bash shell
  • Compatibility (from SKILL.md): Designed for Claude Code
  • Pre-approved tools (allowed-tools): Read, Write, Edit, Bash(databricks:*), Bash(jq:*), Bash(python3:*), Bash(bash:*), Glob, mcp__databricks-workspace-mcp__clusters_events, mcp__databricks-workspace-mcp__clusters_list, mcp__databricks-workspace-mcp__pipelines_get

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Before a destructive op (the hook does this automatically)
  2. A Delta write conflict (D01, D02)
  3. A broken streaming source (D03, D04, D12)
  4. RocksDB state-store OOM (D05)
  5. Auto Loader schema evolution (D10)
  6. DLT rebuild safety (D08, D09, D11)

What it can do on your machine

Read from SKILL.md and the folder at commit cfae287. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Write
    • Edit
    • Bash(databricks:*)
    • Bash(jq:*)
    • Bash(python3:*)
    • Bash(bash:*)
    • Glob
    • mcp__databricks-workspace-mcp__clusters_events
    • mcp__databricks-workspace-mcp__clusters_list

    …and 1 more on the same allowed-tools line.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 2 files in scripts/ (Python and Shell), which the agent can run.

    Shell commands in SKILL.md call:

    • bash
    • python3
    • git

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • docs.databricks.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Designed for Claude Code

    From compatibility in the SKILL.md frontmatter.

Context cost

Databricks Streaming Guardian loads about 3.5k tokens when it runs, and up to ~22k if it reads all its reference files. Until then it costs about 240 tokens; SKILL.md has 1,366 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~240
When it runs · the whole SKILL.md, loaded when a task matches
~3.5k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~22k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from jeremylongshore/tons-of-skills-marketplace at commit cfae287, republished under its MIT licence (© jeremylongshore). 1,366 words, ~3,546 tokens.

Download SKILL.mdSave it as .claude/skills/databricks-streaming-guardian/SKILL.md (or your agent's skills folder). This skill also uses 13 other files; get the full folder from GitHub.
name
databricks-streaming-guardian
description
Guard production Databricks data pipelines — Delta Lake, Liquid Clustering, Structured Streaming, Auto Loader, and DLT — against the twelve foot-guns that fire at scale: OPTIMIZE/auto-compaction conflicts, Liquid-Clustering merge conflicts, VACUUM breaking a streaming checkpoint, RocksDB OOM, Auto Loader schema-evolution stops, and DLT refresh data loss. Includes a PreToolUse hook that blocks DROP/CREATE-OR-REPLACE/VACUUM on a table with active streaming consumers. Use when a Delta MERGE/OPTIMIZE fails with a concurrency exception, a stream breaks after VACUUM or a table replace, an Auto Loader stream stops on a new column, a DLT refresh drops data, or before running a destructive op on a streamed-from table. Trigger with "ConcurrentAppendException", "ConcurrentDeleteDeleteException", "DELTA_FILE_NOT_FOUND", "streaming checkpoint broke", "vacuum broke my stream", "autoloader UnknownFieldException", "dlt full refresh".
allowed-tools
Read, Write, Edit, Bash(databricks:*), Bash(jq:*), Bash(python3:*), Bash(bash:*), Glob, mcp__databricks-workspace-mcp__clusters_events, mcp__databricks-workspace-mcp__clusters_list, mcp__databricks-workspace-mcp__pipelines_get
compatibility
Designed for Claude Code
version
2.28.0
author
Jeremy Longshore <jeremy@intentsolutions.io>
license
MIT
tags
saas, databricks, streaming, delta, data-ops

Databricks Streaming Guardian

The data-ops spine of the pack. Delta Lake, Liquid Clustering, Structured Streaming, and DLT each ship a different set of foot-guns that fire most visibly when production data flows through them at scale — and most of them are documented platform decisions that surprise engineers, not bugs. This skill's job is friction at trigger time (a hook that blocks the genuinely-irreversible op) plus deterministic recovery when something already broke.

Overview

Twelve foot-guns, grouped by the surface that triggers them. Eleven are owned outright (D01–D10, D12); the twelfth — D11, DLT rebuild cost — is shared with databricks-cost-leak-hunter: this skill checks the rebuild cost as part of pre-refresh safety, that skill owns ongoing cost optimization.

Delta write conflicts. D01 ConcurrentDeleteDeleteException — a manual OPTIMIZE colliding with auto-compaction, which is silently enabled on any table touched by MERGE/UPDATE/DELETE. D02 ConcurrentAppendException after moving to Liquid Clustering — LC keeps file-set-level writer conflicts; a fan-out MERGE breaks unless its predicate is narrowed to the clustering keys.

Streaming + checkpoint. D03 DELTA_FILE_NOT_FOUND_DETAILED — VACUUM deletes files the checkpoint pins to. D04 silent checkpoint corruption / reset to batch 0. D05 RocksDB state-store off-heap OOM (the heap looks fine while off-heap state pins multi-GB). D12 DIFFERENT_DELTA_TABLE_READ_BY_STREAMING_SOURCE — CREATE OR REPLACE mints a new UUID and kills every active consumer.

Migration + evolution. D06 Liquid-Clustering migration's hidden full-rewrite cost + downstream partition-predicate breakage. D07 time travel breaking silently when VACUUM crosses the retention boundary. D10 Auto Loader UnknownFieldException stopping the stream on every new column.

DLT. D08 the @dlt.table thread race (out-of-order registration). D09 full refresh silently dropping data from a non-replayable source. D11 the rebuild cost multiplier (checked before a full refresh; ongoing DLT cost is databricks-cost-leak-hunter's job).

The hook (AP02/AP06 — this pack's only blocking hook). A PreToolUse hook (hooks/streaming-guard-hook.py) intercepts a Bash command that runs DROP TABLE, CREATE OR REPLACE TABLE, or VACUUM against a table and — only when it confirms via system.streaming.query_progress that an active stream reads that table — blocks it with a message naming the consumers and the pain. It is precise by design: it matches only real SQL-execution surfaces (never a git commit mentioning "drop table"), and it fails open — if it cannot verify consumers, it allows rather than false-block. Blocking is reserved for the genuinely irreversible.

Deterministic work lives in scripts/; deep knowledge in references/; the Liquid-Clustering predicate rewrite in the merge-rewriter subagent. Two data planes: the databricks-workspace-mcp control plane (cluster/pipeline events) and the CLI Statement Execution API for system.* reads. Either absent → advisory mode on pasted input.

Prerequisites

  • databricks-workspace-mcp registered — for clusters_events (RocksDB OOM correlation) and pipelines_get (DLT event log). Absent → advisory mode.
  • Databricks CLI authenticated + jq, and DATABRICKS_WAREHOUSE_ID set — for the system.streaming.query_progress reads the hook and recovery flows use. The hook fails open (allows) if these are absent, so it never false-blocks.
  • The hook is a plugin-level PreToolUse hook — it runs on Bash commands once the pack is installed. It is silent on everything except a confirmed-unsafe destructive op.

Instructions

Pick the flow by symptom. Always name the exact, searchable Databricks error string — ConcurrentAppendException, ConcurrentDeleteDeleteException, DELTA_FILE_NOT_FOUND_DETAILED, DIFFERENT_DELTA_TABLE_READ_BY_STREAMING_SOURCE, UnknownFieldException — even when the user paraphrases it or gives a short form; the full code is what an operator greps logs and docs for.

Step 1: Before a destructive op (the hook does this automatically)

Running DROP TABLE / CREATE OR REPLACE TABLE / VACUUM on a table? The hook checks for active streaming consumers first and blocks if any exist. To check manually, query system.streaming.query_progress for a stream whose source_description names the table. If consumers exist: do NOT CREATE OR REPLACE (use ALTER/in-place — D12) and do NOT VACUUM below the consumers' checkpoint lag (D03/D07). See ${CLAUDE_SKILL_DIR}/references/checkpoint-recovery.md.

Step 2: A Delta write conflict (D01, D02)
  • ConcurrentDeleteDeleteException (D01) — before a manual OPTIMIZE, probe the table for auto-compaction:

    bash
    bash "${CLAUDE_SKILL_DIR}/scripts/pre-optimize-check.sh" --table main.sales.orders

    If it reports COLLISION RISK, don't run manual OPTIMIZE (or disable auto-compaction first). Details: ${CLAUDE_SKILL_DIR}/references/concurrency-conflicts.md.

  • ConcurrentAppendException on a Liquid-Clustering table (D02) — hand the failing MERGE to the merge-rewriter subagent; it fetches the target's clustering keys via DESCRIBE DETAIL and narrows the ON predicate so writers touch disjoint file sets.

Step 3: A broken streaming source (D03, D04, D12)

Map the symptom to the failure class and its exact error code, then get the recovery tier from the decision tree:

  • file-not-found → DELTA_FILE_NOT_FOUND_DETAILED (VACUUM deleted pinned files — D03)
  • uuid-changed → DIFFERENT_DELTA_TABLE_READ_BY_STREAMING_SOURCE (CREATE OR REPLACE minted a new UUID — D12)
  • checkpoint-reset → silent batchId regression / checkpoint corruption (D04)
  • transient → a restartable blip with an intact checkpoint
bash
python3 "${CLAUDE_SKILL_DIR}/scripts/recover-streaming-source.py" \
  --failure file-not-found --time-travel yes    # or uuid-changed / checkpoint-reset / transient

It echoes the canonical error code and recommends SAFE_RESTART / REPROCESS_FROM_OFFSET / RESTORE_FROM_TIME_TRAVEL / FULL_RESET_BACKFILL with the data-loss tradeoff stated up front. Name that full code in your answer — not just the short class. The full three-tier reasoning is in ${CLAUDE_SKILL_DIR}/references/checkpoint-recovery.md.

Step 4: RocksDB state-store OOM (D05)

A driver/executor OOM while the JVM heap looks healthy points at off-heap RocksDB state. Correlate the OOM to state size with clusters_events, then bound the memory and enable changelog checkpointing per ${CLAUDE_SKILL_DIR}/references/rocksdb-state-store-tuning.md.

Show full SKILL.md (565 more words)Show less
Step 5: Auto Loader schema evolution (D10)

A stream stopping with UnknownFieldException on a new column is the default addNewColumns mode. Choose the mode deliberately (evolve-and-restart vs rescue's silent widening) and pin types with schemaHints per ${CLAUDE_SKILL_DIR}/references/autoloader-schema-evolution.md.

Step 6: DLT rebuild safety (D08, D09, D11)

Before a DLT full refresh, confirm every source is replayable (a Kafka topic past retention or a truncate-and-load source loses data on refresh — D09) and that @dlt.table registration is deterministic (the thread race — D08). Read the DLT event log with pipelines_get; the checklist is in ${CLAUDE_SKILL_DIR}/references/dlt-rebuild-safety.md.

Output

  • A hook decision — a destructive op on a streamed-from table is blocked with the active consumers named and the pain (D12 / D03-D07) explained; everything else passes silently.
  • A pre-OPTIMIZE verdict — SAFE or COLLISION RISK (auto-compaction on) with the disable-or-serialize fix.
  • A rewritten MERGE — the LC clustering-key-scoped predicate (from merge-rewriter) that stops ConcurrentAppendException.
  • A recovery recommendation — the recovery tier + steps + the data-loss risk, for the specific failure class.
  • A tuning / mode / refresh-safety recommendation — RocksDB bounds, Auto Loader mode, or the DLT full-refresh checklist, from the matching reference.

Error Handling

ErrorCauseSolution
ConcurrentDeleteDeleteExceptionManual OPTIMIZE races auto-compaction (D01)Run pre-optimize-check.sh; don't manually OPTIMIZE an auto-compacted table, or disable auto-compact first.
ConcurrentAppendException on an LC tableMERGE predicate not scoped to clustering keys (D02)Route the MERGE to merge-rewriter; narrow the ON predicate to the clustering keys.
DELTA_FILE_NOT_FOUND_DETAILEDVACUUM deleted checkpoint-pinned files (D03)recover-streaming-source.py --failure file-not-found; restore via time travel if in retention, else reprocess/reset.
DIFFERENT_DELTA_TABLE_READ_BY_STREAMING_SOURCECREATE OR REPLACE minted a new UUID (D12)The old checkpoint is dead — new checkpoint + backfill; the hook prevents this going forward.
Driver OOM, heap looks fineOff-heap RocksDB state (D05)Bound state-store memory + changelog checkpointing; size state with a watermark.
UnknownFieldException, stream stoppedAuto Loader addNewColumns default (D10)Restart to evolve (idempotent sink), or choose rescue/schemaHints deliberately.
Hook allowed a destructive op with a warningCould not verify consumers (no CLI/warehouse)Advisory — the hook fails open; verify system.streaming.query_progress manually before running it.

Examples

Example 1: "About to CREATE OR REPLACE a table other jobs stream from."

The PreToolUse hook fires, confirms 2 active consumers via system.streaming.query_progress, and blocks with: "CREATE OR REPLACE mints a new UUID → both consumers die with DIFFERENT_DELTA_TABLE_READ_BY_STREAMING_SOURCE; use ALTER / in-place."

Example 2: "My MERGE into a Liquid-Clustering table fails with ConcurrentAppendException."

The merge-rewriter subagent reads the target's clustering keys via DESCRIBE DETAIL and rewrites the ON predicate to include them, so concurrent writers touch disjoint file sets — the exception stops without serializing the jobs.

Example 3: "My stream died with DELTA_FILE_NOT_FOUND after a VACUUM."

recover-streaming-source.py --failure file-not-found --time-travel yes → RESTORE_FROM_TIME_TRAVEL (no data loss): restore the source to a pre-VACUUM version, restart on the existing checkpoint, then align VACUUM retention with the checkpoint lag.

Example 4: "Before I OPTIMIZE this table."

pre-optimize-check.sh --table main.sales.orders reports COLLISION RISK because delta.autoOptimize.autoCompact is on — so the skill recommends letting auto-compaction do it, or disabling it for the maintenance window first.

Resources

© jeremylongshore, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 13 other files (scripts, references) in skills/.curated/databricks-streaming-guardian of jeremylongshore/tons-of-skills-marketplace.

  • SKILL.md
  • agents/merge-rewriter.md
  • docs/ADR.md
  • docs/ONE-PAGER.md
  • docs/PRD.md
  • eval-spec.yaml
  • hooks/streaming-guard-hook.py
  • references/autoloader-schema-evolution.md
  • references/checkpoint-recovery.md
  • references/concurrency-conflicts.md
  • references/dlt-rebuild-safety.md
  • references/rocksdb-state-store-tuning.md
  • scripts/pre-optimize-check.sh
  • scripts/recover-streaming-source.py

Open the folder on GitHubat commit cfae287

Compare with similar skills

Databricks Streaming Guardian next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Databricks Streaming Guardian compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Databricks Streaming Guardian this skilljeremylongshore/tons-of-skills-marketplace2.8k—~3.5kAutomated safety check: PassMIT
Git PR Flowwgzhao/Addax1.4k—~216Automated safety check: PassApache-2.0
Optimizing Databricks SQLAltimateAI/data-engineering-skills128—~6.7kAutomated safety check: PassMIT
Replication Driven Researchbrycewang-stanford/Auto-Empirical-Research-Skills4.6k—~1.7kAutomated safety check: PassCustom licence
DB SculptorEliasOulkadi/shokunin114—~3.1kAutomated safety check: NotesMIT
Batch Processing Clinical Textmaziyarpanahi/openmed5.5k—~2.2kAutomated safety check: PassApache-2.0

Similar skills

  • Git PR Flow

    wgzhao/Addax

    Standard branch, commit, and pull-request workflow for this repo.

    1.4k GitHub stars~216 tokensUpdated 2 days ago
    DevelopmentAuto-check passed
  • Optimizing Databricks SQL

    AltimateAI/data-engineering-skills

    Analyze DBSQL queries, including SQL embedded in notebooks (spark.sql(...), %sql cells), for anti-patterns, lint issues, and performance problems, using Databricks-specific dialect and platform…

    128 GitHub stars~6.7k tokensUpdated 3 days ago
    DatabasesAuto-check passed
  • Replication Driven Research

    brycewang-stanford/Auto-Empirical-Research-Skills

    A skill your agent uses when starting empirical analysis, creating a data pipeline, generating results, or when data or model specifications change.

    4.6k GitHub stars~1.7k tokensUpdated 5 days ago
    Testing & QAAuto-check passed
  • DB Sculptor

    EliasOulkadi/shokunin

    Design database schemas with Prisma/Drizzle, PostgreSQL index strategy (B-tree, GIN, GiST, BRIN, Hash), query optimization (EXPLAIN ANALYZE), migration safety (expand/contract, zero-downtime), and…

    114 GitHub stars~3.1k tokensUpdated 6 days ago
    DatabasesAuto-check: notes
  • Batch Processing Clinical Text

    maziyarpanahi/openmed

    Run large-scale batch NER, PII extraction, or de-identification over many clinical notes on-device with OpenMed, with sharding, checkpointing, resumability, and append-only JSONL output.

    5.5k GitHub stars~2.2k tokensUpdated today
    DatabasesAuto-check passed
  • Airflow State Store

    astronomer/agents

    Persists task and asset state across retries and DAG runs using Airflow 3.3's AIP-103 key/value stores (taskstatestore, assetstatestore) and the crash-safe ResumableJobMixin.

    451 GitHub stars~6.1k tokensUpdated 3 days ago
    Data & AnalyticsAuto-check passed

More from jeremylongshore/tons-of-skills-marketplace

All 3,342 skills in this repo
  • Performing Security Code Review

    jeremylongshore/tons-of-skills-marketplace

    Execute this skill enables AI assistant to conduct a security-focused code review using the security-agent plugin.

    2.8k GitHub starsUsed in 2 repos~1.3k tokens
    Auto-check: notes
  • Adapting Transfer Learning Models

    jeremylongshore/tons-of-skills-marketplace

    Build this skill automates the adaptation of pre-trained machine learning models using transfer learning techniques.

    2.8k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Agent Context Loader

    jeremylongshore/tons-of-skills-marketplace

    Execute proactive auto-loading: automatically detects and loads agents.md files.

    2.8k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Aggregating Performance Metrics

    jeremylongshore/tons-of-skills-marketplace

    Aggregate and centralize performance metrics from applications, systems, databases, caches, and services.

    2.8k GitHub stars~1.2k tokensUpdated today
    Auto-check passed
  • Analyzing Capacity Planning

    jeremylongshore/tons-of-skills-marketplace

    Execute this skill enables AI assistant to analyze capacity requirements and plan for future growth.

    2.8k GitHub stars~947 tokensUpdated today
    Auto-check passed
  • Analyzing Database Indexes

    jeremylongshore/tons-of-skills-marketplace

    Process use when you need to work with database indexing. An agent skill from jeremylongshore/tons-of-skills-marketplace.

    2.8k GitHub stars~2k tokensUpdated today
    Auto-check passed

Works with

Questions about Databricks Streaming Guardian

What does Databricks Streaming Guardian do?

Guard production Databricks data pipelines — Delta Lake, Liquid Clustering, Structured Streaming, Auto Loader, and DLT — against the twelve foot-guns that fire at scale: OPTIMIZE/auto-compaction…. Databricks Streaming Guardian is an agent skill from jeremylongshore/tons-of-skills-marketplace. Guard production Databricks data pipelines — Delta Lake, Liquid Clustering, Structured Streaming, Auto Loader, and DLT — against the twelve foot-guns that fire at scale: OPTIMIZE/auto-compaction conflicts, Liquid-Clustering merge conflicts, VACUUM breaking a streaming checkpoint, RocksDB OOM, Auto Loader schema-evolution stops, and DLT refresh data loss.

When should I use Databricks Streaming Guardian?

Databricks Streaming Guardian fits situations like: A Delta MERGE/OPTIMIZE fails with a concurrency exception; A stream breaks after VACUUM; A table replace; an Auto Loader stream stops on a new column.

How do I install Databricks Streaming Guardian in Claude Code?

Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill databricks-streaming-guardian -a claude-code`. Or copy the skill folder (skills/.curated/databricks-streaming-guardian in jeremylongshore/tons-of-skills-marketplace) into .claude/skills/databricks-streaming-guardian in your project. Claude Code loads it when a task matches its description.

How do I install Databricks Streaming Guardian in Codex?

Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill databricks-streaming-guardian -a codex`. Or copy the skill folder (skills/.curated/databricks-streaming-guardian in jeremylongshore/tons-of-skills-marketplace) into .agents/skills/databricks-streaming-guardian in your project. Codex loads it when a task matches its description.

Can I use Databricks Streaming Guardian in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill databricks-streaming-guardian -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/databricks-streaming-guardian, .gemini/skills/databricks-streaming-guardian, .github/skills/databricks-streaming-guardian and .opencode/skills/databricks-streaming-guardian in your project.

What does Databricks Streaming Guardian need to run?

Going by SKILL.md and its folder, Databricks Streaming Guardian needs Python and a shell for the scripts in its folder and the command-line tools its instructions call (bash, python3 and git). Our summary lists: Python 3; A Bash shell. Its frontmatter pre-approves these tools: Read, Write, Edit, Bash(databricks:*), Bash(jq:*), Bash(python3:*), Bash(bash:*), Glob, mcp__databricks-workspace-mcp__clusters_events, mcp__databricks-workspace-mcp__clusters_list, mcp__databricks-workspace-mcp__pipelines_get. Compatibility (from SKILL.md): Designed for Claude Code.

Does Databricks Streaming Guardian access the network?

SKILL.md names 1 domain. As links in the text: docs.databricks.com. This is read from the text; nothing was executed.

Is Databricks Streaming Guardian safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Databricks Streaming Guardian use?

Databricks Streaming Guardian is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Databricks Streaming Guardian use?

About 3.5k tokens (SKILL.md is roughly 14k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 19k tokens, read only when the agent opens those files.

What are the alternatives to Databricks Streaming Guardian?

Skills that share tags, products or a category with Databricks Streaming Guardian: Git PR Flow (wgzhao/Addax, 1.4k stars), Optimizing Databricks SQL (AltimateAI/data-engineering-skills, 128 stars), Replication Driven Research (brycewang-stanford/Auto-Empirical-Research-Skills, 4.6k stars) and DB Sculptor (EliasOulkadi/shokunin, 114 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Databricks Streaming Guardian?

jeremylongshore (a GitHub user) maintains it in jeremylongshore/tons-of-skills-marketplace, which has 2,827 GitHub stars. The repository holds 3,342 skills in this directory. The repository was last updated on October 10, 2026.

Source: jeremylongshore/tons-of-skills-marketplace on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.