Git PR Flow
wgzhao/Addax
Standard branch, commit, and pull-request workflow for this repo.
Agent skill
by jeremylongshore in jeremylongshore/tons-of-skills-marketplace
Guard production Databricks data pipelines — Delta Lake, Liquid Clustering, Structured Streaming, Auto Loader, and DLT — against the twelve foot-guns that fire at scale: OPTIMIZE/auto-compaction…
$ npx skills add jeremylongshore/tons-of-skills-marketplace --skill databricks-streaming-guardian -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install jeremylongshore/tons-of-skills-marketplace databricks-streaming-guardian --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/jeremylongshore/tons-of-skills-marketplace.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/.curated/databricks-streaming-guardian .claude/skills/databricks-streaming-guardian && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "databricks-streaming-guardian" agent skill from https://github.com/jeremylongshore/tons-of-skills-marketplace/tree/main/skills/.curated/databricks-streaming-guardian into .claude/skills/databricks-streaming-guardian/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "databricks-streaming-guardian", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/jeremylongshore/tons-of-skills-marketplace/tree/main/skills/.curated/databricks-streaming-guardianType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add jeremylongshore/tons-of-skills-marketplace --skill databricks-streaming-guardian -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install jeremylongshore/tons-of-skills-marketplace databricks-streaming-guardian --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jeremylongshore/tons-of-skills-marketplace.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/.curated/databricks-streaming-guardian .agents/skills/databricks-streaming-guardian && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "databricks-streaming-guardian" agent skill from https://github.com/jeremylongshore/tons-of-skills-marketplace/tree/main/skills/.curated/databricks-streaming-guardian into .agents/skills/databricks-streaming-guardian/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "databricks-streaming-guardian", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add jeremylongshore/tons-of-skills-marketplace --skill databricks-streaming-guardian -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install jeremylongshore/tons-of-skills-marketplace databricks-streaming-guardian --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jeremylongshore/tons-of-skills-marketplace.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/.curated/databricks-streaming-guardian .cursor/skills/databricks-streaming-guardian && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "databricks-streaming-guardian" agent skill from https://github.com/jeremylongshore/tons-of-skills-marketplace/tree/main/skills/.curated/databricks-streaming-guardian into .cursor/skills/databricks-streaming-guardian/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "databricks-streaming-guardian", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/jeremylongshore/tons-of-skills-marketplace.git --path skills/.curated/databricks-streaming-guardian--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add jeremylongshore/tons-of-skills-marketplace --skill databricks-streaming-guardian -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install jeremylongshore/tons-of-skills-marketplace databricks-streaming-guardian --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jeremylongshore/tons-of-skills-marketplace.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/.curated/databricks-streaming-guardian .gemini/skills/databricks-streaming-guardian && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "databricks-streaming-guardian" agent skill from https://github.com/jeremylongshore/tons-of-skills-marketplace/tree/main/skills/.curated/databricks-streaming-guardian into .gemini/skills/databricks-streaming-guardian/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "databricks-streaming-guardian", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install jeremylongshore/tons-of-skills-marketplace databricks-streaming-guardianInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add jeremylongshore/tons-of-skills-marketplace --skill databricks-streaming-guardian -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/jeremylongshore/tons-of-skills-marketplace.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/.curated/databricks-streaming-guardian .github/skills/databricks-streaming-guardian && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "databricks-streaming-guardian" agent skill from https://github.com/jeremylongshore/tons-of-skills-marketplace/tree/main/skills/.curated/databricks-streaming-guardian into .github/skills/databricks-streaming-guardian/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "databricks-streaming-guardian", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add jeremylongshore/tons-of-skills-marketplace --skill databricks-streaming-guardian -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install jeremylongshore/tons-of-skills-marketplace databricks-streaming-guardian --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jeremylongshore/tons-of-skills-marketplace.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/.curated/databricks-streaming-guardian .opencode/skills/databricks-streaming-guardian && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "databricks-streaming-guardian" agent skill from https://github.com/jeremylongshore/tons-of-skills-marketplace/tree/main/skills/.curated/databricks-streaming-guardian into .opencode/skills/databricks-streaming-guardian/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "databricks-streaming-guardian", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
databricks-streaming-guardianGuard production Databricks data pipelines — Delta Lake, Liquid Clustering, Structured Streaming, Auto Loader, and DLT — against the twelve foot-guns that fire at scale: OPTIMIZE/auto-compaction…
Databricks Streaming Guardian is an agent skill from jeremylongshore/tons-of-skills-marketplace. Guard production Databricks data pipelines — Delta Lake, Liquid Clustering, Structured Streaming, Auto Loader, and DLT — against the twelve foot-guns that fire at scale: OPTIMIZE/auto-compaction conflicts, Liquid-Clustering merge conflicts, VACUUM breaking a streaming checkpoint, RocksDB OOM, Auto Loader schema-evolution stops, and DLT refresh data loss. Includes a PreToolUse hook that blocks DROP/CREATE-OR-REPLACE/VACUUM on a table with active streaming consumers. Use when a Delta MERGE/OPTIMIZE fails with a…
Its SKILL.md is about 3.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 18 other files, including scripts and reference files (for example `agents/merge-rewriter.md`, `docs/ADR.md` and `docs/ONE-PAGER.md`). Compatibility notes: Designed for Claude Code
It sits in Development, covering Database administration, Git workflow and Data pipelines and ETL. It works with Databricks. The repository describes itself as: Model-agnostic agent-skills platform with a harness-free canonical layer, verified adapters, and the ccpi package manager. Explore at tonsofskills.com. The licence is MIT.
6 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit cfae287. It shows what the files ask for, not the result of running them.
Pre-approves these tools, so the agent can use them without asking each time:
ReadWriteEditBash(databricks:*)Bash(jq:*)Bash(python3:*)Bash(bash:*)Globmcp__databricks-workspace-mcp__clusters_eventsmcp__databricks-workspace-mcp__clusters_list…and 1 more on the same allowed-tools line.
From allowed-tools in the SKILL.md frontmatter.
Ships 2 files in scripts/ (Python and Shell), which the agent can run.
Shell commands in SKILL.md call:
bashpython3gitFrom the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
docs.databricks.comFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Designed for Claude Code
From compatibility in the SKILL.md frontmatter.
Databricks Streaming Guardian loads about 3.5k tokens when it runs, and up to ~22k if it reads all its reference files. Until then it costs about 240 tokens; SKILL.md has 1,366 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from jeremylongshore/tons-of-skills-marketplace at commit cfae287, republished under its MIT licence (© jeremylongshore). 1,366 words, ~3,546 tokens.
.claude/skills/databricks-streaming-guardian/SKILL.md (or your agent's skills folder). This skill also uses 13 other files; get the full folder from GitHub.The data-ops spine of the pack. Delta Lake, Liquid Clustering, Structured Streaming, and DLT each ship a different set of foot-guns that fire most visibly when production data flows through them at scale — and most of them are documented platform decisions that surprise engineers, not bugs. This skill's job is friction at trigger time (a hook that blocks the genuinely-irreversible op) plus deterministic recovery when something already broke.
Twelve foot-guns, grouped by the surface that triggers them. Eleven are owned
outright (D01–D10, D12); the twelfth — D11, DLT rebuild cost — is shared with
databricks-cost-leak-hunter: this skill checks the rebuild cost as part of
pre-refresh safety, that skill owns ongoing cost optimization.
Delta write conflicts. D01 ConcurrentDeleteDeleteException — a manual
OPTIMIZE colliding with auto-compaction, which is silently enabled on any table
touched by MERGE/UPDATE/DELETE. D02 ConcurrentAppendException after moving
to Liquid Clustering — LC keeps file-set-level writer conflicts; a fan-out MERGE
breaks unless its predicate is narrowed to the clustering keys.
Streaming + checkpoint. D03 DELTA_FILE_NOT_FOUND_DETAILED — VACUUM deletes
files the checkpoint pins to. D04 silent checkpoint corruption / reset to batch 0.
D05 RocksDB state-store off-heap OOM (the heap looks fine while off-heap state pins
multi-GB). D12 DIFFERENT_DELTA_TABLE_READ_BY_STREAMING_SOURCE — CREATE OR REPLACE mints a new UUID and kills every active consumer.
Migration + evolution. D06 Liquid-Clustering migration's hidden full-rewrite
cost + downstream partition-predicate breakage. D07 time travel breaking silently
when VACUUM crosses the retention boundary. D10 Auto Loader
UnknownFieldException stopping the stream on every new column.
DLT. D08 the @dlt.table thread race (out-of-order registration). D09 full
refresh silently dropping data from a non-replayable source. D11 the rebuild cost
multiplier (checked before a full refresh; ongoing DLT cost is
databricks-cost-leak-hunter's job).
The hook (AP02/AP06 — this pack's only blocking hook). A PreToolUse hook
(hooks/streaming-guard-hook.py) intercepts a Bash command that runs DROP TABLE,
CREATE OR REPLACE TABLE, or VACUUM against a table and — only when it confirms
via system.streaming.query_progress that an active stream reads that table —
blocks it with a message naming the consumers and the pain. It is precise by
design: it matches only real SQL-execution surfaces (never a git commit mentioning
"drop table"), and it fails open — if it cannot verify consumers, it allows
rather than false-block. Blocking is reserved for the genuinely irreversible.
Deterministic work lives in scripts/; deep knowledge in references/; the
Liquid-Clustering predicate rewrite in the merge-rewriter subagent. Two data
planes: the databricks-workspace-mcp control plane (cluster/pipeline events) and
the CLI Statement Execution API for system.* reads. Either absent → advisory mode
on pasted input.
databricks-workspace-mcp registered — for clusters_events (RocksDB OOM
correlation) and pipelines_get (DLT event log). Absent → advisory mode.jq, and DATABRICKS_WAREHOUSE_ID set —
for the system.streaming.query_progress reads the hook and recovery flows use.
The hook fails open (allows) if these are absent, so it never false-blocks.PreToolUse hook — it runs on Bash commands once
the pack is installed. It is silent on everything except a confirmed-unsafe
destructive op.Pick the flow by symptom. Always name the exact, searchable Databricks error
string — ConcurrentAppendException, ConcurrentDeleteDeleteException,
DELTA_FILE_NOT_FOUND_DETAILED, DIFFERENT_DELTA_TABLE_READ_BY_STREAMING_SOURCE,
UnknownFieldException — even when the user paraphrases it or gives a short form;
the full code is what an operator greps logs and docs for.
Running DROP TABLE / CREATE OR REPLACE TABLE / VACUUM on a table? The hook
checks for active streaming consumers first and blocks if any exist. To check
manually, query system.streaming.query_progress for a stream whose
source_description names the table. If consumers exist: do NOT CREATE OR REPLACE (use ALTER/in-place — D12) and do NOT VACUUM below the consumers'
checkpoint lag (D03/D07). See
${CLAUDE_SKILL_DIR}/references/checkpoint-recovery.md.
ConcurrentDeleteDeleteException (D01) — before a manual OPTIMIZE, probe
the table for auto-compaction:
bash "${CLAUDE_SKILL_DIR}/scripts/pre-optimize-check.sh" --table main.sales.ordersIf it reports COLLISION RISK, don't run manual OPTIMIZE (or disable
auto-compaction first). Details:
${CLAUDE_SKILL_DIR}/references/concurrency-conflicts.md.
ConcurrentAppendException on a Liquid-Clustering table (D02) — hand the
failing MERGE to the merge-rewriter subagent; it fetches the target's
clustering keys via DESCRIBE DETAIL and narrows the ON predicate so writers
touch disjoint file sets.
Map the symptom to the failure class and its exact error code, then get the recovery tier from the decision tree:
file-not-found → DELTA_FILE_NOT_FOUND_DETAILED (VACUUM deleted pinned files — D03)uuid-changed → DIFFERENT_DELTA_TABLE_READ_BY_STREAMING_SOURCE (CREATE OR REPLACE minted a new UUID — D12)checkpoint-reset → silent batchId regression / checkpoint corruption (D04)transient → a restartable blip with an intact checkpointpython3 "${CLAUDE_SKILL_DIR}/scripts/recover-streaming-source.py" \
--failure file-not-found --time-travel yes # or uuid-changed / checkpoint-reset / transientIt echoes the canonical error code and recommends SAFE_RESTART /
REPROCESS_FROM_OFFSET / RESTORE_FROM_TIME_TRAVEL / FULL_RESET_BACKFILL with the
data-loss tradeoff stated up front. Name that full code in your answer — not just
the short class. The full
three-tier reasoning is in
${CLAUDE_SKILL_DIR}/references/checkpoint-recovery.md.
A driver/executor OOM while the JVM heap looks healthy points at off-heap RocksDB
state. Correlate the OOM to state size with clusters_events, then bound the
memory and enable changelog checkpointing per
${CLAUDE_SKILL_DIR}/references/rocksdb-state-store-tuning.md.
A stream stopping with UnknownFieldException on a new column is the default
addNewColumns mode. Choose the mode deliberately (evolve-and-restart vs
rescue's silent widening) and pin types with schemaHints per
${CLAUDE_SKILL_DIR}/references/autoloader-schema-evolution.md.
Before a DLT full refresh, confirm every source is replayable (a Kafka topic past
retention or a truncate-and-load source loses data on refresh — D09) and that
@dlt.table registration is deterministic (the thread race — D08). Read the DLT
event log with pipelines_get; the checklist is in
${CLAUDE_SKILL_DIR}/references/dlt-rebuild-safety.md.
merge-rewriter) that stops ConcurrentAppendException.| Error | Cause | Solution |
|---|---|---|
ConcurrentDeleteDeleteException | Manual OPTIMIZE races auto-compaction (D01) | Run pre-optimize-check.sh; don't manually OPTIMIZE an auto-compacted table, or disable auto-compact first. |
ConcurrentAppendException on an LC table | MERGE predicate not scoped to clustering keys (D02) | Route the MERGE to merge-rewriter; narrow the ON predicate to the clustering keys. |
DELTA_FILE_NOT_FOUND_DETAILED | VACUUM deleted checkpoint-pinned files (D03) | recover-streaming-source.py --failure file-not-found; restore via time travel if in retention, else reprocess/reset. |
DIFFERENT_DELTA_TABLE_READ_BY_STREAMING_SOURCE | CREATE OR REPLACE minted a new UUID (D12) | The old checkpoint is dead — new checkpoint + backfill; the hook prevents this going forward. |
| Driver OOM, heap looks fine | Off-heap RocksDB state (D05) | Bound state-store memory + changelog checkpointing; size state with a watermark. |
UnknownFieldException, stream stopped | Auto Loader addNewColumns default (D10) | Restart to evolve (idempotent sink), or choose rescue/schemaHints deliberately. |
| Hook allowed a destructive op with a warning | Could not verify consumers (no CLI/warehouse) | Advisory — the hook fails open; verify system.streaming.query_progress manually before running it. |
The PreToolUse hook fires, confirms 2 active consumers via
system.streaming.query_progress, and blocks with: "CREATE OR REPLACE mints a
new UUID → both consumers die with DIFFERENT_DELTA_TABLE_READ_BY_STREAMING_SOURCE;
use ALTER / in-place."
The merge-rewriter subagent reads the target's clustering keys via DESCRIBE DETAIL and rewrites the ON predicate to include them, so concurrent writers
touch disjoint file sets — the exception stops without serializing the jobs.
recover-streaming-source.py --failure file-not-found --time-travel yes →
RESTORE_FROM_TIME_TRAVEL (no data loss): restore the source to a pre-VACUUM version,
restart on the existing checkpoint, then align VACUUM retention with the checkpoint lag.
pre-optimize-check.sh --table main.sales.orders reports COLLISION RISK because
delta.autoOptimize.autoCompact is on — so the skill recommends letting
auto-compaction do it, or disabling it for the maintenance window first.
${CLAUDE_SKILL_DIR}/references/concurrency-conflicts.md — Delta OCC, auto-compaction collisions (D01), Liquid-Clustering writer conflicts (D02).${CLAUDE_SKILL_DIR}/references/checkpoint-recovery.md — the three-tier streaming checkpoint recovery (D03/D04/D12).${CLAUDE_SKILL_DIR}/references/rocksdb-state-store-tuning.md — bounded off-heap state + changelog checkpointing (D05).${CLAUDE_SKILL_DIR}/references/autoloader-schema-evolution.md — the schema-evolution modes + schemaHints (D10).${CLAUDE_SKILL_DIR}/references/dlt-rebuild-safety.md — DLT thread race, full-refresh data loss, tier cost (D08/D09/D11).${CLAUDE_SKILL_DIR}/scripts/pre-optimize-check.sh — auto-compaction collision probe.${CLAUDE_SKILL_DIR}/scripts/recover-streaming-source.py — 4-way recovery decision tree.${CLAUDE_SKILL_DIR}/hooks/streaming-guard-hook.py — the PreToolUse block for destructive ops on streamed-from tables.${CLAUDE_SKILL_DIR}/agents/merge-rewriter.md — rewrites a MERGE predicate for Liquid Clustering.© jeremylongshore, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 13 other files (scripts, references) in skills/.curated/databricks-streaming-guardian of jeremylongshore/tons-of-skills-marketplace.
Open the folder on GitHubat commit cfae287
Databricks Streaming Guardian next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Databricks Streaming Guardian this skilljeremylongshore/tons-of-skills-marketplace | 2.8k | — | ~3.5k | Automated safety check: Pass | MIT | |
| Git PR Flowwgzhao/Addax | 1.4k | — | ~216 | Automated safety check: Pass | Apache-2.0 | |
| Optimizing Databricks SQLAltimateAI/data-engineering-skills | 128 | — | ~6.7k | Automated safety check: Pass | MIT | |
| Replication Driven Researchbrycewang-stanford/Auto-Empirical-Research-Skills | 4.6k | — | ~1.7k | Automated safety check: Pass | Custom licence | |
| DB SculptorEliasOulkadi/shokunin | 114 | — | ~3.1k | Automated safety check: Notes | MIT | |
| Batch Processing Clinical Textmaziyarpanahi/openmed | 5.5k | — | ~2.2k | Automated safety check: Pass | Apache-2.0 |
wgzhao/Addax
Standard branch, commit, and pull-request workflow for this repo.
AltimateAI/data-engineering-skills
Analyze DBSQL queries, including SQL embedded in notebooks (spark.sql(...), %sql cells), for anti-patterns, lint issues, and performance problems, using Databricks-specific dialect and platform…
brycewang-stanford/Auto-Empirical-Research-Skills
A skill your agent uses when starting empirical analysis, creating a data pipeline, generating results, or when data or model specifications change.
EliasOulkadi/shokunin
Design database schemas with Prisma/Drizzle, PostgreSQL index strategy (B-tree, GIN, GiST, BRIN, Hash), query optimization (EXPLAIN ANALYZE), migration safety (expand/contract, zero-downtime), and…
maziyarpanahi/openmed
Run large-scale batch NER, PII extraction, or de-identification over many clinical notes on-device with OpenMed, with sharding, checkpointing, resumability, and append-only JSONL output.
astronomer/agents
Persists task and asset state across retries and DAG runs using Airflow 3.3's AIP-103 key/value stores (taskstatestore, assetstatestore) and the crash-safe ResumableJobMixin.
jeremylongshore/tons-of-skills-marketplace
Execute this skill enables AI assistant to conduct a security-focused code review using the security-agent plugin.
jeremylongshore/tons-of-skills-marketplace
Build this skill automates the adaptation of pre-trained machine learning models using transfer learning techniques.
jeremylongshore/tons-of-skills-marketplace
Execute proactive auto-loading: automatically detects and loads agents.md files.
jeremylongshore/tons-of-skills-marketplace
Aggregate and centralize performance metrics from applications, systems, databases, caches, and services.
jeremylongshore/tons-of-skills-marketplace
Execute this skill enables AI assistant to analyze capacity requirements and plan for future growth.
jeremylongshore/tons-of-skills-marketplace
Process use when you need to work with database indexing. An agent skill from jeremylongshore/tons-of-skills-marketplace.
Works with
Categories
Guard production Databricks data pipelines — Delta Lake, Liquid Clustering, Structured Streaming, Auto Loader, and DLT — against the twelve foot-guns that fire at scale: OPTIMIZE/auto-compaction…. Databricks Streaming Guardian is an agent skill from jeremylongshore/tons-of-skills-marketplace. Guard production Databricks data pipelines — Delta Lake, Liquid Clustering, Structured Streaming, Auto Loader, and DLT — against the twelve foot-guns that fire at scale: OPTIMIZE/auto-compaction conflicts, Liquid-Clustering merge conflicts, VACUUM breaking a streaming checkpoint, RocksDB OOM, Auto Loader schema-evolution stops, and DLT refresh data loss.
Databricks Streaming Guardian fits situations like: A Delta MERGE/OPTIMIZE fails with a concurrency exception; A stream breaks after VACUUM; A table replace; an Auto Loader stream stops on a new column.
Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill databricks-streaming-guardian -a claude-code`. Or copy the skill folder (skills/.curated/databricks-streaming-guardian in jeremylongshore/tons-of-skills-marketplace) into .claude/skills/databricks-streaming-guardian in your project. Claude Code loads it when a task matches its description.
Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill databricks-streaming-guardian -a codex`. Or copy the skill folder (skills/.curated/databricks-streaming-guardian in jeremylongshore/tons-of-skills-marketplace) into .agents/skills/databricks-streaming-guardian in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill databricks-streaming-guardian -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/databricks-streaming-guardian, .gemini/skills/databricks-streaming-guardian, .github/skills/databricks-streaming-guardian and .opencode/skills/databricks-streaming-guardian in your project.
Going by SKILL.md and its folder, Databricks Streaming Guardian needs Python and a shell for the scripts in its folder and the command-line tools its instructions call (bash, python3 and git). Our summary lists: Python 3; A Bash shell. Its frontmatter pre-approves these tools: Read, Write, Edit, Bash(databricks:*), Bash(jq:*), Bash(python3:*), Bash(bash:*), Glob, mcp__databricks-workspace-mcp__clusters_events, mcp__databricks-workspace-mcp__clusters_list, mcp__databricks-workspace-mcp__pipelines_get. Compatibility (from SKILL.md): Designed for Claude Code.
SKILL.md names 1 domain. As links in the text: docs.databricks.com. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Databricks Streaming Guardian is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 3.5k tokens (SKILL.md is roughly 14k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 19k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Databricks Streaming Guardian: Git PR Flow (wgzhao/Addax, 1.4k stars), Optimizing Databricks SQL (AltimateAI/data-engineering-skills, 128 stars), Replication Driven Research (brycewang-stanford/Auto-Empirical-Research-Skills, 4.6k stars) and DB Sculptor (EliasOulkadi/shokunin, 114 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
jeremylongshore (a GitHub user) maintains it in jeremylongshore/tons-of-skills-marketplace, which has 2,827 GitHub stars. The repository holds 3,342 skills in this directory. The repository was last updated on October 10, 2026.
Source: jeremylongshore/tons-of-skills-marketplace on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.