Agent skill

Databricks

by rocky-data in rocky-data/rocky

Databricks REST API and SQL reference for Rocky's warehouse adapter.

Apache-2.0Auto-check passedBackend & APIs

Install Databricks

skills CLI
$ npx skills add rocky-data/rocky --skill databricks -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install rocky-data/rocky databricks --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/rocky-data/rocky.git skills-src && mkdir -p .claude/skills && cp -r skills-src/engine/.claude/skills/databricks .claude/skills/databricks && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
databricks
GitHub stars
304
Token cost
~2k tokens
SKILL.md length
339 words
Files
1
Skills in repo
22
Repo updated
First seen
Licence
Apache-2.0

At a glance

Databricks REST API and SQL reference for Rocky's warehouse adapter.

  • Implementing SQL execution
  • SKILL.md covers SQL Statement Execution API, Authentication, Unity Catalog APIs and SQL Statements Rocky Must…, plus 2 more sections
  • Needs DATABRICKS_TOKEN and DATABRICKS_CLIENT_SECRET
  • Unity Catalog management

What it does

Databricks is an agent skill from rocky-data/rocky. Databricks REST API and SQL reference for Rocky's warehouse adapter. Use when implementing SQL execution, Unity Catalog management, workspace bindings, authentication, permission reconciliation, or any Databricks integration in the rocky-databricks crate.

Its SKILL.md is about 2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Backend & APIs, covering SQL, Accounting and bookkeeping and REST APIs. It works with Databricks and SQL. The repository describes itself as: A SQL transformation engine that type-checks your whole pipeline and catches breaking changes before they run — branches, replay, column-level lineage, compile-time contracts… The licence is Apache-2.0.

When your agent uses it

  • Implementing SQL execution
  • Unity Catalog management
  • Workspace bindings
  • Permission reconciliation

Example prompts

  • “Use the databricks skill to databrick REST API and SQL reference for Rocky's warehouse adapter”
  • “/databricks”

Requirements

  • A credential in DATABRICKS_TOKEN
  • A credential in DATABRICKS_CLIENT_SECRET

What it can do on your machine

Read from SKILL.md and the folder at commit 0cb7c7b. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are sql and json).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • DATABRICKS_TOKEN
    • DATABRICKS_CLIENT_SECRET

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Databricks loads about 2k tokens when it runs. Until then it costs about 67 tokens; SKILL.md has 339 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~67
When it runs · the whole SKILL.md, loaded when a task matches
~2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from rocky-data/rocky at commit 0cb7c7b, republished under its Apache-2.0 licence (© rocky-data). 339 words, ~2,021 tokens.

Download SKILL.mdSave it as .claude/skills/databricks/SKILL.md (or your agent's skills folder).
name
databricks
description
Databricks REST API and SQL reference for Rocky's warehouse adapter. Use when implementing SQL execution, Unity Catalog management, workspace bindings, authentication, permission reconciliation, or any Databricks integration in the rocky-databricks crate.

Databricks API Reference for Rocky

SQL Statement Execution API

Primary execution path for all SQL. No SDK needed — pure REST.

Submit Statement
POST https://{host}/api/2.0/sql/statements
Authorization: Bearer {token}
Content-Type: application/json

{
  "warehouse_id": "{warehouse_id}",
  "statement": "SELECT * FROM catalog.schema.table LIMIT 10",
  "wait_timeout": "30s",
  "disposition": "INLINE",
  "format": "JSON_ARRAY"
}

warehouse_id is extracted from the HTTP path: /sql/1.0/warehouses/{warehouse_id}

Response (immediate if fast):

json
{
  "statement_id": "abc-123",
  "status": { "state": "SUCCEEDED" },
  "manifest": {
    "schema": {
      "columns": [
        { "name": "col1", "type_name": "STRING", "position": 0 }
      ]
    },
    "total_row_count": 10
  },
  "result": {
    "data_array": [["value1"], ["value2"]]
  }
}
Poll Statement (if not immediately complete)
GET https://{host}/api/2.0/sql/statements/{statement_id}
Authorization: Bearer {token}

States: PENDING → RUNNING → SUCCEEDED | FAILED | CANCELED | CLOSED

Poll strategy: 100ms → 200ms → 500ms → 1s → 2s (exponential backoff, cap at 2s)

Cancel Statement
POST https://{host}/api/2.0/sql/statements/{statement_id}/cancel
Important Notes
  • wait_timeout: "0s" returns immediately with PENDING — useful for fire-and-forget
  • wait_timeout: "30s" waits up to 30s inline before returning — avoids polling for fast queries
  • disposition: "INLINE" returns data in response body (good for small results)
  • disposition: "EXTERNAL_LINKS" returns presigned URLs for large results (future: Arrow Flight)
  • Max statement size: 100KB
  • Max concurrent statements per warehouse: varies by warehouse size

Authentication

PAT (Personal Access Token)
Authorization: Bearer {DATABRICKS_TOKEN}
OAuth M2M (Service Principal)

Token request:

POST https://{host}/oidc/v1/token
Content-Type: application/x-www-form-urlencoded

grant_type=client_credentials&
client_id={DATABRICKS_CLIENT_ID}&
client_secret={DATABRICKS_CLIENT_SECRET}&
scope=all-apis

Response:

json
{
  "access_token": "eyJ...",
  "token_type": "Bearer",
  "expires_in": 3600
}

Implementation:

  • Cache the token
  • Refresh when expires_in is within 60s of expiry
  • Use the access_token as Authorization: Bearer {access_token}
Auto-Detection Logic
if DATABRICKS_TOKEN is set and non-empty:
    use PAT auth
else if DATABRICKS_CLIENT_ID and DATABRICKS_CLIENT_SECRET are set:
    use OAuth M2M
else:
    error: no auth configured

Unity Catalog APIs

Catalog Isolation
PATCH https://{host}/api/2.1/unity-catalog/catalogs/{catalog_name}
Authorization: Bearer {token}
Content-Type: application/json

{
  "isolation_mode": "ISOLATED"
}
Workspace Bindings

Get current bindings:

GET https://{host}/api/2.1/unity-catalog/bindings/catalog/{catalog_name}

Response:

json
{
  "bindings": [
    { "workspace_id": 12345, "binding_type": "BINDING_TYPE_READ_WRITE" }
  ]
}

Update bindings (add/remove):

PATCH https://{host}/api/2.1/unity-catalog/bindings/catalog/{catalog_name}
Content-Type: application/json

{
  "add": [
    { "workspace_id": 67890, "binding_type": "BINDING_TYPE_READ_WRITE" }
  ],
  "remove": [
    { "workspace_id": 11111 }
  ]
}

SQL Statements Rocky Must Generate

Catalog Lifecycle
sql
-- Create
CREATE CATALOG IF NOT EXISTS <catalog>

-- Tag (keys/values come from [governance.tags] in rocky.toml; 'managed_by' is always set)
ALTER CATALOG <catalog> SET TAGS (
    'managed_by' = '<pipeline_name>'
    -- plus any tags declared under [governance.tags]
)

-- Inspect
DESCRIBE CATALOG <catalog>
-- Returns rows: (info_name, info_value) — check for 'Catalog Name' row

-- Discover managed catalogs
SELECT catalog_name
FROM system.information_schema.catalog_tags
WHERE tag_name = 'managed_by' AND tag_value = '<pipeline_name>'
Schema Lifecycle
sql
-- Create
CREATE SCHEMA IF NOT EXISTS <catalog>.<schema>

-- Tag (keys/values come from [governance.tags]; Rocky always sets 'managed_by')
ALTER SCHEMA <catalog>.<schema> SET TAGS (
    'layer' = 'raw',
    'connector' = '<connector>',
    'managed_by' = '<pipeline_name>'
    -- plus any tags declared under [governance.tags]
)

-- List schemas
SHOW SCHEMAS IN <catalog>
Incremental Copy (Core Operation)

Rocky reads the prior watermark (MAX(<timestamp_col>) from the previous run) out of its redb state store and threads it into the generated SQL as a literal — it does not subquery the target table. See rocky-core/src/sql_gen.rs; a regression test asserts the generated SQL does not contain SELECT COALESCE(MAX(...)).

sql
-- Full refresh
SELECT * FROM <source_catalog>.<source_schema>.<table>

-- Incremental (append rows newer than the stored watermark)
SELECT * FROM <source_catalog>.<source_schema>.<table>
WHERE _fivetran_synced > TIMESTAMP '<prior_watermark_literal>'

After the copy, Rocky re-queries MAX(<timestamp_col>) from the source and records it as the next watermark.

Schema Drift Detection
sql
-- Get column info for comparison
DESCRIBE TABLE <catalog>.<schema>.<table>
-- Returns rows: (col_name, data_type, comment)

-- On a SAFE type widening (e.g. INT -> BIGINT): evolve in place
ALTER TABLE <target_catalog>.<target_schema>.<table> ALTER COLUMN <col> TYPE <wider_type>

-- On an UNSAFE type change: drop, then full refresh on next run
DROP TABLE IF EXISTS <target_catalog>.<target_schema>.<table>

The safe-vs-unsafe decision is drift.rs::is_safe_type_widening().

Permission Reconciliation
sql
-- Inspect current grants
SHOW GRANTS ON CATALOG <catalog>
-- Returns rows: (principal, action_type, object_type, object_name)

SHOW GRANTS ON SCHEMA <catalog>.<schema>

-- Grant (principal always backtick-quoted)
GRANT BROWSE ON CATALOG <catalog> TO `<principal>`
GRANT USE CATALOG ON CATALOG <catalog> TO `<principal>`
GRANT SELECT ON CATALOG <catalog> TO `<principal>`
GRANT USE SCHEMA ON SCHEMA <catalog>.<schema> TO `<principal>`

-- Revoke
REVOKE BROWSE ON CATALOG <catalog> FROM `<principal>`

-- Supported permission types for reconciliation:
-- BROWSE, USE CATALOG, USE SCHEMA, SELECT, MANAGE, MODIFY
-- Skip these (non-managed): OWNERSHIP, ALL PRIVILEGES, CREATE SCHEMA
Data Quality Checks
sql
-- Single table row count
SELECT COUNT(*) FROM <catalog>.<schema>.<table>

-- Batched row counts (UNION ALL, batches of 200)
SELECT 'cat1' AS c, 'sch1' AS s, 'tbl1' AS t, COUNT(*) AS cnt FROM cat1.sch1.tbl1
UNION ALL
SELECT 'cat1' AS c, 'sch1' AS s, 'tbl2' AS t, COUNT(*) AS cnt FROM cat1.sch1.tbl2
UNION ALL
...

-- Column introspection (batched by schema)
SELECT lower(table_schema), lower(table_name), lower(column_name)
FROM <catalog>.information_schema.columns
WHERE table_schema IN ('schema1', 'schema2', ...)
ORDER BY table_schema, table_name, ordinal_position

Validation Rules

SQL identifiers (catalogs, schemas, tables):

^[a-zA-Z0-9_]+$

Reject anything that doesn't match. Never use format!() with unvalidated strings.

Principal names (for GRANT/REVOKE):

^[a-zA-Z0-9_ \-\.@]+$

Always wrap in backticks: `principal_name`

Error Handling

Common Databricks errors to handle:

  • TEMPORARILY_UNAVAILABLE (503) — Retry with exponential backoff
  • INVALID_PARAMETER_VALUE — Bad SQL or missing object
  • RESOURCE_DOES_NOT_EXIST — Table/catalog/schema not found
  • PERMISSION_DENIED — Missing privileges
  • InvalidOperationHandle — Statement expired, re-submit
  • Rate limiting — Warehouse concurrency limit reached, back off

Retry strategy: 3 attempts, exponential backoff (1s → 3s → 9s), only on transient errors (503, rate limit, InvalidOperationHandle).

© rocky-data, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in engine/.claude/skills/databricks of rocky-data/rocky.

Open the folder on GitHubat commit 0cb7c7b

Compare with similar skills

Databricks next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Databricks compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Databricks this skillrocky-data/rocky304—~2kAutomated safety check: PassApache-2.0
Databricks Coredatabricks/databricks-agent-skills345—~1.9kAutomated safety check: PassCustom licence
Openobserve APIfcakyon/claude-codex-settings1.2k1 repos~4.1kAutomated safety check: PassApache-2.0
NpgsqlrestNpgsqlRest/NpgsqlRest132—~7kAutomated safety check: NotesMIT
Cookbook Authenticationdatabricks-solutions/databricks-apps-cookbook183—~1.8kAutomated safety check: PassCustom licence
Databricks Python SDKdatabricks/databricks-agent-skills3451 repos~4.6kAutomated safety check: PassCustom licence

Similar skills

  • Databricks Core

    databricks/databricks-agent-skills

    Official

    Databricks CLI operations and the parent/entry-point skill for Databricks CLI use: authentication, profile selection, and bundles.

    345 GitHub stars~1.9k tokensUpdated yesterday
    Backend & APIsAuto-check passed
  • Openobserve API

    fcakyon/claude-codex-settings

    This skill should be used when user asks to "query OpenObserve", "create OpenObserve dashboard", "edit OpenObserve panel", "fetch OpenObserve logs", "run OpenObserve search", "list OpenObserve…

    1.2k GitHub starsUsed in 1 repo~4.1k tokens
    Backend & APIsAuto-check passed
  • Npgsqlrest

    NpgsqlRest/NpgsqlRest

    Build and modify REST APIs with NpgsqlRest — exposing PostgreSQL as HTTP endpoints from two sources (database functions/procedures/tables/views, and plain .sql files), driven by SQL comment…

    132 GitHub stars~7k tokensUpdated 4 days ago
    DatabasesAuto-check: notes
  • Cookbook Authentication

    databricks-solutions/databricks-apps-cookbook

    Implement authentication in Databricks Apps: get current user info, on-behalf-of-user (OBO) auth, service principal auth, and OAuth patterns.

    183 GitHub stars~1.8k tokensUpdated 3 days ago
    Backend & APIsAuto-check passed
  • Databricks Python SDK

    databricks/databricks-agent-skills

    Official

    Databricks development guidance including Python SDK, Databricks Connect, CLI, and REST API.

    345 GitHub starsUsed in 1 repo~4.6k tokens
    Backend & APIsAuto-check passed
  • Databricks CLI

    aehrc/pathling

    Expert guidance for using the Databricks CLI to manage Databricks workspaces, clusters, jobs, pipelines, Unity Catalog, SQL warehouses, serving endpoints, secrets, bundles, and all other Databricks…

    137 GitHub stars~2.1k tokensUpdated today
    Game DevelopmentAuto-check passed

More from rocky-data/rocky

All 22 skills in this repo
  • Fivetran

    rocky-data/rocky

    Fivetran REST API reference for Rocky's source adapter. An agent skill from rocky-data/rocky.

    304 GitHub stars~914 tokensUpdated today
    Auto-check passed
  • Rocky Codegen

    rocky-data/rocky

    Rocky CLI JSON-output schema cascade. An agent skill from rocky-data/rocky.

    304 GitHub stars~1.9k tokensUpdated today
    Auto-check passed
  • Rocky Dev

    rocky-data/rocky

    Top-level router for Rocky development tasks. An agent skill from rocky-data/rocky.

    304 GitHub stars~2.1k tokensUpdated today
    Auto-check passed
  • Rocky Dsl Change

    rocky-data/rocky

    Rocky DSL (.rocky file) cross-subproject cascade. An agent skill from rocky-data/rocky.

    304 GitHub stars~1.3k tokensUpdated today
    Auto-check passed
  • Rocky New Adapter

    rocky-data/rocky

    Adding a new warehouse or source adapter crate to the Rocky engine.

    304 GitHub stars~2k tokensUpdated today
    Auto-check passed
  • Rocky New CLI Command

    rocky-data/rocky

    End-to-end checklist for adding a new Rocky CLI subcommand across the engine, JSON schema export, Dagster Pydantic types, Dagster resource wiring, and VS Code extension command.

    304 GitHub stars~1.4k tokensUpdated today
    Auto-check passed

Works with

Questions about Databricks

What does Databricks do?

Databricks REST API and SQL reference for Rocky's warehouse adapter. Databricks is an agent skill from rocky-data/rocky. Databricks REST API and SQL reference for Rocky's warehouse adapter.

When should I use Databricks?

Databricks fits situations like: implementing SQL execution; unity Catalog management; workspace bindings; permission reconciliation.

How do I install Databricks in Claude Code?

Run `npx skills add rocky-data/rocky --skill databricks -a claude-code`. Or copy the skill folder (engine/.claude/skills/databricks in rocky-data/rocky) into .claude/skills/databricks in your project. Claude Code loads it when a task matches its description.

How do I install Databricks in Codex?

Run `npx skills add rocky-data/rocky --skill databricks -a codex`. Or copy the skill folder (engine/.claude/skills/databricks in rocky-data/rocky) into .agents/skills/databricks in your project. Codex loads it when a task matches its description.

Can I use Databricks in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add rocky-data/rocky --skill databricks -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/databricks, .gemini/skills/databricks, .github/skills/databricks and .opencode/skills/databricks in your project.

What does Databricks need to run?

Going by SKILL.md and its folder, Databricks needs credentials named DATABRICKS_TOKEN and DATABRICKS_CLIENT_SECRET. Our summary lists: A credential in DATABRICKS_TOKEN; A credential in DATABRICKS_CLIENT_SECRET.

Does Databricks access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Databricks safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Databricks use?

Databricks is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Databricks use?

About 2k tokens (SKILL.md is roughly 8.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Databricks?

Skills that share tags, products or a category with Databricks: Databricks Core (databricks/databricks-agent-skills, 345 stars), Openobserve API (fcakyon/claude-codex-settings, 1.2k stars), Npgsqlrest (NpgsqlRest/NpgsqlRest, 132 stars) and Cookbook Authentication (databricks-solutions/databricks-apps-cookbook, 183 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Databricks?

rocky-data (a GitHub organization) maintains it in rocky-data/rocky, which has 304 GitHub stars. The repository holds 22 skills in this directory. The repository was last updated on October 7, 2026.

Source: rocky-data/rocky on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.