Ade Bench Cross DB Tasks
dbt-labs/ade-bench
A skill your agent uses when authoring or debugging ade-bench tasks that must run on both DuckDB and Snowflake, including shared project migrations, setup patches, and solution patches
Canonical rocky.toml authoring reference. An agent skill from rocky-data/rocky.
$ npx skills add rocky-data/rocky --skill rocky-config -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install rocky-data/rocky rocky-config --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/rocky-data/rocky.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/rocky-config .claude/skills/rocky-config && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "rocky-config" agent skill from https://github.com/rocky-data/rocky/tree/main/.agents/skills/rocky-config into .claude/skills/rocky-config/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "rocky-config", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/rocky-data/rocky/tree/main/.agents/skills/rocky-configType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add rocky-data/rocky --skill rocky-config -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install rocky-data/rocky rocky-config --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/rocky-data/rocky.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.agents/skills/rocky-config .agents/skills/rocky-config && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "rocky-config" agent skill from https://github.com/rocky-data/rocky/tree/main/.agents/skills/rocky-config into .agents/skills/rocky-config/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "rocky-config", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add rocky-data/rocky --skill rocky-config -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install rocky-data/rocky rocky-config --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/rocky-data/rocky.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.agents/skills/rocky-config .cursor/skills/rocky-config && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "rocky-config" agent skill from https://github.com/rocky-data/rocky/tree/main/.agents/skills/rocky-config into .cursor/skills/rocky-config/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "rocky-config", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/rocky-data/rocky.git --path .agents/skills/rocky-config--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add rocky-data/rocky --skill rocky-config -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install rocky-data/rocky rocky-config --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/rocky-data/rocky.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.agents/skills/rocky-config .gemini/skills/rocky-config && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "rocky-config" agent skill from https://github.com/rocky-data/rocky/tree/main/.agents/skills/rocky-config into .gemini/skills/rocky-config/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "rocky-config", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install rocky-data/rocky rocky-configInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add rocky-data/rocky --skill rocky-config -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/rocky-data/rocky.git skills-src && mkdir -p .github/skills && cp -r skills-src/.agents/skills/rocky-config .github/skills/rocky-config && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "rocky-config" agent skill from https://github.com/rocky-data/rocky/tree/main/.agents/skills/rocky-config into .github/skills/rocky-config/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "rocky-config", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add rocky-data/rocky --skill rocky-config -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install rocky-data/rocky rocky-config --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/rocky-data/rocky.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.agents/skills/rocky-config .opencode/skills/rocky-config && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "rocky-config" agent skill from https://github.com/rocky-data/rocky/tree/main/.agents/skills/rocky-config into .opencode/skills/rocky-config/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "rocky-config", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
rocky-configCanonical rocky.toml authoring reference. An agent skill from rocky-data/rocky.
Rocky Config is an agent skill from rocky-data/rocky. Canonical rocky.toml authoring reference. Use when writing or reviewing a Rocky pipeline config — covers the 4 pipeline types (replication, transformation, quality, snapshot), adapter variants (duckdb/databricks/snowflake/fivetran), minimal-config defaults, env-var substitution, governance, checks, hooks, and the ${VAR:-default} syntax.
Its SKILL.md is about 6.9k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Databases, covering Data warehousing. It works with Databricks, DuckDB and Snowflake. The repository describes itself as: A SQL transformation engine that type-checks your whole pipeline and catches breaking changes before they run — branches, replay, column-level lineage, compile-time contracts… The licence is Apache-2.0.
Read from SKILL.md and the folder at commit 365aebf. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are toml and sql).
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
DATABRICKS_TOKENDATABRICKS_CLIENT_SECRETSNOWFLAKE_OAUTH_TOKENSNOWFLAKE_PASSWORDCLICKHOUSE_PASSWORDMSSQL_PASSWORDMSSQL_ACCESS_TOKENAZURE_CLIENT_SECRETFIVETRAN_API_KEYFIVETRAN_API_SECRETFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Rocky Config loads about 6.9k tokens when it runs. Until then it costs about 88 tokens; SKILL.md has 1,535 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from rocky-data/rocky at commit 365aebf, republished under its Apache-2.0 licence (© rocky-data). 1,535 words, ~6,928 tokens.
.claude/skills/rocky-config/SKILL.md (or your agent's skills folder).Rocky reads one config file — rocky.toml — for everything: adapters, pipelines, governance, state backend, cache. The Rust source of truth is engine/crates/rocky-core/src/config.rs. This skill is the canonical authoring reference.
rocky.toml from scratch (POC, example, production)[source] / [warehouse] — those are pre-Phase-2 and no longer work)Two mandatory sections ([adapter] + at least one [pipeline.<name>]) plus optional global blocks:
[adapter] # Warehouse / source connection
# …
[pipeline.<name>] # One or more pipelines — discriminated by `type` (default: replication)
# …
# Optional globals:
[state] # Embedded state store backend
[cache.schemas] # Schema (DESCRIBE) cache: enabled, ttl_seconds, replicate, trusted_max_age_seconds, strict_sources
[cost] # Cost model for `rocky optimize`
[hook.<event>] # Lifecycle hooks (one per event)
# Governance (tags, grants, workspace bindings) is NOT a top-level table:
# it lives under [pipeline.<name>.target.governance]. A top-level
# [governance] is refused (deny_unknown_fields).Every string value supports ${VAR_NAME} at parse time, with optional default:
token = "${DATABRICKS_TOKEN}" # hard-required
namespace = "${ROCKY_NAMESPACE:-default}" # default when unset
workspace = "${WORKSPACE_IDS:-}" # default to empty stringSubstitution happens in rocky-core/src/config.rs before serde sees the value.
An unnamed [adapter] with a type key auto-wraps as adapter.default. Pipeline adapter refs default to "default" — so you can omit adapter = "default" lines everywhere.
Top-level adapter fields are strict (deny_unknown_fields — typos are parse errors). Adapter-specific keys Rocky doesn't model go under a nested [adapter.<name>.extra] table, which passes through untouched for a custom or process adapter to read.
kind fieldkind declares the role an [adapter.*] block plays. Two valid values: "data" (warehouse read/write) and "discovery" (metadata enumeration).
| Adapter type | kind rule |
|---|---|
databricks, snowflake, bigquery, postgres, redshift, clickhouse, sqlserver | Optional — defaults to "data". Setting "discovery" is a parse error. |
fivetran, airbyte, iceberg, manual | Required — must be "discovery". Omitting it is a parse error: these adapters have no data path. |
duckdb | Optional — absent means "register both roles" (the common DuckDB case). Setting "data" or "discovery" narrows to a single role. |
Requiring kind on discovery-only adapter types is deliberate: a reader should be able to tell from the raw config alone that [adapter.fivetran] is metadata-only, without knowing the Rust trait surface.
[adapter]
type = "duckdb"
path = "playground.duckdb" # omit for in-memory; required if also used for discovery[adapter]
type = "databricks"
host = "${DATABRICKS_HOST}" # e.g. dbc-xxxx.cloud.databricks.com (no https://)
http_path = "${DATABRICKS_HTTP_PATH}" # /sql/1.0/warehouses/<warehouse_id>
[adapter.auth]
token = "${DATABRICKS_TOKEN}" # PAT (tried first)
# client_id = "${DATABRICKS_CLIENT_ID}" # OAuth M2M (fallback)
# client_secret = "${DATABRICKS_CLIENT_SECRET}"[adapter]
type = "snowflake"
account = "${SNOWFLAKE_ACCOUNT}"
username = "${SNOWFLAKE_USER}"
[adapter.auth]
# OAuth (pre-supplied token, highest priority):
# token = "${SNOWFLAKE_OAUTH_TOKEN}"
# RS256 key-pair JWT (preferred for service principals):
private_key_path = "${SNOWFLAKE_KEY_PATH}"
# Password (lowest priority):
# password = "${SNOWFLAKE_PASSWORD}"sslmode)[adapter]
type = "postgres" # or "redshift" (beta)
host = "${PGHOST}" # host or host:port (default 5432 / 5439)
database = "analytics" # the connected database = the only valid catalog
username = "${PGUSER}"
password = "${PGPASSWORD}"
# timeout_secs = 300 # connect timeout + statement_timeout
[adapter.extra] # unknown keys are refused
sslmode = "verify-full" # disable | prefer (default) | require | verify-full
# sslrootcert = "/etc/ssl/rds-ca.pem" # extra PEM roots under verify-full
# port = 6543 # wins over host:port
# max_connections = 8
# merge_mode = "on_conflict" # postgres < 15 only; needs a unique index on unique_key
# late_binding_views = true # redshift only: views WITH NO SCHEMA BINDINGRedshift models can set table attributes in the sidecar [redshift] block (dist_style, dist_key, sort_key, sort_style); rocky compile validates it (E052 / W052) and any other adapter refuses it at SQL generation. IAM auth for Redshift is not built in yet.
[adapter]
type = "clickhouse"
host = "${CLICKHOUSE_HOST}" # host or host:port (default 8123, 8443 with secure)
username = "${CLICKHOUSE_USER}" # default "default"
password = "${CLICKHOUSE_PASSWORD}"
# database = "default" # session default database for unqualified names
# timeout_secs = 300 # per statement; also max_execution_time
[adapter.extra] # unknown keys are refused
secure = true # HTTPS, certificate always verified
# ca_cert = "/etc/ssl/ch-ca.pem" # extra PEM roots (needs secure = true)
# port = 9443 # wins over host:portA Rocky schema is a ClickHouse database: set every model's catalog = "" (a non-empty catalog is refused). Models can set [clickhouse] (engine = a parameterless MergeTree-family name, order_by = columns, partition_by = a column or fn(column)); rocky compile validates it (E053 / W053). merge and incremental with unique_key are refused (E053, no MERGE), as are snapshots (E049), UDFs (E051) and materialized_view.
[adapter]
type = "sqlserver"
host = "${MSSQL_HOST}" # host, host,port or host:port (default 1433)
database = "analytics" # the connected database = the only valid catalog
# exactly ONE auth method:
username = "${MSSQL_USER}" # SQL auth (not on Fabric)
password = "${MSSQL_PASSWORD}"
# oauth_token = "${MSSQL_ACCESS_TOKEN}" # Entra ID access token (not refreshed)
# client_id = "${AZURE_CLIENT_ID}" # Entra ID service principal, with extra.tenant_id
# client_secret = "${AZURE_CLIENT_SECRET}"
[adapter.extra] # unknown keys are refused
# encrypt = "mandatory" # mandatory (default) | strict (TDS 8.0) | optional
# trust_server_certificate = true # local/test servers only
# ca_cert = "/etc/ssl/corp-ca.pem"
# flavor = "fabric" # Fabric Warehouse renderings
# tenant_id = "${AZURE_TENANT_ID}"
# port = 1433
# max_connections = 8Snapshots, materialized_view, UDFs (E051) and regex_match checks are refused on SQL Server.
[adapter.fivetran]
type = "fivetran"
kind = "discovery" # required: fivetran has no data path
destination_id = "${FIVETRAN_DESTINATION_ID}"
api_key = "${FIVETRAN_API_KEY}"
api_secret = "${FIVETRAN_API_SECRET}"Use this block in pipeline.*.source.discovery.adapter to let Rocky query the Fivetran REST API for the list of schemas to sync. Data itself flows through whichever warehouse adapter is referenced by pipeline.*.source.adapter (usually Databricks or Snowflake — the destination Fivetran populates).
[adapter.warehouse]
type = "databricks"
# …
[adapter.source]
type = "fivetran"
kind = "discovery"
# …
[pipeline.raw]
[pipeline.raw.source]
adapter = "source" # ref by name
[pipeline.raw.target]
adapter = "warehouse"Pipelines are discriminated by type. Default is "replication", so a pipeline block with no type is a replication pipeline.
type | What it does | Canonical use |
|---|---|---|
replication (default) | Copies tables from source to target with schema-pattern discovery + incremental/full-refresh strategy | Raw/Bronze layer |
transformation | Runs user SQL models against the warehouse and materializes them | Silver/Gold layer |
quality | Runs data quality checks against existing tables | Standalone QA runs |
snapshot | SCD-Type-2 snapshots with history tracking | Slowly-changing dimensions |
All four share execution, checks, depends_on, and target.adapter. They differ in what source/target shapes look like.
[pipeline.raw]
# type = "replication" # optional, this is the default
strategy = "incremental" # or "full_refresh" or "merge"
timestamp_column = "_fivetran_synced"
# merge_keys = ["id"] # required when strategy = "merge"
# merge_keys_fallback = ["id"] # used when merge_keys is unset
metadata_columns = [
{ name = "_loaded_by", type = "STRING", value = "'rocky'" },
]
[pipeline.raw.source]
# adapter = "default" # optional — first adapter by default
[pipeline.raw.source.schema_pattern]
prefix = "src__"
separator = "__"
components = ["client", "regions...", "connector"]
# "name" = single segment
# "name..." = variable-length (1+)
# Reserved: `table` and `id` are not allowed as component names
# (they're consumed by `--filter` and `[[table_overrides]]`).
[pipeline.raw.target]
catalog_template = "{client}_warehouse"
schema_template = "staging__{regions}__{connector}"
# Per-`(connector, table)` overrides on top of the pipeline defaults.
# Per-field most-specific-match-wins: each field is supplied by the
# most-specific matching rule that sets it; unrelated rules don't
# clobber it. See `engine/crates/rocky-core/src/config.rs` ::
# `TableOverride` for the full surface.
# Skip diagnostic tables on every connector.
[[pipeline.raw.table_overrides]]
match.table = "_diagnostics_*" # `*`/`?` glob in TOML only — CLI literals
enabled = false
# `pii_users` on one specific connector needs a composite merge key.
[[pipeline.raw.table_overrides]]
match.connector = "stripe_main" # matches conn.id OR conn.schema
match.table = "pii_users"
merge_keys = ["user_id", "tenant_id"]
# `audit_log` is append-only — keep it on incremental even when
# the pipeline default is `merge`.
[[pipeline.raw.table_overrides]]
match.connector = "stripe_main"
match.table = "audit_log"
strategy = "incremental"
timestamp_column = "occurred_at"CLI symmetry: --filter table=<literal> filters to one table at a
time (no globs at the CLI; shell quoting is too fragile). The
connector-level --filter id=<conn_id> keys remain unchanged.
[pipeline.silver]
type = "transformation"
models = "models/silver/**" # glob, relative to rocky.toml; default "models/**"
depends_on = ["raw"]
# There is no contracts key: a model's contract is auto-discovered as the
# sibling `<model>.contract.toml`, or passed at the CLI via `--contracts <dir>`.
# A transformation target takes an adapter ref (plus optional governance)
# and nothing else — `catalog`/`schema` here are parse errors
# (`deny_unknown_fields`). Each model names its own destination in its
# sidecar `[target]`:
#
# # models/silver/dim_customers.toml
# [target]
# catalog = "analytics"
# schema = "marts"
# table = "dim_customers"
[pipeline.silver.target]
adapter = "warehouse"Declare the external tables the models read, with optional dbt-style freshness
checked by rocky freshness (exit 1 on error / runtime_error):
[[pipeline.silver.sources]]
schema = "raw" # catalog optional (two-part names)
table = "orders"
[pipeline.silver.sources.freshness]
loaded_at_field = "_loaded_at" # DATE / TIMESTAMP column
warn_after = "12h" # <N>s | <N>h | <N>d; at least one of the two
error_after = "24h" # must be >= warn_after (else E050)
filter = "status <> 'test'" # optional WHERE predicate; no `;`[pipeline.qa]
type = "quality"
depends_on = ["silver"]
# A quality target is an adapter ref only. The tables to check are listed
# in `[[pipeline.qa.tables]]` (omit `table` to check the whole schema).
[pipeline.qa.target]
adapter = "warehouse"
[[pipeline.qa.tables]]
catalog = "analytics"
schema = "marts"
# Quality runs execute `row_count`, `custom`, and `assertions` only. Other
# check kinds are inert here — `rocky validate` warns (V034).
[pipeline.qa.checks]
row_count = true[pipeline.dim_history]
type = "snapshot"
unique_key = ["customer_id"] # required: row identity in the source table
updated_at = "updated_at" # required: change-detection column
depends_on = ["silver"]
[pipeline.dim_history.source]
catalog = "analytics"
schema = "marts"
table = "dim_customers"
[pipeline.dim_history.target]
catalog = "analytics_history"
schema = "snapshots"
table = "dim_customers_history" # required: explicit history tableThe parser applies sane defaults — keep configs lean by omitting anything that matches the default:
| Field | Default | Omit unless |
|---|---|---|
pipeline.type | "replication" | You need transformation/quality/snapshot |
adapter = "default" (in pipeline source/target) | First adapter | Multi-adapter config |
[state]\nbackend = "local" | local (embedded redb) | Using S3, Valkey, or tiered state sync |
auto_create_catalogs = false | false | You want Rocky to CREATE CATALOG |
auto_create_schemas = false | false | You want Rocky to CREATE SCHEMA |
Model name in sidecar .toml | filename stem | You want a different logical name |
Model target.table | name | Renaming on write |
Directory-level target | models/_defaults.toml inherited | Overriding per-model |
Every pipeline type accepts a [checks] block, but execution differs:
replication runs the full set; quality runs row_count, custom, and
assertions only; transformation, snapshot, and load run no pipeline-level
checks today (executed_check_kinds in config.rs is the source of
truth). rocky validate warns (V034) on a check the pipeline type never
executes.
[pipeline.<name>.checks]
enabled = true
row_count = true # source vs target row count
column_match = true # source vs target column list
freshness = { threshold_seconds = 86400 } # max staleness of newest row
null_rate = { columns = ["email"], threshold = 0.05, sample_percent = 10 }
[[pipeline.<name>.checks.custom]]
name = "no_future_dates"
sql = "SELECT COUNT(*) FROM {target} WHERE created_at > CURRENT_TIMESTAMP()"
threshold = 0 # max failing rowsEnforced on load pipelines: the load runs into a staging table, validates the landed schema against the contract, and promotes to the target only on pass — a violation drops staging and fails, so non-conforming data never lands. Types compare in Rocky's normalized vocabulary, so one contract ports across warehouses.
[pipeline.<name>.contract]
required_columns = [
{ name = "id", type = "BIGINT", nullable = false },
]
protected_columns = ["id", "email"] # may not be removed from source
allowed_type_changes = [
{ from = "INT", to = "BIGINT" }, # widening allowlist
][pipeline.<name>.execution]
concurrency = 16
fail_fast = false
error_rate_abort_pct = 50
table_retries = 1
# adaptive_concurrency is planned but not yet a config field.
# The AIMD throttle primitive exists but isn't wired.Governance is configured per pipeline target, under [pipeline.<name>.target.governance]. There is no top-level [governance] table; one is refused at load (deny_unknown_fields on RockyConfig).
[pipeline.bronze.target.governance]
auto_create_catalogs = true
auto_create_schemas = true
[pipeline.bronze.target.governance.tags]
managed_by = "rocky"
[[pipeline.bronze.target.governance.grants]]
principal = "data-readers"
permissions = ["BROWSE", "USE CATALOG", "USE SCHEMA", "SELECT"]
[pipeline.bronze.target.governance.isolation]
enabled = true
workspace_ids = "${WORKSPACE_IDS:-}"Permissions handled: BROWSE, USE CATALOG, USE SCHEMA, SELECT, MANAGE, MODIFY. Skipped: OWNERSHIP, ALL PRIVILEGES, CREATE SCHEMA (non-managed).
# Embedded redb (default — no config needed)
[state]
backend = "local"
# S3-backed state sync
[state]
backend = "s3"
bucket = "my-rocky-state"
prefix = "prod/"
region = "us-east-1"
# Tiered: local redb + S3 for durability
[state]
backend = "tiered"
# … S3 fields plus local path
# State-file namespacing (opt-in, default "none")
[state]
backend = "local"
namespacing = "pipeline" # each pipeline gets its own .rocky-state/<pipeline>.redbnamespacing (StateNamespacing, default "none") controls state-file fan-out. redb allows one writer per file, so running one rocky run per pipeline/client against the single global <models>/.rocky-state.redb serializes them on one lock. "pipeline" gives each pipeline its own <models>/.rocky-state/<pipeline>.redb. "none" is byte-identical to omitting the key. For per-client fan-out use the per-invocation --state-namespace <key> flag (overrides this config); an explicit --state-path disables namespacing for that run.
[run] + per-model [skip])# Opt-in model-skip gate — default OFF (omit the block to keep old behavior)
[run]
skip_unchanged = true # master switch (also via the --skip-unchanged flag)
skip_rowcount_fallback = false # default; allow COUNT(*) when no timestamp column (weaker signal)
lag_tolerance_seconds = 0 # default; any MAX(ts) movement forces a rebuild[run] (RunConfig) tunes the --skip-unchanged gate: skip re-materializing a transformation model whose logic and every upstream's data both appear unchanged since the last successful build. It is a best-effort optimization, not a result-equivalence guarantee — every field defaults to no-skip, and any missing/unreadable/ambiguous input rebuilds (fail-safe).
Not skip-eligible (always rebuild):
CURRENT_TIMESTAMP / NOW(), RANDOM(), UUID(), CURRENT_USER, CURRENT_CATALOG, ANY_VALUE, ARRAY_AGG, unordered LIMIT, or any unknown function.FROM (…), IN (SELECT …), EXISTS, scalar sub-selects), PIVOT / UNNEST / nested-join table-factors, and set operations (UNION / INTERSECT / EXCEPT).content_addressed / time_interval strategies (full_refresh is eligible).--force-rebuild bypasses the gate entirely.
# Per-model override (in the model's .toml sidecar)
name = "fct_orders"
[skip]
eligible = true # Some(false) = always build; Some(true) = eligible; unset = auto
deterministic = true # owner asserts SQL is pure → re-eligible despite the static scan[skip] (SkipConfig in rocky-core/src/models.rs) is a per-model sidecar block. Both fields are Option<bool> — unset means "trust the automatic rules."
--defer / --defer-to <schema> are runtime-only dev flags (not config): build the --model-selected models locally and resolve unbuilt upstream ref()s to a prod/defer schema. The rewrite parses model SQL with the Databricks dialect, so SELECT * EXCEPT(...), trailing-comma selects, and STRUCT(...) literals can't be rewritten — run those without --defer.
[reuse]) — EXPERIMENTAL, do not enable in production[reuse]
enabled = false # default; preview-only, NOT live-verified — leave off in production[reuse] (ReuseConfig) is a preview surface scoped to the Databricks–Iceberg content-addressed write path only (no DuckDB / Snowflake / BigQuery), and it is not yet live-verified against a warehouse. When enabled = true, a successful run only populates an input-match index + provenance record; the reuse decision path only ever resolves to BUILD today (an ONLY-BUILD posture) — a fail-closed verdict is computed but nothing is reused, since the live point-to reuse is not yet wired/verified. It attests an input-logic match + byte-identity of the recorded bytes, never that a fresh re-run would reproduce them. Default-off keeps rocky run byte- and cost-identical. The per-invocation --no-reuse flag forces every model to build. Provenance is auditable per docs/.../guides/verify-a-run.md.
[cache.schemas]
enabled = true # default; false for strict CI (every typecheck hits the warehouse)
ttl_seconds = 86400 # default 24h; lower for high-DDL-churn teams
replicate = false # default; true to ship the cache through state_sync
# trusted_max_age_seconds = 3600 # unset by default; entries younger than this make a missing source column E041
# strict_sources = true # default false; every W041 (seed / old cache entry) becomes E041[cache.schemas] is the only [cache] table (CacheConfig in config.rs): it stores DESCRIBE TABLE results in state.redb so leaf models typecheck against real warehouse types without a live round-trip on every compile. There is no [cache] valkey_url key: ValkeyCacheConfig exists as a type but is not wired into RockyConfig, and a [cache] table with any other key is refused (deny_unknown_fields). The Valkey tier is a [state] backend, not a cache setting.
rocky optimize)[cost]
storage_cost_per_gb_month = 0.023
compute_cost_per_dbu = 0.40
warehouse_size = "Medium"
min_history_runs = 5[hook.on_pipeline_start]
command = "scripts/notify.sh"
timeout_ms = 5000
on_failure = "warn" # default; or "abort" (stop the pipeline) / "ignore" (silent)
[hook.on_pipeline_fail]
url = "${SLACK_WEBHOOK_URL}" # webhook instead of command
template = "default" # built-in presetEvents: on_pipeline_start, on_pipeline_complete, on_pipeline_fail, on_model_start, on_model_complete, on_model_fail, on_check_fail, on_drift_detected.
.sql + .toml)Model files live under models/ and use a sidecar pattern:
# models/marts/dim_customers.toml
name = "dim_customers"
depends_on = ["stg_customers"]
[strategy]
type = "merge" # e.g. full_refresh, merge, delete_insert, time_interval, view, snapshot; not incremental (E037)
unique_key = ["customer_id"]
# update_columns = ["name", "email"] # omit for UPDATE SET *
[target]
catalog = "analytics"
schema = "marts"
table = "dim_customers"
[[sources]] # optional: declare source tables
catalog = "analytics"
schema = "staging"
table = "customers"-- models/marts/dim_customers.sql
-- Pure SQL. No Jinja. No templating.
SELECT customer_id, name, email, updated_at
FROM {{ analytics.staging.customers }} -- Rocky expands refs at compile time, not via JinjaDirectory-level defaults via models/<dir>/_defaults.toml:
[target]
catalog = "analytics"
schema = "marts"SQL identifiers (catalog, schema, table, tenant, region, source names) must match ^[a-zA-Z0-9_]+$. Rocky rejects anything else. Principal names for GRANT/REVOKE allow the broader ^[a-zA-Z0-9_ \-\.@]+$ pattern and are always wrapped in backticks in generated SQL.
For transformation models that need partition-by-partition execution:
# models/fact_events.toml
[strategy]
type = "time_interval"
time_column = "event_date"
granularity = "day" # hour | day | month | year
lookback = 7 # re-run last 7 partitions
batch_size = 4 # partitions per concurrent batch
first_partition = "2024-01-01"Model SQL uses @start_date and @end_date placeholders — Rocky substitutes per partition:
SELECT * FROM raw.events
WHERE event_date >= @start_date AND event_date < @end_dateCLI: rocky run --partition KEY / --from KEY --to KEY / --latest / --missing / --lookback N / --parallel N.
See examples/playground/pocs/00-foundations/01-replication-basics/rocky.toml for the minimal DuckDB replication case (schema-pattern routing), examples/playground/pocs/00-foundations/00-playground-default/rocky.toml for the minimal transformation case (model DAG), or engine/examples/multi-layer/rocky.toml for a full Bronze/Silver/Gold setup.
The following keys do not work anymore — they were the pre-Phase-2 config shape and will be rejected by the parser:
| ❌ Legacy | ✅ Current |
|---|---|
[source] top-level | [pipeline.<name>.source] |
[warehouse] | [adapter] |
[replication] top-level | [pipeline.<name>] with strategy = "incremental" |
[checks] top-level | [pipeline.<name>.checks] |
[target] top-level | [pipeline.<name>.target] |
If you see any of these in a config you're editing, the config is stale — migrate it.
engine/crates/rocky-core/src/config.rs — Rust source of truth for every fieldengine/AGENTS.md — "Configuration" section with full annotated exampleeditors/vscode/schemas/rocky-config.schema.json — JSON Schema for IDE autocompletion (autogenerated; keep in sync with config.rs)examples/playground/AGENTS.md — POC-specific minimal-config idioms© rocky-data, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .agents/skills/rocky-config of rocky-data/rocky.
Open the folder on GitHubat commit 365aebf
Rocky Config next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Rocky Config this skillrocky-data/rocky | 304 | — | ~6.9k | Automated safety check: Pass | Apache-2.0 | |
| Ade Bench Cross DB Tasksdbt-labs/ade-bench | 125 | — | ~2k | Automated safety check: Pass | Apache-2.0 | |
| Altimate Data Warehouse DelegateAltimateAI/data-engineering-skills | 128 | — | ~1.4k | Automated safety check: Pass | MIT | |
| SQL Queriesw95/awesome-claude-corporate-skills | 244 | 3 repos | ~2.8k | Automated safety check: Pass | MIT | |
| Connect Snowflakeai-analyst-lab/ai-analyst | 304 | — | ~1.4k | Automated safety check: Notes | MIT | |
| Databricks Icebergdatabricks/databricks-agent-skills | 345 | — | ~2.6k | Automated safety check: Pass | Custom licence |
dbt-labs/ade-bench
A skill your agent uses when authoring or debugging ade-bench tasks that must run on both DuckDB and Snowflake, including shared project migrations, setup patches, and solution patches
AltimateAI/data-engineering-skills
Delegates dbt and warehouse tasks such as lineage, migrations and cost attribution to the altimate-code CLI agent and relays its answer back.
w95/awesome-claude-corporate-skills
Write correct, performant SQL across all major data warehouse dialects (Snowflake, BigQuery, Databricks, PostgreSQL, etc.).
ai-analyst-lab/ai-analyst
Query the live/remote Snowflake warehouse instead of the local practice copy.
databricks/databricks-agent-skills
Apache Iceberg tables on Databricks — Managed Iceberg tables, External Iceberg Reads (fka Uniform), Compatibility Mode, Iceberg REST Catalog (IRC), Iceberg v3, Snowflake interop, PyIceberg, OSS…
aws/agent-toolkit-for-aws
Guides an end-to-end data-warehouse migration to Amazon Redshift — discovery, schema/SQL/stored-procedure/macro/script conversion, data migration, validation, performance comparison, and reporting.
rocky-data/rocky
Fivetran REST API reference for Rocky's source adapter. An agent skill from rocky-data/rocky.
rocky-data/rocky
Databricks REST API and SQL reference for Rocky's warehouse adapter.
rocky-data/rocky
Rocky CLI JSON-output schema cascade. An agent skill from rocky-data/rocky.
rocky-data/rocky
Top-level router for Rocky development tasks. An agent skill from rocky-data/rocky.
rocky-data/rocky
Rocky DSL (.rocky file) cross-subproject cascade. An agent skill from rocky-data/rocky.
rocky-data/rocky
Adding a new warehouse or source adapter crate to the Rocky engine.
Works with
Categories
Canonical rocky.toml authoring reference. An agent skill from rocky-data/rocky. Rocky Config is an agent skill from rocky-data/rocky.toml authoring reference.
Rocky Config fits situations like: reviewing a Rocky pipeline config — covers the 4 pipeline types (replication; adapter variants (duckdb/databricks/snowflake/fivetran); minimal-config defaults; env-var substitution.
Run `npx skills add rocky-data/rocky --skill rocky-config -a claude-code`. Or copy the skill folder (.agents/skills/rocky-config in rocky-data/rocky) into .claude/skills/rocky-config in your project. Claude Code loads it when a task matches its description.
Run `npx skills add rocky-data/rocky --skill rocky-config -a codex`. Or copy the skill folder (.agents/skills/rocky-config in rocky-data/rocky) into .agents/skills/rocky-config in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add rocky-data/rocky --skill rocky-config -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/rocky-config, .gemini/skills/rocky-config, .github/skills/rocky-config and .opencode/skills/rocky-config in your project.
Going by SKILL.md and its folder, Rocky Config needs credentials named DATABRICKS_TOKEN, DATABRICKS_CLIENT_SECRET, SNOWFLAKE_OAUTH_TOKEN and SNOWFLAKE_PASSWORD. Our summary lists: A credential in DATABRICKS_TOKEN; A credential in DATABRICKS_CLIENT_SECRET.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Rocky Config is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 6.9k tokens (SKILL.md is roughly 28k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Rocky Config: Ade Bench Cross DB Tasks (dbt-labs/ade-bench, 125 stars), Altimate Data Warehouse Delegate (AltimateAI/data-engineering-skills, 128 stars), SQL Queries (w95/awesome-claude-corporate-skills, 244 stars) and Connect Snowflake (ai-analyst-lab/ai-analyst, 304 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
rocky-data (a GitHub organization) maintains it in rocky-data/rocky, which has 304 GitHub stars. The repository holds 22 skills in this directory. The repository was last updated on October 10, 2026.
Source: rocky-data/rocky on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.