Agent skill

Rocky Config

by rocky-data in rocky-data/rocky

Canonical rocky.toml authoring reference. An agent skill from rocky-data/rocky.

Apache-2.0Auto-check passedDatabases

Install Rocky Config

skills CLI
$ npx skills add rocky-data/rocky --skill rocky-config -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install rocky-data/rocky rocky-config --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/rocky-data/rocky.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/rocky-config .claude/skills/rocky-config && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
rocky-config
GitHub stars
304
Token cost
~6.9k tokens
SKILL.md length
1,535 words
Files
1
Skills in repo
22
Repo updated
First seen
Licence
Apache-2.0

At a glance

Canonical rocky.toml authoring reference. An agent skill from rocky-data/rocky.

  • Reviewing a Rocky pipeline config — covers the 4 pipeline types (replication
  • SKILL.md covers When to use this skill, Top-level structure, Env var substitution and Adapter variants, plus 12 more sections
  • Needs DATABRICKS_TOKEN and DATABRICKS_CLIENT_SECRET
  • Adapter variants (duckdb/databricks/snowflake/fivetran)

What it does

Rocky Config is an agent skill from rocky-data/rocky. Canonical rocky.toml authoring reference. Use when writing or reviewing a Rocky pipeline config — covers the 4 pipeline types (replication, transformation, quality, snapshot), adapter variants (duckdb/databricks/snowflake/fivetran), minimal-config defaults, env-var substitution, governance, checks, hooks, and the ${VAR:-default} syntax.

Its SKILL.md is about 6.9k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Databases, covering Data warehousing. It works with Databricks, DuckDB and Snowflake. The repository describes itself as: A SQL transformation engine that type-checks your whole pipeline and catches breaking changes before they run — branches, replay, column-level lineage, compile-time contracts… The licence is Apache-2.0.

When your agent uses it

  • Reviewing a Rocky pipeline config — covers the 4 pipeline types (replication
  • Adapter variants (duckdb/databricks/snowflake/fivetran)
  • Minimal-config defaults
  • Env-var substitution

Example prompts

  • “/rocky-config”

Requirements

  • A credential in DATABRICKS_TOKEN
  • A credential in DATABRICKS_CLIENT_SECRET

What it can do on your machine

Read from SKILL.md and the folder at commit 365aebf. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are toml and sql).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • DATABRICKS_TOKEN
    • DATABRICKS_CLIENT_SECRET
    • SNOWFLAKE_OAUTH_TOKEN
    • SNOWFLAKE_PASSWORD
    • CLICKHOUSE_PASSWORD
    • MSSQL_PASSWORD
    • MSSQL_ACCESS_TOKEN
    • AZURE_CLIENT_SECRET
    • FIVETRAN_API_KEY
    • FIVETRAN_API_SECRET

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Rocky Config loads about 6.9k tokens when it runs. Until then it costs about 88 tokens; SKILL.md has 1,535 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~88
When it runs · the whole SKILL.md, loaded when a task matches
~6.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from rocky-data/rocky at commit 365aebf, republished under its Apache-2.0 licence (© rocky-data). 1,535 words, ~6,928 tokens.

Download SKILL.mdSave it as .claude/skills/rocky-config/SKILL.md (or your agent's skills folder).
name
rocky-config
description
Canonical `rocky.toml` authoring reference. Use when writing or reviewing a Rocky pipeline config — covers the 4 pipeline types (replication, transformation, quality, snapshot), adapter variants (duckdb/databricks/snowflake/fivetran), minimal-config defaults, env-var substitution, governance, checks, hooks, and the ${VAR:-default} syntax.

rocky.toml authoring reference

Rocky reads one config file — rocky.toml — for everything: adapters, pipelines, governance, state backend, cache. The Rust source of truth is engine/crates/rocky-core/src/config.rs. This skill is the canonical authoring reference.

When to use this skill

  • Writing a new rocky.toml from scratch (POC, example, production)
  • Migrating an old config to current shape (watch for [source] / [warehouse] — those are pre-Phase-2 and no longer work)
  • Deciding which pipeline type to use (replication vs transformation vs quality vs snapshot)
  • Looking up an adapter's required fields (databricks needs host + http_path, snowflake needs auth variants, etc.)
  • Adding governance, hooks, state sync, or checks to an existing config

Top-level structure

Two mandatory sections ([adapter] + at least one [pipeline.<name>]) plus optional global blocks:

toml
[adapter]                # Warehouse / source connection
# …

[pipeline.<name>]        # One or more pipelines — discriminated by `type` (default: replication)
# …

# Optional globals:
[state]                  # Embedded state store backend
[cache.schemas]          # Schema (DESCRIBE) cache: enabled, ttl_seconds, replicate, trusted_max_age_seconds, strict_sources
[cost]                   # Cost model for `rocky optimize`
[hook.<event>]           # Lifecycle hooks (one per event)
# Governance (tags, grants, workspace bindings) is NOT a top-level table:
# it lives under [pipeline.<name>.target.governance]. A top-level
# [governance] is refused (deny_unknown_fields).

Env var substitution

Every string value supports ${VAR_NAME} at parse time, with optional default:

toml
token     = "${DATABRICKS_TOKEN}"              # hard-required
namespace = "${ROCKY_NAMESPACE:-default}"      # default when unset
workspace = "${WORKSPACE_IDS:-}"               # default to empty string

Substitution happens in rocky-core/src/config.rs before serde sees the value.

Adapter variants

An unnamed [adapter] with a type key auto-wraps as adapter.default. Pipeline adapter refs default to "default" — so you can omit adapter = "default" lines everywhere.

Top-level adapter fields are strict (deny_unknown_fields — typos are parse errors). Adapter-specific keys Rocky doesn't model go under a nested [adapter.<name>.extra] table, which passes through untouched for a custom or process adapter to read.

The kind field

kind declares the role an [adapter.*] block plays. Two valid values: "data" (warehouse read/write) and "discovery" (metadata enumeration).

Adapter typekind rule
databricks, snowflake, bigquery, postgres, redshift, clickhouse, sqlserverOptional — defaults to "data". Setting "discovery" is a parse error.
fivetran, airbyte, iceberg, manualRequired — must be "discovery". Omitting it is a parse error: these adapters have no data path.
duckdbOptional — absent means "register both roles" (the common DuckDB case). Setting "data" or "discovery" narrows to a single role.

Requiring kind on discovery-only adapter types is deliberate: a reader should be able to tell from the raw config alone that [adapter.fivetran] is metadata-only, without knowing the Rust trait surface.

DuckDB (no credentials)
toml
[adapter]
type = "duckdb"
path = "playground.duckdb"        # omit for in-memory; required if also used for discovery
Databricks (PAT first, OAuth M2M fallback)
toml
[adapter]
type      = "databricks"
host      = "${DATABRICKS_HOST}"                 # e.g. dbc-xxxx.cloud.databricks.com (no https://)
http_path = "${DATABRICKS_HTTP_PATH}"            # /sql/1.0/warehouses/<warehouse_id>

[adapter.auth]
token         = "${DATABRICKS_TOKEN}"            # PAT (tried first)
# client_id     = "${DATABRICKS_CLIENT_ID}"      # OAuth M2M (fallback)
# client_secret = "${DATABRICKS_CLIENT_SECRET}"
Snowflake (OAuth > RS256 key-pair > password priority)
toml
[adapter]
type     = "snowflake"
account  = "${SNOWFLAKE_ACCOUNT}"
username = "${SNOWFLAKE_USER}"

[adapter.auth]
# OAuth (pre-supplied token, highest priority):
# token = "${SNOWFLAKE_OAUTH_TOKEN}"

# RS256 key-pair JWT (preferred for service principals):
private_key_path = "${SNOWFLAKE_KEY_PATH}"

# Password (lowest priority):
# password = "${SNOWFLAKE_PASSWORD}"
PostgreSQL / Redshift (password auth, libpq-style sslmode)
toml
[adapter]
type     = "postgres"                            # or "redshift" (beta)
host     = "${PGHOST}"                           # host or host:port (default 5432 / 5439)
database = "analytics"                           # the connected database = the only valid catalog
username = "${PGUSER}"
password = "${PGPASSWORD}"
# timeout_secs = 300                             # connect timeout + statement_timeout

[adapter.extra]                                  # unknown keys are refused
sslmode = "verify-full"                          # disable | prefer (default) | require | verify-full
# sslrootcert = "/etc/ssl/rds-ca.pem"            # extra PEM roots under verify-full
# port = 6543                                    # wins over host:port
# max_connections = 8
# merge_mode = "on_conflict"                     # postgres < 15 only; needs a unique index on unique_key
# late_binding_views = true                      # redshift only: views WITH NO SCHEMA BINDING

Redshift models can set table attributes in the sidecar [redshift] block (dist_style, dist_key, sort_key, sort_style); rocky compile validates it (E052 / W052) and any other adapter refuses it at SQL generation. IAM auth for Redshift is not built in yet.

ClickHouse (beta; HTTP interface, user/password, TLS)
toml
[adapter]
type     = "clickhouse"
host     = "${CLICKHOUSE_HOST}"                  # host or host:port (default 8123, 8443 with secure)
username = "${CLICKHOUSE_USER}"                  # default "default"
password = "${CLICKHOUSE_PASSWORD}"
# database = "default"                           # session default database for unqualified names
# timeout_secs = 300                             # per statement; also max_execution_time

[adapter.extra]                                  # unknown keys are refused
secure = true                                    # HTTPS, certificate always verified
# ca_cert = "/etc/ssl/ch-ca.pem"                 # extra PEM roots (needs secure = true)
# port = 9443                                    # wins over host:port

A Rocky schema is a ClickHouse database: set every model's catalog = "" (a non-empty catalog is refused). Models can set [clickhouse] (engine = a parameterless MergeTree-family name, order_by = columns, partition_by = a column or fn(column)); rocky compile validates it (E053 / W053). merge and incremental with unique_key are refused (E053, no MERGE), as are snapshots (E049), UDFs (E051) and materialized_view.

SQL Server / Azure SQL / Fabric Warehouse (beta)
toml
[adapter]
type     = "sqlserver"
host     = "${MSSQL_HOST}"                       # host, host,port or host:port (default 1433)
database = "analytics"                           # the connected database = the only valid catalog
# exactly ONE auth method:
username = "${MSSQL_USER}"                       # SQL auth (not on Fabric)
password = "${MSSQL_PASSWORD}"
# oauth_token = "${MSSQL_ACCESS_TOKEN}"          # Entra ID access token (not refreshed)
# client_id = "${AZURE_CLIENT_ID}"               # Entra ID service principal, with extra.tenant_id
# client_secret = "${AZURE_CLIENT_SECRET}"

[adapter.extra]                                  # unknown keys are refused
# encrypt = "mandatory"                          # mandatory (default) | strict (TDS 8.0) | optional
# trust_server_certificate = true                # local/test servers only
# ca_cert = "/etc/ssl/corp-ca.pem"
# flavor = "fabric"                              # Fabric Warehouse renderings
# tenant_id = "${AZURE_TENANT_ID}"
# port = 1433
# max_connections = 8

Snapshots, materialized_view, UDFs (E051) and regex_match checks are refused on SQL Server.

Fivetran (discovery-only)
toml
[adapter.fivetran]
type           = "fivetran"
kind           = "discovery"                     # required: fivetran has no data path
destination_id = "${FIVETRAN_DESTINATION_ID}"
api_key        = "${FIVETRAN_API_KEY}"
api_secret     = "${FIVETRAN_API_SECRET}"

Use this block in pipeline.*.source.discovery.adapter to let Rocky query the Fivetran REST API for the list of schemas to sync. Data itself flows through whichever warehouse adapter is referenced by pipeline.*.source.adapter (usually Databricks or Snowflake — the destination Fivetran populates).

Named adapters (multi-adapter configs)
toml
[adapter.warehouse]
type = "databricks"
# …

[adapter.source]
type = "fivetran"
kind = "discovery"
# …

[pipeline.raw]
[pipeline.raw.source]
adapter = "source"                 # ref by name

[pipeline.raw.target]
adapter = "warehouse"

Pipeline types

Pipelines are discriminated by type. Default is "replication", so a pipeline block with no type is a replication pipeline.

typeWhat it doesCanonical use
replication (default)Copies tables from source to target with schema-pattern discovery + incremental/full-refresh strategyRaw/Bronze layer
transformationRuns user SQL models against the warehouse and materializes themSilver/Gold layer
qualityRuns data quality checks against existing tablesStandalone QA runs
snapshotSCD-Type-2 snapshots with history trackingSlowly-changing dimensions

All four share execution, checks, depends_on, and target.adapter. They differ in what source/target shapes look like.

Replication pipeline
toml
[pipeline.raw]
# type = "replication"               # optional, this is the default
strategy         = "incremental"     # or "full_refresh" or "merge"
timestamp_column = "_fivetran_synced"
# merge_keys     = ["id"]            # required when strategy = "merge"
# merge_keys_fallback = ["id"]       # used when merge_keys is unset
metadata_columns = [
    { name = "_loaded_by", type = "STRING", value = "'rocky'" },
]

[pipeline.raw.source]
# adapter = "default"                # optional — first adapter by default
[pipeline.raw.source.schema_pattern]
prefix     = "src__"
separator  = "__"
components = ["client", "regions...", "connector"]
# "name"    = single segment
# "name..." = variable-length (1+)
# Reserved: `table` and `id` are not allowed as component names
# (they're consumed by `--filter` and `[[table_overrides]]`).

[pipeline.raw.target]
catalog_template = "{client}_warehouse"
schema_template  = "staging__{regions}__{connector}"

# Per-`(connector, table)` overrides on top of the pipeline defaults.
# Per-field most-specific-match-wins: each field is supplied by the
# most-specific matching rule that sets it; unrelated rules don't
# clobber it. See `engine/crates/rocky-core/src/config.rs` ::
# `TableOverride` for the full surface.

# Skip diagnostic tables on every connector.
[[pipeline.raw.table_overrides]]
match.table = "_diagnostics_*"        # `*`/`?` glob in TOML only — CLI literals
enabled     = false

# `pii_users` on one specific connector needs a composite merge key.
[[pipeline.raw.table_overrides]]
match.connector = "stripe_main"        # matches conn.id OR conn.schema
match.table     = "pii_users"
merge_keys      = ["user_id", "tenant_id"]

# `audit_log` is append-only — keep it on incremental even when
# the pipeline default is `merge`.
[[pipeline.raw.table_overrides]]
match.connector  = "stripe_main"
match.table      = "audit_log"
strategy         = "incremental"
timestamp_column = "occurred_at"

CLI symmetry: --filter table=<literal> filters to one table at a time (no globs at the CLI; shell quoting is too fragile). The connector-level --filter id=<conn_id> keys remain unchanged.

Transformation pipeline
toml
[pipeline.silver]
type       = "transformation"
models     = "models/silver/**"   # glob, relative to rocky.toml; default "models/**"
depends_on = ["raw"]

# There is no contracts key: a model's contract is auto-discovered as the
# sibling `<model>.contract.toml`, or passed at the CLI via `--contracts <dir>`.

# A transformation target takes an adapter ref (plus optional governance)
# and nothing else — `catalog`/`schema` here are parse errors
# (`deny_unknown_fields`). Each model names its own destination in its
# sidecar `[target]`:
#
#   # models/silver/dim_customers.toml
#   [target]
#   catalog = "analytics"
#   schema  = "marts"
#   table   = "dim_customers"
[pipeline.silver.target]
adapter = "warehouse"

Declare the external tables the models read, with optional dbt-style freshness checked by rocky freshness (exit 1 on error / runtime_error):

toml
[[pipeline.silver.sources]]
schema = "raw"                     # catalog optional (two-part names)
table  = "orders"

[pipeline.silver.sources.freshness]
loaded_at_field = "_loaded_at"     # DATE / TIMESTAMP column
warn_after      = "12h"            # <N>s | <N>h | <N>d; at least one of the two
error_after     = "24h"            # must be >= warn_after (else E050)
filter          = "status <> 'test'"   # optional WHERE predicate; no `;`
Quality pipeline
toml
[pipeline.qa]
type       = "quality"
depends_on = ["silver"]

# A quality target is an adapter ref only. The tables to check are listed
# in `[[pipeline.qa.tables]]` (omit `table` to check the whole schema).
[pipeline.qa.target]
adapter = "warehouse"

[[pipeline.qa.tables]]
catalog = "analytics"
schema  = "marts"

# Quality runs execute `row_count`, `custom`, and `assertions` only. Other
# check kinds are inert here — `rocky validate` warns (V034).
[pipeline.qa.checks]
row_count = true
Snapshot pipeline
toml
[pipeline.dim_history]
type       = "snapshot"
unique_key = ["customer_id"]     # required: row identity in the source table
updated_at = "updated_at"        # required: change-detection column
depends_on = ["silver"]

[pipeline.dim_history.source]
catalog = "analytics"
schema  = "marts"
table   = "dim_customers"

[pipeline.dim_history.target]
catalog = "analytics_history"
schema  = "snapshots"
table   = "dim_customers_history"   # required: explicit history table

Minimal-config defaults (omit these)

The parser applies sane defaults — keep configs lean by omitting anything that matches the default:

FieldDefaultOmit unless
pipeline.type"replication"You need transformation/quality/snapshot
adapter = "default" (in pipeline source/target)First adapterMulti-adapter config
[state]\nbackend = "local"local (embedded redb)Using S3, Valkey, or tiered state sync
auto_create_catalogs = falsefalseYou want Rocky to CREATE CATALOG
auto_create_schemas = falsefalseYou want Rocky to CREATE SCHEMA
Model name in sidecar .tomlfilename stemYou want a different logical name
Model target.tablenameRenaming on write
Directory-level targetmodels/_defaults.toml inheritedOverriding per-model

Checks

Every pipeline type accepts a [checks] block, but execution differs: replication runs the full set; quality runs row_count, custom, and assertions only; transformation, snapshot, and load run no pipeline-level checks today (executed_check_kinds in config.rs is the source of truth). rocky validate warns (V034) on a check the pipeline type never executes.

toml
[pipeline.<name>.checks]
enabled      = true
row_count    = true                                      # source vs target row count
column_match = true                                      # source vs target column list
freshness    = { threshold_seconds = 86400 }             # max staleness of newest row
null_rate    = { columns = ["email"], threshold = 0.05, sample_percent = 10 }

[[pipeline.<name>.checks.custom]]
name      = "no_future_dates"
sql       = "SELECT COUNT(*) FROM {target} WHERE created_at > CURRENT_TIMESTAMP()"
threshold = 0                                            # max failing rows

Contracts

Enforced on load pipelines: the load runs into a staging table, validates the landed schema against the contract, and promotes to the target only on pass — a violation drops staging and fails, so non-conforming data never lands. Types compare in Rocky's normalized vocabulary, so one contract ports across warehouses.

toml
[pipeline.<name>.contract]
required_columns = [
    { name = "id", type = "BIGINT", nullable = false },
]
protected_columns    = ["id", "email"]                   # may not be removed from source
allowed_type_changes = [
    { from = "INT", to = "BIGINT" },                     # widening allowlist
]

Execution & adaptive concurrency

toml
[pipeline.<name>.execution]
concurrency           = 16
fail_fast             = false
error_rate_abort_pct  = 50
table_retries         = 1
# adaptive_concurrency is planned but not yet a config field.
# The AIMD throttle primitive exists but isn't wired.

Governance (Databricks Unity Catalog)

Governance is configured per pipeline target, under [pipeline.<name>.target.governance]. There is no top-level [governance] table; one is refused at load (deny_unknown_fields on RockyConfig).

toml
[pipeline.bronze.target.governance]
auto_create_catalogs = true
auto_create_schemas  = true

[pipeline.bronze.target.governance.tags]
managed_by = "rocky"

[[pipeline.bronze.target.governance.grants]]
principal   = "data-readers"
permissions = ["BROWSE", "USE CATALOG", "USE SCHEMA", "SELECT"]

[pipeline.bronze.target.governance.isolation]
enabled       = true
workspace_ids = "${WORKSPACE_IDS:-}"

Permissions handled: BROWSE, USE CATALOG, USE SCHEMA, SELECT, MANAGE, MODIFY. Skipped: OWNERSHIP, ALL PRIVILEGES, CREATE SCHEMA (non-managed).

State backend

toml
# Embedded redb (default — no config needed)
[state]
backend = "local"

# S3-backed state sync
[state]
backend = "s3"
bucket  = "my-rocky-state"
prefix  = "prod/"
region  = "us-east-1"

# Tiered: local redb + S3 for durability
[state]
backend = "tiered"
# … S3 fields plus local path

# State-file namespacing (opt-in, default "none")
[state]
backend = "local"
namespacing = "pipeline"   # each pipeline gets its own .rocky-state/<pipeline>.redb

namespacing (StateNamespacing, default "none") controls state-file fan-out. redb allows one writer per file, so running one rocky run per pipeline/client against the single global <models>/.rocky-state.redb serializes them on one lock. "pipeline" gives each pipeline its own <models>/.rocky-state/<pipeline>.redb. "none" is byte-identical to omitting the key. For per-client fan-out use the per-invocation --state-namespace <key> flag (overrides this config); an explicit --state-path disables namespacing for that run.

Show full SKILL.md (587 more words)Show less

Run-skip gate ([run] + per-model [skip])

toml
# Opt-in model-skip gate — default OFF (omit the block to keep old behavior)
[run]
skip_unchanged = true            # master switch (also via the --skip-unchanged flag)
skip_rowcount_fallback = false   # default; allow COUNT(*) when no timestamp column (weaker signal)
lag_tolerance_seconds = 0        # default; any MAX(ts) movement forces a rebuild

[run] (RunConfig) tunes the --skip-unchanged gate: skip re-materializing a transformation model whose logic and every upstream's data both appear unchanged since the last successful build. It is a best-effort optimization, not a result-equivalence guarantee — every field defaults to no-skip, and any missing/unreadable/ambiguous input rebuilds (fail-safe).

Not skip-eligible (always rebuild):

  • Non-deterministic SQL: CURRENT_TIMESTAMP / NOW(), RANDOM(), UUID(), CURRENT_USER, CURRENT_CATALOG, ANY_VALUE, ARRAY_AGG, unordered LIMIT, or any unknown function.
  • Models whose lineage isn't provably complete: CTEs, subqueries (FROM (…), IN (SELECT …), EXISTS, scalar sub-selects), PIVOT / UNNEST / nested-join table-factors, and set operations (UNION / INTERSECT / EXCEPT).
  • content_addressed / time_interval strategies (full_refresh is eligible).

--force-rebuild bypasses the gate entirely.

toml
# Per-model override (in the model's .toml sidecar)
name = "fct_orders"

[skip]
eligible = true        # Some(false) = always build; Some(true) = eligible; unset = auto
deterministic = true   # owner asserts SQL is pure → re-eligible despite the static scan

[skip] (SkipConfig in rocky-core/src/models.rs) is a per-model sidecar block. Both fields are Option<bool> — unset means "trust the automatic rules."

--defer / --defer-to <schema> are runtime-only dev flags (not config): build the --model-selected models locally and resolve unbuilt upstream ref()s to a prod/defer schema. The rewrite parses model SQL with the Databricks dialect, so SELECT * EXCEPT(...), trailing-comma selects, and STRUCT(...) literals can't be rewritten — run those without --defer.

Reuse ([reuse]) — EXPERIMENTAL, do not enable in production

toml
[reuse]
enabled = false   # default; preview-only, NOT live-verified — leave off in production

[reuse] (ReuseConfig) is a preview surface scoped to the Databricks–Iceberg content-addressed write path only (no DuckDB / Snowflake / BigQuery), and it is not yet live-verified against a warehouse. When enabled = true, a successful run only populates an input-match index + provenance record; the reuse decision path only ever resolves to BUILD today (an ONLY-BUILD posture) — a fail-closed verdict is computed but nothing is reused, since the live point-to reuse is not yet wired/verified. It attests an input-logic match + byte-identity of the recorded bytes, never that a fresh re-run would reproduce them. Default-off keeps rocky run byte- and cost-identical. The per-invocation --no-reuse flag forces every model to build. Provenance is auditable per docs/.../guides/verify-a-run.md.

Cache

toml
[cache.schemas]
enabled     = true     # default; false for strict CI (every typecheck hits the warehouse)
ttl_seconds = 86400    # default 24h; lower for high-DDL-churn teams
replicate   = false    # default; true to ship the cache through state_sync
# trusted_max_age_seconds = 3600  # unset by default; entries younger than this make a missing source column E041
# strict_sources = true           # default false; every W041 (seed / old cache entry) becomes E041

[cache.schemas] is the only [cache] table (CacheConfig in config.rs): it stores DESCRIBE TABLE results in state.redb so leaf models typecheck against real warehouse types without a live round-trip on every compile. There is no [cache] valkey_url key: ValkeyCacheConfig exists as a type but is not wired into RockyConfig, and a [cache] table with any other key is refused (deny_unknown_fields). The Valkey tier is a [state] backend, not a cache setting.

Cost model (for rocky optimize)

toml
[cost]
storage_cost_per_gb_month = 0.023
compute_cost_per_dbu      = 0.40
warehouse_size            = "Medium"
min_history_runs          = 5

Hooks

toml
[hook.on_pipeline_start]
command    = "scripts/notify.sh"
timeout_ms = 5000
on_failure = "warn"                  # default; or "abort" (stop the pipeline) / "ignore" (silent)

[hook.on_pipeline_fail]
url        = "${SLACK_WEBHOOK_URL}"   # webhook instead of command
template   = "default"                # built-in preset

Events: on_pipeline_start, on_pipeline_complete, on_pipeline_fail, on_model_start, on_model_complete, on_model_fail, on_check_fail, on_drift_detected.

Model sidecar files (.sql + .toml)

Model files live under models/ and use a sidecar pattern:

toml
# models/marts/dim_customers.toml
name       = "dim_customers"
depends_on = ["stg_customers"]

[strategy]
type       = "merge"                 # e.g. full_refresh, merge, delete_insert, time_interval, view, snapshot; not incremental (E037)
unique_key = ["customer_id"]
# update_columns = ["name", "email"] # omit for UPDATE SET *

[target]
catalog = "analytics"
schema  = "marts"
table   = "dim_customers"

[[sources]]                          # optional: declare source tables
catalog = "analytics"
schema  = "staging"
table   = "customers"
sql
-- models/marts/dim_customers.sql
-- Pure SQL. No Jinja. No templating.
SELECT customer_id, name, email, updated_at
FROM {{ analytics.staging.customers }}        -- Rocky expands refs at compile time, not via Jinja

Directory-level defaults via models/<dir>/_defaults.toml:

toml
[target]
catalog = "analytics"
schema  = "marts"

Validation

SQL identifiers (catalog, schema, table, tenant, region, source names) must match ^[a-zA-Z0-9_]+$. Rocky rejects anything else. Principal names for GRANT/REVOKE allow the broader ^[a-zA-Z0-9_ \-\.@]+$ pattern and are always wrapped in backticks in generated SQL.

Time-interval (partition-keyed) materialization

For transformation models that need partition-by-partition execution:

toml
# models/fact_events.toml
[strategy]
type              = "time_interval"
time_column       = "event_date"
granularity       = "day"                   # hour | day | month | year
lookback          = 7                       # re-run last 7 partitions
batch_size        = 4                       # partitions per concurrent batch
first_partition   = "2024-01-01"

Model SQL uses @start_date and @end_date placeholders — Rocky substitutes per partition:

sql
SELECT * FROM raw.events
WHERE event_date >= @start_date AND event_date < @end_date

CLI: rocky run --partition KEY / --from KEY --to KEY / --latest / --missing / --lookback N / --parallel N.

Full canonical example

See examples/playground/pocs/00-foundations/01-replication-basics/rocky.toml for the minimal DuckDB replication case (schema-pattern routing), examples/playground/pocs/00-foundations/00-playground-default/rocky.toml for the minimal transformation case (model DAG), or engine/examples/multi-layer/rocky.toml for a full Bronze/Silver/Gold setup.

What NOT to write (pre-Phase-2 legacy)

The following keys do not work anymore — they were the pre-Phase-2 config shape and will be rejected by the parser:

❌ Legacy✅ Current
[source] top-level[pipeline.<name>.source]
[warehouse][adapter]
[replication] top-level[pipeline.<name>] with strategy = "incremental"
[checks] top-level[pipeline.<name>.checks]
[target] top-level[pipeline.<name>.target]

If you see any of these in a config you're editing, the config is stale — migrate it.

Reference

  • engine/crates/rocky-core/src/config.rs — Rust source of truth for every field
  • engine/AGENTS.md — "Configuration" section with full annotated example
  • editors/vscode/schemas/rocky-config.schema.json — JSON Schema for IDE autocompletion (autogenerated; keep in sync with config.rs)
  • examples/playground/AGENTS.md — POC-specific minimal-config idioms

© rocky-data, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/rocky-config of rocky-data/rocky.

Open the folder on GitHubat commit 365aebf

Compare with similar skills

Rocky Config next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Rocky Config compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Rocky Config this skillrocky-data/rocky304—~6.9kAutomated safety check: PassApache-2.0
Ade Bench Cross DB Tasksdbt-labs/ade-bench125—~2kAutomated safety check: PassApache-2.0
Altimate Data Warehouse DelegateAltimateAI/data-engineering-skills128—~1.4kAutomated safety check: PassMIT
SQL Queriesw95/awesome-claude-corporate-skills2443 repos~2.8kAutomated safety check: PassMIT
Connect Snowflakeai-analyst-lab/ai-analyst304—~1.4kAutomated safety check: NotesMIT
Databricks Icebergdatabricks/databricks-agent-skills345—~2.6kAutomated safety check: PassCustom licence

Similar skills

  • Ade Bench Cross DB Tasks

    dbt-labs/ade-bench

    Official

    A skill your agent uses when authoring or debugging ade-bench tasks that must run on both DuckDB and Snowflake, including shared project migrations, setup patches, and solution patches

    125 GitHub stars~2k tokensUpdated 10 days ago
    DatabasesAuto-check passed
  • Altimate Data Warehouse Delegate

    AltimateAI/data-engineering-skills

    Delegates dbt and warehouse tasks such as lineage, migrations and cost attribution to the altimate-code CLI agent and relays its answer back.

    128 GitHub stars~1.4k tokensUpdated 2 days ago
    DatabasesAuto-check passed
  • SQL Queries

    w95/awesome-claude-corporate-skills

    Write correct, performant SQL across all major data warehouse dialects (Snowflake, BigQuery, Databricks, PostgreSQL, etc.).

    244 GitHub starsUsed in 3 repos~2.8k tokens
    DatabasesAuto-check passed
  • Connect Snowflake

    ai-analyst-lab/ai-analyst

    Query the live/remote Snowflake warehouse instead of the local practice copy.

    304 GitHub stars~1.4k tokensUpdated 9 days ago
    DatabasesAuto-check: notes
  • Databricks Iceberg

    databricks/databricks-agent-skills

    Official

    Apache Iceberg tables on Databricks — Managed Iceberg tables, External Iceberg Reads (fka Uniform), Compatibility Mode, Iceberg REST Catalog (IRC), Iceberg v3, Snowflake interop, PyIceberg, OSS…

    345 GitHub stars~2.6k tokensUpdated today
    DatabasesAuto-check passed
  • Migrating To Amazon Redshift

    aws/agent-toolkit-for-aws

    Official

    Guides an end-to-end data-warehouse migration to Amazon Redshift — discovery, schema/SQL/stored-procedure/macro/script conversion, data migration, validation, performance comparison, and reporting.

    2.8k GitHub stars~2.7k tokensUpdated today
    DatabasesAuto-check: notes

More from rocky-data/rocky

All 22 skills in this repo
  • Fivetran

    rocky-data/rocky

    Fivetran REST API reference for Rocky's source adapter. An agent skill from rocky-data/rocky.

    304 GitHub stars~914 tokensUpdated today
    Auto-check passed
  • Databricks

    rocky-data/rocky

    Databricks REST API and SQL reference for Rocky's warehouse adapter.

    304 GitHub stars~2k tokensUpdated today
    Auto-check passed
  • Rocky Codegen

    rocky-data/rocky

    Rocky CLI JSON-output schema cascade. An agent skill from rocky-data/rocky.

    304 GitHub stars~1.9k tokensUpdated today
    Auto-check passed
  • Rocky Dev

    rocky-data/rocky

    Top-level router for Rocky development tasks. An agent skill from rocky-data/rocky.

    304 GitHub stars~2.1k tokensUpdated today
    Auto-check passed
  • Rocky Dsl Change

    rocky-data/rocky

    Rocky DSL (.rocky file) cross-subproject cascade. An agent skill from rocky-data/rocky.

    304 GitHub stars~1.3k tokensUpdated today
    Auto-check passed
  • Rocky New Adapter

    rocky-data/rocky

    Adding a new warehouse or source adapter crate to the Rocky engine.

    304 GitHub stars~2k tokensUpdated today
    Auto-check passed

Categories

Questions about Rocky Config

What does Rocky Config do?

Canonical rocky.toml authoring reference. An agent skill from rocky-data/rocky. Rocky Config is an agent skill from rocky-data/rocky.toml authoring reference.

When should I use Rocky Config?

Rocky Config fits situations like: reviewing a Rocky pipeline config — covers the 4 pipeline types (replication; adapter variants (duckdb/databricks/snowflake/fivetran); minimal-config defaults; env-var substitution.

How do I install Rocky Config in Claude Code?

Run `npx skills add rocky-data/rocky --skill rocky-config -a claude-code`. Or copy the skill folder (.agents/skills/rocky-config in rocky-data/rocky) into .claude/skills/rocky-config in your project. Claude Code loads it when a task matches its description.

How do I install Rocky Config in Codex?

Run `npx skills add rocky-data/rocky --skill rocky-config -a codex`. Or copy the skill folder (.agents/skills/rocky-config in rocky-data/rocky) into .agents/skills/rocky-config in your project. Codex loads it when a task matches its description.

Can I use Rocky Config in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add rocky-data/rocky --skill rocky-config -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/rocky-config, .gemini/skills/rocky-config, .github/skills/rocky-config and .opencode/skills/rocky-config in your project.

What does Rocky Config need to run?

Going by SKILL.md and its folder, Rocky Config needs credentials named DATABRICKS_TOKEN, DATABRICKS_CLIENT_SECRET, SNOWFLAKE_OAUTH_TOKEN and SNOWFLAKE_PASSWORD. Our summary lists: A credential in DATABRICKS_TOKEN; A credential in DATABRICKS_CLIENT_SECRET.

Does Rocky Config access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Rocky Config safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Rocky Config use?

Rocky Config is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Rocky Config use?

About 6.9k tokens (SKILL.md is roughly 28k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Rocky Config?

Skills that share tags, products or a category with Rocky Config: Ade Bench Cross DB Tasks (dbt-labs/ade-bench, 125 stars), Altimate Data Warehouse Delegate (AltimateAI/data-engineering-skills, 128 stars), SQL Queries (w95/awesome-claude-corporate-skills, 244 stars) and Connect Snowflake (ai-analyst-lab/ai-analyst, 304 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Rocky Config?

rocky-data (a GitHub organization) maintains it in rocky-data/rocky, which has 304 GitHub stars. The repository holds 22 skills in this directory. The repository was last updated on October 10, 2026.

Source: rocky-data/rocky on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.