Agent skill

Rust Bench Criterion

by rocky-data in rocky-data/rocky

Writing criterion benchmarks for the Rocky engine and wiring them to the perf-gated CI workflow.

Apache-2.0Auto-check passedDatabases

Install Rust Bench Criterion

skills CLI
$ npx skills add rocky-data/rocky --skill rust-bench-criterion -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install rocky-data/rocky rust-bench-criterion --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/rocky-data/rocky.git skills-src && mkdir -p .claude/skills && cp -r skills-src/engine/.claude/skills/rust-bench-criterion .claude/skills/rust-bench-criterion && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
rust-bench-criterion
GitHub stars
304
Token cost
~2.5k tokens
SKILL.md length
1,077 words
Files
1
Skills in repo
22
Repo updated
First seen
Licence
Apache-2.0

At a glance

Writing criterion benchmarks for the Rocky engine and wiring them to the perf-gated CI workflow.

  • Works in 4 steps: Benches only run when the perf label is… → The workflow hardcodes three cargo bench… → Nothing compares your numbers to… → …
  • Adding a new benchmark
  • SKILL.md covers The benches and the CI gate, Adding a new criterion bench…, Adding a bench in a new crate and When to add a perf-labelled PR, plus 2 more sections
  • Calls cargo

What it does

Rust Bench Criterion is an agent skill from rocky-data/rocky. Writing criterion benchmarks for the Rocky engine and wiring them to the perf-gated CI workflow. Use when adding a new benchmark, debugging a bench CI alert, or deciding whether a change needs the perf PR label.

Its SKILL.md is about 2.5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Databases. It works with Rust and SQL. The repository describes itself as: A SQL transformation engine that type-checks your whole pipeline and catches breaking changes before they run — branches, replay, column-level lineage, compile-time contracts… The licence is Apache-2.0.

When your agent uses it

  • Adding a new benchmark
  • Debugging a bench CI alert
  • Deciding whether a change needs the perf PR label

Example prompts

  • “/rust-bench-criterion”

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Benches only run when the perf label is on the PR. Adding a bench doesn't automatically run it on PR. To exercise it, add the perf label.
  2. The workflow hardcodes three cargo bench invocations — compile -p rocky-cli, state_store -p rocky-core, dag_execute -p rocky-core. A new…
  3. Nothing compares your numbers to anything. No baseline, no alert threshold, no PR comment, no pass/fail. To see a regression you download…
  4. Binary startup bench assumes cargo run works from the working directory — it's advisory, not load-bearing; don't panic if it reports noisy…

What it can do on your machine

Read from SKILL.md and the folder at commit 46be77e. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • cargo

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Rust Bench Criterion loads about 2.5k tokens when it runs. Until then it costs about 59 tokens; SKILL.md has 1,077 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~59
When it runs · the whole SKILL.md, loaded when a task matches
~2.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from rocky-data/rocky at commit 46be77e, republished under its Apache-2.0 licence (© rocky-data). 1,077 words, ~2,542 tokens.

Download SKILL.mdSave it as .claude/skills/rust-bench-criterion/SKILL.md (or your agent's skills folder).
name
rust-bench-criterion
description
Writing criterion benchmarks for the Rocky engine and wiring them to the perf-gated CI workflow. Use when adding a new benchmark, debugging a bench CI alert, or deciding whether a change needs the `perf` PR label.

Criterion benchmarks in Rocky

The benches and the CI gate

There are three criterion benches in the engine, and the workflow runs all three:

CrateBenchFileIn engine-bench.yml?
rocky-clicompilebenches/compile.rsyes
rocky-corestate_storebenches/state_store.rsyes
rocky-coredag_executebenches/dag_execute.rsyes

Decided (#1947, 2026-09-17): dag_execute is not redundant with compile's dag_resolution group — dag_resolution benches the synchronous graph algorithms (topological_sort + execution_layers); dag_execute benches DagExecutor::execute itself (the tokio runtime, the semaphore, JoinSet intra-layer fan-out). Different surfaces, so it's wired in with sample_size(10) to keep it cheap. As with the other two, this buys "it still compiles and runs on CI hardware" — see the no-comparison-step note below — not a perf gate. The correctness property this bench demonstrates (parallel beats sequential) is separately pinned by a real, non-flaky assertion: max_concurrency_bounds_intra_layer_fan_out in rocky-core/src/dag_executor.rs's test module counts peak concurrent nodes via an AtomicUsize high-water mark instead of timing anything, and runs on every cargo test, not just perf-labeled PRs.

compile.rs has four groups:

  • cold_compile — full compile() over synthetic projects of 10 / 100 / 1,000 models (plus 10,000 in release mode only — debug builds with DuckDB C++ would take ~30 s per iter and dominate CI).
  • dag_resolution — topological_sort + execution_layers on diamond-shaped DAGs of 10 / 100 / 500 / 1000 nodes.
  • sql_generation — generate_create_table_as_sql and generate_insert_sql for 10 / 100 / 500 replication plans.
  • binary_startup — subprocess startup time for cargo run -p rocky -- --version.

CI wiring — .github/workflows/engine-bench.yml:

yaml
on:
  pull_request:
    types: [labeled, synchronize]
    paths:
      - 'engine/crates/**'
      - 'engine/Cargo.toml'
      - 'engine/Cargo.lock'
      - '.github/workflows/engine-bench.yml'

jobs:
  bench:
    if: contains(github.event.pull_request.labels.*.name, 'perf')
    # ...
    steps:
      - name: Run benchmarks
        run: |
          cargo bench --bench compile -p rocky-cli -- --output-format bencher | tee bench-output.txt
          cargo bench --bench state_store -p rocky-core -- --output-format bencher | tee -a bench-output.txt
          cargo bench --bench dag_execute -p rocky-core -- --output-format bencher | tee -a bench-output.txt
      - name: Upload benchmark results
        uses: actions/upload-artifact@...
        with:
          name: bench-output
          path: engine/bench-output.txt
          if-no-files-found: error

There is no comparison step, and there never has been one that worked. The workflow's own comment says why: it runs only on perf-labeled PRs, with no push trigger, so there is no main-branch baseline to compare against. github-action-benchmark needs a gh-pages history branch, and this repo has none. The run uploads raw bencher output as a downloadable artifact and stops there.

The rules you actually need to know:

  1. Benches only run when the perf label is on the PR. Adding a bench doesn't automatically run it on PR. To exercise it, add the perf label.
  2. The workflow hardcodes three cargo bench invocations — compile -p rocky-cli, state_store -p rocky-core, dag_execute -p rocky-core. A new [[bench]] is not picked up until you add another invocation for it; adding a [[bench]] target to a Cargo.toml alone does nothing in CI.
  3. Nothing compares your numbers to anything. No baseline, no alert threshold, no PR comment, no pass/fail. To see a regression you download the bench-output artifact and read it, or you run the bench locally on both revisions yourself. Treat the CI run as evidence the bench still COMPILES AND RUNS, not as a perf gate.
  4. Binary startup bench assumes cargo run works from the working directory — it's advisory, not load-bearing; don't panic if it reports noisy numbers.

Adding a new criterion bench to an existing crate

  1. Add the dep. In the crate's Cargo.toml:

    toml
    [dev-dependencies]
    criterion = { version = "0.5", features = ["html_reports"] }
    tempfile = "3"  # if you need fixtures
    
    [[bench]]
    name = "my_bench"
    harness = false

    harness = false is required for criterion (it supplies its own main).

  2. Create crates/<crate>/benches/my_bench.rs. The canonical shape, modelled on benches/compile.rs:

    rust
    //! Short doc of what this measures and how to run it.
    //!
    //! Run with: `cargo bench --bench my_bench -p <crate>`
    
    use criterion::{BenchmarkId, Criterion, criterion_group, criterion_main};
    
    fn bench_happy_path(c: &mut Criterion) {
        let mut group = c.benchmark_group("happy_path");
        group.sample_size(10);  // ← drop from the default 100 for expensive iters
    
        for size in [10, 100, 1_000] {
            let input = build_input(size);
            group.bench_with_input(
                BenchmarkId::new("variant_name", size),
                &input,
                |b, input| {
                    b.iter(|| {
                        my_function(input);
                    });
                },
            );
        }
        group.finish();
    }
    
    criterion_group!(benches, bench_happy_path);
    criterion_main!(benches);
    
    fn build_input(size: usize) -> Input { /* ... */ }
  3. Decide what scales to measure. Rocky's pattern is to sweep an n dimension (models, rows, nodes) with a few 10× steps (10 / 100 / 1000). Steps > 10× are rare. If one step takes > 30 s on a debug build, gate it behind cfg!(debug_assertions) so it only runs in release — see the 10_000 gate in bench_cold_compile.

  4. Reduce sample_size for expensive iters. The default 100 samples × ~30 s per iter = 50 minutes. group.sample_size(10) is the right call for anything that touches DuckDB, the full compiler, or a real filesystem fixture.

  5. Name groups and IDs for a reader, not for history. Nothing stores results between runs, so renaming a group loses nothing and costs nothing — there is no baseline to break. Criterion's own target/criterion/ directory is local and per-checkout, so it is not history either: deleting engine/target destroys no CI perf record, because none is kept. Name them clearly because a human reads them in the artifact.

  6. Test locally before adding the perf label.

    bash
    cd engine
    cargo bench --bench my_bench -p <crate>

    Criterion writes HTML reports to target/criterion/<group>/<id>/report/index.html — open those to confirm the numbers look sane before pushing.

Show full SKILL.md (418 more words)Show less

Adding a bench in a new crate

The CI workflow currently hardcodes three invocations (--bench compile -p rocky-cli, --bench state_store -p rocky-core, --bench dag_execute -p rocky-core). If you add a bench in, say, rocky-sql, the workflow won't invoke it. You have two options:

  1. Add another invocation to the workflow (preferred) — append another cargo bench ... | tee -a bench-output.txt line to the Run benchmarks step's run: block so its output lands in the same uploaded artifact (there is no github-action-benchmark step to duplicate — see the no-comparison-step note above). Review with Hugo since it extends the perf budget.
  2. Put the bench in rocky-cli/benches/compile.rs — acceptable if the bench is logically "compile-adjacent" and you can drive it through the existing compile entrypoint. Not acceptable for benchmarks that need to import from a crate rocky-cli doesn't already depend on.

When to add a perf-labelled PR

Rocky's benchmarks are not free — the CI job builds DuckDB with swap and takes substantial time. Add the perf label when:

  • You changed anything in the cold_compile / dag_resolution / sql_generation hot paths.
  • You touched rocky-core/src/{dag,ir,sql_gen,mmap,intern}.rs.
  • You added or removed a dep that sits in the compile hot path.
  • Hugo asks for it in review.

You do not need the label for:

  • Docs / CLAUDE.md / skills / comments-only changes.
  • Changes scoped to a single adapter crate's API client (those are I/O bound, not CPU bound).
  • Tests, fixture regens, dagster integration changes.

Interpreting a regression comment

The github-action-benchmark action posts a comment when a group/ID exceeds 120% of its stored baseline. Triage order:

  1. Re-run the bench once — CI jitter is real, and sample_size(10) amplifies it.
  2. Open the criterion HTML reports from the bench artifact — the mean / median / p95 often tells a different story from the bencher summary.
  3. Profile locally if the regression reproduces — cargo flamegraph --bench compile -p rocky-cli -- --bench <group>/<id> is the usual entry point (requires cargo-flamegraph installed).
  4. Decide fix-vs-accept — for a ≤ 125% regression with a good reason (correctness fix, new feature, dep bump), a PR comment explaining the trade-off is fine. For anything else, fix before merge.

fail-on-alert: false means the comment is advisory. Don't ignore it. A silently-accepted regression accumulates across multiple PRs into a real perf cliff.

  • rust-clippy-triage — clippy warnings on bench code count the same as anywhere else. Benches are --all-targets.
  • rust-async-tokio — benchmarking async code needs tokio::runtime::Runtime::new().unwrap().block_on(...) around the .iter().
  • rust-dep-hygiene — adding criterion to a new crate uses [dev-dependencies], which are not inherited from [workspace.dependencies] in the same way as runtime deps; pin the version in the sub-crate as benches/compile.rs does.

© rocky-data, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in engine/.claude/skills/rust-bench-criterion of rocky-data/rocky.

Open the folder on GitHubat commit 46be77e

Compare with similar skills

Rust Bench Criterion next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Rust Bench Criterion compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Rust Bench Criterion this skillrocky-data/rocky304—~2.5kAutomated safety check: PassApache-2.0
Diesel Guardayarotsky/diesel-guard121—~3.1kAutomated safety check: PassMIT
SeekDB Code Reviewoceanbase/seekdb3.1k—~2.1kAutomated safety check: PassApache-2.0
Webapp Buildersidequery/sidemantic129—~5.5kAutomated safety check: PassAGPL-3.0
Querying Tempotempoxyz/tidx107—~3.1kAutomated safety check: PassMIT
Paro Optimizerzunor/paro105—~979Automated safety check: PassApache-2.0

Similar skills

  • Diesel Guard

    ayarotsky/diesel-guard

    Lints Diesel and SQLx Postgres migrations for unsafe schema changes that lock tables or cause downtime, and authors custom Rhai checks.

    121 GitHub stars~3.1k tokensUpdated 9 days ago
    DatabasesAuto-check passed
  • SeekDB Code Review

    oceanbase/seekdb

    Reviews seekdb pull requests and diffs for real defects in correctness, resources, concurrency, security and tests, reporting only Blocker or Major findings.

    3.1k GitHub stars~2.1k tokensUpdated today
    DevelopmentAuto-check passed
  • Webapp Builder

    sidequery/sidemantic

    Build interactive analytics webapps, demos, dashboards, or embedded app surfaces from Sidemantic semantic models using copyable component primitives and deterministic query inspection.

    129 GitHub stars~5.5k tokensUpdated today
    DatabasesAuto-check passed
  • Querying Tempo

    tempoxyz/tidx

    Query indexed Tempo chain data via tidx HTTP API and CLI. An agent skill from tempoxyz/tidx.

    107 GitHub stars~3.1k tokensUpdated today
    DatabasesAuto-check passed
  • Paro Optimizer

    zunor/paro

    Design, refactor and diagnose Paro's staged optimizer, using EXPLAIN COMPILE for planning and EXPLAIN ANALYZE for execution.

    105 GitHub stars~979 tokensUpdated today
    DatabasesAuto-check passed
  • Paro Benchmark

    zunor/paro

    Run Paro engineering performance gates and exploratory cold/warm, cross-engine or operator comparisons.

    105 GitHub stars~1.8k tokensUpdated today
    DatabasesAuto-check passed

More from rocky-data/rocky

All 22 skills in this repo
  • Fivetran

    rocky-data/rocky

    Fivetran REST API reference for Rocky's source adapter. An agent skill from rocky-data/rocky.

    304 GitHub stars~914 tokensUpdated today
    Auto-check passed
  • Databricks

    rocky-data/rocky

    Databricks REST API and SQL reference for Rocky's warehouse adapter.

    304 GitHub stars~2k tokensUpdated today
    Auto-check passed
  • Rocky Codegen

    rocky-data/rocky

    Rocky CLI JSON-output schema cascade. An agent skill from rocky-data/rocky.

    304 GitHub stars~1.9k tokensUpdated today
    Auto-check passed
  • Rocky Dev

    rocky-data/rocky

    Top-level router for Rocky development tasks. An agent skill from rocky-data/rocky.

    304 GitHub stars~2.1k tokensUpdated today
    Auto-check passed
  • Rocky Dsl Change

    rocky-data/rocky

    Rocky DSL (.rocky file) cross-subproject cascade. An agent skill from rocky-data/rocky.

    304 GitHub stars~1.3k tokensUpdated today
    Auto-check passed
  • Rocky New Adapter

    rocky-data/rocky

    Adding a new warehouse or source adapter crate to the Rocky engine.

    304 GitHub stars~2k tokensUpdated today
    Auto-check passed

Works with

Questions about Rust Bench Criterion

What does Rust Bench Criterion do?

Writing criterion benchmarks for the Rocky engine and wiring them to the perf-gated CI workflow. Rust Bench Criterion is an agent skill from rocky-data/rocky. Writing criterion benchmarks for the Rocky engine and wiring them to the perf-gated CI workflow.

When should I use Rust Bench Criterion?

Rust Bench Criterion fits situations like: adding a new benchmark; debugging a bench CI alert; deciding whether a change needs the perf PR label.

How do I install Rust Bench Criterion in Claude Code?

Run `npx skills add rocky-data/rocky --skill rust-bench-criterion -a claude-code`. Or copy the skill folder (engine/.claude/skills/rust-bench-criterion in rocky-data/rocky) into .claude/skills/rust-bench-criterion in your project. Claude Code loads it when a task matches its description.

How do I install Rust Bench Criterion in Codex?

Run `npx skills add rocky-data/rocky --skill rust-bench-criterion -a codex`. Or copy the skill folder (engine/.claude/skills/rust-bench-criterion in rocky-data/rocky) into .agents/skills/rust-bench-criterion in your project. Codex loads it when a task matches its description.

Can I use Rust Bench Criterion in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add rocky-data/rocky --skill rust-bench-criterion -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/rust-bench-criterion, .gemini/skills/rust-bench-criterion, .github/skills/rust-bench-criterion and .opencode/skills/rust-bench-criterion in your project.

What does Rust Bench Criterion need to run?

Going by SKILL.md and its folder, Rust Bench Criterion needs the command-line tools its instructions call (cargo).

Does Rust Bench Criterion access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Rust Bench Criterion safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Rust Bench Criterion use?

Rust Bench Criterion is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Rust Bench Criterion use?

About 2.5k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Rust Bench Criterion?

Skills that share tags, products or a category with Rust Bench Criterion: Diesel Guard (ayarotsky/diesel-guard, 121 stars), SeekDB Code Review (oceanbase/seekdb, 3.1k stars), Webapp Builder (sidequery/sidemantic, 129 stars) and Querying Tempo (tempoxyz/tidx, 107 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Rust Bench Criterion?

rocky-data (a GitHub organization) maintains it in rocky-data/rocky, which has 304 GitHub stars. The repository holds 22 skills in this directory. The repository was last updated on October 8, 2026.

Source: rocky-data/rocky on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.