Agent skill

Benchmark

by nguyenyou in nguyenyou/scalex

Run scalex performance benchmarks, profiling, and timing analysis.

MITAuto-check passedDevelopment

Install Benchmark

skills CLI
$ npx skills add nguyenyou/scalex --skill benchmark -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install nguyenyou/scalex benchmark --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/nguyenyou/scalex.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/benchmark .claude/skills/benchmark && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
benchmark
GitHub stars
109
Token cost
~3.4k tokens
SKILL.md length
1,043 words
Files
3 (incl. scripts)
Skills in repo
3
Repo updated
First seen
Licence
MIT

At a glance

Run scalex performance benchmarks, profiling, and timing analysis.

  • Works in 8 steps: timings to identify which phase to… → BENCH_EXPORT=before.json bench.sh to… → If deeper analysis needed:… → …
  • The user asks to benchmark scalex
  • SKILL.md covers Overview, Decision guide, Layer 1: --timings flag and Layer 2: hyperfine benchmarks…, plus 7 more sections
  • Runs Shell scripts from its folder; calls brew and xcrun

What it does

Benchmark is an agent skill from nguyenyou/scalex. Run scalex performance benchmarks, profiling, and timing analysis. Use this skill whenever the user asks to benchmark scalex, measure performance, profile index/query times, compare before/after performance of a change, investigate bottlenecks, or mentions "benchmark", "perf", "how fast", "timing", "hyperfine", "profile", "flame graph", "profiling", "--timings", "slow", "bottleneck", "regression", "memory", "heap", "GC", "allocation". Also use proactively after implementing performance improvements to verify…

Its SKILL.md is about 3.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files, including scripts (for example `scripts/bench-compare.sh` and `scripts/bench.sh`).

It sits in Development, covering Performance optimization. The repository describes itself as: Scala code intelligence for coding agents. Zero Build Server. Zero Compilation. Just answers. The licence is MIT.

When your agent uses it

  • The user asks to benchmark scalex
  • Measure performance
  • Profile index/query times
  • Compare before/after performance of a change

Example prompts

  • “benchmark”
  • “how fast”
  • “timing”
  • “/benchmark”

Requirements

  • A Bash shell

Workflow steps

8 steps, taken from the first numbered list in SKILL.md.

  1. timings to identify which phase to optimize
  2. BENCH_EXPORT=before.json bench.sh to capture baseline
  3. If deeper analysis needed: async-profiler flame graph or JFR
  4. Make the change
  5. timings to verify phase improvement
  6. BENCH_EXPORT=after.json bench.sh to capture new numbers
  7. bench-compare.sh before.json after.json to check for regressions
  8. Microbenchmarks if isolating a specific function

What it can do on your machine

Read from SKILL.md and the folder at commit 9098af8. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 2 files in scripts/ (Shell), which the agent can run.

    Shell commands in SKILL.md call:

    • brew
    • xcrun

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Benchmark loads about 3.4k tokens when it runs. Until then it costs about 168 tokens; SKILL.md has 1,043 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~168
When it runs · the whole SKILL.md, loaded when a task matches
~3.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from nguyenyou/scalex at commit 9098af8, republished under its MIT licence (© nguyenyou). 1,043 words, ~3,434 tokens.

Download SKILL.mdSave it as .claude/skills/benchmark/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
benchmark
description
Run scalex performance benchmarks, profiling, and timing analysis. Use this skill whenever the user asks to benchmark scalex, measure performance, profile index/query times, compare before/after performance of a change, investigate bottlenecks, or mentions "benchmark", "perf", "how fast", "timing", "hyperfine", "profile", "flame graph", "profiling", "--timings", "slow", "bottleneck", "regression", "memory", "heap", "GC", "allocation". Also use proactively after implementing performance improvements to verify gains. Covers 6 layers: built-in --timings, hyperfine benchmarks, async-profiler flame graphs, JFR recording, microbenchmarks, and memory profiling.

Overview

Scalex has a multi-layered profiling and benchmarking system. Pick the right layer for the situation:

LayerToolWhen to useWorks in native?
1. --timingsBuilt-in flagQuick phase breakdown, first look at any perf questionYes
2. hyperfinebench.shReproducible before/after comparison with statisticsYes
3. async-profilerprofiling/profile.shDeep CPU/alloc/lock flame graphs to find hotspotsJVM only
4. JFRprofiling/scalex.jfcGC pressure, file I/O patterns, thread utilizationJVM only
5. Microbenchmarkssrc/bench.scalaIsolate per-function cost with warmup + statisticsJVM only
6. Memory profilingbench.sh memoryHeap usage, GC pressure, peak memory across scenariosJVM only

Decision guide

"Where is time spent?" → Start with --timings (Layer 1)

"Is this change faster?" → Use hyperfine before/after (Layer 2), optionally with bench-compare.sh

"Why is parsing slow?" → async-profiler CPU flame graph (Layer 3)

"Why are allocations high?" → async-profiler alloc or JFR ObjectAllocationSample (Layer 3/4)

"Is there GC pressure?" → JFR (Layer 4)

"How fast is extractSymbols on one file?" → Microbenchmark (Layer 5)

"How much memory does indexing use?" → Memory profiling (Layer 6)

"Is there a memory leak or GC regression?" → Memory profiling before/after (Layer 6)


Layer 1: --timings flag

The fastest way to see where time goes. Works in both JVM and native image. Prints to stderr.

bash
# Cold index phase breakdown
rm -rf benchmark/scala3/.scalex
./scalex index benchmark/scala3 --timings

# Warm index
./scalex index benchmark/scala3 --timings

# Query with bloom/text-search breakdown
./scalex refs benchmark/scala3 Compiler --timings

# JVM mode
./mill run index benchmark/scala3 --timings
Phases reported

Index phases: git-ls-files, cache-load, oid-compare, parse, cache-save, and lazy build-* lookup phases

Query phases (refs/imports/coverage): bloom-screen, text-search

Durations are exclusive per thread: nested phases are subtracted from their parent. command and render expose query and output work. request-total measures elapsed application request time, including otherwise uninstrumented work, but not runtime startup before the entry point. Use hyperfine for complete process latency. Concurrent phases can overlap, so summing phases is not a wall-clock measurement. Batch mode reports the shared index load, then a separate request total for each query.

Reading the output
Timings:
  git-ls-files          12.3 ms  ( 1%)
  cache-load            45.2 ms  ( 5%)
  oid-compare            3.1 ms  ( 0%)
  parse                782.0 ms  (80%)
  index-build           89.4 ms  ( 9%)
  cache-save            42.1 ms  ( 4%)
  request-total        974.1 ms
  • parse > 70%: Scalameta parsing dominates — look at parallelism, parser options, or reducing parsed file count
  • cache-load > 20%: Index deserialization — check bloom skip, file size, buffering
  • cache-save > 10%: Index serialization — check if save is running unnecessarily (when parsedCount == 0)
  • bloom-screen > text-search: Bloom filter is slow — check expected element count, FPP
  • text-search >> bloom-screen: Text scanning dominates — check candidate count (bloom too permissive?)

Layer 2: hyperfine benchmarks (bench.sh)

Reproducible, statistical benchmarks using hyperfine against the Scala 3 compiler repo (~17.7K files).

Prerequisites
  • hyperfine installed (brew install hyperfine)
  • Native scalex binary built (./build-native.sh)
  • The scala3 repo is cloned automatically on first run
Running
bash
# Full suite (cold + warm + query + diverse + timings)
.agents/skills/benchmark/scripts/bench.sh

# Individual modes
.agents/skills/benchmark/scripts/bench.sh cold
.agents/skills/benchmark/scripts/bench.sh warm
.agents/skills/benchmark/scripts/bench.sh query
.agents/skills/benchmark/scripts/bench.sh diverse    # miss, heavy refs, fuzzy, grep, hierarchy
.agents/skills/benchmark/scripts/bench.sh timings    # --timings output for cold/warm/refs
.agents/skills/benchmark/scripts/bench.sh memory     # heap usage, GC pressure (JVM only)

# Custom runs/binary
BENCH_RUNS=10 SCALEX_BIN=./target/scalex .agents/skills/benchmark/scripts/bench.sh
Before/after comparison
bash
# 1. Benchmark current state
BENCH_EXPORT=benchmark/results/before.json .agents/skills/benchmark/scripts/bench.sh

# 2. Make changes, rebuild
./build-native.sh

# 3. Benchmark new state
BENCH_EXPORT=benchmark/results/after.json .agents/skills/benchmark/scripts/bench.sh

# 4. Compare (flags >5% regressions)
.agents/skills/benchmark/scripts/bench-compare.sh benchmark/results/before.json benchmark/results/after.json

bench-compare.sh exits non-zero if any benchmark regressed >5%.

Typical ranges (Apple M3 Max)
MetricRangeBottleneck
Cold index3-5sScalameta parsing (CPU-bound, parallel)
Warm index0.8-1.0sOID compare + index load
Query (any)1.2-1.5sIndex deserialization from disk
refs (heavy symbol)2-4sText search across candidate files

Layer 3: async-profiler flame graphs

Reveals call-stack-level CPU hotspots. No code changes needed — JVM agent only.

Prerequisites
bash
brew install async-profiler
# Or set AP_HOME to your installation
Running
bash
# CPU flame graph of cold index
./profiling/profile.sh benchmark/scala3

# Wall-clock (includes I/O wait — useful for parallelStream bottlenecks)
./profiling/profile.sh benchmark/scala3 wall

# Allocation hotspots (where objects are created)
./profiling/profile.sh benchmark/scala3 alloc

# Lock contention (parallelStream synchronization)
./profiling/profile.sh benchmark/scala3 lock

Output: profiling/profile-<event>.html — open in browser for interactive flame graph.

What to look for
  • CPU: Wide bars in Scalameta parsing → specific parser methods. Wide bars in parallelStream infrastructure → overhead from small task granularity.
  • Alloc: Hot allocation sites → potential for object reuse or structural changes.
  • Lock: Contention in ConcurrentLinkedQueue.add() → batch results instead of per-item add.
  • Wall: I/O wait in Files.readAllLines or Files.readString → potential for memory-mapped I/O.

Layer 4: JFR (Java Flight Recorder)

Built into the pinned JDK. Near-zero overhead. Best for GC, file I/O, and thread analysis.

Running
bash
# Record with custom config
SCALEX_JAVA_OPTS="-XX:StartFlightRecording=filename=profiling/scalex.jfr,settings=profiling/scalex.jfc,duration=60s" \
  ./mill run index benchmark/scala3

# Quick summary
jfr summary profiling/scalex.jfr

# Specific events
jfr print --events jdk.ObjectAllocationSample profiling/scalex.jfr | head -100
jfr print --events jdk.GarbageCollection profiling/scalex.jfr
jfr print --events jdk.FileRead profiling/scalex.jfr | head -50
jfr print --events jdk.ThreadPark profiling/scalex.jfr | head -50

# GUI analysis
open profiling/scalex.jfr  # Opens in JDK Mission Control
Custom config

profiling/scalex.jfc is tuned for scalex — it enables allocation sampling, GC events, file I/O (>1ms threshold), and thread parking/monitor events.


Layer 5: Microbenchmarks (src/bench.scala)

Isolate per-function costs with warmup and statistical measurement.

Running
bash
# Specific benchmark
./mill bench.run extract-single benchmark/scala3
./mill bench.run bloom-build benchmark/scala3
./mill bench.run persistence-load benchmark/scala3
./mill bench.run search benchmark/scala3
./mill bench.run refs benchmark/scala3

# All benchmarks
./mill bench.run all benchmark/scala3

# Custom warmup/iterations
./mill bench.run extract-single benchmark/scala3 --warmup 3 --iterations 10
Available benchmarks
BenchmarkWhat it measures
extract-singleextractSymbols on the largest file
extract-batchextractSymbols on 100 files (sequential AND parallel)
bloom-buildbuildBloomFilterFromSource on a large source
persistence-loadIndexPersistence.load with and without bloom deserialization
searchWorkspaceIndex.search("Compiler") warm
refsfindReferences("Phase") warm
index-coldFull cold index including map building

Reports: mean, median, p99, stddev, min, max per benchmark.


Show full SKILL.md (407 more words)Show less

Layer 6: Memory profiling (bench.sh memory)

Measures heap usage, GC pressure, and peak memory across three scenarios: cold index (full parse), warm index (cache load), and refs query. Uses JVM GC logging via -Xlog:gc* — requires Mill (JVM mode), not native binary.

Running
bash
# Full memory profile (cold + warm + refs)
.agents/skills/benchmark/scripts/bench.sh memory

Uses the checked-in Mill launcher. Does not require hyperfine or native binary.

Output

Reports per-scenario: phase timings (from --timings), then memory stats:

--- Cold index (full parse, ~17.7k files) ---
  git-ls-files            60.7 ms  ( 1%)
  parse                 5576.5 ms  (94%)
  cache-save             228.7 ms  ( 4%)
  total                 5952.2 ms

  Peak pre-GC heap:        790 MB
  Heap at exit (used):     676 MB
  Heap at exit (committed): 1184 MB
  GC pauses:               34
  Total GC pause time:     185.3 ms

Ends with a summary table:

=== Memory Summary ===

Scenario             Peak Heap      Exit Used    Exit Commit     GC #
--------             ---------      ---------    -----------     ----
Cold index              790 MB       676.2 MB      1184.0 MB       34
Warm index              256 MB       271.8 MB       584.0 MB        2
refs Phase               53 MB        29.7 MB        56.0 MB        0
Reading the output
  • Peak pre-GC heap: Highest live heap before any GC — the true high-water mark. This is the number that determines minimum -Xmx for constrained environments.
  • Heap at exit (used): Retained heap at process exit — the steady-state footprint of the loaded index + results.
  • Heap at exit (committed): OS-committed memory — what the JVM actually reserved. Higher than "used" because G1 keeps headroom.
  • GC pauses: Number of young-gen collections. High count during cold index is normal (Scalameta ASTs are short-lived).
  • Total GC pause time: Sum of all GC pauses. Should be small relative to wall time (<5%).
What to watch for
  • Cold peak > 1 GB: parallelStream is creating too many concurrent ASTs. Consider batching files in chunks.
  • Warm used > 400 MB: Index deserialization is holding too much data. Check if lazy maps are working.
  • Refs used > 100 MB: Text search is accumulating results. Check if bloom filtering is effective.
  • GC time > 10% of wall time: GC pressure is impacting performance. Look at allocation hotspots (async-profiler alloc, Layer 3).
Baseline ranges (scala3 corpus, 17.7k files, ~38k symbols)
ScenarioPeak HeapExit UsedGC Pauses
Cold index700-900 MB600-700 MB30-70
Warm index200-300 MB250-300 MB1-3
refs query30-60 MB25-35 MB0-1
Ad-hoc memory profiling

For one-off measurements without the script:

bash
# Run any scalex command with GC logging to a temp file
SCALEX_JAVA_OPTS="-Xlog:gc*=info:file=/tmp/scalex-gc.log" \
  ./mill run overview benchmark/scala3

# Check peak heap
grep -o '[0-9]*M->' /tmp/scalex-gc.log | sed 's/M->//' | sort -n | tail -1
# Check exit heap
grep "garbage-first" /tmp/scalex-gc.log | tail -1

Native image profiling

async-profiler and JFR don't work with GraalVM native images. Options:

bash
# macOS Instruments (Time Profiler)
xcrun xctrace record --template "Time Profiler" --launch -- ./scalex index scala3

# --timings always works
./scalex index scala3 --timings

Performance budget (from AGENTS.md)

When evaluating changes:

  • Index times: Accept <5% regression
  • Index size: 0% growth for non-index features, <10% if schema changes
  • Query latency: No regression
  • Feature gate: "Is this better than grep, or does it introduce a worst case?"

The bench-compare.sh script automates the regression check against these thresholds.

Typical workflow for a performance change

  1. --timings to identify which phase to optimize
  2. BENCH_EXPORT=before.json bench.sh to capture baseline
  3. If deeper analysis needed: async-profiler flame graph or JFR
  4. Make the change
  5. --timings to verify phase improvement
  6. BENCH_EXPORT=after.json bench.sh to capture new numbers
  7. bench-compare.sh before.json after.json to check for regressions
  8. Microbenchmarks if isolating a specific function

© nguyenyou, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (scripts) in .agents/skills/benchmark of nguyenyou/scalex.

  • SKILL.md
  • scripts/bench-compare.sh
  • scripts/bench.sh

Open the folder on GitHubat commit 9098af8

Compare with similar skills

Benchmark next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Benchmark compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Benchmark this skillnguyenyou/scalex109—~3.4kAutomated safety check: PassMIT
Code Review ChecklistshareAI-lab/learn-claude-code78k5 repos~1.1kAutomated safety check: PassMIT
LLM Torch Profiler Analysissgl-project/sglang37k2 repos~6.4kAutomated safety check: PassApache-2.0
Pycrazyguitar/pysheeet8.2k—~886Automated safety check: PassMIT
Cmux Debugging Guidemanaflow-ai/cmux28k1 repos~1.1kAutomated safety check: PassCustom licence
Electron Heap Snapshot Analysiskeybase/client9.3k—~875Automated safety check: PassBSD-3-Clause

Similar skills

  • Code Review Checklist

    shareAI-lab/learn-claude-code

    Reviews code against a five-part checklist covering security, correctness, performance, maintainability and testing, and reports findings in a fixed format.

    78k GitHub starsUsed in 5 repos~1.1k tokens
    DevelopmentAuto-check passed
  • LLM Torch Profiler Analysis

    sgl-project/sglang

    Unified LLM torch-profiler triage skill for sglang, vllm, TensorRT-LLM, and TokenSpeed.

    37k GitHub starsUsed in 2 repos~6.4k tokens
    DevelopmentAuto-check passed
  • Py

    crazyguitar/pysheeet

    Comprehensive Python programming reference covering syntax, concurrency, networking, databases, ML/LLM development, and HPC.

    8.2k GitHub stars~886 tokensUpdated 2 days ago
    DevelopmentAuto-check passed
  • Cmux Debugging Guide

    manaflow-ai/cmux

    Covers debug logging, the Debug menu, profiling rules and runtime pitfalls for working on the cmux macOS terminal app.

    28k GitHub starsUsed in 1 repo~1.1k tokens
    DevelopmentAuto-check passed
  • Analyzes V8, Chrome and Electron .heapsnapshot files with Node scripts to find memory leaks, detached DOM nodes and the retainer paths that keep objects alive.

    9.3k GitHub stars~875 tokensUpdated today
    DevelopmentAuto-check passed
  • Runs controlled JMH experiments on the Caffeine cache to find shared contention and hot-path waste, then reviews correctness and returns a reviewable patch.

    18k GitHub stars~2.6k tokensUpdated today
    DevelopmentAuto-check: notes

More from nguyenyou/scalex

  • Scalex

    nguyenyou/scalex

    Explore and navigate Git-tracked Scala 2/3 and Java source with Scalex.

    109 GitHub stars~2k tokensUpdated 21 days ago
    Auto-check passed
  • Capture Banners

    nguyenyou/scalex

    Re-render Scalex banner and OG image PNGs from their HTML source files using Chrome DevTools MCP.

    109 GitHub stars~586 tokensUpdated 21 days ago
    Auto-check passed

Categories

Questions about Benchmark

What does Benchmark do?

Run scalex performance benchmarks, profiling, and timing analysis. Benchmark is an agent skill from nguyenyou/scalex. Run scalex performance benchmarks, profiling, and timing analysis.

When should I use Benchmark?

Benchmark fits situations like: the user asks to benchmark scalex; measure performance; profile index/query times; compare before/after performance of a change.

How do I install Benchmark in Claude Code?

Run `npx skills add nguyenyou/scalex --skill benchmark -a claude-code`. Or copy the skill folder (.agents/skills/benchmark in nguyenyou/scalex) into .claude/skills/benchmark in your project. Claude Code loads it when a task matches its description.

How do I install Benchmark in Codex?

Run `npx skills add nguyenyou/scalex --skill benchmark -a codex`. Or copy the skill folder (.agents/skills/benchmark in nguyenyou/scalex) into .agents/skills/benchmark in your project. Codex loads it when a task matches its description.

Can I use Benchmark in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add nguyenyou/scalex --skill benchmark -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/benchmark, .gemini/skills/benchmark, .github/skills/benchmark and .opencode/skills/benchmark in your project.

What does Benchmark need to run?

Going by SKILL.md and its folder, Benchmark needs a shell for the scripts in its folder and the command-line tools its instructions call (brew and xcrun). Our summary lists: A Bash shell.

Does Benchmark access the network?

SKILL.md names 1 domain. As links in the text: github.com. This is read from the text; nothing was executed.

Is Benchmark safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Benchmark use?

Benchmark is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Benchmark use?

About 3.4k tokens (SKILL.md is roughly 14k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Benchmark?

Skills that share tags, products or a category with Benchmark: Code Review Checklist (shareAI-lab/learn-claude-code, 78k stars), LLM Torch Profiler Analysis (sgl-project/sglang, 37k stars), Py (crazyguitar/pysheeet, 8.2k stars) and Cmux Debugging Guide (manaflow-ai/cmux, 28k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Benchmark?

nguyenyou (a GitHub user) maintains it in nguyenyou/scalex, which has 109 GitHub stars. The repository holds 3 skills in this directory. The repository was last updated on September 18, 2026.

Source: nguyenyou/scalex on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.