Agent skill

Perf Labs Perf

by perf-labs in perf-labs/perf

Performance engineering on x86-64 Linux with the perf-labs/perf toolkit (perf benchmark, perf profile, perf analyze, perf view, perf plot, perf compare, perf info) and system perf.

MITAuto-check: notes

Install Perf Labs Perf

skills CLI
$ npx skills add perf-labs/perf --skill perf-labs-perf -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install perf-labs/perf perf-labs-perf --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
perf-labs-perf
GitHub stars
110
Token cost
~5.3k tokens
SKILL.md length
2,438 words
Files
59
Skills in repo
1
Repo updated
First seen
Licence
MIT

At a glance

Performance engineering on x86-64 Linux with the perf-labs/perf toolkit (perf benchmark, perf profile, perf analyze, perf view, perf plot, perf compare, perf info) and system perf.

  • Works in 6 steps: Frame → Baseline first → Hypothesis, one at a time → …
  • Optimizing code
  • SKILL.md covers Non-negotiable rules, Preflight (do this once per…, Which tool answers which… and Workflow, plus 7 more sections
  • Runs Rust scripts from its folder; calls ruff

What it does

Perf Labs Perf is an agent skill from perf-labs/perf. Performance engineering on x86-64 Linux with the perf-labs/perf toolkit (perf benchmark, perf profile, perf analyze, perf view, perf plot, perf compare, perf info) and system perf. Use when profiling or optimizing code, measuring cycles/latency/IPC/cache/TLB/branch behaviour, benchmarking a function, region or asm snippet, attributing a slowdown with top-down counters, or deciding whether a change is a real speedup. Triggers on "perf benchmark", "perf profile", "perf info", "perf stat", "perf record", "rdpmc"…

Its SKILL.md is about 5.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 63 other files (for example `.github/workflows/linux.yml`, `README.md` and `bin/README.md`).

It works with Linux. The licence is MIT.

When your agent uses it

  • Optimizing code
  • Measuring cycles/latency/IPC/cache/TLB/branch behaviour
  • Benchmarking a function
  • Attributing a slowdown with top-down counters

Example prompts

  • “perf benchmark”
  • “perf profile”
  • “perf info”
  • “/perf-labs-perf”

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Frame
  2. Baseline first
  3. Hypothesis, one at a time
  4. Isolate
  5. Attribute
  6. Verify the fix

What it can do on your machine

Read from SKILL.md and the folder at commit 6d6d902. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Rust, from the files we listed), which the agent can run.

    Shell commands in SKILL.md call:

    • ruff

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Perf Labs Perf loads about 5.3k tokens when it runs. Until then it costs about 159 tokens; SKILL.md has 2,438 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~159
When it runs · the whole SKILL.md, loaded when a task matches
~5.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteRuns commands with sudoSKILL.md:49
    echo 2 | sudo tee /sys/devices/{cpu_core,cpu_atom}/rdpmc

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from perf-labs/perf at commit 6d6d902, republished under its MIT licence (© perf-labs). 2,438 words, ~5,279 tokens.

Download SKILL.mdSave it as .claude/skills/perf-labs-perf/SKILL.md (or your agent's skills folder). This skill also uses 58 other files; get the full folder from GitHub.
name
perf-labs-perf
description
Performance engineering on x86-64 Linux with the perf-labs/perf toolkit (perf benchmark, perf profile, perf analyze, perf view, perf plot, perf compare, perf info) and system perf. Use when profiling or optimizing code, measuring cycles/latency/IPC/cache/TLB/branch behaviour, benchmarking a function, region or asm snippet, attributing a slowdown with top-down counters, or deciding whether a change is a real speedup. Triggers on "perf benchmark", "perf profile", "perf info", "perf stat", "perf record", "rdpmc", "topdown", "IPC", "cycle count", "benchmark", "profile", "is it faster", "cache/TLB bound", "PERF_LABEL".

Performance engineering with perf

You are an experienced performance engineer. You do not guess, you do not hand-wave a benchmark, and you never report a number you did not measure. You work the loop: frame the question → measure a baseline → form one hypothesis → isolate it with an experiment → attribute the cycles → verify the fix with a test.

Non-negotiable rules

  1. Never report an unmeasured claim. "This should be faster because the loop is unrolled" is a hypothesis. perf benchmark rows are evidence.
  2. Always give the unit and the denominator. cycles/operations, ns/operation, IPC, p50/p99 — never a bare number.
  3. Compare like with like. Same mode, same config, same CPU, same binary except the change. Use perf compare to decide whether a difference is real; do not eyeball two tables.
  4. A/B the environment too. If a change is under ~3%, suspect the machine (frequency scaling, migrations, neighbours) before the code. Re-run, pin, then judge.
  5. Attribute before optimizing. "Which bound is it?" (front-end, back-end, bad speculation, retiring) comes before "which line is slow?".
  6. State the uncertainty. Sample count, spread (p10..p99), and whether the effect cleared perf compare's significance test. If it did not, say "no measurable difference".
  7. Do not change the workload to flatter it. Cache/TLB/branch state is part of the question — vary it deliberately and say which state you measured.
  8. Leave the machine as you found it. Restore affinity, priority, NUMA binding; delete scratch binaries you built under /tmp.

Preflight (do this once per session, in this order)

sh
uname -r                                   # 6.x+ required
perf info cpu                              # topology, TSC freq, L1i/L1d/L2/L3
ls /sys/devices/{cpu_core,cpu_atom}/rdpmc  # user-space rdpmc
cat /sys/devices/system/cpu/cpu0/cpufreq/scaling_governor 2>/dev/null

If rdpmc reads as 0, every hardware event needs a perf_event_open syscall instead of rdpmc:

sh
echo 2 | sudo tee /sys/devices/{cpu_core,cpu_atom}/rdpmc

Without it, duration_time (rdtsc) still works; hardware events do not. On hybrid CPUs, events only exist on one core type — perf benchmark pins itself to a PMU-capable CPU, and if you drive perf stat by hand you must taskset -c <p-core> yourself.

perf info a.out lists every label and function in a binary, with addresses, before you write a single target name.

Which tool answers which question

QuestionCommand
What does this op cost in isolation?perf benchmark 'imul eax, 42' --mode latency -e cycles,duration_time
What does this function/region cost?perf benchmark a.out:fizz_buzz -m latency -e cycles,instructions
Latency or throughput bound?-m latency vs -m throughput (same target, two answers)
Does it depend on its input?--data.arg0=1 / --data.arg0=[1,3,5]
What does a program cost with real arguments?perf benchmark /usr/bin/tree:main -m latency -- /path/to/folder (and --env K=V)
Branch mispredicts?--config.branch=predictable vs unpredictable
Cache bound?--config.dcache=hot,warm,cool,cold (and icache)
TLB bound?--config.dtlb=hot,cold (and itlb)
Which microarchitectural bound?-e 'topdown-*'
Where does a whole program spend time?perf profile -f hot -e cycles -- ./a.out, else system perf record + perf report
Which instructions of this target are hot?perf analyze a.out:fizz_buzz -- perf.data (one numbered row per instruction, per-ip events joined)
Which state do these instructions run with?perf analyze a.out:fizz_buzz (data.<reg> per instruction) and --filter on it
How long is a label of an assembly source?perf benchmark foo.s:foo..bar
Is this change actually faster?perf benchmark ... -o data/ for both, then perf compare -- data/
What is the noise floor?perf view --stat min,median,p10,p90,p99 -- data/

Workflow

1. Frame

Write down the question as a number: "how many cycles does fizz_buzz(n) take per call at n=1e6", not "is fizz_buzz slow". Decide the unit of work (operations in perf benchmark is one call / one loop / one snippet execution — state which).

2. Baseline first
sh
perf benchmark a.out:fizz_buzz -m latency -e duration_time,cycles,instructions

Read the emitted row as a self-describing record: mode, iterations, samples, operations, the config.* columns and the event columns. Compute IPC = instructions/cycles and cycles/operations. Sanity-check samples (≈100 by default) and iterations (auto-calibrated, so the loop is long enough to dominate the harness).

3. Hypothesis, one at a time

Ask in this order, and stop at the first "yes":

  • IPC far below the machine width (~4-6 on a modern core)? → dependency-bound or port-bound, not throughput-bound.
  • cycles/operations not close to a round number (the documented latency, or the next power of two)? → something else is in the dependency chain.
  • Big gap between -m latency and -m throughput? → per-call overhead (latency-bound), or the loop hides the cost (throughput-bound).
  • Big gap between --config.branch=predictable and unpredictable? → branch misses dominate; look at branch-misses and layout/alignment.
  • Big gap between dcache=hot and dcache=cold? → the working set does not fit L1 (or the code/data streams fight over L1).
  • Big gap between dtlb=hot and dtlb=cold? → TLB misses; huge pages or fewer pages touched will help.
  • topdown-be-bound high? → memory; fe-bound → front-end/decoder/ITLB; bad-spec → branches; retiring high with low IPC → issue width or a dependency chain.
4. Isolate

Turn the hypothesis into one controlled sweep, one axis at a time:

sh
# input dependence
perf benchmark a.out:fizz_buzz --data.arg0=[1,3,5] -m latency -e cycles

# branch predictability, cache and TLB state, alignment
perf benchmark a.out:fizz_buzz --config.branch=predictable,unpredictable
perf benchmark a.out:fizz_buzz --config.dcache=hot,cold
perf benchmark a.out:fizz_buzz --config.dtlb=hot,cold
perf benchmark a.out:fizz_buzz --config.code=1,32

# regions, to attribute cost inside a function
perf benchmark a.out:hot_begin..hot_end -m latency -e cycles

# a program's entry point, called the way a shell would call it
perf benchmark /usr/bin/tree:main -m latency --env LANG=C.UTF-8 -- /path/to/folder

Put a region label around the suspect code with PERF_LABEL(name) (lib/perf/perf.h for C/C++, lib/perf/perf.rs, lib/perf/perf.zig); labels emit zero instructions, and perf benchmark/perf profile turn them into counter-reading trampolines at startup. A foo_begin/foo_end pair is one region; foo_begin..foo_end is the target.

5. Attribute
sh
perf benchmark a.out:fizz_buzz -m latency -e 'topdown-*'

Read the four level-1 slots (they sum to ~100% of slots) and drill into the one that dominates. Cross-check with raw counters — top-down says which, counters say how much. The slots are a documented alias, not a hardware guarantee: on a CPU whose PMU does not export them this fails with unknown event 'topdown-retiring' rather than reporting zeros, so confirm with perf list | grep topdown first and fall back to the raw counters if the host has no top-down PMU:

SignalMeaning
IPC ≈ 0.3, latency ≫ throughputdependency chain, one load-use or FP latency
IPC ≈ 1one dependent chain per cycle
branch-misses ≫ 0 per branchunpredictable control flow; try inlining/order
cache-misses high and dcache=cold ≫ hotworking set > L1; shrink or block it
L1-dcache-load-misses ≫ LLC-load-missesL1 capacity/conflict, not DRAM
big cold vs hot gap with high retiringthe core retires fast, the load is the bound
dtlb=cold gap ≫ dtlb=hot gappage-walk bound; fewer/huger pages
top-down fe-bound + small itlb=cold gapdecoder/branch-density, not TLB
6. Verify the fix

Change the code, re-run the identical command, and let the statistics decide:

sh
perf benchmark a.old:fizz_buzz -n old --mode latency --event cycles -o data/
perf benchmark a.new:fizz_buzz -n new --mode latency --event cycles -o data/
perf compare -- data/

perf compare runs a two-sided z-test on the arithmetic mean and on the geometric mean and requires both (p = max(p_mean, p_gmean)) to reject at --alpha (default 0.05) before a change is significant. That is what keeps two runs of the same binary from showing up as a win. With no -e it compares every event per operation (cycles/operations), never the raw totals: two runs do a different number of operations, so their counters are never comparable.

Reading perf benchmark output

Columns are the identity of the run followed by what it was measured with:

file  name  mode  iterations  samples  operations  config.*  data.*  cycles  instructions
  • mode: latency (one sample per call) or throughput (one sample for the whole loop). Both are wanted; they answer different questions.
  • iterations/samples: harness trip count and samples collected; the trip count auto-calibrates to a target relative standard error, so a stable row has a stable iterations.
  • operations: denominator for every ratio (cycles/operations).
  • A null counter means it could not be read — never read it as 0.
  • config.backend.<name>.* records the resolved backend and its parameters, so a row can be replayed exactly.

perf view shows time,file,name,mode,samples,duration_time/operations by default — the per-operation cost, not the raw counter — and aggregates (-s min,median,p10,p50,p90,p99,max); -e picks other columns or expressions, -g '' -s '' gives raw rows. perf plot charts <event>/operations by default (ecdf), so plot the per-operation cost, not the raw counter.

Live tracking of a real binary

sh
perf info a.out                                   # what is trackable
perf profile -e cycles,branch-misses -- ./a.out --work 100
perf profile -f fizz_buzz -e cycles -o profile.json -- ./a.out

perf profile patches addresses at startup (ptrace + rdpmc trampolines); the binary on disk is untouched. With no -f it tracks every function and label, which answers "what does this process spend cycles in" without sampling bias. perf info <file> is the one place that lists what is trackable.

Per-instruction view of a target

sh
perf analyze a.out:fizz_buzz                                  # every state
perf analyze a.out:fizz_buzz -- perf.data                     # + per-ip events
perf analyze a.out:fizz_buzz --filter 'latency > 4'           # only those instructions
perf analyze a.out:fizz_buzz --filter '15 in `data.rdi`'      # only that state
perf analyze a.out:fizz_buzz -e assembly,latency              # pick the columns
perf analyze a.out:fizz_buzz -e 'index,assembly,data*'         # or the state columns
perf analyze a.out:fizz_buzz -e instructions/cycles -- perf.data
perf analyze foo.s:foo..bar                                   # an assembly source
perf analyze a.out:fizz_buzz --data.rdi=15                    # a concrete state
perf analyze a.out:fizz_buzz --setup init --teardown fini       # with set-up
perf analyze a.out:fizz_buzz -e assembly | llvm-mca -mcpu=alderlake  # or into llvm-mca

perf analyze never runs anything. index numbers the instructions 0, 1, 2, ...; it is a column like any other — in the default selection, and printed first when it is selected, so a run can be counted directly. The target is explored symbolically, so all states are analyzed: the rows are every instruction the target disassembles to (the whole function or region, followed through its branches), and the columns are what the explored states held — data.<reg> for every register a state pins (the arguments and whatever --data constrains, the very values perf benchmark measures) and data.<addr> for an address a state reads or writes. Every data.* cell is a list of what the states held ([15] when they agree, [0, 1, 1073741825] when they do not), and the registers the exploration only had to pin to keep going are not data and are left out. --filter takes a pandas query over any column (size, latency, data.rdi, ...), so instructions can be selected by the state they run with — in is how you test a state column (15 in \data.rdi`); -e/--eventpicks the columns to show — any of them, with* expanding a pattern (-e data*, -e '*') and an expression allowed (-e instructions/cyclesover the per-ip counters joined from--); a name that is not in the result is an error, and the default is file,name,index,address,encoding,size,latency,throughput,assembly,data*. Asked for assemblyalone, the table's header is.intel_syntax, so it pipes straight into llvm-mca: perf analyze a.out:fizz_buzz -e assembly | llvm-mca -mcpu=alderlake. --datapins the explored state to concrete values (same meaning asperf benchmark --data), and --setup/--teardownrun around the target, exactly as inperf benchmark`.

The result is one table: data given after -- is joined by ip, so a perf.data turns the table into "cycles per instruction"; runs without ip (fizz_buzz.json, profile.json) only contribute their file,name identity.

Only the target's own instructions are listed: the harness's timing reads, cache/TLB steering, register priming and call sequence are never attributed to the target. Use perf analyze to attribute a measured hot spot to instructions; use perf benchmark for the per-operation cost of one.

Show full SKILL.md (840 more words)Show less

System perf, and when to use it instead

The perf-labs tools are for isolated cost. For whole-program behaviour, profile with system perf. One perf dispatches both: it runs perf-<command> when that script exists and otherwise falls through to linux-perf, so perf stat, perf record, perf report, perf annotate, perf script, perf c2c, perf mem, perf lock, perf sched and perf probe all work next to perf benchmark and perf analyze.

sh
taskset -c 4 perf stat -e cycles,instructions,cache-misses,branch-misses ./a.out
perf record -g -F 999 -e cycles:u -o perf.data -- ./a.out
perf report --stdio -g graph,4000 --sort symbol
perf annotate --stdio -s symbol.dso
perf c2c record -g -o c2c.data -- ./a.out

Pitfalls seen in real reviews

  • Latency measured as throughput (or the reverse). Compare like modes.
  • Both modes in one run, reading one column: perf benchmark 'imul eax, 42' -e cycles | perf view -s p99 mixes the latency and throughput rows. Pass --mode latency explicitly.
  • Per-call overhead mistaken for loop cost: always look at latency vs throughput before blaming the operation.
  • Percentiles from one sample are noise. samples is 100 by default (--config.samples=N); a row that wants more is a row to re-run.
  • Sweeping two axes and reading the corner: --config.dcache=hot,cold is fine; --config.dcache=[{L1d:100},{L1d:0}] with --data sweeps is a factorial explosion. One axis at a time.
  • Cache state you did not choose: default is a dcache/dtlb sweep, so the default row is one point of a sweep, not "the" number. Say which tier.
  • Attributing harness cost to the target: it is subtracted differentially, but only the target's own instructions are ever listed, so a long call/ret or a big prologue still shows up as the target's cost. Compare against a neighbouring label before blaming a function boundary.
  • icache=cold looking like icache=hot: x86-64 has no user-mode way to flush the instruction cache (clflushopt only reaches the data hierarchy), so icache steers the code's data-hierarchy line and its instruction translation only. Do not read an icache row as "the code is out of L1i"; it is the itlb row and front-end (topdown-fe-bound) that say anything about the instruction stream.
  • Region spanning labels that moved: perf info warns when foo_end precedes foo_begin; the span between them is still measured, but the region is not what the source suggests.
  • Turbos/scaling governor moving the baseline between two runs: re-measure the baseline in the same session as the candidate.
  • Statistical noise called a speedup: perf compare, not a diff of medians.

Reporting

Report like an engineer who wants to be believed:

fizz_buzz(n=1e6), latency mode, pinned cpu 4, 100 samples, 26732 iterations
baseline   10.00 cycles/op   p10 9  p50 10  p90 11
optimized   6.00 cycles/op   p10 5  p50 6  p90 7
perf compare: -40.0% [-41.2, -38.6] p<1e-4 -> significant
topdown: retiring 62% -> 71% (bad-spec 21% -> 9%): the mispredicts are gone

Include: the exact commands, the unit, sample count and spread, the significance verdict, the top-down/attribution evidence, and what is still unexplained. If a question cannot be answered with the counters available, say so and name the experiment that would answer it.

Working on this repository

  • src/perf/core.py — event resolution, perf_event_open, RDPMC, affinity/ priority/NUMA guards. src/perf/bench.py — symbolic exploration, harness JIT, cache/TLB/branch steering, run calibration, data-page mapping. src/perf/code.py — perf analyze (numbered instructions, one row per explored state, list-valued data.*). src/perf/exec.py — ELF loading, relocation, to_object. asm_labels in info.py maps an assembly source's labels to their position and size, which bench.py turns into a snippet behind perf benchmark foo.s:foo..bar. src/perf/arch/x86_64.py — harness templates, counter reads, eviction/priming asm, page_runs (the page clusters a TLB mprotect covers and bench.py maps up front). src/perf/prof.py — ptrace detours, ring buffer, shadow stack. src/perf/info.py — perf info (cpu topology, labels, functions). src/perf/data.py — result-frame schema, perf.data parsing, query with in over list columns. src/perf/comp.py — the CLT test. src/perf/plot.py — charts, and the sixel backend.
  • Each command is one script named perf-<command>, which system perf dispatches to from perf <command>, so perf benchmark runs bin/perf-benchmark. A script is self-contained: its own parser, its own command, and only the helpers it uses (.perfconfig for its own section, the result loading/formatting its input needs). There is no shared CLI module — keep it that way, and keep a command from reaching into another.
  • A target is one CODE argument everywhere: FILE:TARGET for a func or a begin..end region, FILE:LABEL for an assembly source, and a raw snippet when there is no FILE:. perf info FILE is the only way to list a file's targets, and a target that is not found prints them before it exits. The same value is the only positional argument of the python API (perf.benchmark(code=...), perf.analyze(code=...), perf.to_object(code=...)), as a string ("a.out:fizz_buzz"), a [file, target] pair, or an asm snippet. There is no file=/target=/asm= form; do not add one back.
  • Result-frame identity columns are data._IDENTITY_COLUMNS (file, name, mode); columns that are never metrics are data._NON_METRIC_COLUMNS. Add new columns there, not at every use site.
  • Only the license header is kept: no comments and no docstrings anywhere in src, tests or bin/perf-*. Name things well instead.
  • Development loop: pytest, ruff check src tests bin, ruff format --check src tests bin. ruff is pinned to a release series in pyproject.toml; a ruff upgrade and the reformat it causes go in one commit. tests/README.md explains what is tested and what the suite enforces; example.md has worked, verified invocations of every command if a flag needs showing rather than describing. The ```py blocks in every *.md are format-checked by the suite, so keep them ruff-format clean.
  • studies/x86_64/** are the worked notebook analyses (one per level-1 top-down slot); treat them as the reference for expected numbers and method. studies/README.md explains the method. lib/README.md documents the PERF_LABEL annotations.

© perf-labs, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 58 other files in the repository root of perf-labs/perf.

  • SKILL.md
  • .github/workflows/linux.yml
  • .perfconfig
  • LICENSE
  • README.md
  • bin/README.md
  • bin/perf-analyze
  • bin/perf-benchmark
  • bin/perf-compare
  • bin/perf-info
  • bin/perf-plot
  • bin/perf-profile
  • bin/perf-view
  • lib/README.md
  • lib/perf/perf.h
  • lib/perf/perf.rs
  • … and 43 more

Open the folder on GitHubat commit 6d6d902

Compare with similar skills

Perf Labs Perf next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Perf Labs Perf compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Perf Labs Perf this skillperf-labs/perf110—~5.3kAutomated safety check: NotesMIT
Configuring Horizoncoollabsio/coolify63k4 repos~898Automated safety check: PassMIT
Model Usageopenclaw/openclaw392k1 repos~637Automated safety check: PassMIT
Engine Whats Newflutter/flutter179k—~978Automated safety check: PassBSD-3-Clause
Openclaw Live Updateropenclaw/openclaw392k—~3.7kAutomated safety check: PassMIT
Upgrade Browserflutter/flutter179k—~1.1kAutomated safety check: PassBSD-3-Clause

Similar skills

  • Configuring Horizon

    coollabsio/coolify

    A skill your agent uses whenever the user mentions Horizon by name in a Laravel context.

    63k GitHub starsUsed in 4 repos~898 tokens
    Backend & APIsAuto-check passed
  • Model Usage

    openclaw/openclaw

    Summarize CodexBar local cost logs by model for Codex or Claude, including current or full breakdowns.

    392k GitHub starsUsed in 1 repo~637 tokens
    Auto-check passed
  • Engine Whats New

    flutter/flutter

    Generates the "what's new" release summary and diff file for changes in the Flutter engine (//engine/src/flutter) between two releases (e.g., 3.47 vs 3.44).

    179k GitHub stars~978 tokensUpdated today
    MobileAuto-check passed
  • Openclaw Live Updater

    openclaw/openclaw

    Maintain the canonical live OpenClaw main checkout, macOS LaunchAgent-managed Gateway, local macOS app, exact-head main CI, and recurring full release validation.

    392k GitHub stars~3.7k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Upgrade Browser

    flutter/flutter

    Upgrade browser versions (Chrome or Firefox) in the Flutter Web Engine and/or Framework tests.

    179k GitHub stars~1.1k tokensUpdated today
    MobileAuto-check passed
  • K8s Security Policies

    Cybereason-Public/owLSM

    Comprehensive guide for implementing NetworkPolicy, PodSecurityPolicy, RBAC, and Pod Security Standards in Kubernetes.

    280 GitHub starsUsed in 12 repos~2k tokens
    Backend & APIsAuto-check passed

Works with

Questions about Perf Labs Perf

What does Perf Labs Perf do?

Performance engineering on x86-64 Linux with the perf-labs/perf toolkit (perf benchmark, perf profile, perf analyze, perf view, perf plot, perf compare, perf info) and system perf. Perf Labs Perf is an agent skill from perf-labs/perf. Performance engineering on x86-64 Linux with the perf-labs/perf toolkit (perf benchmark, perf profile, perf analyze, perf view, perf plot, perf compare, perf info) and system perf.

When should I use Perf Labs Perf?

Perf Labs Perf fits situations like: optimizing code; measuring cycles/latency/IPC/cache/TLB/branch behaviour; benchmarking a function; attributing a slowdown with top-down counters.

How do I install Perf Labs Perf in Claude Code?

Run `npx skills add perf-labs/perf --skill perf-labs-perf -a claude-code`. Or copy the skill folder (the perf-labs/perf repository) into .claude/skills/perf-labs-perf in your project. Claude Code loads it when a task matches its description.

How do I install Perf Labs Perf in Codex?

Run `npx skills add perf-labs/perf --skill perf-labs-perf -a codex`. Or copy the skill folder (the perf-labs/perf repository) into .agents/skills/perf-labs-perf in your project. Codex loads it when a task matches its description.

Can I use Perf Labs Perf in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add perf-labs/perf --skill perf-labs-perf -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/perf-labs-perf, .gemini/skills/perf-labs-perf, .github/skills/perf-labs-perf and .opencode/skills/perf-labs-perf in your project.

What does Perf Labs Perf need to run?

Going by SKILL.md and its folder, Perf Labs Perf needs Rust for the scripts in its folder and the command-line tools its instructions call (ruff).

Does Perf Labs Perf access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Perf Labs Perf safe to install?

Our automated static check of SKILL.md found notes only (runs commands with sudo), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Perf Labs Perf use?

Perf Labs Perf is published under the MIT licence (from the LICENSE file in the skill folder). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Perf Labs Perf use?

About 5.3k tokens (SKILL.md is roughly 21k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Perf Labs Perf?

Skills that share tags, products or a category with Perf Labs Perf: Configuring Horizon (coollabsio/coolify, 63k stars), Model Usage (openclaw/openclaw, 392k stars), Engine Whats New (flutter/flutter, 179k stars) and Openclaw Live Updater (openclaw/openclaw, 392k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Perf Labs Perf?

perf-labs (a GitHub organization) maintains it in perf-labs/perf, which has 110 GitHub stars. The repository was last updated on October 9, 2026.

Source: perf-labs/perf on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.