Agent skill

mecatl Offline Performance Optimization

by stacklok in stacklok/mecatl

Runs mecatl's offline benchmark and scenario harness to measure, profile with pprof, optimize and prove a performance win with benchstat, then adds a regression benchmark.

Apache-2.0Auto-check passedDevelopment

Install mecatl Offline Performance Optimization

skills CLI
$ npx skills add stacklok/mecatl --skill perf-optimization -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install stacklok/mecatl perf-optimization --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/stacklok/mecatl.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/perf-optimization .claude/skills/perf-optimization && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
perf-optimization
GitHub stars
218
Token cost
~2.2k tokens
SKILL.md length
963 words
Files
2 (incl. references)
Skills in repo
7
Repo updated
First seen
Licence
Apache-2.0

At a glance

Runs mecatl's offline benchmark and scenario harness to measure, profile with pprof, optimize and prove a performance win with benchstat, then adds a regression benchmark.

  • Works in 5 steps: Measure — pick the benchmark that… → Profile — pinpoint the site, do NOT guess → Design — the smallest change at the… → …
  • Investigating why a mecatl code path is slow or allocating heavily
  • SKILL.md covers The harness (where the numbers…, What to gate on (allocs-first), Workflow and Common pitfalls, plus 2 more sections
  • Calls go and git

What it does

This skill covers the offline side of mecatl's performance work: measure, profile, optimize, prove, then guard with a regression test. It runs task bench and task perf:scenarios, takes a memory profile into pprof to find the real hotspot, and uses benchstat to run an A/B comparison that proves a change actually helped rather than relying on a hypothesis. Allocation counts are checked before anything else, following an allocs-first gating discipline, and profile-guided optimization (PGO) can be wired in as part of the setup.

The discipline named in the skill includes following the profile rather than a guess, keeping pure-performance changes byte-identical to the behavior they replace, mutation-testing any cache guard that was added, and skipping an abstraction that turns out to be the wrong one rather than forcing it to fit. A playbook reference file carries the detailed steps. For diagnosing a running harness through its live performance MCP server, a separate perf-mcp-interpretation skill applies instead; this one is specific to mecatl's own benchmarks, not general Go profiling.

When your agent uses it

  • Investigating why a mecatl code path is slow or allocating heavily
  • Pinpointing a hotspot in a Go benchmark with pprof
  • Proving a performance fix with a benchstat A/B comparison
  • Adding a regression benchmark after fixing a performance issue

Example prompts

  • “Profile the scenario benchmark and find the allocation hotspot in this function.”
  • “Run benchstat to prove this change actually reduced allocs per op.”
  • “Add a regression benchmark so this slowdown doesn't come back.”

Requirements

  • A checkout of the mecatl repository with its Go benchmark and scenario harness

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Measure — pick the benchmark that actually exercises the path
  2. Profile — pinpoint the site, do NOT guess
  3. Design — the smallest change at the proven hotspot
  4. Prove — benchstat before/after, count >= 10
  5. Guard — keep behaviour identical, prove the guard isn't vacuous

What it can do on your machine

Read from SKILL.md and the folder at commit e731897. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • go
    • git

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use git, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

mecatl Offline Performance Optimization loads about 2.2k tokens when it runs, and up to ~3.5k if it reads all its reference files. Until then it costs about 197 tokens; SKILL.md has 963 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~197
When it runs · the whole SKILL.md, loaded when a task matches
~2.2k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~3.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from stacklok/mecatl at commit e731897, republished under its Apache-2.0 licence (© stacklok). 963 words, ~2,161 tokens.

Download SKILL.mdSave it as .claude/skills/perf-optimization/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
perf-optimization
description
Profile-driven performance optimization of mecatl using the offline benchmark + scenario harness. Use when optimizing allocations or latency, reducing allocs/op or memory, profiling a Go benchmark, pinpointing a hotspot with pprof, proving a win with benchstat, adding a regression benchmark, wiring profile-guided optimization (PGO), or investigating "why is this slow / allocating" or a suspected perf regression. Covers task bench, task perf:scenarios, memprofile -> pprof, benchstat A/B, allocs-first gating, PGO setup, and the discipline (follow the profile not the hypothesis; keep pure-perf changes byte-identical; mutation-test cache guards; skip the wrong abstraction). NOT for the live perf MCP server (use perf-mcp-interpretation) or non-mecatl Go profiling.
metadata.author
mecatl
metadata.tags
performance, profiling, benchmarks, pprof, benchstat, allocations

Perf optimization (mecatl offline harness)

The companion to the regression-tracking design in docs/perf-tracking.md. This skill is the offline benchmark/scenario workflow: measure → profile → optimize → prove → guard. For diagnosing a running harness via the perf MCP server, use the perf-mcp-interpretation skill instead — different tool, different signals.

The harness (where the numbers come from)

  • task bench — hot-path microbenchmarks (engine/prompt, engine/governance, engine/agent). BENCHCOUNT default 10. Offline (mockllm + memfs). Not part of task test. The engine is its own Go module (ADR 0036), so task bench/task fuzz run these as cd engine && go test … ./prompt/ ./governance/ ./agent/. An ad-hoc re-run on an engine package must do the same: cd engine && go test -bench=… ./agent/ (an explicit ./engine/agent/ path also resolves via the committed go.work, but the ./... wildcard does not cross the module boundary).
  • task perf:scenarios — five whole-loop scenarios. Four live in perf/scenarios/ (single-session-long, team-fanout, background-subagents, compaction-cycle); the fifth (tui-scrollback) lives in cmd/mecatui/ui/, and the task runs both packages — so go test ./perf/scenarios/ alone gives only four. BENCHCOUNT default 6. MECATL_PERF_JSON=.scratch/x.json writes the KPI JSON.
  • Both are deterministic and offline — never reach for a live model/network.
  • PGO/bench captures over engine/ packages depend on the active go.work (the committed workspace wires ./engine in alongside the root) — a GOWORK=off invocation resolves the engine module standalone and won't see the host repo.

What to gate on (allocs-first)

allocs/op, bytes/op, goroutine-delta, and cache-hit-rate are deterministic and portable — these are the real signal. ns/op and rss_* are machine-specific — treat as advisory shape, never the headline. A perf change is done only when allocs/op moves in benchstat; a wall-clock-only "win" on a shared machine is noise.

Workflow

1. Measure — pick the benchmark that actually exercises the path

Run the relevant benchmark and confirm it reflects the workload you care about. A benchmark whose per-iteration setup dwarfs the code under test, or that is all-miss by construction, will hide a real win. If no benchmark covers the path, add one first (match the existing bench_test.go style: for b.Loop(), results to a package-level sink, offline).

2. Profile — pinpoint the site, do NOT guess

Capture an allocation profile on the benchmark and follow it to a file:line:

sh
go test -run='^$' -bench=BenchmarkX -benchmem -memprofile=.scratch/x.mprof -count=3 ./pkg/
go tool pprof -alloc_space   -top -nodecount=25 .scratch/x.mprof   # bytes
go tool pprof -alloc_objects -top -nodecount=25 .scratch/x.mprof   # object count
go tool pprof -list=FuncName .scratch/x.mprof                      # line-level

Save profiles under .scratch/ (repo rule — never /tmp). Follow the profile to the real site. The hypothesis is often wrong (see the playbook: the TUI hotspot was the string join, not SetContent as assumed). Let -list show you the exact lines.

3. Design — the smallest change at the proven hotspot

Optimize only what the profile proves is hot. Weigh the win against the real cost: a microsecond saved once per turn is invisible next to an LLM round-trip. The wrong abstraction is worse than the allocation — if the clean seam doesn't exist or the fix adds stateful invalidation surface for a marginal gain, it is a legitimate NO-GO. Say so and skip it rather than forcing it.

4. Prove — benchstat before/after, count >= 10
sh
go test -run='^$' -bench=BenchmarkX -benchmem -count=10 ./pkg/ > .scratch/before.txt
# ... apply the change ...
go test -run='^$' -bench=BenchmarkX -benchmem -count=10 ./pkg/ > .scratch/after.txt
benchstat .scratch/before.txt .scratch/after.txt   # go install golang.org/x/perf/cmd/benchstat@latest

allocs/op / B/op must drop with a statistically significant delta. Re-run task perf:scenarios and confirm the scenario KPI moved in the expected direction. Update the baseline snapshot in docs/perf-tracking.md.

Show full SKILL.md (460 more words)Show less
5. Guard — keep behaviour identical, prove the guard isn't vacuous
  • Pure-perf changes must be byte-identical. For the TUI, task test:golden must stay green with zero golden diffs — you change HOW, never WHAT. Caveat: golden refresh runs -update first, so a plain go test ./... (no -update) against the committed goldens is what actually catches a regression; don't rely on the refresh step to catch a stale serve.
  • Any cache/oracle you add must be mutation-tested. Temporarily break it (force the stale/wrong path), confirm a test goes red, then restore (back up with cp, restore with cp — never git checkout, it wipes uncommitted work). A guard that stays green when the behaviour is broken is worse than none.
  • A perf change touching an EXPORTED engine/ symbol trips the api-compat CI gate (task api:check, ADR 0036/0037) — a failure mode a perf optimizer wouldn't expect. Keep pure-perf changes byte-identical to the engine's public surface and it never fires; if the surface legitimately changed, run task api:update and add an engine/CHANGELOG.md entry classified per engine/COMPATIBILITY.md.

Common pitfalls

  • All-miss benchmark hides the win. A streaming bench that mutates state every iteration never hits the cache you added — it's the worst-case floor, flat by design. Add a steady-state bench for the cache-hit path to show the real win.
  • Setup swamps the signal. Build fixtures outside b.Loop(); confirm with a quick profile that setup isn't the dominant allocator.
  • Chasing ns/op on a shared machine. Gate on allocs; ns is advisory.
  • Memoizing across the byte-stable prompt prefix. Any prompt-inventory cache must produce byte-identical output or it tanks the provider cache-hit rate — the exact thing perf-tracking exists to protect. Usually not worth it (see playbook).

Complementary: PGO (free compiler-level wins)

Beyond hand-optimizing a hot path, Profile-Guided Optimization lets the compiler optimize from a CPU profile (typically 2–14% CPU — but here sub-1% of wall-clock, since cost is network-dominated: a free set-and-forget win, not a latency feature). The mechanism is already wired:

  • task pgo:collect builds a PROVISIONAL offline profile under .scratch/pgo/.
  • cmd/mecated/default.pgo is the reserved slot — go build's -pgo=auto applies it automatically the moment a profile is dropped there (nothing in the build chain passes -pgo=off).
  • Do NOT commit an offline-collected profile — it trains the compiler on the mockllm path that ships in no production binary (it can't pessimize, but it wastes the one slot). Commit only a real /debug/pprof/profile capture from a running mecated under load.

Full rationale, the per-binary decision, and the production refresh + staleness process live in perf-tracking.md Phase 4 — read it before touching PGO.

See Also

  • references/playbook.md — pprof flag cookbook + two worked examples (a real win and a real NO-GO) showing the discipline end to end.
  • docs/perf-tracking.md — the KPI design, gating posture, baselines, and the full roadmap.
  • perf-mcp-interpretation skill — the live counterpart (running-harness diagnosis via the perf MCP server).

© stacklok, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (references) in .claude/skills/perf-optimization of stacklok/mecatl.

  • SKILL.md
  • references/playbook.md

Open the folder on GitHubat commit e731897

Compare with similar skills

mecatl Offline Performance Optimization next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

mecatl Offline Performance Optimization compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
mecatl Offline Performance Optimization this skillstacklok/mecatl218—~2.2kAutomated safety check: PassApache-2.0
Code Review ChecklistshareAI-lab/learn-claude-code78k5 repos~1.1kAutomated safety check: PassMIT
LLM Torch Profiler Analysissgl-project/sglang37k2 repos~6.4kAutomated safety check: PassApache-2.0
Pycrazyguitar/pysheeet8.2k—~886Automated safety check: PassMIT
Cmux Debugging Guidemanaflow-ai/cmux28k1 repos~1.1kAutomated safety check: PassCustom licence
Electron Heap Snapshot Analysiskeybase/client9.3k—~875Automated safety check: PassBSD-3-Clause

Similar skills

  • Code Review Checklist

    shareAI-lab/learn-claude-code

    Reviews code against a five-part checklist covering security, correctness, performance, maintainability and testing, and reports findings in a fixed format.

    78k GitHub starsUsed in 5 repos~1.1k tokens
    DevelopmentAuto-check passed
  • LLM Torch Profiler Analysis

    sgl-project/sglang

    Unified LLM torch-profiler triage skill for sglang, vllm, TensorRT-LLM, and TokenSpeed.

    37k GitHub starsUsed in 2 repos~6.4k tokens
    DevelopmentAuto-check passed
  • Py

    crazyguitar/pysheeet

    Comprehensive Python programming reference covering syntax, concurrency, networking, databases, ML/LLM development, and HPC.

    8.2k GitHub stars~886 tokensUpdated today
    DevelopmentAuto-check passed
  • Cmux Debugging Guide

    manaflow-ai/cmux

    Covers debug logging, the Debug menu, profiling rules and runtime pitfalls for working on the cmux macOS terminal app.

    28k GitHub starsUsed in 1 repo~1.1k tokens
    DevelopmentAuto-check passed
  • Analyzes V8, Chrome and Electron .heapsnapshot files with Node scripts to find memory leaks, detached DOM nodes and the retainer paths that keep objects alive.

    9.3k GitHub stars~875 tokensUpdated today
    DevelopmentAuto-check passed
  • Runs controlled JMH experiments on the Caffeine cache to find shared contention and hot-path waste, then reviews correctness and returns a reviewable patch.

    18k GitHub stars~2.6k tokensUpdated yesterday
    DevelopmentAuto-check: notes

More from stacklok/mecatl

  • Interviews you about provider, cost, openness and image needs, then designs the models section of a mecatl settings file with aliases, slots and router categories.

    218 GitHub stars~2.7k tokensUpdated today
    Auto-check passed
  • Mecatl Release Cutting

    stacklok/mecatl

    Cuts a tagged mecatl release by dispatching the release-PR workflow, merging the bot's pull request and verifying the tag, images, Helm chart, signed archives and Homebrew formula.

    218 GitHub stars~4k tokensUpdated today
    Auto-check passed
  • mecatl Learning Config

    stacklok/mecatl

    Designs, validates and writes the learning section of a mecatl settings file, covering mode, sensitivity, reflection budgets and validated or evaluated activation.

    218 GitHub stars~3.8k tokensUpdated today
    Auto-check passed
  • Guides reading mecatl's perf MCP data to find why a running harness is slow, leaking goroutines or growing in memory, using cheap reads before any CPU capture.

    218 GitHub stars~2.3k tokensUpdated today
    Auto-check passed
  • Rebuilds the mecak8s image into the local mecatl-dev Kind cluster and builds mecatui, so you can try in-progress mecatl changes against a real Kubernetes deployment.

    218 GitHub stars~927 tokensUpdated today
    Auto-check passed
  • Panel Review

    stacklok/mecatl

    Review completed non-trivial code across four independent axes: Spec, Standards, Test adequacy, and installed Domain specialists.

    218 GitHub stars~5.8k tokensUpdated today
    Auto-check passed

Categories

Questions about mecatl Offline Performance Optimization

What does mecatl Offline Performance Optimization do?

Runs mecatl's offline benchmark and scenario harness to measure, profile with pprof, optimize and prove a performance win with benchstat, then adds a regression benchmark. This skill covers the offline side of mecatl's performance work: measure, profile, optimize, prove, then guard with a regression test. It runs task bench and task perf:scenarios, takes a memory profile into pprof to find the real hotspot, and uses benchstat to run an A/B comparison that proves a change actually helped rather than relying on a hypothesis.

When should I use mecatl Offline Performance Optimization?

mecatl Offline Performance Optimization fits situations like: investigating why a mecatl code path is slow or allocating heavily; pinpointing a hotspot in a Go benchmark with pprof; proving a performance fix with a benchstat A/B comparison; adding a regression benchmark after fixing a performance issue.

How do I install mecatl Offline Performance Optimization in Claude Code?

Run `npx skills add stacklok/mecatl --skill perf-optimization -a claude-code`. Or copy the skill folder (.claude/skills/perf-optimization in stacklok/mecatl) into .claude/skills/perf-optimization in your project. Claude Code loads it when a task matches its description.

How do I install mecatl Offline Performance Optimization in Codex?

Run `npx skills add stacklok/mecatl --skill perf-optimization -a codex`. Or copy the skill folder (.claude/skills/perf-optimization in stacklok/mecatl) into .agents/skills/perf-optimization in your project. Codex loads it when a task matches its description.

Can I use mecatl Offline Performance Optimization in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add stacklok/mecatl --skill perf-optimization -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/perf-optimization, .gemini/skills/perf-optimization, .github/skills/perf-optimization and .opencode/skills/perf-optimization in your project.

What does mecatl Offline Performance Optimization need to run?

Going by SKILL.md and its folder, mecatl Offline Performance Optimization needs the command-line tools its instructions call (go and git). Our summary lists: A checkout of the mecatl repository with its Go benchmark and scenario harness.

Does mecatl Offline Performance Optimization access the network?

SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is mecatl Offline Performance Optimization safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does mecatl Offline Performance Optimization use?

mecatl Offline Performance Optimization is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does mecatl Offline Performance Optimization use?

About 2.2k tokens (SKILL.md is roughly 8.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.3k tokens, read only when the agent opens those files.

What are the alternatives to mecatl Offline Performance Optimization?

Skills that share tags, products or a category with mecatl Offline Performance Optimization: Code Review Checklist (shareAI-lab/learn-claude-code, 78k stars), LLM Torch Profiler Analysis (sgl-project/sglang, 37k stars), Py (crazyguitar/pysheeet, 8.2k stars) and Cmux Debugging Guide (manaflow-ai/cmux, 28k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains mecatl Offline Performance Optimization?

stacklok (a GitHub organization) maintains it in stacklok/mecatl, which has 218 GitHub stars. The repository holds 7 skills in this directory. The repository was last updated on October 6, 2026.

Source: stacklok/mecatl on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.