Agent skill

Perf Investigation

by jolars in jolars/panache

Profile-driven performance work on the panache parser or formatter.

MITAuto-check passedDevelopment

Install Perf Investigation

skills CLI
$ npx skills add jolars/panache --skill perf-investigation -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install jolars/panache perf-investigation --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/jolars/panache.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/perf-investigation .claude/skills/perf-investigation && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
perf-investigation
GitHub stars
237
Token cost
~3.3k tokens
SKILL.md length
1,406 words
Files
1
Skills in repo
10
Repo updated
First seen
Licence
MIT

At a glance

Profile-driven performance work on the panache parser or formatter.

  • Works in 6 steps: Hotspot (function + approximate %)… → Bucket (from §Classify each hotspot). → Median wall-time delta on the relevant… → …
  • Tasks that involve Linting and formatting
  • SKILL.md covers Scope boundaries, Harness — parser, Harness — formatter and Capture a perf profile, plus 7 more sections
  • Calls cargo

What it does

Perf Investigation is an agent skill from jolars/panache. Profile-driven performance work on the panache parser or formatter. Measure first with perf + the right harness; classify hotspots into one of a small set of buckets; apply the matching cheap fix; verify median wall-time moved before committing.

Its SKILL.md is about 3.3k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Development, covering Linting and formatting. It works with Pandoc. The repository describes itself as: Language server, formatter, and linter for Quarto and other Markdown flavors. The licence is MIT.

When your agent uses it

  • Tasks that involve Linting and formatting

Example prompts

  • “/perf-investigation”

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Hotspot (function + approximate %) addressed.
  2. Bucket (from §Classify each hotspot).
  3. Median wall-time delta on the relevant harness (12 runs,
  4. Test / clippy / fmt status (all green or specific exception).
  5. What was tried but reverted (with reason).
  6. Suggested next hotspot, ranked by likely shared root cause.

What it can do on your machine

Read from SKILL.md and the folder at commit dbf4d6d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • cargo

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Perf Investigation loads about 3.3k tokens when it runs. Until then it costs about 66 tokens; SKILL.md has 1,406 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~66
When it runs · the whole SKILL.md, loaded when a task matches
~3.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from jolars/panache at commit dbf4d6d, republished under its MIT licence (© jolars). 1,406 words, ~3,256 tokens.

Download SKILL.mdSave it as .claude/skills/perf-investigation/SKILL.md (or your agent's skills folder).
name
perf-investigation
description
Profile-driven performance work on the panache parser or formatter. Measure first with perf + the right harness; classify hotspots into one of a small set of buckets; apply the matching cheap fix; verify median wall-time moved before committing.

Use this skill when asked to "speed up parsing", "speed up formatting", "look at the parser/formatter hotspots", "fix the regression after <feature> landed", or anything else where the task is measure parser/formatter cost on a real input and recover wall-time.

The buckets, workflow, and verification steps are shared across parser and formatter; only the harness invocation and the hot-file map differ. The "Harness" sections below have both — pick the one that matches the target.

Scope boundaries

  • Verification is end-to-end test green + commonmark_allowlist green
    • clippy/fmt clean. Performance gains do not justify a snapshot diff or a regression in any test.
  • The thread-local pool / scratch-bundle pattern in inline_ir.rs::ScratchEvents is the established shape for amortizing per-call allocations. Don't invent a new pattern; extend that one.
  • Formatting must remain idempotent (format(format(x)) == format(x)). Any formatter perf change that touches emitter shape needs the golden-cases suite green.
  • Parsing must remain CST-lossless. Any parser perf change that touches builder.token() / builder.start_node() shape needs the parser golden snapshots and the conformance allowlist green.

Follow the parser, formatter, and integration-test invariants in the repository's root AGENTS.md.

Harness — parser

Stress doc: pandoc/MANUAL.txt (~300 KB). Small docs hide per-line dispatcher cost behind allocator noise.

CARGO_PROFILE_RELEASE_DEBUG=true cargo build --release \
    --example profile_parse -p panache-parser
for i in $(seq 1 12); do
  taskset -c 0 ./target/release/examples/profile_parse \
      pandoc/MANUAL.txt 200 2>&1 | tail -1
done

taskset -c 0 pins to one core — without it, scheduling jitter on a hybrid-core CPU swamps small wins. Discard the first 2-3 warmup runs; take median of the remaining ~9-10. Per-run variance on a warm machine is ~3-5%; demand at least that big a delta before declaring a fix worked.

Harness — formatter

The repo's formatting bench is cargo bench --bench formatting. For focused hotspot work, set PANACHE_BENCH_DOC to the doc you're investigating and a low PANACHE_BENCH_ITERATIONS:

cd benches/documents && ./download.sh && cd ../..   # first time only
PANACHE_BENCH_DOC=pandoc_manual.md PANACHE_BENCH_ITERATIONS=3 \
    cargo bench --bench formatting

# Or end-to-end on a single doc via the CLI binary, with hyperfine if
# available (more honest than ad-hoc shell loops):
CARGO_PROFILE_RELEASE_DEBUG=true cargo build --release
hyperfine --warmup 3 \
    'taskset -c 0 ./target/release/panache format \
         < pandoc/MANUAL.txt > /dev/null'

Same warmup-discard rule applies.

Capture a perf profile

perf record --call-graph=dwarf -F 999 -o /tmp/panache_perf.data -- \
    ./target/release/examples/profile_parse pandoc/MANUAL.txt 400
perf report --stdio -i /tmp/panache_perf.data \
    --no-children -g none --percent-limit 1.0 | head -40

Always read cpu_core samples, not cpu_atom — on a hybrid-core CPU cpu_atom typically captures only a handful of samples and percentages there are essentially noise. Use --no-children for the flat self-time view; use -g graph,caller,… (or ,callee,…) when you need to find who calls a hot leaf. For inline-frame visibility add --inline.

For flame graphs, the repo already integrates cargo flamegraph:

PANACHE_BENCH_DOC=pandoc_manual.md PANACHE_BENCH_ITERATIONS=3 \
    cargo flamegraph --bench formatting

Classify each hotspot

Every parser/formatter hotspot recovered so far falls into one of these buckets. Identify which one BEFORE editing:

  • Slice-pattern trim — s.trim_*_matches([' ', '\t']) or similar ASCII-set trims show up as core::str::trim_matches / trim_start_matches with MultiCharEqSearcher / CharPredicateSearcher::next_reject in the call stack. Replace with byte-level helpers from parser/utils/helpers.rs (trim_end_newlines, trim_start_spaces_tabs, trim_end_spaces_tabs, is_blank_line).
  • .trim().is_empty() on every line — Unicode whitespace iterator for what is always ASCII. Use is_blank_line(s) instead.
  • Per-line block parser invoked without leading-byte gate — try_parse_* runs on every non-blank line; allocates / scans before realizing the line can't possibly be the construct. Add a cheap byte gate: bytes after up to 3 spaces matches the expected leading byte. Examples that paid off (parser): [ for ref-def + footnote-def, < for HTML block, : for fenced-div + def-marker, =/- for setext underline (next line). Skip when the existing inner check is already byte-cheap (count_blockquote_markers already has one).
  • Per-call String allocation on a no-match path — a try_parse_* function that returns Option<(String, …)> allocates the string even when the caller's outer guard rejects it. Change the signature to return Option<usize> (or Option<&str>) and have the caller build the String only on confirmed match.
  • Per-iteration Vec::new() in a hot loop — .collect::<Vec<_>>() inside an inner loop, or fresh Vec per call to a function that's invoked per range/paragraph/line. Either pool via the scratch bundle pattern in inline_ir.rs::ScratchBundle or hoist + clear() + extend() so capacity is reused across iterations.
  • Per-call malloc for a discardable builder — GreenNodeBuilder::new() in detect_prepared to "try a parse and throw it away" allocates a fresh NodeCache each call. The right fix is splitting the parser function into a separate validate_* that doesn't emit, not pooling the discardable builder (the cache holds Arcs across parses and pooling it across the benchmark loop creates an unrealistic flatter — each iteration after the first hits a warm cache that wouldn't exist in real CLI usage).
  • char-walk where bytes would do — code-span / list-marker-like scanners stepping pos += rest[pos..].chars().next()?.len_utf8() byte-by-byte through plain ASCII. Replace with memchr-style bytes.iter().position(|&b| b == NEEDLE) (the compiler emits vectorized memchr). All Pandoc / CommonMark structural bytes are ASCII, so byte-level scans are losslessness-safe.
  • to_uppercase() / to_lowercase() on ASCII — Unicode case-folding allocates a fresh String. For ASCII-only checks (e.g. Roman numeral validation), case-fold a byte at a time via b & !0x20.
  • Formatter wrapping / line-builder churn — formatter-specific hotspots tend to live in wrapping (crates/panache-formatter/src/formatter/wrapping.rs), inline emission (inlines.rs), and table layout (tables.rs). Common shapes: String allocation per inline span, repeated width recalculation, per-line Vec<String> for column widths. Same buckets as above, just different files.
  • rowan internals (NodeCache::token, Arc::drop_slow, reserve_rehash, ThinArc::from_header_and_iter) — these dominate the residual ~12-15% on parser benchmarks. Proportional to builder.token() / builder.start_node() call count and to the size of the resulting green tree. Reducing them means emitting fewer tokens (e.g. coalescing a per-line TEXT + NEWLINE pair in raw / code blocks into one TEXT token where the formatter doesn't need the split). This is invasive — verify CST snapshots and the formatter round-trip before changing emitter shape. Don't try to pool the NodeCache across parses; it holds Arc'd green nodes (memory leak) and warming it across benchmark iterations creates a misleading result.
Show full SKILL.md (571 more words)Show less

Apply the smallest matching fix

Rules that have paid off:

  • Don't theorize before measuring. Multiple "should-have-helped" changes (pre-sizing line-split vecs via an LF count; byte-gate on BlockQuoteParser::detect_prepared) regressed wall time and were reverted. The intuition wasn't wrong; the cost model was. Always measure.
  • Verify with measurement, not perf-only. A change can drop a symbol from the perf top-25 without moving median wall time (sample relocation, not work elimination). The wall-time median is the truth.
  • One change per commit. Prevents one regression from masking the win of another. Re-run the test suite + clippy + fmt + a fresh taskset -c 0 measurement cycle for each.
  • Revert promptly. If 12 runs after the change don't show a median shift larger than the baseline noise (~3-5%), the fix doesn't pay; revert and pick a different lever. Don't ship pretty-but-flat refactors as perf.

Verify and commit

For every commit:

cargo test --workspace --no-fail-fast
cargo test -p panache-parser --test commonmark commonmark_allowlist
cargo clippy --workspace --all-targets --all-features -- -D warnings
cargo fmt -- --check

For formatter changes also verify the golden-cases suite explicitly:

cargo test --test golden_cases

Then a fresh measurement (taskset, 12 runs, median). The commit message should name the bucket and quote the median delta:

perf(parser|formatter): <bucket> on <call site>

<one-paragraph rationale: profile pointed here, what specifically>
<was wasteful, what the fix replaces it with>

Median wall time on `<harness command>` (12 runs):
~X ms → ~Y ms (~Z%).

Cite the wall-time number even when it's "in the noise" — that's the honest record, and a reviewer can decide whether to ship a noise-floor change at all.

Key files — parser

  • crates/panache-parser/examples/profile_parse.rs — the harness.
  • crates/panache-parser/src/parser/utils/helpers.rs — byte-level trim / blank-line helpers; first place to look for an existing helper before adding a new one.
  • crates/panache-parser/src/parser/inlines/inline_ir.rs — ScratchEvents / ScratchBundle thread-local pool pattern; build_full_plans for per-paragraph IR work.
  • crates/panache-parser/src/parser/block_dispatcher.rs — every block parser's detect_prepared lives here; this is the hot per-line dispatch site.
  • crates/panache-parser/src/parser/inlines/refdef_map.rs — document-wide refdef pre-pass, called once per parse.
  • pandoc/MANUAL.txt — 300 KB stress doc.

Key files — formatter

  • benches/formatting.rs — the bench harness; respects PANACHE_BENCH_DOC and PANACHE_BENCH_ITERATIONS.
  • benches/documents/ — set of stress docs (small, medium_quarto, tables, math, large_authoring, pandoc_manual.md).
  • crates/panache-formatter/src/formatter/ — split by concern (wrapping, inlines, paragraphs, headings, lists, tables, …). Match the file to the construct your hotspot involves.
  • crates/panache-formatter/src/formatter.rs — top-level orchestration.
  • tests/fixtures/cases/ — formatter goldens (UPDATE_EXPECTED=1 to refresh, but verify diffs carefully — the formatter's idempotency invariant means a wrong refresh is a silent regression).

Don't redo / known traps

  • split_lines_inclusive LF pre-count regressed parser wall time. The extra pass over the input cost more than the resize-grow it saved. Don't try this again unless you change the data structure (e.g. thread-local pooled Vec<&'static str> via lifetime transmute) — and even then, prove the win with measurement first.
  • BlockQuoteParser::detect_prepared byte-gate was a noise-level regression. count_blockquote_markers already has its own internal byte-cheap check; layering another gate on top added a tiny cost without saving meaningful work.
  • Don't pool the rowan NodeCache across parses. Holds Arc'd green nodes (LSP memory leak) and produces misleading benchmark numbers (warm cache after iter 1).
  • Don't trust cpu_atom perf samples on a hybrid-core CPU. Read cpu_core data; atom is too few samples to be reliable.
  • Don't add a too-permissive byte-gate. It has no effect. The gate must match the construct's actual first-byte set after indent; verify with the full test suite, not just by reading the parser code.
  • Don't add String::new() in a detect_prepared before the gate. ReferenceDefinitionParser was the canonical example — the multi-line String::new() + push_str ran on every line; the byte gate is what unlocked the 15% wall-time jump.

Report-back format

When done, report:

  1. Hotspot (function + approximate %) addressed.
  2. Bucket (from §Classify each hotspot).
  3. Median wall-time delta on the relevant harness (12 runs, taskset -c 0).
  4. Test / clippy / fmt status (all green or specific exception).
  5. What was tried but reverted (with reason).
  6. Suggested next hotspot, ranked by likely shared root cause.

© jolars, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/perf-investigation of jolars/panache.

Open the folder on GitHubat commit dbf4d6d

Compare with similar skills

Perf Investigation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Perf Investigation compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Perf Investigation this skilljolars/panache237—~3.3kAutomated safety check: PassMIT
Minimizing Ty Ecosystem Changesastral-sh/ruff50k—~4.6kAutomated safety check: PassMIT
Install Anti-Slop Oxlint Rulesdmmulroy/anti-slop5.4k—~2.2kAutomated safety check: PassMIT
Babysit PR To Pass CIsgl-project/sglang37k2 repos~3kAutomated safety check: PassApache-2.0
Rust Best Practicesfarm-fe/farm5.6k3 repos~1.1kAutomated safety check: PassMIT
Summarise Ecosystem Resultsastral-sh/ruff50k—~2.2kAutomated safety check: PassMIT

Similar skills

  • Official

    A skill your agent uses when a user says "minimize this ty ecosystem change", "reproduce this ecosystem result", "investigate a primer difference", "investigate a mypyprimer difference"…

    50k GitHub stars~4.6k tokensUpdated yesterday
    DevelopmentAuto-check passed
  • Installs, updates or migrates the vendored anti-slop Oxlint plugin in a repository, keeping local rule changes and the plugin's license and provenance files.

    5.4k GitHub stars~2.2k tokensUpdated 1 mo ago
    DevelopmentAuto-check passed
  • Babysit PR To Pass CI

    sgl-project/sglang

    Start and persistently pursue a goal to babysit an SGLang pull request until selected GitHub Actions workflows pass on the latest PR head.

    37k GitHub starsUsed in 2 repos~3k tokens
    DevelopmentAuto-check passed
  • Guide for writing idiomatic Rust code based on Apollo GraphQL's best practices handbook.

    5.6k GitHub starsUsed in 3 repos~1.1k tokens
    DevelopmentAuto-check passed
  • Official

    A skill your agent uses when a user says "summarise ecosystem results", "summarize this ty ecosystem report", "what changed in this ecosystem run?", or asks to summarise or summarize ty ecosystem…

    50k GitHub stars~2.2k tokensUpdated yesterday
    DevelopmentAuto-check passed
  • Go Pedantry

    chromedp/chromedp

    This skill should be used when the user is writing Go code and needs guidance on Go-specific pedantry: error wrapping with fmt.Errorf and %w, interface design (accept interfaces return structs)…

    13k GitHub stars~3.7k tokensUpdated 5 days ago
    DevelopmentAuto-check passed

More from jolars/panache

All 10 skills in this repo
  • Add Syntax Construct

    jolars/panache

    Add a new block-level or inline-level syntax construct to Panache's parser and formatter — confirm the pandoc-native shape first, add SyntaxKinds for every byte category, gate it behind an extension…

    237 GitHub stars~1.8k tokensUpdated yesterday
    Auto-check passed
  • Grow Panache's CommonMark spec conformance under Flavor::CommonMark by running every spec.txt example through the shared parser, comparing rendered HTML against the spec's expected HTML…

    237 GitHub stars~1.1k tokensUpdated yesterday
    Auto-check passed
  • External Tools

    jolars/panache

    Work on Panache's delegation of embedded code blocks to third-party formatters and linters (ruff, shfmt, shellcheck, rustfmt, ...) — add or change a preset, fix the offset mapping that translates a…

    237 GitHub stars~1.2k tokensUpdated yesterday
    Auto-check passed
  • Linter Investigation

    jolars/panache

    Investigate panache's linter (and, secondarily, its parser) against a real-world Quarto/Markdown codebase.

    237 GitHub stars~2k tokensUpdated yesterday
    Auto-check passed
  • Math Parser Formatter

    jolars/panache

    Implement or debug Panache's TeX math parser and formatter internals, including the lossless CST, semantic model, diagnostics, and Badness parity.

    237 GitHub stars~1.5k tokensUpdated yesterday
    Auto-check passed
  • Add Lint Rule

    jolars/panache

    Add a new built-in lint rule to the Panache linter — wire it into the registry, gate it on the right extension/flavor, add a regression fixture with focused assertions, and document it.

    237 GitHub stars~2.8k tokensUpdated yesterday
    Auto-check passed

Works with

Categories

Questions about Perf Investigation

What does Perf Investigation do?

Profile-driven performance work on the panache parser or formatter. Perf Investigation is an agent skill from jolars/panache. Profile-driven performance work on the panache parser or formatter.

When should I use Perf Investigation?

Perf Investigation fits situations like: tasks that involve Linting and formatting.

How do I install Perf Investigation in Claude Code?

Run `npx skills add jolars/panache --skill perf-investigation -a claude-code`. Or copy the skill folder (.agents/skills/perf-investigation in jolars/panache) into .claude/skills/perf-investigation in your project. Claude Code loads it when a task matches its description.

How do I install Perf Investigation in Codex?

Run `npx skills add jolars/panache --skill perf-investigation -a codex`. Or copy the skill folder (.agents/skills/perf-investigation in jolars/panache) into .agents/skills/perf-investigation in your project. Codex loads it when a task matches its description.

Can I use Perf Investigation in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jolars/panache --skill perf-investigation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/perf-investigation, .gemini/skills/perf-investigation, .github/skills/perf-investigation and .opencode/skills/perf-investigation in your project.

What does Perf Investigation need to run?

Going by SKILL.md and its folder, Perf Investigation needs the command-line tools its instructions call (cargo).

Does Perf Investigation access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Perf Investigation safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Perf Investigation use?

Perf Investigation is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Perf Investigation use?

About 3.3k tokens (SKILL.md is roughly 13k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Perf Investigation?

Skills that share tags, products or a category with Perf Investigation: Minimizing Ty Ecosystem Changes (astral-sh/ruff, 50k stars), Install Anti-Slop Oxlint Rules (dmmulroy/anti-slop, 5.4k stars), Babysit PR To Pass CI (sgl-project/sglang, 37k stars) and Rust Best Practices (farm-fe/farm, 5.6k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Perf Investigation?

jolars (a GitHub user) maintains it in jolars/panache, which has 237 GitHub stars. The repository holds 10 skills in this directory. The repository was last updated on October 9, 2026.

Source: jolars/panache on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.