Agent skill

Opt Buffer Object Reuse

by Xilinx in Xilinx/mlir-air

Optimization skill — reuse NPU BufferObjects across calls instead of re-allocating/re-writing them.

MITAuto-check passed

Install Opt Buffer Object Reuse

skills CLI
$ npx skills add Xilinx/mlir-air --skill opt-buffer-object-reuse -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Xilinx/mlir-air opt-buffer-object-reuse --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Xilinx/mlir-air.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/opt-buffer-object-reuse .claude/skills/opt-buffer-object-reuse && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
opt-buffer-object-reuse
GitHub stars
150
Token cost
~1.3k tokens
SKILL.md length
574 words
Files
1
Skills in repo
15
Repo updated
First seen
Licence
MIT

At a glance

Optimization skill — reuse NPU BufferObjects across calls instead of re-allocating/re-writing them.

  • Works in 5 steps: Identify the BO slots → Pre-load weight BOs once (B1) → Reuse intermediate BOs (B2) → …
  • SKILL.md covers Purpose, Success criteria, Knowledge base references and Workflow, plus 2 more sections
  • Calls make

What it does

Opt Buffer Object Reuse is an agent skill from Xilinx/mlir-air. Optimization skill — reuse NPU BufferObjects across calls instead of re-allocating/re-writing them. Two mechanics in one class: (B1) per-layer weight BOs pre-loaded once and skipped via staticinputindices, and (B2) intermediate BOs the kernel overwrites, skipped via intermediateindices. Invoked by phase-4-prefill-optimization and phase-5-decode-optimization to cut redundant host↔NPU data movement. Decode amplifies the weight-BO win (weights reused on every token).

Its SKILL.md is about 1.3k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

The licence is MIT.

Example prompts

  • “/opt-buffer-object-reuse”

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Identify the BO slots
  2. Pre-load weight BOs once (B1)
  3. Reuse intermediate BOs (B2)
  4. prefill vs decode
  5. Validate + measure

What it can do on your machine

Read from SKILL.md and the folder at commit bca27e5. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • make

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Opt Buffer Object Reuse loads about 1.3k tokens when it runs. Until then it costs about 124 tokens; SKILL.md has 574 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~124
When it runs · the whole SKILL.md, loaded when a task matches
~1.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Xilinx/mlir-air at commit bca27e5, republished under its MIT licence (© Xilinx). 574 words, ~1,325 tokens.

Download SKILL.mdSave it as .claude/skills/opt-buffer-object-reuse/SKILL.md (or your agent's skills folder).
name
opt-buffer-object-reuse
description
Optimization skill — reuse NPU BufferObjects across calls instead of re-allocating/re-writing them. Two mechanics in one class: (B1) per-layer weight BOs pre-loaded once and skipped via static_input_indices, and (B2) intermediate BOs the kernel overwrites, skipped via intermediate_indices. Invoked by phase-4-prefill-optimization and phase-5-decode-optimization to cut redundant host↔NPU data movement. Decode amplifies the weight-BO win (weights reused on every token).

Purpose

NPU kernels re-allocate and re-upload BufferObjects (BOs) on every call by default. For an N-layer transformer that re-runs the same kernels per layer (prefill) and per token (decode), this is pure redundant host↔NPU traffic. This skill removes it. It is the same optimization the reference ablation isolated as cells A→B (weight BOs) and B→C (intermediate BOs); together they were a multi-second prefill saving and the dominant decode host-side cost.

Two mechanics, one class:

  • B1 — per-layer weight BOs (static_input_indices): allocate each layer's weight BOs once during setup, write them once, and pass static_input_indices=[<weight slots>] on every cache.load_and_run() so the runtime skips re-writing them.
  • B2 — intermediate BOs (intermediate_indices): for buffers the kernel fully overwrites (its own outputs / scratch), pass intermediate_indices=[<output slots>] so the host does not write them before the call.

Success criteria

Applying this skill is "successful" when ALL hold:

  1. Output cosine ≥ 0.99 vs the pre-optimization baseline (BO reuse must not change the math — same kernels, same inputs, only the upload is skipped). Log max_abs / max_rel informational (match the BF16 convention from phase-1-kernel-validation; do not use a tight rtol).
  2. make verify still PASSES (the end-to-end gate; token-set top-k vs HF bf16).
  3. Measured host/wall time is strictly lower than the baseline.

If (1)/(2) regress (NaN or garbage on the 2nd+ call, correct on the 1st) → the BO bookkeeping is wrong; invoke debug-bo-corruption. If (3) shows no gain → the per-call upload wasn't the bottleneck here; document and keep or revert.

Knowledge base references

  • programming_examples/llms/llama_kernel_builder/cache.py — KernelCache.load_and_run, the static_input_indices / intermediate_indices mechanics this skill drives.
  • programming_examples/llms/llama32_1b/multi_launch_builder/* — the worked example of weight + intermediate BO slots passed to a fused ELF.
  • programming_examples/llms/llama32_1b/llama32_1b_inference.py — prepare_runtime / setup where per-layer BOs are allocated once.

Workflow

Step 1: Identify the BO slots

From the kernel group's argument signature, classify each BO slot:

  • weight / LUT slots — written once, read every call → B1 candidates (static_input_indices).
  • kernel-overwritten slots (the kernel's outputs and scratch) → B2 candidates (intermediate_indices).
  • genuine per-call inputs (the activation that changes each call) — leave as normal host-written inputs.
Show full SKILL.md (246 more words)Show less
Step 2: Pre-load weight BOs once (B1)
  • Allocate per-layer weight BOs in prepare_runtime() (setup), keyed per layer (e.g. bo_key=f"kernel_L{layer_idx}"), and write the weights ONCE.
  • On every cache.load_and_run(), pass static_input_indices=[<weight slots>] so the runtime does not re-upload them.
Step 3: Reuse intermediate BOs (B2)
  • For each slot the kernel fully overwrites, pass intermediate_indices=[<output slots>] on cache.load_and_run() so the host does not write that buffer before the call.
Step 4: prefill vs decode

Same mechanic, two contexts (pass which one as the caller's parameter):

  • prefill: weights re-used across the 16 per-layer calls within one pass.
  • decode: weights re-used across every generated token × every layer — the win is much larger (16 layers × ~7 weight tensors × N tokens of upload removed). Static weight BOs are the dominant decode host-side optimization.
Step 5: Validate + measure
  • Run with BO reuse; compare output to the pre-reuse baseline → cosine ≥ 0.99.
  • Re-run make verify → must still PASS.
  • Profile host/wall time → must be strictly lower.

Failure modes

SymptomLikely causeWhere to look
Correct on 1st call, NaN/garbage on 2nd+ callper-layer BO key collision OR static_input_indices slot list wrongInvoke debug-bo-corruption
Output mismatch on the very 1st calla slot marked intermediate is actually read before being writtenRe-classify that slot as a real input (drop it from intermediate_indices)
No host-time reductionthe per-call upload wasn't the bottleneck (kernel-bound)Document; the merge skill (dispatch) or opt-layout-alignment may be the bigger win

For any failure not in the table, invoke superpowers:systematic-debugging.

Update protocol

Append to <model>/docs/development_progress/phase{4,5}_*.md:

## Buffer-object reuse
- B1 weight BOs: applied / skipped (reason)
- B2 intermediate BOs: applied / skipped (reason)
- Host/wall time before: X ms
- Host/wall time after:  Y ms
- Cosine vs baseline: <value>  | make verify: PASS/FAIL

© Xilinx, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/opt-buffer-object-reuse of Xilinx/mlir-air.

Open the folder on GitHubat commit bca27e5

Compare with similar skills

Opt Buffer Object Reuse next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Opt Buffer Object Reuse compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Opt Buffer Object Reuse this skillXilinx/mlir-air150—~1.3kAutomated safety check: PassMIT
Object Altthedaviddias/Front-End-Checklist74k—~429Automated safety check: PassMIT
Bio Variant Calling Joint CallingFreedomIntelligence/OpenClaw-Medical-Skills3.1k1 repos~2.1kAutomated safety check: PassNone
BufferLeoYeAI/openclaw-master-skills2.2k—~7.5kAutomated safety check: PassMIT
Object Storagesickn33/agentic-awesome-skills47k2 repos~2.6kAutomated safety check: PassMIT
Bio Variant Calling Structural Variant CallingFreedomIntelligence/OpenClaw-Medical-Skills3.1k—~1.7kAutomated safety check: PassNone

Similar skills

  • Object Alt

    thedaviddias/Front-End-Checklist

    A skill your agent uses when reviewing rendered HTML, interactive components, or design-system patterns related to Provide alternative text for objects.

    74k GitHub stars~429 tokensUpdated yesterday
    Frontend & DesignAuto-check passed
  • Bio Variant Calling Joint Calling

    FreedomIntelligence/OpenClaw-Medical-Skills

    Joint genotype calling across multiple samples using GATK CombineGVCFs and GenotypeGVCFs.

    3.1k GitHub starsUsed in 1 repo~2.1k tokens
    Research & ScienceAuto-check passed
  • Buffer

    LeoYeAI/openclaw-master-skills

    Buffer API integration with managed authentication. An agent skill from LeoYeAI/openclaw-master-skills.

    2.2k GitHub stars~7.5k tokensUpdated 2 mo ago
    Backend & APIsAuto-check passed
  • Object Storage

    sickn33/agentic-awesome-skills

    Configure object storage with S3, GCS, and MinIO. An agent skill from sickn33/agentic-awesome-skills.

    47k GitHub starsUsed in 2 repos~2.6k tokens
    Backend & APIsAuto-check passed
  • Bio Variant Calling Structural Variant Calling

    FreedomIntelligence/OpenClaw-Medical-Skills

    Call structural variants (SVs) from short-read sequencing using Manta, Delly, and LUMPY.

    3.1k GitHub stars~1.7k tokensUpdated 2 mo ago
    Research & ScienceAuto-check passed
  • Hard Call

    alirezarezvani/claude-skills

    /em:hard-call — Framework for decisions with no good options.

    28k GitHub stars~1.8k tokensUpdated 1 mo ago
    Auto-check passed

More from Xilinx/mlir-air

All 15 skills in this repo
  • Debug Bo Corruption

    Xilinx/mlir-air

    A skill your agent uses when an NPU kernel passes its standalone shape test but produces NaN, garbage, or stale values when invoked as part of a larger pipeline.

    150 GitHub stars~1.5k tokensUpdated today
    Auto-check passed
  • A skill your agent uses when NPU FlashAttention hangs (ERTCMDSTATETIMEOUT) or produces NaN at headdim ≥ 128.

    150 GitHub stars~1.9k tokensUpdated today
    Auto-check passed
  • A skill your agent uses when stitching kernels into a multi-launch ELF and the AIE compiler rejects the merged module (BD exhaustion, channel routing, herd shape conflict, IR validation error, DMA…

    150 GitHub stars~1.8k tokensUpdated today
    Auto-check passed
  • Deploy New LLM

    Xilinx/mlir-air

    Entry point for deploying a new decoder-only LLM on AMD NPU2.

    150 GitHub stars~4.8k tokensUpdated today
    Auto-check passed
  • Opt Layout Alignment

    Xilinx/mlir-air

    Optimization skill — choose activation layouts so consecutive kernels hand off on-device without a host-side transpose.

    150 GitHub stars~1k tokensUpdated today
    Auto-check passed
  • Procedural recipe for fusing multiple air.launch kernels into one multi-launch ELF (single XRT invocation).

    150 GitHub stars~1.8k tokensUpdated today
    Auto-check passed

Questions about Opt Buffer Object Reuse

What does Opt Buffer Object Reuse do?

Optimization skill — reuse NPU BufferObjects across calls instead of re-allocating/re-writing them. Opt Buffer Object Reuse is an agent skill from Xilinx/mlir-air. Optimization skill — reuse NPU BufferObjects across calls instead of re-allocating/re-writing them.

How do I install Opt Buffer Object Reuse in Claude Code?

Run `npx skills add Xilinx/mlir-air --skill opt-buffer-object-reuse -a claude-code`. Or copy the skill folder (.claude/skills/opt-buffer-object-reuse in Xilinx/mlir-air) into .claude/skills/opt-buffer-object-reuse in your project. Claude Code loads it when a task matches its description.

How do I install Opt Buffer Object Reuse in Codex?

Run `npx skills add Xilinx/mlir-air --skill opt-buffer-object-reuse -a codex`. Or copy the skill folder (.claude/skills/opt-buffer-object-reuse in Xilinx/mlir-air) into .agents/skills/opt-buffer-object-reuse in your project. Codex loads it when a task matches its description.

Can I use Opt Buffer Object Reuse in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Xilinx/mlir-air --skill opt-buffer-object-reuse -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/opt-buffer-object-reuse, .gemini/skills/opt-buffer-object-reuse, .github/skills/opt-buffer-object-reuse and .opencode/skills/opt-buffer-object-reuse in your project.

What does Opt Buffer Object Reuse need to run?

Going by SKILL.md and its folder, Opt Buffer Object Reuse needs the command-line tools its instructions call (make).

Does Opt Buffer Object Reuse access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Opt Buffer Object Reuse safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Opt Buffer Object Reuse use?

Opt Buffer Object Reuse is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Opt Buffer Object Reuse use?

About 1.3k tokens (SKILL.md is roughly 5.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Opt Buffer Object Reuse?

Skills that share tags, products or a category with Opt Buffer Object Reuse: Object Alt (thedaviddias/Front-End-Checklist, 74k stars), Bio Variant Calling Joint Calling (FreedomIntelligence/OpenClaw-Medical-Skills, 3.1k stars), Buffer (LeoYeAI/openclaw-master-skills, 2.2k stars) and Object Storage (sickn33/agentic-awesome-skills, 47k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Opt Buffer Object Reuse?

Xilinx (a GitHub organization) maintains it in Xilinx/mlir-air, which has 150 GitHub stars. The repository holds 15 skills in this directory. The repository was last updated on October 7, 2026.

Source: Xilinx/mlir-air on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.