Agent skill

Vllm Metax Attention Upgrade

by MetaX-MACA in MetaX-MACA/vLLM-metax

Review and adapt MetaX attention backends, MLA, sparse indexers, cache layouts, and their kernel wrappers against a target vLLM revision and the actually installed MetaX component APIs.

Apache-2.0Auto-check passedAI & LLM Engineering

Install Vllm Metax Attention Upgrade

skills CLI
$ npx skills add MetaX-MACA/vLLM-metax --skill vllm-metax-attention-upgrade -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install MetaX-MACA/vLLM-metax vllm-metax-attention-upgrade --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/MetaX-MACA/vLLM-metax.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.codex/skills/vllm-metax-attention-upgrade .claude/skills/vllm-metax-attention-upgrade && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
vllm-metax-attention-upgrade
GitHub stars
180
Token cost
~2.3k tokens
SKILL.md length
1,098 words
Files
4 (incl. references)
Skills in repo
6
Repo updated
First seen
Licence
Apache-2.0

At a glance

Review and adapt MetaX attention backends, MLA, sparse indexers, cache layouts, and their kernel wrappers against a target vLLM revision and the actually installed MetaX component APIs.

  • Attention adaptations outside vllmmetax/patch/
  • SKILL.md covers Scope and baseline, Trace the real component…, Review attention semantics end… and Make a complete, bounded…, plus 1 more section
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • Do not apply monkey-patch headers

What it does

Vllm Metax Attention Upgrade is an agent skill from MetaX-MACA/vLLM-metax. Review and adapt MetaX attention backends, MLA, sparse indexers, cache layouts, and their kernel wrappers against a target vLLM revision and the actually installed MetaX component APIs. Verify dispatch, supported configurations, and GPU numerical behavior. Use for attention adaptations outside vllmmetax/patch/; do not apply monkey-patch headers or patch audit requirements.

Its SKILL.md is about 2.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files, including reference files (for example `agents/openai.yaml`, `references/component-contracts.md` and `references/validation.md`).

It sits in AI & LLM Engineering, covering LLM inference and serving. It works with vLLM. The repository describes itself as: Community maintained hardware plugin for vLLM on MetaX GPU. The licence is Apache-2.0.

When your agent uses it

  • Attention adaptations outside vllmmetax/patch/
  • Do not apply monkey-patch headers
  • Patch audit requirements

Example prompts

  • “/vllm-metax-attention-upgrade”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit df0f52b. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Vllm Metax Attention Upgrade loads about 2.3k tokens when it runs, and up to ~5.6k if it reads all its reference files. Until then it costs about 101 tokens; SKILL.md has 1,098 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~101
When it runs · the whole SKILL.md, loaded when a task matches
~2.3k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~5.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from MetaX-MACA/vLLM-metax at commit df0f52b, republished under its Apache-2.0 licence (© MetaX-MACA). 1,098 words, ~2,256 tokens.

Download SKILL.mdSave it as .claude/skills/vllm-metax-attention-upgrade/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
vllm-metax-attention-upgrade
description
Review and adapt MetaX attention backends, MLA, sparse indexers, cache layouts, and their kernel wrappers against a target vLLM revision and the actually installed MetaX component APIs. Verify dispatch, supported configurations, and GPU numerical behavior. Use for attention adaptations outside vllm_metax/patch/; do not apply monkey-patch headers or patch audit requirements.

vLLM-MetaX Attention Upgrade

Preserve attention semantics while adapting to two independent contracts: the target vLLM Python implementation and the installed MetaX kernel stack. Similar names, matching signatures, successful imports, comments, and upstream tests do not establish MetaX compatibility. Validate the actual runtime implementation.

Scope and baseline

  • Cover requested attention backends under vllm_metax/v1/attention/, common MLA layers, sparse indexer CustomOps, and directly related models and wrappers. Follow callers outside these paths when needed; do not expand into unrelated changes.
  • Keep review-only requests read-only. For fixes, preserve staged and unstaged user changes. Do not stage, commit, reinstall dependencies, edit site-packages, or modify the upstream checkout unless the task authorizes that action.
  • This is independent of the patch-upgrade skill. Do not apply patch-specific headers or write attention findings into vllm_metax/patch/AUDIT.md. If a task also changes monkey patches, use the patch workflow only for that portion.
  • Read applicable repository instructions. Identify whether the requested subject is the index, working tree, installed release, or another revision. For staged reviews, export the index to a temporary directory and test that snapshot; unstaged fixes must not silently satisfy its missing dependencies.
  • Read and apply vllm-metax-upgrade-common before compatibility decisions. It owns environment/source discovery, the shared read-only probe, one-time target confirmation and verification evidence rules. Reuse the same established environment record across upgrade skills; do not ask again for an unchanged mapping. Keep the domain-specific workflow below.
  • Inventory every requested file: active registration/callers, upstream counterpart, reason for each MetaX difference, proposed action and validation. Account for retained and inactive adaptations as well as changed ones; do not let a passing subset stand in for the requested review scope.

Trace the real component contracts

Read component-contracts.md for every affected component family before deciding compatibility. Its versioned observations are prompts for revalidation, not unconditional capability rules.

Trace each path from registration/platform selection through metadata builder, layer binding, wrapper, imported callable, and installed Python/native kernel. Include native, flattened, fallback, prefill/decode, and relevant capture/distributed branches.

For each used component entry point, record:

EvidenceWhat to establish
Runtime identityDistribution version, imported file, actual callable module/source, native extension path, active plugin and relevant feature flags.
InputsPositional/keyword arguments, tensor vs tuple, dtype, shape, strides, contiguity, scales, index units, and meanings of omitted/None arguments.
OutputsTensor vs tuple, output and LSE shapes/dtypes, LSE logarithm base, padding, aliasing, in-place writes and empty-row behavior.
Supported combinationsJoint constraints on dtype, QK/V dimensions, sinks, block size, quantization, native MTP, DCP, variable lengths, and graph capture.
Evidence levelSource-inspected, signature-checked, runtime-reproduced, numerically validated, or untested.

Record the GPU model, available devices, Torch/Triton build identities and MACA runtime/ABI evidence with runtime results. NVIDIA architecture checks in upstream code are not substitutes for the executing MetaX device's capabilities.

Inspect inspect.signature, inspect.getsourcefile, extension docstrings and wrapper bodies where available. Trace *args/**kwargs, decorators, lazy symbol resolution, and aliases to the final implementation. An absent Python signature does not imply missing support. An accepted keyword does not prove its semantics are implemented.

Actively challenge comments such as "FA2 cannot do DiffKV", "all sinks supported", "FP8/FP4 share this API", "cache is contiguous", and "DCP supported". Compare claims with executable restrictions and discriminating calls. Do not rewrite platform code solely to resemble upstream, nor remove a fallback solely because upstream added a feature.

Review attention semantics end to end

  • Follow allocation/spec/layout selection through bind_kv_cache, cache insertion, gather, index conversion, kernel reads, and output merge. A logical shape does not prove a physical layout. Check layer/page/head/token strides and storage offsets.
  • Preserve actual query counts, per-request boundaries and phase when reordering. Review uniform MTP, variable query lengths, short prefills, zero-length padding and CPU/device boundary differences. A reshape requires proven uniformity; otherwise flatten or bucket using correct per-token metadata.
  • Distinguish request-relative token indices, paged slot indices, physical flat rows, and chunk-workspace offsets. Preserve original indices if consumers need different mappings. Review valid-count semantics when -1 entries have interior holes.
  • Treat dense-MHA-prefill tokens excluded from the MQA query as a valid subset. Do not run prefill work for absent tokens or return padded heads as real heads.
  • For FP8/INT8, verify storage layout, scale placement/type, query scaling and whether weights already incorporate a scale. Do not infer semantics from dtype names.
  • Check LSE layout/base, neutral empty outputs and distributed ownership before merging. DCP per-token causal lengths must be localized in the correct order relative to MTP expansion. Derive kernel head counts from gathered queries, not just local TP heads.
  • Trace graph metadata lifetime and lazy scheduler state. Captured pointers must refer to runtime-updated buffers. Warmup success is not proof that changed inputs replay correctly. Metadata reuse must respect each installed component's conditions.
Show full SKILL.md (338 more words)Show less

Make a complete, bounded adaptation

Choose between retaining a necessary MetaX difference, updating an interface/semantic mapping, using a verified fallback, or rejecting an unsupported configuration early. For a suspected component defect, reproduce it by calling that component directly in an isolated process before attributing it to the adapter.

  • Update capability declarations and dispatch together: supports_combination, dtype and layout restrictions, supports_out, graph support, and constructors must agree with reachable implementations. Check combinations, not isolated flags.
  • Keep verified fast paths. A local fallback should target the failing case, preserve masks/dtypes/outputs, avoid full-cache copies and hot-path host synchronization, and state why it exists and what evidence would permit removal. Do not promise speedups.
  • Preserve useful algorithm comments. Explain the exact MetaX/upstream difference, affected installed component evidence, index units and removal/revalidation condition. Do not impose monkey-patch headers on these standalone implementations.
  • Treat optional types and metadata contracts as part of compatibility. Use direct is not None narrowing where needed rather than hiding type errors with suppressions.

Validate and report

Use the affected cases from validation.md. Run actual kernels against independent numerical references in the intended environment. Test wrappers and real branch routing in addition to raw APIs. A mocked conversion, isolated metadata test, or successful import is not a full kernel/end-to-end test.

Use the common skill's isolated-validation and evidence rules. Keep the actual attention component under numerical test real, not mocked.

For each finding or fix, retain the trigger, expected/observed behavior, component and source origins, reproduction command, result, and verification limits. Distinguish a new adaptation regression, an existing adapter defect, and an installed component bug. Keep generated runtime scripts/tests in a clearly identified reproducible artifact; do not depend on old /tmp files being present in future sessions.

Run relevant formatting, lint, type and diff checks. If a checker is unavailable in the selected environment, say so rather than claiming it passed. End with concrete changes or review findings, evidence, and untested scope. Place any requested audit alongside the attention work or at a user-selected path, never in the patch audit by default.

© MetaX-MACA, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files (references) in .codex/skills/vllm-metax-attention-upgrade of MetaX-MACA/vLLM-metax.

  • SKILL.md
  • agents/openai.yaml
  • references/component-contracts.md
  • references/validation.md

Open the folder on GitHubat commit df0f52b

Compare with similar skills

Vllm Metax Attention Upgrade next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Vllm Metax Attention Upgrade compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Vllm Metax Attention Upgrade this skillMetaX-MACA/vLLM-metax180—~2.3kAutomated safety check: PassApache-2.0
SageMaker Serving Image Selectionhuggingface/skills11k1 repos~4.6kAutomated safety check: PassApache-2.0
Hugging Face Local Model Evalshuggingface/skills11k2 repos~1.6kAutomated safety check: PassApache-2.0
CI Fails Buildkiteguqiong96/Lvllm4652 repos~349Automated safety check: PassApache-2.0
Gptqmodel Tokenizer NormalizationModelCloud/GPTQModel1.3k—~1.1kAutomated safety check: PassCustom licence
Add Diffusion Modelvllm-project/vllm-omni7.1k—~7kAutomated safety check: PassApache-2.0

Similar skills

  • Official

    Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones.

    11k GitHub starsUsed in 1 repo~4.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Official

    Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.

    11k GitHub starsUsed in 2 repos~1.6k tokens
    AI & LLM EngineeringAuto-check passed
  • CI Fails Buildkite

    guqiong96/Lvllm

    Fetch and diagnose vLLM Buildkite CI failure logs. An agent skill from guqiong96/Lvllm.

    465 GitHub starsUsed in 2 repos~349 tokens
    AI & LLM EngineeringAuto-check passed
  • Diagnose and correct GPT-QModel tokenizer initialization, tokenization normalization, special-token handling, prompt rendering, and chat-template problems.

    1.3k GitHub stars~1.1k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Add Diffusion Model

    vllm-project/vllm-omni

    Add a new diffusion model (text-to-image, text-to-video, image-to-video, text-to-audio, image editing) to vLLM-Omni, including native non-Diffusers ports, reference-parity validation, Cache-DiT…

    7.1k GitHub stars~7k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Add Recipe

    vllm-project/vllm-omni

    Add or update an in-repository vLLM-Omni model recipe with verified task, input, output, hardware, command, feature, and validation contracts.

    7.1k GitHub stars~1.4k tokensUpdated today
    AI & LLM EngineeringAuto-check passed

More from MetaX-MACA/vLLM-metax

  • Vllm Metax Model Upgrade

    MetaX-MACA/vLLM-metax

    Review and upgrade MetaX model support against a target vLLM revision and installed MACA components, recursively including model-dependent attention and kernels.

    180 GitHub stars~3.2k tokensUpdated yesterday
    Auto-check passed
  • Vllm Metax Registry Upgrade

    MetaX-MACA/vLLM-metax

    Review and adapt vllmmetax/registry registrations, quantization configurations, CustomOps and kernel dispatch against a target vLLM revision and installed MetaX APIs.

    180 GitHub stars~1.8k tokensUpdated yesterday
    Auto-check passed
  • Vllm Metax Patch Upgrade

    MetaX-MACA/vLLM-metax

    Audit and adapt monkey patches in vllmmetax/patch/ against a target upstream revision.

    180 GitHub stars~2.8k tokensUpdated yesterday
    Auto-check passed
  • Vllm Metax Upgrade Common

    MetaX-MACA/vLLM-metax

    Establish the shared environment, source/runtime correspondence, target confirmation and validation evidence for MetaX vLLM upgrades.

    180 GitHub stars~1.1k tokensUpdated yesterday
    Auto-check passed
  • Vllm Metax Model Trim

    MetaX-MACA/vLLM-metax

    Trim large MetaX model directories for dummy smoke tests or real-checkpoint loading on limited GPUs.

    180 GitHub stars~2.4k tokensUpdated yesterday
    Auto-check passed

Works with

Questions about Vllm Metax Attention Upgrade

What does Vllm Metax Attention Upgrade do?

Review and adapt MetaX attention backends, MLA, sparse indexers, cache layouts, and their kernel wrappers against a target vLLM revision and the actually installed MetaX component APIs. Vllm Metax Attention Upgrade is an agent skill from MetaX-MACA/vLLM-metax. Review and adapt MetaX attention backends, MLA, sparse indexers, cache layouts, and their kernel wrappers against a target vLLM revision and the actually installed MetaX component APIs.

When should I use Vllm Metax Attention Upgrade?

Vllm Metax Attention Upgrade fits situations like: attention adaptations outside vllmmetax/patch/; do not apply monkey-patch headers; patch audit requirements.

How do I install Vllm Metax Attention Upgrade in Claude Code?

Run `npx skills add MetaX-MACA/vLLM-metax --skill vllm-metax-attention-upgrade -a claude-code`. Or copy the skill folder (.codex/skills/vllm-metax-attention-upgrade in MetaX-MACA/vLLM-metax) into .claude/skills/vllm-metax-attention-upgrade in your project. Claude Code loads it when a task matches its description.

How do I install Vllm Metax Attention Upgrade in Codex?

Run `npx skills add MetaX-MACA/vLLM-metax --skill vllm-metax-attention-upgrade -a codex`. Or copy the skill folder (.codex/skills/vllm-metax-attention-upgrade in MetaX-MACA/vLLM-metax) into .agents/skills/vllm-metax-attention-upgrade in your project. Codex loads it when a task matches its description.

Can I use Vllm Metax Attention Upgrade in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add MetaX-MACA/vLLM-metax --skill vllm-metax-attention-upgrade -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/vllm-metax-attention-upgrade, .gemini/skills/vllm-metax-attention-upgrade, .github/skills/vllm-metax-attention-upgrade and .opencode/skills/vllm-metax-attention-upgrade in your project.

What does Vllm Metax Attention Upgrade need to run?

SKILL.md names no scripts, command-line tools or credentials: Vllm Metax Attention Upgrade is instructions for the agent only. Our summary lists: Python 3.

Does Vllm Metax Attention Upgrade access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Vllm Metax Attention Upgrade safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Vllm Metax Attention Upgrade use?

Vllm Metax Attention Upgrade is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Vllm Metax Attention Upgrade use?

About 2.3k tokens (SKILL.md is roughly 9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3.4k tokens, read only when the agent opens those files.

What are the alternatives to Vllm Metax Attention Upgrade?

Skills that share tags, products or a category with Vllm Metax Attention Upgrade: SageMaker Serving Image Selection (huggingface/skills, 11k stars), Hugging Face Local Model Evals (huggingface/skills, 11k stars), CI Fails Buildkite (guqiong96/Lvllm, 465 stars) and Gptqmodel Tokenizer Normalization (ModelCloud/GPTQModel, 1.3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Vllm Metax Attention Upgrade?

MetaX-MACA (a GitHub organization) maintains it in MetaX-MACA/vLLM-metax, which has 180 GitHub stars. The repository holds 6 skills in this directory. The repository was last updated on October 10, 2026.

Source: MetaX-MACA/vLLM-metax on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.