Agent skill

Vllm Metax Model Upgrade

by MetaX-MACA in MetaX-MACA/vLLM-metax

Review and upgrade MetaX model support against a target vLLM revision and installed MACA components, recursively including model-dependent attention and kernels.

Apache-2.0Auto-check passedAI & LLM Engineering

Install Vllm Metax Model Upgrade

skills CLI
$ npx skills add MetaX-MACA/vLLM-metax --skill vllm-metax-model-upgrade -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install MetaX-MACA/vLLM-metax vllm-metax-model-upgrade --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/MetaX-MACA/vLLM-metax.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.codex/skills/vllm-metax-model-upgrade .claude/skills/vllm-metax-model-upgrade && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
vllm-metax-model-upgrade
GitHub stars
180
Token cost
~3.2k tokens
SKILL.md length
1,548 words
Files
2 (incl. references)
Skills in repo
6
Repo updated
First seen
Licence
Apache-2.0

At a glance

Review and upgrade MetaX model support against a target vLLM revision and installed MACA components, recursively including model-dependent attention and kernels.

  • Works in 4 steps: Inspect direct dependencies and their… → Follow their dependencies until reaching… → Update affected local dependencies and… → …
  • Model support work
  • SKILL.md covers Scope and environment, Diff the source and audit the…, First establish the… and Recursively inventory and…, plus 3 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Vllm Metax Model Upgrade is an agent skill from MetaX-MACA/vLLM-metax. Review and upgrade MetaX model support against a target vLLM revision and installed MACA components, recursively including model-dependent attention and kernels. Establish quantization and cache differences, preserve upstream structure, and distinguish shared upstream bugs from adaptation defects. Use for model support work, not standalone monkey-patch or registry audits.

Its SKILL.md is about 3.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including reference files (for example `references/review-cases.md`).

It sits in AI & LLM Engineering, covering LLM inference and serving. It works with vLLM. The repository describes itself as: Community maintained hardware plugin for vLLM on MetaX GPU. The licence is Apache-2.0.

When your agent uses it

  • Model support work
  • Not standalone monkey-patch
  • Registry audits

Example prompts

  • “/vllm-metax-model-upgrade”

Requirements

  • Python 3

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Inspect direct dependencies and their upstream changes.
  2. Follow their dependencies until reaching a verified stable interface or an installed
  3. Update affected local dependencies and callers together. Include model-private
  4. Inspect transitive consumers of shared helpers before changing their contracts.

What it can do on your machine

Read from SKILL.md and the folder at commit df0f52b. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Vllm Metax Model Upgrade loads about 3.2k tokens when it runs, and up to ~4k if it reads all its reference files. Until then it costs about 100 tokens; SKILL.md has 1,548 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~100
When it runs · the whole SKILL.md, loaded when a task matches
~3.2k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from MetaX-MACA/vLLM-metax at commit df0f52b, republished under its Apache-2.0 licence (© MetaX-MACA). 1,548 words, ~3,239 tokens.

Download SKILL.mdSave it as .claude/skills/vllm-metax-model-upgrade/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
vllm-metax-model-upgrade
description
Review and upgrade MetaX model support against a target vLLM revision and installed MACA components, recursively including model-dependent attention and kernels. Establish quantization and cache differences, preserve upstream structure, and distinguish shared upstream bugs from adaptation defects. Use for model support work, not standalone monkey-patch or registry audits.

vLLM-MetaX Model Upgrade

Preserve the target upstream structure while implementing the verified MACA execution contract. A successful upgrade includes the model's reachable dependencies, not just its top-level Python file. Upstream support declarations and comments are evidence to investigate, not proof of support on MetaX.

Scope and environment

  • Cover the requested models under vllm_metax/models/, their registration and configuration, and recursively their required attention, cache, compressor, indexer, projection, MTP/draft, MoE and kernel wrappers. Keep unrelated components out of scope.
  • Keep review-only requests read-only. Establish whether the subject is the index, worktree, installed package or specified revision. Preserve staged/unstaged changes; do not stage, commit, reinstall dependencies or edit upstream/site-packages unless authorized. For staged validation, export and import the index snapshot explicitly.
  • Read applicable repository instructions.
  • Read and apply vllm-metax-upgrade-common before compatibility decisions. It owns environment/source discovery, the shared read-only probe, one-time target confirmation and verification evidence rules. Reuse the same established environment record across upgrade skills; do not ask again for an unchanged mapping. Keep the domain-specific workflow below.
  • Apart from the common prerequisite above, do not automatically invoke patch, attention or registry skills just because the model calls those modules. Apply a separately requested specialist workflow only to its relevant portion. No monkey-patch headers or patch audit files are required for model adaptations.

Diff the source and audit the upgrade range first

  • Identify the last upstream version represented by the local models and the confirmed target revision. Before editing, diff each local model file against its exact target-upstream counterpart, including explicit mappings between platform directories. Inspect target-only, local-only, changed and removed files; a matching class name does not establish matching behavior.
  • Review every upstream commit in the version range that touches those model files. Read the patch for each commit, then inspect any other production modules changed by that commit and the model's direct callers/dependencies. Decide whether the behavior is inherited from vllm, needs a MetaX merge, needs a coordinated dependency change, is intentionally platform-specific, or belongs to a new unsupported architecture. Do not infer the answer from a commit title or from the final two-tree diff alone.
  • Keep a compact file/commit disposition record with the affected interface, dependent MetaX symbols, chosen action and evidence. Include commits that require no local edit, so a missed upstream change remains visible. Recheck the complete local-vs-target diff after edits for unaccounted differences.
  • When this skill runs before other MetaX upgrade skills, treat the confirmed target vllm source as the reliable API contract. A referenced MetaX module may still raise import or attribute errors because its own upgrade is pending. Record that boundary, continue source comparison and isolated checks, and do not misclassify the unrelated failure as a model regression. Still adapt a dependency when the requested model's changed logic requires it.

First establish the MACA/upstream differences

Before deciding what to copy or remove, build an evidence-backed difference matrix for each requested model family. Include the effective configuration after platform and speculative-decoding rewrites, not only the original HF config.

  • Inspect the official config.json for the exact checkpoint variant and revision before reasoning about its layer layout. Record backbone and MTP/draft counts separately, count layer-indexed arrays (for example layer_types, rope_theta, partial_rotary_factors), and map zero-based checkpoint layer IDs to each part. Do not substitute config-class defaults or mistake the highest layer ID for the number of layers. If the official config is unavailable, state the evidence gap.
  • Verify which config class the target vLLM actually loads and compare the raw checkpoint values with the constructed config after overrides and rewrites. Check whether constructors truncate or transform layer-indexed arrays before model and MTP code indexes them.
ContractRecord for both upstream and MACA
WeightsCheckpoint dtype, runtime dtype, per-channel/block/MX grouping, scale encoding, packing, post-load conversion and excluded modules.
Activations and projectionsCompute/output/accumulation dtype, fused vs separate projections, grouped GEMM, padding, quantization and inverse RoPE order.
KV and indexer cachesLogical dtype vs physical storage dtype, quantization scales, compressed/SWA separation, page layout/stride and index units.
AttentionSelected backend, model-specific attention, prefill/decode/MTP paths, head dimensions, sinks, causal behavior, compression and capability restrictions.
MoE and parallelismActual experts implementation, routing/scaling, shared experts, TP/PP/SP/EP/DCP collectives and supported combinations.
Optional featuresLoRA, MTP/DSpark, graph capture and each feature's real dispatch and restrictions.
ComponentsImported DeepGEMM, FlashAttention/FlashMLA, MCOPLIB and other actual callables, signatures, extension identities and verified semantics.

Separate intentional MetaX differences, inherited upstream behavior, obsolete overrides, adaptation defects and unverified assumptions. State unknowns explicitly. Inspect actual APIs, wrapper bodies and executable capability checks; challenge existing comments. A dtype name, matching signature or NVIDIA architecture predicate is not a MACA contract. Do not overwrite a necessary BF16/INT8 path with upstream FP8/FP4 behavior merely to reduce the diff. Structural alignment must preserve the established MACA semantics.

Recursively inventory and update dependencies

Trace from model registration through config rewriting, layer construction, weight loading, execution dispatch and kernels. Maintain a visited dependency inventory with: local symbol/file, target upstream counterpart, callers, relevant upstream delta, MetaX difference, action and validation evidence.

For each changed upstream model interface or behavior:

  1. Inspect direct dependencies and their upstream changes.
  2. Follow their dependencies until reaching a verified stable interface or an installed component boundary; record that boundary and why no further local change is needed.
  3. Update affected local dependencies and callers together. Include model-private attention (such as DeepSeek V4 attention/compressor/FlashMLA code) even when it lives outside the main model file or a generic attention directory.
  4. Inspect transitive consumers of shared helpers before changing their contracts. For external binary components, verify the installed API instead of silently upgrading or modifying the component.

Account for unchanged, removed, conditional and inactive files. Report blocked paths rather than claiming completion from successful imports of a subset. Read review-cases.md for relevant model-specific checks; its examples are investigation prompts, not permanent support restrictions.

Show full SKILL.md (591 more words)Show less

Keep the implementation easy to diff

  • Preserve upstream class/function boundaries, names, signatures, method order, control flow and file organization wherever MACA semantics allow it. Follow upstream moves when practical; record explicit mappings where platform directories differ.
  • Avoid unrelated refactors, reformatting, renaming, helper extraction and broad defensive scaffolding. Prefer a small change at the corresponding upstream location over a new abstraction that obscures future upgrades.
  • Every retained or newly introduced MetaX ad-hoc must have a nearby NOTE(MetaX) comment explaining the concrete upstream difference, why MACA needs it, and when it can be removed or revalidated. Preserve useful algorithm and layout explanations. Do not use a generic "MetaX modification" marker as the entire explanation.
  • Compare both the local change diff and the complete local-vs-target diff. The latter must expose necessary platform differences instead of being dominated by structural churn. Remove obsolete workarounds only after verifying the replacement path.

Example note (adapt the content to verified evidence):

python
# NOTE(MetaX): The target upstream path quantizes this projection to FP8.
# This MACA path uses BF16 because <verified component constraint>.
# Preserve <layout/scale invariant>; revalidate when <capability> is available.

Investigate upstream before fixing a suspected bug

For every potential bug found during validation, first inspect the equivalent path in the target upstream revision, including dispatch and relevant dependencies. Where feasible run the same discriminating reproduction. Classify the cause as an adaptation regression, existing MetaX defect, shared upstream defect, component issue or environment mismatch. Distinguish source-based suspicion from reproduced upstream failure.

If upstream has or may have the same issue:

  • First look for the smallest model-local fix or workaround that preserves upstream structure and the intended MACA behavior. When making such a fix, add a nearby NOTE(MetaX) explicitly saying the target upstream may also be affected, citing the symbol/revision and evidence or uncertainty. Explain the local workaround and its removal condition. Do not imply it is a MetaX-only adaptation defect.
  • If resolving that shared defect requires changing other modules/components, do not expand the fix into those dependencies for this bug. Add a targeted logger.warning (or logger.warning_once when available) and report the limitation and repair plan. Prefer configuration/construction time over repeated token-time logging. Explain the trigger, consequence and known workaround without claiming the bug is fixed.
  • Preserve existing rejection/capability checks. A warning does not establish support, justify enabling an unsupported path or replace an existing error with silent success.
  • This warning-only boundary concerns repairs to shared upstream bugs. It does not cancel the recursive dependency updates required for the requested model upgrade. Follow explicit user authorization if they separately request the cross-module fix.

Example shared-defect note:

python
# NOTE(MetaX): Target upstream <revision/symbol> also appears to bypass <contract>.
# <Evidence; state if source-inspected only>. Keep this workaround local by <action>.
# Revisit when upstream handles <condition>; do not remove on version alone.

Validate and report

Apply the common skill's isolated-validation and evidence rules to these model checks.

  • Verify registration and fresh-process construction, then weight loading, transformed parameter state, dispatch and affected numerical behavior. Test meaningful boundary cases and unsupported combinations from the difference matrix.
  • Use real installed kernels and independent numerical references where available. Test model wrappers as well as raw APIs. Isolated mocks can test constructor or routing contracts, but do not establish production reachability or GPU correctness.
  • For parallel changes, check collective ownership and residual semantics. If multi-GPU execution is unavailable, distinguish constructor/source evidence from distributed numerical validation. Likewise, a wrapper check is not an end-to-end LoRA/MTP test.
  • Record runtime origins and any process-local optional-dependency workaround. Run relevant formatting, lint, type and diff checks; disclose unavailable checks.
  • Re-review every finding after edits, including prior failures masked by earlier exceptions. Check for concurrent source changes before reporting final results.

Report the target/environment, difference matrix, dependency coverage, concrete changes or findings, upstream-bug attribution, warning-only unresolved cases, reproducible validation and untested scope. No separate audit file is mandatory. Do not call a model fully supported based only on imports, signatures or a passing kernel micro-test.

© MetaX-MACA, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (references) in .codex/skills/vllm-metax-model-upgrade of MetaX-MACA/vLLM-metax.

  • SKILL.md
  • references/review-cases.md

Open the folder on GitHubat commit df0f52b

Compare with similar skills

Vllm Metax Model Upgrade next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Vllm Metax Model Upgrade compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Vllm Metax Model Upgrade this skillMetaX-MACA/vLLM-metax180—~3.2kAutomated safety check: PassApache-2.0
SageMaker Serving Image Selectionhuggingface/skills11k1 repos~4.6kAutomated safety check: PassApache-2.0
Hugging Face Local Model Evalshuggingface/skills11k2 repos~1.6kAutomated safety check: PassApache-2.0
CI Fails Buildkiteguqiong96/Lvllm4652 repos~349Automated safety check: PassApache-2.0
Gptqmodel Tokenizer NormalizationModelCloud/GPTQModel1.3k—~1.1kAutomated safety check: PassCustom licence
Add Diffusion Modelvllm-project/vllm-omni7.1k—~7kAutomated safety check: PassApache-2.0

Similar skills

  • Official

    Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones.

    11k GitHub starsUsed in 1 repo~4.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Official

    Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.

    11k GitHub starsUsed in 2 repos~1.6k tokens
    AI & LLM EngineeringAuto-check passed
  • CI Fails Buildkite

    guqiong96/Lvllm

    Fetch and diagnose vLLM Buildkite CI failure logs. An agent skill from guqiong96/Lvllm.

    465 GitHub starsUsed in 2 repos~349 tokens
    AI & LLM EngineeringAuto-check passed
  • Diagnose and correct GPT-QModel tokenizer initialization, tokenization normalization, special-token handling, prompt rendering, and chat-template problems.

    1.3k GitHub stars~1.1k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Add Diffusion Model

    vllm-project/vllm-omni

    Add a new diffusion model (text-to-image, text-to-video, image-to-video, text-to-audio, image editing) to vLLM-Omni, including native non-Diffusers ports, reference-parity validation, Cache-DiT…

    7.1k GitHub stars~7k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Add Recipe

    vllm-project/vllm-omni

    Add or update an in-repository vLLM-Omni model recipe with verified task, input, output, hardware, command, feature, and validation contracts.

    7.1k GitHub stars~1.4k tokensUpdated today
    AI & LLM EngineeringAuto-check passed

More from MetaX-MACA/vLLM-metax

  • Vllm Metax Registry Upgrade

    MetaX-MACA/vLLM-metax

    Review and adapt vllmmetax/registry registrations, quantization configurations, CustomOps and kernel dispatch against a target vLLM revision and installed MetaX APIs.

    180 GitHub stars~1.8k tokensUpdated yesterday
    Auto-check passed
  • Vllm Metax Attention Upgrade

    MetaX-MACA/vLLM-metax

    Review and adapt MetaX attention backends, MLA, sparse indexers, cache layouts, and their kernel wrappers against a target vLLM revision and the actually installed MetaX component APIs.

    180 GitHub stars~2.3k tokensUpdated yesterday
    Auto-check passed
  • Vllm Metax Patch Upgrade

    MetaX-MACA/vLLM-metax

    Audit and adapt monkey patches in vllmmetax/patch/ against a target upstream revision.

    180 GitHub stars~2.8k tokensUpdated yesterday
    Auto-check passed
  • Vllm Metax Upgrade Common

    MetaX-MACA/vLLM-metax

    Establish the shared environment, source/runtime correspondence, target confirmation and validation evidence for MetaX vLLM upgrades.

    180 GitHub stars~1.1k tokensUpdated yesterday
    Auto-check passed
  • Vllm Metax Model Trim

    MetaX-MACA/vLLM-metax

    Trim large MetaX model directories for dummy smoke tests or real-checkpoint loading on limited GPUs.

    180 GitHub stars~2.4k tokensUpdated yesterday
    Auto-check passed

Works with

Questions about Vllm Metax Model Upgrade

What does Vllm Metax Model Upgrade do?

Review and upgrade MetaX model support against a target vLLM revision and installed MACA components, recursively including model-dependent attention and kernels. Vllm Metax Model Upgrade is an agent skill from MetaX-MACA/vLLM-metax. Review and upgrade MetaX model support against a target vLLM revision and installed MACA components, recursively including model-dependent attention and kernels.

When should I use Vllm Metax Model Upgrade?

Vllm Metax Model Upgrade fits situations like: model support work; not standalone monkey-patch; registry audits.

How do I install Vllm Metax Model Upgrade in Claude Code?

Run `npx skills add MetaX-MACA/vLLM-metax --skill vllm-metax-model-upgrade -a claude-code`. Or copy the skill folder (.codex/skills/vllm-metax-model-upgrade in MetaX-MACA/vLLM-metax) into .claude/skills/vllm-metax-model-upgrade in your project. Claude Code loads it when a task matches its description.

How do I install Vllm Metax Model Upgrade in Codex?

Run `npx skills add MetaX-MACA/vLLM-metax --skill vllm-metax-model-upgrade -a codex`. Or copy the skill folder (.codex/skills/vllm-metax-model-upgrade in MetaX-MACA/vLLM-metax) into .agents/skills/vllm-metax-model-upgrade in your project. Codex loads it when a task matches its description.

Can I use Vllm Metax Model Upgrade in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add MetaX-MACA/vLLM-metax --skill vllm-metax-model-upgrade -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/vllm-metax-model-upgrade, .gemini/skills/vllm-metax-model-upgrade, .github/skills/vllm-metax-model-upgrade and .opencode/skills/vllm-metax-model-upgrade in your project.

What does Vllm Metax Model Upgrade need to run?

SKILL.md names no scripts, command-line tools or credentials: Vllm Metax Model Upgrade is instructions for the agent only. Our summary lists: Python 3.

Does Vllm Metax Model Upgrade access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Vllm Metax Model Upgrade safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Vllm Metax Model Upgrade use?

Vllm Metax Model Upgrade is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Vllm Metax Model Upgrade use?

About 3.2k tokens (SKILL.md is roughly 13k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 797 tokens, read only when the agent opens those files.

What are the alternatives to Vllm Metax Model Upgrade?

Skills that share tags, products or a category with Vllm Metax Model Upgrade: SageMaker Serving Image Selection (huggingface/skills, 11k stars), Hugging Face Local Model Evals (huggingface/skills, 11k stars), CI Fails Buildkite (guqiong96/Lvllm, 465 stars) and Gptqmodel Tokenizer Normalization (ModelCloud/GPTQModel, 1.3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Vllm Metax Model Upgrade?

MetaX-MACA (a GitHub organization) maintains it in MetaX-MACA/vLLM-metax, which has 180 GitHub stars. The repository holds 6 skills in this directory. The repository was last updated on October 10, 2026.

Source: MetaX-MACA/vLLM-metax on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.