Agent skill

Vllm Metax Patch Upgrade

by MetaX-MACA in MetaX-MACA/vLLM-metax

Audit and adapt monkey patches in vllmmetax/patch/ against a target upstream revision.

Apache-2.0Auto-check passedAI & LLM Engineering

Install Vllm Metax Patch Upgrade

skills CLI
$ npx skills add MetaX-MACA/vLLM-metax --skill vllm-metax-patch-upgrade -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install MetaX-MACA/vLLM-metax vllm-metax-patch-upgrade --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/MetaX-MACA/vLLM-metax.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.codex/skills/vllm-metax-patch-upgrade .claude/skills/vllm-metax-patch-upgrade && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
vllm-metax-patch-upgrade
GitHub stars
180
Token cost
~2.8k tokens
SKILL.md length
1,315 words
Files
3 (incl. references)
Skills in repo
6
Repo updated
First seen
Licence
Apache-2.0

At a glance

Audit and adapt monkey patches in vllmmetax/patch/ against a target upstream revision.

  • Works in 7 steps: Establish the baseline and repository… → Inventory every patch and its activation… → Compare upstream behavior and decide per… → …
  • Tasks that involve LLM inference and serving
  • SKILL.md covers Scope, 1. Establish the baseline and…, 2. Inventory every patch and… and 3. Compare upstream behavior…, plus 4 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Vllm Metax Patch Upgrade is an agent skill from MetaX-MACA/vLLM-metax. Audit and adapt monkey patches in vllmmetax/patch/ against a target upstream revision. Decide whether to retain, update, migrate, or remove each patch; maintain patch headers and explanations; validate changes and record evidence. Use only for compatibility work on this directory, not standalone attention backend, model, kernel, or other adaptations elsewhere in vllmmetax.

Its SKILL.md is about 2.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files, including reference files (for example `agents/openai.yaml` and `references/review-patterns.md`).

It sits in AI & LLM Engineering, covering LLM inference and serving. It works with vLLM. The repository describes itself as: Community maintained hardware plugin for vLLM on MetaX GPU. The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve LLM inference and serving

Example prompts

  • “/vllm-metax-patch-upgrade”

Requirements

  • Python 3

Workflow steps

7 steps, taken from the step headings in SKILL.md.

  1. Establish the baseline and repository requirements
  2. Inventory every patch and its activation path
  3. Compare upstream behavior and decide per patch
  4. Make the smallest complete adaptation
  5. Treat explanatory comments as part of the deliverable
  6. Validate according to the change
  7. Finish the audit and handoff

What it can do on your machine

Read from SKILL.md and the folder at commit 9e9b140. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Vllm Metax Patch Upgrade loads about 2.8k tokens when it runs, and up to ~4.2k if it reads all its reference files. Until then it costs about 101 tokens; SKILL.md has 1,315 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~101
When it runs · the whole SKILL.md, loaded when a task matches
~2.8k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~4.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from MetaX-MACA/vLLM-metax at commit 9e9b140, republished under its Apache-2.0 licence (© MetaX-MACA). 1,315 words, ~2,805 tokens.

Download SKILL.mdSave it as .claude/skills/vllm-metax-patch-upgrade/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
vllm-metax-patch-upgrade
description
Audit and adapt monkey patches in vllm_metax/patch/ against a target upstream revision. Decide whether to retain, update, migrate, or remove each patch; maintain patch headers and explanations; validate changes and record evidence. Use only for compatibility work on this directory, not standalone attention backend, model, kernel, or other adaptations elsewhere in vllm_metax.

vLLM-MetaX Patch Upgrade

Preserve necessary MetaX behavior while inheriting the target upstream implementation. Deliver reviewable changes, a complete patch inventory with decisions, and accurate validation results. Explain why each difference exists and when it can be removed.

Scope

Apply this skill only when the requested adaptation concerns vllm_metax/patch/. An implementation adapted from upstream is not automatically a monkey patch in this scope. Standalone work on attention backends, models, kernels, or other directories must use the normal repository workflow, without inheriting this skill's patch headers or patch audit requirements. The common environment workflow applies through the relevant upgrade skill.

For an in-scope patch, inspect upstream code, callers, registrations, and tests outside the directory as needed. Changes there must directly support that patch's adaptation, removal, migration, or validation. If a request spans both patch and non-patch work, apply this skill only to the patch portion; keep unrelated backend or model findings out of vllm_metax/patch/AUDIT.md.

1. Establish the baseline and repository requirements

  • Read applicable AGENTS.md files, vllm_metax/patch/README.md, patch templates, and existing audit records. Follow current repository requirements and user instructions.
  • Inspect working-tree and staged changes; preserve existing work. By default, do not commit, stage changes, alter installed dependencies, or modify the upstream checkout.
  • Read and apply vllm-metax-upgrade-common before compatibility decisions. It owns environment/source discovery, the shared read-only probe, one-time target confirmation and verification evidence rules. Reuse the same established environment record across upgrade skills; do not ask again for an unchanged mapping. Keep the domain-specific workflow below.

2. Inventory every patch and its activation path

Start from the vllm_metax/patch/ file list and trace plugin entry points and package initializers. Search beyond @patch: include direct attribute assignments, registry mutations, sys.modules redirects, imported aliases, and disabled or unimported patches.

For each patch, record its file, target symbol, purpose, activation status, upstream location, MetaX differences, decision, evidence, and validation approach. A single file may contain patches with different decisions. Classify templates, shared helpers, and aggregate import modules separately as infrastructure.

3. Compare upstream behavior and decide per patch

Read the original explanation, complete target implementation, relevant helpers, and callers before deciding. AST comparisons can help with signatures and function bodies. Trace dynamic exports, inherited methods, and imported aliases to their definitions; a failed static lookup does not establish that an upstream target was removed.

DecisionEvidence and action
RemoveUpstream fixes the issue, the affected path no longer exists, or an existing hook/redirect already provides the behavior. Verify callers, then remove the implementation and its activation entries.
MigrateA registry, platform hook, CustomOp, or narrower extension point now supports the requirement. Move to it, verify dispatch, and remove the redundant monkey patch.
UpdateMetaX behavior is still necessary, but upstream signatures, helpers, branches, or contracts changed. Reapply only the required differences to the current implementation.
RetainThe patch remains compatible and addresses a concrete platform, hardware, checkpoint, or product requirement. Record the reason and removal condition.
  • Age, code similarity, disabled status, or an existing PR link alone is not evidence for removal. Inspect history when needed and confirm the fix exists in the target revision.
  • Investigate why upstream deliberately removed old behavior before restoring it from a patch's copied implementation.
  • Separate hardware constraints from software bugs. Without suitable GPU or distributed evidence, do not declare shared-memory, warp, IPC ABI, or communication restrictions resolved. Record the reason for retention and the remaining validation need.
  • Verify behavioral claims in comments against code. For example, a stable sort over misplaced requests does not necessarily preserve order across an entire region.

4. Make the smallest complete adaptation

  • Prefer existing upstream registration or extension points. A small wrapper should delegate unaffected behavior to the original implementation and preserve explicit inputs.
  • When a full replacement is necessary, preserve the target upstream name, signature, return contract, decorators, and unchanged code. Mark only the required MetaX differences. Do not let an old function copy suppress new validation, model support, or resource cleanup.
  • Preserve descriptor semantics, Triton decorator order, lazy imports, and initialization timing. Check whether from ... import ... already bound the old object; successful patch import does not prove every caller uses the replacement.
  • Follow repository templates for @patch. Use allow_missing=True only when adding an intentional compatibility attribute, never to hide renamed targets or typos.
  • Inspect argument semantics rather than forwarding by name alone. For example, scale=None may select dynamic quantization while a supplied scale selects static quantization.
  • Add or update platform-specific registry entries without replacing the entire table. Compatibility defaults must preserve explicit user settings. Avoid unnecessary mutation of caller-owned dictionaries or configuration objects.
  • On removal, check references, import order, and overlapping replacements. Keep unrelated cleanup outside the task.

For quantization, tokenizer, batch ordering, registration, allocator, or kernel patches, consult review-patterns.md as needed. Its examples guide inspection; they are not fixed decisions for future upstream versions.

Show full SKILL.md (523 more words)Show less

5. Treat explanatory comments as part of the deliverable

Follow the current patch README. Every Python file in the patch directory, including initializers and shared utilities, should have the dedicated header. Preserve license and copyright notices. Templates may retain fill-in placeholders; active files must contain meaningful descriptions.

python
# -----------------------------------------------------------------------------
# Note: Describe the concrete issue, trigger, and reason for the MetaX difference.
#
# Affected versions: State evidence-backed affected versions and the reviewed revision.
#
# Remove at: Give a verifiable condition, such as upstream support or a resolved limit.
# -----------------------------------------------------------------------------
  • Do not substitute scattered Verified against, Remove when, or docstrings for these fields. Describe infrastructure honestly without inventing an upstream defect.
  • Preserve existing algorithm explanations, examples, parameter descriptions, and return semantics. Update them when implementation changes; do not delete useful explanations merely to shorten the file.
  • For complex sorting or kernel patches, explain input/output contracts, invariants, key indices and mappings, the rationale, and the actual differences from upstream.
  • Distinguish correctness requirements, platform restrictions, performance policy, and implementation choices. Preserving non-decode request order does not mean one specific request must precede another for numerical correctness.
  • Include a discriminating input/output example for boundary behavior. Do not promise unmeasured performance gains. Match the source language and write for future maintainers.

6. Validate according to the change

  • Run relevant lint, formatting, syntax, and diff checks. For comment/docstring-only changes, compare ASTs after removing docstrings to confirm executable behavior is unchanged; a full model suite is unnecessary.
  • For behavioral changes, choose the least expensive tests that expose real regressions: boundary arguments, explicit settings, nested configurations, serialization, request/state alignment during reordering, and preserved upstream behavior.
  • Run import smoke tests in fresh processes to avoid duplicate-patch failures. Verify both successful target import and that actual callers reach the replacement.
  • Recheck environment correspondence after dependency installation, checkout changes, or changes to the interpreter, working directory, or import path. In the actual test process, record package __file__/__path__, relevant extension-module paths, and the active MetaX plugin origin; compare them with the preflight plan.
  • For quantization, attention, or kernel changes, compare actual GPU output with a reference when the environment permits. Cover relevant boundary lengths, dtypes, layouts, and masks. Performance claims require benchmarks; communication claims require suitable distributed runs.
  • Reuse nearby tests. If the root conftest introduces unrelated dependencies, an isolated test directory may use --confcutdir; state which fixtures this bypasses.
  • Separate patch failures from environment ABI or optional-dependency failures. A narrowly scoped, process-local workaround may isolate unrelated paths. Do not silently modify the environment or report isolated tests as complete integration validation.
  • Record the interpreter, commands, results, temporary workarounds, and untested scope. Mocked dispatch tests do not replace real kernel numerical validation.

7. Finish the audit and handoff

Update the existing AUDIT.md or repository-designated review record. Account for every inventory entry and ensure decisions match the final code. Include:

  • The environment correspondence table: interpreter/venv, both local checkouts and commits/dirty state, both installed distributions and package locations, effective import origins, content-comparison results, and any intentional source/wheel differences.
  • Each retain/update/migrate/remove decision, concrete upstream evidence, and removal condition.
  • Important behavioral differences, reproducible validation commands, results, and untested scope.

Conclude with the completed changes, key behavior changes, validation results, and audit location. If the environment blocks validation, state what was completed and what remains unverified. Do not equate source inspection or partial success with full model validation. Once appropriate checks pass, expand testing only for new changes, failures, or unresolved concerns.

© MetaX-MACA, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (references) in .codex/skills/vllm-metax-patch-upgrade of MetaX-MACA/vLLM-metax.

  • SKILL.md
  • agents/openai.yaml
  • references/review-patterns.md

Open the folder on GitHubat commit 9e9b140

Compare with similar skills

Vllm Metax Patch Upgrade next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Vllm Metax Patch Upgrade compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Vllm Metax Patch Upgrade this skillMetaX-MACA/vLLM-metax180—~2.8kAutomated safety check: PassApache-2.0
SageMaker Serving Image Selectionhuggingface/skills11k1 repos~4.6kAutomated safety check: PassApache-2.0
Hugging Face Local Model Evalshuggingface/skills11k2 repos~1.6kAutomated safety check: PassApache-2.0
CI Fails Buildkiteguqiong96/Lvllm4652 repos~349Automated safety check: PassApache-2.0
Gptqmodel Tokenizer NormalizationModelCloud/GPTQModel1.3k—~1.1kAutomated safety check: PassCustom licence
Add Diffusion Modelvllm-project/vllm-omni7.1k—~7kAutomated safety check: PassApache-2.0

Similar skills

  • Official

    Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones.

    11k GitHub starsUsed in 1 repo~4.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Official

    Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.

    11k GitHub starsUsed in 2 repos~1.6k tokens
    AI & LLM EngineeringAuto-check passed
  • CI Fails Buildkite

    guqiong96/Lvllm

    Fetch and diagnose vLLM Buildkite CI failure logs. An agent skill from guqiong96/Lvllm.

    465 GitHub starsUsed in 2 repos~349 tokens
    AI & LLM EngineeringAuto-check passed
  • Diagnose and correct GPT-QModel tokenizer initialization, tokenization normalization, special-token handling, prompt rendering, and chat-template problems.

    1.3k GitHub stars~1.1k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Add Diffusion Model

    vllm-project/vllm-omni

    Add a new diffusion model (text-to-image, text-to-video, image-to-video, text-to-audio, image editing) to vLLM-Omni, including native non-Diffusers ports, reference-parity validation, Cache-DiT…

    7.1k GitHub stars~7k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Add Recipe

    vllm-project/vllm-omni

    Add or update an in-repository vLLM-Omni model recipe with verified task, input, output, hardware, command, feature, and validation contracts.

    7.1k GitHub stars~1.4k tokensUpdated today
    AI & LLM EngineeringAuto-check passed

More from MetaX-MACA/vLLM-metax

  • Vllm Metax Model Upgrade

    MetaX-MACA/vLLM-metax

    Review and upgrade MetaX model support against a target vLLM revision and installed MACA components, recursively including model-dependent attention and kernels.

    180 GitHub stars~3.2k tokensUpdated today
    Auto-check passed
  • Vllm Metax Registry Upgrade

    MetaX-MACA/vLLM-metax

    Review and adapt vllmmetax/registry registrations, quantization configurations, CustomOps and kernel dispatch against a target vLLM revision and installed MetaX APIs.

    180 GitHub stars~1.8k tokensUpdated today
    Auto-check passed
  • Vllm Metax Attention Upgrade

    MetaX-MACA/vLLM-metax

    Review and adapt MetaX attention backends, MLA, sparse indexers, cache layouts, and their kernel wrappers against a target vLLM revision and the actually installed MetaX component APIs.

    180 GitHub stars~2.3k tokensUpdated today
    Auto-check passed
  • Vllm Metax Upgrade Common

    MetaX-MACA/vLLM-metax

    Establish the shared environment, source/runtime correspondence, target confirmation and validation evidence for MetaX vLLM upgrades.

    180 GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Vllm Metax Model Trim

    MetaX-MACA/vLLM-metax

    Trim large MetaX model directories for dummy smoke tests or real-checkpoint loading on limited GPUs.

    180 GitHub stars~2.4k tokensUpdated today
    Auto-check passed

Works with

Questions about Vllm Metax Patch Upgrade

What does Vllm Metax Patch Upgrade do?

Audit and adapt monkey patches in vllmmetax/patch/ against a target upstream revision. Vllm Metax Patch Upgrade is an agent skill from MetaX-MACA/vLLM-metax. Audit and adapt monkey patches in vllmmetax/patch/ against a target upstream revision.

When should I use Vllm Metax Patch Upgrade?

Vllm Metax Patch Upgrade fits situations like: tasks that involve LLM inference and serving.

How do I install Vllm Metax Patch Upgrade in Claude Code?

Run `npx skills add MetaX-MACA/vLLM-metax --skill vllm-metax-patch-upgrade -a claude-code`. Or copy the skill folder (.codex/skills/vllm-metax-patch-upgrade in MetaX-MACA/vLLM-metax) into .claude/skills/vllm-metax-patch-upgrade in your project. Claude Code loads it when a task matches its description.

How do I install Vllm Metax Patch Upgrade in Codex?

Run `npx skills add MetaX-MACA/vLLM-metax --skill vllm-metax-patch-upgrade -a codex`. Or copy the skill folder (.codex/skills/vllm-metax-patch-upgrade in MetaX-MACA/vLLM-metax) into .agents/skills/vllm-metax-patch-upgrade in your project. Codex loads it when a task matches its description.

Can I use Vllm Metax Patch Upgrade in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add MetaX-MACA/vLLM-metax --skill vllm-metax-patch-upgrade -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/vllm-metax-patch-upgrade, .gemini/skills/vllm-metax-patch-upgrade, .github/skills/vllm-metax-patch-upgrade and .opencode/skills/vllm-metax-patch-upgrade in your project.

What does Vllm Metax Patch Upgrade need to run?

SKILL.md names no scripts, command-line tools or credentials: Vllm Metax Patch Upgrade is instructions for the agent only. Our summary lists: Python 3.

Does Vllm Metax Patch Upgrade access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Vllm Metax Patch Upgrade safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Vllm Metax Patch Upgrade use?

Vllm Metax Patch Upgrade is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Vllm Metax Patch Upgrade use?

About 2.8k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.4k tokens, read only when the agent opens those files.

What are the alternatives to Vllm Metax Patch Upgrade?

Skills that share tags, products or a category with Vllm Metax Patch Upgrade: SageMaker Serving Image Selection (huggingface/skills, 11k stars), Hugging Face Local Model Evals (huggingface/skills, 11k stars), CI Fails Buildkite (guqiong96/Lvllm, 465 stars) and Gptqmodel Tokenizer Normalization (ModelCloud/GPTQModel, 1.3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Vllm Metax Patch Upgrade?

MetaX-MACA (a GitHub organization) maintains it in MetaX-MACA/vLLM-metax, which has 180 GitHub stars. The repository holds 6 skills in this directory. The repository was last updated on October 10, 2026.

Source: MetaX-MACA/vLLM-metax on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.