Agent skill

Vllm Omni Npu Model Runner Upgrade

by vllm-project in vllm-project/vllm-omni

Upgrade vllm-omni NPU model runners (OmniNPUModelRunner, NPUARModelRunner, NPUGenerationModelRunner) to align with the latest vllm-ascend NPUModelRunner while preserving omni-specific logic.

Apache-2.0Auto-check passedAI & LLM Engineering

Install Vllm Omni Npu Model Runner Upgrade

skills CLI
$ npx skills add vllm-project/vllm-omni --skill vllm-omni-npu-model-runner-upgrade -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install vllm-project/vllm-omni vllm-omni-npu-model-runner-upgrade --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/vllm-project/vllm-omni.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/vllm-omni-npu-upgrade .claude/skills/vllm-omni-npu-model-runner-upgrade && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
vllm-omni-npu-model-runner-upgrade
GitHub stars
7.1k
Token cost
~3k tokens
SKILL.md length
871 words
Files
4 (incl. references)
Skills in repo
20
Repo updated
First seen
Licence
Apache-2.0

At a glance

Upgrade vllm-omni NPU model runners (OmniNPUModelRunner, NPUARModelRunner, NPUGenerationModelRunner) to align with the latest vllm-ascend NPUModelRunner while preserving omni-specific logic.

  • Works in 12 steps: Preparation → Analyze Omni-Specific Logic → Update Base Class (OmniNPUModelRunner) → …
  • Tasks that involve LLM inference and serving
  • SKILL.md covers Overview, File Structure, Inheritance Hierarchy and Omni-Specific Comment Markers, plus 5 more sections
  • Calls python and git

What it does

Vllm Omni Npu Model Runner Upgrade is an agent skill from vllm-project/vllm-omni. Upgrade vllm-omni NPU model runners (OmniNPUModelRunner, NPUARModelRunner, NPUGenerationModelRunner) to align with the latest vllm-ascend NPUModelRunner while preserving omni-specific logic.

Its SKILL.md is about 3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files, including reference files (for example `references/gpu-to-npu-translation.md`, `references/omni-specific-blocks.md` and `references/workflow-checklist.md`).

It sits in AI & LLM Engineering, covering LLM inference and serving. It works with vLLM. The repository describes itself as: A framework for efficient model inference with omni-modality models. The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve LLM inference and serving

Example prompts

  • “/vllm-omni-npu-model-runner-upgrade”

Requirements

  • Python 3

Workflow steps

12 steps, taken from the step headings in SKILL.md.

  1. Preparation
  2. Analyze Omni-Specific Logic
  3. Update Base Class (OmniNPUModelRunner)
  4. Update AR Model Runner
  5. Update Generation Model Runner
  6. Update Imports
  7. Sync GPU-Side Omni Changes
  8. Validation
  9. Forward Context Differences
  10. Graph Wrapper Differences
  11. Buffer Creation
  12. Attention Metadata

What it can do on your machine

Read from SKILL.md and the folder at commit 096988d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python
    • git

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use git, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Vllm Omni Npu Model Runner Upgrade loads about 3k tokens when it runs, and up to ~11k if it reads all its reference files. Until then it costs about 56 tokens; SKILL.md has 871 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~56
When it runs · the whole SKILL.md, loaded when a task matches
~3k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~11k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from vllm-project/vllm-omni at commit 096988d, republished under its Apache-2.0 licence (© vllm-project). 871 words, ~2,993 tokens.

Download SKILL.mdSave it as .claude/skills/vllm-omni-npu-model-runner-upgrade/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
vllm-omni-npu-model-runner-upgrade
description
Upgrade vllm-omni NPU model runners (OmniNPUModelRunner, NPUARModelRunner, NPUGenerationModelRunner) to align with the latest vllm-ascend NPUModelRunner while preserving omni-specific logic.

vLLM-Omni NPU Model Runner Upgrade Skill

Overview

This skill guides the process of upgrading vllm-omni's NPU model runners to align with the latest vllm-ascend codebase while preserving omni-specific enhancements. The NPU runners are designed to run omni multimodal models (like Qwen3-Omni, Bagel, MiMoAudio) on Ascend NPUs.

File Structure

NPU Model Runner Files
vllm-omni/vllm_omni/platforms/npu/worker/
├── __init__.py
├── npu_model_runner.py           # OmniNPUModelRunner (base class)
├── npu_ar_model_runner.py        # NPUARModelRunner (autoregressive)
├── npu_ar_worker.py              # AR worker
├── npu_generation_model_runner.py # NPUGenerationModelRunner (diffusion/non-AR)
└── npu_generation_worker.py      # Generation worker
GPU Reference Files (for omni-specific logic sync)
vllm-omni/vllm_omni/worker/
├── __init__.py
├── gpu_model_runner.py           # OmniGPUModelRunner
├── gpu_ar_model_runner.py        # GPUARModelRunner
├── gpu_ar_worker.py
├── gpu_generation_model_runner.py
├── gpu_generation_worker.py
├── mixins.py
├── base.py
└── gpu_memory_utils.py
vllm-ascend Reference Files
vllm-ascend/vllm_ascend/worker/
├── model_runner_v1.py            # NPUModelRunner (base class to copy from)
├── npu_input_batch.py
├── block_table.py
├── pcp_utils.py
└── worker.py

Inheritance Hierarchy

                    GPUModelRunner (vllm)
                         |
        +----------------+----------------+
        |                                 |
  OmniGPUModelRunner              NPUModelRunner (vllm-ascend)
  (vllm_omni/worker)              (vllm_ascend/worker)
        |                                 |
        +----------- OmniNPUModelRunner --+
                     (multiple inheritance)
                            |
            +---------------+---------------+
            |                               |
    NPUARModelRunner            NPUGenerationModelRunner
    (autoregressive)            (non-autoregressive/diffusion)

Omni-Specific Comment Markers

Omni-specific logic is marked with comment blocks:

python
# -------------------------------------- Omni-new -------------------------------------------------
# ... omni-specific code ...
# -------------------------------------- Omni-new -------------------------------------------------

Or simpler variations:

python
#  -------------------------------------- Omni-new -------------------------------------------------
#  ------------------------------------------------------------------------------------------------

Important:

  • Always preserve and add these markers when modifying code.
  • The reference documents (references/omni-specific-blocks.md) may not be up-to-date. Always grep for Omni-new in the GPU implementations to find the authoritative list of omni-specific blocks.
  • When you discover new omni-specific code that is not documented in the references, please update the reference files.

Key Methods Requiring Attention

OmniNPUModelRunner (npu_model_runner.py)
MethodDescriptionOmni-Specific Logic
load_modelLoad model and initialize talker_mtpUses ACLGraphWrapper instead of CUDAGraphWrapper, initializes talker buffers
_dummy_runWarmup/profiling runtalker_mtp dummy forward, extract_multimodal_outputs
_model_forwardForward pass wrapperInjects model_kwargs_extra, wraps with OmniOutput, NPU-specific graph updates
_talker_mtp_forwardTalker MTP forward for Qwen3-OmniUses set_ascend_forward_context
NPUARModelRunner (npu_ar_model_runner.py)
MethodDescriptionOmni-Specific Logic
__init__Initialize with KV transfer managerOmniKVTransferManager setup
execute_modelMain inference entryKV transfer handling, _update_states override, extract_multimodal_outputs
sample_tokensToken samplingHidden states extraction, multimodal outputs processing, OmniModelRunnerOutput
_resolve_global_request_idRequest ID resolutionFor disaggregated inference
NPUGenerationModelRunner (npu_generation_model_runner.py)
MethodDescriptionOmni-Specific Logic
_update_request_statesUpdate request states for async chunkasync_chunk handling
execute_modelGeneration forwardasync_chunk, seq_token_counts, _run_generation_model
sample_tokensOutput processingmultimodal output packaging to OmniModelRunnerOutput
_dummy_runDummy run overridemodel_kwargs initialization, multimodal extraction
_run_generation_modelRun generation modelCalls _model_forward with sampler

Upgrade Workflow

Step 1: Preparation
  1. Identify target versions(Use gh cli to check):

    • We're using vllm-omni main branch
    • Check the last release of vllm-omni
    • Target vllm-ascend version(Just directly use the local latest vllm-ascend code)
  2. Check GPU-side changes (since last release):

    bash
    cd /root/vllm-workspace/vllm-omni
    git log --oneline --since="<last-release-date>" -- vllm_omni/worker/
  3. Read latest vllm-ascend code:

    • We don't track vllm-ascend changes - just directly use the latest code from /root/vllm-workspace/vllm-ascend/vllm_ascend/worker/model_runner_v1.py
    • Copy the relevant methods and re-insert omni-specific blocks
Step 2: Analyze Omni-Specific Logic

For each NPU model runner file:

  1. Extract existing omni-specific blocks:

    bash
    grep -n "Omni-new" vllm_omni/platforms/npu/worker/npu_model_runner.py
  2. Document each omni block:

    • Which method it belongs to
    • What functionality it provides
    • Dependencies on other omni code
Step 3: Update Base Class (OmniNPUModelRunner)

Note: Always check the GPU implementation gpu_model_runner.py for any new omni logic not yet documented in references.

  1. Read the latest vllm-ascend NPUModelRunner.load_model

  2. Copy the method, keeping the structure

  3. Re-insert omni-specific logic (check GPU gpu_model_runner.py for authoritative list):

    • Replace CUDAGraphWrapper with ACLGraphWrapper
    • Keep talker_mtp initialization
    • Preserve buffer allocations for talker
    • Check for any new omni blocks added since last sync
  4. Update _dummy_run:

    • Copy from vllm-ascend
    • Compare with GPU _dummy_run for omni-specific blocks
    • Re-insert all Omni-new marked code from GPU version
  5. Update _model_forward:

    • Keep the omni wrapper logic
    • Update NPU-specific parts (graph params, SP all-gather)
    • Check GPU version for any new omni logic
Step 4: Update AR Model Runner
  1. Compare with GPU gpu_ar_model_runner.py for any new omni features

  2. Copy execute_model from vllm-ascend

  3. Re-insert omni blocks (reference references/omni-specific-blocks.md, but note it may be incomplete):

    • IMPORTANT: Always check the GPU implementation gpu_ar_model_runner.py for all Omni-new marked code blocks
    • The reference doc may not include newly added omni logic - treat it as a starting point, not exhaustive
    • When discovering new omni code blocks, please update references/omni-specific-blocks.md
    • Common omni blocks include but are not limited to: KV transfer, multimodal outputs, sampling_metadata handling, etc.
  4. Update sample_tokens (also compare with GPU implementation):

    • Compare with gpu_ar_model_runner.py's sample_tokens method
    • Identify all Omni-new marked code blocks
    • Ensure NPU version includes all omni-specific logic
Show full SKILL.md (311 more words)Show less
Step 5: Update Generation Model Runner

Note: Generation model runner may have unique omni logic for diffusion/non-AR models.

  1. Compare with GPU gpu_generation_model_runner.py - grep for all Omni-new blocks

  2. Update execute_model:

    • Check GPU version for all omni-specific blocks
    • Keep async_chunk handling
    • Keep seq_token_counts injection
    • Update forward/context setup from vllm-ascend
    • Look for any new omni logic not documented in references
  3. Update _dummy_run:

    • Copy from vllm-ascend base
    • Compare with GPU _dummy_run if exists
    • Re-insert all omni-specific logic
Step 6: Update Imports

Check and update imports at the top of each file:

python
# Common vllm-ascend imports
from vllm_ascend.ascend_forward_context import get_forward_context, set_ascend_forward_context
from vllm_ascend.attention.attention_v1 import AscendAttentionState
from vllm_ascend.attention.utils import using_paged_attention
from vllm_ascend.compilation.acl_graph import ACLGraphWrapper, update_full_graph_params
from vllm_ascend.ops.rotary_embedding import update_cos_sin
from vllm_ascend.utils import enable_sp, lmhead_tp_enable
from vllm_ascend.worker.model_runner_v1 import SEQ_LEN_WITH_MAX_PA_WORKSPACE, NPUModelRunner

# Omni-specific imports
from vllm_omni.model_executor.models.output_templates import OmniOutput
from vllm_omni.worker.gpu_model_runner import OmniGPUModelRunner
from vllm_omni.outputs import OmniModelRunnerOutput
from vllm_omni.distributed.omni_connectors.kv_transfer_manager import OmniKVTransferManager
Step 7: Sync GPU-Side Omni Changes
  1. Check recent GPU worker changes:

    bash
    git diff <from-tag>..<to-tag> -- vllm_omni/worker/gpu_model_runner.py
    git diff <from-tag>..<to-tag> -- vllm_omni/worker/gpu_ar_model_runner.py
  2. Identify new omni features that need to be ported to NPU

  3. Apply corresponding changes to NPU runners

Step 8: Validation
  1. Run type checking:

    bash
    cd /root/vllm-workspace/vllm-omni
    python -m py_compile vllm_omni/platforms/npu/worker/npu_model_runner.py
    python -m py_compile vllm_omni/platforms/npu/worker/npu_ar_model_runner.py
    python -m py_compile vllm_omni/platforms/npu/worker/npu_generation_model_runner.py
  2. Run import test:

    bash
    python -c "from vllm_omni.platforms.npu.worker import *"
  3. Run model serving test (if hardware available):

    bash
    vllm serve <model-path> --trust-remote-code

Common Pitfalls

1. Forward Context Differences
  • GPU uses set_forward_context
  • NPU uses set_ascend_forward_context
  • Parameters may differ slightly
2. Graph Wrapper Differences
  • GPU: CUDAGraphWrapper
  • NPU: ACLGraphWrapper
  • Constructor parameters may differ
3. Buffer Creation
  • GPU: _make_buffer returns different structure
  • NPU: May need numpy=True/False parameter
4. Attention Metadata
  • GPU: Uses vllm attention metadata builders
  • NPU: Uses AscendCommonAttentionMetadata
5. Sampling
  • GPU: Uses vllm sampler
  • NPU: Uses AscendSampler

Checklist Before Commit

  • All omni-specific comment markers preserved
  • New omni logic from GPU side synced
  • Imports updated to latest vllm-ascend
  • No CUDAGraphWrapper references in NPU code
  • set_ascend_forward_context used instead of set_forward_context
  • ACLGraphWrapper used for talker_mtp wrapping
  • Type hints match vllm-ascend signatures
  • No duplicate code blocks
  • Python syntax valid (py_compile passes)

Reference Files for Comparison

When upgrading, keep these files open for reference:

  1. vllm-ascend NPUModelRunner: /root/vllm-workspace/vllm-ascend/vllm_ascend/worker/model_runner_v1.py
  2. vllm GPUModelRunner: /root/vllm-workspace/vllm/vllm/v1/worker/gpu_model_runner.py
  3. vllm-omni OmniGPUModelRunner: /root/vllm-workspace/vllm-omni/vllm_omni/worker/gpu_model_runner.py

© vllm-project, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files (references) in .claude/skills/vllm-omni-npu-upgrade of vllm-project/vllm-omni.

  • SKILL.md
  • references/gpu-to-npu-translation.md
  • references/omni-specific-blocks.md
  • references/workflow-checklist.md

Open the folder on GitHubat commit 096988d

Compare with similar skills

Vllm Omni Npu Model Runner Upgrade next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Vllm Omni Npu Model Runner Upgrade compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Vllm Omni Npu Model Runner Upgrade this skillvllm-project/vllm-omni7.1k—~3kAutomated safety check: PassApache-2.0
SageMaker Serving Image Selectionhuggingface/skills11k1 repos~4.6kAutomated safety check: PassApache-2.0
Hugging Face Local Model Evalshuggingface/skills11k2 repos~1.6kAutomated safety check: PassApache-2.0
CI Fails Buildkiteguqiong96/Lvllm4652 repos~349Automated safety check: PassApache-2.0
Gptqmodel Tokenizer NormalizationModelCloud/GPTQModel1.3k—~1.1kAutomated safety check: PassCustom licence
Vllm Metax Model UpgradeMetaX-MACA/vLLM-metax180—~3.2kAutomated safety check: PassApache-2.0

Similar skills

  • Official

    Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones.

    11k GitHub starsUsed in 1 repo~4.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Official

    Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.

    11k GitHub starsUsed in 2 repos~1.6k tokens
    AI & LLM EngineeringAuto-check passed
  • CI Fails Buildkite

    guqiong96/Lvllm

    Fetch and diagnose vLLM Buildkite CI failure logs. An agent skill from guqiong96/Lvllm.

    465 GitHub starsUsed in 2 repos~349 tokens
    AI & LLM EngineeringAuto-check passed
  • Diagnose and correct GPT-QModel tokenizer initialization, tokenization normalization, special-token handling, prompt rendering, and chat-template problems.

    1.3k GitHub stars~1.1k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Vllm Metax Model Upgrade

    MetaX-MACA/vLLM-metax

    Review and upgrade MetaX model support against a target vLLM revision and installed MACA components, recursively including model-dependent attention and kernels.

    180 GitHub stars~3.2k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Triton Kernel Writing

    guqiong96/Lvllm

    Write or review Triton kernels for vLLM, with practical guidance for generated-code inspection, launch grids, indexing, specialization, tuning, and representative performance validation.

    465 GitHub starsUsed in 1 repo~831 tokens
    AI & LLM EngineeringAuto-check passed

More from vllm-project/vllm-omni

All 20 skills in this repo
  • Diffusion Perf Opt

    vllm-project/vllm-omni

    Diagnose and optimize vLLM Omni diffusion workloads, especially Wan/Qwen/Flux-style image and video generation.

    7.1k GitHub stars~7.5k tokensUpdated today
    Auto-check passed
  • Precheck PR

    vllm-project/vllm-omni

    Self-check your branch before creating a PR — catch dead code, prevent new model-specific Python examples, verify accuracy/perf claims, validate PR title format, and confirm merge readiness.

    7.1k GitHub stars~1.4k tokensUpdated today
    Auto-check passed
  • Quantization

    vllm-project/vllm-omni

    Work on vLLM-Omni quantization for diffusion, autoregressive, omni, or multi-stage models.

    7.1k GitHub stars~1.4k tokensUpdated today
    Auto-check passed
  • Review PR

    vllm-project/vllm-omni

    Review pull requests and local branches for vllm-project/vllm-omni with a frozen snapshot, module-design ownership, feature-design overlays, targeted validation, and concise evidence-backed findings.

    7.1k GitHub stars~3.8k tokensUpdated today
    Auto-check passed
  • H3 Prompt Writing

    vllm-project/vllm-omni

    Write MiniMax H3 video generation prompts for T2VA, I2VA, FL2VA, L2VA, and Ref2VA.

    7.1k GitHub starsUsed in 6 repos~744 tokens
    Auto-check passed
  • Add Diffusion Model

    vllm-project/vllm-omni

    Add a new diffusion model (text-to-image, text-to-video, image-to-video, text-to-audio, image editing) to vLLM-Omni, including native non-Diffusers ports, reference-parity validation, Cache-DiT…

    7.1k GitHub stars~7k tokensUpdated today
    Auto-check passed

Works with

Questions about Vllm Omni Npu Model Runner Upgrade

What does Vllm Omni Npu Model Runner Upgrade do?

Upgrade vllm-omni NPU model runners (OmniNPUModelRunner, NPUARModelRunner, NPUGenerationModelRunner) to align with the latest vllm-ascend NPUModelRunner while preserving omni-specific logic. Vllm Omni Npu Model Runner Upgrade is an agent skill from vllm-project/vllm-omni. Upgrade vllm-omni NPU model runners (OmniNPUModelRunner, NPUARModelRunner, NPUGenerationModelRunner) to align with the latest vllm-ascend NPUModelRunner while preserving omni-specific logic.

When should I use Vllm Omni Npu Model Runner Upgrade?

Vllm Omni Npu Model Runner Upgrade fits situations like: tasks that involve LLM inference and serving.

How do I install Vllm Omni Npu Model Runner Upgrade in Claude Code?

Run `npx skills add vllm-project/vllm-omni --skill vllm-omni-npu-model-runner-upgrade -a claude-code`. Or copy the skill folder (.claude/skills/vllm-omni-npu-upgrade in vllm-project/vllm-omni) into .claude/skills/vllm-omni-npu-model-runner-upgrade in your project. Claude Code loads it when a task matches its description.

How do I install Vllm Omni Npu Model Runner Upgrade in Codex?

Run `npx skills add vllm-project/vllm-omni --skill vllm-omni-npu-model-runner-upgrade -a codex`. Or copy the skill folder (.claude/skills/vllm-omni-npu-upgrade in vllm-project/vllm-omni) into .agents/skills/vllm-omni-npu-model-runner-upgrade in your project. Codex loads it when a task matches its description.

Can I use Vllm Omni Npu Model Runner Upgrade in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add vllm-project/vllm-omni --skill vllm-omni-npu-model-runner-upgrade -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/vllm-omni-npu-model-runner-upgrade, .gemini/skills/vllm-omni-npu-model-runner-upgrade, .github/skills/vllm-omni-npu-model-runner-upgrade and .opencode/skills/vllm-omni-npu-model-runner-upgrade in your project.

What does Vllm Omni Npu Model Runner Upgrade need to run?

Going by SKILL.md and its folder, Vllm Omni Npu Model Runner Upgrade needs the command-line tools its instructions call (python and git). Our summary lists: Python 3.

Does Vllm Omni Npu Model Runner Upgrade access the network?

SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Vllm Omni Npu Model Runner Upgrade safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Vllm Omni Npu Model Runner Upgrade use?

Vllm Omni Npu Model Runner Upgrade is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Vllm Omni Npu Model Runner Upgrade use?

About 3k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 7.7k tokens, read only when the agent opens those files.

What are the alternatives to Vllm Omni Npu Model Runner Upgrade?

Skills that share tags, products or a category with Vllm Omni Npu Model Runner Upgrade: SageMaker Serving Image Selection (huggingface/skills, 11k stars), Hugging Face Local Model Evals (huggingface/skills, 11k stars), CI Fails Buildkite (guqiong96/Lvllm, 465 stars) and Gptqmodel Tokenizer Normalization (ModelCloud/GPTQModel, 1.3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Vllm Omni Npu Model Runner Upgrade?

vllm-project (a GitHub organization) maintains it in vllm-project/vllm-omni, which has 7,119 GitHub stars. The repository holds 20 skills in this directory. The repository was last updated on October 11, 2026.

Source: vllm-project/vllm-omni on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.