SageMaker Serving Image Selection
huggingface/skills
Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones.
Upgrade vllm-omni NPU model runners (OmniNPUModelRunner, NPUARModelRunner, NPUGenerationModelRunner) to align with the latest vllm-ascend NPUModelRunner while preserving omni-specific logic.
$ npx skills add vllm-project/vllm-omni --skill vllm-omni-npu-model-runner-upgrade -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install vllm-project/vllm-omni vllm-omni-npu-model-runner-upgrade --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/vllm-project/vllm-omni.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/vllm-omni-npu-upgrade .claude/skills/vllm-omni-npu-model-runner-upgrade && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "vllm-omni-npu-model-runner-upgrade" agent skill from https://github.com/vllm-project/vllm-omni/tree/main/.claude/skills/vllm-omni-npu-upgrade into .claude/skills/vllm-omni-npu-model-runner-upgrade/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "vllm-omni-npu-model-runner-upgrade", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/vllm-project/vllm-omni/tree/main/.claude/skills/vllm-omni-npu-upgradeType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add vllm-project/vllm-omni --skill vllm-omni-npu-model-runner-upgrade -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install vllm-project/vllm-omni vllm-omni-npu-model-runner-upgrade --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/vllm-project/vllm-omni.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.claude/skills/vllm-omni-npu-upgrade .agents/skills/vllm-omni-npu-model-runner-upgrade && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "vllm-omni-npu-model-runner-upgrade" agent skill from https://github.com/vllm-project/vllm-omni/tree/main/.claude/skills/vllm-omni-npu-upgrade into .agents/skills/vllm-omni-npu-model-runner-upgrade/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "vllm-omni-npu-model-runner-upgrade", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add vllm-project/vllm-omni --skill vllm-omni-npu-model-runner-upgrade -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install vllm-project/vllm-omni vllm-omni-npu-model-runner-upgrade --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/vllm-project/vllm-omni.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.claude/skills/vllm-omni-npu-upgrade .cursor/skills/vllm-omni-npu-model-runner-upgrade && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "vllm-omni-npu-model-runner-upgrade" agent skill from https://github.com/vllm-project/vllm-omni/tree/main/.claude/skills/vllm-omni-npu-upgrade into .cursor/skills/vllm-omni-npu-model-runner-upgrade/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "vllm-omni-npu-model-runner-upgrade", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/vllm-project/vllm-omni.git --path .claude/skills/vllm-omni-npu-upgrade--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add vllm-project/vllm-omni --skill vllm-omni-npu-model-runner-upgrade -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install vllm-project/vllm-omni vllm-omni-npu-model-runner-upgrade --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/vllm-project/vllm-omni.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.claude/skills/vllm-omni-npu-upgrade .gemini/skills/vllm-omni-npu-model-runner-upgrade && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "vllm-omni-npu-model-runner-upgrade" agent skill from https://github.com/vllm-project/vllm-omni/tree/main/.claude/skills/vllm-omni-npu-upgrade into .gemini/skills/vllm-omni-npu-model-runner-upgrade/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "vllm-omni-npu-model-runner-upgrade", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install vllm-project/vllm-omni vllm-omni-npu-model-runner-upgradeInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add vllm-project/vllm-omni --skill vllm-omni-npu-model-runner-upgrade -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/vllm-project/vllm-omni.git skills-src && mkdir -p .github/skills && cp -r skills-src/.claude/skills/vllm-omni-npu-upgrade .github/skills/vllm-omni-npu-model-runner-upgrade && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "vllm-omni-npu-model-runner-upgrade" agent skill from https://github.com/vllm-project/vllm-omni/tree/main/.claude/skills/vllm-omni-npu-upgrade into .github/skills/vllm-omni-npu-model-runner-upgrade/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "vllm-omni-npu-model-runner-upgrade", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add vllm-project/vllm-omni --skill vllm-omni-npu-model-runner-upgrade -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install vllm-project/vllm-omni vllm-omni-npu-model-runner-upgrade --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/vllm-project/vllm-omni.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.claude/skills/vllm-omni-npu-upgrade .opencode/skills/vllm-omni-npu-model-runner-upgrade && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "vllm-omni-npu-model-runner-upgrade" agent skill from https://github.com/vllm-project/vllm-omni/tree/main/.claude/skills/vllm-omni-npu-upgrade into .opencode/skills/vllm-omni-npu-model-runner-upgrade/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "vllm-omni-npu-model-runner-upgrade", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
vllm-omni-npu-model-runner-upgradeUpgrade vllm-omni NPU model runners (OmniNPUModelRunner, NPUARModelRunner, NPUGenerationModelRunner) to align with the latest vllm-ascend NPUModelRunner while preserving omni-specific logic.
Vllm Omni Npu Model Runner Upgrade is an agent skill from vllm-project/vllm-omni. Upgrade vllm-omni NPU model runners (OmniNPUModelRunner, NPUARModelRunner, NPUGenerationModelRunner) to align with the latest vllm-ascend NPUModelRunner while preserving omni-specific logic.
Its SKILL.md is about 3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files, including reference files (for example `references/gpu-to-npu-translation.md`, `references/omni-specific-blocks.md` and `references/workflow-checklist.md`).
It sits in AI & LLM Engineering, covering LLM inference and serving. It works with vLLM. The repository describes itself as: A framework for efficient model inference with omni-modality models. The licence is Apache-2.0.
12 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 096988d. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
pythongitFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use git, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Vllm Omni Npu Model Runner Upgrade loads about 3k tokens when it runs, and up to ~11k if it reads all its reference files. Until then it costs about 56 tokens; SKILL.md has 871 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from vllm-project/vllm-omni at commit 096988d, republished under its Apache-2.0 licence (© vllm-project). 871 words, ~2,993 tokens.
.claude/skills/vllm-omni-npu-model-runner-upgrade/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.This skill guides the process of upgrading vllm-omni's NPU model runners to align with the latest vllm-ascend codebase while preserving omni-specific enhancements. The NPU runners are designed to run omni multimodal models (like Qwen3-Omni, Bagel, MiMoAudio) on Ascend NPUs.
vllm-omni/vllm_omni/platforms/npu/worker/
├── __init__.py
├── npu_model_runner.py # OmniNPUModelRunner (base class)
├── npu_ar_model_runner.py # NPUARModelRunner (autoregressive)
├── npu_ar_worker.py # AR worker
├── npu_generation_model_runner.py # NPUGenerationModelRunner (diffusion/non-AR)
└── npu_generation_worker.py # Generation workervllm-omni/vllm_omni/worker/
├── __init__.py
├── gpu_model_runner.py # OmniGPUModelRunner
├── gpu_ar_model_runner.py # GPUARModelRunner
├── gpu_ar_worker.py
├── gpu_generation_model_runner.py
├── gpu_generation_worker.py
├── mixins.py
├── base.py
└── gpu_memory_utils.pyvllm-ascend/vllm_ascend/worker/
├── model_runner_v1.py # NPUModelRunner (base class to copy from)
├── npu_input_batch.py
├── block_table.py
├── pcp_utils.py
└── worker.py GPUModelRunner (vllm)
|
+----------------+----------------+
| |
OmniGPUModelRunner NPUModelRunner (vllm-ascend)
(vllm_omni/worker) (vllm_ascend/worker)
| |
+----------- OmniNPUModelRunner --+
(multiple inheritance)
|
+---------------+---------------+
| |
NPUARModelRunner NPUGenerationModelRunner
(autoregressive) (non-autoregressive/diffusion)Omni-specific logic is marked with comment blocks:
# -------------------------------------- Omni-new -------------------------------------------------
# ... omni-specific code ...
# -------------------------------------- Omni-new -------------------------------------------------Or simpler variations:
# -------------------------------------- Omni-new -------------------------------------------------
# ------------------------------------------------------------------------------------------------Important:
references/omni-specific-blocks.md) may not be up-to-date. Always grep for Omni-new in the GPU implementations to find the authoritative list of omni-specific blocks.| Method | Description | Omni-Specific Logic |
|---|---|---|
load_model | Load model and initialize talker_mtp | Uses ACLGraphWrapper instead of CUDAGraphWrapper, initializes talker buffers |
_dummy_run | Warmup/profiling run | talker_mtp dummy forward, extract_multimodal_outputs |
_model_forward | Forward pass wrapper | Injects model_kwargs_extra, wraps with OmniOutput, NPU-specific graph updates |
_talker_mtp_forward | Talker MTP forward for Qwen3-Omni | Uses set_ascend_forward_context |
| Method | Description | Omni-Specific Logic |
|---|---|---|
__init__ | Initialize with KV transfer manager | OmniKVTransferManager setup |
execute_model | Main inference entry | KV transfer handling, _update_states override, extract_multimodal_outputs |
sample_tokens | Token sampling | Hidden states extraction, multimodal outputs processing, OmniModelRunnerOutput |
_resolve_global_request_id | Request ID resolution | For disaggregated inference |
| Method | Description | Omni-Specific Logic |
|---|---|---|
_update_request_states | Update request states for async chunk | async_chunk handling |
execute_model | Generation forward | async_chunk, seq_token_counts, _run_generation_model |
sample_tokens | Output processing | multimodal output packaging to OmniModelRunnerOutput |
_dummy_run | Dummy run override | model_kwargs initialization, multimodal extraction |
_run_generation_model | Run generation model | Calls _model_forward with sampler |
Identify target versions(Use gh cli to check):
Check GPU-side changes (since last release):
cd /root/vllm-workspace/vllm-omni
git log --oneline --since="<last-release-date>" -- vllm_omni/worker/Read latest vllm-ascend code:
/root/vllm-workspace/vllm-ascend/vllm_ascend/worker/model_runner_v1.pyFor each NPU model runner file:
Extract existing omni-specific blocks:
grep -n "Omni-new" vllm_omni/platforms/npu/worker/npu_model_runner.pyDocument each omni block:
Note: Always check the GPU implementation gpu_model_runner.py for any new omni logic not yet documented in references.
Read the latest vllm-ascend NPUModelRunner.load_model
Copy the method, keeping the structure
Re-insert omni-specific logic (check GPU gpu_model_runner.py for authoritative list):
CUDAGraphWrapper with ACLGraphWrapperUpdate _dummy_run:
_dummy_run for omni-specific blocksOmni-new marked code from GPU versionUpdate _model_forward:
Compare with GPU gpu_ar_model_runner.py for any new omni features
Copy execute_model from vllm-ascend
Re-insert omni blocks (reference references/omni-specific-blocks.md, but note it may be incomplete):
gpu_ar_model_runner.py for all Omni-new marked code blocksreferences/omni-specific-blocks.mdUpdate sample_tokens (also compare with GPU implementation):
gpu_ar_model_runner.py's sample_tokens methodOmni-new marked code blocksNote: Generation model runner may have unique omni logic for diffusion/non-AR models.
Compare with GPU gpu_generation_model_runner.py - grep for all Omni-new blocks
Update execute_model:
seq_token_counts injectionUpdate _dummy_run:
_dummy_run if existsCheck and update imports at the top of each file:
# Common vllm-ascend imports
from vllm_ascend.ascend_forward_context import get_forward_context, set_ascend_forward_context
from vllm_ascend.attention.attention_v1 import AscendAttentionState
from vllm_ascend.attention.utils import using_paged_attention
from vllm_ascend.compilation.acl_graph import ACLGraphWrapper, update_full_graph_params
from vllm_ascend.ops.rotary_embedding import update_cos_sin
from vllm_ascend.utils import enable_sp, lmhead_tp_enable
from vllm_ascend.worker.model_runner_v1 import SEQ_LEN_WITH_MAX_PA_WORKSPACE, NPUModelRunner
# Omni-specific imports
from vllm_omni.model_executor.models.output_templates import OmniOutput
from vllm_omni.worker.gpu_model_runner import OmniGPUModelRunner
from vllm_omni.outputs import OmniModelRunnerOutput
from vllm_omni.distributed.omni_connectors.kv_transfer_manager import OmniKVTransferManagerCheck recent GPU worker changes:
git diff <from-tag>..<to-tag> -- vllm_omni/worker/gpu_model_runner.py
git diff <from-tag>..<to-tag> -- vllm_omni/worker/gpu_ar_model_runner.pyIdentify new omni features that need to be ported to NPU
Apply corresponding changes to NPU runners
Run type checking:
cd /root/vllm-workspace/vllm-omni
python -m py_compile vllm_omni/platforms/npu/worker/npu_model_runner.py
python -m py_compile vllm_omni/platforms/npu/worker/npu_ar_model_runner.py
python -m py_compile vllm_omni/platforms/npu/worker/npu_generation_model_runner.pyRun import test:
python -c "from vllm_omni.platforms.npu.worker import *"Run model serving test (if hardware available):
vllm serve <model-path> --trust-remote-codeset_forward_contextset_ascend_forward_contextCUDAGraphWrapperACLGraphWrapper_make_buffer returns different structureAscendCommonAttentionMetadataAscendSamplerCUDAGraphWrapper references in NPU codeset_ascend_forward_context used instead of set_forward_contextACLGraphWrapper used for talker_mtp wrappingWhen upgrading, keep these files open for reference:
/root/vllm-workspace/vllm-ascend/vllm_ascend/worker/model_runner_v1.py/root/vllm-workspace/vllm/vllm/v1/worker/gpu_model_runner.py/root/vllm-workspace/vllm-omni/vllm_omni/worker/gpu_model_runner.py© vllm-project, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 3 other files (references) in .claude/skills/vllm-omni-npu-upgrade of vllm-project/vllm-omni.
Open the folder on GitHubat commit 096988d
Vllm Omni Npu Model Runner Upgrade next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Vllm Omni Npu Model Runner Upgrade this skillvllm-project/vllm-omni | 7.1k | — | ~3k | Automated safety check: Pass | Apache-2.0 | |
| SageMaker Serving Image Selectionhuggingface/skills | 11k | 1 repos | ~4.6k | Automated safety check: Pass | Apache-2.0 | |
| Hugging Face Local Model Evalshuggingface/skills | 11k | 2 repos | ~1.6k | Automated safety check: Pass | Apache-2.0 | |
| CI Fails Buildkiteguqiong96/Lvllm | 465 | 2 repos | ~349 | Automated safety check: Pass | Apache-2.0 | |
| Gptqmodel Tokenizer NormalizationModelCloud/GPTQModel | 1.3k | — | ~1.1k | Automated safety check: Pass | Custom licence | |
| Vllm Metax Model UpgradeMetaX-MACA/vLLM-metax | 180 | — | ~3.2k | Automated safety check: Pass | Apache-2.0 |
huggingface/skills
Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones.
huggingface/skills
Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.
guqiong96/Lvllm
Fetch and diagnose vLLM Buildkite CI failure logs. An agent skill from guqiong96/Lvllm.
ModelCloud/GPTQModel
Diagnose and correct GPT-QModel tokenizer initialization, tokenization normalization, special-token handling, prompt rendering, and chat-template problems.
MetaX-MACA/vLLM-metax
Review and upgrade MetaX model support against a target vLLM revision and installed MACA components, recursively including model-dependent attention and kernels.
guqiong96/Lvllm
Write or review Triton kernels for vLLM, with practical guidance for generated-code inspection, launch grids, indexing, specialization, tuning, and representative performance validation.
vllm-project/vllm-omni
Diagnose and optimize vLLM Omni diffusion workloads, especially Wan/Qwen/Flux-style image and video generation.
vllm-project/vllm-omni
Self-check your branch before creating a PR — catch dead code, prevent new model-specific Python examples, verify accuracy/perf claims, validate PR title format, and confirm merge readiness.
vllm-project/vllm-omni
Work on vLLM-Omni quantization for diffusion, autoregressive, omni, or multi-stage models.
vllm-project/vllm-omni
Review pull requests and local branches for vllm-project/vllm-omni with a frozen snapshot, module-design ownership, feature-design overlays, targeted validation, and concise evidence-backed findings.
vllm-project/vllm-omni
Write MiniMax H3 video generation prompts for T2VA, I2VA, FL2VA, L2VA, and Ref2VA.
vllm-project/vllm-omni
Add a new diffusion model (text-to-image, text-to-video, image-to-video, text-to-audio, image editing) to vLLM-Omni, including native non-Diffusers ports, reference-parity validation, Cache-DiT…
Works with
Categories
Upgrade vllm-omni NPU model runners (OmniNPUModelRunner, NPUARModelRunner, NPUGenerationModelRunner) to align with the latest vllm-ascend NPUModelRunner while preserving omni-specific logic. Vllm Omni Npu Model Runner Upgrade is an agent skill from vllm-project/vllm-omni. Upgrade vllm-omni NPU model runners (OmniNPUModelRunner, NPUARModelRunner, NPUGenerationModelRunner) to align with the latest vllm-ascend NPUModelRunner while preserving omni-specific logic.
Vllm Omni Npu Model Runner Upgrade fits situations like: tasks that involve LLM inference and serving.
Run `npx skills add vllm-project/vllm-omni --skill vllm-omni-npu-model-runner-upgrade -a claude-code`. Or copy the skill folder (.claude/skills/vllm-omni-npu-upgrade in vllm-project/vllm-omni) into .claude/skills/vllm-omni-npu-model-runner-upgrade in your project. Claude Code loads it when a task matches its description.
Run `npx skills add vllm-project/vllm-omni --skill vllm-omni-npu-model-runner-upgrade -a codex`. Or copy the skill folder (.claude/skills/vllm-omni-npu-upgrade in vllm-project/vllm-omni) into .agents/skills/vllm-omni-npu-model-runner-upgrade in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add vllm-project/vllm-omni --skill vllm-omni-npu-model-runner-upgrade -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/vllm-omni-npu-model-runner-upgrade, .gemini/skills/vllm-omni-npu-model-runner-upgrade, .github/skills/vllm-omni-npu-model-runner-upgrade and .opencode/skills/vllm-omni-npu-model-runner-upgrade in your project.
Going by SKILL.md and its folder, Vllm Omni Npu Model Runner Upgrade needs the command-line tools its instructions call (python and git). Our summary lists: Python 3.
SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Vllm Omni Npu Model Runner Upgrade is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 3k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 7.7k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Vllm Omni Npu Model Runner Upgrade: SageMaker Serving Image Selection (huggingface/skills, 11k stars), Hugging Face Local Model Evals (huggingface/skills, 11k stars), CI Fails Buildkite (guqiong96/Lvllm, 465 stars) and Gptqmodel Tokenizer Normalization (ModelCloud/GPTQModel, 1.3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
vllm-project (a GitHub organization) maintains it in vllm-project/vllm-omni, which has 7,119 GitHub stars. The repository holds 20 skills in this directory. The repository was last updated on October 11, 2026.
Source: vllm-project/vllm-omni on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.