Agent skill

Ascend Model Adapter for vLLM

by vllm-project in vllm-project/vllm-ascend

Adapts and debugs Hugging Face or local models to run on vLLM with Ascend NPU, validates them by serving, and delivers the result as one signed commit.

Apache-2.0Auto-check passedAI & LLM Engineering

Install Ascend Model Adapter for vLLM

skills CLI
$ npx skills add vllm-project/vllm-ascend --skill vllm-ascend-model-adapter -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install vllm-project/vllm-ascend vllm-ascend-model-adapter --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/vllm-project/vllm-ascend.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/vllm-ascend-model-adapter .claude/skills/vllm-ascend-model-adapter && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
vllm-ascend-model-adapter
GitHub stars
2.9k
Token cost
~2.2k tokens
SKILL.md length
994 words
Files
6 (incl. references)
Skills in repo
2
Repo updated
First seen
Licence
Apache-2.0

At a glance

Adapts and debugs Hugging Face or local models to run on vLLM with Ascend NPU, validates them by serving, and delivers the result as one signed commit.

  • Works in 8 steps: Collect context → Analyze model first → Choose adaptation strategy (new-model… → …
  • Bringing a new model architecture to vLLM on Ascend NPU
  • SKILL.md covers Overview, Read order, Hard constraints and Execution playbook, plus 1 more section
  • Calls git

What it does

This skill covers both models that vllm-ascend already supports and new architectures not yet registered in vLLM. Work happens in /vllm-workspace/vllm and /vllm-workspace/vllm-ascend, validation uses a direct vllm serve started from /workspace on port 8000 by default, and the transformers package is never upgraded. The agent first collects context and analyzes the config.json, processor, modeling and tokenizer files.

By default it tries to validate ACLGraph, expert parallel, flashcomm1, MTP and multimodal features, marking MoE-only checks as not applicable for other models and recording the reason when a feature cannot be enabled. Dummy weights are encouraged for speed but never accepted as the only evidence, so a real-weight run is mandatory. Reference files cover the workflow checklist, troubleshooting, fp8 on NPU, multimodal and ACLGraph lessons, and deliverables. The final output is a single signed commit with compact docs in Chinese.

When your agent uses it

  • Bringing a new model architecture to vLLM on Ascend NPU
  • Debugging a startup or inference failure of a model on vllm-ascend
  • Checking an fp8 checkpoint on NPU
  • Delivering a model adaptation as one signed commit

Example prompts

  • “Adapt the model in /models/my-model to run on vllm-ascend and validate it with real weights.”
  • “vllm serve fails on startup for this checkpoint; debug it using the troubleshooting guide.”
  • “Check whether expert parallel and ACLGraph work for this MoE model and report the evidence.”

Requirements

  • A vLLM and vllm-ascend workspace on Ascend NPU hardware
  • Model weights available locally, by default under /models
  • A git repository for the final signed commit

Workflow steps

8 steps, taken from the step headings in SKILL.md.

  1. Collect context
  2. Analyze model first
  3. Choose adaptation strategy (new-model capable)
  4. Implement minimal code changes (in implementation roots)
  5. Two-stage validation on Ascend (direct run)
  6. Validate inference and features
  7. Backport, generate artifacts, and commit in delivery repo
  8. Prepare handoff artifacts

What it can do on your machine

Read from SKILL.md and the folder at commit ea01a44. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • git

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use git, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Ascend Model Adapter for vLLM loads about 2.2k tokens when it runs, and up to ~7.6k if it reads all its reference files. Until then it costs about 64 tokens; SKILL.md has 994 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~64
When it runs · the whole SKILL.md, loaded when a task matches
~2.2k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~7.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from vllm-project/vllm-ascend at commit ea01a44, republished under its Apache-2.0 licence (© vllm-project). 994 words, ~2,180 tokens.

Download SKILL.mdSave it as .claude/skills/vllm-ascend-model-adapter/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.
name
vllm-ascend-model-adapter
description
Adapt and debug existing or new models for vLLM on Ascend NPU. Implement in /vllm-workspace/vllm and /vllm-workspace/vllm-ascend, validate via direct vllm serve from /workspace, and deliver one signed commit in the current repo.

vLLM Ascend Model Adapter

Overview

Adapt Hugging Face or local models to run on vllm-ascend with minimal changes, deterministic validation, and single-commit delivery. This skill is for both already-supported models and new architectures not yet registered in vLLM.

Read order

  1. Start with references/workflow-checklist.md.
  2. Read references/multimodal-ep-aclgraph-lessons.md (feature-first checklist).
  3. If startup/inference fails, read references/troubleshooting.md.
  4. If checkpoint is fp8-on-NPU, read references/fp8-on-npu-lessons.md.
  5. Before handoff, read references/deliverables.md.

Hard constraints

  • Never upgrade transformers.
  • Primary implementation roots are fixed by Dockerfile:
    • /vllm-workspace/vllm
    • /vllm-workspace/vllm-ascend
  • Start vllm serve from /workspace with direct command by default.
  • Default API port is 8000 unless user explicitly asks otherwise.
  • Feature-first default: try best to validate ACLGraph / EP / flashcomm1 / MTP / multimodal out-of-box.
  • --enable-expert-parallel and flashcomm1 checks are MoE-only; for non-MoE models mark as not-applicable with evidence.
  • If any feature cannot be enabled, keep evidence and explain reason in final report.
  • Do not rely on PYTHONPATH=<modified-src>:$PYTHONPATH unless debugging fallback is strictly needed.
  • Keep code changes minimal and focused on the target model.
  • Final deliverable commit must be one single signed commit in the current working repo (git commit -sm ...).
  • Keep final docs in Chinese and compact.
  • Dummy-first is encouraged for speed, but dummy is NOT fully equivalent to real weights.
  • Never sign off adaptation using dummy-only evidence; real-weight gate is mandatory.

Execution playbook

1) Collect context
  • Confirm model path (default /models/<model-name>; if environment differs, confirm with user explicitly).
  • Confirm implementation roots (/vllm-workspace/vllm, /vllm-workspace/vllm-ascend).
  • Confirm delivery root (the current git repo where the final commit is expected).
  • Confirm runtime import path points to /vllm-workspace/* install.
  • Use default expected feature set: ACLGraph + EP + flashcomm1 + MTP + multimodal (if model has VL capability).
  • User requirements extend this baseline, not replace it.
2) Analyze model first
  • Inspect config.json, processor files, modeling files, tokenizer files.
  • Identify architecture class, attention variant, quantization type, and multimodal requirements.
  • Check state-dict key prefixes (and safetensors index) to infer mapping needs.
  • Decide whether support already exists in vllm/model_executor/models/registry.py.
3) Choose adaptation strategy (new-model capable)
  • Reuse existing vLLM architecture if compatible.
  • If architecture is missing or incompatible, implement native support:
    • add model adapter under vllm/model_executor/models/;
    • add processor under vllm/transformers_utils/processors/ when needed;
    • register architecture in vllm/model_executor/models/registry.py;
    • implement explicit weight loading/remap rules (including fp8 scale pairing, KV/QK norm sharding, rope variants).
  • If remote code needs newer transformers symbols, do not upgrade dependency.
  • If unavoidable, copy required modeling files from sibling transformers source and keep scope explicit.
  • If failure is backend-specific (kernel/op/platform), patch minimal required code in /vllm-workspace/vllm-ascend.
4) Implement minimal code changes (in implementation roots)
  • Touch only files required for this model adaptation.
  • Keep weight mapping explicit and auditable.
  • Avoid unrelated refactors.
5) Two-stage validation on Ascend (direct run)
  • Run from /workspace with --load-format dummy.
  • Goal: fast validate architecture path / operator path / API path.
  • Do not treat Application startup complete as pass by itself; request smoke is mandatory.
  • Require at least:
    • startup readiness (/v1/models 200),
    • one text request 200,
    • if VL model, one text+image request 200,
    • ACLGraph evidence where expected.
Stage B: real-weight mandatory gate (must pass before sign-off)
  • Remove --load-format dummy and validate with real checkpoint.
  • Goal: validate real-only risks:
    • weight key mapping,
    • fp8/fp4 dequantization path,
    • KV/QK norm sharding with real tensor shapes,
    • load-time/runtime stability.
  • Require HTTP 200 and non-empty output before declaring success.
  • Do not pass Stage B on startup-only evidence.
Show full SKILL.md (448 more words)Show less
6) Validate inference and features
  • Send GET /v1/models first.
  • Send at least one OpenAI-compatible text request.
  • For multimodal models, require at least one text+image request.
  • Validate architecture registration and loader path with logs (no unresolved architecture, no fatal missing-key errors).
  • Try feature-first validation: EP + ACLGraph path first; eager path as fallback/isolation.
  • If startup succeeds but first request crashes (false-ready), treat as runtime failure and continue root-cause isolation.
  • For torch._dynamo + interpolate + NPU contiguous failures on VL paths, try TORCHDYNAMO_DISABLE=1 as diagnostic/stability fallback.
  • For multimodal processor API mismatch (for example skip_tensor_conversion signature mismatch), use text-only isolation (--limit-mm-per-prompt set image/video/audio to 0) to separate processor issues from core weight loading issues.
  • Capacity baseline by default (single machine): max-model-len=128k + max-num-seqs=16.
  • Then expand concurrency (e.g., 32/64) if requested or feasible.
7) Backport, generate artifacts, and commit in delivery repo
  • If implementation happened in /vllm-workspace/*, backport minimal final diff to current working repo.
  • Generate test config YAML at tests/e2e/models/configs/<ModelName>.yaml following the schema of existing configs (must include model_name, hardware, tasks with accuracy metrics, and num_fewshot). Use accuracy results from evaluation to populate metric values.
  • Generate tutorial markdown at docs/source/tutorials/models/<ModelName>.md following the standard template (Introduction, Supported Features, Environment Preparation with docker tabs, Deployment with serve script, Functional Verification with curl example, Accuracy Evaluation, Performance). Fill in model-specific details: HF path, hardware requirements, TP size, max-model-len, served-model-name, sample curl, and accuracy table.
  • Update docs/source/tutorials/models/index.md to include the new tutorial.
  • Confirm test config YAML and tutorial doc are included in the staged files.
  • Commit code changes once (single signed commit).
8) Prepare handoff artifacts
  • Write comprehensive Chinese analysis report.
  • Write compact Chinese runbook for server startup and validation commands.
  • Include feature status matrix (supported / unsupported / checkpoint-missing / not-applicable).
  • Include dummy-vs-real validation matrix and explicit non-equivalence notes.
  • Include changed-file list, key logs, and final commit hash.
  • Post the SKILL.md content (or a link to it) as a comment on the originating GitHub issue to document the AI-assisted workflow.

Quality gate before final answer

  • Service starts successfully from /workspace with direct command.
  • OpenAI-compatible inference request succeeds (not startup-only).
  • Key feature set is attempted and reported: ACLGraph / EP / flashcomm1 / MTP / multimodal.
  • Capacity baseline (128k + bs16) result is reported, or explicit reason why not feasible.
  • Dummy stage evidence is present (if used), and real-weight stage evidence is present (mandatory).
  • Test config YAML exists at tests/e2e/models/configs/<ModelName>.yaml and follows the established schema (model_name, hardware, tasks, num_fewshot).
  • Tutorial doc exists at docs/source/tutorials/models/<ModelName>.md and follows the standard template (Introduction, Supported Features, Environment Preparation, Deployment, Functional Verification, Accuracy Evaluation, Performance).
  • Tutorial index at docs/source/tutorials/models/index.md includes the new model entry.
  • Exactly one signed commit contains all code changes in current working repo.
  • Final response includes commit hash, file paths, key commands, known limits, and failure reasons where applicable.

© vllm-project, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 5 other files (references) in .agents/skills/vllm-ascend-model-adapter of vllm-project/vllm-ascend.

  • SKILL.md
  • references/deliverables.md
  • references/fp8-on-npu-lessons.md
  • references/multimodal-ep-aclgraph-lessons.md
  • references/troubleshooting.md
  • references/workflow-checklist.md

Open the folder on GitHubat commit ea01a44

Compare with similar skills

Ascend Model Adapter for vLLM next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Ascend Model Adapter for vLLM compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Ascend Model Adapter for vLLM this skillvllm-project/vllm-ascend2.9k—~2.2kAutomated safety check: PassApache-2.0
Hugging Face Local Model Evalshuggingface/skills11k2 repos~1.6kAutomated safety check: PassApache-2.0
Vllm Deploy Dockervllm-project/vllm-skills102—~2.5kAutomated safety check: NotesApache-2.0
Jetson PackageNVIDIA/skills3.6k1 repos~1.8kAutomated safety check: PassApache-2.0
SageMaker Serving Image Selectionhuggingface/skills11k1 repos~4.6kAutomated safety check: PassApache-2.0
Dstack Prototypingdstackai/dstack2.3k—~1.6kAutomated safety check: PassMPL-2.0

Similar skills

  • Official

    Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.

    11k GitHub starsUsed in 2 repos~1.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Vllm Deploy Docker

    vllm-project/vllm-skills

    Deploy vLLM using Docker (pre-built images or build-from-source) with NVIDIA GPU support and run the OpenAI-compatible server.

    102 GitHub stars~2.5k tokensUpdated 6 mo ago
    AI & LLM EngineeringAuto-check: notes
  • Jetson Package

    NVIDIA/skills

    Official

    Pick Jetson-compatible containers, vLLM runtime images, and Jetson AI Lab PyPI indexes; maps Orin SM 8.7 vs Thor SM 11.0 and JetPack-specific package choices.

    3.6k GitHub starsUsed in 1 repo~1.8k tokens
    AI & LLM EngineeringAuto-check passed
  • Official

    Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones.

    11k GitHub starsUsed in 1 repo~4.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Dstack Prototyping

    dstackai/dstack

    Use with the dstack skill for model-serving work when the image, serving command, resources, backend/fleet choice, or service behavior is not proven.

    2.3k GitHub stars~1.6k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Graphsignal

    graphsignal/graphsignal

    Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.

    257 GitHub stars~6.3k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed

More from vllm-project/vllm-ascend

  • Ascend Release Manager for vLLM

    vllm-project/vllm-ascend

    Runs the end-to-end vLLM Ascend release process: opens the release checklist and feedback issues, scans for release-blocking bugs and test coverage gaps, and generates release notes and announcements.

    2.9k GitHub stars~7.2k tokensUpdated today
    Auto-check passed

Questions about Ascend Model Adapter for vLLM

What does Ascend Model Adapter for vLLM do?

Adapts and debugs Hugging Face or local models to run on vLLM with Ascend NPU, validates them by serving, and delivers the result as one signed commit. This skill covers both models that vllm-ascend already supports and new architectures not yet registered in vLLM. Work happens in /vllm-workspace/vllm and /vllm-workspace/vllm-ascend, validation uses a direct vllm serve started from /workspace on port 8000 by default, and the transformers package is never upgraded.

When should I use Ascend Model Adapter for vLLM?

Ascend Model Adapter for vLLM fits situations like: bringing a new model architecture to vLLM on Ascend NPU; debugging a startup or inference failure of a model on vllm-ascend; checking an fp8 checkpoint on NPU; delivering a model adaptation as one signed commit.

How do I install Ascend Model Adapter for vLLM in Claude Code?

Run `npx skills add vllm-project/vllm-ascend --skill vllm-ascend-model-adapter -a claude-code`. Or copy the skill folder (.agents/skills/vllm-ascend-model-adapter in vllm-project/vllm-ascend) into .claude/skills/vllm-ascend-model-adapter in your project. Claude Code loads it when a task matches its description.

How do I install Ascend Model Adapter for vLLM in Codex?

Run `npx skills add vllm-project/vllm-ascend --skill vllm-ascend-model-adapter -a codex`. Or copy the skill folder (.agents/skills/vllm-ascend-model-adapter in vllm-project/vllm-ascend) into .agents/skills/vllm-ascend-model-adapter in your project. Codex loads it when a task matches its description.

Can I use Ascend Model Adapter for vLLM in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add vllm-project/vllm-ascend --skill vllm-ascend-model-adapter -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/vllm-ascend-model-adapter, .gemini/skills/vllm-ascend-model-adapter, .github/skills/vllm-ascend-model-adapter and .opencode/skills/vllm-ascend-model-adapter in your project.

What does Ascend Model Adapter for vLLM need to run?

Going by SKILL.md and its folder, Ascend Model Adapter for vLLM needs the command-line tools its instructions call (git). Our summary lists: A vLLM and vllm-ascend workspace on Ascend NPU hardware; Model weights available locally, by default under /models; A git repository for the final signed commit.

Does Ascend Model Adapter for vLLM access the network?

SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Ascend Model Adapter for vLLM safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Ascend Model Adapter for vLLM use?

Ascend Model Adapter for vLLM is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Ascend Model Adapter for vLLM use?

About 2.2k tokens (SKILL.md is roughly 8.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 5.5k tokens, read only when the agent opens those files.

What are the alternatives to Ascend Model Adapter for vLLM?

Skills that share tags, products or a category with Ascend Model Adapter for vLLM: Hugging Face Local Model Evals (huggingface/skills, 11k stars), Vllm Deploy Docker (vllm-project/vllm-skills, 102 stars), Jetson Package (NVIDIA/skills, 3.6k stars) and SageMaker Serving Image Selection (huggingface/skills, 11k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Ascend Model Adapter for vLLM?

vllm-project (a GitHub organization) maintains it in vllm-project/vllm-ascend, which has 2,942 GitHub stars. The repository holds 2 skills in this directory. The repository was last updated on October 10, 2026.

Source: vllm-project/vllm-ascend on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.