Agent skill

Vllm Ascend Server

by ascend-ai-coding in ascend-ai-coding/awesome-ascend-skills

Deploy vLLM inference services on Ascend NPU servers with automatic model detection and optimized configuration.

No licenceAuto-check passedAI & LLM Engineering

Install Vllm Ascend Server

skills CLI
$ npx skills add ascend-ai-coding/awesome-ascend-skills --skill vllm-ascend-server -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install ascend-ai-coding/awesome-ascend-skills vllm-ascend-server --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/ascend-ai-coding/awesome-ascend-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/inference/vllm-ascend-server .claude/skills/vllm-ascend-server && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
vllm-ascend-server
GitHub stars
174
Token cost
~2.7k tokens
SKILL.md length
508 words
Files
24 (incl. references)
Skills in repo
70
Repo updated
First seen
Licence
None found

At a glance

Deploy vLLM inference services on Ascend NPU servers with automatic model detection and optimized configuration.

  • Works in 8 steps: Platform Confirmation → Environment Checks → 5: NPU Availability Check → …
  • Deploy vLLM server on NPU
  • SKILL.md covers Overview, Workflow Summary, Phase 0: Platform Confirmation and Phase 1: Environment Checks, plus 8 more sections
  • Calls docker, curl and pip

What it does

Vllm Ascend Server is an agent skill from ascend-ai-coding/awesome-ascend-skills. Deploy vLLM inference services on Ascend NPU servers with automatic model detection and optimized configuration. Supports local and remote deployment across bare metal, containers, and Docker images. Handles model discovery, quantization auto-detection, tensor parallelism configuration, graph/eager mode selection, and service health verification. Use when users need to: (1) Start or deploy vLLM server on NPU, (2) Launch LLM inference service, (3) Configure multi-card tensor parallel deployment, (4) Enable…

Its SKILL.md is about 2.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 26 other files, including reference files (for example `references/environment-variables.md`, `references/features.md` and `references/graph-mode.md`).

It sits in AI & LLM Engineering, covering LLM inference and serving. It works with vLLM and Docker. The repository describes itself as: A comprehensive knowledge base for Huawei Ascend NPU development, structured as distributed Agent Skills. https://ascend-ai-coding.github.io/awesome-ascend-skills/.

When your agent uses it

  • Deploy vLLM server on NPU
  • Launch LLM inference service
  • Configure multi-card tensor parallel deployment
  • Enable speculative decoding (Eagle)

Example prompts

  • “/vllm-ascend-server”

Requirements

  • Docker

Workflow steps

8 steps, taken from the step headings in SKILL.md.

  1. Platform Confirmation
  2. Environment Checks
  3. 5: NPU Availability Check
  4. Model Discovery
  5. Configuration Selection
  6. Generate Configuration
  7. Execute
  8. Verify

What it can do on your machine

Read from SKILL.md and the folder at commit 62a4ecb. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • docker
    • curl
    • pip
    • ssh

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • docs.vllm.ai

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Vllm Ascend Server loads about 2.7k tokens when it runs, and up to ~22k if it reads all its reference files. Until then it costs about 163 tokens; SKILL.md has 508 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~163
When it runs · the whole SKILL.md, loaded when a task matches
~2.7k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~22k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

Without a licence we can't republish the file, so here is its outline and opening line. It has 508 words (~2,670 tokens).

“This skill deploys vLLM inference services on Ascend NPU servers with automatic model detection, quantization handling, and performance optimization.”

— opening of SKILL.md by ascend-ai-coding
name
vllm-ascend-server

Read the full SKILL.md on GitHub

Files

SKILL.md and 23 other files (references) in skills/inference/vllm-ascend-server of ascend-ai-coding/awesome-ascend-skills.

  • SKILL.md
  • references/environment-variables.md
  • references/features.md
  • references/graph-mode.md
  • references/launch-templates/docker.md
  • references/launch-templates/health-check.md
  • references/launch-templates/offline-inference.md
  • references/launch-templates/online-serving.md
  • references/launch-templates/speculative-decoding.md
  • references/launch_templates.md
  • references/model_configs/deepseek-v3.yaml
  • references/model_configs/glm-4.x.yaml
  • references/model_configs/qwen2.5-vl.yaml
  • references/model_configs/qwen3-235b-a22b.yaml
  • references/model_configs/qwen3-30b.yaml
  • references/model_configs/qwen3-8b.yaml
  • references/model_configs/qwen3-embedding.yaml
  • references/model_configs/qwen3-reranker.yaml
  • … and 6 more

Open the folder on GitHubat commit 62a4ecb

Compare with similar skills

Vllm Ascend Server next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Vllm Ascend Server compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Vllm Ascend Server this skillascend-ai-coding/awesome-ascend-skills174—~2.7kAutomated safety check: PassNone
Dstack Prototypingdstackai/dstack2.3k—~1.6kAutomated safety check: PassMPL-2.0
Ascend Model Adapter for vLLMvllm-project/vllm-ascend2.9k—~2.2kAutomated safety check: PassApache-2.0
vLLM Model ServingOrchestra-Research/AI-Research-SKILLs13k6 repos~2.3kAutomated safety check: PassMIT
Serving LLMs On Instinctamd/skills398—~4kAutomated safety check: NotesMIT
Serving LLMs On Epycamd/skills398—~5.7kAutomated safety check: NotesMIT

Similar skills

  • Dstack Prototyping

    dstackai/dstack

    Use with the dstack skill for model-serving work when the image, serving command, resources, backend/fleet choice, or service behavior is not proven.

    2.3k GitHub stars~1.6k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Ascend Model Adapter for vLLM

    vllm-project/vllm-ascend

    Adapts and debugs Hugging Face or local models to run on vLLM with Ascend NPU, validates them by serving, and delivers the result as one signed commit.

    2.9k GitHub stars~2.2k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • vLLM Model Serving

    Orchestra-Research/AI-Research-SKILLs

    Deploys LLMs with vLLM for high-throughput serving, covering the OpenAI-compatible server, offline batch inference, monitoring and a Docker rollout.

    13k GitHub starsUsed in 6 repos~2.3k tokens
    AI & LLM EngineeringAuto-check passed
  • Serves AI models on AMD Instinct GPU hardware using vLLM. An agent skill from amd/skills.

    398 GitHub stars~4k tokensUpdated today
    AI & LLM EngineeringAuto-check: notes
  • Serves an LLM on a supported AMD EPYC server CPU using vLLM with zentorch, in Docker, Podman, or conda.

    398 GitHub stars~5.7k tokensUpdated today
    AI & LLM EngineeringAuto-check: notes
  • Vllm Deploy Docker

    vllm-project/vllm-skills

    Deploy vLLM using Docker (pre-built images or build-from-source) with NVIDIA GPU support and run the OpenAI-compatible server.

    103 GitHub stars~2.5k tokensUpdated 6 mo ago
    AI & LLM EngineeringAuto-check: notes

More from ascend-ai-coding/awesome-ascend-skills

All 70 skills in this repo
  • Ascend Dmi

    ascend-ai-coding/awesome-ascend-skills

    当用户需要对华为昇腾 NPU 进行硬件层面的管理、测试或诊断时使用此 skill。典型场景: - 查看 NPU 卡的状态、温度、利用率 - 测试内存带宽(h2d/d2h/d2d/p2p) - 跑算力/功耗基准测试(TFLOPS、TOPS) - 诊断 NPU 硬件故障或做健康检查 - 对 NPU 卡做压力测试(aicore、内存) - 复位/恢复卡住或异常的 NPU 卡 典型用户问题(即使不提…

    174 GitHub stars~1.7k tokensUpdated today
    Auto-check passed
  • Ascendc

    ascend-ai-coding/awesome-ascend-skills

    End-to-end AscendC custom operator development for Ascend NPU in an ascend-kernel (csrc/ops + build.sh + torchnpu PyTorch custom op) project.

    174 GitHub stars~3.5k tokensUpdated today
    Auto-check passed
  • Atc Model Converter

    ascend-ai-coding/awesome-ascend-skills

    Complete toolkit for Huawei Ascend NPU model conversion and end-to-end inference adaptation.

    174 GitHub stars~4.6k tokensUpdated today
    Auto-check passed
  • External Cannbot Ops Pypto Op Develop

    ascend-ai-coding/awesome-ascend-skills

    当需要编写 PyPTO 算子实现时使用此 skill。基于需求规格、设计方案和参考实现,生成完整可运行的 PyPTO 算子实现与配套测试、文档。Triggers: 实现算子、写 kernel、编写实现、写 impl、算子编码、开始编码、code the op、写 test、生成测试、写实现代码、op develop、kernel 实现。

    174 GitHub stars~2.1k tokensUpdated today
    Auto-check passed
  • External Gitcode Ascend Megatron Change Analyzer

    ascend-ai-coding/awesome-ascend-skills

    Analyze official Megatron-LM commits, PRs, and branch change sets to identify feature evolution, candidate breaking changes, and migration-relevant events.

    174 GitHub stars~1.2k tokensUpdated today
    Auto-check passed
  • External Gitcode Ascend Megatron Commit Tracker

    ascend-ai-coding/awesome-ascend-skills

    Track and normalize change requests against the official Megatron-LM repository by branch, PR, commit, commit range, or time window.

    174 GitHub stars~1.1k tokensUpdated today
    Auto-check passed

Works with

Questions about Vllm Ascend Server

What does Vllm Ascend Server do?

Deploy vLLM inference services on Ascend NPU servers with automatic model detection and optimized configuration. Vllm Ascend Server is an agent skill from ascend-ai-coding/awesome-ascend-skills. Deploy vLLM inference services on Ascend NPU servers with automatic model detection and optimized configuration.

When should I use Vllm Ascend Server?

Vllm Ascend Server fits situations like: deploy vLLM server on NPU; launch LLM inference service; configure multi-card tensor parallel deployment; enable speculative decoding (Eagle).

How do I install Vllm Ascend Server in Claude Code?

Run `npx skills add ascend-ai-coding/awesome-ascend-skills --skill vllm-ascend-server -a claude-code`. Or copy the skill folder (skills/inference/vllm-ascend-server in ascend-ai-coding/awesome-ascend-skills) into .claude/skills/vllm-ascend-server in your project. Claude Code loads it when a task matches its description.

How do I install Vllm Ascend Server in Codex?

Run `npx skills add ascend-ai-coding/awesome-ascend-skills --skill vllm-ascend-server -a codex`. Or copy the skill folder (skills/inference/vllm-ascend-server in ascend-ai-coding/awesome-ascend-skills) into .agents/skills/vllm-ascend-server in your project. Codex loads it when a task matches its description.

Can I use Vllm Ascend Server in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ascend-ai-coding/awesome-ascend-skills --skill vllm-ascend-server -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/vllm-ascend-server, .gemini/skills/vllm-ascend-server, .github/skills/vllm-ascend-server and .opencode/skills/vllm-ascend-server in your project.

What does Vllm Ascend Server need to run?

Going by SKILL.md and its folder, Vllm Ascend Server needs the command-line tools its instructions call (docker, curl, pip and ssh). Our summary lists: Docker.

Does Vllm Ascend Server access the network?

SKILL.md names 1 domain. As links in the text: docs.vllm.ai. This is read from the text; nothing was executed.

Is Vllm Ascend Server safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Vllm Ascend Server use?

No licence was found for Vllm Ascend Server or its repository. Without one, default copyright applies: ask the author before reusing or redistributing it.

How many tokens does Vllm Ascend Server use?

About 2.7k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 19k tokens, read only when the agent opens those files.

What are the alternatives to Vllm Ascend Server?

Skills that share tags, products or a category with Vllm Ascend Server: Dstack Prototyping (dstackai/dstack, 2.3k stars), Ascend Model Adapter for vLLM (vllm-project/vllm-ascend, 2.9k stars), vLLM Model Serving (Orchestra-Research/AI-Research-SKILLs, 13k stars) and Serving LLMs On Instinct (amd/skills, 398 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Vllm Ascend Server?

ascend-ai-coding (a GitHub organization) maintains it in ascend-ai-coding/awesome-ascend-skills, which has 174 GitHub stars. The repository holds 70 skills in this directory. The repository was last updated on October 8, 2026.

Source: ascend-ai-coding/awesome-ascend-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.