Gptqmodel Tokenizer Normalization
ModelCloud/GPTQModel
Diagnose and correct GPT-QModel tokenizer initialization, tokenization normalization, special-token handling, prompt rendering, and chat-template problems.
Runs the end-to-end vLLM Ascend release process: opens the release checklist and feedback issues, scans for release-blocking bugs and test coverage gaps, and generates release notes and announcements.
$ npx skills add vllm-project/vllm-ascend --skill vllm-ascend-release -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install vllm-project/vllm-ascend vllm-ascend-release --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/vllm-project/vllm-ascend.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/vllm-ascend-release .claude/skills/vllm-ascend-release && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "vllm-ascend-release" agent skill from https://github.com/vllm-project/vllm-ascend/tree/main/.agents/skills/vllm-ascend-release into .claude/skills/vllm-ascend-release/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "vllm-ascend-release", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/vllm-project/vllm-ascend/tree/main/.agents/skills/vllm-ascend-releaseType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add vllm-project/vllm-ascend --skill vllm-ascend-release -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install vllm-project/vllm-ascend vllm-ascend-release --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/vllm-project/vllm-ascend.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.agents/skills/vllm-ascend-release .agents/skills/vllm-ascend-release && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "vllm-ascend-release" agent skill from https://github.com/vllm-project/vllm-ascend/tree/main/.agents/skills/vllm-ascend-release into .agents/skills/vllm-ascend-release/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "vllm-ascend-release", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add vllm-project/vllm-ascend --skill vllm-ascend-release -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install vllm-project/vllm-ascend vllm-ascend-release --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/vllm-project/vllm-ascend.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.agents/skills/vllm-ascend-release .cursor/skills/vllm-ascend-release && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "vllm-ascend-release" agent skill from https://github.com/vllm-project/vllm-ascend/tree/main/.agents/skills/vllm-ascend-release into .cursor/skills/vllm-ascend-release/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "vllm-ascend-release", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/vllm-project/vllm-ascend.git --path .agents/skills/vllm-ascend-release--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add vllm-project/vllm-ascend --skill vllm-ascend-release -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install vllm-project/vllm-ascend vllm-ascend-release --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/vllm-project/vllm-ascend.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.agents/skills/vllm-ascend-release .gemini/skills/vllm-ascend-release && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "vllm-ascend-release" agent skill from https://github.com/vllm-project/vllm-ascend/tree/main/.agents/skills/vllm-ascend-release into .gemini/skills/vllm-ascend-release/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "vllm-ascend-release", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install vllm-project/vllm-ascend vllm-ascend-releaseInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add vllm-project/vllm-ascend --skill vllm-ascend-release -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/vllm-project/vllm-ascend.git skills-src && mkdir -p .github/skills && cp -r skills-src/.agents/skills/vllm-ascend-release .github/skills/vllm-ascend-release && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "vllm-ascend-release" agent skill from https://github.com/vllm-project/vllm-ascend/tree/main/.agents/skills/vllm-ascend-release into .github/skills/vllm-ascend-release/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "vllm-ascend-release", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add vllm-project/vllm-ascend --skill vllm-ascend-release -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install vllm-project/vllm-ascend vllm-ascend-release --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/vllm-project/vllm-ascend.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.agents/skills/vllm-ascend-release .opencode/skills/vllm-ascend-release && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "vllm-ascend-release" agent skill from https://github.com/vllm-project/vllm-ascend/tree/main/.agents/skills/vllm-ascend-release into .opencode/skills/vllm-ascend-release/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "vllm-ascend-release", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
vllm-ascend-releaseRuns the end-to-end vLLM Ascend release process: opens the release checklist and feedback issues, scans for release-blocking bugs and test coverage gaps, and generates release notes and announcements.
After confirming the GitHub CLI is authenticated with repo and workflow scopes against vllm-project/vllm-ascend, and that Ascend NPU hardware or CI is available for functional testing, the skill gathers the release version, branch, target date and release manager, determines the previous version from the existing release list, opens a community feedback issue from a bundled template, and generates the release checklist issue from another template with a dedicated script.
A bug-triage phase scans issues since the last release, and bundled scripts also check nightly build status and test coverage, so the release manager sees critical bugs and gaps before proceeding. Further scripts fetch merged commits, generate the release announcement, and update version references and checklist sections across the repository, keeping human review at decision points such as confirming the gathered release information and judging flagged bugs, rather than automating the whole process end to end.
9 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 03f31cf. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 8 files in scripts/ (Python), which the agent can run.
Shell commands in SKILL.md call:
ghpythongitaptbrewyumuvFrom the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
vllm-ascend.readthedocs.iodocs.vllm.aipypi.orgFrom URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
GITHUB_TOKENFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Ascend Release Manager for vLLM loads about 7.2k tokens when it runs, and up to ~12k if it reads all its reference files. Until then it costs about 59 tokens; SKILL.md has 2,001 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from vllm-project/vllm-ascend at commit 03f31cf, republished under its Apache-2.0 licence (© vllm-project). 2,001 words, ~7,189 tokens.
.claude/skills/vllm-ascend-release/SKILL.md (or your agent's skills folder). This skill also uses 12 other files; get the full folder from GitHub.This skill manages the complete end-to-end release process for vLLM Ascend, from creating the release checklist issue to final release announcement. It automates repetitive tasks while ensuring human oversight at critical decision points.
Use this skill when:
gh) authenticated with write access to vllm-project/vllm-ascenduv for running scriptsBefore starting the release process, verify that gh CLI is installed and authenticated:
# Check if gh is installed
gh --version
# If not installed, install gh CLI:
# Ubuntu/Debian
apt install gh -y
# macOS
brew install gh
# OpenEuler
yum install gh -y
# Check authentication status
gh auth status
# If not authenticated, login with:
gh auth loginExpected output for gh auth status:
github.com
✓ Logged in to github.com account <username> (keyring)
- Active account: true
- Git operations protocol: https
- Token: gho_****
- Token scopes: 'gist', 'read:org', 'repo', 'workflow'Required scopes: repo (for creating issues, PRs, releases) and workflow (for triggering CI workflows).
┌─────────────────────────────────────────────────────────────────────────────┐
│ vLLM Ascend Release Process │
├─────────────────────────────────────────────────────────────────────────────┤
│ │
│ Phase 1: Initialization │
│ ├── Determine version & branch │
│ ├── Create feedback issue │
│ └── Create release checklist issue │
│ │
│ Phase 2: Bug Triage │
│ ├── Scan open bugs │
│ ├── Identify release-blocking bugs │
│ └── Update checklist with bug list │
│ │
│ Phase 3: PR Management │
│ ├── Identify must-merge PRs │
│ └── Update checklist with PR list │
│ │
│ Phase 4: Test Coverage Analysis │
│ ├── Scan PRs for features/models without tests │
│ ├── Check previous feedback issue status │
│ └── Update checklist with items needing manual testing │
│ │
│ Phase 5: Nightly Status │
│ ├── Get latest Nightly-A3 and Nightly-A2 runs │
│ ├── Analyze failures with extract_and_analyze.py │
│ └── Update checklist with nightly status table │
│ │
│ Phase 6: Release Notes (invoke existing skill) │
│ ├── Generate release notes via vllm-ascend-release-note-writer │
│ └── Create release notes PR │
│ │
│ Phase 7: Documentation & Artifacts │
│ └── Update version references (Docker/wheel built by CI automatically) │
│ │
│ Phase 8: Release Execution (requires human review) │
│ ├── Human review & approval │
│ ├── Merge release notes PR │
│ ├── Create GitHub release │
│ └── Verify automated pipelines (PyPI, Docker, ReadTheDocs) │
│ │
│ Phase 9: WeChat Article (微信公众号推文) │
│ ├── Collect release statistics (commits, contributors) │
│ ├── Generate WeChat article from template │
│ └── Review and publish to WeChat official account │
│ │
└─────────────────────────────────────────────────────────────────────────────┘Prompt the user for:
v0.15.0rc1, v0.15.0main2026.03.15# Get the latest release tag
gh release list --repo vllm-project/vllm-ascend --limit 5
# Or check existing tags
git tag --sort=-creatordate | head -10Create a community feedback issue for the release:
gh issue create --repo vllm-project/vllm-ascend \
--title "[Feedback]: v${VERSION} Release Feedback" \
--body "$(cat templates/feedback-issue-template.md)" \
--label "feedback"Use the template in templates/release-checklist-template.md:
# Generate the checklist from template
python scripts/generate_checklist.py \
--version ${VERSION} \
--branch ${BRANCH} \
--date ${DATE} \
--manager ${MANAGER} \
--feedback-issue ${FEEDBACK_ISSUE_NUMBER} \
--output release-checklist.md
# Create the issue
gh issue create --repo vllm-project/vllm-ascend \
--title "[Release]: Release checklist for ${VERSION}" \
--body-file release-checklist.md \
--label "release"Run the issue scanning script to browse all issues since the last release:
python scripts/scan_release_bugs.py \
--repo vllm-project/vllm-ascend \
--since-tag ${LAST_VERSION} \
--output issue-scan.mdThe script:
The output is designed for quick human review:
Issues are automatically flagged when they have:
bug, regression, blocker, priority:high, criticalAfter manual review, add important bugs to the release checklist:
python scripts/update_checklist_section.py \
--issue-number ${CHECKLIST_ISSUE} \
--section "Bug need Solve" \
--content-file bug-list.mdScan for PRs that should be included in the release:
# 1. [Priority] List open PRs/issues in the release milestone
gh pr list --repo vllm-project/vllm-ascend \
--state open \
--search "milestone:${VERSION}" \
--json number,title,url,labels
gh issue list --repo vllm-project/vllm-ascend \
--state open \
--search "milestone:${VERSION}" \
--json number,title,url,labels
# 2. List open PRs with release-related labels
gh pr list --repo vllm-project/vllm-ascend \
--state open \
--label "release-blocker" \
--json number,title,url
# 3. List PRs merged since last release
gh pr list --repo vllm-project/vllm-ascend \
--state merged \
--search "merged:>${LAST_RELEASE_DATE}" \
--json number,title,mergedAtPriority Order:
release-blocker label - critical items that must be mergedUpdate the checklist with PRs that need to be merged:
python scripts/update_checklist_section.py \
--issue-number ${CHECKLIST_ISSUE} \
--section "PR need Merge" \
--content-file pr-list.mdCI already covers most test cases. Manual testing is only needed for:
Run the test coverage scanner:
python scripts/scan_test_coverage.py \
--repo vllm-project/vllm-ascend \
--since-tag ${LAST_VERSION} \
--feedback-issue ${PREVIOUS_FEEDBACK_ISSUE} \
--output test-coverage-analysis.mdThis script:
The output categorizes items:
Features/Models Needing Manual Testing:
Previous Feedback Status:
For items identified above, perform manual testing:
#### Manual Testing Required
- [ ] Model: Kimi K2.5 - Basic inference works
- [ ] Model: GLM-5 - Multimodal features work
- [ ] Feature: Expert parallel with 8 GPUs
- [ ] Feedback: User reported slow startup (verify fixed)python scripts/update_checklist_section.py \
--issue-number ${CHECKLIST_ISSUE} \
--section "Functional Test" \
--content-file test-results.mdGet the latest Nightly-A3 and Nightly-A2 CI runs and analyze failures:
python scripts/scan_nightly_status.py \
--repo vllm-project/vllm-ascend \
--output nightly-status.mdThis script:
extract_and_analyze.py (from main2main-error-analysis skill) for failed runsThe output includes:
| Workflow | Status | Failed Jobs | Code Bugs | Env Flakes | Run |
|---|---|---|---|---|---|
| Nightly-A3 | ✅ success | 0/15 | 0 | 0 | #123 |
| Nightly-A2 | ❌ failure | 3/12 | 2 | 1 | #124 |
For failed runs, it also shows:
python scripts/update_checklist_section.py \
--issue-number ${CHECKLIST_ISSUE} \
--section "Nightly Status" \
--content-file nightly-status.mdThis phase handles the complete release notes writing process, from fetching commits to producing the final release notes.
Fetch all commits between the previous and current version:
# Create output directory
mkdir -p output/${VERSION}
# Fetch commits with contributor statistics
uv run python scripts/fetch_commits.py \
--owner vllm-project \
--repo vllm-ascend \
--base-tag ${LAST_VERSION} \
--head-tag ${NEW_VERSION} \
--stats \
--output output/${VERSION}/0-current-raw-commits.md \
--stats-output output/${VERSION}/0-contributor-stats.mdThe script outputs:
0-current-raw-commits.md: Raw commit list for analysis0-contributor-stats.md: Contributor statistics including new contributorsCreate a CSV file to analyze each commit:
# Create analysis workspace
touch output/${VERSION}/1-commit-analysis-draft.csvThe CSV should have headers:
| Column | Description |
|---|---|
title | Commit title |
pr number | PR number |
user facing impact/summary | What users should know |
category | Highlights/Features/Performance/etc. |
decision | include/exclude/merge |
reason | Why this decision |
Create the initial draft following the category order:
## v${VERSION} - ${DATE}
This is the first release candidate of v${VERSION} for vLLM Ascend.
Please follow the [official doc](https://docs.vllm.ai/projects/ascend/en/latest) to get started.
### Highlights
(Top 3-5 most impactful changes)
### Features
(New functionality)
### Hardware and Operator Support
(New hardware/operators)
### Performance
(Performance improvements)
### Dependencies
(Version upgrades)
### Deprecation & Breaking Changes
(Breaking changes)
### Documentation
(Doc updates)
### Others
(Bug fixes, misc)
### Known Issue
(Known limitations)Save drafts to:
output/${VERSION}/2-highlights-note-draft.md - Initial draftoutput/${VERSION}/3-highlights-note-edit.md - Reviewed/edited versionInclusion Criteria:
Writing Tips:
gh pr view <number> --repo vllm-project/vllm-ascend[#12345](https://github.com/vllm-project/vllm-ascend/pull/12345)Reference:
references/ref-past-release-notes-highlight.md for style examplesAfter release notes are finalized:
# Create branch
git checkout -b release/${VERSION}
# Make changes (see Phase 6 for full list)
# ...
# Create PR
gh pr create --repo vllm-project/vllm-ascend \
--title "Release ${VERSION}" \
--body "Release notes and version updates for ${VERSION}" \
--label "release"| File | Update Required |
|---|---|
README.md | Getting Started version, Branch section |
README.zh.md | Same as above (Chinese) |
docs/source/faqs.md | Feedback issue link |
docs/source/user_guide/release_notes.md | Add new release notes |
docs/source/community/versioning_policy.md | Compatibility matrix, release window |
docs/source/community/contributors.md | New contributors |
mkdocs.yml | Package version (in extra: block) |
.github/workflows/schedule_image_build_and_push.yaml | Config |
python scripts/update_version_references.py \
--version ${VERSION} \
--vllm-version ${VLLM_VERSION} \
--feedback-issue ${FEEDBACK_ISSUE_URL}Before executing the release, verify:
⚠️ Human Review Required: Before executing the release, ensure all previous phases have been reviewed and approved by the release manager. This step requires explicit human confirmation.
Current Approach (Manual): For now, execute release steps manually through GitHub UI or CLI after human review:
Future Approach (Automated): Once the release process is mature and well-tested, consider:
workflow_dispatch)Manual Execution Commands (for reference):
# 1. Merge release notes PR (after human review)
gh pr merge ${RELEASE_PR_NUMBER} --repo vllm-project/vllm-ascend --squash
# 2. Create GitHub release
gh release create ${VERSION} \
--repo vllm-project/vllm-ascend \
--title "vLLM Ascend ${VERSION}" \
--notes-file release-notes.md \
--target main
# 3. Verify automated pipelines (no action needed - CI handles these)
# - Docker image: quay.io/ascend/vllm-ascend:${VERSION}
# - PyPI package: https://pypi.org/project/vllm-ascend/${VERSION}
# - ReadTheDocs: https://app.readthedocs.org/dashboard/
# 4. Upload 310P wheel if applicable
gh release upload ${VERSION} \
--repo vllm-project/vllm-ascend \
vllm_ascend-${VERSION}-310p-*.whl# 1. Broadcast release (prepare announcement)
python scripts/generate_announcement.py \
--version ${VERSION} \
--release-notes release-notes.md \
--output announcement.md
# 2. Close release checklist issue
gh issue close ${CHECKLIST_ISSUE} \
--repo vllm-project/vllm-ascend \
--comment "Release ${VERSION} completed successfully!"After release notes are finalized and the release is completed, generate a WeChat article for community broadcast.
The WeChat article follows a structured format with emojis for visual appeal:
| Section | Emoji | Description | Recommended Items |
|---|---|---|---|
| Opening Paragraph | 🎉 | Version announcement + positioning + core highlights summary | 1 paragraph |
| Statistics | 🥳 | Number of commits, new contributors | 1 line |
| Core Highlights | 💥 | Top 2-3 most important features/optimizations | 2-3 items |
| New Features | 🆕 | New functionality, models, operators | 3-5 items |
| Performance | 🚀 | Performance improvements (include metrics when available) | 2-4 items |
| Refactoring | 🔨 | Code refactoring, dependency upgrades | 1-3 items |
| Bug Fixes | 🐞 | Important bug fixes | 3-5 items |
| Quality/Testing | 🛡️ | Test coverage, CI/CD improvements | 0-2 items |
| Documentation | 📄 | Documentation updates (can combine into 1 item) | 1 item |
| Links | ➡️ | Source code, quick start, installation guide | 3 links |
vLLM Ascend ${VERSION}版本发布🎉 此版本是针对vLLM v${VLLM_VERSION}系列版本首个RC版本,[1-2句核心亮点描述]。
🥳 本版本共计${COMMITS_COUNT}个commits,新增${NEW_CONTRIBUTORS_COUNT}位新开发者!
💥 [核心亮点1]
💥 [核心亮点2]
🆕 [新特性1]
🆕 [新特性2]
🆕 [新特性3]
🚀 [性能优化1,最好包含具体数据如"提升X%"]
🚀 [性能优化2]
🔨 [重构/依赖升级1]
🔨 [重构/依赖升级2]
🐞 修复 [重要bug1]
🐞 修复 [重要bug2]
🐞 修复 [重要bug3]
🛡️ [质量/测试改进]
📄 [文档更新汇总]
➡️ 源码地址:https://github.com/vllm-project/vllm-ascend/releases/tag/${VERSION}
➡️ 快速体验:https://vllm-ascend.readthedocs.io/en/latest/quick_start.html
➡️ 安装指南:https://vllm-ascend.readthedocs.io/en/latest/installation.htmlImportant: WeChat articles are typically published after the release is complete. Always fetch the release note directly from the release tag, as it contains the most accurate and up-to-date information including the precise new contributor count.
# Fetch release note from release tag (recommended - most accurate source)
gh release view ${VERSION} --repo vllm-project/vllm-ascend --json body,name,tagName
# The release body contains:
# - Highlights, Features, Performance, Documentation sections
# - Bug fixes (Others section)
# - Dependencies and Known Issues
# - New Contributors list with exact countWhy use release tag instead of other sources:
Opening Paragraph:
Content Selection:
Language Style:
Statistics from Release Tag:
git rev-list --count ${LAST_VERSION}..${VERSION}vLLM Ascend v0.18.0rc1版本发布🎉 此版本是针对vLLM v0.18.0系列版本首个RC版本,重点完成了C8(INT8 KV cache)对GQA attention模型的支持,以及性能优化、问题修复等。
🥳 本版本新增9位新开发者,感谢社区开发者的持续贡献!
💥 C8(INT8 KV cache)支持GQA attention模型,同时适配DeepSeek-V3.1 PD分离场景
💥 DeepSeek模型通过新MLA算子支持Ascend 950PR&950DT 系列产品
🆕 Flash Comm V1支持VL模型的MLA,解除多模态服务限制
🆕 支持speculative decoding中target和draft模型使用不同attention backend
🆕 VL MoE模型支持SP,`sp_threshold`替换为vLLM原生`sp_min_token_num`
🆕 Qwen VL模型支持`w8a8_mxfp8`量化
🚀 Triton算子重编译优化,提升算子性能
🚀 Qwen3.5/Qwen3-Next GDN prefill路径优化,预构建chunk metadata减少h2d同步开销
🚀 FIA prefill context merge路径简化,提升运行时效率
🐞 TorchNPU 和 triton-ascend 依赖版本更新,请参考官方release note
🐞 修复PD分离场景decode节点因DP节点shape不对齐导致卡住的问题
🐞 修复单卡部署多实例显存 OOM 问题
🐞 修复 DeepSeek v3.1 C8在MTP + full decode + full graph模式下的问题
🐞 修复`AscendModelSlimConfig`中量化配置key映射导致的权重加载报错问题
📄 更新Kimi-K2.5、GLM-4.7、DeepSeek-V3.2、MiniMax-M2.5及PD分离部署文档
➡️ 源码地址:
https://github.com/vllm-project/vllm-ascend/releases/tag/v0.18.0rc1
➡️ 快速体验:
https://vllm-ascend.readthedocs.io/en/v0.18.0/quick_start.html
➡️ 安装指南:
https://docs.vllm.ai/projects/ascend/en/v0.18.0/installation.htmlFetches all commits between two tags and generates contributor statistics.
Arguments:
--owner: Repository owner (default: vllm-project)--repo: Repository name (default: vllm-ascend)--base-tag: Base tag (older version, e.g., v0.14.0)--head-tag: Head tag (newer version, e.g., v0.15.0rc1)--output: Output file for commits (default: 0-current-raw-commits.md)--stats: Generate contributor statistics--stats-output: Output file for statistics (default: 0-contributor-stats.md)--sort: Sort mode (chronological/alphabetical/reverse)--include-date: Include commit date in output--token: GitHub token (or use GITHUB_TOKEN env var)Output:
Generates the release checklist issue body from template.
Arguments:
--version: Release version (e.g., v0.15.0rc1)--branch: Release branch (default: main)--date: Target release date--manager: Release manager GitHub username--feedback-issue: Feedback issue number--output: Output file pathScans GitHub issues since the last release for human review.
Arguments:
--repo: Repository (default: vllm-project/vllm-ascend)--since-tag: Previous release tag (including rc versions)--state: Issue state filter (open, closed, all; default: all)--output: Output file pathOutput: Markdown report with:
Identifies features/models that need manual testing.
Arguments:
--repo: Repository (default: vllm-project/vllm-ascend)--since-tag: Previous release tag--feedback-issue: Previous release feedback issue number (optional)--output: Output file pathOutput: Markdown report with:
Scans Nightly CI status for release readiness.
Arguments:
--repo: Repository (default: vllm-project/vllm-ascend)--output: Output file pathOutput: Markdown report with:
Dependencies:
main2main-error-analysis/scripts/extract_and_analyze.py for detailed analysisUpdates a specific section of the release checklist issue.
Arguments:
--issue-number: Release checklist issue number--section: Section name to update--content-file: File containing new content--append: Append to section instead of replaceUpdates version references across documentation files.
Arguments:
--version: New version--vllm-version: Compatible vLLM version--feedback-issue: Feedback issue URLGenerates release announcement for broadcasting.
Arguments:
--version: Release version--release-notes: Release notes file--output: Output file pathThe release checklist issue template (see file for full template).
The feedback collection issue template.
List of files that need version updates and their update patterns.
Past release notes examples for style and category reference. Use this as a guide when writing new release notes to maintain consistency in:
| Issue | Solution |
|---|---|
| GitHub API rate limit | Use authenticated requests, implement backoff |
| Test timeout | Increase timeout, check hardware availability |
| Model not found | Verify model path, check storage |
| CI failure | Check CI logs, retry or fix |
If the release process fails midway:
Human Oversight: This skill automates tasks but requires human approval at key decision points (bug prioritization, test results review, release approval).
Idempotency: Most scripts can be re-run safely. Issue updates use section replacement.
Rollback: If a release needs to be rolled back:
Communication: Keep the community informed through the feedback issue and release checklist.
Testing: Always run functional tests before release, even for RC versions.
© vllm-project, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 12 other files (scripts, references) in .agents/skills/vllm-ascend-release of vllm-project/vllm-ascend.
Open the folder on GitHubat commit 03f31cf
Ascend Release Manager for vLLM next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Ascend Release Manager for vLLM this skillvllm-project/vllm-ascend | 2.9k | — | ~7.2k | Automated safety check: Pass | Apache-2.0 | |
| Gptqmodel Tokenizer NormalizationModelCloud/GPTQModel | 1.3k | — | ~1.1k | Automated safety check: Pass | Custom licence | |
| Vllm Upstream Deduppytorch/test-infra | 113 | — | ~1.3k | Automated safety check: Pass | Custom licence | |
| Vllm Daily PR Issue Trackerascend-ai-coding/awesome-ascend-skills | 174 | — | ~731 | Automated safety check: Pass | None | |
| Code Create Staged Planopen-thoughts/OpenThoughts-Agent | 301 | — | ~1.5k | Automated safety check: Pass | Apache-2.0 | |
| SageMaker Serving Image Selectionhuggingface/skills | 11k | 1 repos | ~4.6k | Automated safety check: Pass | Apache-2.0 |
ModelCloud/GPTQModel
Diagnose and correct GPT-QModel tokenizer initialization, tokenization normalization, special-token handling, prompt rendering, and chat-template problems.
pytorch/test-infra
Review vLLM-routed torch-nightly root causes against existing upstream vLLM issues using read-only search results, and emit a validated upstream-checks artifact for the filer.
ascend-ai-coding/awesome-ascend-skills
Track daily PRs and Issues from vllm-project/vllm and vllm-project/vllm-ascend, filter by model (DeepSeek/Qwen/GLM/MiniMax/Kimi) and tech topics (PD disaggregation, MTP, quantization, graph mode…
open-thoughts/OpenThoughts-Agent
DESIGN a non-trivial codebase change (Harbor / MarinSkyRL / vLLM / OT-Agent / LLaMA-Factory) as a dependency-ordered STAGED PLAN before writing code — a feature port, a multi-step fix with parity…
huggingface/skills
Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones.
Orchestra-Research/AI-Research-SKILLs
Deploys LLMs with vLLM for high-throughput serving, covering the OpenAI-compatible server, offline batch inference, monitoring and a Docker rollout.
vllm-project/vllm-ascend
Adapts and debugs Hugging Face or local models to run on vLLM with Ascend NPU, validates them by serving, and delivers the result as one signed commit.
Categories
Runs the end-to-end vLLM Ascend release process: opens the release checklist and feedback issues, scans for release-blocking bugs and test coverage gaps, and generates release notes and announcements. After confirming the GitHub CLI is authenticated with repo and workflow scopes against vllm-project/vllm-ascend, and that Ascend NPU hardware or CI is available for functional testing, the skill gathers the release version, branch, target date and release manager, determines the previous version from the existing release list, opens a community feedback issue from a bundled template, and generates the release checklist issue from another template with a dedicated script.
Ascend Release Manager for vLLM fits situations like: starting a new vLLM Ascend RC or stable release cycle; scanning for release-blocking bugs and nightly test failures before cutting a release; generating release notes and a version-bump announcement for vLLM Ascend.
Run `npx skills add vllm-project/vllm-ascend --skill vllm-ascend-release -a claude-code`. Or copy the skill folder (.agents/skills/vllm-ascend-release in vllm-project/vllm-ascend) into .claude/skills/vllm-ascend-release in your project. Claude Code loads it when a task matches its description.
Run `npx skills add vllm-project/vllm-ascend --skill vllm-ascend-release -a codex`. Or copy the skill folder (.agents/skills/vllm-ascend-release in vllm-project/vllm-ascend) into .agents/skills/vllm-ascend-release in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add vllm-project/vllm-ascend --skill vllm-ascend-release -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/vllm-ascend-release, .gemini/skills/vllm-ascend-release, .github/skills/vllm-ascend-release and .opencode/skills/vllm-ascend-release in your project.
Going by SKILL.md and its folder, Ascend Release Manager for vLLM needs Python for the scripts in its folder, the command-line tools its instructions call (gh, python, git, apt, brew and yum) and credentials named GITHUB_TOKEN. Our summary lists: The GitHub CLI authenticated with repo and workflow scopes; Python with uv; Access to Ascend NPU hardware or CI.
SKILL.md names 3 domains. In commands or code: vllm-ascend.readthedocs.io, docs.vllm.ai and pypi.org; the agent is likely to contact these when it follows the instructions. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Ascend Release Manager for vLLM is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 7.2k tokens (SKILL.md is roughly 29k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 5k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Ascend Release Manager for vLLM: Gptqmodel Tokenizer Normalization (ModelCloud/GPTQModel, 1.3k stars), Vllm Upstream Dedup (pytorch/test-infra, 113 stars), Vllm Daily PR Issue Tracker (ascend-ai-coding/awesome-ascend-skills, 174 stars) and Code Create Staged Plan (open-thoughts/OpenThoughts-Agent, 301 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
vllm-project (a GitHub organization) maintains it in vllm-project/vllm-ascend, which has 2,949 GitHub stars. The repository holds 2 skills in this directory. The repository was last updated on October 11, 2026.
Source: vllm-project/vllm-ascend on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.