Agent skill

Ascend Release Manager for vLLM

by vllm-project in vllm-project/vllm-ascend

Runs the end-to-end vLLM Ascend release process: opens the release checklist and feedback issues, scans for release-blocking bugs and test coverage gaps, and generates release notes and announcements.

Apache-2.0Auto-check passedAI & LLM Engineering

Install Ascend Release Manager for vLLM

skills CLI
$ npx skills add vllm-project/vllm-ascend --skill vllm-ascend-release -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install vllm-project/vllm-ascend vllm-ascend-release --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/vllm-project/vllm-ascend.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/vllm-ascend-release .claude/skills/vllm-ascend-release && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
vllm-ascend-release
GitHub stars
2.9k
Token cost
~7.2k tokens
SKILL.md length
2,001 words
Files
13 (incl. scripts, references)
Skills in repo
2
Repo updated
First seen
Licence
Apache-2.0

At a glance

Runs the end-to-end vLLM Ascend release process: opens the release checklist and feedback issues, scans for release-blocking bugs and test coverage gaps, and generates release notes and announcements.

  • Works in 9 steps: Initialization → Bug Triage → PR Management → …
  • Starting a new vLLM Ascend RC or stable release cycle
  • SKILL.md covers Overview, When to Use This Skill, Prerequisites and Workflow Overview, plus 6 more sections
  • Runs Python scripts from its folder; calls gh, python and git; reaches vllm-ascend.readthedocs.io and docs.vllm.ai; needs GITHUB_TOKEN

What it does

After confirming the GitHub CLI is authenticated with repo and workflow scopes against vllm-project/vllm-ascend, and that Ascend NPU hardware or CI is available for functional testing, the skill gathers the release version, branch, target date and release manager, determines the previous version from the existing release list, opens a community feedback issue from a bundled template, and generates the release checklist issue from another template with a dedicated script.

A bug-triage phase scans issues since the last release, and bundled scripts also check nightly build status and test coverage, so the release manager sees critical bugs and gaps before proceeding. Further scripts fetch merged commits, generate the release announcement, and update version references and checklist sections across the repository, keeping human review at decision points such as confirming the gathered release information and judging flagged bugs, rather than automating the whole process end to end.

When your agent uses it

  • Starting a new vLLM Ascend RC or stable release cycle
  • Scanning for release-blocking bugs and nightly test failures before cutting a release
  • Generating release notes and a version-bump announcement for vLLM Ascend

Example prompts

  • “Start the release process for vLLM Ascend v0.15.0rc1.”
  • “Scan for critical bugs and test coverage gaps before we cut this release.”
  • “Generate the release announcement and bump version references for v0.15.0.”

Requirements

  • The GitHub CLI authenticated with repo and workflow scopes
  • Python with uv
  • Access to Ascend NPU hardware or CI

Workflow steps

9 steps, taken from the step headings in SKILL.md.

  1. Initialization
  2. Bug Triage
  3. PR Management
  4. Test Coverage Analysis
  5. Nightly Status
  6. Release Notes
  7. Documentation & Artifacts
  8. Release Execution
  9. WeChat Article (微信公众号推文)

What it can do on your machine

Read from SKILL.md and the folder at commit 03f31cf. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 8 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • gh
    • python
    • git
    • apt
    • brew
    • yum
    • uv

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • vllm-ascend.readthedocs.io
    • docs.vllm.ai
    • pypi.org

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • GITHUB_TOKEN

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Ascend Release Manager for vLLM loads about 7.2k tokens when it runs, and up to ~12k if it reads all its reference files. Until then it costs about 59 tokens; SKILL.md has 2,001 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~59
When it runs · the whole SKILL.md, loaded when a task matches
~7.2k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~12k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from vllm-project/vllm-ascend at commit 03f31cf, republished under its Apache-2.0 licence (© vllm-project). 2,001 words, ~7,189 tokens.

Download SKILL.mdSave it as .claude/skills/vllm-ascend-release/SKILL.md (or your agent's skills folder). This skill also uses 12 other files; get the full folder from GitHub.
name
vllm-ascend-release
description
End-to-end release management skill for vLLM Ascend. Creates release checklist issues, identifies critical bugs, runs functional tests, invokes release note generation, and guides through the complete release process.

vLLM Ascend Release Skill

Overview

This skill manages the complete end-to-end release process for vLLM Ascend, from creating the release checklist issue to final release announcement. It automates repetitive tasks while ensuring human oversight at critical decision points.

When to Use This Skill

Use this skill when:

  • Starting a new release cycle (RC or stable)
  • The release manager needs to track release progress
  • Preparing release artifacts (notes, documentation, tests)

Prerequisites

  • GitHub CLI (gh) authenticated with write access to vllm-project/vllm-ascend
  • Access to Ascend NPU hardware for functional testing (or CI infrastructure)
  • Python environment with uv for running scripts
Verify GitHub CLI Installation

Before starting the release process, verify that gh CLI is installed and authenticated:

bash
# Check if gh is installed
gh --version

# If not installed, install gh CLI:
# Ubuntu/Debian
apt install gh -y

# macOS
brew install gh

# OpenEuler
yum install gh -y

# Check authentication status
gh auth status

# If not authenticated, login with:
gh auth login

Expected output for gh auth status:

github.com
  ✓ Logged in to github.com account <username> (keyring)
  - Active account: true
  - Git operations protocol: https
  - Token: gho_****
  - Token scopes: 'gist', 'read:org', 'repo', 'workflow'

Required scopes: repo (for creating issues, PRs, releases) and workflow (for triggering CI workflows).

Workflow Overview

┌─────────────────────────────────────────────────────────────────────────────┐
│                         vLLM Ascend Release Process                         │
├─────────────────────────────────────────────────────────────────────────────┤
│                                                                             │
│  Phase 1: Initialization                                                    │
│  ├── Determine version & branch                                             │
│  ├── Create feedback issue                                                  │
│  └── Create release checklist issue                                         │
│                                                                             │
│  Phase 2: Bug Triage                                                        │
│  ├── Scan open bugs                                                         │
│  ├── Identify release-blocking bugs                                         │
│  └── Update checklist with bug list                                         │
│                                                                             │
│  Phase 3: PR Management                                                     │
│  ├── Identify must-merge PRs                                                │
│  └── Update checklist with PR list                                          │
│                                                                             │
│  Phase 4: Test Coverage Analysis                                            │
│  ├── Scan PRs for features/models without tests                             │
│  ├── Check previous feedback issue status                                   │
│  └── Update checklist with items needing manual testing                     │
│                                                                             │
│  Phase 5: Nightly Status                                                    │
│  ├── Get latest Nightly-A3 and Nightly-A2 runs                              │
│  ├── Analyze failures with extract_and_analyze.py                           │
│  └── Update checklist with nightly status table                             │
│                                                                             │
│  Phase 6: Release Notes (invoke existing skill)                             │
│  ├── Generate release notes via vllm-ascend-release-note-writer             │
│  └── Create release notes PR                                                │
│                                                                             │
│  Phase 7: Documentation & Artifacts                                         │
│  └── Update version references (Docker/wheel built by CI automatically)     │
│                                                                             │
│  Phase 8: Release Execution (requires human review)                         │
│  ├── Human review & approval                                                │
│  ├── Merge release notes PR                                                 │
│  ├── Create GitHub release                                                  │
│  └── Verify automated pipelines (PyPI, Docker, ReadTheDocs)                 │
│                                                                             │
│  Phase 9: WeChat Article (微信公众号推文)                                     │
│  ├── Collect release statistics (commits, contributors)                     │
│  ├── Generate WeChat article from template                                  │
│  └── Review and publish to WeChat official account                          │
│                                                                             │
└─────────────────────────────────────────────────────────────────────────────┘

Phase 1: Initialization

1.1 Gather Release Information

Prompt the user for:

  • Release Version: e.g., v0.15.0rc1, v0.15.0
  • Release Branch: typically main
  • Target Release Date: e.g., 2026.03.15
  • Release Manager: GitHub username
1.2 Determine Previous Version
bash
# Get the latest release tag
gh release list --repo vllm-project/vllm-ascend --limit 5

# Or check existing tags
git tag --sort=-creatordate | head -10
1.3 Create Feedback Issue

Create a community feedback issue for the release:

bash
gh issue create --repo vllm-project/vllm-ascend \
  --title "[Feedback]: v${VERSION} Release Feedback" \
  --body "$(cat templates/feedback-issue-template.md)" \
  --label "feedback"
1.4 Create Release Checklist Issue

Use the template in templates/release-checklist-template.md:

bash
# Generate the checklist from template
python scripts/generate_checklist.py \
  --version ${VERSION} \
  --branch ${BRANCH} \
  --date ${DATE} \
  --manager ${MANAGER} \
  --feedback-issue ${FEEDBACK_ISSUE_NUMBER} \
  --output release-checklist.md

# Create the issue
gh issue create --repo vllm-project/vllm-ascend \
  --title "[Release]: Release checklist for ${VERSION}" \
  --body-file release-checklist.md \
  --label "release"

Phase 2: Bug Triage

2.1 Scan Issues Since Last Release

Run the issue scanning script to browse all issues since the last release:

bash
python scripts/scan_release_bugs.py \
  --repo vllm-project/vllm-ascend \
  --since-tag ${LAST_VERSION} \
  --output issue-scan.md

The script:

  1. Gets the release date of the previous version (including rc versions)
  2. Fetches all issues created since that date
  3. Generates a report with:
    • Flagged issues: Automatically flagged based on engagement or keywords
    • All open issues: Quick browse table with titles
    • Recently closed issues: May be relevant for release notes
2.2 Human Review Process

The output is designed for quick human review:

  1. Check flagged issues first - these have high engagement or concerning keywords
  2. Browse the open issues table - scan titles, click to investigate if needed
  3. Review closed issues - identify fixes that should be highlighted in release notes
2.3 Issue Flagging Criteria

Issues are automatically flagged when they have:

  • High reactions (≥5) or many comments (≥5)
  • Labels: bug, regression, blocker, priority:high, critical
  • Keywords in title: crash, hang, freeze, oom, error, fail, etc.
2.4 Update Checklist

After manual review, add important bugs to the release checklist:

bash
python scripts/update_checklist_section.py \
  --issue-number ${CHECKLIST_ISSUE} \
  --section "Bug need Solve" \
  --content-file bug-list.md

Phase 3: PR Management

3.1 Identify Must-Merge PRs

Scan for PRs that should be included in the release:

bash
# 1. [Priority] List open PRs/issues in the release milestone
gh pr list --repo vllm-project/vllm-ascend \
  --state open \
  --search "milestone:${VERSION}" \
  --json number,title,url,labels

gh issue list --repo vllm-project/vllm-ascend \
  --state open \
  --search "milestone:${VERSION}" \
  --json number,title,url,labels

# 2. List open PRs with release-related labels
gh pr list --repo vllm-project/vllm-ascend \
  --state open \
  --label "release-blocker" \
  --json number,title,url

# 3. List PRs merged since last release
gh pr list --repo vllm-project/vllm-ascend \
  --state merged \
  --search "merged:>${LAST_RELEASE_DATE}" \
  --json number,title,mergedAt

Priority Order:

  1. PRs/Issues in the release milestone - these are explicitly targeted for this release
  2. PRs with release-blocker label - critical items that must be merged
  3. Recently merged PRs - for tracking what's already included
3.2 Update Checklist

Update the checklist with PRs that need to be merged:

bash
python scripts/update_checklist_section.py \
  --issue-number ${CHECKLIST_ISSUE} \
  --section "PR need Merge" \
  --content-file pr-list.md

Phase 4: Test Coverage Analysis

4.1 Identify Features/Models Needing Testing

CI already covers most test cases. Manual testing is only needed for:

  • New features merged without test cases
  • New models added due to environment constraints (e.g., CI doesn't have the model)
  • Issues reported in the previous release's feedback

Run the test coverage scanner:

bash
python scripts/scan_test_coverage.py \
  --repo vllm-project/vllm-ascend \
  --since-tag ${LAST_VERSION} \
  --feedback-issue ${PREVIOUS_FEEDBACK_ISSUE} \
  --output test-coverage-analysis.md

This script:

  1. Scans PRs merged since the last release
  2. Identifies features/models without corresponding test files
  3. Checks the previous feedback issue for unresolved problems
4.2 Review the Analysis

The output categorizes items:

Features/Models Needing Manual Testing:

  • New model support (e.g., Kimi K2.5, GLM-5)
  • Features that couldn't be tested in CI

Previous Feedback Status:

  • Unresolved issues from the feedback thread
  • Items that need manual verification
4.3 Manual Testing Checklist

For items identified above, perform manual testing:

markdown
#### Manual Testing Required

- [ ] Model: Kimi K2.5 - Basic inference works
- [ ] Model: GLM-5 - Multimodal features work
- [ ] Feature: Expert parallel with 8 GPUs
- [ ] Feedback: User reported slow startup (verify fixed)
4.4 Update Checklist with Results
bash
python scripts/update_checklist_section.py \
  --issue-number ${CHECKLIST_ISSUE} \
  --section "Functional Test" \
  --content-file test-results.md

Phase 5: Nightly Status

5.1 Analyze Nightly CI Runs

Get the latest Nightly-A3 and Nightly-A2 CI runs and analyze failures:

bash
python scripts/scan_nightly_status.py \
  --repo vllm-project/vllm-ascend \
  --output nightly-status.md

This script:

  1. Fetches the latest Nightly-A3 and Nightly-A2 workflow runs
  2. Calls extract_and_analyze.py (from main2main-error-analysis skill) for failed runs
  3. Extracts and categorizes errors:
    • Code Bugs: Real issues that need fixing
    • Environment Flakes: Transient issues (network, disk, etc.)
5.2 Review Output

The output includes:

WorkflowStatusFailed JobsCode BugsEnv FlakesRun
Nightly-A3✅ success0/1500#123
Nightly-A2❌ failure3/1221#124

For failed runs, it also shows:

  • Code bugs that need fixing before release
  • Failed test cases
  • Environment flakes (informational)
5.3 Update Checklist
bash
python scripts/update_checklist_section.py \
  --issue-number ${CHECKLIST_ISSUE} \
  --section "Nightly Status" \
  --content-file nightly-status.md

Phase 6: Release Notes

This phase handles the complete release notes writing process, from fetching commits to producing the final release notes.

6.1 Fetch Commits

Fetch all commits between the previous and current version:

bash
# Create output directory
mkdir -p output/${VERSION}

# Fetch commits with contributor statistics
uv run python scripts/fetch_commits.py \
  --owner vllm-project \
  --repo vllm-ascend \
  --base-tag ${LAST_VERSION} \
  --head-tag ${NEW_VERSION} \
  --stats \
  --output output/${VERSION}/0-current-raw-commits.md \
  --stats-output output/${VERSION}/0-contributor-stats.md

The script outputs:

  • 0-current-raw-commits.md: Raw commit list for analysis
  • 0-contributor-stats.md: Contributor statistics including new contributors
6.2 Analyze Commits

Create a CSV file to analyze each commit:

bash
# Create analysis workspace
touch output/${VERSION}/1-commit-analysis-draft.csv

The CSV should have headers:

ColumnDescription
titleCommit title
pr numberPR number
user facing impact/summaryWhat users should know
categoryHighlights/Features/Performance/etc.
decisioninclude/exclude/merge
reasonWhy this decision
6.3 Draft Release Notes

Create the initial draft following the category order:

markdown
## v${VERSION} - ${DATE}

This is the first release candidate of v${VERSION} for vLLM Ascend.
Please follow the [official doc](https://docs.vllm.ai/projects/ascend/en/latest) to get started.

### Highlights
(Top 3-5 most impactful changes)

### Features
(New functionality)

### Hardware and Operator Support
(New hardware/operators)

### Performance
(Performance improvements)

### Dependencies
(Version upgrades)

### Deprecation & Breaking Changes
(Breaking changes)

### Documentation
(Doc updates)

### Others
(Bug fixes, misc)

### Known Issue
(Known limitations)

Save drafts to:

  • output/${VERSION}/2-highlights-note-draft.md - Initial draft
  • output/${VERSION}/3-highlights-note-edit.md - Reviewed/edited version
6.4 Release Notes Writing Guidelines

Inclusion Criteria:

  • User experience improvements (CLI, error messages)
  • Core features (PD Disaggregation, KVCache, Graph mode, CP/SP, quantization)
  • Breaking changes and deprecations (always include)
  • Significant infrastructure changes
  • Major dependency updates (CANN/torch_npu/triton-ascend)
  • Hardware compatibility expansions (310P, A2, A3)

Writing Tips:

  • Focus on what users should know, not internal details
  • Look up PR descriptions when uncertain: gh pr view <number> --repo vllm-project/vllm-ascend
  • Group related changes together
  • Include PR links: [#12345](https://github.com/vllm-project/vllm-ascend/pull/12345)

Reference:

  • See references/ref-past-release-notes-highlight.md for style examples
6.5 Create Release Notes PR

After release notes are finalized:

bash
# Create branch
git checkout -b release/${VERSION}

# Make changes (see Phase 6 for full list)
# ...

# Create PR
gh pr create --repo vllm-project/vllm-ascend \
  --title "Release ${VERSION}" \
  --body "Release notes and version updates for ${VERSION}" \
  --label "release"

Phase 7: Documentation & Artifacts

7.1 Files to Update
FileUpdate Required
README.mdGetting Started version, Branch section
README.zh.mdSame as above (Chinese)
docs/source/faqs.mdFeedback issue link
docs/source/user_guide/release_notes.mdAdd new release notes
docs/source/community/versioning_policy.mdCompatibility matrix, release window
docs/source/community/contributors.mdNew contributors
mkdocs.ymlPackage version (in extra: block)
.github/workflows/schedule_image_build_and_push.yamlConfig
7.2 Version Update Script
bash
python scripts/update_version_references.py \
  --version ${VERSION} \
  --vllm-version ${VLLM_VERSION} \
  --feedback-issue ${FEEDBACK_ISSUE_URL}

Phase 8: Release Execution

8.1 Pre-Release Checklist

Before executing the release, verify:

  • All P0/P1 bugs resolved or documented as known issues
  • All must-merge PRs merged
  • Functional tests passing
  • Release notes reviewed and approved
  • Documentation updated
  • CI passing on release branch
8.2 Execute Release

⚠️ Human Review Required: Before executing the release, ensure all previous phases have been reviewed and approved by the release manager. This step requires explicit human confirmation.

Current Approach (Manual): For now, execute release steps manually through GitHub UI or CLI after human review:

  1. Merge release notes PR - Review PR, ensure CI passes, then merge via GitHub UI
  2. Create GitHub release - Go to GitHub Releases page, create new release with tag
  3. Verify automated pipelines - Docker image and wheel package are built automatically by CI

Future Approach (Automated): Once the release process is mature and well-tested, consider:

  • Adding a GitHub Actions workflow with manual trigger (workflow_dispatch)
  • Requiring approval from release manager before workflow proceeds
  • Automating all steps below with proper guards

Manual Execution Commands (for reference):

bash
# 1. Merge release notes PR (after human review)
gh pr merge ${RELEASE_PR_NUMBER} --repo vllm-project/vllm-ascend --squash

# 2. Create GitHub release
gh release create ${VERSION} \
  --repo vllm-project/vllm-ascend \
  --title "vLLM Ascend ${VERSION}" \
  --notes-file release-notes.md \
  --target main

# 3. Verify automated pipelines (no action needed - CI handles these)
# - Docker image: quay.io/ascend/vllm-ascend:${VERSION}
# - PyPI package: https://pypi.org/project/vllm-ascend/${VERSION}
# - ReadTheDocs: https://app.readthedocs.org/dashboard/

# 4. Upload 310P wheel if applicable
gh release upload ${VERSION} \
  --repo vllm-project/vllm-ascend \
  vllm_ascend-${VERSION}-310p-*.whl
8.3 Post-Release
bash
# 1. Broadcast release (prepare announcement)
python scripts/generate_announcement.py \
  --version ${VERSION} \
  --release-notes release-notes.md \
  --output announcement.md

# 2. Close release checklist issue
gh issue close ${CHECKLIST_ISSUE} \
  --repo vllm-project/vllm-ascend \
  --comment "Release ${VERSION} completed successfully!"

Phase 9: WeChat Article (微信公众号推文)

After release notes are finalized and the release is completed, generate a WeChat article for community broadcast.

9.1 Article Structure Template

The WeChat article follows a structured format with emojis for visual appeal:

SectionEmojiDescriptionRecommended Items
Opening Paragraph🎉Version announcement + positioning + core highlights summary1 paragraph
Statistics🥳Number of commits, new contributors1 line
Core Highlights💥Top 2-3 most important features/optimizations2-3 items
New Features🆕New functionality, models, operators3-5 items
Performance🚀Performance improvements (include metrics when available)2-4 items
Refactoring🔨Code refactoring, dependency upgrades1-3 items
Bug Fixes🐞Important bug fixes3-5 items
Quality/Testing🛡️Test coverage, CI/CD improvements0-2 items
Documentation📄Documentation updates (can combine into 1 item)1 item
Links➡️Source code, quick start, installation guide3 links
Show full SKILL.md (780 more words)Show less
9.2 Article Template
markdown
vLLM Ascend ${VERSION}版本发布🎉 此版本是针对vLLM v${VLLM_VERSION}系列版本首个RC版本,[1-2句核心亮点描述]。

🥳 本版本共计${COMMITS_COUNT}个commits,新增${NEW_CONTRIBUTORS_COUNT}位新开发者!
💥 [核心亮点1]
💥 [核心亮点2]
🆕 [新特性1]
🆕 [新特性2]
🆕 [新特性3]
🚀 [性能优化1,最好包含具体数据如"提升X%"]
🚀 [性能优化2]
🔨 [重构/依赖升级1]
🔨 [重构/依赖升级2]
🐞 修复 [重要bug1]
🐞 修复 [重要bug2]
🐞 修复 [重要bug3]
🛡️ [质量/测试改进]
📄 [文档更新汇总]

➡️ 源码地址:https://github.com/vllm-project/vllm-ascend/releases/tag/${VERSION}
➡️ 快速体验:https://vllm-ascend.readthedocs.io/en/latest/quick_start.html
➡️ 安装指南:https://vllm-ascend.readthedocs.io/en/latest/installation.html
9.3 Fetch Release Note from Release Tag

Important: WeChat articles are typically published after the release is complete. Always fetch the release note directly from the release tag, as it contains the most accurate and up-to-date information including the precise new contributor count.

bash
# Fetch release note from release tag (recommended - most accurate source)
gh release view ${VERSION} --repo vllm-project/vllm-ascend --json body,name,tagName

# The release body contains:
# - Highlights, Features, Performance, Documentation sections
# - Bug fixes (Others section)
# - Dependencies and Known Issues
# - New Contributors list with exact count

Why use release tag instead of other sources:

  • The release tag's "New Contributors" section is auto-generated by GitHub and is the most accurate
  • Release notes in the tag may have last-minute updates not in the PR
  • Dependencies and Known Issues sections are finalized at release time
9.4 Writing Guidelines
  1. Opening Paragraph:

    • Start with version number and 🎉
    • Describe version positioning (RC/stable, which vLLM version)
    • Highlight 1-2 core themes of this release
  2. Content Selection:

    • Prioritize user-facing features over internal refactoring
    • Include specific performance numbers when available
    • Group related items (e.g., multiple bug fixes for one feature)
    • Highlight breaking changes or dependency upgrades
  3. Language Style:

    • Use concise, active voice
    • Avoid overly technical jargon
    • Keep each item to one line when possible
    • Use "完成支持/适配" for new features, "优化/提升" for performance
  4. Statistics from Release Tag:

    • New contributor count: Count entries in "New Contributors" section of release body
    • For commits count (if needed): git rev-list --count ${LAST_VERSION}..${VERSION}
9.5 Example: v0.18.0rc1
vLLM Ascend v0.18.0rc1版本发布🎉 此版本是针对vLLM v0.18.0系列版本首个RC版本,重点完成了C8(INT8 KV cache)对GQA attention模型的支持,以及性能优化、问题修复等。

🥳 本版本新增9位新开发者,感谢社区开发者的持续贡献!
💥 C8(INT8 KV cache)支持GQA attention模型,同时适配DeepSeek-V3.1 PD分离场景
💥 DeepSeek模型通过新MLA算子支持Ascend 950PR&950DT 系列产品
🆕 Flash Comm V1支持VL模型的MLA,解除多模态服务限制
🆕 支持speculative decoding中target和draft模型使用不同attention backend
🆕 VL MoE模型支持SP,`sp_threshold`替换为vLLM原生`sp_min_token_num`
🆕 Qwen VL模型支持`w8a8_mxfp8`量化
🚀 Triton算子重编译优化,提升算子性能
🚀 Qwen3.5/Qwen3-Next GDN prefill路径优化,预构建chunk metadata减少h2d同步开销
🚀 FIA prefill context merge路径简化,提升运行时效率
🐞 TorchNPU 和 triton-ascend 依赖版本更新,请参考官方release note
🐞 修复PD分离场景decode节点因DP节点shape不对齐导致卡住的问题
🐞 修复单卡部署多实例显存 OOM 问题
🐞 修复 DeepSeek v3.1 C8在MTP + full decode + full graph模式下的问题
🐞 修复`AscendModelSlimConfig`中量化配置key映射导致的权重加载报错问题

📄 更新Kimi-K2.5、GLM-4.7、DeepSeek-V3.2、MiniMax-M2.5及PD分离部署文档

➡️ 源码地址:
https://github.com/vllm-project/vllm-ascend/releases/tag/v0.18.0rc1
➡️ 快速体验:
https://vllm-ascend.readthedocs.io/en/v0.18.0/quick_start.html
➡️ 安装指南:
https://docs.vllm.ai/projects/ascend/en/v0.18.0/installation.html

Script Reference

scripts/fetch_commits.py

Fetches all commits between two tags and generates contributor statistics.

Arguments:

  • --owner: Repository owner (default: vllm-project)
  • --repo: Repository name (default: vllm-ascend)
  • --base-tag: Base tag (older version, e.g., v0.14.0)
  • --head-tag: Head tag (newer version, e.g., v0.15.0rc1)
  • --output: Output file for commits (default: 0-current-raw-commits.md)
  • --stats: Generate contributor statistics
  • --stats-output: Output file for statistics (default: 0-contributor-stats.md)
  • --sort: Sort mode (chronological/alphabetical/reverse)
  • --include-date: Include commit date in output
  • --token: GitHub token (or use GITHUB_TOKEN env var)

Output:

  • Commit list in markdown format with PR links
  • Contributor statistics including new contributors
scripts/generate_checklist.py

Generates the release checklist issue body from template.

Arguments:

  • --version: Release version (e.g., v0.15.0rc1)
  • --branch: Release branch (default: main)
  • --date: Target release date
  • --manager: Release manager GitHub username
  • --feedback-issue: Feedback issue number
  • --output: Output file path
scripts/scan_release_bugs.py

Scans GitHub issues since the last release for human review.

Arguments:

  • --repo: Repository (default: vllm-project/vllm-ascend)
  • --since-tag: Previous release tag (including rc versions)
  • --state: Issue state filter (open, closed, all; default: all)
  • --output: Output file path

Output: Markdown report with:

  • Flagged issues (auto-detected as important)
  • All open issues table for quick browsing
  • Recently closed issues summary
scripts/scan_test_coverage.py

Identifies features/models that need manual testing.

Arguments:

  • --repo: Repository (default: vllm-project/vllm-ascend)
  • --since-tag: Previous release tag
  • --feedback-issue: Previous release feedback issue number (optional)
  • --output: Output file path

Output: Markdown report with:

  • Features/models merged without test coverage
  • Previous feedback issue status (resolved/unresolved)
scripts/scan_nightly_status.py

Scans Nightly CI status for release readiness.

Arguments:

  • --repo: Repository (default: vllm-project/vllm-ascend)
  • --output: Output file path

Output: Markdown report with:

  • Summary table of Nightly-A3 and Nightly-A2 status
  • Code bugs that need fixing (from extract_and_analyze.py)
  • Environment flakes (informational)
  • Failed test cases

Dependencies:

  • Calls main2main-error-analysis/scripts/extract_and_analyze.py for detailed analysis
scripts/update_checklist_section.py

Updates a specific section of the release checklist issue.

Arguments:

  • --issue-number: Release checklist issue number
  • --section: Section name to update
  • --content-file: File containing new content
  • --append: Append to section instead of replace
scripts/update_version_references.py

Updates version references across documentation files.

Arguments:

  • --version: New version
  • --vllm-version: Compatible vLLM version
  • --feedback-issue: Feedback issue URL
scripts/generate_announcement.py

Generates release announcement for broadcasting.

Arguments:

  • --version: Release version
  • --release-notes: Release notes file
  • --output: Output file path

Templates

templates/release-checklist-template.md

The release checklist issue template (see file for full template).

templates/feedback-issue-template.md

The feedback collection issue template.

References

references/version-files.yaml

List of files that need version updates and their update patterns.

references/ref-past-release-notes-highlight.md

Past release notes examples for style and category reference. Use this as a guide when writing new release notes to maintain consistency in:

  • Section ordering and naming
  • Writing style and tone
  • Level of detail for different categories
  • How to describe features, bug fixes, and breaking changes

Error Handling

Common Issues
IssueSolution
GitHub API rate limitUse authenticated requests, implement backoff
Test timeoutIncrease timeout, check hardware availability
Model not foundVerify model path, check storage
CI failureCheck CI logs, retry or fix
Recovery Procedures

If the release process fails midway:

  1. Check the release checklist issue for current state
  2. Resume from the last incomplete step
  3. Update checklist with failure notes
  4. Notify release manager

Important Notes

  1. Human Oversight: This skill automates tasks but requires human approval at key decision points (bug prioritization, test results review, release approval).

  2. Idempotency: Most scripts can be re-run safely. Issue updates use section replacement.

  3. Rollback: If a release needs to be rolled back:

    • Delete the GitHub release
    • Revert the release notes PR
    • Update checklist issue with rollback notes
  4. Communication: Keep the community informed through the feedback issue and release checklist.

  5. Testing: Always run functional tests before release, even for RC versions.

© vllm-project, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 12 other files (scripts, references) in .agents/skills/vllm-ascend-release of vllm-project/vllm-ascend.

  • SKILL.md
  • references/ref-past-release-notes-highlight.md
  • references/version-files.yaml
  • scripts/fetch_commits.py
  • scripts/generate_announcement.py
  • scripts/generate_checklist.py
  • scripts/scan_nightly_status.py
  • scripts/scan_release_bugs.py
  • scripts/scan_test_coverage.py
  • scripts/update_checklist_section.py
  • scripts/update_version_references.py
  • templates/feedback-issue-template.md
  • templates/release-checklist-template.md

Open the folder on GitHubat commit 03f31cf

Compare with similar skills

Ascend Release Manager for vLLM next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Ascend Release Manager for vLLM compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Ascend Release Manager for vLLM this skillvllm-project/vllm-ascend2.9k—~7.2kAutomated safety check: PassApache-2.0
Gptqmodel Tokenizer NormalizationModelCloud/GPTQModel1.3k—~1.1kAutomated safety check: PassCustom licence
Vllm Upstream Deduppytorch/test-infra113—~1.3kAutomated safety check: PassCustom licence
Vllm Daily PR Issue Trackerascend-ai-coding/awesome-ascend-skills174—~731Automated safety check: PassNone
Code Create Staged Planopen-thoughts/OpenThoughts-Agent301—~1.5kAutomated safety check: PassApache-2.0
SageMaker Serving Image Selectionhuggingface/skills11k1 repos~4.6kAutomated safety check: PassApache-2.0

Similar skills

  • Diagnose and correct GPT-QModel tokenizer initialization, tokenization normalization, special-token handling, prompt rendering, and chat-template problems.

    1.3k GitHub stars~1.1k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Vllm Upstream Dedup

    pytorch/test-infra

    Review vLLM-routed torch-nightly root causes against existing upstream vLLM issues using read-only search results, and emit a validated upstream-checks artifact for the filer.

    113 GitHub stars~1.3k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Vllm Daily PR Issue Tracker

    ascend-ai-coding/awesome-ascend-skills

    Track daily PRs and Issues from vllm-project/vllm and vllm-project/vllm-ascend, filter by model (DeepSeek/Qwen/GLM/MiniMax/Kimi) and tech topics (PD disaggregation, MTP, quantization, graph mode…

    174 GitHub stars~731 tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Code Create Staged Plan

    open-thoughts/OpenThoughts-Agent

    DESIGN a non-trivial codebase change (Harbor / MarinSkyRL / vLLM / OT-Agent / LLaMA-Factory) as a dependency-ordered STAGED PLAN before writing code — a feature port, a multi-step fix with parity…

    301 GitHub stars~1.5k tokensUpdated 12 days ago
    AI & LLM EngineeringAuto-check passed
  • Official

    Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones.

    11k GitHub starsUsed in 1 repo~4.6k tokens
    AI & LLM EngineeringAuto-check passed
  • vLLM Model Serving

    Orchestra-Research/AI-Research-SKILLs

    Deploys LLMs with vLLM for high-throughput serving, covering the OpenAI-compatible server, offline batch inference, monitoring and a Docker rollout.

    13k GitHub starsUsed in 5 repos~2.3k tokens
    AI & LLM EngineeringAuto-check passed

More from vllm-project/vllm-ascend

  • Ascend Model Adapter for vLLM

    vllm-project/vllm-ascend

    Adapts and debugs Hugging Face or local models to run on vLLM with Ascend NPU, validates them by serving, and delivers the result as one signed commit.

    2.9k GitHub stars~2.2k tokensUpdated today
    Auto-check passed

Works with

Questions about Ascend Release Manager for vLLM

What does Ascend Release Manager for vLLM do?

Runs the end-to-end vLLM Ascend release process: opens the release checklist and feedback issues, scans for release-blocking bugs and test coverage gaps, and generates release notes and announcements. After confirming the GitHub CLI is authenticated with repo and workflow scopes against vllm-project/vllm-ascend, and that Ascend NPU hardware or CI is available for functional testing, the skill gathers the release version, branch, target date and release manager, determines the previous version from the existing release list, opens a community feedback issue from a bundled template, and generates the release checklist issue from another template with a dedicated script.

When should I use Ascend Release Manager for vLLM?

Ascend Release Manager for vLLM fits situations like: starting a new vLLM Ascend RC or stable release cycle; scanning for release-blocking bugs and nightly test failures before cutting a release; generating release notes and a version-bump announcement for vLLM Ascend.

How do I install Ascend Release Manager for vLLM in Claude Code?

Run `npx skills add vllm-project/vllm-ascend --skill vllm-ascend-release -a claude-code`. Or copy the skill folder (.agents/skills/vllm-ascend-release in vllm-project/vllm-ascend) into .claude/skills/vllm-ascend-release in your project. Claude Code loads it when a task matches its description.

How do I install Ascend Release Manager for vLLM in Codex?

Run `npx skills add vllm-project/vllm-ascend --skill vllm-ascend-release -a codex`. Or copy the skill folder (.agents/skills/vllm-ascend-release in vllm-project/vllm-ascend) into .agents/skills/vllm-ascend-release in your project. Codex loads it when a task matches its description.

Can I use Ascend Release Manager for vLLM in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add vllm-project/vllm-ascend --skill vllm-ascend-release -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/vllm-ascend-release, .gemini/skills/vllm-ascend-release, .github/skills/vllm-ascend-release and .opencode/skills/vllm-ascend-release in your project.

What does Ascend Release Manager for vLLM need to run?

Going by SKILL.md and its folder, Ascend Release Manager for vLLM needs Python for the scripts in its folder, the command-line tools its instructions call (gh, python, git, apt, brew and yum) and credentials named GITHUB_TOKEN. Our summary lists: The GitHub CLI authenticated with repo and workflow scopes; Python with uv; Access to Ascend NPU hardware or CI.

Does Ascend Release Manager for vLLM access the network?

SKILL.md names 3 domains. In commands or code: vllm-ascend.readthedocs.io, docs.vllm.ai and pypi.org; the agent is likely to contact these when it follows the instructions. This is read from the text; nothing was executed.

Is Ascend Release Manager for vLLM safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Ascend Release Manager for vLLM use?

Ascend Release Manager for vLLM is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Ascend Release Manager for vLLM use?

About 7.2k tokens (SKILL.md is roughly 29k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 5k tokens, read only when the agent opens those files.

What are the alternatives to Ascend Release Manager for vLLM?

Skills that share tags, products or a category with Ascend Release Manager for vLLM: Gptqmodel Tokenizer Normalization (ModelCloud/GPTQModel, 1.3k stars), Vllm Upstream Dedup (pytorch/test-infra, 113 stars), Vllm Daily PR Issue Tracker (ascend-ai-coding/awesome-ascend-skills, 174 stars) and Code Create Staged Plan (open-thoughts/OpenThoughts-Agent, 301 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Ascend Release Manager for vLLM?

vllm-project (a GitHub organization) maintains it in vllm-project/vllm-ascend, which has 2,949 GitHub stars. The repository holds 2 skills in this directory. The repository was last updated on October 11, 2026.

Source: vllm-project/vllm-ascend on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.