Universal VLM (vision-language understanding model) training guide for Huawei Ascend NPU using MindSpeed-MM.

No licenceAuto-check passedAI & LLM Engineering

Install Mindspeed Mm Vlm

skills CLI
$ npx skills add ascend-ai-coding/awesome-ascend-skills --skill mindspeed-mm-vlm -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install ascend-ai-coding/awesome-ascend-skills mindspeed-mm-vlm --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/ascend-ai-coding/awesome-ascend-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/training/mindspeed-mm/mindspeed-mm-vlm .claude/skills/mindspeed-mm-vlm && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
mindspeed-mm-vlm
GitHub stars
174
Token cost
~5.2k tokens
SKILL.md length
1,612 words
Files
4 (incl. references)
Skills in repo
70
Repo updated
First seen
Licence
None found

At a glance

Universal VLM (vision-language understanding model) training guide for Huawei Ascend NPU using MindSpeed-MM.

  • Works in 8 steps: Model-Specific Dependencies → Weight Download and Conversion → Dataset Preparation → …
  • Fine-tuning any multimodal understanding model on Ascend NPU
  • SKILL.md covers Prerequisites, Supported VLM Models, How to Train Any VLM Model and Quick Start: Qwen2.5VL-3B…, plus 5 more sections
  • Calls pip, bash and git; reaches gitcode.com and github.com

What it does

Mindspeed Mm Vlm is an agent skill from ascend-ai-coding/awesome-ascend-skills. Universal VLM (vision-language understanding model) training guide for Huawei Ascend NPU using MindSpeed-MM. Covers all three framework patterns (Megatron, FSDP2, Custom trainers), weight conversion, dataset preparation (MLLM JSON format), fine-tuning, inference, and evaluation. Supports Qwen2.5VL, Qwen2VL, Qwen3VL, InternVL2.5/3/3.5, GLM4.1V, GLM4.5V, DeepSeekVL2, DeepSeekOCR, Ming, and more. Use when training or fine-tuning any multimodal understanding model on Ascend NPU.

Its SKILL.md is about 5.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files, including reference files (for example `references/data-format.md`, `references/model-delta-cards.md` and `references/per-model-deps.md`).

It sits in AI & LLM Engineering, covering Fine-tuning and Computer vision. It works with NVIDIA AI Platform and Qwen. The repository describes itself as: A comprehensive knowledge base for Huawei Ascend NPU development, structured as distributed Agent Skills. https://ascend-ai-coding.github.io/awesome-ascend-skills/.

When your agent uses it

  • Fine-tuning any multimodal understanding model on Ascend NPU
  • Tasks that involve Fine-tuning
  • Tasks that involve Computer vision

Example prompts

  • “/mindspeed-mm-vlm”

Requirements

  • Python 3
  • Docker

Workflow steps

8 steps, taken from the step headings in SKILL.md.

  1. Model-Specific Dependencies
  2. Weight Download and Conversion
  3. Dataset Preparation
  4. Configuration Files
  5. Training Script Configuration
  6. Launch Training
  7. Inference Verification (Optional)
  8. Evaluation (Optional)

What it can do on your machine

Read from SKILL.md and the folder at commit 62a4ecb. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pip
    • bash
    • git
    • python3
    • huggingface-cli
    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • gitcode.com
    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Mindspeed Mm Vlm loads about 5.2k tokens when it runs, and up to ~10k if it reads all its reference files. Until then it costs about 124 tokens; SKILL.md has 1,612 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~124
When it runs · the whole SKILL.md, loaded when a task matches
~5.2k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~10k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

Without a licence we can't republish the file, so here is its outline and opening line. It has 1,612 words (~5,174 tokens).

“This Skill guides users through training multimodal understanding (VLM) models on Huawei Ascend NPU using MindSpeed-MM. It uses Qwen2.5VL-3B as the flagship example and covers the end-to-end fine-tuning workflow.”

— opening of SKILL.md by ascend-ai-coding
name
mindspeed-mm-vlm
keywords
mindspeed-mm, vlm, multimodal understanding, qwen2.5vl, qwen3vl, internvl, glm4v, deepseekvl, finetune, fine-tuning, vision-language, megatron, fsdp2

Read the full SKILL.md on GitHub

Files

SKILL.md and 3 other files (references) in skills/training/mindspeed-mm/mindspeed-mm-vlm of ascend-ai-coding/awesome-ascend-skills.

  • SKILL.md
  • references/data-format.md
  • references/model-delta-cards.md
  • references/per-model-deps.md

Open the folder on GitHubat commit 62a4ecb

Compare with similar skills

Mindspeed Mm Vlm next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Mindspeed Mm Vlm compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Mindspeed Mm Vlm this skillascend-ai-coding/awesome-ascend-skills174—~5.2kAutomated safety check: PassNone
Tinker Training Costsundial-org/skills153—~1.2kAutomated safety check: PassNone
Tao Finetune ClipNVIDIA/skills3.6k—~4kAutomated safety check: NotesApache-2.0
Tao Finetune Huggingface ModelNVIDIA/skills3.6k—~4.9kAutomated safety check: NotesApache-2.0
Defect Image Generation with Cosmos AnomalyGenNVIDIA/skills3.6k—~5kAutomated safety check: NotesApache-2.0
Fix Art IssuesOpenPipe/ART11k—~840Automated safety check: NotesApache-2.0

Similar skills

  • Tinker Training Cost

    sundial-org/skills

    Calculate training costs for Tinker fine-tuning jobs. An agent skill from sundial-org/skills.

    153 GitHub stars~1.2k tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check passed
  • Tao Finetune Clip

    NVIDIA/skills

    Official

    CLIP vision-language model for image-text retrieval, zero-shot classification, embedding extraction, ONNX export, and TensorRT deployment.

    3.6k GitHub stars~4k tokensUpdated today
    AI & LLM EngineeringAuto-check: notes
  • Official

    Fine-tune any HuggingFace CV / VLM / LLM model on local NVIDIA GPUs inside an NGC PyTorch container when no dedicated TAO model skill matches.

    3.6k GitHub stars~4.9k tokensUpdated today
    AI & LLM EngineeringAuto-check: notes
  • Orchestrates defect image generation for PCBA, metal surface and glass inspection with NVIDIA Cosmos AnomalyGen on OSMO, from cold-start Day 0 to real-photo Day 1 labeling.

    3.6k GitHub stars~5k tokensUpdated today
    AI & LLM EngineeringAuto-check: notes
  • Fix Art Issues

    OpenPipe/ART

    Fix a GitHub issue on OpenPipe/ART and open a PR. An agent skill from OpenPipe/ART.

    11k GitHub stars~840 tokensUpdated today
    AI & LLM EngineeringAuto-check: notes
  • Trains and evaluates several WiFi-signal-based pose and sensing models, from unsupervised pose estimation to domain adaptation and publishing.

    97k GitHub stars~1.3k tokensUpdated today
    AI & LLM EngineeringAuto-check: notes

More from ascend-ai-coding/awesome-ascend-skills

All 70 skills in this repo
  • Ascend Dmi

    ascend-ai-coding/awesome-ascend-skills

    当用户需要对华为昇腾 NPU 进行硬件层面的管理、测试或诊断时使用此 skill。典型场景: - 查看 NPU 卡的状态、温度、利用率 - 测试内存带宽(h2d/d2h/d2d/p2p) - 跑算力/功耗基准测试(TFLOPS、TOPS) - 诊断 NPU 硬件故障或做健康检查 - 对 NPU 卡做压力测试(aicore、内存) - 复位/恢复卡住或异常的 NPU 卡 典型用户问题(即使不提…

    174 GitHub stars~1.7k tokensUpdated today
    Auto-check passed
  • Ascendc

    ascend-ai-coding/awesome-ascend-skills

    End-to-end AscendC custom operator development for Ascend NPU in an ascend-kernel (csrc/ops + build.sh + torchnpu PyTorch custom op) project.

    174 GitHub stars~3.5k tokensUpdated today
    Auto-check passed
  • Atc Model Converter

    ascend-ai-coding/awesome-ascend-skills

    Complete toolkit for Huawei Ascend NPU model conversion and end-to-end inference adaptation.

    174 GitHub stars~4.6k tokensUpdated today
    Auto-check passed
  • External Cannbot Ops Pypto Op Develop

    ascend-ai-coding/awesome-ascend-skills

    当需要编写 PyPTO 算子实现时使用此 skill。基于需求规格、设计方案和参考实现,生成完整可运行的 PyPTO 算子实现与配套测试、文档。Triggers: 实现算子、写 kernel、编写实现、写 impl、算子编码、开始编码、code the op、写 test、生成测试、写实现代码、op develop、kernel 实现。

    174 GitHub stars~2.1k tokensUpdated today
    Auto-check passed
  • External Gitcode Ascend Megatron Change Analyzer

    ascend-ai-coding/awesome-ascend-skills

    Analyze official Megatron-LM commits, PRs, and branch change sets to identify feature evolution, candidate breaking changes, and migration-relevant events.

    174 GitHub stars~1.2k tokensUpdated today
    Auto-check passed
  • External Gitcode Ascend Megatron Commit Tracker

    ascend-ai-coding/awesome-ascend-skills

    Track and normalize change requests against the official Megatron-LM repository by branch, PR, commit, commit range, or time window.

    174 GitHub stars~1.1k tokensUpdated today
    Auto-check passed

Questions about Mindspeed Mm Vlm

What does Mindspeed Mm Vlm do?

Universal VLM (vision-language understanding model) training guide for Huawei Ascend NPU using MindSpeed-MM. Mindspeed Mm Vlm is an agent skill from ascend-ai-coding/awesome-ascend-skills. Universal VLM (vision-language understanding model) training guide for Huawei Ascend NPU using MindSpeed-MM.

When should I use Mindspeed Mm Vlm?

Mindspeed Mm Vlm fits situations like: fine-tuning any multimodal understanding model on Ascend NPU; tasks that involve Fine-tuning; tasks that involve Computer vision.

How do I install Mindspeed Mm Vlm in Claude Code?

Run `npx skills add ascend-ai-coding/awesome-ascend-skills --skill mindspeed-mm-vlm -a claude-code`. Or copy the skill folder (skills/training/mindspeed-mm/mindspeed-mm-vlm in ascend-ai-coding/awesome-ascend-skills) into .claude/skills/mindspeed-mm-vlm in your project. Claude Code loads it when a task matches its description.

How do I install Mindspeed Mm Vlm in Codex?

Run `npx skills add ascend-ai-coding/awesome-ascend-skills --skill mindspeed-mm-vlm -a codex`. Or copy the skill folder (skills/training/mindspeed-mm/mindspeed-mm-vlm in ascend-ai-coding/awesome-ascend-skills) into .agents/skills/mindspeed-mm-vlm in your project. Codex loads it when a task matches its description.

Can I use Mindspeed Mm Vlm in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ascend-ai-coding/awesome-ascend-skills --skill mindspeed-mm-vlm -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/mindspeed-mm-vlm, .gemini/skills/mindspeed-mm-vlm, .github/skills/mindspeed-mm-vlm and .opencode/skills/mindspeed-mm-vlm in your project.

What does Mindspeed Mm Vlm need to run?

Going by SKILL.md and its folder, Mindspeed Mm Vlm needs the command-line tools its instructions call (pip, bash, git, python3, huggingface-cli and python). Our summary lists: Python 3; Docker.

Does Mindspeed Mm Vlm access the network?

SKILL.md names 2 domains. In commands or code: gitcode.com and github.com; the agent is likely to contact these when it follows the instructions. This is read from the text; nothing was executed.

Is Mindspeed Mm Vlm safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Mindspeed Mm Vlm use?

No licence was found for Mindspeed Mm Vlm or its repository. Without one, default copyright applies: ask the author before reusing or redistributing it.

How many tokens does Mindspeed Mm Vlm use?

About 5.2k tokens (SKILL.md is roughly 21k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 5.2k tokens, read only when the agent opens those files.

What are the alternatives to Mindspeed Mm Vlm?

Skills that share tags, products or a category with Mindspeed Mm Vlm: Tinker Training Cost (sundial-org/skills, 153 stars), Tao Finetune Clip (NVIDIA/skills, 3.6k stars), Tao Finetune Huggingface Model (NVIDIA/skills, 3.6k stars) and Defect Image Generation with Cosmos AnomalyGen (NVIDIA/skills, 3.6k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Mindspeed Mm Vlm?

ascend-ai-coding (a GitHub organization) maintains it in ascend-ai-coding/awesome-ascend-skills, which has 174 GitHub stars. The repository holds 70 skills in this directory. The repository was last updated on October 10, 2026.

Source: ascend-ai-coding/awesome-ascend-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.