Agent skill

Npu Torchair Infer

by ascend-ai-coding in ascend-ai-coding/awesome-ascend-skills

Migrate any HuggingFace model to Ascend NPU torchair graph mode (torch.compile) and benchmark it for accuracy and performance against NPU eager and CPU eager.

No licenceAuto-check passedAI & LLM Engineering

Install Npu Torchair Infer

skills CLI
$ npx skills add ascend-ai-coding/awesome-ascend-skills --skill npu-torchair-infer -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install ascend-ai-coding/awesome-ascend-skills npu-torchair-infer --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/ascend-ai-coding/awesome-ascend-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/inference/npu-torchair-infer .claude/skills/npu-torchair-infer && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
npu-torchair-infer
GitHub stars
174
Token cost
~1.3k tokens
SKILL.md length
272 words
Files
7 (incl. scripts, references)
Skills in repo
70
Repo updated
First seen
Licence
None found

At a glance

Migrate any HuggingFace model to Ascend NPU torchair graph mode (torch.compile) and benchmark it for accuracy and performance against NPU eager and CPU eager.

  • Works in 5 steps: Step 0 on CPU — inputs/outputs resolve… → Step 1 NPU eager works. → Step 2 torchair — compiles ⇒ done; else… → …
  • Benchmarking HF models (vision
  • SKILL.md covers Overview, The one rule you cannot break, Main flow (run in order) and Decision tree, plus 3 more sections
  • Runs Python and Shell scripts from its folder; calls python and bash; reaches hf-mirror.com

What it does

Npu Torchair Infer is an agent skill from ascend-ai-coding/awesome-ascend-skills. Migrate any HuggingFace model to Ascend NPU torchair graph mode (torch.compile) and benchmark it for accuracy and performance against NPU eager and CPU eager. Use when running, compiling, or benchmarking HF models (vision, text, image-text encoders such as SigLIP2, DINOv3, ViT, CLIP, Qwen-VL, SAM) on Ascend 910B/CANN with torchnpu and torchair; when a torch.compile graph-mode run on NPU fails (Dynamo TorchRuntimeError, unsupported op, interpolate/contiguous errors); or when comparing torchair vs npueager vs cpu…

Its SKILL.md is about 1.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 8 other files, including scripts and reference files (for example `references/environment.md`, `references/methodology.md` and `references/troubleshooting.md`).

It sits in AI & LLM Engineering, covering Model hubs and datasets. It works with Hugging Face and Qwen. The repository describes itself as: A comprehensive knowledge base for Huawei Ascend NPU development, structured as distributed Agent Skills. https://ascend-ai-coding.github.io/awesome-ascend-skills/.

When your agent uses it

  • Benchmarking HF models (vision
  • Image-text encoders such as SigLIP2
  • SAM) on Ascend 910B/CANN with torchnpu and torchair
  • A torch.compile graph-mode run on NPU fails (Dynamo TorchRuntimeError

Example prompts

  • “/npu-torchair-infer”

Requirements

  • Python 3
  • A Bash shell

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Step 0 on CPU — inputs/outputs resolve (fast).
  2. Step 1 NPU eager works.
  3. Step 2 torchair — compiles ⇒ done; else walk references/troubleshooting.md.
  4. Step 3 benchmark — gate on cosine (fp16) or <1e-3 (fp32).
  5. If inputs could not be inferred, add a small build_inputs override (references/methodology.md).

What it can do on your machine

Read from SKILL.md and the folder at commit 62a4ecb. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 3 files in scripts/ (Python and Shell), which the agent can run.

    Shell commands in SKILL.md call:

    • python
    • bash

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • hf-mirror.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Npu Torchair Infer loads about 1.3k tokens when it runs, and up to ~4.4k if it reads all its reference files. Until then it costs about 150 tokens; SKILL.md has 272 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~150
When it runs · the whole SKILL.md, loaded when a task matches
~1.3k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~4.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

Without a licence we can't republish the file, so here is its outline and opening line. It has 272 words (~1,252 tokens).

“Bring an arbitrary HuggingFace model up on Ascend NPU torchair graph mode, verify it numerically against CPU/NPU-eager, and measure its speedup. The bundled scripts are model-agnostic (config-driven inputs, recursive output comparison), so the same flow applies to the next model…”

— opening of SKILL.md by ascend-ai-coding
name
npu-torchair-infer

Read the full SKILL.md on GitHub

Files

SKILL.md and 6 other files (scripts, references) in skills/inference/npu-torchair-infer of ascend-ai-coding/awesome-ascend-skills.

  • SKILL.md
  • references/environment.md
  • references/methodology.md
  • references/troubleshooting.md
  • scripts/benchmark.py
  • scripts/run_benchmark.sh
  • scripts/torch_air_infer.py

Open the folder on GitHubat commit 62a4ecb

Compare with similar skills

Npu Torchair Infer next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Npu Torchair Infer compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Npu Torchair Infer this skillascend-ai-coding/awesome-ascend-skills174—~1.3kAutomated safety check: PassNone
Qwen Mtp GgufR6410418/Jackrong-llm-finetuning-guide1.7k—~1.7kAutomated safety check: PassMIT
Veomni New ModelByteDance-Seed/VeOmni2.2k—~2kAutomated safety check: PassApache-2.0
Add Modelguoqingbao/xinfer334—~4.2kAutomated safety check: NotesMIT
Resolvealexziskind1/model-shelf130—~792Automated safety check: PassMIT
Check Modelguoqingbao/xinfer334—~3.8kAutomated safety check: PassMIT

Similar skills

  • Qwen Mtp Gguf

    R6410418/Jackrong-llm-finetuning-guide

    Complete agent-ready workflow for Qwen-family MTP or nextn GGUF conversion and release.

    1.7k GitHub stars~1.7k tokensUpdated 3 mo ago
    AI & LLM EngineeringAuto-check passed
  • Veomni New Model

    ByteDance-Seed/VeOmni

    A skill your agent uses when adding support for a new model to VeOmni.

    2.2k GitHub stars~2k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Add Model

    guoqingbao/xinfer

    Adapt and port new LLM model architectures to this xinfer project.

    334 GitHub stars~4.2k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check: notes
  • Resolve

    alexziskind1/model-shelf

    Always resolve Hugging Face models via model-shelf before any download.

    130 GitHub stars~792 tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Check Model

    guoqingbao/xinfer

    Check model compatibility with xinfer before loading. An agent skill from guoqingbao/xinfer.

    334 GitHub stars~3.8k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Test Model

    guoqingbao/xinfer

    Test LLM models served by xinfer for correctness, output quality, and performance.

    334 GitHub stars~2.6k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed

More from ascend-ai-coding/awesome-ascend-skills

All 70 skills in this repo
  • Ascend Dmi

    ascend-ai-coding/awesome-ascend-skills

    当用户需要对华为昇腾 NPU 进行硬件层面的管理、测试或诊断时使用此 skill。典型场景: - 查看 NPU 卡的状态、温度、利用率 - 测试内存带宽(h2d/d2h/d2d/p2p) - 跑算力/功耗基准测试(TFLOPS、TOPS) - 诊断 NPU 硬件故障或做健康检查 - 对 NPU 卡做压力测试(aicore、内存) - 复位/恢复卡住或异常的 NPU 卡 典型用户问题(即使不提…

    174 GitHub stars~1.7k tokensUpdated today
    Auto-check passed
  • Ascendc

    ascend-ai-coding/awesome-ascend-skills

    End-to-end AscendC custom operator development for Ascend NPU in an ascend-kernel (csrc/ops + build.sh + torchnpu PyTorch custom op) project.

    174 GitHub stars~3.5k tokensUpdated today
    Auto-check passed
  • Atc Model Converter

    ascend-ai-coding/awesome-ascend-skills

    Complete toolkit for Huawei Ascend NPU model conversion and end-to-end inference adaptation.

    174 GitHub stars~4.6k tokensUpdated today
    Auto-check passed
  • External Cannbot Ops Pypto Op Develop

    ascend-ai-coding/awesome-ascend-skills

    当需要编写 PyPTO 算子实现时使用此 skill。基于需求规格、设计方案和参考实现,生成完整可运行的 PyPTO 算子实现与配套测试、文档。Triggers: 实现算子、写 kernel、编写实现、写 impl、算子编码、开始编码、code the op、写 test、生成测试、写实现代码、op develop、kernel 实现。

    174 GitHub stars~2.1k tokensUpdated today
    Auto-check passed
  • External Gitcode Ascend Megatron Change Analyzer

    ascend-ai-coding/awesome-ascend-skills

    Analyze official Megatron-LM commits, PRs, and branch change sets to identify feature evolution, candidate breaking changes, and migration-relevant events.

    174 GitHub stars~1.2k tokensUpdated today
    Auto-check passed
  • External Gitcode Ascend Megatron Commit Tracker

    ascend-ai-coding/awesome-ascend-skills

    Track and normalize change requests against the official Megatron-LM repository by branch, PR, commit, commit range, or time window.

    174 GitHub stars~1.1k tokensUpdated today
    Auto-check passed

Questions about Npu Torchair Infer

What does Npu Torchair Infer do?

Migrate any HuggingFace model to Ascend NPU torchair graph mode (torch.compile) and benchmark it for accuracy and performance against NPU eager and CPU eager. Npu Torchair Infer is an agent skill from ascend-ai-coding/awesome-ascend-skills.compile) and benchmark it for accuracy and performance against NPU eager and CPU eager.

When should I use Npu Torchair Infer?

Npu Torchair Infer fits situations like: benchmarking HF models (vision; image-text encoders such as SigLIP2; SAM) on Ascend 910B/CANN with torchnpu and torchair; A torch.compile graph-mode run on NPU fails (Dynamo TorchRuntimeError.

How do I install Npu Torchair Infer in Claude Code?

Run `npx skills add ascend-ai-coding/awesome-ascend-skills --skill npu-torchair-infer -a claude-code`. Or copy the skill folder (skills/inference/npu-torchair-infer in ascend-ai-coding/awesome-ascend-skills) into .claude/skills/npu-torchair-infer in your project. Claude Code loads it when a task matches its description.

How do I install Npu Torchair Infer in Codex?

Run `npx skills add ascend-ai-coding/awesome-ascend-skills --skill npu-torchair-infer -a codex`. Or copy the skill folder (skills/inference/npu-torchair-infer in ascend-ai-coding/awesome-ascend-skills) into .agents/skills/npu-torchair-infer in your project. Codex loads it when a task matches its description.

Can I use Npu Torchair Infer in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ascend-ai-coding/awesome-ascend-skills --skill npu-torchair-infer -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/npu-torchair-infer, .gemini/skills/npu-torchair-infer, .github/skills/npu-torchair-infer and .opencode/skills/npu-torchair-infer in your project.

What does Npu Torchair Infer need to run?

Going by SKILL.md and its folder, Npu Torchair Infer needs Python and a shell for the scripts in its folder and the command-line tools its instructions call (python and bash). Our summary lists: Python 3; A Bash shell.

Does Npu Torchair Infer access the network?

SKILL.md names 1 domain. In commands or code: hf-mirror.com; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is Npu Torchair Infer safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Npu Torchair Infer use?

No licence was found for Npu Torchair Infer or its repository. Without one, default copyright applies: ask the author before reusing or redistributing it.

How many tokens does Npu Torchair Infer use?

About 1.3k tokens (SKILL.md is roughly 5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3.1k tokens, read only when the agent opens those files.

What are the alternatives to Npu Torchair Infer?

Skills that share tags, products or a category with Npu Torchair Infer: Qwen Mtp Gguf (R6410418/Jackrong-llm-finetuning-guide, 1.7k stars), Veomni New Model (ByteDance-Seed/VeOmni, 2.2k stars), Add Model (guoqingbao/xinfer, 334 stars) and Resolve (alexziskind1/model-shelf, 130 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Npu Torchair Infer?

ascend-ai-coding (a GitHub organization) maintains it in ascend-ai-coding/awesome-ascend-skills, which has 174 GitHub stars. The repository holds 70 skills in this directory. The repository was last updated on October 10, 2026.

Source: ascend-ai-coding/awesome-ascend-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.