Search
AI & LLM Engineering · ascend-ai-coding/awesome-ascend-skills
Skills
Sort:BestMost starsTrending todayTrending this weekTrending this monthNewestRecently updatedName
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 1 | 1.Ascendc End-to-end AscendC custom operator development for Ascend NPU in an ascend-kernel (csrc/ops + build.sh + torchnpu PyTorch custom op) project. | ascend-ai-coding/ | 174 | — | ~3.5k | Automated safety check: Pass | No licence | yesterday |
| 2 | Complete toolkit for Huawei Ascend NPU model conversion and end-to-end inference adaptation. | ascend-ai-coding/ | 174 | — | ~4.6k | Automated safety check: Pass | No licence | yesterday |
| 3 | Analyze official Megatron-LM commits, PRs, and branch change sets to identify feature evolution, candidate breaking changes, and migration-relevant events. | ascend-ai-coding/ | 174 | — | ~1.2k | Automated safety check: Pass | No licence | yesterday |
| 4 | Track and normalize change requests against the official Megatron-LM repository by branch, PR, commit, commit range, or time window. | ascend-ai-coding/ | 174 | — | ~1.1k | Automated safety check: Pass | No licence | yesterday |
| 5 | Map migration-relevant Megatron changes onto the official MindSpeed repository by resolving branch alignment, locating affected subsystems, and identifying concrete adaptation points. | ascend-ai-coding/ | 174 | — | ~1.3k | Automated safety check: Pass | No licence | yesterday |
| 6 | Automatically fetch InferenceX benchmark data and generate daily performance reports for LLM inference on various hardware (NVIDIA, AMD, etc.). | ascend-ai-coding/ | 174 | — | ~1.1k | Automated safety check: Pass | No licence | yesterday |
| 7 | Migrate any HuggingFace model to Ascend NPU torchair graph mode (torch.compile) and benchmark it for accuracy and performance against NPU eager and CPU eager. | ascend-ai-coding/ | 174 | — | ~1.3k | Automated safety check: Pass | No licence | yesterday |
| 8 | 通过 PyTorch torch.distributed 接口测试昇腾 NPU 通信算子性能。支持指定任意 tensor shape、dtype,使用 torchrun 启动,贴近真实训练场景的通信算子测试与性能分析。Use for testing collective communication operators (AllReduce, AllGather, ReduceScatter… | ascend-ai-coding/ | 174 | — | ~2.2k | Automated safety check: Pass | No licence | yesterday |
| 9 | vLLM Ascend plugin for LLM inference serving on Huawei Ascend NPU. | ascend-ai-coding/ | 174 | — | ~2.7k | Automated safety check: Pass | No licence | yesterday |
| 10 | Track daily PRs and Issues from vllm-project/vllm and vllm-project/vllm-ascend, filter by model (DeepSeek/Qwen/GLM/MiniMax/Kimi) and tech topics (PD disaggregation, MTP, quantization, graph mode… | ascend-ai-coding/ | 174 | — | ~731 | Automated safety check: Pass | No licence | yesterday |
| 11 | Interactive online benchmark orchestrator for vLLM inference services using vllm bench serve. | ascend-ai-coding/ | 174 | — | ~5.6k | Automated safety check: Pass | No licence | yesterday |
| 12 | 12.Ais Bench AISBench Benchmark - AI model evaluation tool for Ascend NPU. | ascend-ai-coding/ | 174 | — | ~2.7k | Automated safety check: Pass | No licence | yesterday |
| 13 | Verl 分布式训练服务一键拉起与配置。触发场景:(1) 用户要启动 Verl 训练任务或部署 RLHF/DAPO 训练环境 (2) 在 NPU 集群上拉起 Verl 训练容器 (3) 配置 Ray 集群和 SwanLab 监控 (4) 根据 7 位二进制掩码灵活配置加速特性。支持 Qwen3-8B 等 Megatron 模型的 DAPO 训练全流程。 | ascend-ai-coding/ | 174 | — | ~2k | Automated safety check: Pass | No licence | yesterday |
| 14 | 14.Rl Msprobe 自动化 verl msprobe 精度数据采集;开始前检查/预装 msprobe(pip install mindstudio-probe)。自动识别三种模式:(1) 训练采集——globalprofiler + precisiondebugger stages;(2) 推理采集——vLLM/SGLang rollout dump;(3) 训推一致性——engine patch +… | ascend-ai-coding/ | 174 | — | ~2.4k | Automated safety check: Pass | No licence | yesterday |
| 15 | Ankh 蛋白质语言模型昇腾 NPU 迁移 Skill,适用于 Ankh base/large、Ankh3 large/XL 以及同类基于 HuggingFace Transformers 与 PyTorch 的蛋白模型从 CUDA/GPU 到华为 Ascend NPU 的环境检查、代码适配、权重加载、验证脚本补齐与文档沉淀。 | ascend-ai-coding/ | 174 | — | ~2.1k | Automated safety check: Pass | No licence | yesterday |
| 16 | 昇腾 TensorFlow Community 迁移适配 Skill,适用于将基于 TensorFlow 2.x 的模型原生部署到华为 Ascend NPU,而不经过 TF 到 PyTorch 转换,覆盖 aarch64 源码编译 TF 2.6.5、tfplugin 安装、自动迁移工具使用、手动适配与精度验证。 | ascend-ai-coding/ | 174 | — | ~2.1k | Automated safety check: Pass | No licence | yesterday |
| 17 | Boltz2 蛋白质结构预测模型的昇腾 NPU 迁移与复现 Skill,适用于在华为 Ascend 910、910B、910C 上准备权重、适配 Lightning 和 CUDA only kernel、完成 Boltz2 端到端结构预测推理,并沉淀可复现的环境与验证命令。 | ascend-ai-coding/ | 174 | — | ~2.8k | Automated safety check: Pass | No licence | yesterday |
| 18 | BoltzGen 昇腾 NPU 迁移与复现 Skill,适用于在华为 Ascend NPU 上部署 BoltzGen 生成式蛋白设计与逆折叠流程,覆盖环境准备、权重缓存、cuEquivariance 兼容、源码适配和端到端推理验证。 | ascend-ai-coding/ | 174 | — | ~3.8k | Automated safety check: Pass | No licence | yesterday |
| 19 | DeepFRI 的 TensorFlow 到 PyTorch 转换与昇腾 NPU 迁移 Skill,适用于蛋白质功能预测场景下的 TF 模型分析、PyTorch 重写、权重逐层映射、NPU 推理与精度验证,尤其适合需要在 Ascend 上运行 DeepFRI CNN 或 GCN 路径时使用。 | ascend-ai-coding/ | 174 | — | ~2.4k | Automated safety check: Pass | No licence | yesterday |
| 20 | DeepFRI TensorFlow 原生昇腾 NPU 迁移 Skill,适用于不做 TF 到 PyTorch 转换、而是直接使用 TensorFlow 2.6.5 与 npudevice 在华为 Ascend 上运行 DeepFRI 的场景,覆盖源码编译、tfplugin 安装、代码适配、推理与 CPU 对比验证。 | ascend-ai-coding/ | 174 | — | ~2.1k | Automated safety check: Pass | No licence | yesterday |
| 21 | DiffSBDD 昇腾 NPU 迁移 Skill,适用于将基于等变扩散模型的结构化药物设计项目从 CUDA 迁移到华为 Ascend NPU,覆盖环境搭建、依赖安装、torchscatter 源码编译、代码适配以及 de novo 推理验证。 | ascend-ai-coding/ | 174 | — | ~898 | Automated safety check: Pass | No licence | yesterday |
| 22 | GENERator DNA 序列生成模型的昇腾 NPU 迁移 Skill,适用于将基于 HuggingFace Transformers 的 Causal LM 从 CUDA 迁移到华为 Ascend NPU,覆盖环境搭建、依赖安装、代码适配、多进程处理和 sequence recovery 验证。 | ascend-ai-coding/ | 174 | — | ~827 | Automated safety check: Pass | No licence | yesterday |
| 23 | OligoFormer 昇腾 NPU 迁移 Skill,适用于将基于 PyTorch Transformer 的 siRNA 效能预测模型迁移到华为 Ascend NPU,覆盖环境搭建、RNA-FM 依赖安装、代码适配、推理验证与可选训练流程。 | ascend-ai-coding/ | 174 | — | ~1.3k | Automated safety check: Pass | No licence | yesterday |
| 24 | ProteinBERT 昇腾 NPU 部署与迁移 Skill,适用于将 TensorFlow 或 Keras 版 ProteinBERT 转成基于 PyTorch 与 torchnpu 的实现,覆盖权重转换、embedding 提取、微调训练、注意力可视化和 GPU 与 NPU 精度验证。 | ascend-ai-coding/ | 174 | — | ~1.9k | Automated safety check: Pass | No licence | yesterday |
| 25 | TensorFlow 或 Keras 模型改写到 PyTorch 的通用 Skill,适用于在华为 Ascend NPU 或其他依赖 PyTorch 生态的平台上完成层级映射、权重转换、逐层数值验证和端到端精度对比,尤其适合 ProteinBERT、DeepFRI 这类科学模型的跨框架迁移。 | ascend-ai-coding/ | 174 | — | ~3.4k | Automated safety check: Pass | No licence | yesterday |
| 26 | HuggingFace Diffusers 环境配置指南,用于华为昇腾 NPU。覆盖 CANN 版本检测、PyTorch + torchnpu 安装、Diffusers 库安装及环境验证。当用户需要在昇腾 NPU 上配置 Diffusers 环境时使用。 | ascend-ai-coding/ | 174 | — | ~566 | Automated safety check: Pass | No licence | yesterday |
| 27 | Diffusers Pipeline 推理指南,用于华为昇腾 NPU。覆盖环境预检、通用 Pipeline 推理(图像/视频模型)、内存优化(CPU offload、attention slicing、VAE slicing)、LoRA 加载与融合、多卡推理和按版本检索 Diffusers API。用户一旦提到在昇腾 NPU 上运行 FLUX、SDXL、Wan、CogVideoX 等… | ascend-ai-coding/ | 174 | — | ~3.2k | Automated safety check: Pass | No licence | yesterday |
| 28 | 基于 PyTorch 框架的昇腾 NPU 模型推理融合算子优化技能。分析模型代码,识别可替换为 torchnpu 融合算子的计算模式,生成替换方案。触发场景:torchnpu 融合算子替换、MoE/Attention/FFN/Norm 等模块的推理算子适配、torchnpu API 使用咨询。基于仓库已有模型的融合算子经验,按计算语义推荐最佳方案。 | ascend-ai-coding/ | 174 | — | ~1.7k | Automated safety check: Pass | No licence | yesterday |
| 29 | 基于 PyTorch 框架的昇腾 NPU 模型推理适配与部署基线技能。从 HF 链接或本地模型代码出发,按 cann-recipes-infer 仓库规范适配到 ModelRunner 推理框架,输出可运行的标准模型目录和性能基线数据。触发场景:新模型适配到昇腾 NPU 推理框架、已有模型的部署基线采集、模型迁移和初始跑通验证。 | ascend-ai-coding/ | 174 | — | ~1.8k | Automated safety check: Pass | No licence | yesterday |
| 30 | 从用户 PyTorch/Python 代码中提取算子实现,构建为算子任务格式的标准化 任务文件。支持两种模式:单 case(单一自包含 .py,getinputs 返回单组)和 多 case(.py + 同名 .json 配对,getinputgroups 返回多组)。 | ascend-ai-coding/ | 174 | — | ~1.4k | Automated safety check: Pass | No licence | yesterday |
| 31 | 能完成昇腾NPU驱动和固件安装部署,实现安装包正则匹配提取、按需添加可执行权限、Python+Shell双重包校验、系统依赖先验后装、适配CentOS/RHEL/Ubuntu/Debian系统,适用于昇腾NPU驱动和固件安装部署。 | ascend-ai-coding/ | 174 | — | ~663 | Automated safety check: Notes | No licence | yesterday |
| 32 | 根据设计文档生成 AscendC 算子完整代码实现并完成框架适配。TRIGGER when: 设计文档已完成,需要生成 ophost/opkernel 代码、注册到 PyTorch 框架、编译测试。关键词:代码生成、ophost、opkernel、tiling、kernel、框架适配、算子注册。 | ascend-ai-coding/ | 174 | — | ~2k | Automated safety check: Pass | No licence | yesterday |
| 33 | Ascend C 算子 mssanitizer 内存检测分析技能。用于检测和分析算子内存问题:非法内存访问、非法释放、内存泄漏、UB地址越界,生成问题报告。自动识别算子工程类型(ops算子仓用GE IR模式,自定义算子用Python模式)。触发关键词:mssanitizer、内存检测、内存泄漏、非法访问、illegal free、内存错误。 | ascend-ai-coding/ | 174 | — | ~3.4k | Automated safety check: Pass | No licence | yesterday |
| 34 | 自动生成 ATB 到 ACLNN 算子替换的详细设计文档。接收用户提供的 ATB 和 ACLNN 接口文档链接, 输出包含参数映射、开发自测、风险评估的 7 章结构化设计文档。 | ascend-ai-coding/ | 174 | — | ~2.4k | Automated safety check: Pass | No licence | yesterday |
| 35 | Generate migration deliverables for bringing relevant Megatron changes into MindSpeed after branch alignment and impact mapping are complete. | ascend-ai-coding/ | 174 | — | ~1.3k | Automated safety check: Pass | No licence | yesterday |
| 36 | Verl 单异步 DAPO 训练配置生成器。触发场景:(1) 启动单异步 DAPO 训练 (2) 生成训练脚本 (3) 配置特性参数 (4) 训练前检查。特性策略:用户未指定时默认开启性能特性(flashattn/dynamicbatch/removepadding/gradientcheckpointing),显存特性(offload/recompute)默认关闭。OOM… | ascend-ai-coding/ | 174 | — | ~1.2k | Automated safety check: Pass | No licence | yesterday |
| 37 | 昇腾 NPU 平台 vLLM 大模型推理服务一键部署。触发:用户说'部署 模型名'、'NPU 部署模型'、'vllm serve'。流程:SSH检查 → NPU检查 → 配置发现(必须验证) → 用户确认 → 部署 → cron监控 → 验证。约束:(1) 配置必须从官方文档验证,禁止猜测;(2) 后台启动必须用cron监控,禁止手动轮询。支持… | ascend-ai-coding/ | 174 | — | ~1.2k | Automated safety check: Pass | No licence | yesterday |
| 38 | 面向 Ascend PyTorch Profiler / msprof DB(如 ascendpytorchprofiler.db、msprof.db)的 SQL 分析技能。将自然语言问题(算子耗时、通信、下发、调度、schema/table 查询)转为安全可执行 SQL,并按需从官方文档提取表结构详情。 | ascend-ai-coding/ | 174 | — | ~1.4k | Automated safety check: Pass | No licence | yesterday |
| 39 | 模型层 Tensor 打点与精度对比工具。用于在模型 forward 过程中捕获模型各层中间 tensor,实现 GPU/NPU 精度对比调试。支持 vLLM、SGLang 推理框架。When to use: When you need to debug precision issues between GPU and NPU,or validate layer-wise tensor… | ascend-ai-coding/ | 174 | — | ~2.4k | Automated safety check: Pass | No licence | yesterday |
| 40 | MindSpeed-MM multimodal model suite environment setup guide for Huawei Ascend NPU. | ascend-ai-coding/ | 174 | — | ~3.1k | Automated safety check: Pass | No licence | yesterday |
| 41 | MindSpeed-MM skill router and model index for Huawei Ascend NPU. | ascend-ai-coding/ | 174 | — | ~3.5k | Automated safety check: Pass | No licence | yesterday |
| 42 | Universal VLM (vision-language understanding model) training guide for Huawei Ascend NPU using MindSpeed-MM. | ascend-ai-coding/ | 174 | — | ~5.2k | Automated safety check: Pass | No licence | yesterday |
| 43 | MindSpeed-MM weight conversion guide using mm-convert CLI tool. | ascend-ai-coding/ | 174 | — | ~2.1k | Automated safety check: Pass | No licence | yesterday |
| 44 | Generates an executable, end-to-end VERL reinforcement learning quickstart runbook for Ascend/NPU (docker image, dataset preprocessing, model setup, mainppo training, and examples/run.sh flow). | ascend-ai-coding/ | 174 | — | ~592 | Automated safety check: Pass | No licence | yesterday |
| 45 | Deploy vLLM inference services on Ascend NPU servers with automatic model detection and optimized configuration. | ascend-ai-coding/ | 174 | — | ~2.7k | Automated safety check: Pass | No licence | yesterday |