Search
By ascend-ai-coding
Skills
Sort:BestMost starsTrending todayTrending this weekTrending this monthNewestRecently updatedName
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 1 | 当用户需要对华为昇腾 NPU 进行硬件层面的管理、测试或诊断时使用此 skill。典型场景: - 查看 NPU 卡的状态、温度、利用率 - 测试内存带宽(h2d/d2h/d2d/p2p) - 跑算力/功耗基准测试(TFLOPS、TOPS) - 诊断 NPU 硬件故障或做健康检查 - 对 NPU 卡做压力测试(aicore、内存) - 复位/恢复卡住或异常的 NPU 卡 典型用户问题(即使不提… | ascend-ai-coding/ | 174 | — | ~1.7k | Automated safety check: Pass | No licence | today |
| 2 | 2.Ascendc End-to-end AscendC custom operator development for Ascend NPU in an ascend-kernel (csrc/ops + build.sh + torchnpu PyTorch custom op) project. | ascend-ai-coding/ | 174 | — | ~3.5k | Automated safety check: Pass | No licence | today |
| 3 | Complete toolkit for Huawei Ascend NPU model conversion and end-to-end inference adaptation. | ascend-ai-coding/ | 174 | — | ~4.6k | Automated safety check: Pass | No licence | today |
| 4 | 当需要编写 PyPTO 算子实现时使用此 skill。基于需求规格、设计方案和参考实现,生成完整可运行的 PyPTO 算子实现与配套测试、文档。Triggers: 实现算子、写 kernel、编写实现、写 impl、算子编码、开始编码、code the op、写 test、生成测试、写实现代码、op develop、kernel 实现。 | ascend-ai-coding/ | 174 | — | ~2.1k | Automated safety check: Pass | No licence | today |
| 5 | Analyze official Megatron-LM commits, PRs, and branch change sets to identify feature evolution, candidate breaking changes, and migration-relevant events. | ascend-ai-coding/ | 174 | — | ~1.2k | Automated safety check: Pass | No licence | today |
| 6 | Track and normalize change requests against the official Megatron-LM repository by branch, PR, commit, commit range, or time window. | ascend-ai-coding/ | 174 | — | ~1.1k | Automated safety check: Pass | No licence | today |
| 7 | Map migration-relevant Megatron changes onto the official MindSpeed repository by resolving branch alignment, locating affected subsystems, and identifying concrete adaptation points. | ascend-ai-coding/ | 174 | — | ~1.3k | Automated safety check: Pass | No licence | today |
| 8 | Automates GitCode open-source repo merge workflow: commit → push → issue → PR → pipeline → review → /lgtm & /approve → merge. | ascend-ai-coding/ | 174 | — | ~1.4k | Automated safety check: Pass | No licence | today |
| 9 | Analyze closed GitHub issues to create troubleshooting case studies with root cause analysis and lessons learned. | ascend-ai-coding/ | 174 | — | ~1.4k | Automated safety check: Pass | No licence | today |
| 10 | Automatically fetch InferenceX benchmark data and generate daily performance reports for LLM inference on various hardware (NVIDIA, AMD, etc.). | ascend-ai-coding/ | 174 | — | ~1.1k | Automated safety check: Pass | No licence | today |
| 11 | 11.Npu Smi Huawei Ascend NPU npu-smi command reference. An agent skill from ascend-ai-coding/awesome-ascend-skills. | ascend-ai-coding/ | 174 | — | ~2k | Automated safety check: Pass | No licence | today |
| 12 | 12.Create PR Creates GitHub pull requests with properly formatted titles that pass the check-pr-title CI validation. | ascend-ai-coding/ | 174 | 1 repo | ~1.2k | Automated safety check: Pass | No licence | today |
| 13 | Migrate any HuggingFace model to Ascend NPU torchair graph mode (torch.compile) and benchmark it for accuracy and performance against NPU eager and CPU eager. | ascend-ai-coding/ | 174 | — | ~1.3k | Automated safety check: Pass | No licence | today |
| 14 | 通过 PyTorch torch.distributed 接口测试昇腾 NPU 通信算子性能。支持指定任意 tensor shape、dtype,使用 torchrun 启动,贴近真实训练场景的通信算子测试与性能分析。Use for testing collective communication operators (AllReduce, AllGather, ReduceScatter… | ascend-ai-coding/ | 174 | — | ~2.2k | Automated safety check: Pass | No licence | today |
| 15 | 15.Vllm Ascend vLLM Ascend plugin for LLM inference serving on Huawei Ascend NPU. | ascend-ai-coding/ | 174 | — | ~2.7k | Automated safety check: Pass | No licence | today |
| 16 | Track daily PRs and Issues from vllm-project/vllm and vllm-project/vllm-ascend, filter by model (DeepSeek/Qwen/GLM/MiniMax/Kimi) and tech topics (PD disaggregation, MTP, quantization, graph mode… | ascend-ai-coding/ | 174 | — | ~731 | Automated safety check: Pass | No licence | today |
| 17 | 17.Hccl Test HCCL (Huawei Collective Communication Library) performance testing for Ascend NPU clusters. | ascend-ai-coding/ | 174 | — | ~2.2k | Automated safety check: Warn | No licence | today |
| 18 | Interactive online benchmark orchestrator for vLLM inference services using vllm bench serve. | ascend-ai-coding/ | 174 | — | ~5.6k | Automated safety check: Pass | No licence | today |
| 19 | 19.Ais Bench AISBench Benchmark - AI model evaluation tool for Ascend NPU. | ascend-ai-coding/ | 174 | — | ~2.7k | Automated safety check: Pass | No licence | today |
| 20 | Create Docker containers for Huawei Ascend NPU development with proper device mappings and volume mounts. | ascend-ai-coding/ | 174 | — | ~1.3k | Automated safety check: Pass | No licence | today |
| 21 | Analyze Huawei Ascend NPU profiling data to discover hidden performance anomalies and produce a detailed model architecture report reverse-engineered from profiling. | ascend-ai-coding/ | 174 | — | ~5.8k | Automated safety check: Pass | No licence | today |
| 22 | Verl 分布式训练服务一键拉起与配置。触发场景:(1) 用户要启动 Verl 训练任务或部署 RLHF/DAPO 训练环境 (2) 在 NPU 集群上拉起 Verl 训练容器 (3) 配置 Ray 集群和 SwanLab 监控 (4) 根据 7 位二进制掩码灵活配置加速特性。支持 Qwen3-8B 等 Megatron 模型的 DAPO 训练全流程。 | ascend-ai-coding/ | 174 | — | ~2k | Automated safety check: Pass | No licence | today |
| 23 | 昇腾 NPU 单算子性能基准测试 Skill;当前版本只做现有环境检查、CANN 版本识别、用户确认后执行 benchmark,不负责修复或安装环境。 | ascend-ai-coding/ | 174 | — | ~617 | Automated safety check: Pass | No licence | today |
| 24 | Skill for analyzing communication performance bottlenecks and detecting slow/fast rank issues in Ascend NPU systems. | ascend-ai-coding/ | 174 | — | ~590 | Automated safety check: Pass | No licence | today |
| 25 | 25.Rl Msprobe 自动化 verl msprobe 精度数据采集;开始前检查/预装 msprobe(pip install mindstudio-probe)。自动识别三种模式:(1) 训练采集——globalprofiler + precisiondebugger stages;(2) 推理采集——vLLM/SGLang rollout dump;(3) 训推一致性——engine patch +… | ascend-ai-coding/ | 174 | — | ~2.4k | Automated safety check: Pass | No licence | today |
| 26 | AI for Science 场景下的昇腾 NPU Profiling 采集与性能分析 Skill,用于在华为 Ascend NPU 上使用 torchnpu.profiler 采集 L0、L1、L2 级性能数据,分析训练或推理中的算子耗时、调用栈、内存与瓶颈,并指导后续调优。 | ascend-ai-coding/ | 174 | — | ~3k | Automated safety check: Pass | No licence | today |
| 27 | Ankh 蛋白质语言模型昇腾 NPU 迁移 Skill,适用于 Ankh base/large、Ankh3 large/XL 以及同类基于 HuggingFace Transformers 与 PyTorch 的蛋白模型从 CUDA/GPU 到华为 Ascend NPU 的环境检查、代码适配、权重加载、验证脚本补齐与文档沉淀。 | ascend-ai-coding/ | 174 | — | ~2.1k | Automated safety check: Pass | No licence | today |
| 28 | 昇腾 TensorFlow Community 迁移适配 Skill,适用于将基于 TensorFlow 2.x 的模型原生部署到华为 Ascend NPU,而不经过 TF 到 PyTorch 转换,覆盖 aarch64 源码编译 TF 2.6.5、tfplugin 安装、自动迁移工具使用、手动适配与精度验证。 | ascend-ai-coding/ | 174 | — | ~2.1k | Automated safety check: Pass | No licence | today |
| 29 | Boltz2 蛋白质结构预测模型的昇腾 NPU 迁移与复现 Skill,适用于在华为 Ascend 910、910B、910C 上准备权重、适配 Lightning 和 CUDA only kernel、完成 Boltz2 端到端结构预测推理,并沉淀可复现的环境与验证命令。 | ascend-ai-coding/ | 174 | — | ~2.8k | Automated safety check: Pass | No licence | today |
| 30 | BoltzGen 昇腾 NPU 迁移与复现 Skill,适用于在华为 Ascend NPU 上部署 BoltzGen 生成式蛋白设计与逆折叠流程,覆盖环境准备、权重缓存、cuEquivariance 兼容、源码适配和端到端推理验证。 | ascend-ai-coding/ | 174 | — | ~3.8k | Automated safety check: Pass | No licence | today |
| 31 | DeepFRI 的 TensorFlow 到 PyTorch 转换与昇腾 NPU 迁移 Skill,适用于蛋白质功能预测场景下的 TF 模型分析、PyTorch 重写、权重逐层映射、NPU 推理与精度验证,尤其适合需要在 Ascend 上运行 DeepFRI CNN 或 GCN 路径时使用。 | ascend-ai-coding/ | 174 | — | ~2.4k | Automated safety check: Pass | No licence | today |
| 32 | DeepFRI TensorFlow 原生昇腾 NPU 迁移 Skill,适用于不做 TF 到 PyTorch 转换、而是直接使用 TensorFlow 2.6.5 与 npudevice 在华为 Ascend 上运行 DeepFRI 的场景,覆盖源码编译、tfplugin 安装、代码适配、推理与 CPU 对比验证。 | ascend-ai-coding/ | 174 | — | ~2.1k | Automated safety check: Pass | No licence | today |
| 33 | DiffSBDD 昇腾 NPU 迁移 Skill,适用于将基于等变扩散模型的结构化药物设计项目从 CUDA 迁移到华为 Ascend NPU,覆盖环境搭建、依赖安装、torchscatter 源码编译、代码适配以及 de novo 推理验证。 | ascend-ai-coding/ | 174 | — | ~898 | Automated safety check: Pass | No licence | today |
| 34 | GENERator DNA 序列生成模型的昇腾 NPU 迁移 Skill,适用于将基于 HuggingFace Transformers 的 Causal LM 从 CUDA 迁移到华为 Ascend NPU,覆盖环境搭建、依赖安装、代码适配、多进程处理和 sequence recovery 验证。 | ascend-ai-coding/ | 174 | — | ~827 | Automated safety check: Pass | No licence | today |
| 35 | OligoFormer 昇腾 NPU 迁移 Skill,适用于将基于 PyTorch Transformer 的 siRNA 效能预测模型迁移到华为 Ascend NPU,覆盖环境搭建、RNA-FM 依赖安装、代码适配、推理验证与可选训练流程。 | ascend-ai-coding/ | 174 | — | ~1.3k | Automated safety check: Pass | No licence | today |
| 36 | ProteinBERT 昇腾 NPU 部署与迁移 Skill,适用于将 TensorFlow 或 Keras 版 ProteinBERT 转成基于 PyTorch 与 torchnpu 的实现,覆盖权重转换、embedding 提取、微调训练、注意力可视化和 GPU 与 NPU 精度验证。 | ascend-ai-coding/ | 174 | — | ~1.9k | Automated safety check: Pass | No licence | today |
| 37 | TensorFlow 或 Keras 模型改写到 PyTorch 的通用 Skill,适用于在华为 Ascend NPU 或其他依赖 PyTorch 生态的平台上完成层级映射、权重转换、逐层数值验证和端到端精度对比,尤其适合 ProteinBERT、DeepFRI 这类科学模型的跨框架迁移。 | ascend-ai-coding/ | 174 | — | ~3.4k | Automated safety check: Pass | No licence | today |
| 38 | HuggingFace Diffusers 环境配置指南,用于华为昇腾 NPU。覆盖 CANN 版本检测、PyTorch + torchnpu 安装、Diffusers 库安装及环境验证。当用户需要在昇腾 NPU 上配置 Diffusers 环境时使用。 | ascend-ai-coding/ | 174 | — | ~566 | Automated safety check: Pass | No licence | today |
| 39 | Diffusers Pipeline 推理指南,用于华为昇腾 NPU。覆盖环境预检、通用 Pipeline 推理(图像/视频模型)、内存优化(CPU offload、attention slicing、VAE slicing)、LoRA 加载与融合、多卡推理和按版本检索 Diffusers API。用户一旦提到在昇腾 NPU 上运行 FLUX、SDXL、Wan、CogVideoX 等… | ascend-ai-coding/ | 174 | — | ~3.2k | Automated safety check: Pass | No licence | today |
| 40 | 基于 PyTorch 框架的昇腾 NPU 模型推理融合算子优化技能。分析模型代码,识别可替换为 torchnpu 融合算子的计算模式,生成替换方案。触发场景:torchnpu 融合算子替换、MoE/Attention/FFN/Norm 等模块的推理算子适配、torchnpu API 使用咨询。基于仓库已有模型的融合算子经验,按计算语义推荐最佳方案。 | ascend-ai-coding/ | 174 | — | ~1.7k | Automated safety check: Pass | No licence | today |
| 41 | 基于 PyTorch 框架的昇腾 NPU 模型推理适配与部署基线技能。从 HF 链接或本地模型代码出发,按 cann-recipes-infer 仓库规范适配到 ModelRunner 推理框架,输出可运行的标准模型目录和性能基线数据。触发场景:新模型适配到昇腾 NPU 推理框架、已有模型的部署基线采集、模型迁移和初始跑通验证。 | ascend-ai-coding/ | 174 | — | ~1.8k | Automated safety check: Pass | No licence | today |
| 42 | Ascend C 算子卡死/崩溃调试技能。用于处理程序无法运行完的场景:(1) 程序卡死/挂起/超时,Kernel 无响应,(2) 程序崩溃(Segmentation Fault、Abort),(3) Buffer 冲突/死锁导致的核心挂起,(4) 需要解析 plog 日志定位卡死/崩溃位置。触发关键词:卡死、挂起、超时、崩溃、hang、crash、deadlock、Segmentation… | ascend-ai-coding/ | 174 | — | ~502 | Automated safety check: Pass | No licence | today |
| 43 | Ascend C 开发资源检索技能。通过本地 API 文档索引、示例代码映射和在线文档兜底搜索定位开发资料,优先查本地、缺失时再查在线。当需要查询 API 用法、示例代码、兼容性信息、官方资料入口或定位文档来源时使用。 | ascend-ai-coding/ | 174 | — | ~989 | Automated safety check: Pass | No licence | today |
| 44 | Ascend C 算子精度调试技能,提供精度问题诊断和解决方法。触发:输出异常(全为0、随机值、未初始化)、精度验证失败(rtol/atol 不达标)、FP16 精度差于预期、Cast 后数据错误、需要排查流水线同步(EnQue/DeQue)或 DataCopy 对齐问题。 | ascend-ai-coding/ | 174 | — | ~2.2k | Automated safety check: Pass | No licence | today |
| 45 | Ascend C 算子运行时错误调试技能。用于处理算子运行时问题:(1) aclnn 返回错误码(161xxx/361xxx/561xxx,包括环境配置、Tiling、Kernel 查找等错误),(2) 需要解析 plog 日志定位问题。触发关键词:运行时错误、错误码、Tiling错误、Kernel查找失败、环境变量、plog。 | ascend-ai-coding/ | 174 | — | ~528 | Automated safety check: Pass | No licence | today |
| 46 | NPU 性能采集与分析,用于采集算子性能数据、定位性能瓶颈并给出优化建议。当用户在算子开发过程中提到"上板性能"、"算子性能测试"、"硬件性能验证"、"NPU性能采集"、"NPU profiling"等场景时触发。 | ascend-ai-coding/ | 174 | — | ~2k | Automated safety check: Pass | No licence | today |
| 47 | 当需要设计 PyPTO 算子实现方案时使用此 skill。基于算子规格与相关上下文,生成 DESIGN.md(含 API 映射、Tiling 策略、Loop 结构)。Triggers: 生成设计方案、生成 design、设计这个算子、写 DESIGN.md、算子设计、API 映射、Tiling 策略、tiling strategy、Loop 结构、数据切分、怎么切分数据、怎么做… | ascend-ai-coding/ | 174 | — | ~2.4k | Automated safety check: Pass | No licence | today |
| 48 | 从用户 PyTorch/Python 代码中提取算子实现,构建为算子任务格式的标准化 任务文件。支持两种模式:单 case(单一自包含 .py,getinputs 返回单组)和 多 case(.py + 同名 .json 配对,getinputgroups 返回多组)。 | ascend-ai-coding/ | 174 | — | ~1.4k | Automated safety check: Pass | No licence | today |