Search

CUDA

282 skills found, page 6.
Search results
#SkillRepositoryStarsUsed inTokensAuto-checkLicenceUpdated
241

Operational guide for enabling Megatron FSDP in Megatron-Bridge, including config knobs, code anchors, pitfalls, and verification.

NVIDIA/skills3.6k—~973Automated safety check: PassApache-2.0yesterday
242

Techniques for reducing peak GPU memory in Megatron Bridge — expandable segments, PEFT + SP input re-gather, parallelism resizing, activation recompute, CPU offloading constraints, and common OOM…

NVIDIA/skills3.6k—~3.6kAutomated safety check: PassApache-2.0yesterday
243

MoE expert-parallel communication overlap in Megatron Bridge.

NVIDIA/skills3.6k—~1.9kAutomated safety check: PassApache-2.0yesterday
244

Representative, point-in-time MoE training playbooks by hardware and model family.

NVIDIA/skills3.6k—~2kAutomated safety check: PassApache-2.0yesterday
245

Long-context MoE training guidance for Megatron Bridge. An agent skill from NVIDIA/skills.

NVIDIA/skills3.6k—~1.2kAutomated safety check: PassApache-2.0yesterday
246

Evidence-gated workflow for MoE performance optimization in Megatron Bridge.

NVIDIA/skills3.6k—~3.3kAutomated safety check: PassApache-2.0yesterday
247

Practical guidance for training MoE VLMs in Megatron Bridge.

NVIDIA/skills3.6k—~1.3kAutomated safety check: PassApache-2.0yesterday
248

Operational guide for enabling TP, DP, and PP communication overlap in Megatron-Bridge, including config knobs, code anchors, pitfalls, and verification.

NVIDIA/skills3.6k—~924Automated safety check: PassApache-2.0yesterday
249

Optimize transformer attention with Flash Attention — 2-4x speedup, 10-20x memory reduction for long sequences on CUDA GPUs.

AlexAI-MCP/hermes-CCC135—~1.2kAutomated safety check: PassMIT6 mo ago
250

Diagnose and fix CoreWeave GPU scheduling, pod, and networking errors.

jeremylongshore/tons-of-skills-marketplace2.8k—~1.1kAutomated safety check: PassMITyesterday
251

Upgrade CoreWeave deployments and migrate between GPU types.

jeremylongshore/tons-of-skills-marketplace2.8k—~1.9kAutomated safety check: PassMITyesterday
252

GPU optimization for consumer NVIDIA GPUs (8-24GB VRAM) covering mixed precision, gradient checkpointing, XGBoost GPU, CuPy/cuDF migration, and torch.compile.

Mathews-Tom/armory329—~3.5kAutomated safety check: NotesMIT5 days ago
253

Use this repo skill for DiscoArt image generation, configuration/prompt scheduling, CLI, Jina serving, Docker runtime planning, and troubleshooting.

VectorSpaceLab/AREX-Skill331—~1.3kAutomated safety check: PassUnknown1 mo ago
254

Use this operating skill for Tencent-Hunyuan/HunyuanVideo text-to-video setup, checkpoint layout, inference commands, Gradio launch, and CUDA/FP8/xDiT troubleshooting.

VectorSpaceLab/AREX-Skill331—~1.1kAutomated safety check: PassUnknown1 mo ago
255

Use this repo skill for Make-It-3D single-image 3D creation, including CUDA asset setup, alpha-image validation, coarse NeRF optimization, refinement, rendering, export, and troubleshooting.

VectorSpaceLab/AREX-Skill331—~1.4kAutomated safety check: PassApache-2.01 mo ago
256

Use this quip-miner sub-skill for CPU, CUDA, Metal, Modal, and QPU mining commands, backend dependencies, QPU budgets, unified streaming, and PoW/mempool scheduling.

VectorSpaceLab/AREX-Skill331—~753Automated safety check: PassAGPL-3.01 mo ago
257

A skill your agent uses for Motional nuPlan autonomous-driving planning workflows: dataset and map access, scenario filtering, planner implementation, open- or closed-loop simulation, metrics…

VectorSpaceLab/AREX-Skill331—~1.2kAutomated safety check: PassUnknown1 mo ago
258

A skill your agent uses for Anomalib benchmark pipelines, tiled ensemble workflows, and advanced pipeline orchestration helpers while keeping experimental execution paths explicit.

VectorSpaceLab/AREX-Skill331—~952Automated safety check: PassApache-2.01 mo ago
259

Use this repo skill for QuIP quip-miner, the Substrate-integrated quantum mining CLI, when configuring or operating CPU/CUDA/Metal/Modal/QPU miners, managing hybrid wallets/bootstrap/identity…

VectorSpaceLab/AREX-Skill331—~1.4kAutomated safety check: PassAGPL-3.01 mo ago
260

Use this sub-skill for Torch-TensorRT runtime performance controls, CUDA Graphs, output allocation, caches, TensorRT-RTX runtime settings, mutable modules, refit, weight streaming, and benchmark…

VectorSpaceLab/AREX-Skill331—~1kAutomated safety check: PassBSD-3-Clause1 mo ago
261

Use this repo skill for Microsoft Swin-Transformer image-classification model, config, data, checkpoint, SimMIM, Swin-MoE, and optional CUDA acceleration workflows.

VectorSpaceLab/AREX-Skill331—~1.2kAutomated safety check: PassMIT1 mo ago
262

A skill your agent uses for Torch-TensorRT tasks: compiling PyTorch models with TensorRT, dynamic-shape/export workflows, runtime optimization, Triton/C++/distributed deployment, debugging…

VectorSpaceLab/AREX-Skill331—~1.5kAutomated safety check: PassBSD-3-Clause1 mo ago
263

Generate novel crystal structures and molecules using ADiT (All-atom Diffusion Transformer), a unified latent diffusion model.

learningmatter-mit/AtomisticSkills176—~1.4kAutomated safety check: PassMIT3 days ago
264

AI for Science 场景下的昇腾 NPU Profiling 采集与性能分析 Skill,用于在华为 Ascend NPU 上使用 torchnpu.profiler 采集 L0、L1、L2 级性能数据,分析训练或推理中的算子耗时、调用栈、内存与瓶颈,并指导后续调优。

ascend-ai-coding/awesome-ascend-skills174—~3kAutomated safety check: PassNo licenceyesterday
265

Ankh 蛋白质语言模型昇腾 NPU 迁移 Skill,适用于 Ankh base/large、Ankh3 large/XL 以及同类基于 HuggingFace Transformers 与 PyTorch 的蛋白模型从 CUDA/GPU 到华为 Ascend NPU 的环境检查、代码适配、权重加载、验证脚本补齐与文档沉淀。

ascend-ai-coding/awesome-ascend-skills174—~2.1kAutomated safety check: PassNo licenceyesterday
266

Boltz2 蛋白质结构预测模型的昇腾 NPU 迁移与复现 Skill,适用于在华为 Ascend 910、910B、910C 上准备权重、适配 Lightning 和 CUDA only kernel、完成 Boltz2 端到端结构预测推理,并沉淀可复现的环境与验证命令。

ascend-ai-coding/awesome-ascend-skills174—~2.8kAutomated safety check: PassNo licenceyesterday
267

BoltzGen 昇腾 NPU 迁移与复现 Skill,适用于在华为 Ascend NPU 上部署 BoltzGen 生成式蛋白设计与逆折叠流程,覆盖环境准备、权重缓存、cuEquivariance 兼容、源码适配和端到端推理验证。

ascend-ai-coding/awesome-ascend-skills174—~3.8kAutomated safety check: PassNo licenceyesterday
268

DiffSBDD 昇腾 NPU 迁移 Skill,适用于将基于等变扩散模型的结构化药物设计项目从 CUDA 迁移到华为 Ascend NPU,覆盖环境搭建、依赖安装、torchscatter 源码编译、代码适配以及 de novo 推理验证。

ascend-ai-coding/awesome-ascend-skills174—~898Automated safety check: PassNo licenceyesterday
269

GENERator DNA 序列生成模型的昇腾 NPU 迁移 Skill,适用于将基于 HuggingFace Transformers 的 Causal LM 从 CUDA 迁移到华为 Ascend NPU,覆盖环境搭建、依赖安装、代码适配、多进程处理和 sequence recovery 验证。

ascend-ai-coding/awesome-ascend-skills174—~827Automated safety check: PassNo licenceyesterday
270

OligoFormer 昇腾 NPU 迁移 Skill,适用于将基于 PyTorch Transformer 的 siRNA 效能预测模型迁移到华为 Ascend NPU,覆盖环境搭建、RNA-FM 依赖安装、代码适配、推理验证与可选训练流程。

ascend-ai-coding/awesome-ascend-skills174—~1.3kAutomated safety check: PassNo licenceyesterday
271
271.Tao SetupOfficial

One-time session setup and orchestration map for the TAO skill bank.

NVIDIA/skills3.6k—~1.8kAutomated safety check: WarnApache-2.0yesterday
272

A skill your agent uses when training or debugging a neural net in PyTorch — the forward/loss/backward/step loop and its silent bugs, mixed precision (AMP), AdamW/LR schedules, DDP/FSDP/ZeRO…

ericrisco/rsc-harness180—~3.4kAutomated safety check: PassMITyesterday
273

Install, manage, and run ComfyUI instances. An agent skill from sundial-org/awesome-openclaw-skills.

sundial-org/awesome-openclaw-skills663—~1.5kAutomated safety check: PassNo licence7 mo ago
274
274.Cuda

CUDA C/C++ skill for NVIDIA GPU kernel programming. An agent skill from mohitmishra786/low-level-dev-skills.

mohitmishra786/low-level-dev-skills252—~1.9kAutomated safety check: PassMIT3 mo ago
275

CUDA debugging skill for GPU program correctness. An agent skill from mohitmishra786/low-level-dev-skills.

mohitmishra786/low-level-dev-skills252—~1.5kAutomated safety check: PassMIT3 mo ago
276

CUDA profiling skill for NVIDIA GPU performance analysis. An agent skill from mohitmishra786/low-level-dev-skills.

mohitmishra786/low-level-dev-skills252—~1.6kAutomated safety check: NotesMIT3 mo ago
277

HIP and ROCm skill for AMD GPU programming. An agent skill from mohitmishra786/low-level-dev-skills.

mohitmishra786/low-level-dev-skills252—~1.6kAutomated safety check: NotesMIT3 mo ago
278

Triton language skill for Python GPU kernel authoring. An agent skill from mohitmishra786/low-level-dev-skills.

mohitmishra786/low-level-dev-skills252—~1.8kAutomated safety check: PassMIT3 mo ago
279

Operate, configure, benchmark, and troubleshoot llama.cpp across CPU, Metal, CUDA, HIP/ROCm, Vulkan, SYCL, and hybrid or multi-GPU systems.

magnus919/agent-skills115—~2.3kAutomated safety check: PassMITyesterday
280

使用本地 BiRefNet GGUF 模型完成图片或视频抠图、人物抠图、主体分割和背景移除,并输出透明 PNG、MOV 或 WebM。适用于用户提到图片抠图、照片去背景、人像透明图、视频抠图、透明视频、BiRefNet、JPG/PNG/BMP/WebP 图片,或 MP4/MOV/WebM 视频的场景;无需 Python、PyTorch 或 CUDA。

aiskillstore/marketplace433—~748Automated safety check: PassMITyesterday
281
281.Uma

Run structure relaxation and phonon calculations using Meta's UMA (Universal Materials Accelerator) via fairchem

lamm-mit/scienceclaw246—~3.5kAutomated safety check: PassApache-2.01 mo ago
282
282.Tao Run On DockerOfficial

The Docker execution platform for TAO jobs — a local daemon or a remote GPU box via DOCKERHOST=ssh://user@host.

NVIDIA/skills3.6k—~5kAutomated safety check: WarnApache-2.0yesterday