Cuda Cpp Kernel
vipshop/cache-dit
A skill your agent uses when writing, debugging, porting, reviewing, or optimizing CUDA C++ or PTX kernels; investigating CUDA Runtime or Driver API behavior; profiling kernels with Nsight Systems…
GPU optimization workflow using uipc.profile, uipc.profile.nsight, and Nsight Compute CLI.
$ npx skills add spiriMirror/libuipc --skill gpu-optimization -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install spiriMirror/libuipc gpu-optimization --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/spiriMirror/libuipc.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.cursor/skills/gpu-optimization .claude/skills/gpu-optimization && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "gpu-optimization" agent skill from https://github.com/spiriMirror/libuipc/tree/main/.cursor/skills/gpu-optimization into .claude/skills/gpu-optimization/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gpu-optimization", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/spiriMirror/libuipc/tree/main/.cursor/skills/gpu-optimizationType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add spiriMirror/libuipc --skill gpu-optimization -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install spiriMirror/libuipc gpu-optimization --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/spiriMirror/libuipc.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.cursor/skills/gpu-optimization .agents/skills/gpu-optimization && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "gpu-optimization" agent skill from https://github.com/spiriMirror/libuipc/tree/main/.cursor/skills/gpu-optimization into .agents/skills/gpu-optimization/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gpu-optimization", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add spiriMirror/libuipc --skill gpu-optimization -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install spiriMirror/libuipc gpu-optimization --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/spiriMirror/libuipc.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.cursor/skills/gpu-optimization .cursor/skills/gpu-optimization && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "gpu-optimization" agent skill from https://github.com/spiriMirror/libuipc/tree/main/.cursor/skills/gpu-optimization into .cursor/skills/gpu-optimization/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gpu-optimization", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/spiriMirror/libuipc.git --path .cursor/skills/gpu-optimization--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add spiriMirror/libuipc --skill gpu-optimization -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install spiriMirror/libuipc gpu-optimization --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/spiriMirror/libuipc.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.cursor/skills/gpu-optimization .gemini/skills/gpu-optimization && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "gpu-optimization" agent skill from https://github.com/spiriMirror/libuipc/tree/main/.cursor/skills/gpu-optimization into .gemini/skills/gpu-optimization/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gpu-optimization", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install spiriMirror/libuipc gpu-optimizationInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add spiriMirror/libuipc --skill gpu-optimization -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/spiriMirror/libuipc.git skills-src && mkdir -p .github/skills && cp -r skills-src/.cursor/skills/gpu-optimization .github/skills/gpu-optimization && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "gpu-optimization" agent skill from https://github.com/spiriMirror/libuipc/tree/main/.cursor/skills/gpu-optimization into .github/skills/gpu-optimization/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gpu-optimization", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add spiriMirror/libuipc --skill gpu-optimization -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install spiriMirror/libuipc gpu-optimization --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/spiriMirror/libuipc.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.cursor/skills/gpu-optimization .opencode/skills/gpu-optimization && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "gpu-optimization" agent skill from https://github.com/spiriMirror/libuipc/tree/main/.cursor/skills/gpu-optimization into .opencode/skills/gpu-optimization/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gpu-optimization", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
gpu-optimizationGPU optimization workflow using uipc.profile, uipc.profile.nsight, and Nsight Compute CLI.
GPU Optimization is an agent skill from spiriMirror/libuipc. GPU optimization workflow using uipc.profile, uipc.profile.nsight, and Nsight Compute CLI. Use when profiling, optimizing, or benchmarking CUDA kernels.
Its SKILL.md is about 3.6k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in AI & LLM Engineering, covering GPU and accelerator computing. It works with CUDA and C++. The repository describes itself as: A Modern Python and C++20 Library of Unified Incremental Potential Contact. The licence is Apache-2.0.
4 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 9c748a7. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
pythonFrom the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
huggingface.coFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
GPU Optimization loads about 3.6k tokens when it runs. Until then it costs about 42 tokens; SKILL.md has 1,048 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from spiriMirror/libuipc at commit 9c748a7, republished under its Apache-2.0 licence (© spiriMirror). 1,048 words, ~3,647 tokens.
.claude/skills/gpu-optimization/SKILL.md (or your agent's skills folder).Workflow for profiling and optimizing CUDA kernels in libuipc using uipc.profile + uipc.profile.nsight + Nsight Compute CLI (ncu).
For building and installing, see cmake-workflow.
Build type: Always use Release for benchmarking and profiling. Debug is too slow; RelWithDebInfo adds debug info overhead that skews results.
.ncu-rep files — they are binary, for Nsight Compute GUI only.Debug builds — results are meaningless.timer_frames.json for the actual timer hierarchy — it is dynamic and varies per scene/run.# List available benchmark scenes
python -m uipc.cli.benchmark list
# Run a benchmark (headless, timer-based); --scene accepts multiple names
python -m uipc.cli.benchmark run --scene <scene> [<scene2> ...] --frames <N> --output <dir>
# Profile with Nsight Compute (kernel-level GPU metrics)
python -m uipc.cli.benchmark profile --scene <scene> --frames <N> --output <dir> --count <K> --skip <S> --ncu-set <set>
# Analyze bottlenecks from a benchmark result directory
python -m uipc.cli.benchmark analyze <result_dir> --ncu-csv <csv_path>
# Compare two benchmark result directories
python -m uipc.cli.benchmark compare <before_dir> <after_dir> --output <out_dir>profile flags:
--ncu-set default — basic metrics (SM%, occupancy, registers). Use full for duration/memory but requires admin/root.--count <K> — profile first K kernel launches.--skip <S> — skip initial warmup/init kernels.--ncu-path <path> — explicit path to ncu executable. Auto-detected via NCU_PATH env var, PATH, or default NVIDIA install locations.Note: The CLI profile defaults to --ncu-set default. The Python API nsight.run() defaults to ncu_set='full'. Be explicit to avoid confusion.
Run both run (for SimulationStats timer data) and profile (for Nsight Compute kernel metrics). Both are needed to make good optimization decisions.
Benchmark report (report/report.md from run):
Newton Iteration: 45%, Line Search: 30%)Nsight Compute report (<scene>_report.md / <scene>_report.json from profile):
| File | Source | What it tells you |
|---|---|---|
report/report.md | run | Wall-clock time breakdown by stage — what matters most |
timer_frames.json | run | Per-frame timer tree as JSON — source of truth for hierarchy |
<scene>_report.md | profile | Kernel GPU metrics — what's inefficient |
<scene>_report.json | profile | Structured kernel data with source_hint and optimization_hints |
<scene>.ncu-rep | profile | Binary — do NOT read, for Nsight Compute GUI only |
A kernel is worth optimizing only if it's both inefficient AND in a hot stage.
timer_frames.json or report/report.md to find which stages take the most wall-clock time (e.g., Detect DCD Candidates: 40%).StacklessBVH::* -> collision detection stage). Use the source_hint field in the JSON report.source_hint in the JSON report to find the CUDA source..cu file and optimize the kernel.Release (see cmake-workflow), re-benchmark, and compare.If the benchmark report shows a hot stage but you need more detail to locate the bottleneck within it, add Timer scopes in the C++ source:
#include <uipc/common/timer.h>
void MySystem::do_something()
{
{
Timer timer{"MySystem::phase_A"};
// ... GPU work ...
}
{
Timer timer{"MySystem::phase_B"};
// ... GPU work ...
}
}The Timer is scoped — it starts on construction and stops on destruction. Nested timers form a tree. The names appear in timer_frames.json and report/report.md, letting you drill down into which sub-phase of a hot stage is the actual bottleneck before running the more expensive ncu profiling.
The timer tree is dynamic — it varies per run depending on which simulation features are active (contact, friction, animation, etc.). The actual tree for any run is stored in timer_frames.json (written by run). Always read that file for the real hierarchy.
Below is a representative example with all features enabled, showing the typical nesting from sim_engine_do_advance.cu, global_linear_system.cu, and linear_pcg.cu:
Pipeline engine/
├── Rebuild Scene engine/
│ └── Update Diff Parm diff_sim/ (conditional)
└── Simulation engine/
├── Clear External Forces external_force/ (conditional)
├── Step Animation animator/ (conditional)
├── Compute External Force Accel. external_force/ (conditional)
├── Detect DCD Candidates collision_detection/ (conditional)
├── Newton Iteration engine/ (LOOP)
│ ├── Detect DCD Candidates collision_detection/ (iter > 0)
│ ├── Compute DyTopo Effect dytopo_effect_system/ (conditional)
│ │ ├── Assemble Dytopo Effect
│ │ ├── Convert Dytopo Matrix
│ │ └── Distribute Dytopo Effect
│ ├── Solve Global Linear System linear_system/
│ │ ├── Build Linear System
│ │ └── Solve Linear System
│ │ └── PCG
│ │ ├── Apply Preconditioner
│ │ └── SpMV (per PCG iteration)
│ └── Line Search line_search/
│ ├── Detect Trajectory Cand. collision_detection/
│ ├── Compute Energy line_search/ (initial E0)
│ ├── Filter CCD TOI collision_detection/
│ ├── Compute CFL Condition contact_system/
│ └── Line Search Iteration line_search/ (LOOP)
│ ├── Filter Contact Cand. collision_detection/
│ └── Compute Energy line_search/
└── Update Velocity time_integrator/Parent timer durations include their children. E.g., if Newton Iteration is 80% of frame time, PCG and Line Search durations are already counted inside that 80%.
All backend kernels are named __global__ functions (defined in anonymous namespaces) launched with raw <<<>>> — kernel symbols are directly readable in ncu reports, e.g.:
InfoStacklessBVHSimplexTrajectoryFilter_detect_k1_kernel
FEMLineSearchReporter_step_forward_kernel
StacklessBVH::... (older helpers may still appear as <Class>::<method>)Naming convention: <OwningClass>_<original_function>[_kN]_kernel (_kN = Nth launch point in the same function). To find the source of a kernel: search the kernel name (or its suffix) in src/backends/cuda/.
Legacy note: before the named-kernel migration, kernels were launched via muda::ParallelFor().apply(N, lambda) and appeared as parallel_for_kernel<Class::method()::lambda> (with #N for the Nth lambda in a function). The _shorten_kernel_name() function in nsight.py was written for that scheme; with named kernels it passes the readable name through unchanged.
Buffer operations (buffer_fill_kernel<T>, buffer_copy_kernel<T> in cuda_tool) are memory operations, usually not optimization targets.
uipc.profile — benchmark runner (timer-based, in-process)
uipc.profile.nsight — Nsight Compute profiler (kernel-level, ncu subprocess)The primary input for both is a World. A Scene is also accepted as a convenience shortcut (a temporary Engine + World is created internally).
from uipc import Scene, Engine, World
from uipc.assets import load
scene = Scene(Scene.default_config())
load('cube_ground', scene)
engine = Engine('cuda', 'my_workspace')
world = World(engine)
world.init(scene)uipc.profile — benchmark runnerfrom uipc import profile
# Simple: benchmark 10 frames from the world's current frame
result = profile.run(world, num_frames=10, name='baseline', output_dir='bench')
print(result['summary'])
# Flexible: mix warmup and benchmarking
with profile.session(world, name='baseline', output_dir='bench') as s:
s.advance(50) # warmup 50 frames (no stats)
s.profile(10) # benchmark 10 frames (collect stats)
print(s.result['summary'])
# Compare two saved benchmark directories
md = profile.compare('bench/baseline', 'bench/optimized', output_dir='comparison')
print(md)
# Load a saved result for later analysis
data = profile.load_result('bench/baseline')uipc.profile.nsight — Nsight Compute profilerfrom uipc.profile import nsight
# Simple: profile 2 frames under ncu
result = nsight.run(world, num_frames=2, name='cube_ground',
output_dir='ncu_results', ncu_set='default')
# Flexible: mix warmup and profiling
with nsight.session(world, name='cube_ground') as s:
s.profile(10) # profiles from world.frame()
print(s.result)When a World is passed, nsight automatically calls world.dump() and uses world.recover() in the subprocess for instant state restoration (no replay cost). The engine workspace is obtained via WorldVisitor(world).engine().workspace().
SimulationStats visualization toolsAvailable on the stats object in benchmark results (result['stats']):
stats.profiler_heatmap() — sunburst chart of timer breakdownstats.system_dependency_graph(workspace) — directed graph of backend system dependenciesstats.plot(keys, metric='duration') — per-frame line/bar chartstats.to_markdown(keys) — Markdown table of per-frame valuesAll relative to src/backends/cuda/:
finite_element/ — FEM constitutions, gradient/Hessian assemblyaffine_body/ — Affine body dynamicslinear_system/ — PCG solver, SpMV, preconditionerscollision_detection/ — BVH, trajectory filtering, DCDcontact_system/ — Contact forces (IPC barrier)global_geometry/ — Vertex management, bounding boxestime_integrator/ — Time integration, velocity updatedytopo_effect_system/ — Dynamic topology effectsline_search/ — Line search energy evaluationexternal_force/ — External force computationanimator/ — Animation steppingdiff_sim/ — Differentiable simulationcoupling_system/ — Multi-body couplingimplicit_geometry/ — Implicit geometry representationsinter_primitive_effect_system/ — Inter-primitive effectsnewton_tolerance/ — Newton convergence toleranceengine/ — Core pipeline orchestrationutils/ — Shared utilitiesWhen the bottleneck report identifies a hot kernel:
__launch_bounds__), adjust block size.muda::ParallelFor.__launch_bounds__(blockSize, minBlocksPerSM).Scenes are loaded from HuggingFace: MuGdxy/uipc-assets. Each asset has a scene.py with build_scene(scene). The asset module is at assets/init.py.
run)<output_dir>/<scene_name>/
benchmark.json # metadata (name, wall_time, num_frames, summary)
timer_frames.json # per-frame timer tree (JSON) — read this for hierarchy
workspace_<name>/ # engine workspace (contains systems.json)
report/
report.md # summary with charts — AGENT: read this
profiler_heatmap.svg # sunburst chart
system_deps.svg # system dependency graph (if workspace exists)
*.svg # per-timer chartsprofile)<output_dir>/
<scene>_report.md # AGENT: read this — kernel hotspot table
<scene>_report.json # AGENT: read this — structured metrics + source_hints
<scene>.ncu-rep # binary — do NOT read
<scene>.csv # raw CSV (already parsed into reports)JSON report per kernel:
{
"name": "StacklessBVH::calcExtNodeSplitMetrics",
"launches": 3,
"total_duration_ns": 123456.0,
"registers_per_thread": 40,
"avg_sm_pct": 100.0,
"avg_occupancy_pct": 48.0,
"source_hint": "src/backends/cuda/collision_detection/",
"optimization_hints": ["Low occupancy - consider reducing registers"],
"full_names": ["void muda::parallel_for_kernel<...>"]
}© spiriMirror, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .cursor/skills/gpu-optimization of spiriMirror/libuipc.
Open the folder on GitHubat commit 9c748a7
GPU Optimization next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| GPU Optimization this skillspiriMirror/libuipc | 335 | — | ~3.6k | Automated safety check: Pass | Apache-2.0 | |
| Cuda Cpp Kernelvipshop/cache-dit | 1.3k | — | ~2.3k | Automated safety check: Pass | Apache-2.0 | |
| Add Jit Kernelguqiong96/Lsglang | 143 | 1 repos | ~10k | Automated safety check: Pass | Apache-2.0 | |
| At Dispatch V2intel/torch-xpu-ops | 115 | 3 repos | ~2.2k | Automated safety check: Pass | Apache-2.0 | |
| Add Jit Kernelsgl-project/sglang | 37k | — | ~13k | Automated safety check: Pass | Apache-2.0 | |
| Cudamohitmishra786/low-level-dev-skills | 253 | — | ~1.9k | Automated safety check: Pass | MIT |
vipshop/cache-dit
A skill your agent uses when writing, debugging, porting, reviewing, or optimizing CUDA C++ or PTX kernels; investigating CUDA Runtime or Driver API behavior; profiling kernels with Nsight Systems…
guqiong96/Lsglang
Step-by-step tutorial for adding a new lightweight JIT CUDA kernel to sglang's jitkernel module
intel/torch-xpu-ops
Convert PyTorch ATDISPATCH macros to ATDISPATCHV2 format in ATen C++ code.
sgl-project/sglang
Step-by-step tutorial for adding a new lightweight JIT CUDA kernel to sglang.kernels JIT infrastructure and public operator groups
mohitmishra786/low-level-dev-skills
CUDA C/C++ skill for NVIDIA GPU kernel programming. An agent skill from mohitmishra786/low-level-dev-skills.
PaddlePaddle/Paddle
A skill your agent uses when needing to compile, rebuild, or install Paddle from source after code changes.
spiriMirror/libuipc
Review a libuipc pull request end-to-end: checkout the PR, summarize changes, list files for the human reviewer, and perform a domain-aware AI review covering physics correctness, backend…
spiriMirror/libuipc
Build and test the libuipc project using XMake. An agent skill from spiriMirror/libuipc.
spiriMirror/libuipc
Build and test the libuipc project using CMake. An agent skill from spiriMirror/libuipc.
spiriMirror/libuipc
Documentation style guide and rules for creating documentation
spiriMirror/libuipc
Fix a GitHub issue with proper branch, testing, and PR workflow
spiriMirror/libuipc
Fix a pull request based on review feedback. An agent skill from spiriMirror/libuipc.
Categories
GPU optimization workflow using uipc.profile, uipc.profile.nsight, and Nsight Compute CLI. GPU Optimization is an agent skill from spiriMirror/libuipc.nsight, and Nsight Compute CLI.
GPU Optimization fits situations like: benchmarking CUDA kernels; tasks that involve GPU and accelerator computing.
Run `npx skills add spiriMirror/libuipc --skill gpu-optimization -a claude-code`. Or copy the skill folder (.cursor/skills/gpu-optimization in spiriMirror/libuipc) into .claude/skills/gpu-optimization in your project. Claude Code loads it when a task matches its description.
Run `npx skills add spiriMirror/libuipc --skill gpu-optimization -a codex`. Or copy the skill folder (.cursor/skills/gpu-optimization in spiriMirror/libuipc) into .agents/skills/gpu-optimization in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add spiriMirror/libuipc --skill gpu-optimization -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/gpu-optimization, .gemini/skills/gpu-optimization, .github/skills/gpu-optimization and .opencode/skills/gpu-optimization in your project.
Going by SKILL.md and its folder, GPU Optimization needs the command-line tools its instructions call (python). Our summary lists: Python 3.
SKILL.md names 1 domain. As links in the text: huggingface.co. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
GPU Optimization is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 3.6k tokens (SKILL.md is roughly 15k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with GPU Optimization: Cuda Cpp Kernel (vipshop/cache-dit, 1.3k stars), Add Jit Kernel (guqiong96/Lsglang, 143 stars), At Dispatch V2 (intel/torch-xpu-ops, 115 stars) and Add Jit Kernel (sgl-project/sglang, 37k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
spiriMirror (a GitHub organization) maintains it in spiriMirror/libuipc, which has 335 GitHub stars. The repository holds 16 skills in this directory. The repository was last updated on October 3, 2026.
Source: spiriMirror/libuipc on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.