Graphsignal
graphsignal/graphsignal
Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.
A skill your agent uses when something is wrong: Search() hangs, all evaluations return INVALIDSCORE, scores aren't improving, every config returns the same number, ptxas errors fill the log, CV% is…
$ npx skills add NVIDIA/CompileIQ --skill compileiq-debug -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install NVIDIA/CompileIQ compileiq-debug --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/NVIDIA/CompileIQ.git skills-src && mkdir -p .claude/skills && cp -r skills-src/agent-skills/compileiq-debug .claude/skills/compileiq-debug && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "compileiq-debug" agent skill from https://github.com/NVIDIA/CompileIQ/tree/main/agent-skills/compileiq-debug into .claude/skills/compileiq-debug/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "compileiq-debug", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/NVIDIA/CompileIQ/tree/main/agent-skills/compileiq-debugType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add NVIDIA/CompileIQ --skill compileiq-debug -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install NVIDIA/CompileIQ compileiq-debug --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/CompileIQ.git skills-src && mkdir -p .agents/skills && cp -r skills-src/agent-skills/compileiq-debug .agents/skills/compileiq-debug && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "compileiq-debug" agent skill from https://github.com/NVIDIA/CompileIQ/tree/main/agent-skills/compileiq-debug into .agents/skills/compileiq-debug/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "compileiq-debug", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add NVIDIA/CompileIQ --skill compileiq-debug -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install NVIDIA/CompileIQ compileiq-debug --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/CompileIQ.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/agent-skills/compileiq-debug .cursor/skills/compileiq-debug && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "compileiq-debug" agent skill from https://github.com/NVIDIA/CompileIQ/tree/main/agent-skills/compileiq-debug into .cursor/skills/compileiq-debug/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "compileiq-debug", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/NVIDIA/CompileIQ.git --path agent-skills/compileiq-debug--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add NVIDIA/CompileIQ --skill compileiq-debug -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install NVIDIA/CompileIQ compileiq-debug --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/CompileIQ.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/agent-skills/compileiq-debug .gemini/skills/compileiq-debug && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "compileiq-debug" agent skill from https://github.com/NVIDIA/CompileIQ/tree/main/agent-skills/compileiq-debug into .gemini/skills/compileiq-debug/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "compileiq-debug", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install NVIDIA/CompileIQ compileiq-debugInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add NVIDIA/CompileIQ --skill compileiq-debug -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/NVIDIA/CompileIQ.git skills-src && mkdir -p .github/skills && cp -r skills-src/agent-skills/compileiq-debug .github/skills/compileiq-debug && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "compileiq-debug" agent skill from https://github.com/NVIDIA/CompileIQ/tree/main/agent-skills/compileiq-debug into .github/skills/compileiq-debug/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "compileiq-debug", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add NVIDIA/CompileIQ --skill compileiq-debug -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install NVIDIA/CompileIQ compileiq-debug --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/CompileIQ.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/agent-skills/compileiq-debug .opencode/skills/compileiq-debug && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "compileiq-debug" agent skill from https://github.com/NVIDIA/CompileIQ/tree/main/agent-skills/compileiq-debug into .opencode/skills/compileiq-debug/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "compileiq-debug", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
compileiq-debugA skill your agent uses when something is wrong: Search() hangs, all evaluations return INVALIDSCORE, scores aren't improving, every config returns the same number, ptxas errors fill the log, CV% is…
Compileiq Debug is an agent skill from NVIDIA/CompileIQ, published by the product's own GitHub organization. Use when something is wrong: Search() hangs, all evaluations return INVALIDSCORE, scores aren't improving, every config returns the same number, ptxas errors fill the log, CV% is too high, or a winning ACF candidate needs NCU profiling to explain. Symptom-indexed table on top. Triggers on "compileiq hang", "socket timeout", "INVALIDSCORE", "not converging", "every score is the same", "TypeError fromhex", "ncu profile", "register spill", "ptxas error", "not in expected format", "high cv".
Its SKILL.md is about 2.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including scripts (for example `scripts/diagnose_csv.py`).
It sits in AI & LLM Engineering. It works with NVIDIA AI Platform and CUDA. The repository describes itself as: An Optimizer for Nvidia Compilers. The licence is Apache-2.0.
3 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 743aca4. It shows what the files ask for, not the result of running them.
Pre-approves these tools, so the agent can use them without asking each time:
BashReadFrom allowed-tools in the SKILL.md frontmatter.
Ships 1 file in scripts/ (Python), which the agent can run.
Shell commands in SKILL.md call:
pythonghbashFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use gh, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Compileiq Debug loads about 2.9k tokens when it runs. Until then it costs about 128 tokens; SKILL.md has 1,038 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check noted patterns worth knowing about, such as sudo or a known installer.
allowed-tools: Bash, ReadAutomated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from NVIDIA/CompileIQ at commit 743aca4, republished under its Apache-2.0 licence (© NVIDIA). 1,038 words, ~2,877 tokens.
.claude/skills/compileiq-debug/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.A symptom-indexed cheat sheet. Find the row that matches what the user is seeing, follow the first action, then dig into the matching detail section.
| Symptom | Most-likely cause | First action |
|---|---|---|
| Search hangs on first eval; socket timeout | Search space too large for default CIQ_SOCKET_TIMEOUT=20; OR release fetch slow/blocked; OR forkserver issue | Raise CIQ_SOCKET_TIMEOUT=120; if still hangs, CIQ_SEARCH_SPACES_DIR=<local mirror>; if still hangs, CIQ_PROCESS_MODE=spawn. Do NOT symlink BLAS — that fix is obsolete |
Every eval returns INVALID_SCORE | Return-type mismatch (tuple vs scalar) / correctness gate / task_timeout too tight | Search.sample(1) + call objective by hand; check num_objectives vs return shape; raise task_timeout |
| Every eval returns the same score | ACF not reaching the compiler — framework cache is hiding it | Apply Debug-pack O0 ACF by hand; if the score doesn't regress, fix cache-bust (TRITON_ALWAYS_COMPILE=1, HELION_SKIP_CACHE=1, fresh TRITON_CACHE_DIR, drop FlashInfer cubin packages) |
TypeError: fromhex() … not dict | Legacy bytes.fromhex(config_blob) in objective | Replace with save_compiler_config(acf_path, config) — see compileiq-author-objective |
"not in expected format" | Objective returned wrong shape | num_objectives must equal len(return_tuple); scalar return only when num_objectives=1 |
| Convergence stalled (best score flat) | Pool too small for space; mutate_rate too low; or kernel near-optimal | Raise pool_size; raise mutate_rate; sample diversity with Search.sample(20) |
| Increasing invalid rate over generations | Mutation arm spreading; compiler version drift mid-run | CIQ_KEEP_CACHE=1, re-run, inspect failing configs offline |
| CV% > 10% on validation | Unlocked clocks / thermal throttling / GPU contention | Lock GPU + memory clocks; pin CUDA_VISIBLE_DEVICES; watch nvidia-smi dmon for thermals |
| Need to know why a winning ACF helps | Profile with NCU | See NCU section below |
Not BLAS. Do not send users to symlink
libblas.so.
Current shipped binaries link only libm/libc/libstdc++/libgcc_s and
do not require BLAS/LAPACK.
Real causes today, in order of frequency:
CIQ_SOCKET_TIMEOUT too low. Default is 20 seconds, which is fine for
small search spaces but fails on big ones. Raise to 120 first; raise to
300+ for very large spaces.PtxasSearchSpace().retrieve() downloads from github.com. On a corporate
firewall this can stall. Pre-stage the mirror:gh release download search-spaces-latest -R NVIDIA/CompileIQ -D /shared/mirror
export CIQ_SEARCH_SPACES_DIR=/shared/mirrorforkserver is unsupported on the host. Set CIQ_PROCESS_MODE=spawn.
IsoMultiProcessWorker already uses fork by default.Sanity-check the shape before assuming the worst:
sample = tuner.sample(1)[0]
score = objective(sample)
print(type(score), score)Common shape mismatches:
num_objectives=1 but objective returns a tuple (latency,). Drop the
trailing comma.num_objectives=2 but objective returns a scalar.task_timeout is shorter than a clean compile takes; raise it.Almost always a framework cache serving a stale binary. Run the O0/O3 canary
from the Debug pack to confirm — see compileiq-booster-pack for the exact
test. If O0 doesn't regress vs baseline, the ACF is not reaching PTXAS. Fix:
| Framework | Cache-bust |
|---|---|
| Triton | TRITON_ALWAYS_COMPILE=1 + unique TRITON_CACHE_DIR per eval |
| Helion | HELION_SKIP_CACHE=1 |
| FlashInfer | Confirm flashinfer_cubin and flashinfer_jit_cache packages are absent (docs/flashinfer_booster.md:56-64) |
| Raw nvcc | Clean the build dir between candidates |
Legacy pattern from the pre-2026 skill set:
# OLD — DO NOT USE
def objective(config_blob):
with open(tmp_path, "wb") as f:
f.write(bytes.fromhex(config_blob))
...Replace with:
from compileiq.utils.helpers import save_compiler_config
def objective(config: str):
save_compiler_config(tmp_path, config)
...save_compiler_config does the bytes.fromhex internally. See
compileiq-author-objective for the full pattern.
The objective returned a shape CompileIQ's core doesn't expect. Rules:
num_objectives=1: objective must return a single scalar (int | float).
Not a 1-tuple, not a list.num_objectives>=2: objective must return a tuple or list of that length.result = objective(sample)
assert (
(search_config.num_objectives == 1 and isinstance(result, (int, float)))
or (search_config.num_objectives > 1 and len(result) == search_config.num_objectives)
), f"shape mismatch: {result!r} vs num_objectives={search_config.num_objectives}"Three causes, in order:
pool_size = max(2 * num_objectives + 1, 32) is the auto-derived floor — for spaces with >1k design points, raise to 64-128.mutate_rate=0.25. Raise to 0.3-0.5 if the search is converging on the first generation.Search.sample(20) and timing each sample by hand — if the spread is <5%, the search space is shallow.Probably a mutation arm spreading a structurally-bad config across the
population. Re-run with CIQ_KEEP_CACHE=1 so the failing configs are
preserved at ~/.cache/compileiq/, then replay them by hand to identify the
common factor.
If
cv = std/mean > 10%, validation can't tell the signal from the noise.
Fixes, in order:
compileiq-run-search for the
nvidia-smi --lock-*-clocks snippet).CUDA_VISIBLE_DEVICES=<gpu> so the validation has the GPU to itself.nvidia-smi dmon -i <gpu> -s pucvm for thermal throttling events.cudaEvent to NVBench (entropy-based stopping criterion, cold-cache between samples).When a search misbehaves, re-run with:
export CIQ_KEEP_CACHE=1The cache at ~/.cache/compileiq/ is preserved after the run. You can:
Quick pandas snippet:
import pandas as pd
df = pd.read_csv("results.csv")
df["score_numeric"] = pd.to_numeric(df["score_1"], errors="coerce")
gen_summary = df.groupby("generation").agg(
n=("score_numeric", "size"),
invalid=("score_numeric", lambda s: s.isna().sum() + (s > 1e10).sum()),
best=("score_numeric", "min"),
)
print(gen_summary)If invalid doesn't decrease across generations, your search is structurally
broken — try the O0/O3 canary in compileiq-author-objective.
For an automated version: python scripts/diagnose_csv.py results.csv.
Profile only after a validated ACF candidate exists. Don't profile every config — it's slow.
# Baseline (no ACF)
ncu --set full -o baseline -f --kernel-name my_kernel python bench.py
# ACF-applied — match the injection your objective uses
# Raw PTXAS:
ncu --set full -o opt -f --kernel-name my_kernel \
bash -c 'PTXAS_OPTIONS="--apply-controls=best.acf" python bench.py'
# NVCC build:
nvcc -Xptxas --apply-controls=best.acf bench.cu -o bench && \
ncu --set full -o opt -f --kernel-name my_kernel ./bench
# Diff
ncu --import baseline.ncu-rep --import opt.ncu-rep --csv --page raw > diff.csv| Metric | What it means |
|---|---|
sm__throughput.avg.pct_of_peak_sustained_elapsed | Compute throughput |
gpu__compute_memory_throughput.avg.pct_of_peak_sustained_elapsed | Memory throughput |
sm__warps_active.avg.pct_of_peak_sustained_active | Achieved occupancy |
launch__registers_per_thread | Register pressure |
l2__throughput.avg.pct_of_peak_sustained_elapsed | L2 pressure |
If the ACF moved any of these meaningfully, that's the mechanism. If none of them moved but the win is real, look at lower-level metrics (warp stalls, issue slot utilization) — those are harder to interpret but often the answer.
ptxas -v -arch=sm_100 --apply-controls best.acf kernel.ptx 2>&1 \
| grep -E "registers|spill|stack"Reports Used N registers, X bytes stack frame, Y bytes spill stores, Z bytes spill loads. If Y + Z goes up vs baseline, the ACF traded
register pressure for memory traffic — sometimes a real win, sometimes not.
Investigate before shipping.
python scripts/diagnose_csv.py --self-testSynthesizes a small results.csv covering each pathology (clean convergence,
rising invalid rate, stalled best-score) and asserts the heuristic
classifications match.
ldd). Carrying it forward sends users on a wild goose chase.COMMON_PTXAS_ERRORS dict mapping individual ptxas error
strings to fixes — users get INVALID_SCORE instead, no need to recognize
specific messages.compileiq-bootstrap.validate_objective_function introspection helper — replaced by
Search.sample(1) + the Debug-pack O0/O3 canary.compileiq-booster-pack or
compileiq-author-objective.compileiq-bootstrap.compileiq-run-search.© NVIDIA, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 1 other file (scripts) in agent-skills/compileiq-debug of NVIDIA/CompileIQ.
Open the folder on GitHubat commit 743aca4
Compileiq Debug next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Compileiq Debug this skillNVIDIA/CompileIQ | 137 | — | ~2.9k | Automated safety check: Notes | Apache-2.0 | |
| Graphsignalgraphsignal/graphsignal | 257 | — | ~6.2k | Automated safety check: Pass | Apache-2.0 | |
| LLM Torch Profiler Trace AnalysisBBuf/AI-Infra-Auto-Driven-SKILLS | 900 | — | ~2.8k | Automated safety check: Pass | None | |
| Optimize OpCVCUDA/CV-CUDA | 2.7k | — | ~834 | Automated safety check: Pass | Custom licence | |
| Cutlass SkillslowlyC/agent-gpu-skills | 169 | — | ~1.3k | Automated safety check: Pass | MIT | |
| Setup Workshop Nemoclawbrevdev/workshop-build-an-agent | 143 | — | ~5.2k | Automated safety check: Pass | Apache-2.0 |
graphsignal/graphsignal
Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.
BBuf/AI-Infra-Auto-Driven-SKILLS
Analyzes Torch Profiler traces from SGLang, vLLM and TensorRT-LLM servers into kernel attribution, overlap and fusion tables.
CVCUDA/CV-CUDA
Drive a single-operator optimization campaign per .agents/guidance/OPTIMIZATIONGUIDELINES.md, with a deterministically enforced definition-of-done and versioned MR summary.
slowlyC/agent-gpu-skills
Write, debug, and optimize CUTLASS, CuTe, and CuTeDSL GPU kernels from local upstream source, examples, and headers.
brevdev/workshop-build-an-agent
Set up the NVIDIA "Build an Agent" DevX workshop as a working JupyterLab environment from INSIDE a locked-down OpenShell/NemoClaw sandbox, and hand the user the token URL + access commands.
Orchestra-Research/AI-Research-SKILLs
Runs LLM inference on CPU, Apple Silicon, and consumer GPUs without NVIDIA hardware.
NVIDIA/CompileIQ
Use BEFORE running a full CompileIQ search. An agent skill from NVIDIA/CompileIQ.
NVIDIA/CompileIQ
A skill your agent uses when starting a fresh CompileIQ project, hitting a socket timeout, or before running any other compileiq- skill.
NVIDIA/CompileIQ
A skill your agent uses when composing the Search(...) call and calling .start().
NVIDIA/CompileIQ
Use AFTER a Search has completed and BEFORE claiming any speedup or shipping an ACF.
NVIDIA/CompileIQ
A skill your agent uses when writing the objectivefunction= passed to Search().
NVIDIA/CompileIQ
A skill your agent uses when picking the searchspace= argument for Search().
Works with
Categories
A skill your agent uses when something is wrong: Search() hangs, all evaluations return INVALIDSCORE, scores aren't improving, every config returns the same number, ptxas errors fill the log, CV% is…. Compileiq Debug is an agent skill from NVIDIA/CompileIQ, published by the product's own GitHub organization. Use when something is wrong: Search() hangs, all evaluations return INVALIDSCORE, scores aren't improving, every config returns the same number, ptxas errors fill the log, CV% is too high, or a winning ACF candidate needs NCU profiling to explain.
Compileiq Debug fits situations like: something is wrong: Search() hangs; all evaluations return INVALIDSCORE; scores arent improving; every config returns the same number.
Run `npx skills add NVIDIA/CompileIQ --skill compileiq-debug -a claude-code`. Or copy the skill folder (agent-skills/compileiq-debug in NVIDIA/CompileIQ) into .claude/skills/compileiq-debug in your project. Claude Code loads it when a task matches its description.
Run `npx skills add NVIDIA/CompileIQ --skill compileiq-debug -a codex`. Or copy the skill folder (agent-skills/compileiq-debug in NVIDIA/CompileIQ) into .agents/skills/compileiq-debug in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NVIDIA/CompileIQ --skill compileiq-debug -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/compileiq-debug, .gemini/skills/compileiq-debug, .github/skills/compileiq-debug and .opencode/skills/compileiq-debug in your project.
Going by SKILL.md and its folder, Compileiq Debug needs Python for the scripts in its folder and the command-line tools its instructions call (python, gh and bash). Our summary lists: Python 3. Its frontmatter pre-approves these tools: Bash, Read.
SKILL.md contains no URLs. Its commands use gh, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Compileiq Debug is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.9k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Compileiq Debug: Graphsignal (graphsignal/graphsignal, 257 stars), LLM Torch Profiler Trace Analysis (BBuf/AI-Infra-Auto-Driven-SKILLS, 900 stars), Optimize Op (CVCUDA/CV-CUDA, 2.7k stars) and Cutlass Skill (slowlyC/agent-gpu-skills, 169 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
NVIDIA (a GitHub organization, an official publisher) maintains it in NVIDIA/CompileIQ, which has 137 GitHub stars. The repository holds 7 skills in this directory. The repository was last updated on September 23, 2026.
Source: NVIDIA/CompileIQ on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.