Sglang Prod Incident Triage
sgl-project/sglang
Replay-first debug flow for SGLang serving problems. An agent skill from sgl-project/sglang.
A skill your agent uses when composing the Search(...) call and calling .start().
$ npx skills add NVIDIA/CompileIQ --skill compileiq-run-search -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install NVIDIA/CompileIQ compileiq-run-search --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/NVIDIA/CompileIQ.git skills-src && mkdir -p .claude/skills && cp -r skills-src/agent-skills/compileiq-run-search .claude/skills/compileiq-run-search && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "compileiq-run-search" agent skill from https://github.com/NVIDIA/CompileIQ/tree/main/agent-skills/compileiq-run-search into .claude/skills/compileiq-run-search/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "compileiq-run-search", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/NVIDIA/CompileIQ/tree/main/agent-skills/compileiq-run-searchType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add NVIDIA/CompileIQ --skill compileiq-run-search -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install NVIDIA/CompileIQ compileiq-run-search --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/CompileIQ.git skills-src && mkdir -p .agents/skills && cp -r skills-src/agent-skills/compileiq-run-search .agents/skills/compileiq-run-search && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "compileiq-run-search" agent skill from https://github.com/NVIDIA/CompileIQ/tree/main/agent-skills/compileiq-run-search into .agents/skills/compileiq-run-search/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "compileiq-run-search", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add NVIDIA/CompileIQ --skill compileiq-run-search -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install NVIDIA/CompileIQ compileiq-run-search --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/CompileIQ.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/agent-skills/compileiq-run-search .cursor/skills/compileiq-run-search && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "compileiq-run-search" agent skill from https://github.com/NVIDIA/CompileIQ/tree/main/agent-skills/compileiq-run-search into .cursor/skills/compileiq-run-search/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "compileiq-run-search", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/NVIDIA/CompileIQ.git --path agent-skills/compileiq-run-search--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add NVIDIA/CompileIQ --skill compileiq-run-search -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install NVIDIA/CompileIQ compileiq-run-search --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/CompileIQ.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/agent-skills/compileiq-run-search .gemini/skills/compileiq-run-search && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "compileiq-run-search" agent skill from https://github.com/NVIDIA/CompileIQ/tree/main/agent-skills/compileiq-run-search into .gemini/skills/compileiq-run-search/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "compileiq-run-search", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install NVIDIA/CompileIQ compileiq-run-searchInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add NVIDIA/CompileIQ --skill compileiq-run-search -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/NVIDIA/CompileIQ.git skills-src && mkdir -p .github/skills && cp -r skills-src/agent-skills/compileiq-run-search .github/skills/compileiq-run-search && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "compileiq-run-search" agent skill from https://github.com/NVIDIA/CompileIQ/tree/main/agent-skills/compileiq-run-search into .github/skills/compileiq-run-search/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "compileiq-run-search", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add NVIDIA/CompileIQ --skill compileiq-run-search -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install NVIDIA/CompileIQ compileiq-run-search --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/CompileIQ.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/agent-skills/compileiq-run-search .opencode/skills/compileiq-run-search && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "compileiq-run-search" agent skill from https://github.com/NVIDIA/CompileIQ/tree/main/agent-skills/compileiq-run-search into .opencode/skills/compileiq-run-search/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "compileiq-run-search", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
compileiq-run-searchA skill your agent uses when composing the Search(...) call and calling .start().
Compileiq Run Search is an agent skill from NVIDIA/CompileIQ, published by the product's own GitHub organization. Use when composing the Search(...) call and calling .start(). Covers the four worker classes (MultiProcessWorker / IsoMultiProcessWorker / RayWorker / AsyncWorker) and when to pick each, SearchConfiguration sizing rules, dumpresults checkpointing, trackerconfig choice (Disabled / Loguru / MLflow), numworkers/tasktimeout semantics, and GPU clock locking for stable measurements. Triggers on "Search()", "tuner.start()", "poolsize", "numworkers", "tasktimeout", "IsoMultiProcessWorker", "RayWorker", "dumpresults"…
Its SKILL.md is about 2.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including scripts (for example `scripts/smoke_search.py`).
It sits in AI & LLM Engineering. It works with MLflow and CUDA. The repository describes itself as: An Optimizer for Nvidia Compilers. The licence is Apache-2.0.
3 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 743aca4. It shows what the files ask for, not the result of running them.
Pre-approves these tools, so the agent can use them without asking each time:
BashReadFrom allowed-tools in the SKILL.md frontmatter.
Ships 1 file in scripts/ (Python), which the agent can run.
Shell commands in SKILL.md call:
pythonFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Compileiq Run Search loads about 2.2k tokens when it runs. Until then it costs about 142 tokens; SKILL.md has 651 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check noted patterns worth knowing about, such as sudo or a known installer.
sudo nvidia-smi -pm 1sudo nvidia-smi --lock-gpu-clocks=$MAX_GPU,$MAX_GPU --lock-memory-clocks=$MAX_MEM,$MAX_MEMa CI container or a shared cluster where sudo isn't available, skipallowed-tools: Bash, ReadAutomated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from NVIDIA/CompileIQ at commit 743aca4, republished under its Apache-2.0 licence (© NVIDIA). 651 words, ~2,183 tokens.
.claude/skills/compileiq-run-search/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.After you have an objective function (from compileiq-author-objective) and
a search space (from compileiq-search-space), this skill helps you choose
the worker, size the configuration, and run the search safely.
Search(...) and call .start().Pass either a built-in WorkerTypes enum value or the worker class itself to
Search(worker_type=...):
from compileiq.types import WorkerTypes
from compileiq.worker import (
MultiProcessWorker, # default
IsoMultiProcessWorker, # spawns fresh process per task; kill-safe
RayWorker, # distributed
AsyncWorker, # asyncio for async def objectives
)| Situation | Worker class | Why |
|---|---|---|
| GPU kernel that may hang, OOM, or leak CUDA context | IsoMultiProcessWorker | One fresh process per task; parent kills on task_timeout. Defaults to fork. (docs/workers.md:42) |
| Triton mixed example on Blackwell-class GPUs | WorkerTypes.ISOLATED + CIQ_PROCESS_MODE=spawn | Isolates each evaluation and avoids leaking illegal memory access state across runs. |
| Fast (<100ms), stateless objective | MultiProcessWorker (default) | Reuses a pool; lower overhead. Defaults to forkserver. |
| Multi-node / multi-GPU cluster | RayWorker | User must set up Ray cluster + install compileiq on every worker. Both num_workers and task_timeout are ignored. (docs/workers.md:79-91) |
I/O-bound async def objective | AsyncWorker | Concurrency, not parallelism. Rare for GPU work. |
Default recommendation for compiler tuning of GPU kernels:
IsoMultiProcessWorker with task_timeout between 30s (small kernels) and
180s (large attention / XLA HLO).
Reference: compileiq/types.py:473-615. Defaults auto-derive; only set what
you must.
from compileiq.types import SearchConfiguration, ProblemType
config = SearchConfiguration(
problem_type=ProblemType.MIN, # MIN for latency; MAX for throughput
generations=10, # required, > 0
pool_size=15, # > 5; auto-derives if omitted
# cull_size auto-derives to 75% of pool, rounded down to even
# mutate_rate defaults to 0.25
# num_objectives defaults to 1
# normalize defaults to False (set True for cross-GPU runs)
)| Knob | Default | When to override |
|---|---|---|
generations | required | 10 for initial exploration; 20-40 for a deep run. |
pool_size | auto (≥32) | 15 for tiny spaces; 32 for ≥1k design points; 64-128 for ≥10k. |
cull_size | 75% of pool, even | Almost never override directly. |
mutate_rate | 0.25 | Raise to 0.3-0.5 only if convergence stalls in early gens. |
num_objectives | 1 | Must equal len(return_tuple) from the objective. |
normalize | False | True when running across heterogeneous nodes or GPUs. |
Sanity rule of thumb: if pool_size * generations < 50, you are exploring,
not optimizing. If > 2000, you are probably overfitting to measurement noise
— compileiq-validate-result will earn its keep there.
from pathlib import Path
from compileiq.ciq import Search
from compileiq.search_spaces.compilers import PtxasSearchSpace
from compileiq.tracker import LoguruTrackerConfig
tuner = Search(
objective_function=objective,
search_space=PtxasSearchSpace(version="13.3", variant="att"),
search_config=config,
worker_type=IsoMultiProcessWorker, # or WorkerTypes.ISOLATED
tracker_config=LoguruTrackerConfig(sink="optimization.log"),
dump_results=Path("results.csv"), # ALWAYS set this
cache_folder=None, # default ~/.cache/compileiq
disable_progress_bar=False,
exit_on_failure=True,
debug=False,
)Always set dump_results=Path(...). CSV is flushed every batch, so a crashed
or killed run leaves recoverable state.
results = tuner.start(num_workers=4, task_timeout=120)num_workers: ignored by workers where respects_num_workers=False
(RayWorker, AsyncWorker); CompileIQ emits the warning
"num_workers is not supported by <WorkerName>" (compileiq/ciq.py:449-451)
so users recognize it.task_timeout: ignored where supports_timeout=False (RayWorker).
Critical for IsoMultiProcessWorker — without it a hung config wedges that
branch.SearchResult. Don't process inline; hand off to
compileiq-validate-result.from compileiq.tracker import DisabledTrackerConfig, LoguruTrackerConfig, MLflowTrackerConfigDisabledTrackerConfig() — default, no overhead. Fine for one-off runs.LoguruTrackerConfig(sink="optimization.log", level="INFO") —
recommended for serious campaigns. Negligible overhead.MLflowTrackerConfig(experiment_name="...", tracking_uri="...", run_name="...")
— when integrating with ML Ops; creates a nested MLflow run per evaluation.Search.sample(n) returns n randomly sampled parameter dicts from the
search space without running the search. Use it to:
Search).sample = tuner.sample(1)[0]
print(sample)
print(objective(sample)) # should return a real float, not raiseStable measurements need locked clocks. Lock before tuner.start(),
unlock via atexit. Requires sudo.
sudo nvidia-smi -pm 1
MAX_GPU=$(nvidia-smi --query-gpu=clocks.max.graphics --format=csv,noheader,nounits | head -1)
MAX_MEM=$(nvidia-smi --query-gpu=clocks.max.memory --format=csv,noheader,nounits | head -1)
sudo nvidia-smi --lock-gpu-clocks=$MAX_GPU,$MAX_GPU --lock-memory-clocks=$MAX_MEM,$MAX_MEMimport atexit, subprocess
def unlock():
subprocess.run(["sudo", "nvidia-smi", "--reset-gpu-clocks", "--reset-memory-clocks"],
check=False)
atexit.register(unlock)Inside a CI container or a shared cluster where sudo isn't available, skip this; report higher CV% to the validation skill so it knows to compensate.
python scripts/smoke_search.pyRuns a 2-generation search on x**2 + y with MultiProcessWorker and
verifies results.get_best_result() returns a dict with score_1 and params.
task_timeout with IsoMultiProcessWorker is the most
common reason a search hangs for hours. The worker will kill a stuck
process but only after task_timeout elapses.forkserver issues on some hosts manifest as EOFError or "Broken pipe"
on the first eval. Set CIQ_PROCESS_MODE=spawn.num_workers > num_gpus is fine for fast CPU-side objectives but
oversubscribes GPUs for kernel objectives. For GPU kernels: pin
CUDA_VISIBLE_DEVICES inside the objective and set
num_workers = num_gpus..start() returns: compileiq-validate-result.compileiq-debug.© NVIDIA, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 1 other file (scripts) in agent-skills/compileiq-run-search of NVIDIA/CompileIQ.
Open the folder on GitHubat commit 743aca4
Compileiq Run Search next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Compileiq Run Search this skillNVIDIA/CompileIQ | 138 | — | ~2.2k | Automated safety check: Notes | Apache-2.0 | |
| Sglang Prod Incident Triagesgl-project/sglang | 37k | 3 repos | ~2.1k | Automated safety check: Pass | Apache-2.0 | |
| Paddle BuildPaddlePaddle/Paddle | 24k | — | ~1k | Automated safety check: Pass | Apache-2.0 | |
| Paddle Design CompilerPaddlePaddle/Paddle | 24k | — | ~3.6k | Automated safety check: Pass | Apache-2.0 | |
| Fastllm Triton Opsztxz16/fastllm | 5.1k | — | ~1.8k | Automated safety check: Pass | Apache-2.0 | |
| Cuda Index Widthpytorch/pytorch | 104k | — | ~1.6k | Automated safety check: Pass | Custom licence |
sgl-project/sglang
Replay-first debug flow for SGLang serving problems. An agent skill from sgl-project/sglang.
PaddlePaddle/Paddle
A skill your agent uses when needing to compile, rebuild, or install Paddle from source after code changes.
PaddlePaddle/Paddle
A skill your agent uses when working with Paddle 3.0 compiler full pipeline: SOT (Symbolic Opcode Translator) for bytecode-level dy2st graph capture, PIR (Paddle IR) for SSA-based intermediate…
ztxz16/fastllm
Guide for adding Triton-backed CUDA operators to FastLLM. An agent skill from ztxz16/fastllm.
pytorch/pytorch
Choose 32-bit vs 64-bit index math in PyTorch CUDA kernels. An agent skill from pytorch/pytorch.
JimLiu/science-skills
Biohub ESMFold2 / ESMFold2-Fast all-atom co-folding (Candido et al.
NVIDIA/CompileIQ
Use BEFORE running a full CompileIQ search. An agent skill from NVIDIA/CompileIQ.
NVIDIA/CompileIQ
A skill your agent uses when starting a fresh CompileIQ project, hitting a socket timeout, or before running any other compileiq- skill.
NVIDIA/CompileIQ
A skill your agent uses when something is wrong: Search() hangs, all evaluations return INVALIDSCORE, scores aren't improving, every config returns the same number, ptxas errors fill the log, CV% is…
NVIDIA/CompileIQ
Use AFTER a Search has completed and BEFORE claiming any speedup or shipping an ACF.
NVIDIA/CompileIQ
A skill your agent uses when writing the objectivefunction= passed to Search().
NVIDIA/CompileIQ
A skill your agent uses when picking the searchspace= argument for Search().
Categories
A skill your agent uses when composing the Search(...) call and calling .start(). Compileiq Run Search is an agent skill from NVIDIA/CompileIQ, published by the product's own GitHub organization.start().
Compileiq Run Search fits situations like: composing the Search(...) call and calling .start(); isoMultiProcessWorker.
Run `npx skills add NVIDIA/CompileIQ --skill compileiq-run-search -a claude-code`. Or copy the skill folder (agent-skills/compileiq-run-search in NVIDIA/CompileIQ) into .claude/skills/compileiq-run-search in your project. Claude Code loads it when a task matches its description.
Run `npx skills add NVIDIA/CompileIQ --skill compileiq-run-search -a codex`. Or copy the skill folder (agent-skills/compileiq-run-search in NVIDIA/CompileIQ) into .agents/skills/compileiq-run-search in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NVIDIA/CompileIQ --skill compileiq-run-search -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/compileiq-run-search, .gemini/skills/compileiq-run-search, .github/skills/compileiq-run-search and .opencode/skills/compileiq-run-search in your project.
Going by SKILL.md and its folder, Compileiq Run Search needs Python for the scripts in its folder and the command-line tools its instructions call (python). Our summary lists: Python 3. Its frontmatter pre-approves these tools: Bash, Read.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found notes only (runs commands with sudo; pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Compileiq Run Search is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.2k tokens (SKILL.md is roughly 8.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Compileiq Run Search: Sglang Prod Incident Triage (sgl-project/sglang, 37k stars), Paddle Build (PaddlePaddle/Paddle, 24k stars), Paddle Design Compiler (PaddlePaddle/Paddle, 24k stars) and Fastllm Triton Ops (ztxz16/fastllm, 5.1k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
NVIDIA (a GitHub organization, an official publisher) maintains it in NVIDIA/CompileIQ, which has 138 GitHub stars. The repository holds 7 skills in this directory. The repository was last updated on September 23, 2026.
Source: NVIDIA/CompileIQ on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.