Official agent skill

Compileiq Run Search

by NVIDIA in NVIDIA/CompileIQ

A skill your agent uses when composing the Search(...) call and calling .start().

OfficialApache-2.0Auto-check: notesAI & LLM Engineering

Install Compileiq Run Search

skills CLI
$ npx skills add NVIDIA/CompileIQ --skill compileiq-run-search -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install NVIDIA/CompileIQ compileiq-run-search --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/NVIDIA/CompileIQ.git skills-src && mkdir -p .claude/skills && cp -r skills-src/agent-skills/compileiq-run-search .claude/skills/compileiq-run-search && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
compileiq-run-search
GitHub stars
138
Token cost
~2.2k tokens
SKILL.md length
651 words
Files
2 (incl. scripts)
Skills in repo
7
Repo updated
First seen
Licence
Apache-2.0

At a glance

A skill your agent uses when composing the Search(...) call and calling .start().

  • Works in 3 steps: Confirm the search space resolves at all… → Eyeball that the dicts have the keys… → Feed a single sample into the objective…
  • Composing the Search(...) call and calling .start()
  • SKILL.md covers When, Worker selection, SearchConfiguration sizing and Search(...) constructor —…, plus 7 more sections
  • Runs Python scripts from its folder; calls python

What it does

Compileiq Run Search is an agent skill from NVIDIA/CompileIQ, published by the product's own GitHub organization. Use when composing the Search(...) call and calling .start(). Covers the four worker classes (MultiProcessWorker / IsoMultiProcessWorker / RayWorker / AsyncWorker) and when to pick each, SearchConfiguration sizing rules, dumpresults checkpointing, trackerconfig choice (Disabled / Loguru / MLflow), numworkers/tasktimeout semantics, and GPU clock locking for stable measurements. Triggers on "Search()", "tuner.start()", "poolsize", "numworkers", "tasktimeout", "IsoMultiProcessWorker", "RayWorker", "dumpresults"…

Its SKILL.md is about 2.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including scripts (for example `scripts/smoke_search.py`).

It sits in AI & LLM Engineering. It works with MLflow and CUDA. The repository describes itself as: An Optimizer for Nvidia Compilers. The licence is Apache-2.0.

When your agent uses it

  • Composing the Search(...) call and calling .start()
  • IsoMultiProcessWorker

Example prompts

  • “Search()”
  • “tuner.start()”
  • “poolsize”
  • “/compileiq-run-search”

Requirements

  • Python 3
  • Pre-approved tools (allowed-tools): Bash, Read

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. Confirm the search space resolves at all (cheaper than the bootstrap
  2. Eyeball that the dicts have the keys your objective expects.
  3. Feed a single sample into the objective by hand to verify it runs.

What it can do on your machine

Read from SKILL.md and the folder at commit 743aca4. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Bash
    • Read

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Compileiq Run Search loads about 2.2k tokens when it runs. Until then it costs about 142 tokens; SKILL.md has 651 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~142
When it runs · the whole SKILL.md, loaded when a task matches
~2.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteRuns commands with sudoSKILL.md:174
    sudo nvidia-smi -pm 1
  • NoteRuns commands with sudoSKILL.md:177
    sudo nvidia-smi --lock-gpu-clocks=$MAX_GPU,$MAX_GPU --lock-memory-clocks=$MAX_MEM,$MAX_MEM
  • NoteRuns commands with sudoSKILL.md:188
    a CI container or a shared cluster where sudo isn't available, skip
  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Bash, Read

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from NVIDIA/CompileIQ at commit 743aca4, republished under its Apache-2.0 licence (© NVIDIA). 651 words, ~2,183 tokens.

Download SKILL.mdSave it as .claude/skills/compileiq-run-search/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
compileiq-run-search
description
Use when composing the Search(...) call and calling .start(). Covers the four worker classes (MultiProcessWorker / IsoMultiProcessWorker / RayWorker / AsyncWorker) and when to pick each, SearchConfiguration sizing rules, dump_results checkpointing, tracker_config choice (Disabled / Loguru / MLflow), num_workers/task_timeout semantics, and GPU clock locking for stable measurements. Triggers on "Search()", "tuner.start()", "pool_size", "num_workers", "task_timeout", "IsoMultiProcessWorker", "RayWorker", "dump_results", "MLflow", "GPU clocks".
allowed-tools
Bash, Read
when_to_use
- About to instantiate Search() and call .start(). - Search converges too slow / too fast and the user is unsure how to size. - Search is hanging on a few…
license
Apache-2.0
metadata.version
1.0.0
metadata.author
NVIDIA CompileIQ
metadata.domain
compiler-optimization
paths
**/*.py

After you have an objective function (from compileiq-author-objective) and a search space (from compileiq-search-space), this skill helps you choose the worker, size the configuration, and run the search safely.

When

  • About to instantiate Search(...) and call .start().
  • Search is converging too fast or too slow and the user is unsure how to re-size pool/generations.
  • Search hangs on individual configs and the worker doesn't kill them.
  • Scaling out from one GPU to a Ray cluster.

Worker selection

Pass either a built-in WorkerTypes enum value or the worker class itself to Search(worker_type=...):

python
from compileiq.types import WorkerTypes
from compileiq.worker import (
    MultiProcessWorker,    # default
    IsoMultiProcessWorker, # spawns fresh process per task; kill-safe
    RayWorker,             # distributed
    AsyncWorker,           # asyncio for async def objectives
)
SituationWorker classWhy
GPU kernel that may hang, OOM, or leak CUDA contextIsoMultiProcessWorkerOne fresh process per task; parent kills on task_timeout. Defaults to fork. (docs/workers.md:42)
Triton mixed example on Blackwell-class GPUsWorkerTypes.ISOLATED + CIQ_PROCESS_MODE=spawnIsolates each evaluation and avoids leaking illegal memory access state across runs.
Fast (<100ms), stateless objectiveMultiProcessWorker (default)Reuses a pool; lower overhead. Defaults to forkserver.
Multi-node / multi-GPU clusterRayWorkerUser must set up Ray cluster + install compileiq on every worker. Both num_workers and task_timeout are ignored. (docs/workers.md:79-91)
I/O-bound async def objectiveAsyncWorkerConcurrency, not parallelism. Rare for GPU work.

Default recommendation for compiler tuning of GPU kernels: IsoMultiProcessWorker with task_timeout between 30s (small kernels) and 180s (large attention / XLA HLO).

SearchConfiguration sizing

Reference: compileiq/types.py:473-615. Defaults auto-derive; only set what you must.

python
from compileiq.types import SearchConfiguration, ProblemType

config = SearchConfiguration(
    problem_type=ProblemType.MIN,   # MIN for latency; MAX for throughput
    generations=10,                  # required, > 0
    pool_size=15,                    # > 5; auto-derives if omitted
    # cull_size auto-derives to 75% of pool, rounded down to even
    # mutate_rate defaults to 0.25
    # num_objectives defaults to 1
    # normalize defaults to False (set True for cross-GPU runs)
)
KnobDefaultWhen to override
generationsrequired10 for initial exploration; 20-40 for a deep run.
pool_sizeauto (≥32)15 for tiny spaces; 32 for ≥1k design points; 64-128 for ≥10k.
cull_size75% of pool, evenAlmost never override directly.
mutate_rate0.25Raise to 0.3-0.5 only if convergence stalls in early gens.
num_objectives1Must equal len(return_tuple) from the objective.
normalizeFalseTrue when running across heterogeneous nodes or GPUs.

Sanity rule of thumb: if pool_size * generations < 50, you are exploring, not optimizing. If > 2000, you are probably overfitting to measurement noise — compileiq-validate-result will earn its keep there.

Search(...) constructor — every relevant kwarg

python
from pathlib import Path
from compileiq.ciq import Search
from compileiq.search_spaces.compilers import PtxasSearchSpace
from compileiq.tracker import LoguruTrackerConfig

tuner = Search(
    objective_function=objective,
    search_space=PtxasSearchSpace(version="13.3", variant="att"),
    search_config=config,
    worker_type=IsoMultiProcessWorker,                 # or WorkerTypes.ISOLATED
    tracker_config=LoguruTrackerConfig(sink="optimization.log"),
    dump_results=Path("results.csv"),                  # ALWAYS set this
    cache_folder=None,                                  # default ~/.cache/compileiq
    disable_progress_bar=False,
    exit_on_failure=True,
    debug=False,
)

Always set dump_results=Path(...). CSV is flushed every batch, so a crashed or killed run leaves recoverable state.

start(...) semantics

python
results = tuner.start(num_workers=4, task_timeout=120)
  • num_workers: ignored by workers where respects_num_workers=False (RayWorker, AsyncWorker); CompileIQ emits the warning "num_workers is not supported by <WorkerName>" (compileiq/ciq.py:449-451) so users recognize it.
  • task_timeout: ignored where supports_timeout=False (RayWorker). Critical for IsoMultiProcessWorker — without it a hung config wedges that branch.
  • Returns a SearchResult. Don't process inline; hand off to compileiq-validate-result.
Show full SKILL.md (261 more words)Show less

Tracker choice (one-line each)

python
from compileiq.tracker import DisabledTrackerConfig, LoguruTrackerConfig, MLflowTrackerConfig
  • DisabledTrackerConfig() — default, no overhead. Fine for one-off runs.
  • LoguruTrackerConfig(sink="optimization.log", level="INFO") — recommended for serious campaigns. Negligible overhead.
  • MLflowTrackerConfig(experiment_name="...", tracking_uri="...", run_name="...") — when integrating with ML Ops; creates a nested MLflow run per evaluation.

Search.sample(n) returns n randomly sampled parameter dicts from the search space without running the search. Use it to:

  1. Confirm the search space resolves at all (cheaper than the bootstrap round-trip; uses the in-memory state of Search).
  2. Eyeball that the dicts have the keys your objective expects.
  3. Feed a single sample into the objective by hand to verify it runs.
python
sample = tuner.sample(1)[0]
print(sample)
print(objective(sample))   # should return a real float, not raise

GPU clock locking (operator-level)

Stable measurements need locked clocks. Lock before tuner.start(), unlock via atexit. Requires sudo.

bash
sudo nvidia-smi -pm 1
MAX_GPU=$(nvidia-smi --query-gpu=clocks.max.graphics --format=csv,noheader,nounits | head -1)
MAX_MEM=$(nvidia-smi --query-gpu=clocks.max.memory --format=csv,noheader,nounits | head -1)
sudo nvidia-smi --lock-gpu-clocks=$MAX_GPU,$MAX_GPU --lock-memory-clocks=$MAX_MEM,$MAX_MEM
python
import atexit, subprocess
def unlock():
    subprocess.run(["sudo", "nvidia-smi", "--reset-gpu-clocks", "--reset-memory-clocks"],
                   check=False)
atexit.register(unlock)

Inside a CI container or a shared cluster where sudo isn't available, skip this; report higher CV% to the validation skill so it knows to compensate.

Self-test

bash
python scripts/smoke_search.py

Runs a 2-generation search on x**2 + y with MultiProcessWorker and verifies results.get_best_result() returns a dict with score_1 and params.

Gotchas

  • Forgetting task_timeout with IsoMultiProcessWorker is the most common reason a search hangs for hours. The worker will kill a stuck process but only after task_timeout elapses.
  • forkserver issues on some hosts manifest as EOFError or "Broken pipe" on the first eval. Set CIQ_PROCESS_MODE=spawn.
  • num_workers > num_gpus is fine for fast CPU-side objectives but oversubscribes GPUs for kernel objectives. For GPU kernels: pin CUDA_VISIBLE_DEVICES inside the objective and set num_workers = num_gpus.
  • Don't put GPU-clock lock calls inside the objective. They require sudo and are per-host operator setup, not per-eval.

Next

  • After .start() returns: compileiq-validate-result.
  • If something's wrong: compileiq-debug.

© NVIDIA, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (scripts) in agent-skills/compileiq-run-search of NVIDIA/CompileIQ.

  • SKILL.md
  • scripts/smoke_search.py

Open the folder on GitHubat commit 743aca4

Compare with similar skills

Compileiq Run Search next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Compileiq Run Search compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Compileiq Run Search this skillNVIDIA/CompileIQ138—~2.2kAutomated safety check: NotesApache-2.0
Sglang Prod Incident Triagesgl-project/sglang37k3 repos~2.1kAutomated safety check: PassApache-2.0
Paddle BuildPaddlePaddle/Paddle24k—~1kAutomated safety check: PassApache-2.0
Paddle Design CompilerPaddlePaddle/Paddle24k—~3.6kAutomated safety check: PassApache-2.0
Fastllm Triton Opsztxz16/fastllm5.1k—~1.8kAutomated safety check: PassApache-2.0
Cuda Index Widthpytorch/pytorch104k—~1.6kAutomated safety check: PassCustom licence

Similar skills

  • Sglang Prod Incident Triage

    sgl-project/sglang

    Replay-first debug flow for SGLang serving problems. An agent skill from sgl-project/sglang.

    37k GitHub starsUsed in 3 repos~2.1k tokens
    AI & LLM EngineeringAuto-check passed
  • Paddle Build

    PaddlePaddle/Paddle

    A skill your agent uses when needing to compile, rebuild, or install Paddle from source after code changes.

    24k GitHub stars~1k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Paddle Design Compiler

    PaddlePaddle/Paddle

    A skill your agent uses when working with Paddle 3.0 compiler full pipeline: SOT (Symbolic Opcode Translator) for bytecode-level dy2st graph capture, PIR (Paddle IR) for SSA-based intermediate…

    24k GitHub stars~3.6k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Fastllm Triton Ops

    ztxz16/fastllm

    Guide for adding Triton-backed CUDA operators to FastLLM. An agent skill from ztxz16/fastllm.

    5.1k GitHub stars~1.8k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Cuda Index Width

    pytorch/pytorch

    Choose 32-bit vs 64-bit index math in PyTorch CUDA kernels. An agent skill from pytorch/pytorch.

    104k GitHub stars~1.6k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Esmfold2

    JimLiu/science-skills

    Biohub ESMFold2 / ESMFold2-Fast all-atom co-folding (Candido et al.

    227 GitHub starsUsed in 4 repos~2.5k tokens
    AI & LLM EngineeringAuto-check passed

More from NVIDIA/CompileIQ

  • Compileiq Booster Pack

    NVIDIA/CompileIQ

    Official

    Use BEFORE running a full CompileIQ search. An agent skill from NVIDIA/CompileIQ.

    138 GitHub stars~2.1k tokensUpdated 15 days ago
    Auto-check: notes
  • Compileiq Bootstrap

    NVIDIA/CompileIQ

    Official

    A skill your agent uses when starting a fresh CompileIQ project, hitting a socket timeout, or before running any other compileiq- skill.

    138 GitHub stars~1.3k tokensUpdated 15 days ago
    Auto-check: notes
  • Compileiq Debug

    NVIDIA/CompileIQ

    Official

    A skill your agent uses when something is wrong: Search() hangs, all evaluations return INVALIDSCORE, scores aren't improving, every config returns the same number, ptxas errors fill the log, CV% is…

    138 GitHub stars~2.9k tokensUpdated 15 days ago
    Auto-check: notes
  • Official

    Use AFTER a Search has completed and BEFORE claiming any speedup or shipping an ACF.

    138 GitHub stars~2.2k tokensUpdated 15 days ago
    Auto-check: notes
  • Official

    A skill your agent uses when writing the objectivefunction= passed to Search().

    138 GitHub stars~2.3k tokensUpdated 15 days ago
    Auto-check: notes
  • Compileiq Search Space

    NVIDIA/CompileIQ

    Official

    A skill your agent uses when picking the searchspace= argument for Search().

    138 GitHub stars~1.9k tokensUpdated 15 days ago
    Auto-check: notes

Works with

Questions about Compileiq Run Search

What does Compileiq Run Search do?

A skill your agent uses when composing the Search(...) call and calling .start(). Compileiq Run Search is an agent skill from NVIDIA/CompileIQ, published by the product's own GitHub organization.start().

When should I use Compileiq Run Search?

Compileiq Run Search fits situations like: composing the Search(...) call and calling .start(); isoMultiProcessWorker.

How do I install Compileiq Run Search in Claude Code?

Run `npx skills add NVIDIA/CompileIQ --skill compileiq-run-search -a claude-code`. Or copy the skill folder (agent-skills/compileiq-run-search in NVIDIA/CompileIQ) into .claude/skills/compileiq-run-search in your project. Claude Code loads it when a task matches its description.

How do I install Compileiq Run Search in Codex?

Run `npx skills add NVIDIA/CompileIQ --skill compileiq-run-search -a codex`. Or copy the skill folder (agent-skills/compileiq-run-search in NVIDIA/CompileIQ) into .agents/skills/compileiq-run-search in your project. Codex loads it when a task matches its description.

Can I use Compileiq Run Search in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NVIDIA/CompileIQ --skill compileiq-run-search -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/compileiq-run-search, .gemini/skills/compileiq-run-search, .github/skills/compileiq-run-search and .opencode/skills/compileiq-run-search in your project.

What does Compileiq Run Search need to run?

Going by SKILL.md and its folder, Compileiq Run Search needs Python for the scripts in its folder and the command-line tools its instructions call (python). Our summary lists: Python 3. Its frontmatter pre-approves these tools: Bash, Read.

Does Compileiq Run Search access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Compileiq Run Search safe to install?

Our automated static check of SKILL.md found notes only (runs commands with sudo; pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Compileiq Run Search use?

Compileiq Run Search is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Compileiq Run Search use?

About 2.2k tokens (SKILL.md is roughly 8.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Compileiq Run Search?

Skills that share tags, products or a category with Compileiq Run Search: Sglang Prod Incident Triage (sgl-project/sglang, 37k stars), Paddle Build (PaddlePaddle/Paddle, 24k stars), Paddle Design Compiler (PaddlePaddle/Paddle, 24k stars) and Fastllm Triton Ops (ztxz16/fastllm, 5.1k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Compileiq Run Search?

NVIDIA (a GitHub organization, an official publisher) maintains it in NVIDIA/CompileIQ, which has 138 GitHub stars. The repository holds 7 skills in this directory. The repository was last updated on September 23, 2026.

Source: NVIDIA/CompileIQ on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.