Official agent skill

Compileiq Validate Result

by NVIDIA in NVIDIA/CompileIQ

Use AFTER a Search has completed and BEFORE claiming any speedup or shipping an ACF.

OfficialApache-2.0Auto-check: notesDocuments & Office

Install Compileiq Validate Result

skills CLI
$ npx skills add NVIDIA/CompileIQ --skill compileiq-validate-result -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install NVIDIA/CompileIQ compileiq-validate-result --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/NVIDIA/CompileIQ.git skills-src && mkdir -p .claude/skills && cp -r skills-src/agent-skills/compileiq-validate-result .claude/skills/compileiq-validate-result && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
compileiq-validate-result
GitHub stars
137
Token cost
~2.2k tokens
SKILL.md length
578 words
Files
2 (incl. scripts)
Skills in repo
7
Repo updated
First seen
Licence
Apache-2.0

At a glance

Use AFTER a Search has completed and BEFORE claiming any speedup or shipping an ACF.

  • Works in 7 steps: Load the CSV → Extract candidates → Re-measure on fresh cache (the actual… → …
  • Validate result
  • SKILL.md covers When, Steps, CLI helper and Self-test, plus 2 more sections
  • Runs Python scripts from its folder; calls python

What it does

Compileiq Validate Result is an agent skill from NVIDIA/CompileIQ, published by the product's own GitHub organization. Use AFTER a Search has completed and BEFORE claiming any speedup or shipping an ACF. Loads the dumpresults CSV, extracts top-K candidates (single-objective) or the Pareto front (multi-objective), re-measures each against the no-ACF baseline with 100+ trials on fresh caches, runs Welch's t-test plus Cohen's d, rejects three classic false-positive patterns (lucky-min / higher-variance / multiple-comparisons-of-N), and saves the validated winner as best.acf. Triggers on "validate result", "extract best config"…

Its SKILL.md is about 2.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including scripts (for example `scripts/welch_validate.py`).

It sits in Documents & Office, covering CSV and tabular files. The repository describes itself as: An Optimizer for Nvidia Compilers. The licence is Apache-2.0.

When your agent uses it

  • Validate result
  • Extract best config
  • Is my speedup real

Example prompts

  • “s t-test plus Cohen”
  • “validate result”
  • “extract best config”
  • “/compileiq-validate-result”

Requirements

  • Python 3
  • Pre-approved tools (allowed-tools): Bash, Read

Workflow steps

7 steps, taken from the step headings in SKILL.md.

  1. Load the CSV
  2. Extract candidates
  3. Re-measure on fresh cache (the actual validation)
  4. Statistical gate — the ship rule
  5. Three false-positive patterns to actively check
  6. Save the validated winner
  7. Reproducibility log

What it can do on your machine

Read from SKILL.md and the folder at commit 743aca4. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Bash
    • Read

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Compileiq Validate Result loads about 2.2k tokens when it runs. Until then it costs about 161 tokens; SKILL.md has 578 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~161
When it runs · the whole SKILL.md, loaded when a task matches
~2.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Bash, Read

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from NVIDIA/CompileIQ at commit 743aca4, republished under its Apache-2.0 licence (© NVIDIA). 578 words, ~2,167 tokens.

Download SKILL.mdSave it as .claude/skills/compileiq-validate-result/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
compileiq-validate-result
description
Use AFTER a Search has completed and BEFORE claiming any speedup or shipping an ACF. Loads the dump_results CSV, extracts top-K candidates (single-objective) or the Pareto front (multi-objective), re-measures each against the no-ACF baseline with 100+ trials on fresh caches, runs Welch's t-test plus Cohen's d, rejects three classic false-positive patterns (lucky-min / higher-variance / multiple-comparisons-of-N), and saves the validated winner as best.acf. Triggers on "validate result", "extract best config", "Welch's t-test", "is my speedup real", "save best ACF", "pareto front", "claim speedup", "ship config".
allowed-tools
Bash, Read
when_to_use
- tuner.start() returned and there's a results CSV. - User wants to ship an ACF to production. - Reported speedup feels too good to be true. Don't use when…
license
Apache-2.0
metadata.version
1.0.0
metadata.author
NVIDIA CompileIQ
metadata.domain
compiler-optimization
paths
**/*.csv, **/*.acf, **/*.py

compileiq-validate-result

The score CompileIQ reports during a search uses N=5-15 trials per evaluation and a shared cache. That's appropriate for the search loop but wildly insufficient for shipping. This skill is the gate before any ACF goes to production.

When

  • tuner.start() has returned and there's a dump_results= CSV on disk.
  • User wants to claim a speedup or ship an ACF.
  • A reported speedup feels too clean — validate it.

Steps

1. Load the CSV
python
from compileiq.results import SearchResult

results = SearchResult.from_csv("results.csv", problem_type="min", clear_duplicates=True)
df = results.get_results()
print(f"{len(df)} evaluations across {df['generation'].max()+1} generations")
2. Extract candidates

Single-objective:

python
best = results.get_best_result()
# dict: {metadata, generation, score_1, params, [norm_score_1]}
score = best.get("score_1", best.get("score"))   # legacy defensiveness
acf_hex = best["params"]                          # hex string

score_1 (with underscore-one) is the canonical key — it matches the multi-objective convention score_N. Older code sometimes uses plain score; the fallback above handles both shapes.

For top-K:

python
import pandas as pd
df_valid = df[pd.to_numeric(df["score_1"], errors="coerce") < 1e10]
top_k = df_valid.nsmallest(5, "score_1")          # nlargest for MAX problems

Multi-objective:

python
front = results.pareto_front()   # raises if num_objectives == 1
for candidate in front:
    print(candidate["score_1"], candidate["score_2"], candidate["params"])

Mixed user+compiler search space: results carry separate keys — best["user_space"] for the user-side knobs, best["params"] for the ACF hex. Save both.

3. Re-measure on fresh cache (the actual validation)
StageWarmupTrialsCacheGPU clocks
Optimization (during tuner.start())5-255-15per-evalrecommended locked
Validation≥50≥100per-measurementmust be locked

Both the baseline (no ACF) and each top-K candidate are re-measured at validation N. The optimization-time measurement is too noisy to ship from.

4. Statistical gate — the ship rule
python
import numpy as np
from scipy import stats

def validate_speedup(baseline_ms: np.ndarray, optimized_ms: np.ndarray) -> dict:
    t, p = stats.ttest_ind(baseline_ms, optimized_ms, equal_var=False)   # Welch's
    b_mean, b_std = baseline_ms.mean(),  baseline_ms.std(ddof=1)
    o_mean, o_std = optimized_ms.mean(), optimized_ms.std(ddof=1)
    pooled = np.sqrt((b_std**2 + o_std**2) / 2)
    d = (b_mean - o_mean) / pooled if pooled > 0 else 0.0
    return {
        "speedup_mean":  b_mean / o_mean,
        "speedup_median": np.median(baseline_ms) / np.median(optimized_ms),
        "p_value":       float(p),
        "cohens_d":      float(d),
        "significant":   bool(p < 0.05 and o_mean < b_mean and d > 0.2),
        "baseline":  {"mean": b_mean, "std": b_std,
                      "p5": np.percentile(baseline_ms,  5),
                      "p95": np.percentile(baseline_ms, 95)},
        "optimized": {"mean": o_mean, "std": o_std,
                      "p5": np.percentile(optimized_ms,  5),
                      "p95": np.percentile(optimized_ms, 95)},
    }

Ship rule: p_value < 0.05 AND cohens_d > 0.2 (preferably > 0.5) AND optimized.mean < baseline.mean. Anything weaker, do not claim a speedup.

5. Three false-positive patterns to actively check
#PatternSymptomCauseCheckDisposition
1Lucky-minOptimized min is lower but mean is equal or worseOptimizer picked a config that occasionally runs fastCompare means, not minimums; reject if optimized.mean ≥ baseline.meanReject.
2Higher-varianceOptimized p5-p95 range is wider than baseline with same meanACF didn't speed anything up; just spread the distributionCompute (p95 - p5) for both; reject if optimized range is materially wider (>25%)Reject.
3Multiple-comparisonsBest of 500 evaluations looks 2-5% faster but doesn't reproduceWith 500 evals some will look good by chanceRe-measure top-K on a fresh cache and fresh trials; reject candidates that don't surviveReject.
6. Save the validated winner
python
from compileiq.utils.helpers import save_compiler_config
save_compiler_config("best.acf", best["params"])

# Mixed search spaces: persist the user_space knobs separately
if "user_space" in best:
    import json
    Path("best.user_space.json").write_text(json.dumps(best["user_space"], indent=2))
Show full SKILL.md (259 more words)Show less
7. Reproducibility log

Append one row per candidate decision to validation-log.csv. Fields, per docs/flashinfer_booster.md:135-148:

  • timestamp (UTC ISO 8601)
  • ACF filename + sha256
  • manifest / release version
  • benchmark command
  • GPU model + driver version
  • CTK version (nvcc release)
  • ptxas, nvcc paths + versions
  • framework version or commit (Triton / Helion / FlashInfer / cuTeDSL)
  • input shape
  • baseline mean ± std
  • candidate mean ± std
  • p-value
  • Cohen's d
  • decision: KEPT or REJECTED:<reason>

The scripts/welch_validate.py helper records the timing/statistical fields, ACF hash, benchmark commands, GPU/toolchain metadata, and common environment variables automatically. Pass --manifest, --framework, and --input-shape for workload-specific fields the helper cannot infer.

CLI helper

bash
python scripts/welch_validate.py \
    --acf best.acf \
    --baseline-cmd "python bench.py --routine matmul" \
    --opt-cmd "PTXAS_OPTIONS='--apply-controls=best.acf' python bench.py --routine matmul" \
    --trials 100 --warmup 50 \
    --score-regex 'mean: ([0-9.]+)' \
    --manifest booster-packs-YYYY.MM.DD \
    --framework "flashinfer <version>" \
    --input-shape "routine=matmul, M=..., N=..., K=..." \
    --output validation-log.csv

Prints KEPT or REJECTED:<reason> and appends a row to the log. Also importable: from welch_validate import validate_speedup.

Self-test

bash
python scripts/welch_validate.py --self-test

Synthesizes two identical normal distributions, asserts the statistical gate returns significant=False. Then differs them, asserts significant=True. Catches misconfigured scipy/numpy before a real validation.

Gotchas

  • pareto_front() raises if num_objectives == 1. Guard with if results.num_scores > 1: or use try/except.
  • score_1 vs score. Current API is score_1. Some older results exporters used plain score. The defensive read pattern best.get("score_1", best.get("score")) handles both.
  • Don't validate on the same cache the search used. With CIQ_KEEP_CACHE=1 active during search, validation must explicitly wipe ~/.cache/compileiq or use a fresh TRITON_CACHE_DIR and HELION_SKIP_CACHE=1. Otherwise the optimization-time numbers re-appear and you're not validating anything.
  • Validation N is independent of optimization N. Even if the search used N=5 per evaluation, validation needs N ≥ 100. Don't try to be clever and reuse search-time samples.

Next

  • If the validated speedup ships: commit best.acf and validation-log.csv.
  • If validation fails: compileiq-debug for diagnosis.
  • For more thorough exploration: re-run compileiq-run-search with bigger pool_size/generations.

© NVIDIA, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (scripts) in agent-skills/compileiq-validate-result of NVIDIA/CompileIQ.

  • SKILL.md
  • scripts/welch_validate.py

Open the folder on GitHubat commit 743aca4

Compare with similar skills

Compileiq Validate Result next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Compileiq Validate Result compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Compileiq Validate Result this skillNVIDIA/CompileIQ137—~2.2kAutomated safety check: NotesApache-2.0
Data Table Managern8n-io/n8n207k—~2.3kAutomated safety check: PassCustom licence
Instrument Data To Allotropeaws-samples/amazon-bedrock-agents-healthcare-lifesciences2742 repos~2.7kAutomated safety check: PassApache-2.0
Abuse Hunternexu-io/harness-engineering-guide663—~1.9kAutomated safety check: PassMIT
Markitshift-labs-ai/markit1.3k—~299Automated safety check: PassMIT
Sector Analysttradermonty/claude-trading-skills3k1 repos~2.3kAutomated safety check: PassMIT

Similar skills

  • Official

    Load before calling data-tables or parse-file. An agent skill from n8n-io/n8n.

    207k GitHub stars~2.3k tokensUpdated today
    Documents & OfficeAuto-check passed
  • Instrument Data To Allotrope

    aws-samples/amazon-bedrock-agents-healthcare-lifesciences

    Official

    Convert laboratory instrument output files (PDF, CSV, Excel, TXT) to Allotrope Simple Model (ASM) JSON format or flattened 2D CSV.

    274 GitHub starsUsed in 2 repos~2.7k tokens
    Documents & OfficeAuto-check passed
  • Abuse Hunter

    nexu-io/harness-engineering-guide

    Detect and investigate bulk registration abuse on SaaS platforms.

    663 GitHub stars~1.9k tokensUpdated 5 mo ago
    Documents & OfficeAuto-check passed
  • Markit

    shift-labs-ai/markit

    Convert files and URLs to Markdown. An agent skill from shift-labs-ai/markit.

    1.3k GitHub stars~299 tokensUpdated 1 mo ago
    Documents & OfficeAuto-check passed
  • Sector Analyst

    tradermonty/claude-trading-skills

    This skill should be used when analyzing sector rotation patterns and market cycle positioning.

    3k GitHub starsUsed in 1 repo~2.3k tokens
    Documents & OfficeAuto-check passed
  • Uap Release Analyzer

    ckpxgfnksd-max/uap-release-analyzer

    Inventory, extract, and analyze tranches of declassified UAP/UFO files — including war.gov/UFO/ "PURSUE" releases, FBI Vault, NARA boxes, and AARO publications.

    155 GitHub stars~2.6k tokensUpdated 5 mo ago
    Documents & OfficeAuto-check passed

More from NVIDIA/CompileIQ

  • Compileiq Booster Pack

    NVIDIA/CompileIQ

    Official

    Use BEFORE running a full CompileIQ search. An agent skill from NVIDIA/CompileIQ.

    137 GitHub stars~2.1k tokensUpdated 14 days ago
    Auto-check: notes
  • Compileiq Bootstrap

    NVIDIA/CompileIQ

    Official

    A skill your agent uses when starting a fresh CompileIQ project, hitting a socket timeout, or before running any other compileiq- skill.

    137 GitHub stars~1.3k tokensUpdated 14 days ago
    Auto-check: notes
  • Compileiq Debug

    NVIDIA/CompileIQ

    Official

    A skill your agent uses when something is wrong: Search() hangs, all evaluations return INVALIDSCORE, scores aren't improving, every config returns the same number, ptxas errors fill the log, CV% is…

    137 GitHub stars~2.9k tokensUpdated 14 days ago
    Auto-check: notes
  • Compileiq Run Search

    NVIDIA/CompileIQ

    Official

    A skill your agent uses when composing the Search(...) call and calling .start().

    137 GitHub stars~2.2k tokensUpdated 14 days ago
    Auto-check: notes
  • Official

    A skill your agent uses when writing the objectivefunction= passed to Search().

    137 GitHub stars~2.3k tokensUpdated 14 days ago
    Auto-check: notes
  • Compileiq Search Space

    NVIDIA/CompileIQ

    Official

    A skill your agent uses when picking the searchspace= argument for Search().

    137 GitHub stars~1.9k tokensUpdated 14 days ago
    Auto-check: notes

Questions about Compileiq Validate Result

What does Compileiq Validate Result do?

Use AFTER a Search has completed and BEFORE claiming any speedup or shipping an ACF. Compileiq Validate Result is an agent skill from NVIDIA/CompileIQ, published by the product's own GitHub organization. Use AFTER a Search has completed and BEFORE claiming any speedup or shipping an ACF.

When should I use Compileiq Validate Result?

Compileiq Validate Result fits situations like: validate result; extract best config; is my speedup real.

How do I install Compileiq Validate Result in Claude Code?

Run `npx skills add NVIDIA/CompileIQ --skill compileiq-validate-result -a claude-code`. Or copy the skill folder (agent-skills/compileiq-validate-result in NVIDIA/CompileIQ) into .claude/skills/compileiq-validate-result in your project. Claude Code loads it when a task matches its description.

How do I install Compileiq Validate Result in Codex?

Run `npx skills add NVIDIA/CompileIQ --skill compileiq-validate-result -a codex`. Or copy the skill folder (agent-skills/compileiq-validate-result in NVIDIA/CompileIQ) into .agents/skills/compileiq-validate-result in your project. Codex loads it when a task matches its description.

Can I use Compileiq Validate Result in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NVIDIA/CompileIQ --skill compileiq-validate-result -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/compileiq-validate-result, .gemini/skills/compileiq-validate-result, .github/skills/compileiq-validate-result and .opencode/skills/compileiq-validate-result in your project.

What does Compileiq Validate Result need to run?

Going by SKILL.md and its folder, Compileiq Validate Result needs Python for the scripts in its folder and the command-line tools its instructions call (python). Our summary lists: Python 3. Its frontmatter pre-approves these tools: Bash, Read.

Does Compileiq Validate Result access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Compileiq Validate Result safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Compileiq Validate Result use?

Compileiq Validate Result is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Compileiq Validate Result use?

About 2.2k tokens (SKILL.md is roughly 8.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Compileiq Validate Result?

Skills that share tags, products or a category with Compileiq Validate Result: Data Table Manager (n8n-io/n8n, 207k stars), Instrument Data To Allotrope (aws-samples/amazon-bedrock-agents-healthcare-lifesciences, 274 stars), Abuse Hunter (nexu-io/harness-engineering-guide, 663 stars) and Markit (shift-labs-ai/markit, 1.3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Compileiq Validate Result?

NVIDIA (a GitHub organization, an official publisher) maintains it in NVIDIA/CompileIQ, which has 137 GitHub stars. The repository holds 7 skills in this directory. The repository was last updated on September 23, 2026.

Source: NVIDIA/CompileIQ on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.