Official agent skill

Warp Compile Time Optimizer

by NVIDIA in NVIDIA/skills

A skill your agent uses when compile time or startup time is the problem in code that uses Warp: a request to improve, optimize, or cut compile times; an app that is slow to start or stalls at the…

OfficialApache-2.0Auto-check: notesAI & LLM Engineering

Install Warp Compile Time Optimizer

skills CLI
$ npx skills add NVIDIA/skills --skill warp-compile-time-optimizer -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install NVIDIA/skills warp-compile-time-optimizer --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/warp-compile-time-optimizer .claude/skills/warp-compile-time-optimizer && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
warp-compile-time-optimizer
GitHub stars
3.5k
Token cost
~3.5k tokens
SKILL.md length
1,701 words
Files
107 (incl. scripts, references)
Skills in repo
380
Repo updated
First seen
Licence
Apache-2.0

At a glance

A skill your agent uses when compile time or startup time is the problem in code that uses Warp: a request to improve, optimize, or cut compile times; an app that is slow to start or stalls at the…

  • Works in 6 steps: Find the real cost, and confirm it is… → Ask about runtime tradeoffs when needed → Diagnose from the measurement, not from… → …
  • Startup time is the problem in code that uses Warp: a request to improve
  • SKILL.md covers Start with a runnable command, Compilation model, When a module's options are… and Preserve behavior, plus 4 more sections
  • Runs Python scripts from its folder; calls python

What it does

Warp Compile Time Optimizer is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Use when compile time or startup time is the problem in code that uses Warp: a request to improve, optimize, or cut compile times; an app that is slow to start or stalls at the first wp.launch; seconds of compiling before real work begins; JIT modules recompiling on every run or every CI job. Only applies when the code being optimized uses Warp kernels. Not for steady-state kernel runtime, memory, correctness, building Warp itself from source, or nvcc/C++ build times.

Its SKILL.md is about 3.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 114 other files, including scripts and reference files (for example `BENCHMARK.md`, `config/skillspector-baseline.yaml` and `evals/config.yml`). Compatibility notes: Requires Python 3.10+ and an installed warp-lang package. A CUDA device is needed to diagnose CUDA-specific mechanisms.

It sits in AI & LLM Engineering. It works with CUDA and C++. The repository describes itself as: Agent Skills for NVIDIA products — install into Claude Code, Codex, and other coding agents to run Physical AI, robotics, simulation, CUDA, and RAG workflows end to end. The licence is Apache-2.0.

When your agent uses it

  • Startup time is the problem in code that uses Warp: a request to improve
  • Cut compile times
  • An app that is slow to start
  • Stalls at the first wp.launch

Example prompts

  • “/warp-compile-time-optimizer”

Requirements

  • Python 3
  • Compatibility (from SKILL.md): Requires Python 3.10+ and an installed warp-lang package. A CUDA device is needed to diagnose CUDA-specific mechanisms.
  • Pre-approved tools (allowed-tools): Bash, Read, Edit, Write, Glob, Grep, env

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Find the real cost, and confirm it is compilation
  2. Ask about runtime tradeoffs when needed
  3. Diagnose from the measurement, not from reading the source
  4. Choose module boundaries deliberately
  5. Verify
  6. Report results

What it can do on your machine

Read from SKILL.md and the folder at commit 0e0d506. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Bash
    • Read
    • Edit
    • Write
    • Glob
    • Grep
    • env

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python, from the files we listed), which the agent can run.

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Requires Python 3.10+ and an installed warp-lang package. A CUDA device is needed to diagnose CUDA-specific mechanisms.

    From compatibility in the SKILL.md frontmatter.

Context cost

Warp Compile Time Optimizer loads about 3.5k tokens when it runs, and up to ~14k if it reads all its reference files. Until then it costs about 125 tokens; SKILL.md has 1,701 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~125
When it runs · the whole SKILL.md, loaded when a task matches
~3.5k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~14k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Bash, Read, Edit, Write, Glob, Grep, env

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from NVIDIA/skills at commit 0e0d506, republished under its Apache-2.0 licence (© NVIDIA). 1,701 words, ~3,495 tokens.

Download SKILL.mdSave it as .claude/skills/warp-compile-time-optimizer/SKILL.md (or your agent's skills folder). This skill also uses 106 other files; get the full folder from GitHub.
name
warp-compile-time-optimizer
description
Use when compile time or startup time is the problem in code that uses Warp: a request to improve, optimize, or cut compile times; an app that is slow to start or stalls at the first wp.launch; seconds of compiling before real work begins; JIT modules recompiling on every run or every CI job. Only applies when the code being optimized uses Warp kernels. Not for steady-state kernel runtime, memory, correctness, building Warp itself from source, or nvcc/C++ build times.
allowed-tools
Bash, Read, Edit, Write, Glob, Grep, env
compatibility
Requires Python 3.10+ and an installed warp-lang package. A CUDA device is needed to diagnose CUDA-specific mechanisms.
license
Apache-2.0
metadata.version
0.1.0
metadata.author
Warp Team <warp-python@nvidia.com>
metadata.tags
warp, compilation, cold-start, startup-latency, kernel-cache, gpu

Warp cold-start compile time

Start with a runnable command

The probe runs the target in a subprocess, writes only to temporary directories, and needs no network, external tool servers, or Warp checkout.

ScriptPurposeArguments
scripts/warp_compile_probe.pyMeasure isolated cold/warm compilation and launches.measure [OPTIONS] -- COMMAND...; use --help.

Use run_script("scripts/warp_compile_probe.py", args=[...]) when supported; otherwise use the Python command below. The target must run to completion.

Compilation model

Warp compiles modules, not individual kernels. A module's identity is:

text
(live kernel & function set) x (module options) x (CUDA block_dim) x (generic instances)

Each identity requires code generation and native compilation for the full module.

Cold-start cost is roughly:

text
number of distinct module identities you touch  x  size of each module

Reduce it in two ways:

  1. Stop identity churn. A module that changes after loading compiles again.
  2. Stop source duplication. A module used at three block dimensions compiles every kernel three times.

Deleting one kernel from a module that still builds saves only part of one build. Removing an unnecessary module identity saves the full build.

When neither applies, overlap independent CUDA builds (CS-13). This changes when work happens, not how much is compiled, so judge it on elapsed time.

When a module's options are fixed

Options have two deadlines:

  1. Module creation (Python import). The module copies enable_backward, max_unroll, lineinfo, deterministic, deterministic_max_records, and compile_time_trace from warp.config. Setting a global later is silently ignored by that module. default_grid_stride is the exception.
  2. First load. Before compilation, change the existing module with wp.set_module_options() or wp.get_module(name).options. Changing an option after load creates a new identity and rebuilds the module (CS-3).
Missed deadlineSymptomCost
wp.config.* set after importhash unchanged, option silently absentthe entire benefit, invisibly
module options set after loada second hash, module builds twiceone extra build, visible in the trace

After changing an option, confirm the hash moved for every target module. An unchanged hash means the option never arrived.

Preserve behavior

Change how Warp compiles the code, not the workload.

Do not delete or merge kernels to claim a gain. Apparently redundant stages may preserve ownership, aliasing, retained outputs, numerical boundaries, or API behavior. Fix duplication at the module level.

Preserve every launch and its order, dimensions, dtypes, devices, block dimensions, gradients, numerical modes, dynamic/plugin behavior, and public API signatures. Keep kernel names when moving definitions to module scope because logs, cache artifacts, and external tools expose them.

Instructions

1. Find the real cost, and confirm it is compilation

Ask what command the user actually waits on, then measure it cold:

bash
python scripts/warp_compile_probe.py measure --samples 3 \
    --json baseline.json -- <the user's command>

The probe gives each sample private WARP_CACHE_PATH, WARP_CACHE_ROOT, and CUDA_CACHE_PATH directories, enables module timers, and records launches. It creates those directories under the system temporary directory and removes them itself, so isolating a sample never requires writing a cache into the project or deleting anything to re-measure. Isolate a hand-rolled sample the same way. Never clear a live cache with wp.clear_kernel_cache() or wp.clear_lto_cache(); clearing is not isolated and can disrupt other processes.

Read the probe output before source. If compilation is a small part of wall time, report the real bottleneck and stop. For libraries and tests, use the smallest command that compiles the workload's modules.

Modules that each compiled once, with no repeated hashes, block-dimension variants, or LTO, have no structural churn. This rules out redundant builds, not oversized builds; still check cache reuse (CS-2), backward codegen (CS-10), unrolling (CS-11), the precompiled header (CS-12), and overlap when several CUDA modules remain (CS-13).

Every sample also re-runs the command against the cache it just populated. If warm module work is not near zero, diagnose cache reuse (CS-2) before changing module structure.

2. Ask about runtime tradeoffs when needed

Apply ordering fixes, lifecycle grouping, and option hoists without asking. Ask before changing fast_math, max_unroll, or a MathDx/tile implementation:

Some of these knobs cut compile time but can make the compiled kernels slower or change numerics. Are you optimizing a fast edit-run loop (where slower kernels are usually fine), or production startup (where they usually are not)?

If the user is unavailable:

  • Leave numerics-changing options and implementation swaps alone.
  • Before removing a capability such as backward codegen, test whether it is used; report what is removed, the measured benefit, and how to revert.
  • Set global wp.config.* options at application entry points, not in library code.

Record declined options and their measured benefits in the step 6 ledger.

Match the scope of the change to the scope of the evidence

A profile supports a change to the measured application, not every consumer of a shared library. Repository searches also miss out-of-tree and future callers. For example, a forward-only application does not justify disabling gradients inside a solver library that another application differentiates through.

Scope the option to the measured process, before importing the library:

python
import warp as wp

wp.config.enable_backward = False   # must precede the library import

import the_library

This also reaches every module the application loads. CS-10 covers the silent import-order trap. If only a library change works, send its maintainers the measurement and let them decide the contract.

3. Diagnose from the measurement, not from reading the source

The probe prints every compiled module identity with its name, hash, device, and block dimension, then names which modules built more than once. Match what you see:

What the probe showsWhat it meansWhere to look
One module name, several hashesIdentity churn: its kernel set, options, or generic instances changed after it first loadedCS-1, CS-3, CS-6
One module name, several block_dim valuesThe whole module is recompiled per block dimension (CUDA)CS-5
Many one-kernel modules in one featureFixed per-module cost repeatedCS-4
A hash-named module per kernelmodule="unique" used on stable kernelsCS-9
Big gap between module time and native compile time, plus .lto artifactsMathDx/LTO setupCS-7
(compiled) on a run that should have been warmCache is not being reusedCS-2
Modules load, then "Failed to find module"Concurrent CPU JIT first-use raceCS-8
Large generated source, no rebuild problemUnroll budgetCS-11
Adjoint code in a module nothing differentiatesBackward codegenCS-10
Compiles slow across the board, or a few small modules on CUDA below toolkit 13The precompiled header is turned off, or is not paying for itselfCS-12
Several independent modules, each built once, overlap_factor near 1.0Builds are running one at a time; parallel loading is off by defaultCS-13
An option you set changed nothing, and that module's hash is unchangedIt was assigned after the module was created, so it never arrived"When a module's options are fixed"
No row above firesNothing is being built redundantly; the cost is the size of the builds themselvesStep 6

references/mechanisms.md has one section per mechanism: how to confirm it, the fix, its limits, and its failure mode. Read only the sections selected by the measurement.

Show full SKILL.md (605 more words)Show less
4. Choose module boundaries deliberately

Group kernels in one module only when they share:

  • lifecycle: they are defined, loaded, and invalidated together;
  • option set: they need the same fast_math, enable_backward, max_unroll, and MathDx settings;
  • stable block dimension on CUDA.

Kernels with the same lifecycle but different stable block dimensions should not share a module because each would compile twice. Separate kernels with independent lifecycles too.

Kernels whose block dimension varies at runtime (chosen from input size, say) have no stable mapping, so keep them in their own module rather than dragging a whole shared module into an extra variant.

Prefer the least invasive change that removes a build. Ordering fixes and option hoists are cheaper and safer than re-architecting module layout; regroup only when fixed per-module cost or block-dimension duplication dominates.

wp.set_module_options() targets its calling Python module, not kernels with an explicit module="pkg.name". Either use a real Python module or update the named module before it loads:

python
wp.set_module_options({"enable_backward": False})  # at module scope

wp.get_module("pkg.name").options.update({"enable_backward": False})

Do not pass wp.get_module() to wp.set_module_options(module=...), or use @wp.kernel(module_options={...}) without module="unique". Per-kernel enable_backward=False has a tile-module exception covered by CS-10. Confirm the module hash after every option change.

5. Verify
bash
python scripts/warp_compile_probe.py measure --samples 3 \
    --json candidate.json -- <the same command>
python scripts/warp_compile_probe.py compare baseline.json candidate.json

compare rejects changed launch topology and treats a result inside max(1% of baseline, 2 x baseline MAD) as inconclusive.

For BUILDS OVERLAPPED, judge scheduling changes on compile elapsed rather than summed module timers. The required warm pass supplies that clock. See references/measurement.md.

Then check what the probe cannot see:

  • Diff numeric output and run the project's tests or entry points.
  • After changing enable_backward or boundaries, verify a gradient path.
  • After changing fast_math, max_unroll, MathDx, or an implementation, benchmark steady-state runtime.
  • Exercise uncovered dynamic kernels, dtypes, profiles, and modes.
6. Report results

Read "Reporting results" in references/measurement.md. Report:

Describe every optimization in plain language: name the behavior, the evidence, and the effect. For example, write "moved module options before the first load to avoid a redundant rebuild," not "applied CS-3." Treat CS-* labels as internal navigation aids, not user-facing explanations.

OptionMeasuredWhy not takenTo take it
enable_backward=False on pkg.solver−38% colda live tape traverses these kernelsset at the entry point, then re-check adjoints
max_unroll=4−2%, inside noisechanges generated code for no measured gain—
  • before/after medians and sample counts;
  • the reduction and residual cost, sized against the original complaint;
  • mechanisms fixed, ruled out, measured and declined, or not reached;
  • tradeoffs and behavior not verified;
  • the ledger above for declined or incompletely verified options;
  • the measured value of relaxing any constraint that blocked a fix.

State only what the evidence supports. "No structural churn" does not mean "optimal" or "irreducible." Measure declined levers when practical; label any estimate untested. Before reporting no available fix, check CS-13. For a possible module split, first measure a one-kernel module with the same options to establish the repeated fixed cost.

Troubleshooting

Run the target command directly before debugging the probe. See references/measurement.md for cache/noise issues and references/mechanisms.md for mechanism-specific failures.

Limitations

Two rules override any gain:

  • Isolate both Warp and CUDA caches for every cold sample.
  • Keep max_workers <= 1 when a load can target CPU, including device=None and mixed device lists. Concurrent CPU first loads can lose kernels; retries do not make them safe. CUDA-only loading is unaffected.

Measurements are environment-specific: cold times move with CPU, GPU, driver, toolchain, and Warp version. Mechanisms transfer; numbers do not. Read the known unknowns in references/mechanisms.md before making broad claims.

Reference files

  • references/mechanisms.md: the thirteen compile-time mechanisms, each with its confirming signal, fix, applicability limits, and failure mode. Read the sections your measurement points to.
  • references/measurement.md: measurement protocol, what each metric does and does not mean, reporting guidance, log examples, and manual measurement.

© NVIDIA, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 106 other files (scripts, references) in skills/warp-compile-time-optimizer of NVIDIA/skills.

  • SKILL.md
  • BENCHMARK.md
  • config/skillspector-baseline.yaml
  • evals/config.yml
  • evals/evals.json
  • evals/files/c01_particle_preprocessing/pyproject.toml
  • evals/files/c01_particle_preprocessing/src/particle_prep/__init__.py
  • evals/files/c01_particle_preprocessing/src/particle_prep/__main__.py
  • evals/files/c01_particle_preprocessing/src/particle_prep/pipeline.py
  • evals/files/c01_particle_preprocessing/src/particle_prep/plugins.py
  • evals/files/c01_particle_preprocessing/src/particle_prep/stages.py
  • evals/files/c04_batch_signals/pyproject.toml
  • evals/files/c04_batch_signals/src/batch_signals
  • … and 94 more

Open the folder on GitHubat commit 0e0d506

Compare with similar skills

Warp Compile Time Optimizer next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Warp Compile Time Optimizer compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Warp Compile Time Optimizer this skillNVIDIA/skills3.5k—~3.5kAutomated safety check: NotesApache-2.0
Paddle BuildPaddlePaddle/Paddle24k—~1kAutomated safety check: PassApache-2.0
Fastllm Triton Opsztxz16/fastllm5.1k—~1.8kAutomated safety check: PassApache-2.0
Ako4allTongmingLAIC/AKO4ALL369—~4kAutomated safety check: PassMIT
Cuda Cpp Kernelvipshop/cache-dit1.3k—~2.3kAutomated safety check: PassApache-2.0
Paddle Op DevPaddlePaddle/Paddle24k—~1.3kAutomated safety check: PassApache-2.0

Similar skills

  • Paddle Build

    PaddlePaddle/Paddle

    A skill your agent uses when needing to compile, rebuild, or install Paddle from source after code changes.

    24k GitHub stars~1k tokensUpdated 7 days ago
    AI & LLM EngineeringAuto-check passed
  • Fastllm Triton Ops

    ztxz16/fastllm

    Guide for adding Triton-backed CUDA operators to FastLLM. An agent skill from ztxz16/fastllm.

    5.1k GitHub stars~1.8k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Ako4all

    TongmingLAIC/AKO4ALL

    Drive an agentic loop that iteratively optimizes a GPU kernel for maximum speedup.

    369 GitHub stars~4k tokensUpdated 22 days ago
    AI & LLM EngineeringAuto-check passed
  • Cuda Cpp Kernel

    vipshop/cache-dit

    A skill your agent uses when writing, debugging, porting, reviewing, or optimizing CUDA C++ or PTX kernels; investigating CUDA Runtime or Driver API behavior; profiling kernels with Nsight Systems…

    1.3k GitHub stars~2.3k tokensUpdated 8 days ago
    AI & LLM EngineeringAuto-check passed
  • Paddle Op Dev

    PaddlePaddle/Paddle

    PaddlePaddle (飞桨) C++ 算子开发指南。提供从 YAML 配置、InferMeta 函数、Kernel 实现、Python API 封装、单元测试到编译验证的完整算子开发流程指导。在以下场景使用此 skill:(1) 为 Paddle 框架新增 C++ 算子 (2) 修改或调试已有 Paddle 算子 (3) 编写算子的 YAML…

    24k GitHub stars~1.3k tokensUpdated 7 days ago
    AI & LLM EngineeringAuto-check passed
  • Embedded AI Deployment

    matlab/agent-skills-playground

    Deploy AI models to embedded hardware using MathWorks tools (MATLAB, Simulink, Embedded Coder).

    181 GitHub starsUsed in 1 repo~3.4k tokens
    AI & LLM EngineeringAuto-check passed

More from NVIDIA/skills

All 380 skills in this repo
  • Official

    A skill your agent uses when the user wants to deploy, run, debug, tear down, or call the REST API of the RTVI-CV 2D detection / tracking microservice.

    3.5k GitHub starsUsed in 1 repo~4.5k tokens
    Auto-check passed
  • Official

    Generates, validates, compares and explains HOLOLINK_def.svh macro files for the HSB IP, using bundled Python scripts and asking before it writes anything.

    3.5k GitHub stars~2.9k tokensUpdated today
    Auto-check passed
  • Official

    Runs and validates an end-to-end Mission Control demo in a locally installed Isaac Sim, with a Nova Carter robot driven through a Python server.

    3.5k GitHub stars~4.8k tokensUpdated today
    Auto-check passed
  • Orchestrates defect image generation for PCBA, metal surface and glass inspection with NVIDIA Cosmos AnomalyGen on OSMO, from cold-start Day 0 to real-photo Day 1 labeling.

    3.5k GitHub stars~5k tokensUpdated today
    Auto-check: notes
  • Orchestrates video data augmentation and auto-labeling workflows on OSMO, from flow selection and preflight checks to submission, monitoring and output download.

    3.5k GitHub stars~4.7k tokensUpdated today
    Auto-check: notes
  • Official

    Runs NVIDIA TAO Data Services KPI analysis on object detection results, comparing predictions to ground truth and writing per-class precision, recall and AP to a CSV.

    3.5k GitHub stars~2.7k tokensUpdated today
    Auto-check: notes

Works with

Questions about Warp Compile Time Optimizer

What does Warp Compile Time Optimizer do?

A skill your agent uses when compile time or startup time is the problem in code that uses Warp: a request to improve, optimize, or cut compile times; an app that is slow to start or stalls at the…. Warp Compile Time Optimizer is an agent skill from NVIDIA/skills, published by the product's own GitHub organization.launch; seconds of compiling before real work begins; JIT modules recompiling on every run or every CI job.

When should I use Warp Compile Time Optimizer?

Warp Compile Time Optimizer fits situations like: startup time is the problem in code that uses Warp: a request to improve; cut compile times; an app that is slow to start; stalls at the first wp.launch.

How do I install Warp Compile Time Optimizer in Claude Code?

Run `npx skills add NVIDIA/skills --skill warp-compile-time-optimizer -a claude-code`. Or copy the skill folder (skills/warp-compile-time-optimizer in NVIDIA/skills) into .claude/skills/warp-compile-time-optimizer in your project. Claude Code loads it when a task matches its description.

How do I install Warp Compile Time Optimizer in Codex?

Run `npx skills add NVIDIA/skills --skill warp-compile-time-optimizer -a codex`. Or copy the skill folder (skills/warp-compile-time-optimizer in NVIDIA/skills) into .agents/skills/warp-compile-time-optimizer in your project. Codex loads it when a task matches its description.

Can I use Warp Compile Time Optimizer in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NVIDIA/skills --skill warp-compile-time-optimizer -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/warp-compile-time-optimizer, .gemini/skills/warp-compile-time-optimizer, .github/skills/warp-compile-time-optimizer and .opencode/skills/warp-compile-time-optimizer in your project.

What does Warp Compile Time Optimizer need to run?

Going by SKILL.md and its folder, Warp Compile Time Optimizer needs Python for the scripts in its folder and the command-line tools its instructions call (python). Our summary lists: Python 3. Its frontmatter pre-approves these tools: Bash, Read, Edit, Write, Glob, Grep, env. Compatibility (from SKILL.md): Requires Python 3.10+ and an installed warp-lang package. A CUDA device is needed to diagnose CUDA-specific mechanisms..

Does Warp Compile Time Optimizer access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Warp Compile Time Optimizer safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Warp Compile Time Optimizer use?

Warp Compile Time Optimizer is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Warp Compile Time Optimizer use?

About 3.5k tokens (SKILL.md is roughly 14k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 11k tokens, read only when the agent opens those files.

What are the alternatives to Warp Compile Time Optimizer?

Skills that share tags, products or a category with Warp Compile Time Optimizer: Paddle Build (PaddlePaddle/Paddle, 24k stars), Fastllm Triton Ops (ztxz16/fastllm, 5.1k stars), Ako4all (TongmingLAIC/AKO4ALL, 369 stars) and Cuda Cpp Kernel (vipshop/cache-dit, 1.3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Warp Compile Time Optimizer?

NVIDIA (a GitHub organization, an official publisher) maintains it in NVIDIA/skills, which has 3,534 GitHub stars. The repository holds 380 skills in this directory. The repository was last updated on October 7, 2026.

Source: NVIDIA/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.