Agent skill

Cuda Kernel Optimizer

by KernelFlow-ops in KernelFlow-ops/cuda-optimized-skill

Iteratively optimize a CUDA/CUTLASS/Triton kernel only when strict on-device compilation, correctness, timing, and NCU evidence gates pass.

MITAuto-check passedAI & LLM Engineering

Install Cuda Kernel Optimizer

skills CLI
$ npx skills add KernelFlow-ops/cuda-optimized-skill --skill cuda-kernel-optimizer -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install KernelFlow-ops/cuda-optimized-skill cuda-kernel-optimizer --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/KernelFlow-ops/cuda-optimized-skill.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/cuda-kernel-optimizer .claude/skills/cuda-kernel-optimizer && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
cuda-kernel-optimizer
GitHub stars
212
Token cost
~4.3k tokens
SKILL.md length
1,637 words
Files
53 (incl. scripts, references)
Skills in repo
1
Repo updated
First seen
Licence
MIT

At a glance

Iteratively optimize a CUDA/CUTLASS/Triton kernel only when strict on-device compilation, correctness, timing, and NCU evidence gates pass.

  • Works in 5 steps: Strict hardware gate → Initialize the run folder → Seed best with a baseline benchmark → …
  • Tasks that involve GPU and accelerator computing
  • SKILL.md covers What this skill does, Key point, Inputs the skill expects from… and The loop at a glance, plus 10 more sections
  • Calls python

What it does

Cuda Kernel Optimizer is an agent skill from KernelFlow-ops/cuda-optimized-skill. Iteratively optimize a CUDA/CUTLASS/Triton kernel only when strict on-device compilation, correctness, timing, and NCU evidence gates pass.

Its SKILL.md is about 4.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 56 other files, including scripts and reference files (for example `examples/walkthrough.md`, `references/method_registry.json` and `references/metric_registry.json`).

It sits in AI & LLM Engineering, covering GPU and accelerator computing. It works with CUDA. The repository describes itself as: A CUDA kernel optimization toolkit for validation, benchmarking, Nsight Compute profiling, bottleneck analysis, and iterative tuning. It helps improve custom GPU operators with… The licence is MIT.

When your agent uses it

  • Tasks that involve GPU and accelerator computing

Example prompts

  • “/cuda-kernel-optimizer”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Strict hardware gate
  2. Initialize the run folder
  3. Seed best with a baseline benchmark
  4. Iteration loop (up to N iterations)
  5. Final summary

What it can do on your machine

Read from SKILL.md and the folder at commit 114a6cb. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 10 files in scripts/, which the agent can run.

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Cuda Kernel Optimizer loads about 4.3k tokens when it runs, and up to ~37k if it reads all its reference files. Until then it costs about 40 tokens; SKILL.md has 1,637 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~40
When it runs · the whole SKILL.md, loaded when a task matches
~4.3k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~37k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from KernelFlow-ops/cuda-optimized-skill at commit 114a6cb, republished under its MIT licence (© KernelFlow-ops). 1,637 words, ~4,270 tokens.

Download SKILL.mdSave it as .claude/skills/cuda-kernel-optimizer/SKILL.md (or your agent's skills folder). This skill also uses 52 other files; get the full folder from GitHub.
name
cuda-kernel-optimizer
description
Iteratively optimize a CUDA/CUTLASS/Triton kernel only when strict on-device compilation, correctness, timing, and NCU evidence gates pass.

CUDA Kernel Iterative Optimizer (v3)

What this skill does

Given:

  • a baseline kernel file (.cu for CUDA / CUTLASS, or .py for Triton),
  • a reference Python file (exposes reference(**kwargs) — same contract as benchmark.py --ref),
  • optional atol/rtol and workload_model(**dims) -> {flops, bytes_min} in the reference,
  • kernel dimension arguments (e.g. --M=4096 --N=4096 --K=4096),
  • optional iteration count N (default 3), ncu_num (default 5), and branches (default 4),
  • optional --compile-jobs auto|N (default auto) and --numerics-mode reference|strict|approximate (default reference),

the skill runs a roofline-guided, branch-and-select iterative optimization loop and produces a timestamped directory of per-iteration artifacts plus a final summary.

Key point

  1. Evidence-driven axis budget: use a real workload roofline when ref.py exposes workload_model; otherwise label the result as a bottleneck-gap heuristic and do not claim near_peak.
  2. Branch-and-Select: each iteration generates K candidate kernels (hyperparameter/implementation variants), benchmarks all, selects champion.
  3. Ablation attribution: after selecting champion, each method is individually ablated to determine its actual contribution.
  4. SASS verification: cuobjdump --dump-sass confirms claimed optimizations actually appear in generated code.
  5. Every iteration produces an adaptive ncu report on the champion kernel; only CPU compilation is parallel.
  6. No evidence, no generation: the requested iteration count is a maximum. Never write candidates until the hardware, correctness, timing, and NCU gates pass.

Inputs the skill expects from the user

Before starting, confirm you have:

  1. Baseline operator file, e.g. ./gemm.cu or ./gemm_triton.py
  2. Reference file, e.g. ./ref.py (required — correctness validation depends on it)
  3. Dimensions — kernel-signature scalars like --M=4096 --N=4096 --K=4096
  4. Iteration count N (default 3)
  5. ncu_num — how many top metrics to extract per axis (default 5)
  6. branches — how many hyperparameter variants per iteration (default 4)

benchmark.py is bundled at scripts/benchmark.py; all scripts default to it automatically.

If any of these are missing, ask the user once — briefly — then proceed.

The loop at a glance

0. hardware_gate      → env.json (GPU runtime + tools, at most 3 attempts)
1. init run folder    → run_YYYYMMDD_HHMMSS/
2. copy baseline      → baseline/ + bench once to seed `best`
3. for i in 1..N:
     a. profile best_kernel with adaptive ncu (full/light by duration)  → iterv{i}/best_input.ncu-rep
     b. extract top compute/mem/latency            → ncu_top.json
     c. roofline.py: compute real roofline or bottleneck gaps → roofline.json + axis_budget
     d. Claude picks methods (b_axis per axis, cap=2) → analysis.md (CoT)
     e. Claude writes K branch kernels (same methods, diff hyperparams)
     f. contract gate, then branch_explore.py: parallel CPU compile, serial GPU bench all K → select champion
     g. if all branches FAIL any strict gate: stop the run
     h. ncu profile champion (adaptive, same bundle) → iterv{i}/kernel.ncu-rep
     i. ablate.py: single-method rollback bench    → attribution.json
     j. sass_check.py: verify SASS signatures      → sass_check.json
     k. update state with attribution + SASS results
4. emit summary.md

Steps (a), (b), (c), (f), (h), (i), (j) are scripted and reproducible; GPU timing itself is noisy. Steps (d) and (e) are where Claude thinks — follow the reasoning rules in references/optimization_catalog.md and references/ncu_metrics_guide.md.


Step 0 — Strict hardware gate

Run the strict gate before creating a run or generating any kernel:

bash
python <skill>/scripts/hardware_gate.py --out ./env.json --backend cuda --attempts 3

It searches PATH, CUDA environment roots, /usr/local/cuda*, /opt/cuda*, and common Nsight Compute locations. It must create a CUDA context and resolve backend dependencies, ncu, and (for CUDA/CUTLASS) nvcc plus cuobjdump. The bundled benchmark requires PyTorch. Retry discovery at most three times. If generation_allowed is false, stop immediately; benchmark-only degradation is forbidden. --diagnostic reports the environment but never authorizes generation.

Step 0b — Preflight the baseline + ref contract

bash
python <skill>/scripts/preflight.py \
  --baseline ./gemm.cu \
  --ref      ./ref.py \
  --dims     '{"M":4096,"N":4096,"K":4096}'

Validates baseline and reference contracts. On failure, surface errors directly to the user. orchestrate.py setup runs this automatically.

Step 1 — Initialize the run folder

bash
python <skill>/scripts/state.py init \
  --baseline ./gemm.cu \
  --ref ./ref.py \
  --iterations 3 \
  --ncu-num 5 \
  --branches 4 \
  --dims '{"M":4096,"N":4096,"K":4096}' \
  --env ./env.json

Creates ./run_YYYYMMDD_HHMMSS/ next to the baseline file and writes state.json. Only baseline/ is created initially; iteration directories are created lazily after their pre-generation gates pass:

jsonc
{
  "run_dir": "...",
  "baseline_file": "...",
  "ref_file": "...",
  "best_file": "<baseline>",
  "best_metric_ms": null,
  "best_ncu_rep": null,
  "env": {...},
  "iterations_total": 3,
  "ncu_num": 5,
  "branches": 4,
  "selected_methods": [],
  "effective_methods": [],
  "ineffective_methods": [],
  "implementation_failed_methods": [],
  "dims": {...},
  "history": [],
  "roofline_history": [],
  "frontier": []
}

Schema v5 also records run_status, generation_allowed, stop_reason, stop_stage, hardware_attempts, verified_iterations, and last_verified_iter. A stopped run always includes stop.json; an unverified iteration never increments verified_iterations.

Step 2 — Seed best with a baseline benchmark

bash
python <skill>/scripts/run_iteration.py seed-baseline \
  --state ./run_*/state.json

The baseline must pass compilation, contract, multi-seed correctness, stable positive kernel/reference timing, and NCU collection/import before iteration 1 may open. Store baseline NCU evidence under baseline/. On failure write stop.json, render a stopped summary, and do not create iterv1.

Step 3 — Iteration loop (up to N iterations)

Treat N as a maximum. Before writing any branch, run open-iter; it rechecks hardware and profiles the current best into staging. Only a successful, non-degraded NCU report with at least one parsed metric may publish iterv{i} and its branch directories.

3a. Profile the current best with ncu (adaptive report)
bash
python <skill>/scripts/profile_ncu.py \
  --state ./run_*/state.json \
  --iter $i \
  --which best_input

The default policy uses full below 10 ms and the light/basic metric bundle at 10 ms and above. A full replay that times out or exits with code 11 is retried with the light bundle. ncu_top.json records the selected set, duration, reason, and all attempts.

3b. Compute roofline gaps and axis budgets
bash
python <skill>/scripts/roofline.py \
  --state ./run_*/state.json \
  --iter $i

Reads ncu_top.json + env.json, computes:

  • Δ_c = compute utilization gap
  • Δ_m = bandwidth utilization gap
  • Δ_l = max stall percentage

Writes iterv{i}/roofline.json:

jsonc
{
  "delta_compute": 0.85,
  "delta_memory": 0.60,
  "delta_latency": 0.55,
  "bound": "compute",
  "near_peak": false,
  "axis_budget": {"compute": 1, "memory": 1, "latency": 1}
}

Budget allocation rule: proportional to known Δ values, rounded, cap per axis = 2, total = 3. near_peak/early stop is allowed only when all three gaps are known, a workload model is present, and all Δ < 0.15.

3c. Select methods (Claude reasons here)

Read (in this order):

  1. references/method_registry.json — canonical IDs, priorities, capabilities and relations
  2. references/metric_registry.json — versioned NCU aliases, units and missing-value semantics
  3. references/optimization_catalog.md — only the relevant method/backend cards and conditional archetype packs
  4. iterv{i}/roofline.json — axis budgets and bound classification
  5. iterv{i}/ncu_top.json — current bottleneck metrics
  6. state.json — method history and numerics_mode
  7. The current best_file source code
  8. references/ncu_metrics_guide.md — only metrics relevant to the observed bottleneck

Selection rule — BUDGET-AWARE PRIORITY SCAN:

For each axis with b_axis > 0, scan the catalog from P1 downward. For each priority level, check:

  1. Is method.id already in selected_methods? → skip (already tried)
  2. Does the detected sm_arch meet the method's arch requirement? → skip if not
  3. Does the method's skip condition apply? → skip (record reason in analysis.md)
  4. Does the method's trigger condition match the ncu evidence? → skip if no bottleneck here

Select methods in priority order until b_axis eligible methods are found. If fewer candidates pass all gates, leave the budget under-filled and record the reason; never add an untriggered or unsafe filler.

Produce up to B methods (sum of axis budgets, typically 3). Record concise evidence and decision rationale; do not emit private Chain-of-Thought.

Hard constraints:

  1. Apply typed registry relations: conflicts reject a pair, complements are allowed, and same-budget groups consume one slot.
  2. Methods in ineffective_methods are blocked unless ncu bottleneck has fundamentally changed.
  3. Methods in implementation_failed_methods require explicit acknowledgment of the prior failure.
  4. All methods must pass backend, capability, toolchain and numerical-semantic gates.
  5. Per-axis cap is 2 — no axis can receive more than 2 methods.

Save to iterv{i}/analysis.md using the template in templates/iteration_report.md.

Show full SKILL.md (667 more words)Show less
3d. Generate K branch kernels (Claude writes code)

All K branches share the same method combination from step 3c. They differ in hyperparameters and implementation details:

  • Tile sizes (BLOCK_M, BLOCK_N, BLOCK_K)
  • Pipeline stage count (num_stages)
  • Warp count (num_warps)
  • Implementation variant within a method (e.g., swizzle mode, MMA atom selection)

Write K kernels under iterv{i}/branches/b{1..K}/kernel.<ext>.

3e. Branch explore: compile + benchmark all K
bash
python <skill>/scripts/branch_explore.py \
  --state ./run_*/state.json \
  --iter $i

For the bundled benchmark, compiles CUDA/CUTLASS branches in a bounded CPU pool, waits for all builds, then benchmarks them serially on the ranking GPU in b1..bK order. Triton and unsupported custom benchmarks use the original serial path. Selects champion by (average_ms, branch_index); non-champions are saved to state.frontier.

3f. Stop on validation failure

If all branches fail compilation, contract, correctness, or stable timing, stop the run and do not generate later iterations. Environment discovery alone retries up to three times; kernel validation failures are terminal for this run.

Before GPU work, CUDA branches pass contract_check.py. Its independent compile_pass, contract_pass, correctness_pass, race_safe, and timing_valid states are persisted; a failed contract cannot be timed or selected. Branch results retain requested/realized shapes, OOM attempts, robust CV, and kernel-only/end-to-end timing fields.

3g. Profile champion with ncu (same adaptive bundle)
bash
python <skill>/scripts/profile_ncu.py \
  --state ./run_*/state.json \
  --iter $i \
  --which kernel

Writes iterv{i}/kernel.ncu-rep. The selected metric bundle is kept identical to the first profile in the run so baseline/champion deltas remain comparable. Failure, timeout after full-to-light fallback, degraded output, empty report, CSV import failure, or zero parsed metrics stops the run before promotion.

3h. Ablation attribution
bash
python <skill>/scripts/ablate.py \
  --state ./run_*/state.json \
  --iter $i

For each pre-generated ablation kernel, compile CUDA/CUTLASS versions in the same bounded pool, then benchmark them serially. Missing or failed ablations are inconclusive. Computes attribution:

attribution(m) = ms_without_m - ms_champion

Positive attribution = the method contributed positively. Near-zero or negative = the method was not helpful.

Writes iterv{i}/attribution.json.

3i. SASS verification
bash
python <skill>/scripts/sass_check.py \
  --state ./run_*/state.json \
  --iter $i

Runs the declared verifier on the compiled champion. SASS status is pass|fail|inconclusive|not_applicable|tool_error; empty patterns, Triton and unavailable artifacts are not success. Writes iterv{i}/sass_check.json.

3j. Update global state
bash
python <skill>/scripts/state.py update \
  --state ./run_*/state.json \
  --iter $i \
  --kernel iterv{i}/kernel.<ext> \
  --bench iterv{i}/bench.json \
  --methods-json iterv{i}/methods.json \
  --attribution iterv{i}/attribution.json \
  --sass-check iterv{i}/sass_check.json

Rules:

  • selected_methods += all methods (always)
  • Method enters effective_methods only if: attribution > noise_threshold AND SASS verified
  • Method enters implementation_failed_methods only on a strong method-specific verification failure
  • Method enters ineffective_methods if attribution ≤ noise_threshold and verification is conclusive
  • Method enters unverified_methods if ablation or verification is inconclusive/unavailable
  • If new_ms < best_ms by more than noise_threshold → best_file updated
  • Append record to state.history and state.roofline_history

Step 4 — Final summary

bash
python <skill>/scripts/summarize.py \
  --state ./run_*/state.json \
  --out ./run_*/summary.md

Compilation and numerical policy

--compile-jobs auto caps workers at half the CPUs, four total workers, and available-memory estimates (2 GiB/job for CUDA, 4 GiB/job for CUTLASS), retaining 2 GiB. Resource/OOM failures retry once serially. The build manifest covers the effective source, compiler/toolchain, architecture, flags and include/link inputs; only a matching manifest may be reused. The run-local cache is never a timing cache.

numerics_mode=reference keeps the reference's effective atol/rtol contract. strict admits only bitwise-preserving or explicitly preconditioned methods; approximate requires explicit opt-in and still runs all correctness checks. Unknown NCU metrics and SASS/tool errors are inconclusive, never zero or pass.


Reasoning references

  • references/optimization_catalog.md — Catalog of optimization methods by axis, with algorithmic methods section.
  • references/ncu_metrics_guide.md — How to read ncu output and map bottleneck signatures.
  • references/sass_signatures.json — Expected SASS instruction patterns per method.

Failure modes to watch for

  • Benchmark crashes → check bench.json "error" field.
  • ncu reports no parseable metrics → stop; permissions or launch selection must be fixed first.
  • can_read_counters: false in env.json → retry discovery up to three times, then stop.
  • Triton + @triton.autotune → hard-code config before profiling.
  • Champion chosen but all methods have near-zero attribution → the speedup came from hyperparameter change, not methods. Record in analysis.md.
  • SASS signature missing but kernel is faster → nvcc took a different path. Mark the method unverified unless a strong, method-specific failure is proven; keep the kernel if it's faster.
  • Branch explore: all K branches fail validation → Claude must rewrite with different approach.
  • Early stop triggered → all Δ < 0.15, kernel is near roofline. Report to user.

Strict CLI exit codes are 0 for verified success, 2 for kernel validation/no valid branch, 3 for hardware/tool/NCU failure, and 4 for baseline/reference configuration or static-contract failure. Every failure after state creation must write stop.json, set run_status=stopped, render summary.md, and leave all later iterations uncreated.


Output contract

<baseline-dir>/run_YYYYMMDD_HHMMSS/
├── env.json
├── state.json
├── baseline/
│   ├── <baseline>           (copied)
│   └── bench.json
├── iterv1/
│   ├── kernel.<ext>          (champion)
│   ├── analysis.md           (roofline + methods + CoT)
│   ├── methods.json
│   ├── roofline.json
│   ├── best_input.ncu-rep    (profile of best going INTO this iter)
│   ├── ncu_top.json
│   ├── kernel.ncu-rep        (profile of champion — ALWAYS present)
│   ├── attribution.json
│   ├── sass_check.json
│   ├── bench.json
│   └── branches/
│       ├── b1/ ... b4/       (all branch candidates)
├── iterv2/...
├── iterv3/...
└── summary.md

© KernelFlow-ops, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 52 other files (scripts, references) in skills/cuda-kernel-optimizer of KernelFlow-ops/cuda-optimized-skill.

  • SKILL.md
  • examples/walkthrough.md
  • references/method_registry.json
  • references/metric_registry.json
  • references/ncu_metrics_guide.md
  • references/optimization_catalog.md
  • references/sass_signatures.json
  • scripts/__pycache__/ablate.cpython-313.pyc
  • scripts/__pycache__/benchmark.cpython-310.pyc
  • scripts/__pycache__/benchmark.cpython-313.pyc
  • scripts/__pycache__/branch_explore.cpython-313.pyc
  • scripts/__pycache__/build.cpython-313.pyc
  • scripts/__pycache__/check_env.cpython-313.pyc
  • scripts/__pycache__/contract_check.cpython-313.pyc
  • scripts/__pycache__/gpu_lock.cpython-313.pyc
  • scripts/__pycache__/lint_registry.cpython-313.pyc
  • scripts/__pycache__/orchestrate.cpython-313.pyc
  • … and 36 more

Open the folder on GitHubat commit 114a6cb

Compare with similar skills

Cuda Kernel Optimizer next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Cuda Kernel Optimizer compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Cuda Kernel Optimizer this skillKernelFlow-ops/cuda-optimized-skill212—~4.3kAutomated safety check: PassMIT
MUSA GPU Training Optimizeropen-infra-skills/infra-skills141—~1.7kAutomated safety check: PassApache-2.0
DGX Spark Training Gotchaswshobson/agents40k1 repos~2kAutomated safety check: PassMIT
Ncu Report Skillmit-han-lab/ncu-report-skill244—~2kAutomated safety check: PassMIT
Megatron-LM on SLURMNVIDIA/Megatron-LM18k—~1.8kAutomated safety check: PassApache-2.0
Mamba State-Space ModelsOrchestra-Research/AI-Research-SKILLs13k3 repos~1.8kAutomated safety check: PassMIT

Similar skills

  • MUSA GPU Training Optimizer

    open-infra-skills/infra-skills

    Profiles, benchmarks and tunes AI training workloads on Moore Threads MUSA GPUs with a measurement-first process that keeps model behavior unchanged.

    141 GitHub stars~1.7k tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check passed
  • Preflight checks and diagnosis for ten known failure modes of ML training on NVIDIA DGX Spark's GB10, spanning launch errors, memory, thermals, bandwidth and precision.

    40k GitHub starsUsed in 1 repo~2k tokens
    AI & LLM EngineeringAuto-check passed
  • Ncu Report Skill

    mit-han-lab/ncu-report-skill

    Profile CUDA kernels with Nsight Compute on B200 / sm100. An agent skill from mit-han-lab/ncu-report-skill.

    244 GitHub stars~2k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Megatron-LM on SLURM

    NVIDIA/Megatron-LM

    Official

    Shows how to launch distributed Megatron-LM training on a SLURM cluster: sbatch skeleton, torch.distributed.run setup, CUDA_DEVICE_MAX_CONNECTIONS rules and failure diagnosis.

    18k GitHub stars~1.8k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Mamba State-Space Models

    Orchestra-Research/AI-Research-SKILLs

    Guide to using Mamba selective state-space models for linear-time sequence modeling, from the Mamba block and pretrained checkpoints to Mamba-2 and speed comparisons.

    13k GitHub starsUsed in 3 repos~1.8k tokens
    AI & LLM EngineeringAuto-check passed
  • Lambda Labs GPU Cloud

    Orchestra-Research/AI-Research-SKILLs

    Guide to renting GPUs on Lambda Labs for ML training and inference: on-demand instances, 1-Click Clusters, SSH access, persistent filesystems and alternatives.

    13k GitHub starsUsed in 5 repos~3k tokens
    AI & LLM EngineeringAuto-check: warnings

Works with

Questions about Cuda Kernel Optimizer

What does Cuda Kernel Optimizer do?

Iteratively optimize a CUDA/CUTLASS/Triton kernel only when strict on-device compilation, correctness, timing, and NCU evidence gates pass. Cuda Kernel Optimizer is an agent skill from KernelFlow-ops/cuda-optimized-skill. Iteratively optimize a CUDA/CUTLASS/Triton kernel only when strict on-device compilation, correctness, timing, and NCU evidence gates pass.

When should I use Cuda Kernel Optimizer?

Cuda Kernel Optimizer fits situations like: tasks that involve GPU and accelerator computing.

How do I install Cuda Kernel Optimizer in Claude Code?

Run `npx skills add KernelFlow-ops/cuda-optimized-skill --skill cuda-kernel-optimizer -a claude-code`. Or copy the skill folder (skills/cuda-kernel-optimizer in KernelFlow-ops/cuda-optimized-skill) into .claude/skills/cuda-kernel-optimizer in your project. Claude Code loads it when a task matches its description.

How do I install Cuda Kernel Optimizer in Codex?

Run `npx skills add KernelFlow-ops/cuda-optimized-skill --skill cuda-kernel-optimizer -a codex`. Or copy the skill folder (skills/cuda-kernel-optimizer in KernelFlow-ops/cuda-optimized-skill) into .agents/skills/cuda-kernel-optimizer in your project. Codex loads it when a task matches its description.

Can I use Cuda Kernel Optimizer in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add KernelFlow-ops/cuda-optimized-skill --skill cuda-kernel-optimizer -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/cuda-kernel-optimizer, .gemini/skills/cuda-kernel-optimizer, .github/skills/cuda-kernel-optimizer and .opencode/skills/cuda-kernel-optimizer in your project.

What does Cuda Kernel Optimizer need to run?

Going by SKILL.md and its folder, Cuda Kernel Optimizer needs the command-line tools its instructions call (python). Our summary lists: Python 3.

Does Cuda Kernel Optimizer access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Cuda Kernel Optimizer safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Cuda Kernel Optimizer use?

Cuda Kernel Optimizer is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Cuda Kernel Optimizer use?

About 4.3k tokens (SKILL.md is roughly 17k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 33k tokens, read only when the agent opens those files.

What are the alternatives to Cuda Kernel Optimizer?

Skills that share tags, products or a category with Cuda Kernel Optimizer: MUSA GPU Training Optimizer (open-infra-skills/infra-skills, 141 stars), DGX Spark Training Gotchas (wshobson/agents, 40k stars), Ncu Report Skill (mit-han-lab/ncu-report-skill, 244 stars) and Megatron-LM on SLURM (NVIDIA/Megatron-LM, 18k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Cuda Kernel Optimizer?

KernelFlow-ops (a GitHub user) maintains it in KernelFlow-ops/cuda-optimized-skill, which has 212 GitHub stars. The repository was last updated on September 5, 2026.

Source: KernelFlow-ops/cuda-optimized-skill on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.