MUSA GPU Training Optimizer
open-infra-skills/infra-skills
Profiles, benchmarks and tunes AI training workloads on Moore Threads MUSA GPUs with a measurement-first process that keeps model behavior unchanged.
Iteratively optimize a CUDA/CUTLASS/Triton kernel only when strict on-device compilation, correctness, timing, and NCU evidence gates pass.
$ npx skills add KernelFlow-ops/cuda-optimized-skill --skill cuda-kernel-optimizer -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install KernelFlow-ops/cuda-optimized-skill cuda-kernel-optimizer --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/KernelFlow-ops/cuda-optimized-skill.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/cuda-kernel-optimizer .claude/skills/cuda-kernel-optimizer && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "cuda-kernel-optimizer" agent skill from https://github.com/KernelFlow-ops/cuda-optimized-skill/tree/main/skills/cuda-kernel-optimizer into .claude/skills/cuda-kernel-optimizer/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "cuda-kernel-optimizer", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/KernelFlow-ops/cuda-optimized-skill/tree/main/skills/cuda-kernel-optimizerType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add KernelFlow-ops/cuda-optimized-skill --skill cuda-kernel-optimizer -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install KernelFlow-ops/cuda-optimized-skill cuda-kernel-optimizer --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/KernelFlow-ops/cuda-optimized-skill.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/cuda-kernel-optimizer .agents/skills/cuda-kernel-optimizer && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "cuda-kernel-optimizer" agent skill from https://github.com/KernelFlow-ops/cuda-optimized-skill/tree/main/skills/cuda-kernel-optimizer into .agents/skills/cuda-kernel-optimizer/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "cuda-kernel-optimizer", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add KernelFlow-ops/cuda-optimized-skill --skill cuda-kernel-optimizer -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install KernelFlow-ops/cuda-optimized-skill cuda-kernel-optimizer --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/KernelFlow-ops/cuda-optimized-skill.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/cuda-kernel-optimizer .cursor/skills/cuda-kernel-optimizer && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "cuda-kernel-optimizer" agent skill from https://github.com/KernelFlow-ops/cuda-optimized-skill/tree/main/skills/cuda-kernel-optimizer into .cursor/skills/cuda-kernel-optimizer/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "cuda-kernel-optimizer", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/KernelFlow-ops/cuda-optimized-skill.git --path skills/cuda-kernel-optimizer--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add KernelFlow-ops/cuda-optimized-skill --skill cuda-kernel-optimizer -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install KernelFlow-ops/cuda-optimized-skill cuda-kernel-optimizer --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/KernelFlow-ops/cuda-optimized-skill.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/cuda-kernel-optimizer .gemini/skills/cuda-kernel-optimizer && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "cuda-kernel-optimizer" agent skill from https://github.com/KernelFlow-ops/cuda-optimized-skill/tree/main/skills/cuda-kernel-optimizer into .gemini/skills/cuda-kernel-optimizer/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "cuda-kernel-optimizer", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install KernelFlow-ops/cuda-optimized-skill cuda-kernel-optimizerInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add KernelFlow-ops/cuda-optimized-skill --skill cuda-kernel-optimizer -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/KernelFlow-ops/cuda-optimized-skill.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/cuda-kernel-optimizer .github/skills/cuda-kernel-optimizer && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "cuda-kernel-optimizer" agent skill from https://github.com/KernelFlow-ops/cuda-optimized-skill/tree/main/skills/cuda-kernel-optimizer into .github/skills/cuda-kernel-optimizer/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "cuda-kernel-optimizer", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add KernelFlow-ops/cuda-optimized-skill --skill cuda-kernel-optimizer -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install KernelFlow-ops/cuda-optimized-skill cuda-kernel-optimizer --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/KernelFlow-ops/cuda-optimized-skill.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/cuda-kernel-optimizer .opencode/skills/cuda-kernel-optimizer && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "cuda-kernel-optimizer" agent skill from https://github.com/KernelFlow-ops/cuda-optimized-skill/tree/main/skills/cuda-kernel-optimizer into .opencode/skills/cuda-kernel-optimizer/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "cuda-kernel-optimizer", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
cuda-kernel-optimizerIteratively optimize a CUDA/CUTLASS/Triton kernel only when strict on-device compilation, correctness, timing, and NCU evidence gates pass.
Cuda Kernel Optimizer is an agent skill from KernelFlow-ops/cuda-optimized-skill. Iteratively optimize a CUDA/CUTLASS/Triton kernel only when strict on-device compilation, correctness, timing, and NCU evidence gates pass.
Its SKILL.md is about 4.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 56 other files, including scripts and reference files (for example `examples/walkthrough.md`, `references/method_registry.json` and `references/metric_registry.json`).
It sits in AI & LLM Engineering, covering GPU and accelerator computing. It works with CUDA. The repository describes itself as: A CUDA kernel optimization toolkit for validation, benchmarking, Nsight Compute profiling, bottleneck analysis, and iterative tuning. It helps improve custom GPU operators with… The licence is MIT.
5 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 114a6cb. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 10 files in scripts/, which the agent can run.
Shell commands in SKILL.md call:
pythonFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Cuda Kernel Optimizer loads about 4.3k tokens when it runs, and up to ~37k if it reads all its reference files. Until then it costs about 40 tokens; SKILL.md has 1,637 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from KernelFlow-ops/cuda-optimized-skill at commit 114a6cb, republished under its MIT licence (© KernelFlow-ops). 1,637 words, ~4,270 tokens.
.claude/skills/cuda-kernel-optimizer/SKILL.md (or your agent's skills folder). This skill also uses 52 other files; get the full folder from GitHub.Given:
.cu for CUDA / CUTLASS, or .py for Triton),reference(**kwargs) — same contract as benchmark.py --ref),atol/rtol and workload_model(**dims) -> {flops, bytes_min} in the reference,--M=4096 --N=4096 --K=4096),N (default 3), ncu_num (default 5), and branches (default 4),--compile-jobs auto|N (default auto) and --numerics-mode reference|strict|approximate (default reference),the skill runs a roofline-guided, branch-and-select iterative optimization loop and produces a timestamped directory of per-iteration artifacts plus a final summary.
ref.py exposes workload_model; otherwise label the result as a bottleneck-gap heuristic and do not claim near_peak.cuobjdump --dump-sass confirms claimed optimizations actually appear in generated code.Before starting, confirm you have:
./gemm.cu or ./gemm_triton.py./ref.py (required — correctness validation depends on it)--M=4096 --N=4096 --K=4096N (default 3)ncu_num — how many top metrics to extract per axis (default 5)branches — how many hyperparameter variants per iteration (default 4)
benchmark.pyis bundled atscripts/benchmark.py; all scripts default to it automatically.
If any of these are missing, ask the user once — briefly — then proceed.
0. hardware_gate → env.json (GPU runtime + tools, at most 3 attempts)
1. init run folder → run_YYYYMMDD_HHMMSS/
2. copy baseline → baseline/ + bench once to seed `best`
3. for i in 1..N:
a. profile best_kernel with adaptive ncu (full/light by duration) → iterv{i}/best_input.ncu-rep
b. extract top compute/mem/latency → ncu_top.json
c. roofline.py: compute real roofline or bottleneck gaps → roofline.json + axis_budget
d. Claude picks methods (b_axis per axis, cap=2) → analysis.md (CoT)
e. Claude writes K branch kernels (same methods, diff hyperparams)
f. contract gate, then branch_explore.py: parallel CPU compile, serial GPU bench all K → select champion
g. if all branches FAIL any strict gate: stop the run
h. ncu profile champion (adaptive, same bundle) → iterv{i}/kernel.ncu-rep
i. ablate.py: single-method rollback bench → attribution.json
j. sass_check.py: verify SASS signatures → sass_check.json
k. update state with attribution + SASS results
4. emit summary.mdSteps (a), (b), (c), (f), (h), (i), (j) are scripted and reproducible; GPU timing itself is noisy.
Steps (d) and (e) are where Claude thinks — follow the reasoning rules in references/optimization_catalog.md and references/ncu_metrics_guide.md.
Run the strict gate before creating a run or generating any kernel:
python <skill>/scripts/hardware_gate.py --out ./env.json --backend cuda --attempts 3It searches PATH, CUDA environment roots, /usr/local/cuda*, /opt/cuda*, and common Nsight Compute locations. It must create a CUDA context and resolve backend dependencies, ncu, and (for CUDA/CUTLASS) nvcc plus cuobjdump. The bundled benchmark requires PyTorch. Retry discovery at most three times. If generation_allowed is false, stop immediately; benchmark-only degradation is forbidden. --diagnostic reports the environment but never authorizes generation.
python <skill>/scripts/preflight.py \
--baseline ./gemm.cu \
--ref ./ref.py \
--dims '{"M":4096,"N":4096,"K":4096}'Validates baseline and reference contracts. On failure, surface errors directly to the user. orchestrate.py setup runs this automatically.
python <skill>/scripts/state.py init \
--baseline ./gemm.cu \
--ref ./ref.py \
--iterations 3 \
--ncu-num 5 \
--branches 4 \
--dims '{"M":4096,"N":4096,"K":4096}' \
--env ./env.jsonCreates ./run_YYYYMMDD_HHMMSS/ next to the baseline file and writes state.json. Only baseline/ is created initially; iteration directories are created lazily after their pre-generation gates pass:
{
"run_dir": "...",
"baseline_file": "...",
"ref_file": "...",
"best_file": "<baseline>",
"best_metric_ms": null,
"best_ncu_rep": null,
"env": {...},
"iterations_total": 3,
"ncu_num": 5,
"branches": 4,
"selected_methods": [],
"effective_methods": [],
"ineffective_methods": [],
"implementation_failed_methods": [],
"dims": {...},
"history": [],
"roofline_history": [],
"frontier": []
}Schema v5 also records run_status, generation_allowed, stop_reason,
stop_stage, hardware_attempts, verified_iterations, and
last_verified_iter. A stopped run always includes stop.json; an unverified
iteration never increments verified_iterations.
best with a baseline benchmarkpython <skill>/scripts/run_iteration.py seed-baseline \
--state ./run_*/state.jsonThe baseline must pass compilation, contract, multi-seed correctness, stable positive kernel/reference timing, and NCU collection/import before iteration 1 may open. Store baseline NCU evidence under baseline/. On failure write stop.json, render a stopped summary, and do not create iterv1.
Treat N as a maximum. Before writing any branch, run open-iter; it rechecks hardware and profiles the current best into staging. Only a successful, non-degraded NCU report with at least one parsed metric may publish iterv{i} and its branch directories.
best with ncu (adaptive report)python <skill>/scripts/profile_ncu.py \
--state ./run_*/state.json \
--iter $i \
--which best_inputThe default policy uses full below 10 ms and the light/basic metric bundle
at 10 ms and above. A full replay that times out or exits with code 11 is
retried with the light bundle. ncu_top.json records the selected set,
duration, reason, and all attempts.
python <skill>/scripts/roofline.py \
--state ./run_*/state.json \
--iter $iReads ncu_top.json + env.json, computes:
Δ_c = compute utilization gapΔ_m = bandwidth utilization gapΔ_l = max stall percentageWrites iterv{i}/roofline.json:
{
"delta_compute": 0.85,
"delta_memory": 0.60,
"delta_latency": 0.55,
"bound": "compute",
"near_peak": false,
"axis_budget": {"compute": 1, "memory": 1, "latency": 1}
}Budget allocation rule: proportional to known Δ values, rounded, cap per axis = 2, total = 3. near_peak/early stop is allowed only when all three gaps are known, a workload model is present, and all Δ < 0.15.
Read (in this order):
references/method_registry.json — canonical IDs, priorities, capabilities and relationsreferences/metric_registry.json — versioned NCU aliases, units and missing-value semanticsreferences/optimization_catalog.md — only the relevant method/backend cards and conditional archetype packsiterv{i}/roofline.json — axis budgets and bound classificationiterv{i}/ncu_top.json — current bottleneck metricsstate.json — method history and numerics_modebest_file source codereferences/ncu_metrics_guide.md — only metrics relevant to the observed bottleneckSelection rule — BUDGET-AWARE PRIORITY SCAN:
For each axis with b_axis > 0, scan the catalog from P1 downward. For each priority level, check:
method.id already in selected_methods? → skip (already tried)sm_arch meet the method's arch requirement? → skip if notSelect methods in priority order until b_axis eligible methods are found. If fewer candidates pass all gates, leave the budget under-filled and record the reason; never add an untriggered or unsafe filler.
Produce up to B methods (sum of axis budgets, typically 3). Record concise evidence and decision rationale; do not emit private Chain-of-Thought.
Hard constraints:
ineffective_methods are blocked unless ncu bottleneck has fundamentally changed.implementation_failed_methods require explicit acknowledgment of the prior failure.Save to iterv{i}/analysis.md using the template in templates/iteration_report.md.
All K branches share the same method combination from step 3c. They differ in hyperparameters and implementation details:
Write K kernels under iterv{i}/branches/b{1..K}/kernel.<ext>.
python <skill>/scripts/branch_explore.py \
--state ./run_*/state.json \
--iter $iFor the bundled benchmark, compiles CUDA/CUTLASS branches in a bounded CPU pool, waits for all builds, then benchmarks them serially on the ranking GPU in b1..bK order. Triton and unsupported custom benchmarks use the original serial path. Selects champion by (average_ms, branch_index); non-champions are saved to state.frontier.
If all branches fail compilation, contract, correctness, or stable timing, stop the run and do not generate later iterations. Environment discovery alone retries up to three times; kernel validation failures are terminal for this run.
Before GPU work, CUDA branches pass contract_check.py. Its independent
compile_pass, contract_pass, correctness_pass, race_safe, and
timing_valid states are persisted; a failed contract cannot be timed or
selected. Branch results retain requested/realized shapes, OOM attempts,
robust CV, and kernel-only/end-to-end timing fields.
python <skill>/scripts/profile_ncu.py \
--state ./run_*/state.json \
--iter $i \
--which kernelWrites iterv{i}/kernel.ncu-rep. The selected metric bundle is kept identical
to the first profile in the run so baseline/champion deltas remain comparable.
Failure, timeout after full-to-light fallback, degraded output, empty report,
CSV import failure, or zero parsed metrics stops the run before promotion.
python <skill>/scripts/ablate.py \
--state ./run_*/state.json \
--iter $iFor each pre-generated ablation kernel, compile CUDA/CUTLASS versions in the same bounded pool, then benchmark them serially. Missing or failed ablations are inconclusive. Computes attribution:
attribution(m) = ms_without_m - ms_championPositive attribution = the method contributed positively. Near-zero or negative = the method was not helpful.
Writes iterv{i}/attribution.json.
python <skill>/scripts/sass_check.py \
--state ./run_*/state.json \
--iter $iRuns the declared verifier on the compiled champion. SASS status is pass|fail|inconclusive|not_applicable|tool_error; empty patterns, Triton and unavailable artifacts are not success. Writes iterv{i}/sass_check.json.
python <skill>/scripts/state.py update \
--state ./run_*/state.json \
--iter $i \
--kernel iterv{i}/kernel.<ext> \
--bench iterv{i}/bench.json \
--methods-json iterv{i}/methods.json \
--attribution iterv{i}/attribution.json \
--sass-check iterv{i}/sass_check.jsonRules:
selected_methods += all methods (always)effective_methods only if: attribution > noise_threshold AND SASS verifiedimplementation_failed_methods only on a strong method-specific verification failureineffective_methods if attribution ≤ noise_threshold and verification is conclusiveunverified_methods if ablation or verification is inconclusive/unavailablenew_ms < best_ms by more than noise_threshold → best_file updatedstate.history and state.roofline_historypython <skill>/scripts/summarize.py \
--state ./run_*/state.json \
--out ./run_*/summary.md--compile-jobs auto caps workers at half the CPUs, four total workers, and
available-memory estimates (2 GiB/job for CUDA, 4 GiB/job for CUTLASS), retaining
2 GiB. Resource/OOM failures retry once serially. The build manifest covers the
effective source, compiler/toolchain, architecture, flags and include/link
inputs; only a matching manifest may be reused. The run-local cache is never a
timing cache.
numerics_mode=reference keeps the reference's effective atol/rtol contract.
strict admits only bitwise-preserving or explicitly preconditioned methods;
approximate requires explicit opt-in and still runs all correctness checks.
Unknown NCU metrics and SASS/tool errors are inconclusive, never zero or pass.
references/optimization_catalog.md — Catalog of optimization methods by axis, with algorithmic methods section.references/ncu_metrics_guide.md — How to read ncu output and map bottleneck signatures.references/sass_signatures.json — Expected SASS instruction patterns per method.bench.json "error" field.can_read_counters: false in env.json → retry discovery up to three times, then stop.@triton.autotune → hard-code config before profiling.unverified unless a strong, method-specific failure is proven; keep the kernel if it's faster.Strict CLI exit codes are 0 for verified success, 2 for kernel validation/no
valid branch, 3 for hardware/tool/NCU failure, and 4 for baseline/reference
configuration or static-contract failure. Every failure after state creation
must write stop.json, set run_status=stopped, render summary.md, and leave
all later iterations uncreated.
<baseline-dir>/run_YYYYMMDD_HHMMSS/
├── env.json
├── state.json
├── baseline/
│ ├── <baseline> (copied)
│ └── bench.json
├── iterv1/
│ ├── kernel.<ext> (champion)
│ ├── analysis.md (roofline + methods + CoT)
│ ├── methods.json
│ ├── roofline.json
│ ├── best_input.ncu-rep (profile of best going INTO this iter)
│ ├── ncu_top.json
│ ├── kernel.ncu-rep (profile of champion — ALWAYS present)
│ ├── attribution.json
│ ├── sass_check.json
│ ├── bench.json
│ └── branches/
│ ├── b1/ ... b4/ (all branch candidates)
├── iterv2/...
├── iterv3/...
└── summary.md© KernelFlow-ops, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 52 other files (scripts, references) in skills/cuda-kernel-optimizer of KernelFlow-ops/cuda-optimized-skill.
Open the folder on GitHubat commit 114a6cb
Cuda Kernel Optimizer next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Cuda Kernel Optimizer this skillKernelFlow-ops/cuda-optimized-skill | 212 | — | ~4.3k | Automated safety check: Pass | MIT | |
| MUSA GPU Training Optimizeropen-infra-skills/infra-skills | 141 | — | ~1.7k | Automated safety check: Pass | Apache-2.0 | |
| DGX Spark Training Gotchaswshobson/agents | 40k | 1 repos | ~2k | Automated safety check: Pass | MIT | |
| Ncu Report Skillmit-han-lab/ncu-report-skill | 244 | — | ~2k | Automated safety check: Pass | MIT | |
| Megatron-LM on SLURMNVIDIA/Megatron-LM | 18k | — | ~1.8k | Automated safety check: Pass | Apache-2.0 | |
| Mamba State-Space ModelsOrchestra-Research/AI-Research-SKILLs | 13k | 3 repos | ~1.8k | Automated safety check: Pass | MIT |
open-infra-skills/infra-skills
Profiles, benchmarks and tunes AI training workloads on Moore Threads MUSA GPUs with a measurement-first process that keeps model behavior unchanged.
wshobson/agents
Preflight checks and diagnosis for ten known failure modes of ML training on NVIDIA DGX Spark's GB10, spanning launch errors, memory, thermals, bandwidth and precision.
mit-han-lab/ncu-report-skill
Profile CUDA kernels with Nsight Compute on B200 / sm100. An agent skill from mit-han-lab/ncu-report-skill.
NVIDIA/Megatron-LM
Shows how to launch distributed Megatron-LM training on a SLURM cluster: sbatch skeleton, torch.distributed.run setup, CUDA_DEVICE_MAX_CONNECTIONS rules and failure diagnosis.
Orchestra-Research/AI-Research-SKILLs
Guide to using Mamba selective state-space models for linear-time sequence modeling, from the Mamba block and pretrained checkpoints to Mamba-2 and speed comparisons.
Orchestra-Research/AI-Research-SKILLs
Guide to renting GPUs on Lambda Labs for ML training and inference: on-demand instances, 1-Click Clusters, SSH access, persistent filesystems and alternatives.
Works with
Categories
Iteratively optimize a CUDA/CUTLASS/Triton kernel only when strict on-device compilation, correctness, timing, and NCU evidence gates pass. Cuda Kernel Optimizer is an agent skill from KernelFlow-ops/cuda-optimized-skill. Iteratively optimize a CUDA/CUTLASS/Triton kernel only when strict on-device compilation, correctness, timing, and NCU evidence gates pass.
Cuda Kernel Optimizer fits situations like: tasks that involve GPU and accelerator computing.
Run `npx skills add KernelFlow-ops/cuda-optimized-skill --skill cuda-kernel-optimizer -a claude-code`. Or copy the skill folder (skills/cuda-kernel-optimizer in KernelFlow-ops/cuda-optimized-skill) into .claude/skills/cuda-kernel-optimizer in your project. Claude Code loads it when a task matches its description.
Run `npx skills add KernelFlow-ops/cuda-optimized-skill --skill cuda-kernel-optimizer -a codex`. Or copy the skill folder (skills/cuda-kernel-optimizer in KernelFlow-ops/cuda-optimized-skill) into .agents/skills/cuda-kernel-optimizer in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add KernelFlow-ops/cuda-optimized-skill --skill cuda-kernel-optimizer -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/cuda-kernel-optimizer, .gemini/skills/cuda-kernel-optimizer, .github/skills/cuda-kernel-optimizer and .opencode/skills/cuda-kernel-optimizer in your project.
Going by SKILL.md and its folder, Cuda Kernel Optimizer needs the command-line tools its instructions call (python). Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Cuda Kernel Optimizer is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 4.3k tokens (SKILL.md is roughly 17k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 33k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Cuda Kernel Optimizer: MUSA GPU Training Optimizer (open-infra-skills/infra-skills, 141 stars), DGX Spark Training Gotchas (wshobson/agents, 40k stars), Ncu Report Skill (mit-han-lab/ncu-report-skill, 244 stars) and Megatron-LM on SLURM (NVIDIA/Megatron-LM, 18k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
KernelFlow-ops (a GitHub user) maintains it in KernelFlow-ops/cuda-optimized-skill, which has 212 GitHub stars. The repository was last updated on September 5, 2026.
Source: KernelFlow-ops/cuda-optimized-skill on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.