Docstring
pytorch/pytorch
Write docstrings for PyTorch functions and methods following PyTorch conventions.
Run a bounded TileLang implementation and optimization task from a PyTorch reference or existing kernel through baseline analysis, a written plan, candidate edits, lowering inspection, validation…
$ npx skills add sablin39/tilelang-cuda-skills --skill tilelang-optimization-loop -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install sablin39/tilelang-cuda-skills tilelang-optimization-loop --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/sablin39/tilelang-cuda-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/tilelang-optimization-loop .claude/skills/tilelang-optimization-loop && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "tilelang-optimization-loop" agent skill from https://github.com/sablin39/tilelang-cuda-skills/tree/main/skills/tilelang-optimization-loop into .claude/skills/tilelang-optimization-loop/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tilelang-optimization-loop", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/sablin39/tilelang-cuda-skills/tree/main/skills/tilelang-optimization-loopType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add sablin39/tilelang-cuda-skills --skill tilelang-optimization-loop -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install sablin39/tilelang-cuda-skills tilelang-optimization-loop --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/sablin39/tilelang-cuda-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/tilelang-optimization-loop .agents/skills/tilelang-optimization-loop && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "tilelang-optimization-loop" agent skill from https://github.com/sablin39/tilelang-cuda-skills/tree/main/skills/tilelang-optimization-loop into .agents/skills/tilelang-optimization-loop/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tilelang-optimization-loop", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add sablin39/tilelang-cuda-skills --skill tilelang-optimization-loop -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install sablin39/tilelang-cuda-skills tilelang-optimization-loop --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/sablin39/tilelang-cuda-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/tilelang-optimization-loop .cursor/skills/tilelang-optimization-loop && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "tilelang-optimization-loop" agent skill from https://github.com/sablin39/tilelang-cuda-skills/tree/main/skills/tilelang-optimization-loop into .cursor/skills/tilelang-optimization-loop/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tilelang-optimization-loop", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/sablin39/tilelang-cuda-skills.git --path skills/tilelang-optimization-loop--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add sablin39/tilelang-cuda-skills --skill tilelang-optimization-loop -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install sablin39/tilelang-cuda-skills tilelang-optimization-loop --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/sablin39/tilelang-cuda-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/tilelang-optimization-loop .gemini/skills/tilelang-optimization-loop && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "tilelang-optimization-loop" agent skill from https://github.com/sablin39/tilelang-cuda-skills/tree/main/skills/tilelang-optimization-loop into .gemini/skills/tilelang-optimization-loop/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tilelang-optimization-loop", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install sablin39/tilelang-cuda-skills tilelang-optimization-loopInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add sablin39/tilelang-cuda-skills --skill tilelang-optimization-loop -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/sablin39/tilelang-cuda-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/tilelang-optimization-loop .github/skills/tilelang-optimization-loop && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "tilelang-optimization-loop" agent skill from https://github.com/sablin39/tilelang-cuda-skills/tree/main/skills/tilelang-optimization-loop into .github/skills/tilelang-optimization-loop/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tilelang-optimization-loop", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add sablin39/tilelang-cuda-skills --skill tilelang-optimization-loop -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install sablin39/tilelang-cuda-skills tilelang-optimization-loop --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/sablin39/tilelang-cuda-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/tilelang-optimization-loop .opencode/skills/tilelang-optimization-loop && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "tilelang-optimization-loop" agent skill from https://github.com/sablin39/tilelang-cuda-skills/tree/main/skills/tilelang-optimization-loop into .opencode/skills/tilelang-optimization-loop/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tilelang-optimization-loop", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
tilelang-optimization-loopRun a bounded TileLang implementation and optimization task from a PyTorch reference or existing kernel through baseline analysis, a written plan, candidate edits, lowering inspection, validation…
Tilelang Optimization Loop is an agent skill from sablin39/tilelang-cuda-skills. Run a bounded TileLang implementation and optimization task from a PyTorch reference or existing kernel through baseline analysis, a written plan, candidate edits, lowering inspection, validation, measurement, and promotion. Use for iterative kernel work and reproducible handoffs; use a specialist directly for a single API, debugging, or profiling question.
Its SKILL.md is about 2.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 7 other files, including reference files (for example `references/execution-harness.md`, `references/experiment-record.md` and `references/external-tools.md`).
It sits in AI & LLM Engineering, covering Deep learning. It works with PyTorch. The repository describes itself as: Skills for writing tilelang and debugging with CUDA toolkits.
6 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 502a514. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md.
From the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
github.comFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Tilelang Optimization Loop loads about 2.7k tokens when it runs, and up to ~9.7k if it reads all its reference files. Until then it costs about 97 tokens; SKILL.md has 1,310 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
Without a licence we can't republish the file, so here is its outline and opening line. It has 1,310 words (~2,720 tokens).
“Follow a repeatable experiment cycle in the task workspace. Keep the skill suite and its upstream knowledge separate from generated kernels and results. Resume from evidence that still matches the source, inputs, and environment. This adapts the KDA workflow to…”
SKILL.md and 6 other files (references) in skills/tilelang-optimization-loop of sablin39/tilelang-cuda-skills.
Open the folder on GitHubat commit 502a514
Tilelang Optimization Loop next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Tilelang Optimization Loop this skillsablin39/tilelang-cuda-skills | 145 | — | ~2.7k | Automated safety check: Pass | None | |
| Docstringpytorch/pytorch | 104k | 2 repos | ~2.6k | Automated safety check: Pass | Custom licence | |
| ExecuTorch Cortex-M Backendpytorch/executorch | 5.1k | — | ~872 | Automated safety check: Pass | Custom licence | |
| Qualcomm QNN Backend Developmentpytorch/executorch | 5.1k | — | ~1.8k | Automated safety check: Pass | Custom licence | |
| Pt2 Bug Basherpytorch/pytorch | 104k | — | ~3.5k | Automated safety check: Pass | Custom licence | |
| Formattingbrendanhasz/probflow | 175 | — | ~381 | Automated safety check: Pass | MIT |
pytorch/pytorch
Write docstrings for PyTorch functions and methods following PyTorch conventions.
pytorch/executorch
Developer guide for the Cortex-M (CMSIS-NN) backend in ExecuTorch: quantization pipeline, pass manager, tests and adding new ops.
pytorch/executorch
Helps build, test and extend the Qualcomm AI Engine Direct (QNN) backend in ExecuTorch, with routes for new ops, model export, Buck-vs-CMake parity fixes and per-layer accuracy debugging.
pytorch/pytorch
Debug PyTorch 2 compiler stack failures including Dynamo graph breaks, Inductor codegen errors, AOTAutograd crashes, and accuracy mismatches.
brendanhasz/probflow
Ensure consistent code formatting using the uv package manager and pre-commit.
intel/torch-xpu-ops
Check PyTorch ciflow/xpu (xpu.yml) on the main branch, collect the failing XPU test cases from the most recent completed run(s), analyze the ROOT CAUSE of each failure with AI, and produce a list…
sablin39/tilelang-cuda-skills
Inspect a PyTorch baseline with torch.compile graphs/code and torch.profiler before drafting a custom kernel, then compare TileLang and PyTorch timelines, forward/backward regions, and allocations.
sablin39/tilelang-cuda-skills
Write human-facing TileLang design documents for proposed or completed kernel implementations.
sablin39/tilelang-cuda-skills
Draft, debug, and measure CUDA kernels and host launch workflows.
sablin39/tilelang-cuda-skills
Diagnose TileLang compilation, host argument, memory-liveness, runtime, and numerical failures.
sablin39/tilelang-cuda-skills
Describe model or operator HIR with TileFoundry, inspect topology and placement, run static cost analysis, visualize authored dataflow, and check a runtime implementation against its evaluator.
sablin39/tilelang-cuda-skills
Measure TileLang kernel and workflow performance, choose benchmark or timeline tools, and compare implementations fairly.
Works with
Categories
Run a bounded TileLang implementation and optimization task from a PyTorch reference or existing kernel through baseline analysis, a written plan, candidate edits, lowering inspection, validation…. Tilelang Optimization Loop is an agent skill from sablin39/tilelang-cuda-skills. Run a bounded TileLang implementation and optimization task from a PyTorch reference or existing kernel through baseline analysis, a written plan, candidate edits, lowering inspection, validation, measurement, and promotion.
Tilelang Optimization Loop fits situations like: iterative kernel work and reproducible handoffs; use a specialist directly for a single API; profiling question.
Run `npx skills add sablin39/tilelang-cuda-skills --skill tilelang-optimization-loop -a claude-code`. Or copy the skill folder (skills/tilelang-optimization-loop in sablin39/tilelang-cuda-skills) into .claude/skills/tilelang-optimization-loop in your project. Claude Code loads it when a task matches its description.
Run `npx skills add sablin39/tilelang-cuda-skills --skill tilelang-optimization-loop -a codex`. Or copy the skill folder (skills/tilelang-optimization-loop in sablin39/tilelang-cuda-skills) into .agents/skills/tilelang-optimization-loop in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add sablin39/tilelang-cuda-skills --skill tilelang-optimization-loop -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/tilelang-optimization-loop, .gemini/skills/tilelang-optimization-loop, .github/skills/tilelang-optimization-loop and .opencode/skills/tilelang-optimization-loop in your project.
SKILL.md names no scripts, command-line tools or credentials: Tilelang Optimization Loop is instructions for the agent only. Our summary lists: Python 3.
SKILL.md names 1 domain. As links in the text: github.com. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
No licence was found for Tilelang Optimization Loop or its repository. Without one, default copyright applies: ask the author before reusing or redistributing it.
About 2.7k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 6.9k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Tilelang Optimization Loop: Docstring (pytorch/pytorch, 104k stars), ExecuTorch Cortex-M Backend (pytorch/executorch, 5.1k stars), Qualcomm QNN Backend Development (pytorch/executorch, 5.1k stars) and Pt2 Bug Basher (pytorch/pytorch, 104k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
sablin39 (a GitHub user) maintains it in sablin39/tilelang-cuda-skills, which has 145 GitHub stars. The repository holds 10 skills in this directory. The repository was last updated on September 15, 2026.
Source: sablin39/tilelang-cuda-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.