Agent skill

Tilelang Optimization Loop

by sablin39 in sablin39/tilelang-cuda-skills

Run a bounded TileLang implementation and optimization task from a PyTorch reference or existing kernel through baseline analysis, a written plan, candidate edits, lowering inspection, validation…

No licenceAuto-check passedAI & LLM Engineering

Install Tilelang Optimization Loop

skills CLI
$ npx skills add sablin39/tilelang-cuda-skills --skill tilelang-optimization-loop -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install sablin39/tilelang-cuda-skills tilelang-optimization-loop --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/sablin39/tilelang-cuda-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/tilelang-optimization-loop .claude/skills/tilelang-optimization-loop && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
tilelang-optimization-loop
GitHub stars
145
Token cost
~2.7k tokens
SKILL.md length
1,310 words
Files
7 (incl. references)
Skills in repo
10
Repo updated
First seen
Licence
None found

At a glance

Run a bounded TileLang implementation and optimization task from a PyTorch reference or existing kernel through baseline analysis, a written plan, candidate edits, lowering inspection, validation…

  • Works in 6 steps: Establish the contract and baseline → Write the draft, then an executable plan → Implement and inspect one candidate → …
  • Iterative kernel work and reproducible handoffs
  • SKILL.md covers 1. Establish the contract and…, 2. Write the draft, then an…, 3. Implement and inspect one… and 4. Validate, then measure, plus 2 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Tilelang Optimization Loop is an agent skill from sablin39/tilelang-cuda-skills. Run a bounded TileLang implementation and optimization task from a PyTorch reference or existing kernel through baseline analysis, a written plan, candidate edits, lowering inspection, validation, measurement, and promotion. Use for iterative kernel work and reproducible handoffs; use a specialist directly for a single API, debugging, or profiling question.

Its SKILL.md is about 2.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 7 other files, including reference files (for example `references/execution-harness.md`, `references/experiment-record.md` and `references/external-tools.md`).

It sits in AI & LLM Engineering, covering Deep learning. It works with PyTorch. The repository describes itself as: Skills for writing tilelang and debugging with CUDA toolkits.

When your agent uses it

  • Iterative kernel work and reproducible handoffs
  • Use a specialist directly for a single API
  • Profiling question

Example prompts

  • “/tilelang-optimization-loop”

Requirements

  • Python 3

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Establish the contract and baseline
  2. Write the draft, then an executable plan
  3. Implement and inspect one candidate
  4. Validate, then measure
  5. Record a decision and choose the next experiment
  6. Reproduce and hand off

What it can do on your machine

Read from SKILL.md and the folder at commit 502a514. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Tilelang Optimization Loop loads about 2.7k tokens when it runs, and up to ~9.7k if it reads all its reference files. Until then it costs about 97 tokens; SKILL.md has 1,310 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~97
When it runs · the whole SKILL.md, loaded when a task matches
~2.7k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~9.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

Without a licence we can't republish the file, so here is its outline and opening line. It has 1,310 words (~2,720 tokens).

“Follow a repeatable experiment cycle in the task workspace. Keep the skill suite and its upstream knowledge separate from generated kernels and results. Resume from evidence that still matches the source, inputs, and environment. This adapts the KDA workflow to…”

— opening of SKILL.md by sablin39
name
tilelang-optimization-loop

Read the full SKILL.md on GitHub

Files

SKILL.md and 6 other files (references) in skills/tilelang-optimization-loop of sablin39/tilelang-cuda-skills.

  • SKILL.md
  • references/execution-harness.md
  • references/experiment-record.md
  • references/external-tools.md
  • references/joint-fwd-bwd.md
  • references/kernelwiki.md
  • references/worker-handoff.md

Open the folder on GitHubat commit 502a514

Compare with similar skills

Tilelang Optimization Loop next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Tilelang Optimization Loop compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Tilelang Optimization Loop this skillsablin39/tilelang-cuda-skills145—~2.7kAutomated safety check: PassNone
Docstringpytorch/pytorch104k2 repos~2.6kAutomated safety check: PassCustom licence
ExecuTorch Cortex-M Backendpytorch/executorch5.1k—~872Automated safety check: PassCustom licence
Qualcomm QNN Backend Developmentpytorch/executorch5.1k—~1.8kAutomated safety check: PassCustom licence
Pt2 Bug Basherpytorch/pytorch104k—~3.5kAutomated safety check: PassCustom licence
Formattingbrendanhasz/probflow175—~381Automated safety check: PassMIT

Similar skills

  • Docstring

    pytorch/pytorch

    Write docstrings for PyTorch functions and methods following PyTorch conventions.

    104k GitHub starsUsed in 2 repos~2.6k tokens
    AI & LLM EngineeringAuto-check passed
  • ExecuTorch Cortex-M Backend

    pytorch/executorch

    Developer guide for the Cortex-M (CMSIS-NN) backend in ExecuTorch: quantization pipeline, pass manager, tests and adding new ops.

    5.1k GitHub stars~872 tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Helps build, test and extend the Qualcomm AI Engine Direct (QNN) backend in ExecuTorch, with routes for new ops, model export, Buck-vs-CMake parity fixes and per-layer accuracy debugging.

    5.1k GitHub stars~1.8k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Pt2 Bug Basher

    pytorch/pytorch

    Debug PyTorch 2 compiler stack failures including Dynamo graph breaks, Inductor codegen errors, AOTAutograd crashes, and accuracy mismatches.

    104k GitHub stars~3.5k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Formatting

    brendanhasz/probflow

    Ensure consistent code formatting using the uv package manager and pre-commit.

    175 GitHub stars~381 tokensUpdated 11 days ago
    AI & LLM EngineeringAuto-check passed
  • Xpu CI Health Check

    intel/torch-xpu-ops

    Official

    Check PyTorch ciflow/xpu (xpu.yml) on the main branch, collect the failing XPU test cases from the most recent completed run(s), analyze the ROOT CAUSE of each failure with AI, and produce a list…

    115 GitHub stars~1.5k tokensUpdated today
    AI & LLM EngineeringAuto-check passed

More from sablin39/tilelang-cuda-skills

All 10 skills in this repo
  • Torch Profiling Tilelang Programs

    sablin39/tilelang-cuda-skills

    Inspect a PyTorch baseline with torch.compile graphs/code and torch.profiler before drafting a custom kernel, then compare TileLang and PyTorch timelines, forward/backward regions, and allocations.

    145 GitHub stars~1.3k tokensUpdated 22 days ago
    Auto-check passed
  • Writing Tilelang Design Docs

    sablin39/tilelang-cuda-skills

    Write human-facing TileLang design documents for proposed or completed kernel implementations.

    145 GitHub stars~2.2k tokensUpdated 22 days ago
    Auto-check passed
  • Cuda

    sablin39/tilelang-cuda-skills

    Draft, debug, and measure CUDA kernels and host launch workflows.

    145 GitHub stars~990 tokensUpdated 22 days ago
    Auto-check passed
  • Debugging Tilelang Programs

    sablin39/tilelang-cuda-skills

    Diagnose TileLang compilation, host argument, memory-liveness, runtime, and numerical failures.

    145 GitHub stars~1.6k tokensUpdated 22 days ago
    Auto-check passed
  • Designing With Tilefoundry

    sablin39/tilelang-cuda-skills

    Describe model or operator HIR with TileFoundry, inspect topology and placement, run static cost analysis, visualize authored dataflow, and check a runtime implementation against its evaluator.

    145 GitHub stars~3.4k tokensUpdated 22 days ago
    Auto-check passed
  • Profiling Tilelang Programs

    sablin39/tilelang-cuda-skills

    Measure TileLang kernel and workflow performance, choose benchmark or timeline tools, and compare implementations fairly.

    145 GitHub stars~1.6k tokensUpdated 22 days ago
    Auto-check passed

Works with

Questions about Tilelang Optimization Loop

What does Tilelang Optimization Loop do?

Run a bounded TileLang implementation and optimization task from a PyTorch reference or existing kernel through baseline analysis, a written plan, candidate edits, lowering inspection, validation…. Tilelang Optimization Loop is an agent skill from sablin39/tilelang-cuda-skills. Run a bounded TileLang implementation and optimization task from a PyTorch reference or existing kernel through baseline analysis, a written plan, candidate edits, lowering inspection, validation, measurement, and promotion.

When should I use Tilelang Optimization Loop?

Tilelang Optimization Loop fits situations like: iterative kernel work and reproducible handoffs; use a specialist directly for a single API; profiling question.

How do I install Tilelang Optimization Loop in Claude Code?

Run `npx skills add sablin39/tilelang-cuda-skills --skill tilelang-optimization-loop -a claude-code`. Or copy the skill folder (skills/tilelang-optimization-loop in sablin39/tilelang-cuda-skills) into .claude/skills/tilelang-optimization-loop in your project. Claude Code loads it when a task matches its description.

How do I install Tilelang Optimization Loop in Codex?

Run `npx skills add sablin39/tilelang-cuda-skills --skill tilelang-optimization-loop -a codex`. Or copy the skill folder (skills/tilelang-optimization-loop in sablin39/tilelang-cuda-skills) into .agents/skills/tilelang-optimization-loop in your project. Codex loads it when a task matches its description.

Can I use Tilelang Optimization Loop in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add sablin39/tilelang-cuda-skills --skill tilelang-optimization-loop -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/tilelang-optimization-loop, .gemini/skills/tilelang-optimization-loop, .github/skills/tilelang-optimization-loop and .opencode/skills/tilelang-optimization-loop in your project.

What does Tilelang Optimization Loop need to run?

SKILL.md names no scripts, command-line tools or credentials: Tilelang Optimization Loop is instructions for the agent only. Our summary lists: Python 3.

Does Tilelang Optimization Loop access the network?

SKILL.md names 1 domain. As links in the text: github.com. This is read from the text; nothing was executed.

Is Tilelang Optimization Loop safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Tilelang Optimization Loop use?

No licence was found for Tilelang Optimization Loop or its repository. Without one, default copyright applies: ask the author before reusing or redistributing it.

How many tokens does Tilelang Optimization Loop use?

About 2.7k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 6.9k tokens, read only when the agent opens those files.

What are the alternatives to Tilelang Optimization Loop?

Skills that share tags, products or a category with Tilelang Optimization Loop: Docstring (pytorch/pytorch, 104k stars), ExecuTorch Cortex-M Backend (pytorch/executorch, 5.1k stars), Qualcomm QNN Backend Development (pytorch/executorch, 5.1k stars) and Pt2 Bug Basher (pytorch/pytorch, 104k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Tilelang Optimization Loop?

sablin39 (a GitHub user) maintains it in sablin39/tilelang-cuda-skills, which has 145 GitHub stars. The repository holds 10 skills in this directory. The repository was last updated on September 15, 2026.

Source: sablin39/tilelang-cuda-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.