Paddle Build
PaddlePaddle/Paddle
A skill your agent uses when needing to compile, rebuild, or install Paddle from source after code changes.
Register, extend, or audit instructions in the table-driven T.ptx dialect (python/tvm/backend/cuda/ptx/table.py), and move the table to a newer PTX ISA version.
$ npx skills add mlc-ai/relax --skill tirx-ptx-dialect -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install mlc-ai/relax tirx-ptx-dialect --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/mlc-ai/relax.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/tirx-ptx-dialect .claude/skills/tirx-ptx-dialect && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "tirx-ptx-dialect" agent skill from https://github.com/mlc-ai/relax/tree/mlc/.agents/skills/tirx-ptx-dialect into .claude/skills/tirx-ptx-dialect/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tirx-ptx-dialect", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/mlc-ai/relax/tree/mlc/.agents/skills/tirx-ptx-dialectType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add mlc-ai/relax --skill tirx-ptx-dialect -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install mlc-ai/relax tirx-ptx-dialect --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/mlc-ai/relax.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.agents/skills/tirx-ptx-dialect .agents/skills/tirx-ptx-dialect && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "tirx-ptx-dialect" agent skill from https://github.com/mlc-ai/relax/tree/mlc/.agents/skills/tirx-ptx-dialect into .agents/skills/tirx-ptx-dialect/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tirx-ptx-dialect", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add mlc-ai/relax --skill tirx-ptx-dialect -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install mlc-ai/relax tirx-ptx-dialect --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/mlc-ai/relax.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.agents/skills/tirx-ptx-dialect .cursor/skills/tirx-ptx-dialect && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "tirx-ptx-dialect" agent skill from https://github.com/mlc-ai/relax/tree/mlc/.agents/skills/tirx-ptx-dialect into .cursor/skills/tirx-ptx-dialect/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tirx-ptx-dialect", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/mlc-ai/relax.git --path .agents/skills/tirx-ptx-dialect--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add mlc-ai/relax --skill tirx-ptx-dialect -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install mlc-ai/relax tirx-ptx-dialect --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/mlc-ai/relax.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.agents/skills/tirx-ptx-dialect .gemini/skills/tirx-ptx-dialect && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "tirx-ptx-dialect" agent skill from https://github.com/mlc-ai/relax/tree/mlc/.agents/skills/tirx-ptx-dialect into .gemini/skills/tirx-ptx-dialect/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tirx-ptx-dialect", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install mlc-ai/relax tirx-ptx-dialectInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add mlc-ai/relax --skill tirx-ptx-dialect -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/mlc-ai/relax.git skills-src && mkdir -p .github/skills && cp -r skills-src/.agents/skills/tirx-ptx-dialect .github/skills/tirx-ptx-dialect && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "tirx-ptx-dialect" agent skill from https://github.com/mlc-ai/relax/tree/mlc/.agents/skills/tirx-ptx-dialect into .github/skills/tirx-ptx-dialect/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tirx-ptx-dialect", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add mlc-ai/relax --skill tirx-ptx-dialect -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install mlc-ai/relax tirx-ptx-dialect --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/mlc-ai/relax.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.agents/skills/tirx-ptx-dialect .opencode/skills/tirx-ptx-dialect && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "tirx-ptx-dialect" agent skill from https://github.com/mlc-ai/relax/tree/mlc/.agents/skills/tirx-ptx-dialect into .opencode/skills/tirx-ptx-dialect/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tirx-ptx-dialect", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
tirx-ptx-dialectRegister, extend, or audit instructions in the table-driven T.ptx dialect (python/tvm/backend/cuda/ptx/table.py), and move the table to a newer PTX ISA version.
Tirx Ptx Dialect is an agent skill from mlc-ai/relax. Register, extend, or audit instructions in the table-driven T.ptx dialect (python/tvm/backend/cuda/ptx/table.py), and move the table to a newer PTX ISA version. Use when adding a PTX instruction or qualifier, widening an operand domain, fixing a ptxas certification failure, or checking the table's comments against the PTX ISA document.
Its SKILL.md is about 5.1k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in AI & LLM Engineering. It works with CUDA and Python. The licence is Apache-2.0.
6 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 201fa00. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
pythonFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Tirx Ptx Dialect loads about 5.1k tokens when it runs. Until then it costs about 89 tokens; SKILL.md has 2,452 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from mlc-ai/relax at commit 201fa00, republished under its Apache-2.0 licence (© mlc-ai). 2,452 words, ~5,142 tokens.
.claude/skills/tirx-ptx-dialect/SKILL.md (or your agent's skills folder).T.ptx)T.ptx.<mnemonic>.<qualifiers>(operands...) is not hand-written per
instruction. One data table describes every instruction family; one engine
resolves calls, renders an inline-asm helper, and registers the TVM op; three
generators derive the IDE stub, a coverage table and a helper dump from the
same table. Adding an instruction means adding data, then proving it
against ptxas.
| file | role |
|---|---|
python/tvm/backend/cuda/ptx/table.py | the table: InstructionEntry / ModifierSlot / OperandSlot, check functions, variant enumeration. tvm-free. Read its three dataclass docstrings before editing. |
python/tvm/backend/cuda/ptx/render.py | renders one variant to a __forceinline__ __device__ helper (asm or asm volatile, according to asm_volatile). CBinding = how each TVM dtype binds a register; BRIDGE = three scoped cases covering the two register classes inline asm cannot bind directly (.pred, and the block-local .b8 registers behind the e2m1x2 and st_async_b8reg tokens) — not a general .b8 mechanism. tvm-free. |
python/tvm/backend/cuda/ptx/engine.py | register_table() makes each entry a TVM op tirx.ptx.<name> with a generic codegen; PTXNamespace resolves T.ptx attribute chains and the string form T.ptx["..."]; trace-time coercion and check/imm_check. |
python/tvm/backend/cuda/ptx/gen_stubs.py | python -m tvm.backend.cuda.ptx.gen_stubs -o python/tvm/script/tirx.pyi — the checked-in stub; a test diffs it against the generator. |
python/tvm/backend/cuda/ptx/gen_helpers.py | python -m tvm.backend.cuda.ptx.gen_helpers <family> — dump the exact helper source of every variant without compiling a kernel (it still imports tvm, so it needs the usual built worktree on PYTHONPATH). |
python/tvm/backend/cuda/ptx/gen_coverage.py | markdown coverage table of the whole dialect. |
tests/python/tirx/codegen/test_ptx_dialect.py | per-ISA-section dispatch tests, structural invariants, ptxas certification. |
tests/python/tirx/codegen/test_ptx_cvt.py | one case per registered cvt syntax line (_FORM_CASES); coverage of every cvt entry is asserted. |
tests/python/tirx/codegen/test_ptx_addr.py | T.ptx.addr(base, byte_offset) immediates. |
The namespace itself is installed by python/tvm/backend/cuda/__init__.py::script_namespaces();
nothing there changes when the table grows.
DO NOT introduce a new mechanism in an
InstructionEntrywithout the user's explicit approval. Stop, explain why the existing entry model is insufficient and what codegen or validation behavior the mechanism would add, then wait for approval before implementing it.
The first four are tested mechanically — a new entry that breaks one fails the suite. The last one (what the comments claim) is audited by hand, §4.
mnemonic when the register-group shape or the result structure changes
(scalar vs vector destination, mul vs mul.wide), or when an optional
operand is told apart only by call arity ({, count}, state vs the _
sink); the engine then selects by arity and by which tokens are written.
An operand whose presence is fixed by one written qualifier stays in the
entry with a 0-or-1-register lanes function — {, cache_policy} with
.L2::cache_hint, {, ctaMask} with .multicast::cluster use
lanes=_present_lanes("<slot>"). Variants that only add or remove dotted
qualifiers are optional slots of one entry.@p
wrapper and the BRIDGE boundary conversions. raw_render is the escape
hatch and every user of it must be listed in RAW_ENTRIES inside
test_ptx_single_instruction_invariant.renderings() enumerates the
product of modifiers × operand dtypes × closed immediates × @p (for
entries without a destination) and full certification assembles all of it
through ptxas at cert_arch. Four things get representative coverage
instead of the full product, and each says so in its docstring: an OPEN
immediate is certified at imm_combos's samples (default "0"), which
proves the shape, not the caller's constant; sink masks are walked once at
one representative modifier combination per sink domain (renderings);
address immediates are a small separate axis (_addr_offset_samples);
and the destination-predicated pred= helpers (*_pred_undef /
*_pred_keep) are not enumerated at all — cover those with a
production-shaped test when a call site needs them. Conversely a check()
must never reject a legal form to make certification pass — narrow with
evidence.test_ptx_dispatch_unambiguous); helper names must be unique
(test_ptx_all_variants_render_unique).sm_XX floors, and a NOT REGISTERED:
note with the reason for every syntax line or token deliberately left out.
A fact that comes from ptxas rather than the document is marked MEASURED
and quotes the ptxas message. Section and table numbers follow the ISA
version named in the module docstring (currently PTX ISA 9.4, the CUDA
13.4 developer-preview document, see §5).Open the instruction's section in the ISA document and transcribe, in this order:
name is the table key and must be a Python
identifier. Dotted mnemonics are a family plus single-choice slots
(cvta.to.shared → cvta + slots), or mnemonic="st.bulk" with
name="st_bulk" when the dot is part of the instruction's identity.
Several entries sharing a mnemonic (mov, mbarrier, mma) use
distinct names (mov, mov_pack_2, ...) and the same mnemonic.
Entries that collide with an existing family on the same shape need a
suffix (add_int vs add/add_half) and are told apart by their tokens.ModifierSlot(name, choices, optional)), in asm render order —
exactly the {.qual} positions of the syntax line, each with the ISA's
.qual = { ... } list as choices. Tokens with :: are written as-is
("shared::cta"); users type shared__cta. Python keywords get a
trailing underscore only at the call site (global_).OperandSlot), in PTX operand order, destinations first
because PTX writes them first:rw: "w" destination, "rw" accumulator ("+"), "r" input. A
"w" operand binds "="; it does not block pred= — under a false
predicate the value is undefined unless the caller passes
preserve_dst=True, which switches the binding to "+" (see
test_ptx_predicated_destination_*). Those two helpers are rendered on
demand, not by renderings().kind: "reg" (default), "addr" ([%k], space from space= or the
entry's space slot; set allow_imm_offset=True only for an independent
byte address), "ptr" (raw pointer value), "imm" (text operand:
literal= fixed by the ISA, choices= closed caller set, neither = open
immediate certified at samples; validate open ones with imm_check).dtype: a PTX_TYPE_DTYPES key, a slot name, or a module-level pure
function of the modifier map (_wide_dtype). None means the entry's
type slot. dtypes= narrows/widens the TVM dtype domain independently
(see the relaxed-carrier note in §6 before widening anything).lanes (int or function of the modifier map) for brace-enclosed register
groups; vector=, bracket=, pipe= for the few composite spellings;
sinkable= only where the ISA lets a caller write _ for a lane.check: one pure module-level function mod_map -> error | None per
entry with a one-line docstring (it is surfaced by the stub and the coverage
table). Encode only what the ISA syntax block and notes state; every
restriction that comes from ptxas instead is a separate MEASURED clause.
Never a lambda: frozen dataclasses hash callables by identity.cert_arch: the maximum sm_XX floor over the entry's variants
(default is PTX_ARCH, sm_90). Certifying below a variant's floor makes
ptxas report legal forms as illegal, and that verdict would then get baked
into a check().orders_memory=True for fences/barriers/waits (no address operand but the
"memory" clobber is needed); asm_volatile stays at its default.Worked example — the whole registration of movmatrix (commit a29a5e97ba,
4 files, 40 lines):
# movmatrix per PTX ISA 9.7.16.5.17 -- transpose one distributed m8n8
# matrix whose 16-bit elements are carried by one b32 register per lane.
#
# movmatrix.sync.aligned.m8n8.trans.b16 d, a;
InstructionEntry(
name="movmatrix",
slots=(
ModifierSlot("sync", ("sync",)),
ModifierSlot("aligned", ("aligned",)),
ModifierSlot("shape", ("m8n8",)),
ModifierSlot("trans", ("trans",)),
ModifierSlot("type", ("b16",)),
),
cert_arch="sm_75",
operands=(
OperandSlot("d", rw="w", dtype="b32"),
OperandSlot("a", dtype="b32"),
),
),Put the entry under its ISA chapter banner in _ENTRIES (# PTX ISA 9.7.x — ...), next to the instructions it shares a mnemonic with.
Run from the TVM repo root with the workspace PYTHONPATH (see tir-test).
check() that filters out every
combination, a duplicate name, or a bad allow_imm_offset slot fails at
import.python -c "from tvm.backend.cuda.ptx.table import TABLE, variants; e=TABLE['<name>']; print(len(variants(e)))"
python -m tvm.backend.cuda.ptx.gen_helpers <name> | head -60 # eyeball the asmtest_ptx_<section>_dispatch) asserting the exact emitted line, e.g.
assert "mul.wide.s32 %0, %1, %2;" in src;cvt line → a _FORM_CASES row in test_ptx_cvt.py
(test_cvt_cases_cover_every_registered_entry fails otherwise);raw_render → RAW_ENTRIES;pred=) call site → a production-shaped
codegen test, since certification does not enumerate those helpers;test_codegen_cuda.py::test_megamoe_extracted_intrinsics_codegen
(that test covers one workload, not the dialect).test_ptx_stub_up_to_date fail until you do:python -m tvm.backend.cuda.ptx.gen_stubs -o python/tvm/script/tirx.pyiPTX_CERT): structural invariants + dispatch tests +
a seeded ptxas sample.python -m pytest tests/python/tirx/codegen/test_ptx_dialect.py tests/python/tirx/codegen/test_ptx_cvt.py tests/python/tirx/codegen/test_ptx_addr.py -qtest_ptx_all_variants_render_unique ends with assert total == <N>;
its failure message prints the new total — update the constant and
re-run.-n auto on a many-core box mostly burns memory).
~3.5 min at -n 16, ~6.5 min at -n 8 on a B200-class host for the
~760k-variant PTX ISA 9.4 table (test_ptx_all_variants_render_unique
pins the exact count).PTX_CERT=1 python -m pytest -n 16 -q \
tests/python/tirx/codegen/test_ptx_dialect.py::test_ptx_all_helpers_certify<arch> batch <k> and the ptxas message. To find the
variant: render the failing family with gen_helpers, or assemble the
batch yourself (nvcc -arch=<arch> -ptx then ptxas, map the error line
back to the preceding .entry). Fix by narrowing the domain with a
MEASURED clause, never by loosening the invariant.pre-commit run --files <changed files> and commit as
feat(lower-tirx): support PTX <instruction> (Conventional Commits, see
the workspace CLAUDE.md).The comments are load-bearing (they are what a reviewer checks the data
against), so audit them the way the data is certified: download the ISA HTML
for the version the module docstring names, convert to text, and verify every
section number, quoted sentence, reproduced syntax line, type list and
sm_XX note. Things that have gone wrong before and are worth a targeted
pass: a NOT REGISTERED bullet naming something the table registers a few
hundred lines later; a section number copied from a neighbouring
instruction; a version number attached to the wrong syntax line; "the N-th
syntax line" ordinals; Table NN numbers, which are global and shift between
ISA versions.
nvcc -ptx only shows the
version the front end chose to emit, and -ptx never validates inline
asm:command -v nvcc ptxas && nvcc --version | tail -1 && ptxas --version | tail -1
printf '.version 9.4\n.target sm_107a\n.address_size 64\n.visible .entry k() { ret; }\n' \
| ptxas -arch=sm_107a - -o /dev/null # CUDA 13.4: accepted.version directive above what ptxas implements is refused outright
("Unsupported .version ...; current version is '9.4'"), which is the
quickest way to read a toolkit's ceiling. Then let the certification suite
(which compiles real inline-asm helpers to cubin) be the final word.table.py: "Section and table numbers cite PTX ISA 9.4", URL
docs.nvidia.com/cuda/developer-preview/13.4/parallel-thread-execution/;
once CUDA 13.4 ships, the stable copy is
docs.nvidia.com/cuda/archive/13.4.0/parallel-thread-execution/). The
live docs.nvidia.com/cuda/parallel-thread-execution/ page is whatever
version is newest and its numbering differs; every ISA version stays
available under docs.nvidia.com/cuda/archive/<cuda version>/.MEASURED clause in table.py / render.py names the toolkit it
was taken on (currently CUDA 13.4). A toolkit bump re-measures all of them
(§6 step 5); a clause that names an older toolkit is a bug.nvcc/ptxas; the on-GPU round-trip tests need
a driver that can load a cubin from that toolkit. A newer toolkit with an
older driver can therefore certify a table but not run it.PTX_ARCH (env, default sm_90) is the certification arch for entries
without cert_arch.Do this as one commit series, in this order:
.version (§5). Until then, a new form cannot
be registered: certification would fail, and a check() written against
an older ptxas would silently delete the coverage later.<num> <title>),
and derive a rule(old) -> new function from where the new version
inserted sections; cross-check it by asserting every old 9.7.* entry
maps to an identically titled new entry. Never trust title matching
alone (duplicate titles such as "mov" x2 or "Async Proxy" x2 collide) and
never shift chapters only — a chapter-only shift was the exact mistake
made the first time this table changed versions.table.py,
engine.py, render.py, tests/python/tirx/codegen/test_ptx_*.py and
test_codegen_cuda.py, skipping any block that already carries the new
numbering (_PTX_94_ENTRIES was such a block for 9.3→9.4). Hand-fix the
forms a regex cannot: ranges that span an insertion (9.7.x.a-b), brace
groups (9.7.4.{1,2,3}), sibling shorthand (.12, / .17), global
Table NN numbers, and :line offsets (re-derive against the new
section text from the quoted sentence, or drop the offset and keep the
quote). Then regenerate the stub (§3).cert_arch. Keep the new version's additions
grouped (the _PTX_94_ENTRIES list is the 9.4 group).MEASURED clause on the new toolkit and reword it
to name that toolkit. Registered forms are re-proven by full
certification; excluded forms need a targeted probe each (raw PTX through
ptxas -arch=<a> for grammar claims, an exact force-inlined kernel
through nvcc for crash claims — test_ptx_dialect._certification_kernel
builds that shape). Quote the diagnostic the new ptxas actually prints;
texts change between releases. If a gap has closed, either widen the
domain with certification or say so in the clause — never leave a
sentence that attributes the restriction to a toolkit you no longer use.
ISA §9.4.1 Tables 27/28 are an upper bound and say so ("some combinations
may still be invalid for a particular instruction"); the _cvt_dst_dtypes
/ _cvt_src_dtypes docstrings record the measured gaps.Worked example, PTX ISA 9.3 → 9.4:
9.7.6 Alternate Floating-Point Instructions, so
every 9.7.N with N ≥ 6 became 9.7.(N+1) (comparison 9.7.6→9.7.7,
data movement 9.7.9→9.7.10, fabric 9.7.10→9.7.11, sync
9.7.14→9.7.15, mma 9.7.15→9.7.16, wgmma 9.7.16→9.7.17, tcgen05
9.7.17→9.7.18, misc 9.7.20→9.7.21).applypriority.async.bulk[.tensor] were inserted as
9.7.10.18/19, so 9.7.9.18..27 → 9.7.10.20..29 (+2), and "Overriding
tensor property value" as 9.7.10.28.5.2, so
9.7.9.26.5.2..4 → 9.7.10.28.5.3..5.9.7.18.10.8,
so 9.7.17.10.8.x → 9.7.18.10.9.x and 9.7.17.10.9.x → 9.7.18.10.10.x.52/53 → 59/61,
tensormap new_val validity 33 → 36; the relaxed-typing tables 27/28
kept their numbers.test_ptx_cvt.py, render.py:
cvt as 9.7.9.21, which in 9.3 is cvta) — check every file's baseline
before mapping it..L2::cache_hint diagnostics
on .shared/.local/.volatile, the bare
clusterlaunchcontrol.query_cancel diagnostic, multimem.st.async
accepting 16/32-bit sources for its byte forms, and the bf16 atom
bit-bucket forms compiling again (still withheld pending certification).9.7.6.1-4 (alternate FP
x4 arithmetic, sm_100a/sm_103a), 9.7.10.28.1.3 (report mechanisms),
9.7.18.10.7.2.6 / .3.6 (block16 K=128/256 scale layouts).© mlc-ai, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .agents/skills/tirx-ptx-dialect of mlc-ai/relax.
Open the folder on GitHubat commit 201fa00
Tirx Ptx Dialect next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Tirx Ptx Dialect this skillmlc-ai/relax | 175 | — | ~5.1k | Automated safety check: Pass | Apache-2.0 | |
| Paddle BuildPaddlePaddle/Paddle | 24k | — | ~1k | Automated safety check: Pass | Apache-2.0 | |
| Paddle Design CompilerPaddlePaddle/Paddle | 24k | — | ~3.6k | Automated safety check: Pass | Apache-2.0 | |
| Paddle DebugPaddlePaddle/Paddle | 24k | — | ~1.4k | Automated safety check: Pass | Apache-2.0 | |
| Optimize OpCVCUDA/CV-CUDA | 2.7k | — | ~834 | Automated safety check: Pass | Custom licence | |
| Ako4allTongmingLAIC/AKO4ALL | 369 | — | ~4k | Automated safety check: Pass | MIT |
PaddlePaddle/Paddle
A skill your agent uses when needing to compile, rebuild, or install Paddle from source after code changes.
PaddlePaddle/Paddle
A skill your agent uses when working with Paddle 3.0 compiler full pipeline: SOT (Symbolic Opcode Translator) for bytecode-level dy2st graph capture, PIR (Paddle IR) for SSA-based intermediate…
PaddlePaddle/Paddle
在 Paddle 代码库中定位问题并输出高质量调试报告的专用技能。当遇到以下场景时优先使用:(1) Paddle 框架 bug 调试,(2) 算子实现问题排查,(3) 训练脚本异常诊断,(4) 分布式训练故障定位,(5) CUDA/GPU 相关错误处理,(6) 需要生成结构化调试报告。
CVCUDA/CV-CUDA
Drive a single-operator optimization campaign per .agents/guidance/OPTIMIZATIONGUIDELINES.md, with a deterministically enforced definition-of-done and versioned MR summary.
TongmingLAIC/AKO4ALL
Drive an agentic loop that iteratively optimizes a GPU kernel for maximum speedup.
PaddlePaddle/Paddle
PaddlePaddle (飞桨) C++ 算子开发指南。提供从 YAML 配置、InferMeta 函数、Kernel 实现、Python API 封装、单元测试到编译验证的完整算子开发流程指导。在以下场景使用此 skill:(1) 为 Paddle 框架新增 C++ 算子 (2) 修改或调试已有 Paddle 算子 (3) 编写算子的 YAML…
Categories
Register, extend, or audit instructions in the table-driven T.ptx dialect (python/tvm/backend/cuda/ptx/table.py), and move the table to a newer PTX ISA version. Tirx Ptx Dialect is an agent skill from mlc-ai/relax.py), and move the table to a newer PTX ISA version.
Tirx Ptx Dialect fits situations like: adding a PTX instruction; widening an operand domain; fixing a ptxas certification failure; checking the tables comments against the PTX ISA document.
Run `npx skills add mlc-ai/relax --skill tirx-ptx-dialect -a claude-code`. Or copy the skill folder (.agents/skills/tirx-ptx-dialect in mlc-ai/relax) into .claude/skills/tirx-ptx-dialect in your project. Claude Code loads it when a task matches its description.
Run `npx skills add mlc-ai/relax --skill tirx-ptx-dialect -a codex`. Or copy the skill folder (.agents/skills/tirx-ptx-dialect in mlc-ai/relax) into .agents/skills/tirx-ptx-dialect in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add mlc-ai/relax --skill tirx-ptx-dialect -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/tirx-ptx-dialect, .gemini/skills/tirx-ptx-dialect, .github/skills/tirx-ptx-dialect and .opencode/skills/tirx-ptx-dialect in your project.
Going by SKILL.md and its folder, Tirx Ptx Dialect needs the command-line tools its instructions call (python). Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Tirx Ptx Dialect is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 5.1k tokens (SKILL.md is roughly 21k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Tirx Ptx Dialect: Paddle Build (PaddlePaddle/Paddle, 24k stars), Paddle Design Compiler (PaddlePaddle/Paddle, 24k stars), Paddle Debug (PaddlePaddle/Paddle, 24k stars) and Optimize Op (CVCUDA/CV-CUDA, 2.7k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
mlc-ai (a GitHub organization) maintains it in mlc-ai/relax, which has 175 GitHub stars. The repository was last updated on October 4, 2026.
Source: mlc-ai/relax on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.