Agent skill

Ttir Decomposition For Ttmetal

by tenstorrent in tenstorrent/tt-mlir

Add a new composite op decomposition pattern to the TTMetal pipeline.

Apache-2.0Auto-check passed

Install Ttir Decomposition For Ttmetal

skills CLI
$ npx skills add tenstorrent/tt-mlir --skill ttir-decomposition-for-ttmetal -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install tenstorrent/tt-mlir ttir-decomposition-for-ttmetal --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/tenstorrent/tt-mlir.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/ttir-decomposition-for-ttmetal .claude/skills/ttir-decomposition-for-ttmetal && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
ttir-decomposition-for-ttmetal
GitHub stars
314
Token cost
~2.2k tokens
SKILL.md length
563 words
Files
1
Skills in repo
9
Repo updated
First seen
Licence
Apache-2.0

At a glance

Add a new composite op decomposition pattern to the TTMetal pipeline.

  • Works in 6 steps: Understand the Op → Add a Pattern to DecomposeComposites.cpp → Creating TTIR Ops in Decomposition Code → …
  • The user wants to decompose/lower a high-level TTIR op (e.g
  • SKILL.md covers When to Use, Architecture, Files to Modify and Step-by-Step, plus 1 more section
  • Calls cmake and pytest

What it does

Ttir Decomposition For Ttmetal is an agent skill from tenstorrent/tt-mlir. Add a new composite op decomposition pattern to the TTMetal pipeline. Use when the user wants to decompose/lower a high-level TTIR op (e.g. rmsnorm, sdpa, layernorm, softmax) into primitive TTIR ops (matmul, add, multiply, etc.) for the D2M/TTMetal backend. Also trigger when the user mentions "decomposition pattern", "decompose op for ttmetal", or "lower op to primitives".

Its SKILL.md is about 2.2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It works with C++ and Python. The repository describes itself as: Tenstorrent MLIR compiler. The licence is Apache-2.0.

When your agent uses it

  • The user wants to decompose/lower a high-level TTIR op (e.g
  • The user mentions decomposition pattern
  • Decompose op for ttmetal
  • Lower op to primitives

Example prompts

  • “decomposition pattern”
  • “decompose op for ttmetal”
  • “lower op to primitives”
  • “/ttir-decomposition-for-ttmetal”

Requirements

  • Python 3

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Understand the Op
  2. Add a Pattern to DecomposeComposites.cpp
  3. Creating TTIR Ops in Decomposition Code
  4. Add MLIR Lit Tests
  5. Add Python Builder Tests
  6. Build and Iterate

What it can do on your machine

Read from SKILL.md and the folder at commit 78b7044. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • cmake
    • pytest

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Ttir Decomposition For Ttmetal loads about 2.2k tokens when it runs. Until then it costs about 102 tokens; SKILL.md has 563 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~102
When it runs · the whole SKILL.md, loaded when a task matches
~2.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from tenstorrent/tt-mlir at commit 78b7044, republished under its Apache-2.0 licence (© tenstorrent). 563 words, ~2,165 tokens.

Download SKILL.mdSave it as .claude/skills/ttir-decomposition-for-ttmetal/SKILL.md (or your agent's skills folder).
name
ttir-decomposition-for-ttmetal
description
Add a new composite op decomposition pattern to the TTMetal pipeline. Use when the user wants to decompose/lower a high-level TTIR op (e.g. rms_norm, sdpa, layer_norm, softmax) into primitive TTIR ops (matmul, add, multiply, etc.) for the D2M/TTMetal backend. Also trigger when the user mentions "decomposition pattern", "decompose op for ttmetal", or "lower op to primitives".

TTIR Composite Op Decomposition for TTMetal

Decompose a high-level fused TTIR op into primitive TTIR ops so the D2M/TTMetal backend can lower them individually. The TTNN backend keeps native fused ops; these decomposition patterns only run in the TTMetal pipeline via the unified TTIRDecomposeComposites pass.

When to Use

  • The op has no D2M/TTMetal lowering but can be expressed as a sequence of ops that do (matmul, add, multiply, reduce, reshape, permute, etc.).
  • The TTNN backend already has native support, so it skips the decomposition.

Architecture

All composite decompositions live in a single pass (ttir-decompose-composites) that uses MLIR's greedy pattern rewriter. Each op decomposition is an OpRewritePattern<T> with a configurable benefit level that controls application order. For example, SDPA has higher benefit than softmax so it runs first — the softmax ops it produces are then caught by the softmax pattern on subsequent rewriter iterations.

Files to Modify

FileAction
lib/Dialect/TTIR/Transforms/DecomposeComposites.cppEdit — add a new OpRewritePattern
include/ttmlir/Dialect/TTIR/Transforms/Passes.tdEdit — update description if desired
test/ttmlir/Dialect/TTIR/Transforms/metal_composite_decompositions.mlirEdit — add FileCheck tests
test/python/golden/d2m/test_composite_ops.pyEdit — add Python builder tests

You should NOT need to touch CMakeLists.txt or Passes.td pass registration in the common case. However, when adding a new composite decomposition, verify that ttir-decompose-composites is scheduled in the relevant D2M pipeline in D2MPipelines.cpp, and update that pipeline if necessary.

Step-by-Step

1. Understand the Op

Read the op definition in include/ttmlir/Dialect/TTIR/IR/TTIROps.td. Note tensor shapes, attributes (optional mask, scale, etc.), and the mathematical decomposition into primitives.

2. Add a Pattern to DecomposeComposites.cpp

Open lib/Dialect/TTIR/Transforms/DecomposeComposites.cpp and add a new OpRewritePattern<YourOp> struct. Follow the existing patterns as examples.

Pattern template:

cpp
struct DecomposeYourOpPattern : public OpRewritePattern<YourOp> {
  using OpRewritePattern<YourOp>::OpRewritePattern;

  LogicalResult matchAndRewrite(YourOp op,
                                PatternRewriter &rewriter) const override {
    Location loc = op.getLoc();
    // ... decomposition logic using rewriter.create<T>(...) ...
    rewriter.replaceOp(op, result);
    return success();
  }
};

Then register the pattern in TTIRDecomposeComposites::runOnOperation():

cpp
void runOnOperation() final {
  RewritePatternSet patterns(&getContext());
  patterns.add<DecomposeSDPAPattern>(&getContext(), /*benefit=*/2);
  patterns.add<DecomposeRMSNormPattern>(&getContext(), /*benefit=*/1);
  patterns.add<DecomposeSoftmaxPattern>(&getContext(), /*benefit=*/0);
  patterns.add<DecomposeYourOpPattern>(&getContext(), /*benefit=*/N);  // NEW

  if (failed(applyPatternsGreedily(getOperation(), std::move(patterns)))) {
    signalPassFailure();
  }
}

Benefit ordering: If your decomposition produces ops that another pattern needs to decompose further (e.g. SDPA produces softmax), give your pattern a higher benefit number than the downstream pattern.

Key conventions:

  • Use OpRewritePattern<T> and PatternRewriter, not IRRewriter.
  • Always return success() after rewriter.replaceOp(op, result).
  • Call rewriter.replaceOp(op, result) at the end to replace the original.
3. Creating TTIR Ops in Decomposition Code

Common op creation patterns (use rewriter.create<T>(...)):

cpp
// Elementwise binary (add, multiply, subtract, etc.)
auto add = rewriter.create<AddOp>(loc, resultType, lhs, rhs);

// MatmulOp
auto mm = rewriter.create<MatmulOp>(loc, resultType, a, b);

// SoftmaxOp
auto sm = rewriter.create<SoftmaxOp>(loc, resultType, input,
    rewriter.getSI32IntegerAttr(dim),
    rewriter.getBoolAttr(false));

// FullOp (scalar constant broadcast to shape)
auto full = rewriter.create<FullOp>(loc, resultType,
    rewriter.getF32FloatAttr(value));

// ReshapeOp
SmallVector<int32_t> shapeI32(newShape.begin(), newShape.end());
auto reshape = rewriter.create<ReshapeOp>(loc, newType, input,
    rewriter.getI32ArrayAttr(shapeI32));

// PermuteOp
auto permute = rewriter.create<PermuteOp>(loc, permutedType, input,
    rewriter.getDenseI64ArrayAttr(permutation));

// MeanOp (reduction)
auto mean = rewriter.create<MeanOp>(loc, reducedType, input,
    rewriter.getBoolAttr(/*keep_dim=*/true),
    rewriter.getI32ArrayAttr(reduceDims));

// RsqrtOp (unary)
auto rsqrt = rewriter.create<RsqrtOp>(loc, type, input);
Show full SKILL.md (233 more words)Show less
4. Add MLIR Lit Tests

Add test functions and FileCheck assertions to test/ttmlir/Dialect/TTIR/Transforms/metal_composite_decompositions.mlir.

The file uses a single pass (--ttir-decompose-composites) with multiple check prefixes. Add a new check prefix for your op and a new RUN line:

// RUN: ttmlir-opt --ttir-decompose-composites %s | FileCheck %s --check-prefix=YOUROP

Then add test functions:

// YOUROP-LABEL: func.func @your_op_basic
// YOUROP-NOT: ttir.your_op
// YOUROP: "ttir.multiply"
// YOUROP: "ttir.add"
// YOUROP: return
func.func @your_op_basic(%input: tensor<...xbf16>) -> tensor<...xbf16> {
  %0 = "ttir.your_op"(%input) <{...}> : (...) -> ...
  return %0 : ...
}
5. Add Python Builder Tests

Add tests to test/python/golden/d2m/test_composite_ops.py. This file contains all composite decomposition tests for the TTMetal pipeline.

Follow the existing patterns (SDPA, RMSNorm, softmax) as examples:

python
@pytest.mark.parametrize("shape", [...])
@pytest.mark.parametrize("target", ["ttmetal"])
def test_your_op_decomposition(
    shape: Shape,
    target: str,
    request,
    device,
):
    """Test your_op decomposition for the TTMetal pipeline."""

    def module(builder: TTIRBuilder):
        @builder.func([shape], [torch.float32])
        def your_op(
            in0: Operand,
            builder: TTIRBuilder,
            unit_attrs: Optional[List[str]] = None,
        ):
            return builder.your_op(in0, ..., unit_attrs=unit_attrs)

    compile_and_execute_ttir(
        module,
        target=target,
        **get_request_kwargs(request),
        device=device,
    )

Key points:

  • Always use target="ttmetal" — decompositions only run in the TTMetal pipeline.
  • Use compile_and_execute_ttir from builder.base.builder_apis.
  • Parametrize over the attributes that affect decomposition logic (e.g. causal vs non-causal, with/without weight, different dims).
6. Build and Iterate

Run ./build_and_test.sh or:

bash
source env/activate
cmake --build build

Fix compilation errors, then test with the lit test:

bash
build/bin/ttmlir-opt --ttir-decompose-composites test/ttmlir/Dialect/TTIR/Transforms/metal_composite_decompositions.mlir

And the Python tests:

bash
pytest -svv test/python/golden/d2m/test_composite_ops.py

Reference Implementations

All live in lib/Dialect/TTIR/Transforms/DecomposeComposites.cpp:

  • DecomposeRMSNormPattern (benefit 1): Decomposes rms_norm(x, w, b, eps) into x^2 -> mean -> +eps -> rsqrt -> *x -> *w -> +b.

  • DecomposeSDPAPattern (benefit 2): Decomposes scaled_dot_product_attention(Q, K, V, mask) into Q @ K^T -> *scale -> +mask -> softmax -> @ V, with GQA head expansion via reshape. Produces ttir.softmax ops that the softmax pattern then decomposes.

  • DecomposeSoftmaxPattern (benefit 0): Decomposes softmax(x, dim) into max -> subtract -> exp -> sum -> div (uses ttir.div rather than reciprocal -> multiply to work around a broadcast-multiply bug). When numericStable=false, the max-subtract step is skipped.

  • Inner-min decomposition (in TTIRToD2M): D2MInnerMinDecompositionRewriter in lib/Conversion/TTIRToD2M/TTIRToD2M.cpp rewrites inner-dim ttir.min into neg(max(neg(x))) during conversion (there is no tile_reduce_min kernel). Outer-dim min reductions use the accumulation path with d2m.tile_minimum.

© tenstorrent, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/ttir-decomposition-for-ttmetal of tenstorrent/tt-mlir.

Open the folder on GitHubat commit 78b7044

Compare with similar skills

Ttir Decomposition For Ttmetal next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Ttir Decomposition For Ttmetal compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Ttir Decomposition For Ttmetal this skilltenstorrent/tt-mlir314—~2.2kAutomated safety check: PassApache-2.0
Paddle BuildPaddlePaddle/Paddle24k—~1kAutomated safety check: PassApache-2.0
Fory Releaseapache/fory4.6k—~2.9kAutomated safety check: PassApache-2.0
CodeQL Security Scantrailofbits/skills7.5k—~4.6kAutomated safety check: NotesCC-BY-SA-4.0
pybind11 Release Preparationpybind/pybind1118k—~1.7kAutomated safety check: PassCustom licence
Paddle Eager GraphPaddlePaddle/Paddle24k—~562Automated safety check: PassApache-2.0

Similar skills

  • Paddle Build

    PaddlePaddle/Paddle

    A skill your agent uses when needing to compile, rebuild, or install Paddle from source after code changes.

    24k GitHub stars~1k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Fory Release

    apache/fory

    Prepare an Apache Fory release candidate from a clean release branch, including the version bump, RC tag, JVM staging, ASF source artifacts, SVN upload, and vote email.

    4.6k GitHub stars~2.9k tokensUpdated yesterday
    Auto-check passed
  • CodeQL Security Scan

    trailofbits/skills

    Official

    Scans a codebase for vulnerabilities with CodeQL's data flow and taint tracking in run-all or important-only modes, including data extensions for project-specific sources and sinks.

    7.5k GitHub stars~4.6k tokensUpdated yesterday
    SecurityAuto-check: notes
  • Opens the pybind11 release-preparation pull request: picking the release base, bumping the version in common.h and integrating the changelog, following docs/release.rst.

    18k GitHub stars~1.7k tokensUpdated 2 days ago
    DevelopmentAuto-check passed
  • Paddle Eager Graph

    PaddlePaddle/Paddle

    A skill your agent uses when navigating Paddle eager-mode (dynamic graph) source code, tracing forward/backward execution, debugging autograd issues, understanding PyLayer, or investigating…

    24k GitHub stars~562 tokensUpdated yesterday
    DevelopmentAuto-check passed
  • Onnxtxt

    onnx/onnx

    Read or write ONNX text format ("onnxtxt"). An agent skill from onnx/onnx.

    22k GitHub stars~1.3k tokensUpdated today
    AI & LLM EngineeringAuto-check passed

More from tenstorrent/tt-mlir

All 9 skills in this repo
  • Add Op

    tenstorrent/tt-mlir

    How to add a new operation (op) to the tt-mlir compiler across all layers: TTIR/TTNN dialect definitions, StableHLO composite conversion, TTIR-to-TTNN conversion, EmitC/EmitPy conversions…

    314 GitHub stars~11k tokensUpdated yesterday
    Auto-check passed
  • Add Ttir Builder Op

    tenstorrent/tt-mlir

    Add full builder API support (@tag, @parse, @split) for a TTIR op.

    314 GitHub stars~1.6k tokensUpdated yesterday
    Auto-check passed
  • Run Ops Mlir Snippets

    tenstorrent/tt-mlir

    Compile and optionally execute every func.func in an ops.mlir-style snippet file (or every .mlir file in a directory) using runopsmlirsnippets.py.

    314 GitHub stars~1.8k tokensUpdated yesterday
    Auto-check passed
  • Uplift Ttsim CI

    tenstorrent/tt-mlir

    Uplift the TTSim version used by tt-mlir CI and refresh WH/BH simulator skips.

    314 GitHub stars~3.1k tokensUpdated yesterday
    Auto-check passed
  • Validate a tt-mlir PR against tt-xla by creating a cherry-picked branch and triggering CI.

    314 GitHub stars~4.5k tokensUpdated yesterday
    Auto-check passed
  • Triage Tt Metal Asserts

    tenstorrent/tt-mlir

    Triage a tt-metal uplift diff or digest of TTFATAL validation changes against what tt-mlir guarantees at each optimization level (0: workarounds only, 1: optimizer with DRAM-only fallback, 2: L1…

    314 GitHub stars~7.8k tokensUpdated yesterday
    Auto-check passed

Works with

Questions about Ttir Decomposition For Ttmetal

What does Ttir Decomposition For Ttmetal do?

Add a new composite op decomposition pattern to the TTMetal pipeline. Ttir Decomposition For Ttmetal is an agent skill from tenstorrent/tt-mlir. Add a new composite op decomposition pattern to the TTMetal pipeline.

When should I use Ttir Decomposition For Ttmetal?

Ttir Decomposition For Ttmetal fits situations like: the user wants to decompose/lower a high-level TTIR op (e.g; the user mentions decomposition pattern; decompose op for ttmetal; lower op to primitives.

How do I install Ttir Decomposition For Ttmetal in Claude Code?

Run `npx skills add tenstorrent/tt-mlir --skill ttir-decomposition-for-ttmetal -a claude-code`. Or copy the skill folder (.claude/skills/ttir-decomposition-for-ttmetal in tenstorrent/tt-mlir) into .claude/skills/ttir-decomposition-for-ttmetal in your project. Claude Code loads it when a task matches its description.

How do I install Ttir Decomposition For Ttmetal in Codex?

Run `npx skills add tenstorrent/tt-mlir --skill ttir-decomposition-for-ttmetal -a codex`. Or copy the skill folder (.claude/skills/ttir-decomposition-for-ttmetal in tenstorrent/tt-mlir) into .agents/skills/ttir-decomposition-for-ttmetal in your project. Codex loads it when a task matches its description.

Can I use Ttir Decomposition For Ttmetal in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add tenstorrent/tt-mlir --skill ttir-decomposition-for-ttmetal -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ttir-decomposition-for-ttmetal, .gemini/skills/ttir-decomposition-for-ttmetal, .github/skills/ttir-decomposition-for-ttmetal and .opencode/skills/ttir-decomposition-for-ttmetal in your project.

What does Ttir Decomposition For Ttmetal need to run?

Going by SKILL.md and its folder, Ttir Decomposition For Ttmetal needs the command-line tools its instructions call (cmake and pytest). Our summary lists: Python 3.

Does Ttir Decomposition For Ttmetal access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Ttir Decomposition For Ttmetal safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Ttir Decomposition For Ttmetal use?

Ttir Decomposition For Ttmetal is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Ttir Decomposition For Ttmetal use?

About 2.2k tokens (SKILL.md is roughly 8.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Ttir Decomposition For Ttmetal?

Skills that share tags, products or a category with Ttir Decomposition For Ttmetal: Paddle Build (PaddlePaddle/Paddle, 24k stars), Fory Release (apache/fory, 4.6k stars), CodeQL Security Scan (trailofbits/skills, 7.5k stars) and pybind11 Release Preparation (pybind/pybind11, 18k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Ttir Decomposition For Ttmetal?

tenstorrent (a GitHub organization) maintains it in tenstorrent/tt-mlir, which has 314 GitHub stars. The repository holds 9 skills in this directory. The repository was last updated on October 10, 2026.

Source: tenstorrent/tt-mlir on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.