All skills
nvidia avatar

/tilegym-converting-cutile-to-triton

@2bf003b
by NVIDIA Corporationnvidia/tilegym821 stars
89

Converts cuTile GPU kernels (@ct.kernel) to Triton (@triton.jit). Handles standard in-repo conversion, debugging (cudaErrorIllegalAddress, shape mismatch, numerical mismatch), and mapping cuTile idioms (ct.load/ct.store, ct.Constant, ct.launch) to Triton equivalents. Covers dual-kernel layout flags (e.g. transpose=True/False + autotune grid via META) per translations/advanced-patterns.md. Use when converting, porting, or translating cuTile kernels to Triton, or debugging existing Triton translations.

  • 25 files
  • 237.1 KB
  • CC-BY-4
  • Updated 4 months ago
  • GitHub

Use this Skill: https://skilld.dev/gh/nvidia/tilegym/tilegym-converting-cutile-to-triton

This session only. Nothing lands on disk.

translationsfile-structure.md

≈780 tokens on demand. Your agent reads this file only when SKILL.md points to it.

File Structure & Registration (cuTile → Triton)

Where to place Triton files when converting from cuTile. Inverse of the Triton→cuTile layout.


Directory Structure

Standard Mode (Directory-Based)

When converting from cuTile to Triton, create Triton files under the triton/ mirror of the cuTile path.

There are two top-level layouts depending on whether the op is a first-party TileGym op or an external-framework suite:

TileGym/
├── src/tilegym/
│   ├── ops/                        # First-party TileGym ops (fmha, matmul, softmax, …)
│   │   ├── triton/
│   │   │   ├── add.py              # Triton (target of c2t conversion)
│   │   │   ├── softmax.py
│   │   │   └── layer_norm.py
│   │   └── cutile/
│   │       ├── add.py              # Existing cuTile (source)
│   │       ├── softmax.py
│   │       └── layer_norm.py
│   └── suites/                     # External-framework suites
│       └── <framework>/            # e.g. unsloth, flashinfer
│           ├── triton/
│           │   └── <op>.py         # Triton conversion target
│           └── cutile/
│               └── <op>.py         # Existing cuTile source
└── tests/
    ├── ops/
    │   └── test_<op>.py            # Tests for ops/ kernels
    └── suites/
        └── <framework>/
            └── test_<op>.py        # Tests for suites/ kernels

Path derivation:

# ops/ kernel: swap /cutile/ → /triton/
CUTILE_PATH="src/tilegym/ops/cutile/softmax.py"
TRITON_PATH="${CUTILE_PATH//\/cutile\//\/triton\/}"
# → src/tilegym/ops/triton/softmax.py

# suites/ kernel: same rule
CUTILE_PATH="src/tilegym/suites/<framework>/cutile/<op>.py"
TRITON_PATH="${CUTILE_PATH//\/cutile\//\/triton\/}"
# → src/tilegym/suites/<framework>/triton/<op>.py

mkdir -p $(dirname $TRITON_PATH)

Registration Patterns

Same as the Triton→cuTile skill: register implementations by backend.

from tilegym.backend import register_impl

@register_impl("op_name", backend="triton")
def op_triton(...):
    ...

@register_impl("op_name", backend="cutile")
def op_cutile(...):
    ...

Tests typically parametrize over backend=["triton", "cutile"] so both are exercised.


Multi-Agent / Two-Step Workflow

Step Purpose
Step 1: Convert cuTile → Triton conversion
Step 2: Perf Performance testing & comparison (Triton vs cuTile)

Default: run both steps unless the user asks only for conversion or only for perf.


Common Pitfalls

  • Wrong path: Putting the new Triton file in cutile/ instead of triton/.
  • Leftover cuTile imports: Removing import cuda.tile as ct and all ct.* usage in the new Triton file; use import triton.language as tl and triton.jit only.
  • Launch style: Using ct.launch(stream, grid, kernel, args) in the Triton host; must use <code>kernel[grid](launch_args)</code> and triton.cdiv for grid.

Source: SKILL.md on GitHub

No third-party reports yet.

Signed by skilld at 2bf003b. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub 4 days ago.

Activeupdated 4 months ago
version
1.0.0
tools
[
  "Read",
  "Write",
  "Grep",
  "Glob",
  "Bash"
]
Other metadata
metadata
{
  "author": "TileGym Team <TileGym@nvidia.com>",
  "tags": [
    "cutile",
    "triton",
    "conversion",
    "gpu",
    "kernel"
  ]
}

README badge

README badge for nvidia/tilegym/tilegym-converting-cutile-to-triton