All skills
nvidia avatar

/tilegym-converting-cutile-triton-to-cutile-rs

@48a26cf
by NVIDIA Corporationnvidia/tilegym821 stars
89

Use this skill to convert, port, or translate Triton-TileIR or cuTile-Python GPU kernels to cutile-rs (Rust). The orchestrator runs scripts/preflight.sh, then drives a bounded Agent A -> B -> D -> E pipeline (Agent C is diagnostic, Agent F optional), delegating all kernel/host/correctness/perf work to sub-agents and routing by each stage's single-line VERDICT.

  • 44 files
  • 452.9 KB
  • CC-BY-4
  • Updated 2 months ago
  • GitHub

Use this Skill: https://skilld.dev/gh/nvidia/tilegym/tilegym-converting-cutile-triton-to-cutile-rs

This session only. Nothing lands on disk.

referencesir-diff-checklist.md

≈678 tokens on demand. Your agent reads this file only when SKILL.md points to it.

IR Diff Checklist

When comparing original cudatile IR with cutile-rs generated IR, every item must be checked and resolved.

Pre-check: Correct kernel variant?

  • Confirmed actual kernel name via nsys/torch.profiler (NOT assumed from source code)
  • Confirmed launch config (grid, block) matches between dump and runtime
  • Matched dump file to confirmed kernel name (CUDA_TILE_DUMP_TILEIR dumps ALL variants)

MUST match (correctness + performance critical)

  • optimization_hints on entry — num_cta_in_cga, occupancy
  • strides=[?,1] — last dim must be constant 1. For &Tensor params: auto. For raw ptr: use Array::<{[-1, 1]}> { dims: &[only_dynamic_val] }
  • accumulator type — f32 for MMA, not f16
  • ftof conversion — f32→f16 before store if accumulator is f32
  • mmaf op types — input f16, accumulator f32
  • token count — match original (usually 1 shared make_token)
  • for loop step — matches original step size
  • partition tile shapes — (BM x BK), (BK x BN), (BM x BN) match
  • padding_value — neg_inf or zero where original has it (now supported since commit 4ba3a83)
  • maxf/minf — should NOT have rounding_mode attribute (exact ops)
  • grid size — persistent kernel grid must match (e.g., NUM_SM * occupancy)
  • block size / num_warps — cutile-rs doesn't support simt_num_warps_in_cta, compiler chooses automatically

OK to differ (semantic equivalent)

  • divi rounding<positive_inf> → (a + b - 1) / b — same semantics
  • mini signed → if/else — same semantics
  • get_index_space_shape → manual (dim + tile - 1) / tile — same semantics
  • negative remi correction (xori+andi+select) → Rust % — may differ for negative values but blockId is always non-negative
  • for ... step %assumed_var vs for ... step %raw_var — cutile-rs uses assume-wrapped variables

Expected noise (ignorable)

  • Dead constant <i32: 1> / constant <i32: N> — from macro expansion, use --canonicalize to remove
  • Extra assume bounded<0, ?> on blockId — auto-generated by cutile-rs
  • Extra assume bounded<0, N> on for loop index — auto-generated
  • Extra assume bounded<0, ?> on gridSize — auto-generated

Known gaps (cannot fix, record)

  • make_strided_view — not supported in cutile-rs; use tview.partition(...) / tview.partition_mut(...) (lowers to make_partition_view in IR)
  • tensor_dim_factor hint — may not appear in IR
  • Debug info / source locations — cutile-rs doesn't generate
  • simt_num_warps_in_cta — not supported as optimization hint

Source: SKILL.md on GitHub

No third-party reports yet.

Signed by skilld at 48a26cf. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub 4 days ago.

Activeupdated 2 months ago
metadata
{
  "author": "TileGym Team <TileGym@nvidia.com>"
}

README badge

README badge for nvidia/tilegym/tilegym-converting-cutile-triton-to-cutile-rs