IR Diff Checklist
When comparing original cudatile IR with cutile-rs generated IR, every item must be checked and resolved.
Pre-check: Correct kernel variant?
- Confirmed actual kernel name via nsys/torch.profiler (NOT assumed from source code)
- Confirmed launch config (grid, block) matches between dump and runtime
- Matched dump file to confirmed kernel name (CUDA_TILE_DUMP_TILEIR dumps ALL variants)
MUST match (correctness + performance critical)
-
optimization_hintson entry —num_cta_in_cga,occupancy -
strides=[?,1]— last dim must be constant 1. For&Tensorparams: auto. For raw ptr: useArray::<{[-1, 1]}> { dims: &[only_dynamic_val] } - accumulator type —
f32for MMA, notf16 -
ftofconversion — f32→f16 before store if accumulator is f32 -
mmafop types — input f16, accumulator f32 - token count — match original (usually 1 shared
make_token) - for loop step — matches original step size
- partition tile shapes —
(BM x BK),(BK x BN),(BM x BN)match -
padding_value—neg_inforzerowhere original has it (now supported since commit4ba3a83) -
maxf/minf— should NOT haverounding_modeattribute (exact ops) - grid size — persistent kernel grid must match (e.g.,
NUM_SM * occupancy) - block size / num_warps — cutile-rs doesn't support
simt_num_warps_in_cta, compiler chooses automatically
OK to differ (semantic equivalent)
-
divi rounding<positive_inf>→(a + b - 1) / b— same semantics -
mini signed→if/else— same semantics -
get_index_space_shape→ manual(dim + tile - 1) / tile— same semantics - negative
remicorrection (xori+andi+select) → Rust%— may differ for negative values but blockId is always non-negative -
for ... step %assumed_varvsfor ... step %raw_var— cutile-rs uses assume-wrapped variables
Expected noise (ignorable)
- Dead
constant <i32: 1>/constant <i32: N>— from macro expansion, use--canonicalizeto remove - Extra
assume bounded<0, ?>on blockId — auto-generated by cutile-rs - Extra
assume bounded<0, N>on for loop index — auto-generated - Extra
assume bounded<0, ?>on gridSize — auto-generated
Known gaps (cannot fix, record)
-
make_strided_view— not supported in cutile-rs; usetview.partition(...)/tview.partition_mut(...)(lowers tomake_partition_viewin IR) -
tensor_dim_factorhint — may not appear in IR - Debug info / source locations — cutile-rs doesn't generate
-
simt_num_warps_in_cta— not supported as optimization hint