Critical Rules for cuTile Python → Julia Conversion
1-based indexing everywhere:
ct.bid,ct.num_tilesaxis,dimsfor reductions,permutedimsaxes,ct.extractindices,ct.cataxis — ALL shifted +1 from Python.forloops work in kernels (cuTile 0.2+):for k in Int32(1):nandfor k in Int32(0):n - Int32(1)are fully supported. Step ranges also work:for i in start:step:stop. Thewhilepattern still works butforis preferred for simple iteration.Explicit broadcasting: Python cuTile auto-broadcasts
+,-,*,/between different shapes. Julia requires.+,.-,.*,./for shape-mismatched tiles. Same-shape+/-and scalar*//work without dots.Left-aligned broadcasting: Julia broadcasts from dimension 1 (left), Python/NumPy from last dimension (right). A
(N,)tile cannot broadcast with(M, N). Usereshape(a, (1, N))first.Constants at launch, not signature: Python annotates
param: ct.Constant[int]in kernel signature. Julia uses plainparam::Intin signature and wraps withct.Constant(val)at thect.launchcall site.Kernel must return nothing: Every Julia cuTile kernel must end with
returnorreturn nothing.Column-major memory layout: Julia arrays are column-major. For multi-dimensional data that was row-major in Python, consider transposing the logical layout or using batch-last ordering (e.g.,
(M, K, Batch)instead of Python's(Batch, M, K)).Reduction keeps dims:
sum(tile; dims=2)produces(M, 1)not(M,). Usedropdims(result; dims=2)to remove the singleton.Type names:
ct.float32→Float32,ct.float16→Float16,ct.int32→Int32,ct.bfloat16→BFloat16,ct.tfloat32→ct.TFloat32.Integer types in loops: Loop counters and increments must have matching types. Use
Int32consistently. Preferred:for k in Int32(1):n(handles types automatically). Thewhilepattern also works:k = Int32(1); while k <= n; ...; k += Int32(1); end.ct.launcharg order is positional: Kernel args after the grid inct.launch(kernel, grid, arg1, arg2, ...)map 1:1 to the kernel's parameter list. If the kernel signature is(output, input, ...), you MUST passoutputfirst. Swapping arguments silently produces wrong results (the kernel reads from the output buffer and writes to the input buffer).Element-wise
max/minbetween tiles: Usemax.(a, b)(broadcast syntax), NOTmax(a, b). The non-broadcastmax(a, b)on two tiles is not supported in kernel IR and will fail withIRError: Unsupported function call: max. Similarlymin(a, b)→min.(a, b).IRStructurizer / compiler errors should be reported: If you encounter
IRError,MethodErrormentioningIRStructurizer.BlockArg, or other internal compiler errors, these are bugs in the cuTile.jl compiler pipeline — do not work around them. Write a minimal reproducer and file it upstream.Tile-size limits for
ct.load: TMA-basedct.loadhas hardware limits on how much data can be loaded at once (~16K elements). For large tensors, use chunked or online algorithms that iterate over the data in fixed-size tiles, using eitherct.load/ct.storewith column indices orct.gather/ct.scatterwith index tiles.ct.Constantparameters work as shape arguments: The shape tuple inct.arange(N)andfill(val, (N,))can usect.Constantkernel parameters — cuTile.jl's const-seeded inference pipeline resolves them at compile time. Pass tile sizes asct.Constant(val)at thect.launchcall site and use the corresponding::Intparameter directly in shape tuples. No@evalmetaprogramming needed.function my_kernel(output::ct.TileArray{T, 2}, input::ct.TileArray{T, 2}, TILE_SIZE::Int) where {T} ct.@compiler_options occupancy=2 bid = ct.bid(1) tile = ct.load(input; index=(bid, Int32(1)), shape=(1, TILE_SIZE)) # TILE_SIZE from ct.Constant # ... end ct.launch(my_kernel, grid, output_cu, input_cu, ct.Constant(tile_size))ct.loadorderparameter remaps BOTH shape AND index positions: When usingorder=(2,1,...), theorderdefines a logical-to-physical dimension mapping that applies to both the shape tuple and the index tuple. Iforder=(2,1,3,4), then index position 0 → physical array dim 1, index position 1 → physical array dim 0. You must place tile iterators at the index position that maps to the correct physical dimension.# Array K_jl has physical dimensions (D, S, H, B) # We want: tile TILE_D from D (all of it), tile TILE_N from S (iterate with j) # order=(2,1,3,4) maps: position 0 → physical dim 1 (S), position 1 → physical dim 0 (D) # ✅ CORRECT: j at position 0 (maps to S), 1 at position 1 (maps to D) ct.load(K_jl, (j, 1, head_idx, batch_idx), (TILE_N, TILE_D, 1, 1); order=(2,1,3,4)) # ❌ WRONG: j at position 1 (maps to D!), 1 at position 0 (maps to S — always tile 1!) ct.load(K_jl, (1, j, head_idx, batch_idx), (TILE_N, TILE_D, 1, 1); order=(2,1,3,4))Symptom: First tile (j=1) produces correct results, subsequent tiles read wrong data (zeros from out-of-bounds D, or stale data from always reading the same S tile). Errors grow with loop iteration count.
rsqrtusage: cuTile.jl exportsrsqrt, sorsqrt.(tile)works via broadcast dot syntax.map(ct.rsqrt, tile)also works. For other math functions (exp,log,sqrt,sin,cos,abs), the broadcast dot syntax works fine (e.g.,exp.(tile)) because these are inBase.rsqrtis NOT inBasebut IS exported by cuTile.jl.