All skills
nvidia avatar

/tilegym-monkey-patch-kernels-to-transformers

@8d74158
by NVIDIA Corporationnvidia/tilegym821 stars
89

Integrate TileGym kernels into Hugging Face `transformers` models by replacing the library's submodule(s) and certain class(es)' implementations, and patching certain class(es)' init/forward/load weight methods prior to instantiating models. Used when the user requires integrating TileGym kernels into `transformers` models.

  • 10 files
  • 1.1 MB
  • CC-BY-4
  • Updated 4 months ago
  • GitHub

Use this Skill: https://skilld.dev/gh/nvidia/tilegym/tilegym-monkey-patch-kernels-to-transformers

This session only. Nothing lands on disk.

referenceskernel-inventory-schema.md

≈961 tokens on demand. Your agent reads this file only when SKILL.md points to it.

Transformer Kernel Inventory Schema

Transformer-local kernels must use FlashInfer-Bench-style metadata so agents can inventory, compare, and reuse kernels across auto-kernelize runs.

Schema source of truth:

Directory layout

For each transformer module, reusable generated kernels live in:

src/tilegym/transformers/<submodule_name>/
|- kernel_definitions/
|  |- <kernel_name>.json
|- kernel_solutions/
|  |- <kernel_name>.json
|- kernels/
|  |- <kernel_name>.py
|- modeling_<submodule_name>.py

kernels/<kernel_name>.py contains reusable kernel implementation and thin wrapper code only. Model-specific monkey-patch glue, class replacement, patched forward methods, and checkpoint compatibility logic belong in modeling_<submodule_name>.py or src/tilegym/transformers/monkey_patch.py.

Definition requirements

Use strict FlashInfer Definition metadata:

  • name: include concrete problem information.
  • op_type: general compute category.
  • axes: symbolic const/var dimensions.
  • inputs and outputs: tensor specs with shape and dtype.
  • reference: PyTorch code containing a global run function.
  • tags: use namespaced tags such as model:<name>, stage:prefill, stage:decode, status:draft, status:verified, and fused.

Definitions describe math and interface. They do not describe implementation source files.

Reference provenance

reference is both an executable correctness contract and a provenance pointer. It must:

  • start with one or more # Source: <permalink> comments before any imports or code;
  • point to the precise upstream code region that implements the same compute pattern in transformers, using a GitHub-style permalink with line anchors;
  • point to the Hugging Face Hub model card or remote modeling_*.py code region when the model uses trust_remote_code=True;
  • include multiple # Source: comments for fused kernels whose Definition combines adjacent upstream operations;
  • keep a global run(...) function after the source comments, written in clear PyTorch and matching the Definition inputs and outputs.

Prefer immutable commit permalinks over branch links. The source comments should identify the upstream math or model callsite, not the generated cuTile implementation.

Solution requirements

Use FlashInfer Solution metadata with source paths:

  • name, definition, author, spec, and sources are required.
  • spec.language is cuda-tile for cuTile kernels.
  • spec.entry_point uses {file_path}::{function_name} and points at kernels/<kernel_name>.py.
  • spec.target_hardware lists supported CUDA compute capabilities, for example SM100.
  • sources.path references one or more repo-relative files containing the implementation.

For TileGym in-repo inventory, sources.content is not required. If an external FlashInfer-Bench submission needs embedded file content, materialize it from sources.path at export time.

Agent workflow rules

Explore subagents must return:

  • a list of compute requirements as draft Definition objects, including reference snippets with precise source comments;
  • an inventory of existing reusable kernels as Solution objects.

Candidate proposal must compare Definitions first:

  • exact Definition match: reuse the existing Solution;
  • compatible Definition with layout/signature gap: propose a small adapter;
  • no compatible Solution: create a new Definition, Solution, and dedicated kernel file.

Kept auto-kernelize experiments must check in the Definition, Solution, and kernel implementation. Discarded experiments keep draft metadata under sandbox/ only.

Source: SKILL.md on GitHub

No third-party reports yet.

Signed by skilld at 8d74158. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub 4 days ago.

Activeupdated 4 months ago
Other metadata
compatibility
Verified on Claude Code with Opus-4.6 and onward, CodeX with GPT-5.5 and onward, and Cursor (Agent mode) with GPT-5.3-CodeX and stronger models.
metadata
{
  "author": "TileGym Team <TileGym@nvidia.com>",
  "version": "2026.06.03",
  "tags": [
    "tilegym",
    "transformers",
    "integration",
    "kernel",
    "monkey-patch"
  ]
}

README badge

README badge for nvidia/tilegym/tilegym-monkey-patch-kernels-to-transformers