PyTorch Dispatch Mechanism
This document explains how a Python-level PyTorch call routes through the dispatcher to reach a backend-specific C++ or CUDA implementation.
Note: All file paths in this document are relative to the PyTorch source checkout in the skill cache. The default path is ~/.cache/tilegym/pytorch-source, unless the user chooses another cache directory. All searches must stay within that checkout.
Overview
When you call torch.relu(x) or x.matmul(y), PyTorch doesn't directly call a single function. Instead, it goes through a dispatcher that selects the correct backend implementation based on:
- The operation being called
- The tensor's device (CPU, CUDA, MPS, etc.)
- Whether autograd is tracking gradients
- Other dispatch keys (Batched for vmap, FuncTorchGradWrapper, etc.)
The Dispatch Pipeline
Python call: torch.add(a, b)
│
▼
Python binding (generated or manual)
│
▼
torch::Dispatcher::call() ← C++ dispatcher
│
▼
DispatchKeySet resolution ← examines tensor DispatchKeys
│
▼
Dispatch key chain:
Autograd → ... → CPU/CUDA ← walks through keys in priority order
│
▼
Backend kernel (e.g., at::native::add_cpu or at::native::add_cuda)native_functions.yaml: The Master Registry
Every ATen operation is declared in aten/src/ATen/native/native_functions.yaml. This is the single source of truth for what operations exist and how they dispatch.
Entry Format
- func: operation_name.overload(Tensor self, Tensor other, ...) -> Tensor
variants: function, method # Available as torch.op() and/or tensor.op()
structured: True # Uses structured kernels (optional)
structured_delegate: op.out # Delegates to out= variant (optional)
dispatch:
CPU: op_cpu # CPU implementation function name
CUDA: op_cuda # CUDA implementation function name
SparseCPU: op_sparse_cpu # Sparse CPU variant (optional)
SparseCUDA: op_sparse_cuda # Sparse CUDA variant (optional)
MPS: op_mps # Apple Metal variant (optional)Reading a Real Entry
Example: the add operation:
- func: add.Tensor(Tensor self, Tensor other, *, Scalar alpha=1) -> Tensor
device_check: NoCheck
structured_delegate: add.out
variants: function, method
tags: [canonical, pointwise]This tells us:
- Signature:
add(self, other, alpha=1) -> Tensor - Variants: Available as both
torch.add(a, b)anda.add(b) - Delegation: Implementation delegates to the
add.outvariant - Tags: It's a pointwise operation
The add.out variant:
- func: add.out(Tensor self, Tensor other, *, Scalar alpha=1, Tensor(a!) out) -> Tensor(a!)
device_check: NoCheck
structured: True
dispatch:
CPU, CUDA: add_out
SparseCPU: add_out_sparse_cpu
SparseCUDA: add_out_sparse_cuda
SparseCsrCPU: add_out_sparse_csr_cpu
SparseCsrCUDA: add_out_sparse_csr_cuda
MkldnnCPU: mkldnn_add_out
MPS: add_out_mpsThis tells us:
- Both CPU and CUDA dispatch to a function named
add_out - Sparse tensors have separate implementations
- MKL-DNN has its own path on CPU
- MPS (Apple GPU) has its own path
Common Dispatch Patterns
Pattern 1: Shared implementation (same function for CPU and CUDA)
dispatch:
CPU, CUDA: my_op_impl # Single function handles both, uses TensorIteratorPattern 2: Separate backends
dispatch:
CPU: my_op_cpu
CUDA: my_op_cudaPattern 3: Structured delegation (most common for standard ops)
structured_delegate: my_op.out # Delegates to the out= variantPattern 4: No dispatch key (pure Python or composite)
# No dispatch: key means it's a CompositeImplicitAutograd op
# Implemented once, works on all backends, autograd handled automaticallyPattern 5: CompositeExplicitAutograd
dispatch:
CompositeExplicitAutograd: my_op_impl # Works on all backends, custom autogradDispatchKey System
DispatchKeys determine which implementation gets called. They form an ordered priority chain.
Key DispatchKeys (in priority order)
| DispatchKey | Purpose |
|---|---|
Autograd (AutogradCPU, AutogradCUDA) |
Records operations for backward pass |
Batched |
vmap batching transforms |
Functionalize |
Functional transforms |
ADInplaceOrView |
Tracks inplace/view ops for autograd |
BackendSelect |
Selects between backends when ambiguous |
CPU |
CPU implementation |
CUDA |
CUDA implementation |
MPS |
Apple Metal Performance Shaders |
SparseCPU / SparseCUDA |
Sparse tensor backends |
QuantizedCPU / QuantizedCUDA |
Quantized tensor backends |
How Dispatch Keys Are Resolved
- Each tensor has a
DispatchKeySetbased on its device, layout, and other properties - When an op is called, PyTorch computes the union of all input tensors' key sets
- The dispatcher walks through keys from highest to lowest priority
- The first key that has a registered kernel for this op is called
- That kernel may call
redispatchto continue to the next key
Autograd Dispatch
For most ops, the dispatch chain looks like:
AutogradCUDA → CUDA kernelThe AutogradCUDA wrapper:
- Saves tensors needed for backward
- Calls the actual CUDA kernel via
redispatch - Attaches a
grad_fnto the output tensor
Python-to-C++ Bridge Mechanisms
Mechanism 1: torch._C._VariableFunctions (via _VF)
Used primarily by nn.functional and some nn.Module implementations.
# In torch/_VF.py:
# _VF is a namespace that routes to torch._C._VariableFunctions
# In torch/nn/modules/rnn.py:
result = _VF.lstm(input, hx, self._flat_weights, ...)
# This calls torch._C._VariableFunctions.lstm()
# Which routes to the C++ dispatcherMechanism 2: torch._C direct calls
# In torch/nn/functional.py:
return torch._C._nn.linear(input, weight, bias)
# Direct call to generated C++ bindingMechanism 3: torch.ops namespace
# Calling ops by their registered name:
torch.ops.aten.add(a, b)
torch.ops.aten.mm(a, b)
# Routes through the C++ dispatcherMechanism 4: Python-defined ops (CompositeImplicitAutograd)
Some ops are implemented purely in Python and never touch C++:
# In torch/nn/functional.py:
def multi_head_attention_forward(...):
# Implemented in pure Python using other torch ops
q = linear(query, in_proj_weight, ...)
...Structured Kernels
Modern PyTorch ops use "structured kernels" — a pattern that reduces boilerplate:
- Meta function: Computes output shape and dtype without allocating memory
- Implementation function: Does the actual computation
- func: add.out(Tensor self, Tensor other, *, Scalar alpha=1, Tensor(a!) out) -> Tensor(a!)
structured: True
dispatch:
CPU, CUDA: add_outIn C++:
// Meta function (shape inference):
TORCH_META_FUNC(add) (const Tensor& self, const Tensor& other, const Scalar& alpha) {
// ... compute output shape, set output metadata
}
// CPU implementation:
TORCH_IMPL_FUNC(add_out_cpu) (const Tensor& self, const Tensor& other, ...) {
// ... actual CPU computation
}
// CUDA implementation:
TORCH_IMPL_FUNC(add_out_cuda) (const Tensor& self, const Tensor& other, ...) {
// ... actual CUDA computation
}Code Generation Pipeline
The code generation system converts YAML declarations into C++ dispatch code.
Input Files
aten/src/ATen/native/native_functions.yaml— Op declarationstools/autograd/derivatives.yaml— Autograd backward formulastools/autograd/templates/*.cpp— C++ template files
Generator Scripts
torchgen/gen.py— Main generator for dispatch codetools/autograd/gen_autograd.py— Generates autograd wrapperstools/autograd/gen_variable_type.py— Generates VariableType dispatch
Output (generated during build)
RegisterCPU.cpp,RegisterCUDA.cpp— Backend registrationsFunctions.h,NativeFunctions.h— Function declarationsVariableType_*.cpp— Autograd wrappers- Python bindings (pybind11 code)
Reading Generated Code
Since generated code only exists after building, you can understand it by:
- Reading the YAML entries for your op
- Reading the templates in
tools/autograd/templates/ - Understanding the pattern: YAML entry → template substitution → generated C++
Tracing an Op Through Dispatch: Quick Reference
Given an op name (e.g., lstm):
- Find in YAML:
grep -n "func:.*lstm" aten/src/ATen/native/native_functions.yaml - Read dispatch table: Look at the
dispatch:section - Find CPU impl: Search for the CPU function name in
aten/src/ATen/native/*.cpp - Find CUDA impl: Search for the CUDA function name in
aten/src/ATen/native/cuda/*.cuoraten/src/ATen/native/cudnn/*.cpp - Find autograd:
grep -n "lstm" tools/autograd/derivatives.yaml