Operator / Kernel
算子AdvancedThe basic computational building block of a neural network — like matmul or softmax — and its hardware implementation.
An operator is the smallest unit of computation in a deep learning framework — for example, matrix multiplication, convolution, normalization, or softmax. A model's forward pass is a computation graph made of operators wired together. A “kernel” is the actual implementation code for a given operator on a specific piece of hardware, such as a CUDA kernel that runs on a GPU. How fast a model runs at inference depends heavily on how well these kernels are written, and on whether several small operators can be “fused” into a single kernel to cut down on memory reads and writes. This fusing, and picking the fastest kernel, is largely what TensorRT and torch.compile do; when engineers talk about “operator support” while porting a model to a Chinese domestic chip, they mean whether that hardware has a matching implementation for each operator the model uses.
ExampleFlashAttention fuses several steps inside attention — matrix multiply, softmax, then multiplying by the value matrix — into a single CUDA kernel, substantially cutting memory access and speeding up inference on long sequences.
- Also called
- kernel, CUDA kernel
- Related
- CUDA · FlashAttention · NVIDIA TensorRT · torch.compile · Compute Architecture for Neural Networks (Huawei Ascend, CANN) · CUDA Graphs
- Sources
- PyTorch 文档:Custom C++ and CUDA Operators (Chinese)