Embodied AI Glossary中文

U-Net

Common

A U-shaped convolutional network that downsamples layer by layer, then upsamples back, with matching layers connected directly.

U-Net is a convolutional network Ronneberger and colleagues proposed in 2015 for medical image segmentation. The left half is an encoder that downsamples layer by layer, shrinking resolution while extracting more abstract features; the right half is a decoder that upsamples layer by layer back to the original size; matching levels on the left and right are joined by skip connections, which pass the encoder's features directly to the decoder — drawn out, the shape looks like the letter U. This lets the network see global context without losing precise location, which suits tasks that take one image in and produce a same-sized result. It later became the standard backbone for the denoising network in diffusion models — Stable Diffusion and Stable Video Diffusion both use it — though the diffusion Transformer (DiT), introduced in 2022, has begun to replace it. Diffusion Policy in robotics also commonly uses a 1D temporal-convolution U-Net to denoise action sequences.

ExampleWhen Stable Diffusion generates an image, at every step it feeds the noisy latent into the U-Net, which predicts the noise in it; that noise is subtracted before the next step, and repeating this dozens of times produces a clean image.

Also called
UNet
Related
Convolutional Neural Network · Diffusion Model · Diffusion Transformer · Latent Diffusion Model · Diffusion Policy · Residual Network
Sources
U-Net: Convolutional Networks for Biomedical Image Segmentation (arXiv 1505.04597)
Scalable Diffusion Models with Transformers (arXiv 2212.09748)

See it in the full glossary →