Embodied AI Glossary中文

LoRA

低秩适配Common

Freezing a large model's original weights and fine-tuning it by training only two small matrices inserted alongside them.

LoRA was introduced by Microsoft's Edward Hu and colleagues in 2021 and is the most widely used parameter-efficient fine-tuning method. It freezes the pretrained weight matrix W and adds a side path B·A next to selected layers (commonly the projection matrices inside a Transformer's attention blocks): A and B are two narrow matrices whose rank r is far smaller than the original matrix's dimensions, and only A and B are trained during fine-tuning. After training, B·A can be added straight back into W, adding no extra latency at inference. The paper reports roughly a 10,000-fold reduction in trainable parameters and about a third of the memory requirement compared with fully fine-tuning GPT-3 175B with Adam. For embodied AI, since most VLAs are built on multi-billion-parameter vision-language models, LoRA lets an ordinary lab adapt a model to its own robot and tasks on a single consumer GPU.

ExampleIn the OpenVLA paper, a rank-32 LoRA trains only about 1.4% of the parameters and reaches 68.2% success, close to full fine-tuning's 69.7%. The openpi repository lists π0 LoRA fine-tuning as needing 22.5 GB of memory or more, running on an RTX 4090, versus 70 GB or more for full fine-tuning.

Also called
Low-Rank Adaptation
Related
Parameter-Efficient Fine-Tuning · Full Fine-Tuning · Adapter · Fine-tuning · Backbone Freezing · OpenVLA
Sources
LoRA: Low-Rank Adaptation of Large Language Models (arXiv 2106.09685)
OpenVLA: An Open-Source Vision-Language-Action Model (arXiv 2406.09246)
openpi(Physical Intelligence 官方仓库,显存需求表) (Chinese)

See it in the full glossary →