Embodied AI Glossary中文

Rotary Position Embedding

旋转位置编码RoPEAdvanced

A positional encoding that writes position as a rotation angle applied to a vector, so attention naturally senses relative distance.

Rotary Position Embedding was proposed by Jianlin Su and colleagues at Shenzhen's Zhuiyi Technology in the 2021 RoFormer paper. A Transformer has no inherent sense of token order, so it needs a positional encoding to supply one. RoPE treats every two dimensions of the query and key vectors as a 2D plane and rotates them by an angle set by the token's position, with different dimension-pairs rotating at different rates; as a result, when two tokens' vectors are dotted together, the result depends only on their relative distance. It adds no learnable parameters, attention naturally decays as distance grows, and context length can be extended by adjusting the rotation frequencies. After Meta's LLaMA adopted RoPE, it became the mainstream choice for large language models, and VLMs in the Qwen series, along with VLAs built on them, use it too. Qwen2-VL further introduces multimodal RoPE (M-RoPE), splitting position into three components — time, height, and width — for image and video tokens.

ExampleQwen2-VL's M-RoPE: a text token's three position indices are all the same, equivalent to ordinary RoPE; an image token's time index is fixed while its height and width indices vary with its position in the grid; each subsequent video frame increments the time index by one.

Also called
RoPE, M-RoPE
Related
Positional Encoding · Transformer · Self-Attention · Context Length · Qwen-VL · Llama
Sources
RoFormer: Enhanced Transformer with Rotary Position Embedding (arXiv 2104.09864)
Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution (arXiv 2409.12191)
LLaMA: Open and Efficient Foundation Language Models (arXiv 2302.13971)

See it in the full glossary →