Embodied AI Glossary中文

Mamba

Advanced

A sequence model that replaces attention with an input-dependent state space model, scaling linearly with sequence length.

Mamba was proposed by Albert Gu at Carnegie Mellon and Tri Dao at Princeton in December 2023. A Transformer's self-attention scales with the square of sequence length, which gets expensive for long sequences; a state space model (SSM) instead absorbs input step by step into a fixed-size hidden state, like a recurrent network, scaling linearly — but earlier SSMs had fixed parameters and couldn't decide what to remember or forget based on content. Mamba makes those parameters depend on the current input (making it 'selective'), paired with a GPU-oriented parallel scan algorithm, and the whole architecture uses no attention and no separate MLP block. The paper reports about 5x the inference throughput of a Transformer, with Mamba-3B matching the performance of a Transformer twice its size; in 2024 the same two authors released Mamba-2, 2-8x faster still. In robotics, work such as RoboMamba uses it to cut a VLA's inference latency.

ExampleRoboMamba builds a VLA on a Mamba language-model backbone, learns manipulation skills while fine-tuning only about 0.1% of its parameters, and the paper reports 3x the inference speed of VLA models that existed at the time.

Also called
Selective State Space Model, Selective SSM, Mamba-2
Related
State Space Model · Transformer · Recurrent Neural Network · RoboMamba · Inference Latency · Context Length
Sources
Mamba: Linear-Time Sequence Modeling with Selective State Spaces (arXiv:2312.00752)
Transformers are SSMs (Mamba-2, arXiv:2405.21060)
RoboMamba: Efficient Vision-Language-Action Model for Robotic Reasoning and Manipulation (arXiv:2406.04339)
As of
2024-05

See it in the full glossary →