State Space Model
状态空间模型SSMAdvancedA network that processes long sequences with a hidden state updated over time, with compute growing linearly with sequence length.
The state space model comes from control theory: a hidden state summarizes the past, and every new input updates that state and produces an output according to linear equations. In deep learning, an SSM discretizes those equations and uses them as a network layer, with the parameters learned through training. In 2021, Stanford's Gu, Goel, and Ré proposed S4 (ICLR 2022), which made SSMs work well on sequences tens of thousands of steps long for the first time; in 2023, Gu and Dao's Mamba made the parameters depend on the input itself (calling this 'selective'), letting the model decide what to remember or forget based on content. Compared with a Transformer, an SSM only needs to keep a fixed-size state at inference, so compute grows linearly with sequence length rather than quadratically like attention, which suits long history, high-frequency control, and on-device deployment. Don't confuse this with 'state space' in reinforcement learning, which means the set of all possible states.
ExampleThe Mamba paper reports about 5x the inference throughput of a similarly sized Transformer, with a 3B-parameter Mamba language model beating a Transformer of the same size and matching one twice as large; in embodied AI, RoboMamba uses a Mamba language model in place of a Transformer as the inference backbone for its VLA.
- Also called
- SSM, S4, Selective SSM
- Related
- Mamba · Transformer · Recurrent Neural Network · RoboMamba · Context Length · State Space
- Sources
- Efficiently Modeling Long Sequences with Structured State Spaces (S4, arXiv:2111.00396)
Mamba: Linear-Time Sequence Modeling with Selective State Spaces (arXiv:2312.00752)