In the past few years, sequence modeling has been almost dominated by Transformers. It relies on self attention to allow each position to directly "see" all other positions, which has excellent results, but the cost is that the computational complexity increases with the square of the sequence length - doubling the sequence results in four times the cost, making long texts expensive and slow. Is there a path that preserves long-term memory while also linearly expanding? The State Space Model (SSM) provides an answer.
The state space model originated from control theory: a system maintains a "state" at any given time, and as new inputs arrive, the state updates according to fixed rules and generates outputs based on it. Applying it to a sequence is equivalent to continuously evolving a hidden state on the timeline, compressing historical information into a fixed size vector, and gradually "spitting out" the results. Due to the recursive nature of the update rules, the computational complexity increases linearly with the length of the sequence.
Early SSM had an awkward situation: the update rules were fixed, and the model could not choose "remember what" or "ignore what" based on the current content, resulting in mediocre performance in tasks such as copying and retrieving words that required precise recall. The key modification of Mamba is to make the parameters change with the input - that is, the "selective" mechanism: dynamically determining how much old information to retain and how much new information to write in the state based on the current token. This gives the model a "content perception" ability similar to attention.
Another challenge is efficiency. The selective mechanism disrupted the structure that could have been quickly computed using convolution, so Mamba adopted a hardware aware parallel scanning algorithm to reorganize the original serial recursion on the GPU, fully utilizing the memory hierarchy, and truly improving speed. This is also the reason why it can achieve near linear extension on long sequences.
Its advantage lies in the fact that every time a new token is generated during inference, only a fixed size of state needs to be updated, unlike attention that requires constantly growing KV Cache. Therefore, in long context and streaming generation scenarios, both memory and latency are more friendly. Mamba and its subsequent work, such as the hybrid architecture that combines attention with SSM, have demonstrated competitiveness in ultra long sequence tasks such as language, audio, and genomics.
It needs to be objectively viewed that SSM is not a silver bullet that "replaces Transformer". Pure SSM still has shortcomings in tasks that require precise retrieval and strong correlation reasoning, so the more common approach in the industry is a hybrid architecture: using a small number of attention layers to supplement precise recall, and using SSM layers for efficiency. The fusion of two routes may be the more realistic answer for long sequence modeling.
In summary, Transformer uses attention to "see the big picture", while SSM uses state to "keep track of the flow" - the former is precise but expensive, while the latter is efficient and good at long-term operations. How to choose and combine is one of the most active directions in current sequence modeling.
[Reference source] Comprehensive compilation of state space models and Mamba related research literature and publicly available technical materials.