In the previous article, we talked about how Convolutional Neural Networks (CNN) "understand" a graph. But in the real world, a lot of data is not static images, but a sequential sequence - a paragraph, a speech, a string of prices, a video. For this type of data, "order" itself is information: by shuffling words, the meaning may completely change.
Recurrent Neural Network (RNN) is a model designed for sequences. Its core idea is "memory": every time a new input is read in, the network not only looks at the current moment's information, but also references the hidden state preserved in the previous step (which can be understood as compressed memory of the "previous content"), and then updates this state before passing it on to the next step. So, the information passed down along the timeline like a relay baton.
This design brings two benefits. One is parameter sharing: regardless of the length of the sequence, the same set of weights is used to process each time step, which allows the model to handle inputs of any length and prevents the number of parameters from exploding with increasing sequence length. The second is context aware: because each step carries history, theoretically the model can synthesize the previous context to determine the current meaning. Early machine translation, speech recognition, text generation, and other tasks relied on RNNs and their variants for a long time.
But RNNs also have inherent weaknesses. When the sequence is very long, the gradient will continuously multiply during backpropagation, which can easily become extremely small (gradient disappearance) or extremely large (gradient explosion), causing the model to "forget" information from too long ago and making it difficult to train stably. This is the origin of the "long-distance dependency" problem.
To alleviate this issue, researchers have proposed improved structures with "gating" mechanisms, such as LSTM and GRU, which significantly enhance the ability to model long sequences by selectively remembering or forgetting information. Later on, Transformer, with its attention mechanism, achieved parallel computing and stronger long-range modeling capabilities, gradually replacing RNN as the mainstream. However, the idea of "sequential modeling" represented by RNN is not outdated, and understanding it is still an important part of understanding the evolution of sequential models.
In summary, RNN enabled neural networks to learn "sequential thinking" for the first time. Although it had the limitation of memory decay, it paved the way for later attention mechanisms and Transformers.
[Reference source] Comprehensive compilation of industry information and publicly available materials from research institutions.