The attention mechanism is a core component of modern large language models. It was first used for machine translation by Bahdanau et al. in 2014, and then in their 2017 paper "Attention Is All You Need", the Google team proposed a fully attention based Transformer architecture, laying the foundation for almost all mainstream large models today.
1、 Why do we need attention mechanisms
Before the emergence of attention mechanisms, processing sequence data mainly relied on recurrent neural networks (RNNs) and their variants LSTM and GRU. This type of model processes word elements one by one in chronological order, which has two obvious problems: firstly, it is difficult to perform parallel computing and the training speed is slow; Secondly, when the sequence is very long, early information is easily diluted during transmission, making long-distance dependencies difficult to capture.
The idea of attention mechanism is to allow the model to directly "see" all other word elements in the sequence when processing each word element, and dynamically determine which parts to focus on based on relevance, rather than compressing all information into a fixed length state.
2、 Core concepts: Query, Key, Value
The attention mechanism is usually described by three vectors:
-Query: What is the current keyword 'looking for';
-Key: an index of what each morpheme can provide;
-Value: The actual information carried by each morpheme.
The calculation process is roughly as follows: use Query to dot product with all keys to obtain the correlation score, scale and normalize to weights using Softmax, and then use these weights to weight and sum all values to obtain the output representation of the current word element.
3、 Scaling dot product attention and multi head attention
The standard Scaled Dot Product Attention first divides the dot product result by the root d_k (where d_k is the dimension of the Key) to prevent the Softmax gradient from disappearing due to excessively large values. The formula can be written as: Attention (Q, K, V)=softmax (QK ^ T/sqrt (d_k)) V.
A single attention can only capture one association pattern. Transformer introduces Multi Head Attention, which projects Query, Key, and Value into multiple subspaces for parallel computation, allowing the model to simultaneously focus on different levels of relationships such as syntax, semantics, and position. Finally, the results of each head are concatenated and linearly transformed again.
4、 Self attention and causal mask
When Query, Key, and Value all come from the same sequence, it is called Self Attention. In generative models such as the GPT series, in order to ensure that only the content before it can be seen when predicting the t-th morpheme, an upper triangular mask is applied to the attention score, setting the score of future positions to negative infinity. This mechanism is called a causal mask.
5、 Computational complexity and optimization
The computational complexity of self attention increases by the square of the sequence length (O (n ^ 2)), which is also the main bottleneck of long context processing. Various optimization schemes such as sparse attention, sliding window attention, and linear attention have been proposed in the industry to reduce complexity to approximately linear, thereby supporting longer contextual windows.
6、 Summary
The attention mechanism allows the model to break free from the limitation of fixed length states and dynamically allocate "attention" in the sequence, which is the key to the success of Transformers and large language models. Understanding the computational logic of Query, Key, Value, and multi head attention is the foundation for further learning models such as Transformer, BERT, GPT, etc.
[Reference source]
- Vaswani et al., "Attention Is All You Need", NeurIPS 2017(arXiv:1706.03762)
- Bahdanau et al., "Neural Machine Translation by Jointly Learning to Align and Translate", ICLR 2015(arXiv:1409.0473)
-Comprehensive compilation of industry information that has been publicly released