Back to Home

Introduction to Attention Mechanisms: From QKV to Self Attention, How the Core Engine of Transformer Works

September 25, 2026 at 08:02 AMSource: RunByAI0 comment(s)TechGuide

The attention mechanism is the core design that distinguishes Transformer from early recurrent networks and pushes large models to the present day. It can be said that without attention, there is no GPT. This article explains its ins and outs in the most straightforward way possible.

1、 Why do we need attention

When dealing with a sentence, the meaning of each word depends on the context. For example, "Apple has released a new phone" and "I ate an apple", the same word refers to completely different things. In the early days of using RNN and LSTM to process sequences, information had to be transmitted one by one along time steps, and it was easy to forget as long as it was far away, and it could not be parallelized. Attention has changed its approach: allowing each position to directly look at all positions in the sequence, weighted and summarized according to their degree of correlation, which not only captures long-distance dependencies but also enables parallel computation at once.

2、 What are Q, K, and V exactly

Pay attention to projecting each input vector into three roles: Query represents what the current word wants to ask; Key represents what labels each word can provide; Value represents the true content carried by each word. The most fitting analogy is search engines: you enter a query Q, and the system compares it with the keys of all web pages to calculate its relevance. Then, the system weights the values of each web page based on their relevance to form the answer.

3、 Zoom dot product attention

The standard calculation is divided into four steps: the first step is to dot product Q and all K to obtain the correlation score; Step 2, divide by the root d_k (where d_k is the vector dimension) to perform scaling, in order to prevent the dot product value from becoming too large and the softmax gradient from disappearing; Step three, use softmax to convert the score into a weight that adds up to 1; Step four, use these weights to weight and sum V to obtain the output. The formula states that Attention (Q, K, V) is equal to softmax (the transpose of QK divided by the root d_k) multiplied by V. Scaling is not an optional detail, it is one of the key factors for stable convergence in training.

4、 Self attention and multi head attention

When Q, K, and V all come from the same sequence, it is called self attention: each word in the sentence is looking at all the words in the same sentence, thus obtaining a representation that integrates the context. Multi head attention is the process of parallelizing the above steps multiple times, using different projection matrices for each head to focus on different relationship patterns - some heads focus on syntactic structures, some heads focus on referents, and finally concatenate the results of each head before projecting them. Combining multiple perspectives, the expressive power far exceeds that of a single head.

5、 Cost and Improvement

Attention also comes at a cost: the computational cost increases exponentially with the length of the sequence, making it difficult to handle long texts. So various optimizations emerged: sparse attention, sliding window attention only counts as local connections; FlashAttention has become the default choice for training and inference by reducing video memory usage and improving speed through partitioning and recalculation; Combined with KV Cache, historical keys and values are cached during inference to avoid duplicate calculations.

Summary

The core idea of attention mechanism is actually very simple: not to memorize the order, but to enable each position to retrieve global information as needed. The combination of QKV retrieval, scaling dot product, and multi head parallelism forms the foundation of all major models today. In the next article, we will talk about another inconspicuous yet essential component in Transformer - position encoding.

[Reference source] Comprehensive compilation of industry information released publicly (Transformer original paper Attention Is All You Need).

deep learningTransformerlarge model
Discussion

Comments (0)

No comments yet. Be the first!

Leave a Comment