The words' cat chasing mouse 'and' mouse chasing cat 'are exactly the same, but their meanings are opposite. To understand this difference, the big model must know where each word is located. But the core mechanism of Transformer - self attention, itself is' insensitive 'to order. Position coding was born to address this shortcoming.
1、 Why Attention Needs' Location '
Self attention calculates the correlation between any two positions, essentially a collective weighted sum of inputs. If no additional location information is injected and a sentence is scrambled, the result obtained will hardly change. This is completely unacceptable in language tasks.
2、 From absolute position to relative position
Sinusoidal position encoding: Generate a vector for each position using sin and cos functions of different frequencies, and directly add it to the word embedding. This is the approach adopted in the original Transformer paper.
Learnable positional encoding: treats the vectors of each position as trainable parameters, which is used by BERT and early GPT. The disadvantage is that the number of positions is fixed, making it difficult to handle beyond the training length.
Relative positional encoding: It does not care about "ranking in which position", but about "how far apart two words are", which is usually better in generalization.
3、 RoPE (Rotation Position Encoding)
RoPE's approach is clever: applying positional information in a "rotated" manner to the Query and Key vectors. The specific method is to pair the vectors pairwise and rotate them according to the angle of their positions; When two vectors are dot product, the result naturally carries their relative distance information. This preserves the form of absolute position while naturally expressing relative position.
Because of this property, RoPE has a better foundation in long text extrapolation and has become a common choice for mainstream open-source models such as Llama, Qwen, DeepSeek, etc. A number of length extrapolation improvement methods (such as position interpolation, frequency scaling, etc.) have also been derived around it, which are commonly used to expand the context window.
4、 Its relationship with the 'context window'
The position encoding determines how long the model has "seen" positions. If only a few thousand tokens are seen during training and pulled directly into a very long context, significant degradation often occurs. Therefore, long contextual ability typically requires specialized extrapolation techniques and fine-tuning of long texts.
5、 Inspiration for Users
When dealing with long documents, placing key information at the beginning or end is usually more secure than placing it in the middle.
The ability to search for a needle in a haystack is closely related to the position encoding scheme.
Understanding positional encoding helps explain why the model "forgets" the middle content in long texts.
Conclusion
Position encoding may seem like an engineering detail, but it determines the way big models perceive "order" and "distance". The evolution from sine coding to RoPE also reflects the continuous exploration of large models in pursuing longer context and stronger generalization ability.
[Reference source] Comprehensive compilation of industry information publicly released (RoPE and other public papers and technical documents).