Back to Home

Introduction to Position Encoding: From Sine Waves to RoPE, How Transformer Knows the Order of Words

September 25, 2026 at 08:02 AMSource: RunByAI0 comment(s)TechGuide

The attention mechanism has a natural flaw: it is insensitive to input order. By shuffling the words in a sentence, the result calculated from self attention remains almost unchanged, but "cat chasing dog" and "dog chasing cat" are obviously two different things. So it is necessary to additionally tell the model who is in front and who is behind - this is positional encoding.

1、 Why can't attention see the order

Self attention is essentially a weighted sum that treats all positions equally, and whoever comes first does not affect the computational structure. This brings the benefits of parallel computing, but also means that sequential information must be injected externally.

2、 Learning style and sine style

Early Transformers used a set of deterministic sine and cosine functions to generate position vectors, with different dimensions corresponding to different frequencies, which were then directly added to word vectors. It does not require training and can be extrapolated to longer sequences than during training. Another approach is learning based position embedding, which treats each position as a trainable parameter, but it is limited by the maximum length during training, beyond which it cannot take a value. Both of them add position vectors to word vectors, which is simple and effective, but can start to struggle in the face of long sequences and strong extrapolation requirements.

3、 Relative position and RoPE

People gradually realize that models are often concerned with the relative distance between words, rather than absolute numbering, thus developing relative positional encoding. The most influential one among them is the Rotary Position Embedding (RoPE). Its clever idea lies in not adding an additional position vector, but directly rotating the Query and Key vectors, with the rotation angle related to the position number. In this way, when two vectors do dot product, the result naturally depends only on their relative distance - the relative position information is encoded into the attention score itself. RoPE is compatible with efficient implementations such as linear attention and has good extrapolation capabilities. It has now become a standard feature for mainstream large models such as LLaMA and Qwen.

4、 Extrapolation of Location in Long Context Era

When the context window expands from 4K to 128K or even longer, position encoding becomes a bottleneck again. Common practices include position interpolation, which compresses positions beyond the training length back to the original range; And NTK aware scaling, YaRN, etc., make finer adjustments in the frequency dimension to preserve local discrimination as much as possible. Their goal is the same: to make the model recognize longer text at a lower cost.

Summary

Position encoding solves a seemingly small but actually fatal problem: allowing the attention mechanism of parallel computing to see the order again. From sine functions to RoPE, and now to today's long context extrapolation, the same goal lies behind this evolutionary path - to feed location information steadily into the model in a more efficient way.

[Reference source] Comprehensive compilation of industry information publicly released.

deep learningTransformerlarge model
Discussion

Comments (0)

No comments yet. Be the first!

Leave a Comment