Chapter 4 gave a transformer its sense of order by adding a position vector to each token's vector. Many recent decoder models do it differently. They leave the token vectors alone and instead rotate each query and key vector, inside attention, by angles that depend on its position. This is called rotary position embedding.
Take a vector of even length at position . Group its slots into adjacent pairs: Pair , for , is turned anticlockwise by the angle (in radians)
which means
Pair turns by a full radian for every step in position, and each later pair turns more slowly. These are the waves of different lengths from the positional encoding lesson, used as rotation speeds. A vector at position is not turned at all.
Task: write rotate_positions(vectors, start, base). The list vectors holds query (or key) vectors for consecutive tokens, all of the same even length, and vectors[k] sits at position start + k. Return the rotated vectors, with every number rounded to 4 decimal places.
start matters during generation: if the model has already processed 500 tokens and now computes the vector for one new token, that token sits at position 500, not at 0.
Why rotate? A rotation never changes a vector's length, and once the query at position and the key at position have both been rotated, their dot product depends only on the gap . The attention score sees how far apart two tokens are, not where in the text they happen to sit.