Point attention at its own sequence and every token queries every token, itself included. What comes out is the attention map: one row per token doing the attending, one column per token being attended to, and every row summing to because focus is a fixed budget.
The sequence arrives as token vectors stacked as the rows of , each of width . Two learned matrices, and , both of shape , turn every token into a query and a key:
with the softmax taken across each row.
Task: write self_attention_weights(tokens, w_query, w_key) returning the weight matrix, every entry rounded to 4 decimal places.
[i][j] is how much of token i's attention lands on token j.[[1.0]].and come from the same but through two different matrices, and that is exactly what lets the map be lopsided: token 3 can attend heavily to token 1 while token 1 all but ignores token 3. One shared matrix would force the grid to be symmetric and throw that freedom away.