The decoder is producing one output word, and it has one query vector to spend. In front of it sits the whole input sentence, already projected: a key and a value for every position.
Four stages, in the order the chapter built them:
| stage | what happens |
|---|---|
| score | dot product of the query with each key |
| scale | divide every score by , where is the length of the query |
| softmax | turn the scores into weights that are positive and sum to |
| blend | multiply each value by its weight and add them up |
Task: write attend(query, keys, values) returning the context vector as a plain list of numbers, every entry rounded to 4 decimal places.
keys[i] and values[i] are the key and the value of the same position. There is always at least one position.Keys and values are kept separate for a reason. The keys decide how much of each position you take; the values decide what you take. Blend the keys by mistake and you still get a vector of the right shape — just one that means nothing.