Chapter 3 scores a query against a key with a dot product. The first attention models for translation scored differently. They passed the decoder's current state and each encoder state through a small learned layer of their own, and read a single number out of the result. This is called additive attention.
The decoder state is doing the asking. The encoder states are the input words it can look at.
Score. Two learned matrices project the decoder state and each encoder state into the same number of slots, . The two projections are added slot by slot, each slot is squashed with , and a learned vector of length combines the slots into one number:
Because both projections land in slots, the decoder state and the encoder states do not have to be the same size. A dot product between them would demand that.
Weights. Softmax over the scores gives attention weights that are positive and sum to 1.
Context. The context vector is the weighted sum of the encoder states themselves:
Matrices are given as lists of rows: W[i] is row , so is sum(W[i][k] * x[k] for k in range(len(x))).
Task: write additive_attention(decoder_state, encoder_states, W_dec, W_enc, v). Return a tuple (weights, context): the list of attention weights, one per encoder state, and the context vector, the same size as one encoder state. Round every number to 4 decimal places, and build the context from the unrounded weights.