The encoder has read the French sentence and produced one vector per input word. The decoder is part-way through writing the English, with its own vector for each output position so far. Cross-attention is the only wire between the two stacks, and it is ordinary attention with one change: the query comes from the decoder, the keys and the values from the encoder.
Task: write cross_attention(decoder_states, encoder_outputs, Wq, Wk, Wv), returning one context vector per decoder position, rounded to 4 decimal places.
| argument | shape | what it is |
|---|---|---|
decoder_states | vectors of length | one per output position, written below |
encoder_outputs | vectors of length | one per input word, written below |
Wq | projects a decoder state into a query | |
Wk | projects an encoder output into a key | |
Wv | projects an encoder output into a value |
Project, score, scale, softmax, blend:
Softmax each decoder row across all encoder positions, giving weights that sum to for each , then
Return the context vectors, each of length . There is no mask anywhere in this operation: the encoder has already read the entire input, so every output position is free to draw on all of it.
Nothing here is a new mechanism. Point all three projections at a single sequence instead of two and this identical code is self-attention — which is why one block can serve translation, classification and generation alike.