A translation model is turning the French sentence
lechatnoirdort— 4 tokens
into English. The decoder has produced 3 tokens so far:
theblackcat
Two tables of attention weights are recorded at this moment.
| what it records | |
|---|---|
| Map A | the encoder's self-attention over the French sentence |
| Map B | the decoder's cross-attention — the weights each decoder position produced so far placed on the French positions, stacked into one table |
As always, a row of an attention map is one position doing the attending and a column is a position being attended to.
Which statement about the two maps is correct?
Select all that apply.