A translation transformer is being trained on one pair of sentences.
| stack | sequence | length |
|---|---|---|
| encoder | the French input | tokens |
| decoder | the English target | tokens |
Training hands the decoder the whole English target at once. Each decoder block runs masked self-attention over the English tokens, then cross-attention into the encoder's outputs.
Now look at one position inside one decoder block: the row belonging to English token number , counting from .
Across that position's two attention sublayers combined, how many of its attention weights are non-zero?
Assume single-head attention in both sublayers, and take every unmasked weight to be non-zero — softmax never returns exactly zero for a finite score.
Select all that apply.