A trained decoder-only transformer is run twice, deterministically, with the same weights both times. Its stack is the standard one from this chapter: token embeddings plus positional encoding, then 12 identical blocks, where each block is masked self-attention, add and norm, feed-forward, add and norm.
Take each word below as one token, so both prompts are 6 tokens long:
| position | 1 | 2 | 3 | 4 | 5 | 6 |
|---|---|---|---|---|---|---|
| run A | the | film | was | long | but | brilliant |
| run B | the | film | was | long | and | dull |
The two prompts diverge from position 5 onward.
At the top of the stack each run produces six output vectors, one per position. How many of those six are guaranteed to be identical across the two runs — whatever the model's weights happen to be?
Select all that apply.