A self-attention layer projects each token into three vectors using three learned matrices
and then computes, for the query at position ,
An engineer refactoring the code swaps two lines. The tensor produced by is now handed to the layer as the values, and the tensor produced by is handed to it as the keys. The swap is consistent everywhere, both matrices produce vectors of the same width , and the layer is then trained from a fresh random initialisation on the usual task. So the refactored layer computes
What happens when this layer is trained?
Select all that apply.