One position's vector is passing through the first of a transformer block's two add & norm steps. The block's width is a toy .
| vector | value |
|---|---|
| what arrived at the block for this position, | |
| what the self-attention sublayer produced for this position, |
Layer normalization here is
with the mean and the spread taken across the four slots of the one vector (dividing by ). This layer's two learned numbers are a scale and a shift .
Which vector leaves this add & norm step?
Select all that apply.