Every sublayer in a transformer block hands its result to the same two-step wrapper: add, then normalize. The sublayer's output is added back onto the vectors that went into it — nothing is replaced — and the sum is rescaled so the numbers cannot drift as the stack gets deeper.
Task: write add_and_norm(x, sublayer_out, gamma, beta), returning the wrapped output, rounded to 4 decimal places.
| argument | shape | what it is |
|---|---|---|
x | vectors of length | what went into the sublayer |
sublayer_out | vectors of length | what the sublayer produced |
gamma | length | the learned scale, shared by every position |
beta | length | the learned shift, shared by every position |
Take one position at a time. Write for sublayer_out[i], and work from that position's own slots:
This is the "add & norm" that appears twice in every transformer block. The add is what gives the original vectors a clear path through all ninety-six layers instead of being shrunk away; the norm is what stops a harmless-looking drift of a third per layer compounding into numbers no optimizer can work with.