Self-attention, add and normalize, feed-forward, add and normalize. Four parts, in that order, and that is the entire block — the thing GPT stacks ninety-six times without changing anything about it.
Task: write transformer_block(x, Wq, Wk, Wv, W1, b1, W2, b2) for a single masked attention head, returning the block's output, rounded to 4 decimal places. Both normalizations have their learned scale and shift at the neutral values and , so there is nothing extra to apply.
| argument | shape | what it is |
|---|---|---|
x | vectors of length | the tokens entering the block |
Wq, Wk | queries and keys | |
Wv | values, the same width as x | |
W1, b1 | , length | widen each position to the hidden width |
W2, b2 | , length | squeeze it back to |
1 — Masked self-attention. With , and ,
Row may use positions through only; everything to its right is blocked. Softmax each row over the positions it is allowed to see, then blend the values: .
2 — Add and normalize. .
3 — Feed-forward, each position on its own. , where keeps positive entries and sets negative ones to .
4 — Add and normalize again. The block returns .
works on each position's vector separately, using that vector's own slots:
Round only the numbers you return.
Two things are worth noticing once it runs. The output has exactly the shape of the input, which is the property that lets you stack this ninety-six times. And the second residual adds , not : every sublayer adds to whatever it was handed, so one running workspace travels all the way up the stack.