Inside every transformer block, right after attention, sits the feed-forward sublayer. It takes the vector at one position, widens it, applies a non-linearity, and narrows it back to its original width:
If each position's vector holds numbers and the wide middle holds :
Written out, hidden value is and output value is .
The same weights are used at every position, and each position is handled on its own. No position's output depends on any other position.
Task: write feed_forward(x, w1, b1, w2, b2).
x is a sequence: a list of position vectors, each a list of numbers.w1 and w2 are lists of rows; b1 and b2 are plain lists.