Unrolling a plain RNN draws one copy of the same cell per time step. Each copy receives two things: its own token, and the hidden state handed over by the copy before it. Every copy uses the same weights.
Here the hidden state has units, and each token has already been turned into a single number. At step the cell reads the token and the previous hidden state :
Task: write rnn_forward(tokens, W_h, W_x, b, h0), returning the hidden state after every step — the whole unrolled chain — with each number rounded to 4 decimal places.
| argument | shape | meaning |
|---|---|---|
tokens | list of numbers | one token per time step, in order |
W_h | list of lists | W_h[r][c] is the weight from unit of the previous state into unit of the new one |
W_x | list of numbers | W_x[r] multiplies the token on its way into unit |
b | list of numbers | one bias per unit |
h0 | list of numbers | the hidden state before any token has been read |
[[...h_1...], [...h_2...], ...]. h0 is not part of the output.W_h, W_x and b are used at every step. Nothing in this function depends on the position.Two of the test cases use identical weights and the same two tokens in the opposite order — and they disagree. That asymmetry is the whole point of recurrence: what a step remembers depends on what was read before it, so a sequence is not a bag of tokens.