A plain RNN with a one-number hidden state reads tokens and makes a single prediction at the end.
The forward pass, for :
is a fixed starting value, handed to the network rather than learned. The prediction is the final hidden state , and the loss is its squared error against a target :
Only three numbers are learned — , and — and the same three are used at every one of the steps.
Task: write bptt_grads(tokens, w, u, b, h0, target) returning [dL_dw, dL_du], the gradient of the loss with respect to and with respect to , in that order, rounded to 4 decimal places.
You will need the derivative of the activation,
which, because , is simply at step .
tokens is a list of numbers, with .Nothing in this is a new algorithm. It is the chain rule, run along the unrolled chain instead of down a stack of layers. What it buys you is a direct look at the chapter's central claim: feed it a long sequence with well below and watch how little of the loss's influence is left by the time it reaches the first token.