Chapter 5 ended with a full training run: forward pass, loss, backward pass, update, batch after batch and epoch after epoch. Here you write that loop yourself, for the smallest network that still has a hidden layer, with nothing random in it so every run is reproducible.
The network. An input has numbers. A hidden layer of ReLU neurons feeds a single output with no activation:
In code these are W1[i][j] (the weight from input to hidden neuron ), b1[j], W2[j] (one number per hidden neuron) and b2 (a single number).
The ReLU's slope. When you send gradients back through a hidden neuron, its slope is where and where . That includes being exactly : there the slope is .
The loss. Mean squared error: the average of over whichever examples are being scored. For a batch of rows, each row's prediction receives ; every other gradient follows from the chain rule.
The run.
0 to batch_size - 1, then the next batch_size rows, and so on. If the rows don't divide evenly, the last batch is simply smaller.Task: write train(X, y, W1, b1, W2, b2, lr, batch_size, epochs). X is a list of rows and y a list of targets.
Return a list with one number per epoch: the mean squared error over the whole training set, measured with the weights as they stand at the end of that epoch, rounded to 4 decimal places.