This is the backward pass on a real, if tiny, network, from start to finish: run it forward, score it, then find a gradient for every weight and every bias in it.
The network. inputs feed a hidden layer of neurons with ReLU, which feeds an output layer of neurons with no activation. Both layers are dense and use a connection table: W[i][j] is the weight on the wire from unit i of the layer below to neuron j.
| parameter | size | holds |
|---|---|---|
W1 | W1[i][j]: input → hidden neuron | |
b1 | one bias per hidden neuron | |
W2 | W2[j][m]: hidden neuron → output | |
b2 | one bias per output |
On one example , the forward pass computes
The loss is the mean squared error over the whole batch. Square the error of every output on every example, add them up, and divide by how many there were:
ReLU's slope is where and where , including at exactly .
Task: write backward_pass(inputs, targets, W1, b1, W2, b2), returning
[loss, dW1, db1, dW2, db2]
where loss is and each gradient holds for every entry of the parameter it is named after, laid out exactly like that parameter (dW1 is , db2 has entries, and so on).
inputs has examples of values; targets has the same rows of values.For the first test, the forward pass gives , so and against a target of . The loss is therefore . The gradients are yours to find.
Once these gradients exist, the update is one line per parameter: step each one against its own gradient. They all move together, because every gradient here came from the same forward pass, with the weights as they stood then. A classic bug in hand-written training loops updates
W2first and then sends the error back through the newW2. That hands the hidden layer blame for a network that never made the prediction.