During the backward pass every layer is handed an error signal: one number per neuron, saying how fast the loss changes as that neuron's weighted sum changes. Receiving it is only half of the layer's job. The layer still has to turn it into a gradient for each of its own weights and biases, one per connection, so that every one of them can be updated.
The layer is dense and its weights are stored as a connection table, with one row per input and one column per neuron: weights[i][j] is the weight on the wire from input i to neuron j. With inputs and neurons the table is . On one example, neuron computes
The training step runs a batch of examples, and the loss it minimises is the average of the examples' own losses:
For each example you are given two lists:
inputs[e]: the values that arrived at the layer on that example. The forward pass stored them.errors[e]: the error signals the backward pass delivered for that example alone. Entry is , how fast example 's own loss changes as neuron 's weighted sum on that example changes.Task: write layer_gradients(inputs, errors), returning [weight_grads, bias_grads]:
weight_grads: a table in the same layout as weights. Entry [i][j] is , the gradient of the batch loss .bias_grads: entries, where entry is .Round every value to 4 decimal places. inputs has rows of numbers, and errors has the same rows, each of numbers. A table with a single row or a single column is still a list of lists.
Notice what this function was never given: the layer's weights. Everything it needed was either stored on the way forward or delivered on the way back. That is why a training framework keeps every layer's input in memory until the backward pass has come through. The forward pass cannot throw them away as soon as the next layer has used them.