A hidden layer takes 4 inputs and has 3 ReLU neurons, so it holds weights and biases: 15 parameters in all. Neuron computes the weighted sum
and sends on to the next layer. So is the weight that carries input into neuron .
During the forward pass on one training example, the layer cached its inputs and its weighted sums. Here are the inputs:
| input | ||||
|---|---|---|---|---|
| value |
The two zeros come from neurons in the layer below that fell silent.
The backward pass has already worked its way from the output down to this layer. The last column below is the blame it delivered to each neuron's output, : how much the loss changes per unit nudge of what that neuron sent onward.
| neuron | cached weighted sum | blame arriving at its output |
|---|---|---|
| 1 | ||
| 2 | ||
| 3 |
The backward pass now computes a gradient for every one of the layer's 15 parameters.
How many of the 15 parameters get a gradient of exactly zero on this example?
The answer is a whole number.