A classifier takes 4 input features, passes them through two hidden layers of 16 neurons each, both using ReLU, and ends in an output layer of 3 neurons, one per class, feeding a softmax. Every layer is fully connected, so the network has
weights and biases.
Instead of random initialization, someone sets every weight to and every bias to . The network is then trained with mini-batch gradient descent for many epochs on a varied training set: the examples differ from one another, and all three classes appear.
Backpropagation (Chapter 3) supplies the gradients. In the backward pass each neuron receives a share of the blame, and these three facts describe how it is used:
With a mini-batch, each gradient is the average of these per-example gradients over the batch.
However long it trains, some of the 403 parameters are forced to stay exactly equal to one another. What is the largest number of different values that can ever appear among the network's 403 weights and biases?
The answer is a whole number.