A small network reads greyscale scans of handwritten digits and sorts each one into one of the digits. Each pixel arrives as one input value, so the input layer is numbers wide. Every layer is fully connected to the one above it:
| layer | size |
|---|---|
| input | pixel values |
| hidden layer 1 | neurons |
| hidden layer 2 | neurons |
| output | neurons |
Every neuron in the hidden and output layers has one weight for each neuron (or pixel) in the layer directly below it, plus its own bias. Those weights and biases are the network's learnable parameters.
The model has to shrink to fit on a small device, and the team is weighing four changes. Each would be applied on its own to the original network. Anything a change doesn't mention stays exactly as it is.
Which single change removes the most learnable parameters?
Select all that apply.