A network that reads handwritten digits takes 500 inputs (pixel values), has two hidden layers of 200 and 100 ReLU neurons, and ends in 10 output neurons with no activation. Every layer is fully connected, and to keep the count clean the network has no biases.
| layer | connects | weights |
|---|---|---|
| 1 | inputs neurons | |
| 2 | ||
| 3 | outputs | |
| total |
Making one prediction takes exactly one multiplication per weight, since every weight multiplies the one value that arrives along it. That is multiplications.
Now count one full training step on one example: a forward pass, the loss, a backward pass that finds the gradient of every weight, and then the update, which multiplies each weight's gradient by the learning rate and subtracts it from the weight. Count every multiplication of two numbers, with these rules:
How many multiplications does the whole training step take?
Select all that apply.