A digit classifier has 784 inputs, a hidden layer of 256 ReLU neurons, a second hidden layer of 128 ReLU neurons, and 10 output neurons feeding a softmax. It is trained with cross-entropy, and every layer is fully connected.
Initialization is He: every weight in a layer is drawn independently, with mean and variance , where is the number of inputs each neuron in that layer receives. Biases start at .
The loss includes an L2 penalty on the weights (not the biases):
On the very first mini-batch, before any update, the data part of the loss (the cross-entropy) comes out at .
At initialization, what is the expected value of the penalty, as a multiple of that data loss? In other words, find the expected penalty divided by .
Round your answer to 2 decimal places.