The chapter's rule for initialization is scale the spread by the number of inputs: more inputs means smaller weights, so the signal keeps roughly its size through the whole network. In this problem you check that claim one layer at a time, for two different scaling rules.
Take a stack of fully connected layers with ReLU after every layer. layer_sizes : there are input features, and layer (for ) has neurons, each receiving all values from the layer below. Biases start at zero.
Every weight in layer is drawn independently at random with mean and spread (standard deviation) , set by scheme:
scheme | spread of layer 's weights |
|---|---|
"he" | |
"xavier" |
How the signal's size changes. A neuron in layer computes over its inputs. Treat the terms as independent. Each weight has mean , so
where is the mean square of the values arriving at layer , meaning the average of :
input_spread), so .Task: write signal_spread(layer_sizes, scheme, input_spread), returning [weight_spreads, signal_spreads]. Each list has one entry per layer, to :
weight_spreads: ,signal_spreads: the spread of layer 's values before its ReLU.Round every number to 4 decimal places. layer_sizes holds at least two positive whole numbers, scheme is "he" or "xavier", and input_spread is positive.
Worked through, signal_spread([6, 3, 3], "xavier", 1.0):
[[0.4714, 0.5774], [1.1547, 0.8165]].Every layer multiplies the signal's variance by some factor, and a deep network applies its factor again at every layer. This is the compounding from Chapter 3, and it happens before a single gradient has been computed. A factor of looks harmless for one layer. Over fifty layers it leaves of the variance. Run a deep stack under both schemes and compare the second lists.