Xavier initialization was derived assuming the activation is roughly symmetric around zero — true for tanh, false for ReLU. ReLU throws away every negative value, so about half the signal disappears at each layer, and weights sized by Xavier leave the activations shrinking as the network deepens.
He initialization compensates by doubling the variance. For a normal distribution, draw each weight with mean 0 and:
Task: write he_std(fan_in) returning that standard deviation, rounded to 4 decimal places.
fan_in is how many inputs the layer takes, and is a positive integer.fan_in appears — unlike Xavier, this formula doesn't balance the forward and backward passes against each other; it just keeps the forward signal alive.The 2 is doing all the work, and it's there for a precise reason: ReLU zeroes out half the inputs, which halves the variance, so you start with twice as much to end up with the right amount. This is the default initializer for ReLU networks in every major framework, and it's a large part of why networks deeper than a few layers became trainable at all.