Starting weights matter more than they look. Set them all to zero and every neuron in a layer computes the same thing forever. Set them too large and the signal grows layer after layer until it overflows; too small and it fades to nothing before reaching the output.
Xavier (Glorot) initialization picks the size that keeps the signal steady. For a uniform distribution, draw each weight from [-limit, +limit] where:
Task: write xavier_limit(fan_in, fan_out) returning that limit, rounded to 4 decimal places.
fan_in is how many inputs the layer takes; fan_out is how many outputs it produces. Both are positive integers.The two fans both appear because the formula is a compromise. Keeping the variance of the forward signal steady argues for dividing by fan_in; keeping the backward gradient steady argues for fan_out. Glorot's answer was to split the difference and use their sum.
Notice what the formula implies: a wide layer gets smaller starting weights. Each neuron there is adding up many more inputs, so each one has to contribute less for the total to come out the same size.