A small network has 5 inputs, one hidden layer of 8 ReLU neurons, and 1 output neuron. Someone skipped random initialisation: every weight starts at and every bias at .
Trained with plain gradient descent, one example at a time, this network is stuck in the trap the lesson describes. All 8 hidden neurons compute the same value on every input and receive the same update on every step, so they stay identical forever: eight neurons doing the work of one.
Keeping the same starting weights, the team considers four changes. Each would be made on its own:
| change | what it does |
|---|---|
| Adam | replaces plain gradient descent with Adam, so every weight gets its own adaptive step size |
| L2 penalty | adds an L2 penalty to the loss, so every weight gets its own pull towards zero, set by that weight's size |
| mini-batches | averages the gradients over mini-batches of 32 varied examples instead of using one example at a time |
| dropout | applies dropout at rate to the hidden layer during training |
Which change, on its own, would let the 8 hidden neurons become different from one another?
Select all that apply.