Training has stalled at a flat point: the gradient is zero in every direction. Whether that point is a dead end depends on how the loss curves as you move away from it.
Here is a crude but classic toy model of a network with weights. At a flat point, look along separate directions, one for each weight. Along each direction the loss either curves up (stepping away along it, either way, raises the loss) or curves down (stepping away along it lowers the loss). In the model:
The flat point is then one of the lesson's three shapes:
| what the directions do | shape of the flat point |
|---|---|
| every direction curves up | local minimum |
| at least one curves up and at least one curves down | saddle point |
| every direction curves down | local maximum |
What is the smallest number of weights for which a flat point in this model is more likely to be a saddle point than a local minimum?
Give a whole number.