A model with two weights is trained on one of two loss surfaces:
Both are the same narrow valley: steep across, shallow along, with its bottom at where the loss is . In A the valley floor runs along the axis. B is exactly A turned by 45° about the bottom, so its floor runs along the diagonal .
Each run starts at the matching point: A at , and B at . In both cases that is 1 unit up the steep wall and 7 units along the floor from the bottom, with a starting loss of .
For reference, the gradients are
Three optimizers were raced on valley A. The table shows the number of steps until the loss fell below and stayed there:
| optimizer | settings | steps on A |
|---|---|---|
| plain descent | rate | |
| momentum | memory , rate | |
| Adam | memory settings and , rate |
Momentum keeps a velocity that starts at . Each step it updates and then moves the weights by . Adam is the standard version, with bias correction on.
The same three optimizers, with identical settings, are now raced on valley B from its matching start.
Which prediction for valley B is right?
Select all that apply.