A student codes plain gradient descent for a one-weight model with loss
starting at with a fixed learning rate of . There is no momentum, no schedule and no mini-batch noise, so every step should use the exact gradient. Their training log:
| step | 0 | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 |
|---|---|---|---|---|---|---|---|---|---|
| loss | 9.00 | 5.76 | 3.24 | 1.44 | 0.36 | 0.00 | 0.36 | 1.44 | 3.24 |
The loss falls all the way to zero and then climbs straight back up.
Which diagnosis explains this log?
Select all that apply.