The chapter raced plain descent against momentum and watched the paths. Here you run that race yourself, over several different landscapes, and count how long each optimizer takes to settle.
Two weights are trained on the loss
which is passed in as the symmetric matrix M = [[p, q], [q, r]]. Every M you will be given is a bowl whose single lowest point is , where . When the bowl's axes line up with and . When the valley runs diagonally. The gradient is
Plain descent. Measure both partial derivatives at the current point, then move both weights:
Momentum. Keep a velocity that starts at . On every update, measure the gradient at the current point, then
Here (beta) is how much of the previous motion is kept. At this is exactly plain descent.
Both runs start from the same point start = [x, y], use the same lr, and last exactly 1000 updates. The two runs are independent: momentum starts from start with zero velocity.
When has a run settled? A falling loss is not enough, because momentum's loss can dip below a threshold and then climb back over it. A run's settling count is the smallest number of updates such that the loss after updates, and after every later update up to the 1000th, is strictly below tol. The loss after updates is the loss at start, so a run that begins below tol and stays there has a settling count of . If the loss after the 1000th update is not below tol, the run never settled, and its count is None.
For example, suppose a run's losses after updates were and every later loss stayed below . With tol = 0.05, its settling count is , not : update 2 dipped under, but update 3 climbed back out.
Task: write race_to_settle(M, start, lr, beta, tol), returning [plain_count, momentum_count]. Each entry is a whole number or None.
You can rely on lr , beta and tol .
Nothing in these rules favours either optimizer. Which one wins, and by how much, depends entirely on the landscape and the settings, and that is the reason to race them.