Chapter 4 compares plain descent, momentum and Adam by racing them on the same landscape. Here you build that race.
The track is a stretched bowl. With weights and a curvature for each coordinate,
A large makes that direction steep and a small one makes it gentle. When the curvatures differ a lot, the bowl becomes a narrow valley.
The runners. All three start from the same point. Each takes exactly steps updates using its own learning rate from rates. At every update a runner measures the gradient at its current position and then moves every coordinate at once. Everything below is per coordinate.
"sgd"): "momentum"), with velocity starting at :"adam"), with and starting at , and update number :Task: write optimizer_race(curvatures, start, rates, steps).
curvatures and start are equal-length lists of floats.rates is a dict with keys "sgd", "momentum" and "adam".Return a dict with the same three keys, each mapped to that runner's final loss , rounded to 4 decimal places. A runner whose learning rate is too large for the track still reports its loss, however huge.