Sideways Against Forward — hard Optimizers problem | Incognition
Sideways Against Forward
30 pts · 30 coins
OptimizersHard
Sideways Against Forward
Two weights sit in the kind of narrow valley this chapter keeps returning to: steep across, shallow along.
The along weight runs down the valley floor. Its gradient is +0.5 at every step.
The across weight runs between the steep walls. Every step overshoots the floor and lands on the opposite wall, so its gradient alternates: +6,−6,+6,−6,…
To isolate what each optimizer does with the gradients it is handed, assume both weights keep receiving exactly this pattern for thousands of steps, whichever optimizer is running. Two optimizers are compared.
Plain gradient descent, learning rate 0.01: each weight moves by 0.01×g.
Adam, learning rate 0.001, default memory settings. Each weight keeps its own two memories, both starting at 0. At step t=1,2,3,… it receives gradient g and does
m←0.9m+0.1g(direction memory)
s←0.999s+0.001g2(size memory)
step size=0.001×s/(1−0.999t)∣m/(1−0.9t)∣
The two divisions by 1−0.9t and 1−0.999t are the bias correction from the lesson. The tiny constant that real implementations add to the denominator is negligible here and left out.
For either optimizer, define its sideways-to-forward ratio as
size of the along weight’s stepsize of the across weight’s step
measured once the run has gone on long enough for this ratio to settle.
How many times smaller is Adam's settled sideways-to-forward ratio than plain gradient descent's?