Your team trains an image classifier from a published recipe: SGD with momentum 0.9 at a learning rate of 0.1, cut to a tenth at two milestones. It trains well. The logs show that individual weights' gradients are mostly between and , and that a typical weight moves about per update.
A teammate swaps the optimizer for Adam with its default memory settings ( and , bias correction on) and changes nothing else. The learning rate of and the schedule stay exactly as they were.
Within the first 20 updates the loss shoots upward, and then it becomes NaN.
Which diagnosis is right?
Select all that apply.