A model has been training calmly until one weight, , lands somewhere very steep. For this problem, treat the loss as depending on that weight alone:
The lowest loss is at . The weight is currently , and it is trained by gradient descent with learning rate :
where is the gradient used for the update.
With plain gradient descent, and the weight goes Each update overshoots further than the last, flipping sign and growing, and the run is heading for NaN.
The team switches on gradient clipping with a limit of . Before every update, if the gradient is replaced by with the same sign, so a gradient of becomes and a gradient of becomes . If it is left exactly as it is. The loss, the learning rate and the starting point all stay the same.
With clipping on, what does do in the long run?
Select all that apply.