In Gradient Descent by Hand the loop ran exactly 500 times. Nobody knows in advance how many steps a line will need, so real training loops decide for themselves when to stop. They stop when the nudges stop mattering.
Fit the line to points by gradient descent on the mean squared error. This is the same loop as in the lesson:
A nudge is how far a knob moves in one step: and . The step itself is and .
Task: write fit_until_settled(xs, ys, rate, tol, max_steps).
tol in size (ignore their sign), stop. That step counts.max_steps steps, even if the nudges have not settled.Return a tuple (w, b, steps). w and b are the final values, each rounded to 4 decimal places, and steps is the number of steps taken, as an integer. Return plain Python numbers.
A cap on the number of steps is not optional in real code. With a learning rate that is far too large, the nudges grow instead of shrinking, and without a cap the loop would never end.