The learning rate is the most important dial in training, and the usual way to set it is not theory but a quick experiment. Try a ladder of rates, each a fixed multiple of the last, give each one a short run, and watch three behaviours appear in order: crawling, working, then blowing up.
Run that experiment on a straight-line model , fitted by mean squared error to points. Its gradients are
Task: write rate_sweep(xs, ys, start_rate, factor, max_rates, steps).
start_rate, start_rate * factor, start_rate * factor**2, and so on: at most max_rates of them, in that order.steps steps of gradient descent using all the points at once. In every step, compute both gradients at the current before changing either one. The rate's result is the MSE after its last step.Return a tuple (chosen, losses):
losses is the result of every recorded rate, in the order tried, each rounded to 4 decimal places.chosen is the recorded rate with the lowest result (on a tie, the smaller rate), rounded to 4 decimal places, or None if the very first rate blew up.