Linear regression predicts one number from another with a straight line,
and gradient descent finds and by repeatedly stepping downhill on the mean squared error. Batch, stochastic and mini-batch gradient descent are the same algorithm with one difference: how many rows each update looks at. Your job is a single training function that runs all three.
For a batch of rows, the cost is
and its slopes with respect to the two parameters are
The training recipe
batch_size rows. If the rows don't divide evenly, the last batch is smaller. If batch_size is bigger than the dataset, the only batch is the whole dataset.learning_rate:Task: write fit_line(xs, ys, learning_rate, epochs, batch_size) and return the tuple (b0, b1) after epochs epochs, each rounded to 4 decimal places.
Setting batch_size = len(xs) gives batch gradient descent, batch_size = 1 gives stochastic gradient descent, and anything in between is mini-batch.
Run the same data with all three batch sizes and compare. Stochastic gradient descent makes
len(xs)updates per epoch where batch gradient descent makes one, so it travels much further in a single pass. Where it ends up depends on the data. If every row sits exactly on one straight line, all the one-row steps agree and stochastic gradient descent lands on that line. If the rows are scattered, each one-row step tugs the line toward that single row, so with a fixed learning rate it never comes to rest: it repeats the same tugs every epoch, around a point away from the line batch gradient descent closes in on. The bigger the learning rate, the further away that point is.