A model has a single weight and four training examples with targets
Each example's loss is , so its gradient is . A batch's gradient is the average of its examples' gradients. The best for the whole dataset is the mean target, .
The examples are stored, and fed, in exactly the order , with no shuffling. Starting from with a learning rate of , you make one pass through the data with each of three batch sizes:
| scheme | batches, in order |
|---|---|
| full batch | |
| mini-batch of 2 | , then |
| stochastic | , then , then , then |
Every update uses the gradient at the current value of , recomputed after the previous update.
Where does end up after the one pass, for each scheme?
Select all that apply.