A team trains an image classifier on training examples with mini-batches of , making one weight update per mini-batch. Their loss curve is jumpy, so they measure how noisy the steps are. Holding the weights fixed, they draw many mini-batches at each of three sizes and log one weight's mini-batch gradient:
| batch size | average of the mini-batch gradients | standard deviation of the mini-batch gradients (the batch-to-batch wobble) |
|---|---|---|
| (current) | ||
At the current batch size the wobble is three times the size of the signal underneath it. The team wants it no bigger than the signal itself, so a standard deviation of at most , and plans to get there by raising the batch size and changing nothing else.
Each mini-batch gradient is the average of per-example gradients drawn independently from the training set, so the wobble follows one smooth rule in , including at batch sizes the team did not log. Changing leaves the average at ; only the wobble changes. Every mini-batch is full, with no partial batch at the end of an epoch.
If they switch to the smallest batch size that meets the target, how many weight updates will one epoch contain?
Give a whole number.