Training rarely feeds a model one example at a time, and almost never the whole dataset at once. It works through mini-batches: small consecutive chunks of the (already shuffled) example order.
The shuffling is done for you — order is the list of example positions, already in the order you should use. Your job is the chunking.
Task: write make_batches(order, batch_size, drop_last) returning a list of batches, where each batch is a list of positions taken from order in sequence.
order into consecutive chunks of batch_size. The first batch is the first batch_size entries, the next batch the ones after that, and so on.drop_last decides what happens to it: True throws that short chunk away, False keeps it as a smaller batch.batch_size long is never short, so drop_last leaves it alone.batch_size is larger than the whole list, you get one short batch — or, with drop_last=True, no batches at all.Why anyone would throw data away: a short final batch gives a noisier gradient than the others, and some layers (batch normalisation especially) behave oddly on a batch of one. Dropping it keeps every step identical in shape.