Your GPU can hold 32 training examples at a time, but you want every update to use a bigger batch than that. The standard fix is gradient accumulation:
Done correctly, that one update is identical to the step batch gradient descent would take on all of those examples together. The GPU just never has to hold them all at once.
Each micro-batch reports the average gradient over its own examples, which is what the loss gives when is the number of examples in that micro-batch.
Task: write accumulated_step(params, micro_grads, sizes, lr) and return the parameters after that single update.
params — the current parameter values, a list of floats.micro_grads — one gradient per micro-batch: micro_grads[j][i] is micro-batch j's average gradient for parameter i.sizes — sizes[j] is the number of examples in micro-batch j. They are not always equal: when the examples do not divide evenly, the last micro-batch is shorter than the rest.lr — the learning rate .The update must equal exactly the gradient descent step , with computed over all the examples in all the micro-batches at once. Return the new parameter values, each rounded to 4 decimal places.