A training recipe was tuned for batches of 6 examples with learning rate , but the GPU only has room for 2 examples at once. So the engineer uses gradient accumulation. Each batch goes through as three micro-batches of 2. For each micro-batch, the gradient is computed (the mean of its two examples' gradients) and added into an accumulator. Only after the third micro-batch is the accumulator used to take one step. Then the accumulator is reset to zero and the next batch begins.
The goal is for every step to be exactly the step the recipe's batch of 6 would have taken, where a batch's gradient is the mean of its examples' gradients.
The model has a single weight, , and each example's loss is , whose gradient is
The whole dataset is one batch, split into micro-batches like this:
| micro-batch | ||
|---|---|---|
| A | ||
| A | ||
| B | ||
| B | ||
| C | ||
| C |
Training starts at . Two full passes over the data are run, so the weight is updated twice.
What is the weight after the second update? Round your answer to 3 decimal places.