Batch normalization standardises every feature across the batch, then lets the network rescale and shift it. Running it forwards is a few lines. This problem is about running it backwards.
A batch holds examples with features each, and x[i][j] is feature of example . In training mode, the forward pass works column by column, using statistics computed from this batch:
The variance divides by , not . The rest of the network turns the outputs into a loss , and the backward pass has already delivered the upstream gradient dout[i][j] .
Task: write batchnorm_backward(x, dout, gamma, eps) and return a tuple (dx, dgamma, dbeta):
dx: an list of lists with dx[i][j] ,dgamma: a list of values, ,dbeta: a list of values, .Round every number to 4 decimal places. eps is always positive, and can be as small as 1.