Batch norm, group norm and instance norm all do the same two things: subtract a mean, then divide by a spread. They differ only in which numbers share that mean and spread.
Picture a layer's output for a whole batch as a block of numbers labelled three ways:
| method | one mean and spread for each... | computed over |
|---|---|---|
| batch norm | neuron | every example in the batch, every position |
| group norm | example, and group of neurons | every neuron in that group, every position |
| instance norm | example, and neuron | every position |
Each one then computes with , and afterwards applies its learned scale and shift.
A team is training a network built from plain fully-connected layers. Its inputs are so large that only one example fits in memory, so it trains with a batch size of 1. One hidden layer has 6 neurons, and group norm would split them into two groups of 3. For the current training example, that layer's six outputs are
The team is deciding which normalization to put after this layer.
Looking at the normalized values before the learned scale and shift, which of the three methods passes on anything that depends on the input?
Select all that apply.