A language model is a stack of 24 transformer blocks. They are all built to the same plan and all have width — every vector travelling up the stack has slots — but each block carries its own learned weights.
Inside a block, two sublayers run in order: masked self-attention, then the feed-forward network. Each one is wrapped the same way — its output is added back onto its input, and a layer normalization is applied to the result. Every one of those normalizations has its own learned numbers.
A normalization learns two numbers for every slot it normalizes: a scale and a shift.
Nothing else is being counted here. Leave out the attention projections and the feed-forward matrices entirely.
How many learned parameters do all of this model's layer normalizations hold together?
Answer with the exact whole number.