Layer normalization does two things to each position's vector. It recentres it (subtract the mean) and rescales it (divide by the spread), then applies a learned scale and shift. Many recent language models keep only the rescaling. Their version, RMSNorm, divides each vector by its root mean square and multiplies by a learned gain, one per slot. No mean is subtracted, and there is no shift.
For a vector of length and gains :
Here is a tiny number that prevents a division by zero when every slot is . It sits inside the square root.
Like layer norm, RMSNorm works on each position's vector separately: one root mean square per vector, computed from that vector alone.
Task: write rms_norm(vectors, gain, eps), where vectors is a list of position vectors, all of length , and gain is a list of gains. Return the normalized vectors, with every number rounded to 4 decimal places.
Dropping the recentring saves work on every vector in every layer, and in practice it works about as well. The rescaling is the part that keeps the numbers from drifting through a deep stack.