A discriminator that gets too sharp stops teaching. A tiny change to its input can swing its verdict wildly, and the generator is left with almost no useful signal. One widely used fix is a constraint on the discriminator itself: no layer may stretch any input by more than a factor of 1.
From Linear Algebra, a weight matrix turns the unit circle into an ellipse, and its largest singular value is how far that ellipse reaches from its centre along its longest axis (half that axis's full length): the most can stretch any input of length 1. Dividing every weight by pulls that farthest reach in to exactly 1:
Finding exactly at every training step would be slow, so it is estimated with a few rounds of a cheap procedure. Start from a given vector , with one entry per row of . Each round does two things, in this order:
(A vector's length is the square root of the sum of its squared entries.) After the last round, the estimate is
Task: write spectral_normalize(W, u, steps).
W is a list of rows of any shape, not necessarily square.u has one entry per row of W. It need not have length 1.steps is the number of rounds, at least 1.Return a tuple (sigma, W_norm): the estimate after exactly steps rounds, and W with every entry divided by that estimate. Round sigma and every entry of W_norm to 4 decimal places, dividing by the unrounded estimate.
Real training runs a single round per step and carries
uover to the next step, so the estimate is allowed to be unsettled. Report what the rounds actually produce, not the exact singular value.