A VAE's encoder does not hand over a code. For each input it hands over a centre and a spread for every code entry, and the code is drawn from that little cloud. Training then scores the result with two terms that pull in opposite directions, and one number that sets the balance between them.
In this problem the random draws have already been made, so everything is deterministic. For one input with positions and a code with entries:
| step | rule |
|---|---|
| draw the code | |
| decode | |
| reconstruction term | |
| KL term |
Here is the centre, the spread (always positive), and the draw from the standard shape at the origin (a standard normal), all for code entry ; is the natural logarithm. The decoder is a single linear layer: weights with one row per position and one column per code entry, and bias .
Task: write vae_loss(inputs, centres, spreads, noise, weights, bias, beta).
centres, spreads and noise belongs to row of inputs. The batch may hold a single input.beta.[reconstruction, kl, loss], each rounded to 4 decimal places. Round only at the very end: the loss is built from the unrounded averages.The two terms want different things. Reconstruction would like every input's cloud small and far from the others, so a draw can never be mistaken for a neighbour's. The KL term charges every cloud for straying from one standard shape at the origin. is the exchange rate, and the whole behaviour of a VAE lives in that trade.