A toy VAE, small enough to work by hand.
The data. One-pixel images. Only two exist — a black pixel (value ) and a white pixel (value ) — and they are equally common.
The code. A single number.
The blobs. As in any VAE, the encoder outputs a centre and a spread, and the code is drawn from the blob they describe. To keep the arithmetic exact, this toy uses a flat blob: the code is drawn uniformly from the interval running from to , every value in it equally likely. (A real VAE's blob is bell-shaped; the flatness only makes the sums clean.)
The team kept turning the KL term up until random codes stopped landing in holes. The encoder now outputs:
| image | centre | spread |
|---|---|---|
| black () | ||
| white () |
The encoder is now frozen. The decoder turns a code into a pixel value, and it may be any function of the code at all — assume unlimited capacity, trained to the best it can possibly do. Its reconstruction loss is the squared difference between the decoded pixel and the original pixel, averaged over both images and over every possible draw of the code.
What is the lowest reconstruction loss any decoder can reach with these blobs?
Select all that apply.