Every image in a dataset is a strip of 8 pixels holding one bar: two neighbouring pixels at 1, the other six at 0. The bar can sit anywhere in the strip, but never in the first or last pixel, so nothing below ever spills off an edge.
Two trained autoencoders are compared on this data.
- Crisp always outputs a perfectly sharp bar: two pixels at 1, the rest at 0. Usually the bar is in exactly the right place. On a fraction p of images, it lands one pixel to the right of where it belongs.
- Soft never outputs a sharp bar. On every image it outputs a smear centred exactly on the true bar. For a bar at pixels 4 and 5 it outputs
[0, 0, 0.25, 0.5, 0.5, 0.25, 0, 0]
and it outputs the same shape, shifted along, for a bar anywhere else.
Reconstruction loss is the usual one: subtract, square and average over the 8 positions of an image, then average over all the images.
For which values of p does reconstruction loss score Soft as the better model?