A team decides the many-small-steps business is overhead. The forward process hands out free pairs at any noise level, so why not train on the pair at the very end of the chain and generate in a single shot?
Their setup:
Nothing in the pipeline is broken. Every target is exact, because the real image was in hand the whole time. There is no adversary, so nothing can destabilise. The training loss falls smoothly and settles at a low plateau.
At generation time they draw fresh noise and run the network once. Every sample is a soft, smeared wash of colour with no object in it — and all the samples look much alike.
What did this network actually learn?
Select all that apply.