Two teams train the same architecture on the same images with the same squared-error loss and the same forward process. Only the target differs:
They compare their numbers at step 5 of 1000 — the shallow end of the chain, where the state is
Pixel values and noise are scaled so that both pieces have roughly unit size, and the noise is unrelated to the image.
| team | target | average squared error at step 5 |
|---|---|---|
| A | the clean image | |
| B | the noise |
Team A points at the table: same data, same architecture, same loss function, and a number nearly fifty times smaller. They conclude that predicting the image is the better target.
What do these two numbers actually tell you?
Select all that apply.