A working diffusion setup: one network, steps, trained to predict the noise that was added, and told at every call which step it is on.
An engineer trims it. The same weights run at every step anyway, so the step number looks like a spare wire — the network is rebuilt without that input. Everything else is untouched: the same forward process, the same free pairs, the same squared error against the noise that was actually added, the same architecture otherwise.
Training runs cleanly and the loss falls, but it settles well above where the original version settled. Samples come out muddy: over-smoothed in some regions, still visibly grainy in others.
Which explanation fits what happened?
Select all that apply.