The forward process has nothing to learn. At step it keeps a fraction of what it was handed and mixes in fresh noise, following a fixed schedule . Step keeps , so the signal remaining after steps is the running product
Training would be hopeless if every example meant running the chain times. It doesn't: one calculation lands you at any noise level you like.
Here is the noise you drew — and therefore the answer the denoiser will later be scored against.
Task: write noisy_at_step(x0, betas, noise, t), returning as a plain list, rounded to 4 decimal places.
betas[0] is , the amount of noise added by the first step.x0 and noise are lists of the same length; everything is elementwise.t runs from 0 up to len(betas). At t = 0 nothing has happened yet.Notice that the two coefficients are square roots rather than the fractions themselves. That is what holds the mixture at a constant scale: , so a barely touched image and an almost fully destroyed one arrive at the network with the same typical magnitude. One network handles every noise level partly because of this.