Training a diffusion model repeats one recipe on every batch. For each clean image: pick a step , jump straight to that noise level with the closed formula while keeping the noise you added, ask the network to predict that noise, and score its guess with squared error.
The schedule is a list . The share of signal left at step is
and each pixel of the noisy image at step is
where is the clean pixel and is the noise drawn for it. Within one batch, every image gets its own step, so a single batch mixes nearly clean images with nearly pure noise.
Task: write noise_prediction_loss(images, noises, steps, betas, model).
images: the clean images, each a list of pixel values, all the same length.noises: the noise drawn for each image, with the same shape as images.steps: one step per image, counting from 1.betas: the schedule. betas[0] is , and the list may be longer than any step used.model: the network, as a function. model(x_t, t) takes one noisy image (a list) and its step, and returns the predicted noise as a list of the same length.An image's loss is the mean, over its pixels, of . The batch loss is the mean of the image losses.
Return a tuple (per_image, batch): the list of image losses in order, and the batch loss. Round every value to 4 decimal places, averaging the unrounded image losses.
In real training,
noisesandstepsare drawn at random every time; they are passed in here so the result can be checked. The test models are small stand-in functions such aslambda x, t: [0.5 * v for v in x].