Inpainting means filling in a missing part of an image so that it fits what is around it. A diffusion model trained only to make whole images can already do this. Nothing is retrained; one extra move is added to sampling.
Ordinary sampling starts from pure noise and applies the trained denoising step again and again: . When you already know some pixels of the answer, add this after every step: overwrite each known pixel with its true value noised to the level the image has just reached, using the forward formula
where is that level, is the pixel's true value and is a noise value for that pixel. Unknown pixels keep whatever the denoiser produced. The denoiser therefore always sees known pixels at the same noise level as the rest of the image, and at every step it shapes the unknown region to agree with them.
In full: set . Then for :
denoise_step(x, t), which takes the whole image from level to level .Level 0 means no noise at all (), so after the last step the known pixels hold exactly their true values.
Task: write inpaint(known, mask, x_T, known_noise, alpha_bars, denoise_step).
known: the true image. Only the positions where mask is 1 matter; the others hold placeholder zeros.mask: 1 for a known pixel, 0 for a pixel to fill in.x_T: the starting noise, one value per pixel.known_noise: the value for each pixel, reused at every level.alpha_bars: , so its length is .denoise_step: the trained denoising step, as a function. denoise_step(x, t) takes the whole image as a list and returns a new list.Return the final image , every value rounded to 4 decimal places.
A real sampler draws fresh noise for the known pixels at every step; one fixed draw is used here so the result can be checked. The test denoisers are small stand-ins, such as a step that pulls every pixel part of the way toward the image's average. Like a real denoiser, they let pixels influence one another, which is what lets the known region steer the rest.