A diffusion model is part-way through generating an image from the prompt "a wooden chair". At this step the denoiser is run twice, exactly as classifier-free guidance requires: once with the prompt attached, once with no condition at all. Each run returns a predicted noise pattern the same shape as the image, and we read off the value at two positions in it.
| position | run without the prompt | run with the prompt |
|---|---|---|
| A | ||
| B |
The sampler then combines the two runs using a guidance scale of , starting from the unconditioned prediction and pushing away from it in the direction of the conditioned one:
What does the guided prediction come to at each of the two positions?
Select all that apply.