Slow sampling was diffusion's one real weakness, and distilling many steps into a model that needs only a few is one of the ways that gap has been closed. This is what it does to a stopwatch.
A latent-diffusion tool generates one image on one machine. Three parts do the work, and they do not run the same number of times.
| part | cost of one run |
|---|---|
| text encoder | ms |
| denoiser — one pass over the latent | ms |
| decoder — latent to pixels | ms |
The text encoder runs once per image and the decoder runs once per image, whatever the sampler does in between.
Setup A — the default. sampling steps with classifier-free guidance switched on, so the denoiser is run twice at every step.
Setup B — a distilled sampler. sampling steps, and guidance has been baked into the distilled model, so the denoiser is run once at every step.
How many times faster is setup B at producing one finished image, end to end?
Round your answer to 2 decimal places.