Diffusion reaches an image by a long, winding route of many small denoising steps. Rectified flow, used by several recent image generators, trains on straight lines instead, so that far fewer steps are needed.
Training (already done). Pair a noise sample with a real image . The straight line between them is
so is pure noise and is the image. Along this line the velocity, the change in per unit of , is the same everywhere: . The network is trained with plain squared error to predict that velocity from the point and the time .
Sampling (your job). Start from fresh noise at and follow the network's velocity until , using Euler's method: split the trip into equal steps of size . At each step, ask the network for the velocity at the current point and the current time, then move in a straight line in that direction for :
Task: write flow_sample(noise, velocity, steps).
noise: the starting sample, a list of numbers.velocity: the trained network, as a function. velocity(x, t) takes the current list and the current time, and returns a velocity list of the same length.steps: , at least 1.Return the sample after all steps, every value rounded to 4 decimal places.
The test velocities are small stand-ins. For example, if the whole dataset were the single image
target, a perfectly trained network would outputlambda x, t: [(m - v) / (1 - t) for m, v in zip(target, x)].