You are scoring a generator the way FID does — not image against image, but the whole cloud of generated samples against the whole cloud of real ones.
Here is that idea stripped down to one dimension. A feature network has already reduced every image to a single number. Five real photographs and five generated samples come out as
| set | feature values |
|---|---|
| real | |
| generated |
The score compares the two clouds through their centre and their spread only:
where summarise the real set and the generated one. Use the population standard deviation — divide the squared deviations by , not by . Lower is better, as always.
What score does this generator get?
Round the answer to 2 decimal places.