A GAN is training on photographs of faces. Its discriminator ends, as binary classifiers do, in a single sigmoid unit:
where is the raw score the last layer produces and is the probability the discriminator assigns to "this one is real".
The generator never looks at a real face. Everything it learns arrives as a gradient that leaves the discriminator's verdict and travels backwards into the generator's weights — and on the way out it must pass through that final sigmoid, so the signal from each sample is multiplied by evaluated at that sample.
The team records the discriminator's average verdict on a batch of generated faces at two points in the run:
| checkpoint | mean on the fake batch |
|---|---|
| after 1 epoch | |
| after 20 epochs |
They are delighted with the second row: the judge now catches almost every fake.
Using at each batch's mean verdict as a stand-in for the whole batch, how has the size of the signal reaching the generator changed between the two checkpoints?
Select all that apply.