The generator never sees a real example. Everything it learns arrives as a derivative handed back through the discriminator — move your output this way and I would have been more convinced. So the size of that derivative is the size of the lesson.
Take the smallest possible judge. The generator emits one number per fake; the discriminator's head turns it into a logit and squashes it:
near means "real", near means "fake". There are two ways to write down what the generator wants, and they are not the same function:
| form | what the generator minimises |
|---|---|
| original | |
| non-saturating |
Both are driven down by fooling the judge. What differs is the slope on the way there.
Task: write generator_feedback(outputs, weight, bias), where outputs is the list of numbers the generator produced for a batch of fakes and weight, bias are the head's and . Return four numbers, each rounded to 4 decimal places:
outputs is never empty, and weight may be negative.A large loss and a usable gradient are not the same thing. One of these two forms can report a batch as almost nothing to fix at the precise moment the discriminator is rejecting every single sample in it.