A GAN training step has two halves. First the discriminator is updated on a batch of real examples and a batch of fakes. Then the generator is scored against the discriminator as it stands after that update. Here you will write both losses and the discriminator's update, using the smallest discriminator there is.
The discriminator. The data are single numbers. The discriminator has two parameters, and . It gives an example a raw score (a logit), then squashes that with a sigmoid into its probability that is real:
Discriminator loss. Binary cross-entropy, with label for every real example and label for every fake. Each batch is averaged on its own, and the two averages are added:
Discriminator update. One step of gradient descent on and , with learning rate . For a single example with score and label , the slope of its cross-entropy with respect to is . Because , that example's slope with respect to is , and with respect to it is . The gradient of combines these per-example slopes in exactly the way combines the per-example losses.
Generator loss. The generator wants its fakes to be called real. It does not minimise , which goes flat exactly when the discriminator confidently rejects a fake. It minimises this instead:
is measured with the updated and . The generator's own update is not part of this task: the fakes are handed to you as numbers.
Task: write gan_step(w, b, real, fake, lr). Return a tuple (d_loss, new_w, new_b, g_loss): at the starting parameters, the two parameters after one update, and at the updated parameters, each rounded to 4 decimal places.
Large scores. A confident discriminator can produce scores like , and your losses must still come out finite. These two identities never take the logarithm of a probability that has rounded to 0: