A team trains a network to say which of four tree species a leaf photo shows. The four species are equally common in both the training set and the validation set. The network ends in softmax and is trained with categorical cross-entropy: the loss on one photo is , where is the probability given to the true species, and every logged value is an average over photos. A copy of the network is saved at every logged step.
Run 1, learning rate :
| step | training loss | validation loss |
|---|---|---|
Flat from the start, so a teammate stops run 1 at step and reruns with a learning rate times larger. Nothing else changes.
Run 2, learning rate :
| step | training loss | validation loss |
|---|---|---|
Which reading of the two runs do the logs best support?
Select all that apply.