A classifier's last layer emits one raw score per class, the logits . Softmax turns them into probabilities, and categorical cross-entropy charges the model for the probability it gave the true class :
A batch's loss is the mean of its examples' losses.
Task: write mean_cross_entropy(logits, labels).
logits is a list of rows, one per example, each holding logits.labels holds each example's true class index, counting from 0.Real logits are not tidy. Yours may be anywhere between and . A model that is confidently wrong can give the true class a probability far too small for a computer to store, so it comes out as exactly 0.0. The true loss for that example is still an ordinary finite number (it can run into the hundreds), and your function must return its exact value: no crash, no inf, and no capping it at something smaller.
When a training loss suddenly turns into NaN, this calculation is one of the first places to look: an that overflowed, or a .