The softmax lesson mentioned an optional dial called temperature. Every logit is divided by the same number before softmax:
is plain softmax. A smaller sharpens the probabilities toward the top class, and a larger flattens them toward an even split. Dividing every logit by the same positive number never changes which one is largest, so the model's answers are identical at every temperature. Only its confidence moves.
A finished classifier has scored four photos. Each row holds its logits for three animals:
| photo | cat | dog | fox | true animal |
|---|---|---|---|---|
| 1 | 4.0 | 0.5 | −1.0 | cat |
| 2 | 3.5 | 1.0 | 0.0 | cat |
| 3 | 0.0 | 4.5 | 0.5 | dog |
| 4 | 5.0 | 0.0 | 1.0 | fox |
The loss is categorical cross-entropy. The loss on one photo is , where is the probability given to its true class and is the natural logarithm. The batch loss is the average over the photos.
Task: write best_temperature(batch_logits, true_classes), returning [temperature, loss_at_one, loss_at_best], each rounded to 4 decimal places:
temperature is the value of that gives the lowest batch loss;loss_at_one is the batch loss at ;loss_at_best is the batch loss at that best temperature.Details:
batch_logits[i] is the list of logits for photo i, one per class, and true_classes[i] is the index of its true class. Every photo has the same number of classes, at least .The four photos above are the first test, with true_classes equal to [0, 0, 1, 2].
A best temperature above 1 means the model is overconfident: its mistakes are held too firmly, and the batch pays less once every photo is made more modest. A best temperature below 1 means it is timid about answers it mostly gets right. The answers never change either way, which is why accuracy cannot see any of this, and why cross-entropy, which grades confidence, can.