In distillation the student learns from the teacher's whole spread of probabilities, not just the right answer: cat, but dog was a close second. There's a catch at an ordinary softmax. A confident teacher puts almost all its probability on one class, and the close-second information nearly disappears. So both models' scores are first softened with a temperature :
At this is the ordinary softmax. A larger spreads the probabilities out more evenly.
The loss for one example blends two parts.
The is there because softening shrinks the soft part's gradients by roughly a factor of . Multiplying it back keeps the two parts on a comparable scale whatever temperature you pick.
Task: write distillation_loss(student, teacher, labels, T, alpha).
student and teacher are lists of score lists, one per example.labels holds each example's correct class index.Average the soft part over the batch, and average the hard part over the batch. Return a tuple (soft, hard, loss): the average soft part, the average hard part, and their blend. Round each to 4 decimal places, only at the end: blend the unrounded averages.