Ordinary cross-entropy tells the model the correct class has probability exactly 1 and every other class exactly 0. Chase that target and the model is pushed to be infinitely confident — which makes it brittle, badly calibrated, and prone to memorising noisy labels.
Label smoothing softens the target. Spread a small amount of probability ε evenly across all K classes, and give the rest to the correct one:
Then score the prediction against that softened target with the usual cross-entropy sum:
Task: write label_smoothing_loss(probs, target, epsilon) returning the loss, rounded to 4 decimal places.
probs is the model's predicted probability for each class and sums to 1. target is the index of the correct class. K is len(probs).ε/K as well, which is why its weight is 1 - ε + ε/K and not 1 - ε. The weights then still add up to exactly 1.epsilon = 0 this collapses back to plain cross-entropy — a useful thing to check.The effect is a floor under the loss: no matter how confident the model gets, it can never reach zero, because the target it's chasing was never 1 in the first place.