Most augmentations take one photograph and change it. Mixup does something stranger: it takes two photographs and blends them, pixel by pixel, into a single training image that is part cat and part aeroplane.
The part that catches people out is the label. The rule from the lesson is that an augmentation must never change the answer — and an image that is only three-quarters cat is no longer fully a cat. So the label is blended by the same fraction the pixels were, and the model is trained against that soft target.
Task: write mixup(image_a, label_a, image_b, label_b, lam).
image_a and image_b are grids of the same shape — a list of rows, each entry a grey level from to .label_a and label_b are one-hot vectors over the same classes: [0, 1, 0] is class of three.lam is the mixing fraction , between and . It is the share of A, so every pixel becomes
and the label is blended by exactly the same rule, entry by entry.
[mixed_image, mixed_label] — the grid first, then the label vector. Round every number to 4 decimal places.Worked through, mixup([[0, 100]], [1, 0], [[200, 0]], [0, 1], 0.25):
[[[150.0, 25.0]], [0.25, 0.75]].Note that means the result is mostly B, in the pixels and in the label alike. The two must agree: an image that is 75% aeroplane carries a label that is 75% aeroplane, or you have handed the model a picture and told it the wrong thing about it.
This looks like vandalism and it works remarkably well. The model can no longer answer with total confidence from one memorised patch, because the target it is being scored against is itself a blend.