A segmentation network returns a class for every pixel, so checking its work means laying two grids side by side: the labels an annotator painted, and the labels the network predicted.
The obvious score is the share of pixels it got right, and it has exactly the flaw the failure-modes lesson warned about. If a tumour covers 4% of a scan, a network that paints every pixel "healthy tissue" is 96% right and has found nothing. The enormous background class drowns out the one class anybody cares about.
The standard fix borrows IoU from bounding boxes and applies it to one class at a time. Treat "every pixel the annotator labelled " as one shape and "every pixel the network labelled " as another, and score how well the two shapes overlap:
Averaging that over the classes gives mean IoU, and every class gets an equal vote in it, however few pixels it covers.
Real annotations add one wrinkle. Along a blurred boundary, annotators often refuse to commit and paint the pixel with a special void label instead. A void pixel is not scored at all: whatever the network predicted there neither helps nor hurts, in either score.
Task: write segmentation_scores(truth, prediction, void), returning [pixel_accuracy, mean_iou], each rounded to 4 decimal places.
truth and prediction are grids of the same shape — a list of rows, each entry a whole-number class id.void is the value that marks an unscored pixel in truth. It differs between datasets, so read it from the argument rather than assuming one. The network never predicts it.truth entry is not void. Everything below counts scored pixels only.pixel_accuracy — the share of scored pixels where the prediction equals the truth.mean_iou — the average of over every class that appears at a scored pixel, in the truth or in the prediction.Worked through, segmentation_scores([[0, 0, 1], [0, 255, 1]], [[0, 1, 1], [0, 1, 0]], 255):
[0.6, 0.4167].Pixel accuracy asks how much of the image is right. Mean IoU asks how well each thing in it was found. On a scan that is mostly background those are very different questions, and only the second one notices a missing tumour.