AlexNet's headline ImageNet result in 2012 was a top-5 error of 15.3%. ImageNet has 1000 classes, many of them nearly identical (there are over a hundred dog breeds alone), so the competition let each model name five guesses per photo. A photo only counted as an error if the true class was missing from all five.
You are scoring a small insect-identification CNN the same way. For every photo, its head produces one raw score per class. These are the scores before softmax. Classes are numbered from 0, so row[c] is the score for class c.
Task: write top_k_error(scores, labels, k).
scores is a list of rows, one per photo, each holding one score per class.labels holds each photo's true class index.k is how many guesses each photo gets.k highest-scoring classes, and an error otherwise.Two details settle the edge cases:
k classes score strictly higher than it.k may be larger than the number of classes. Then every class makes the shortlist.With
k = 1this is the ordinary error rate. Allowing more guesses can only lower the number, which is why papers always say which one they are quoting.