A trained image classifier ends like this. The backbone hands over a feature block. Global average pooling turns each of the maps into a single number. One dense layer turns those numbers into class scores — one per ImageNet class.
A student looks at this and says:
"So each of the feature maps is the detector for one class. The dense layer just reads off which map fired hardest."
That picture is wrong. Which statement best explains why?
Select all that apply.