A backbone for a 100-class problem ends with a feature block of
Two heads are on the table. Neither has a hidden layer — both go straight into a dense output layer with one unit per class.
| head | what it does before the dense output layer |
|---|---|
| flatten head | flattens the whole block into one long vector |
| GAP head | global average pooling: averages each feature map down to a single number |
Count only the weights of the dense output layer in each head. Ignore biases, and remember that pooling has nothing to train.
How many weights does switching from the flatten head to the GAP head remove?
Give your answer in millions, rounded to 2 decimal places (so 3,500,000 would be entered as 3.50).