You are adapting a pretrained backbone to your own six-class dataset. The old -class head is thrown away and a new head is bolted on.
The backbone's weights, block by block:
| part | weights |
|---|---|
| stem + block 1 | |
| block 2 | |
| block 3 | |
| block 4 (the top block) |
The backbone ends with global average pooling, which turns each image into a vector of numbers. The new head is two fully connected layers: , then . A fully connected layer gives every output unit one weight per input number, plus one bias of its own.
You follow the usual recipe for a mid-sized dataset: freeze the backbone except its top block, and train that top block together with the new head.
What percentage of the whole model's parameters are you training?
Round your answer to 2 decimal places.