A small classifier has three fully connected layers, mapping 80 input features to 65, then to 64, then to the 10 class scores. So the layers hold , then , then weights. After training, the team checks the magnitude of every weight:
| layer | number of weights | every weight's magnitude lies between |
|---|---|---|
| layer 1 | and | |
| layer 2 | and | |
| layer 3 (the 10 class scores) | and |
Layer 3 receives much larger activations than the other layers do, so its weights settled at a smaller scale.
The team prunes with one rule for the whole network: delete the 8% of weights with the smallest magnitudes, wherever they are, then retrain briefly. Deleted weights stay deleted during the retrain; everything else, including every layer's biases, is free to change.
What does the network do after pruning and the brief retrain?
Select all that apply.