VGG uses nothing but convolutions, and the standard argument for that is familiar: two stacked layers see the same region as one layer, using weights instead of , with an activation function in between.
Look at the point in VGG-16 where a block widens. A block arrives, and the block's first two convolutional layers are both conv3-128 — , filters, "same" padding:
| layer | input depth | filters |
|---|---|---|
first conv3-128 | ||
second conv3-128 |
Compare that pair against a single layer doing the same job: channels in, filters out, covering the same region of the incoming block.
Count the weights in each arrangement, ignoring biases. Which statement is correct?
Select all that apply.