VGG-16 carries 138 million weights, and its conv stack is only about 15 million of them. The rest sit in the dense layers bolted onto the end. Global average pooling later deleted almost all of that by averaging each feature map down to a single number, using no weights at all.
You are given a network description and asked to produce the ledger: what the backbone costs, what each of the two candidate heads costs, and how lopsided the result is.
Task: write weight_ledger(size, depth, stages, hidden, classes).
size × size grid of depth channels.stages is a list of [num_convs, filters] pairs. A stage runs num_convs convolutions one after another — each outputs filters feature maps, with "same" padding so the grid is unchanged — and then a max pool with stride 2, which halves each side. Every stage ends with a pool, including the last. If a side is odd, the pool drops the leftover row and column ().hidden units, then to a dense layer of classes units;classes units;Return [backbone, dense_head, gap_head, share], where the first three are the weight counts as plain integers and share is the fraction of a dense-headed network's total that sits in its head,
rounded to 4 decimal places.
The first test is VGG-16's shape, right down to the block it finishes on. The last number you return is the one that explains why anyone bothered inventing global average pooling — and why a backbone transfers to a new task so much more readily than the thing sitting on top of it.